v1.91.7.0 feat: add functional QA and pre-publication docs checks (#2983)

* feat: add surface-aware exploratory QA and ship documentation gates

* test: preserve delegated QA setup authority after main integration

* fix(qa): clarify exploration order and preserve report artifacts

* test(qa): follow the shared setup reference directly

* refactor(ship): make verification and recovery routes explicit

* test(ship): align evidence and review guards with explicit routes

* fix(workflows): clarify ship recovery and functional QA evidence

* fix(workflows): clarify approval recovery and full QA coverage

* refactor(workflows): order review transactions and clarify ship state

* fix(ship): clarify final verification and fail closed at publication

* fix(evals): attribute native atomic documentation writes

* fix(ship): clarify recovery and documentation lifecycle guidance

* fix(test): preserve observed native placeholder styling in CI

* fix(codex): report watchdog timeouts without a process-exit race

* Checkpoint functional QA implementation and workflow validation repairs

* Fix documentation and shared-review fixture contracts

* docs: clarify judge reuse and evaluation supervision

* test: align review evidence and selected case contracts

* test: verify append-only documentation checkpoints and recovery

* fix: qualify QA workflows and CI validation repairs

* fix: launch shared-libs fixture scripts on Windows

* fix: qualify QA deadlines, fixture isolation, and shard cleanup

* fix: preserve qualified QA and cancellation repairs

* fix: enforce functional fixture authority and share strict event decoding

* fix: retain free-test evidence and explain recovery

* fix: reject malformed native evidence after decoder consolidation

* test: use reliable capture for telemetry privacy filters

* test: refresh measured quick coverage and document validation costs

* Fix native fixture receipts and preserve VM validation evidence

* Align negative judge controls with upstream clarity policy

* Fix report-only QA preparation and public evidence handling

* Clarify QA-only preparation and current-report preservation

* Stream Ship quality judgments with an explicit 64k response contract

* Validate compact judge reasoning locally with supported wire schema

* Align functional QA fixture instructions with evidence acceptance

* Bind native browser diagnostics to execution evidence and align review verdicts

* Preserve native diagnostic line boundaries

* Serialize functional QA evidence from native captures

* Keep large QA evidence fixture payload out of Windows argv
This commit is contained in:
Garry Tan authored and GitHub committed 2026-09-29 06:07:35 -07:00
1 parent 65bfb0ce49
commit dcaea52800
333 files changed
+41755 -7357

No files matched your search

+5 -3
View File
@@ -142,9 +142,11 @@ the same command line:
GSTACK_SESSION_KIND=spawned "$_SS" --skill "document-release" ...
```
gstack itself uses this: `/ship` Step 18 dispatches the `/document-release`
subagent with this prefix so its interactive gates auto-choose instead of
prose-stopping. Deliberately narrow: only `spawned` is honored — `headless`
gstack itself uses this: `/ship` Step 14.5 dispatches the `/document-release`
subagent with this prefix before final commit, verification and publication.
Its ship-owned scope overrides generic spawned auto-choice: risky or uncertain
documentation changes return as blockers for the parent, without interactive
questions or automatic approval. Deliberately narrow: only `spawned` is honored — `headless`
already has `GSTACK_HEADLESS`, and letting an env var force `interactive`
over CI markers would be a misclassification footgun. Empty or other values
are reserved and ignored (fall through to ambient detection). Note that hook
+5 -5
View File
@@ -25,7 +25,7 @@ gstack/
│ ├── gen-agents-digest.ts # Generates the budget-capped instruction-tier digest (agents-digest/)
│ ├── host-config.ts # HostConfig interface + validator
│ ├── host-config-export.ts # Shell bridge for setup script
│ ├── resolvers/ # Template resolver modules (preamble, aside = the Aside driver contract + research, browse = $B fallback setup + command reference, design, design-checklist = renders review/design-checklist.md from lib/design-catalog.ts, review, gbrain, etc.)
│ ├── resolvers/ # Template resolver modules (preamble, aside = the Aside driver contract + research, browse = $B fallback setup + command reference, qa = surface-aware QA/exploration, sections = lazy loading, design, design-checklist = renders review/design-checklist.md from lib/design-catalog.ts, review, gbrain, etc.)
│ ├── skill-check.ts # Health dashboard
│ ├── test-free-shards.ts # Strict parallel free-suite runner (GSTACK_FREE_JOBS, opt-in flaky retry)
│ ├── test-paid-shards.ts # Sharded paid-tier runner (one Bun process per shard)
@@ -43,7 +43,7 @@ gstack/
│ ├── setup-*.test.ts, relink.test.ts, hook-scripts.test.ts # Tier 1: setup linker ownership, retired-skill prune, browser hint, rebuild check + Chromium bootstrap (anchor-sliced from setup), gstack-relink, PreToolUse hooks (free)
│ ├── skill-llm-eval.test.ts # Tier 3: LLM-as-judge (~$0.15/run)
│ └── skill-e2e-*.test.ts # Tier 2: E2E via claude -p (~$3.85/run, split by category)
├── qa-only/ # /qa-only skill (report-only QA, no fixes)
├── qa/, qa-only/ # Surface-aware browser/functional QA; /qa-only reports and proposes tests without product edits
├── plan-design-review/ # /plan-design-review skill (report-only design audit)
├── design-review/ # /design-review skill (design audit + fix loop)
├── ship/ # Ship workflow skill
@@ -65,7 +65,7 @@ gstack/
├── guard/, unfreeze/ # /guard (careful + freeze in one), /unfreeze
├── gstack-upgrade/ # /gstack-upgrade skill + migrations/ (run after ./setup during an upgrade)
├── bin/ # CLI utilities (gstack-render.ts = render a local HTML file through Aside or the engine, gstack-design-detect.ts = probe/scan through a user-installed impeccable engine; gstack-design-md.ts = open DESIGN.md check/convert/tokens/mark; gstack-repo-mode, gstack-slug, gstack-config, gstack-wtree, gstack-evidence, gstack-issue-guard, gstack-relink, gstack-memorable, etc.)
├── document-release/ # /document-release skill (post-ship doc updates + Diataxis coverage map)
├── document-release/ # /document-release skill (every-ship pre-verification audit; standalone doc updates + Diataxis coverage map)
├── document-generate/ # /document-generate skill (Diataxis doc generator: tutorial/how-to/reference/explanation)
├── cso/ # /cso skill (OWASP Top 10 + STRIDE security audit)
├── design-consultation/ # /design-consultation skill (design system from scratch)
@@ -73,7 +73,7 @@ gstack/
├── open-gstack-browser/ # /open-gstack-browser skill (launch GStack Browser)
├── connect-chrome/ # symlink → open-gstack-browser (backwards compat)
├── setup-browser-cookies/, pair-agent/, skillify/ # Fallback-engine skills (cookie import, shared-browser tunnel, codify a /scrape)
├── qa/, qa-only/, scrape/ # Browser skills (with design-review/, canary/, benchmark/) — Aside first via {{ASIDE_SETUP}}, $B when Aside is absent
├── scrape/ # Browser data extraction (with design-review/, canary/, benchmark/); Aside first, $B fallback
├── make-pdf/ # /make-pdf skill + compiled `pdf` binary (embeds lib/aside-render.ts); test/ = unit tests (cli-exit-codes, setup-smoke, render) + e2e/*-gate.test.ts on whichever engine resolves
├── diagram/ # /diagram skill (mermaid → SVG/PNG/.excalidraw through bin/gstack-render.ts + lib/diagram-render)
├── design/ # Design binary CLI (GPT Image API)
@@ -82,7 +82,7 @@ gstack/
│ └── dist/ # Compiled binary
├── agents-digest/ # Committed 2KB instruction-tier rules digest (gstack-AGENTS.md) for rules-reading hosts
├── extension/ # Chrome extension (side panel + activity feed + CSS inspector)
├── lib/ # Shared libraries (aside-render.ts = local-HTML rendering, Aside first, engine fallback; design-catalog.ts = the typed design anti-pattern catalog every design skill renders from; design-detect-contract.ts = detector sentinel vocabulary; design-md.ts = open DESIGN.md reader/writer; dom-dump-script.ts + generated dom-dump.js = rendered-DOM dump for the detector; review-evidence.ts = review-start receipt binding and computed freshness; frontend-scope.ts; claude-bin.ts, error-handling.ts, worktree.ts, egress-receipt.ts, context-bill.ts, redact-engine.ts, tracker-guard.ts, version-source.ts, code-intelligence/)
├── lib/ # Shared libraries (aside-render.ts = local-HTML rendering, Aside first, engine fallback; design-catalog.ts = the typed design anti-pattern catalog every design skill renders from; design-detect-contract.ts = detector sentinel vocabulary; design-md.ts = open DESIGN.md reader/writer; dom-dump-script.ts + generated dom-dump.js = rendered-DOM dump for the detector; review-evidence.ts = review-start receipt binding, computed freshness and shared-code snapshot eligibility; frontend-scope.ts; claude-bin.ts, error-handling.ts, worktree.ts, egress-receipt.ts, context-bill.ts, redact-engine.ts, tracker-guard.ts, version-source.ts, code-intelligence/)
│ └── diagram-render/ # Vendored mermaid + excalidraw runtimes, built into one offline bundle the renderer loads
├── patches/ # bun `patchedDependencies` patches (playwright-core windowsHide)
├── docs/designs/ # Design documents (incl. IMPECCABLE_INTEROP.md = the design detector / catalog / open DESIGN.md record, and fork-port-residual-2026-09/ evaluation evidence)
+86 -17
View File
@@ -127,6 +127,21 @@ on a Mac they drive Aside and on Linux CI they drive the built browse binary,
skipping only when neither exists. The `$B`-driven E2E cases and `browse/test/`
run on every platform as before, so Linux CI proves the fallback engine live.
**Bootstrap dependency retention is opt-in qualification, not the behavior test.**
`qa-bootstrap` still runs its original unpinned Vitest installation and assertions
on macOS and Linux, including documented unsharded commands. Only a Linux paid
shard runner issues the owned retention scope: it binds each actual fixture and
native lifetime, retains locks, package manifests and the installed file/link
inventory, and acknowledges capture before deleting the fixture. The outer
runner also captures evidence when a callback is killed. Incomplete capture
fails qualification and preserves the source fixture as well as partial evidence.
Other platforms explicitly report retention as unavailable and still execute the
native behavior test. A run without the runner-issued scope earns no retained
dependency qualification credit; candidate acceptance requiring that evidence
must use the Linux sharded path and verify every attempt's complete capture,
acknowledgment and cleanup fallback. A passing unsharded or macOS behavior test
does not substitute for that evidence.
**The renderer picks the same way, so the render gates are engine-agnostic.**
`/make-pdf`, `/diagram`, and design previews print and screenshot their local
HTML through `lib/aside-render.ts` / `bin/gstack-render.ts`, which render in
@@ -177,13 +192,35 @@ fallback; unknown files get 75th-percentile pessimism, and both full-suite and
the long pole. Packed
shards get duration-aware walls (`max(base, predicted × 3, files × 5s)`). The
legacy `--shards N --shard i` path keeps stable hash indices. Required CI uses
one duration-packed `--ci-plan`, 20 isolated `--ci-run` machines, and a
`--ci-verify` aggregate. `TREE_MUTATING` is EMPTY:
`gen-skill-docs.ts` has a `main()` guard (imports never regenerate; pinned by
`test/gen-skill-docs-import-purity.test.ts`) and `--out-dir` renders every
host, so all former mutators render into mkdtemps and the trailing serial
shard is gone. The map remains a mechanism — a test that genuinely must write
shared artifacts in place earns a reasoned entry and is serialized again.
one duration-packed `--ci-plan`, 20 ordinary `--ci-run` shards plus a separate
exclusive-fixture shard, and a `--ci-verify` aggregate. CI shards still run on
independent machines without ordering unrelated jobs. The public
`TREE_MUTATING` map now classifies exclusive host-state fixtures; its sole
entry is `test/bootstrap-retention.test.ts`, whose same-UID nondumpable actors
affect host-wide procfs permission checks. Locally, this file runs only after
all parallel shards settle, and cancellation prevents that final phase from
starting. No selected files, retries, budgets, or receipt requirements are
removed. Former generator mutators still render into private output directories;
this serial phase protects process visibility, not in-place doc generation.
Full child output is retained in private files under `.context/free-test-logs/`,
outside each shard's temporary cleanup directory. The runner prints the path at
launch and completion. Losing the log fails the run even when the child exits
successfully. A redirected log directory is rejected before launching a child.
On failure, read that log first: the recovery message distinguishes incomplete
capture, unconfirmed cleanup, deadline expiry and a test/module failure. Fix the
demonstrated cause before rerunning. A focused `bun test` command is offered only
when every failure is attributable to existing selected files; it proves that
repair, not completion of the original selection. Preserve failed attempts when
sharing results, and inspect logs for private data before sharing them.
Before publication, classify new deterministic regressions for quick feedback.
Refresh the timing seed with the existing recorder on fixed inputs; do not edit
source while tests run. Critical boundary controls belong in `QUICK_CORE` when
their feedback cost is justified. Other measured files qualify at two seconds
or less; slow and unmeasured files remain outside quick, not outside full tests.
Report cold setup separately from warm execution, while retaining failed-attempt,
retry and cleanup time in the total cost.
**PTY fixture timing.** Plan-count sessions wake on terminal output or exit,
with at least 250ms between expensive observations and a 2s fallback for
@@ -230,6 +267,21 @@ key must name a living paid test (`test/touchfiles.test.ts`'s reverse
invariant), and `git show <sha>:path` fixtures are banned — vendor the bytes
instead (`test/git-ref-fixture-tripwire.test.ts`).
Functional QA and documentation acceptance require an explicit `EVALS_RUN_ID`;
`GSTACK_EVAL_DIR` alone does not satisfy their evidence-ownership guard. For each
local invocation, supply a fresh ID to the documented detached runner:
```bash
EVALS_RUN_ID="local-$(bun -e 'console.log(crypto.randomUUID())')" bun run eval:bg:pr
```
Neither package scripts nor `gstack-detach` invent this identity. CI's PR/manual
slices, periodic slices and weekly gate census supply an ID bound to the workflow
run, attempt, job and slice. Existing slice artifacts retain per-shard snapshots;
separate always-run `native-captures-<EVALS_RUN_ID>` artifacts retain project/legacy
`e2e-runs` and `evals/qa-callers` evidence for 90 days. These diagnostic artifacts
are not collector results and do not establish that an unfinished test passed.
**Fast PR profile and evidence reuse.** `test:pr` selects the changed cases in
`scripts/test-pr-profile.ts` plus every changed quality judge. `--profile full`
retains the broad census; no case IDs or tier assignments are removed. The plan
@@ -246,8 +298,9 @@ with matching before/after inputs. The audited workflow-judge adapter hashes the
actual expanded prompt, source/fixture/rubric/runner closure, installed SDK,
model parameters and runtime. Missing/unknown inputs force execution. Receipts
are scoped to the same repository and PR, expire after 24 hours, and contain
public scores and provenance rather than prompts or secrets. Only the 14 cases
using `runWorkflowJudge` are eligible; the other 11 quality cases remain fresh.
public scores and provenance rather than prompts or secrets. Of the 17 cases
using `runWorkflowJudge`, 16 are eligible; the cookie workflow's custom input
does not match the cache adapter and stays fresh, as do the other 11 quality cases.
CI supplies the scoped cache/runtime configuration; local runs are fresh by
default. Cached scores must
pass current assertions; reused records retain their original source and time
@@ -325,12 +378,24 @@ including two minutes for cleanup. No per-case budget grows. Overlay wrappers
have a 1,830-second minimum shard wall and run without Bun retries; see the
[overlay contract](OVERLAY_BENCHMARK_CONTRACT.md) for their unchanged work budget.
The quality file reserves 6,400 seconds for all 25 cases and their existing
retry, plus cleanup. Each still has 120 seconds of model work. Its 14 workflow
The quality file reserves 7,180 seconds for all 28 cases and their existing
retry, plus cleanup. Each still has 120 seconds of model work. Its 17 workflow
judges own their deadline and abort signal, with five seconds for terminal
recording inside a ten-second Bun grace; the other 11 retain their existing
120-second Bun timeout. Late responses cannot create records or cache passes.
The ship documentation file reserves 10,920 seconds for five 600-second cases and
eight 300-second fault cases, each with one retry, plus cleanup. The standalone
documentation child retains its 600-second case. The five review/ship explorer
cases reserve 3,270 seconds including their existing retry and finalization grace.
These are whole-file supervision limits, not additional model work per case.
The shared-library path file reserves 3,720 seconds for its three serial
600-second cases, each with one retry, plus 120 seconds for cleanup. Its
registered budget keeps the file in its own shard and binds the expected wall
to both the saved plan and the execution receipt; missing or stale budget
records fail reconciliation. Case deadlines, model budgets and retries do not grow.
`resolvePaidShardBudget(files, overrideMs?)` is the canonical per-job resolver.
Autoplan, each registered finding file, and each overlay wrapper require their
own shard, even with `--files-per-shard` above one. Mixed or multi-file overlay
@@ -341,20 +406,24 @@ Planner entries and execution results record the effective wall,
its source and policy identifier. Custom drivers must resolve each job instead
of passing their ordinary 1800-second default as an explicit Autoplan cap;
their outer controller/detach wall must also cover the allocated work and cleanup.
`eval:bg:pr` and `eval:bg:periodic` have 72000/66000-second outer caps; the PR
The current paid census has 122 files: 61 gate-tier and 103 periodic-tier.
`eval:bg:pr` and `eval:bg:periodic` have 92820/67380-second outer caps; the PR
wrapper covers a full-gate fallback at its default two workers. The broad gate
wrapper reserves 33600 seconds, and release reserves 100000 seconds for both
wrapper reserves 49320 seconds, and release reserves 116700 seconds for both
tiers. Legacy monolithic
`eval:bg`/`eval:bg:all` retain their shorter 5400/7200-second caps and do not
promise two complete Autoplan attempts; use the sharded periodic path for this policy.
Periodic CI plans `--slices 8 --autoplan-slice`: the eighth runs only Autoplan.
When overlays are selected, the seventh is reserved for their serial wrappers;
Periodic CI plans `--slices 9 --autoplan-slice`: the ninth runs only Autoplan.
When overlays are selected, the eighth is reserved for their serial wrappers;
registered finding files are distributed across the remaining ordinary slices
by their supervised walls. Each slice job has a 355-minute cap; Autoplan retains
by their supervised walls. Each slice job has a 360-minute cap; Autoplan retains
its 172-minute shard wall. Reconciliation rejects missing, duplicated or misplaced
registered work and absent budget records. The weekly gate census has a
350-minute cap and PR slices have a 220-minute cap. Free supervision tests
352-minute cap across eight single-worker slices with at most four running at
once. Its longest current work wall is 302 minutes. PR slices retain seven
two-worker slices with a 265-minute cap for their 242-minute work wall plus
setup. Free supervision tests
verify these bounds against the complete current census, configured retries,
and setup reserve. Ordinary paid tiers and the default 1800-second
shard wall remain unchanged; the registered and overlay policies above supply
+83
View File
@@ -27,6 +27,47 @@ Overlay efficacy experiments retain their full fixture/model/arm/trial matrix.
Security cases retain their source, path, socket, process and lease identities.
These are distinct scenario dimensions, not repeated work to delete.
## Functional QA contract map
The deterministic owners below protect the failure boundary; their live partners
prove that an agent follows it. A shared fixture or captured event does not replace
an independent live trial. All free owners run in `bun run test`; quick eligibility
depends on measured duration or an explicit `QUICK_CORE` entry, not this table.
| Contract | Deterministic owner | Necessary live boundary | Host and lane |
| --- | --- | --- | --- |
| CLI/API/webhook QA without browser setup | `qa-functional-fixture`, `qa-functional-evidence`, `qa-lazy-sections` | `skill-e2e-qa-functional`: CLI and webhook report sessions | Linux/macOS free; selected PR gate; Windows only where curated |
| Report-only preserves local and remote authority | `qa-only-capability`, `qa-functional-observer`, `qa-functional-observer-atomic`, `qa-caller-authority` | Independent report-only sessions with synthetic owned endpoints/auth | Linux kernel observation; free callback controls plus selected PR gate |
| Repair reproduces the defect, adds a failing regression and rechecks adjacent behavior | `qa-fix-loop-fixture`, `qa-functional-evidence` | `skill-e2e-qa-functional-fix`: CLI and webhook repair sessions | Free controls plus selected PR gate |
| Review and Ship actually explore | `qa-exploratory-callers`, `qa-caller-report-observer`, `qa-checkpoint-evidence` | `skill-e2e-qa-callers`: actual Review/Ship callers | Free captures plus selected PR gate |
| Smoke expiry preserves required plan checks | `qa-deadline`, `qa-deadline-selection`, `qa-browser-deadline-evidence` | `ship-exploratory-plan-checks` | Free deadline/dispatch controls plus selected PR gate |
| Late changes invalidate affected results | `qa-caller-freshness-order`, `qa-deadline-publication-observer`, `shared-libs-revalidation-prompt` | `ship-exploratory-late-input` and the existing late-input documentation handoff | Free stale-input controls plus selected PR gate |
| Documentation completes before publication and respects protected files | `docsync-authority`, `docsync-atomic-writes`, `docsync-report-interface`, `docsync-lifecycle-interface` | `skill-e2e-ship-docsync`, `skill-e2e-docsync-spawned` | Free state/permission controls; registered gate/periodic scenarios retain their tiers |
| Cancellation drains owned work before another attempt | `shared-libs-cancellation`, `session-runner-stream-lifecycle`, `agent-sdk-runner`, `paid-shard-settlement` | Existing actual shared-library/SDK caller scenarios | Free real-callback/process controls; registered live gate/periodic trials remain independent |
| Missing tools or incomplete results never become verified coverage | `qa-probe-gates`, `qa-supervision-selection`, `test-free-shards`, `test-free-shards-capture`, `paid-shards` | `ship-exploratory-unavailable` and existing reporting-boundary sessions | Free negative controls plus selected PR gate; unsupported hosts remain unexecuted |
Names without a suffix refer to `test/<name>.test.ts`. Keep missing, stale,
duplicate, selected-but-unstarted, malformed/truncated and observer-overflow
controls distinct from legitimate empty selections. File restoration cannot
replace write observation, and a clean local tree cannot prove that an external
request made no mutation. Fixture endpoints and credentials must be synthetic
and owned; specifically authorized functional requests remain permitted.
Functional fixtures register their existing closed command policy as a native
PreToolUse hook, so an unsupported request is refused before execution. The
callback regression invokes the registered command with native hook input,
observes an isolated mutation target and permits the owned webhook positive
control. This is a command boundary, not a sandbox for arbitrary target code.
Its private CLI configuration is outside the observed product tree, and the
fixture's existing cleanup owns both directories.
Review/Ship observations now use the same strict native event decoder as QA
checkpoints and documentation. Caller-specific handoff/freshness interpretation
stays separate. Original missing, orphaned and duplicate-call controls were run
before replacing three incidental error-wording assertions with rejection checks;
the existing positive attribution case still runs, and a completed-ID reuse
negative control prevents incomplete evidence from becoming green.
## Complete inventory, not just the fast subset
At the audited revision, all 1,124 tracked Bun test files partition into 1,010 free
@@ -104,6 +145,41 @@ case/sample inventory and report skips and unavailable platforms separately.
Do not subtract failures from elapsed time or use a smaller selection as proof
that the complete suite got faster.
### Functional-QA cleanup measurement — September 28, 2026
On the same four-CPU Linux machine, using Bun 1.4.0, Node 22.20.0 and Claude
Code 2.1.251, the existing duration recorder measured all 1,113 free files.
The refreshed seed selects 931 files for quick feedback: 90 newly included and
20 newly excluded by measured cost, a net increase of 70. No files remain
unclassified. All 182 slow files remain in the complete suite. The functional
command observer, checkpoint decoder and log-capture controls are explicit
quick-core cases; each measured under two seconds.
| Existing command / attempt | Executed scope | Result | Wall time |
| --- | --- | --- | ---: |
| `bun run test:free --record-durations` | 1,113 files | 29,175 pass, 5 fail, 131 skip | 680.64s |
| `bun run test:quick`, first measured attempt | 931 files | 22,160 pass, 2 fail, 100 skip | 125.39s |
| `bun run test:quick`, repaired attempt | The same 931 files | 22,162 pass, 0 fail, 100 skip | 52.43s |
The profile's five failures came from the machine's Git identity wrapper
overwriting synthetic fixture authors. Running the two affected files with native
Git in the isolated test environment passed all 59 tests in 75.32s; normal checkout
commits retained the configured identity. Both quick attempts used that corrected
environment. Their two telemetry timeouts used Bun's synchronous piped-input
path; the repair reuses the existing file-backed command capture helper without
changing commands, assertions or deadlines. The seed retains observed costs,
including failed attempts; it is a scheduling hint, not a passing receipt.
Cold dependency installation took 0.477s and the integrated build took 3.84s,
separate from warm test execution; CLI installation was not independently timed.
An earlier 63.37s profile was cancelled for a decoder repair, with an additional
scoped browser cleanup, and earns no completion credit. Failed, cancelled and
repair runs are costs, not time removed from the workflow. The quick target of
one minute was met on this machine, but these measurements establish neither a
cross-environment speedup nor full release, live-model or Windows acceptance.
### Earlier component comparisons
Measured component comparisons:
| Workload | Before | After | Coverage retained |
@@ -167,6 +243,13 @@ not a fresh full-census runtime improvement.
## Evidence validity
After integrating main's September 28 Ubicloud improvements, the scheduling seed
uses upstream's CI-environment timings for shared files and preserves the 52
previously measured branch-only entries. These are scheduling hints from two
machines, not a matched performance comparison or acceptance result. Refresh
the whole seed with `bun run test:ubicloud --record-durations` when measuring a
new common baseline; do not infer a speedup by adding these measurements.
Check the executable actually used by each SDK, print-mode and terminal launcher.
A CLI version cached during preflight does not prove the version used by later
sessions if PATH contents change. Use native session-init versions, terminal
+1 -1
View File
@@ -75,5 +75,5 @@ A coverage map written in Diataxis terms gives you a deterministic answer to "di
- **Reference for the skill that implements this:** [`document-generate/SKILL.md`](../document-generate/SKILL.md)
- **Reference for the audit that uses this taxonomy:** [`document-release/SKILL.md`](../document-release/SKILL.md)
- **Tutorial for using `/document-generate`:** [`tutorial-document-generate.md`](./tutorial-document-generate.md)
- **How-to: document a shipped feature:** [`howto-document-a-shipped-feature.md`](./howto-document-a-shipped-feature.md)
- **How-to: document a feature before shipping:** [`howto-document-a-shipped-feature.md`](./howto-document-a-shipped-feature.md)
- **Diataxis homepage:** https://diataxis.fr/ — Procida's canonical reference for the framework
+15 -15
View File
@@ -1,14 +1,14 @@
# How to document a feature you just shipped
# How to document a feature before it ships
This is the post-ship workflow: you merged a PR, the docs are stale, and you want a coverage map plus filled gaps in one pass. You'll run `/document-release` to audit, then `/document-generate` to fill the gaps it finds.
This is the pre-merge documentation workflow: the feature is implemented and you want to audit coverage and fill gaps. `/ship` already runs the relevant documentation audit before publication; use standalone `/document-release` on a committed feature branch to revisit it, then `/document-generate` for missing pages.
## Prerequisites
- gstack installed (`./setup` complete; verify with `which gstack` or by typing `/` in Claude Code and seeing skills listed)
- The branch with your shipped feature is checked out
- A PR exists on GitHub or GitLab (recommended — the workflow updates the PR body with a coverage map)
- The committed feature branch is checked out, before merge
- Optional: an existing GitHub or GitLab PR lets standalone `/document-release` update its body with the coverage map
If no PR exists yet, run `/ship` first to create one; that's what `/document-release` is designed to run against.
No PR is required to audit. `/ship` runs its audit before creating or updating the PR; a standalone invocation without a PR skips the PR-body update.
## Steps
@@ -30,13 +30,13 @@ Coverage map:
FooProcessor ❌ ❌ ❌ ❌
```
Items with zero coverage are **critical gaps**. Items with only reference coverage are **common gaps**. Both land in the PR body as a `### Documentation Debt` subsection so reviewers see them.
Items with zero coverage are **critical gaps**. Items with only reference coverage are **common gaps**. The audit reports both; when a PR exists, it also adds a `### Documentation Debt` subsection for reviewers.
If `/document-release` reports everything is covered, you're done. Skip the rest of this how-to.
### 2. Read the documentation debt section in the PR body
### 2. Read the reported documentation gaps
Open your PR (the skill prints the URL). Scroll to `## Documentation` → `### Documentation Debt`. Each item is tagged with the Diataxis quadrant that would fill it:
Use the audit's coverage map and gap summary. If a PR exists, open `## Documentation` → `### Documentation Debt` in its body. Each item is tagged with the Diataxis quadrant that would fill it:
```
### Documentation Debt
@@ -67,23 +67,23 @@ Re-run `/document-release`:
/document-release
```
The coverage map should now show the previously-flagged entities with green checkmarks in the previously-empty quadrants. The PR body's Documentation Debt section should be empty or reduced to items you intentionally deferred.
The coverage map should now show the previously-flagged entities with green checkmarks in the previously-empty quadrants. Reported documentation debt, including the PR-body section when present, should be empty or reduced to items you intentionally deferred.
## Verification
Open your PR and confirm:
Read the audit output and, when present, the PR body. Confirm:
1. The PR body has a `## Documentation` section with a doc-diff preview.
2. The `### Documentation Debt` subsection lists zero critical gaps (or only items you knowingly deferred).
1. The audit summarizes the docs reviewed and changed; an existing PR has a `## Documentation` section with a doc-diff preview.
2. The reported documentation debt lists zero critical gaps (or only items you knowingly deferred).
3. Each generated doc file in `docs/` opens cleanly and cross-links to siblings (reference → how-to → tutorial → explanation).
4. Run `grep -rE '\]\([^)]*\.md\)' docs/` and verify no link points to a missing file.
If all four check, your PR is ready to land with complete documentation.
These checks complete the documentation pass, with any deferred gaps recorded. They do not replace `/ship`'s code review, tests or final verification.
## Troubleshooting
**`/document-release` reports "No public surface changes detected."**
The diff is internal-only (refactors, tests, infra). No docs are needed. Skip to landing.
There may be no new public surface, but still check affected setup, testing, architecture and workflow instructions. A completed audit can report current documentation; an empty public-surface map alone is not that audit.
**The Diataxis quadrant tag on a gap doesn't match what you'd expect.**
The skill uses an entity taxonomy to decide which quadrants matter (CLI flags want reference + how-to; internal modules want reference + explanation; user-facing features want all four). If you disagree, you can override by hand-editing the docs after generation. The audit is a guide, not a constraint.
@@ -92,7 +92,7 @@ The skill uses an entity taxonomy to decide which quadrants matter (CLI flags wa
Tutorials should hit a working result in 3 steps or fewer. Re-run the skill and ask it to compress, or hand-edit. The Step 8 Quality Self-Review catches some of these but not all.
**You want to document a feature but no PR exists yet.**
Run `/ship` first to create the PR, then this workflow. Without a PR, `/document-release` can still audit but skips the PR-body update.
Run standalone `/document-release` on the committed feature branch; it can audit without a PR and skips the PR-body update. Or run `/ship`, which includes the audit before publication.
**A generated reference doc has hallucinated API signatures.**
File a bug. The skill's Step 1 archaeology is supposed to read implementation files end-to-end, not just signatures, specifically to prevent this. Include the generated text and the actual code so we can trace why the archaeology missed it.
+85
View File
@@ -0,0 +1,85 @@
# QA deadlines
Bounded exploratory QA uses `bin/gstack-qa-deadline` from the installed gstack
runtime. Browser Quick keeps its 30-second limit; browser Full/Regression uses the
15-minute maximum of its 5–15-minute exploration window. Review/ship smoke keeps its
5-minute or 12-probe limit, whichever comes first. The workflow enforces the probe
count; the helper enforces elapsed time. Required plan checks are outside the smoke
guard: run them after smoke with the same checkpoint sequence and finite command
timeouts capped by the caller's remaining deadline. An expired caller deadline
leaves checks not-run; never restart the smoke clock to run them.
Functional Full/Quick/Regression has no default total exploration deadline; Quick
limits scope to success plus the highest-risk changed edge. A caller's stricter
duration or absolute deadline still bounds the run. Standalone mixed runs use
owned `REPORT_DIR/browser` and `REPORT_DIR/functional` directories for their clocks
and checkpoints, with one final report at `REPORT_DIR`. Create those directories
before starting their clocks. Single-surface runs and review/ship smoke keep their
clock and checkpoints at `REPORT_DIR`; fixed caller paths take precedence.
## Command interface
Run the helper with Bun. `FILE` is `deadline.json` inside the invocation-owned,
canonical probe directory; its parent must already exist. Use quoted absolute
paths in place of `GUARD` and `FILE` below.
```text
bun GUARD start FILE SECONDS [EARLIER_UTC]
bun GUARD status FILE
bun GUARD run FILE -- COMMAND ARGS...
```
`start` runs once, immediately before the baseline. It exclusively creates a
versioned, read-only receipt and clamps the selected duration to an earlier caller
deadline when supplied. It does not replace an existing file. `status` reads the
actual clock. `run` checks the same receipt again before launching, then supervises
the command for the remaining time. Replays and minimization use that same deadline;
never replace the receipt or restart the timer to finish more work.
Arguments are passed directly, without shell evaluation. For a permitted script,
the child command is `bash -c 'script'`; keep every probe inside that child rather
than appending an unguarded command after the helper. Missing or malformed state,
symlinked paths and unavailable process containment block dispatch.
The child's stdout/stderr remain its evidence. Guard-owned lines begin with
`QA_DEADLINE ` and contain separate JSON bookkeeping; do not copy them into the
checkpoint's observed program JSON. Mark a refused next probe not-run in the report;
preserve its original checkpoint rather than rewriting it as an observation.
For JSON-emitting probes, `observed` is the decoded child JSON itself, not a `child`
envelope or a mixture of results and guard metadata. For other output, retain the
full child text. Command fields retain the complete outer command, including the
guard invocation; guard diagnostics and interpretations belong in the report.
## Reporting measurements
The configured probe budget is not the total session duration. Child launch/finish
receipts measure guarded command spans; gaps between calls do not measure individual
tool costs or establish how many probes can fit in another run.
`/qa-only` loads its reporting section after probing stops and checks every repeated
finding against the retained evidence before writing. Its in-progress report marks
total session elapsed as unmeasured: the final Write and cleanup have not finished.
Initial charters and final findings use the caller's same report file. Learning notes
and automatic memory obey the same caller-authorized write destinations.
An optional measured interval names its actual start/end receipts and excluded work,
including later report Writes and cleanup; it is not a completed-session measurement.
## Exit and cleanup behavior
- Guard expiry or timeout returns 124. A child can independently return 124 too;
use the guard receipt's event and `timedOut` field to distinguish those cases.
- Guard errors return 2; a missing executable returns 127. Otherwise the child's
status is preserved.
- On Linux/macOS, cleanup covers the command's inherited process group. Detached
or new-session descendants are outside that guarantee, so detached probes are
unsupported. Force-killing the guard itself with SIGKILL also prevents its POSIX
cleanup handler from running.
- On Windows, dedicated nested Jobs contain the probe worker and its descendants,
including children whose immediate parent exits. Failure to initialize this
containment prevents the command from starting.
- Receipt flushing happens after probe cleanup and can take up to five seconds;
blocked or broken output returns 2. This allowance does not extend probe work.
Terminating a probe client does not undo a request already accepted by a service
or stop an already-running browser. Preserve any known partial effects and report
uncertain completion instead of assuming cancellation meant no effect. Report
writing may finish after the exploration deadline, but new probes may not start.
+27 -13
View File
@@ -15,16 +15,16 @@ Detailed guides for every gstack skill — philosophy, workflow, and examples.
| [`/design-review`](#design-review) | **Designer Who Codes** | Live-site visual audit + fix loop. 80-item audit, then fixes what it finds. Atomic commits, before/after screenshots. |
| [`/design-shotgun`](#design-shotgun) | **Design Explorer** | Generate multiple AI design variants, open a comparison board in your browser, and iterate until you approve a direction. Taste memory biases toward your preferences. |
| [`/design-html`](#design-html) | **Design Engineer** | Generates production-quality Pretext-native HTML. Works with approved mockups, CEO plans, design reviews, or from scratch. Text reflows on resize, heights adjust to content. Smart API routing per design type. Framework detection for React/Svelte/Vue. Previews render through your Aside browser. |
| [`/qa`](#qa) | **QA Lead** | Test your app, find bugs, fix them with atomic commits, re-verify. Auto-generates regression tests for every fix. |
| [`/qa-only`](#qa) | **QA Reporter** | Same methodology as /qa but report only. Use when you want a pure bug report without code changes. |
| [`/qa`](#qa) | **QA Lead** | Explore browser and functional behavior (APIs, CLIs, jobs, workers, webhooks), reproduce defects, prove regressions fail before repair, then fix and re-verify. |
| [`/qa-only`](#qa) | **QA Reporter** | Explore the same surfaces and propose regression cases with evidence, without changing product code or tests. |
| [`/scrape`](#browse) | **Browser Data Extractor** | Pull structured data off a web page — tables, lists, prices — in your Aside browser with the page's real logged-in state. Same driver contract as `/browse`. On the fallback browser, a codified browser-skill answers a repeat intent in ~200ms. |
| [`/skillify`](#browse) | **Skill Codifier** | Fallback-browser skill: walks back through your conversation, finds the last `/scrape` prototype, synthesizes script + test + fixture, runs the test, asks before committing. On Aside, durable per-site automation belongs to Aside's own skills. |
| [`/ship`](#ship) | **Release Engineer** | Sync main, run tests, audit coverage, push, open PR. Bootstraps test frameworks if you don't have one. One command. |
| [`/ship`](#ship) | **Release Engineer** | Sync main, run tests, explore changed behavior within a bound, audit coverage and docs before final verification, then push and open or update a PR. Bootstraps test frameworks when appropriate. |
| [`/land-and-deploy`](#land-and-deploy) | **Release Engineer** | Merge the PR, wait for CI and deploy, verify production health. One command from "approved" to "verified in production." |
| [`/canary`](#canary) | **SRE** | Post-deploy monitoring loop. Watches for console errors, performance regressions, and page failures in your Aside browser. |
| [`/benchmark`](#benchmark) | **Performance Engineer** | Baseline page load times, Core Web Vitals, and resource sizes. Compare before/after on every PR. Track trends over time. |
| [`/cso`](#cso) | **Chief Security Officer** | Supported security findings with explicit coverage. Static assessment remains available without catalog profiles; contained runtime/scanner execution requires matching qualified profiles. Runtime-tested bundles authenticate separate external assertions. Project-test completion remains `self_reported` because target code controls the test process; `tested` is reserved for a future target-independent completion witness. |
| [`/document-release`](#document-release) | **Technical Writer** | Update all project docs to match what you just shipped. Catches stale READMEs automatically. |
| [`/document-release`](#document-release) | **Technical Writer** | Audit relevant docs on every ship before final verification; standalone runs can also update docs after a PR exists. Catches stale READMEs and reports unresolved gaps. |
| [`/document-generate`](#document-generate) | **Technical Writer** | Generate Diataxis docs (tutorial / how-to / reference / explanation) for a feature from code. |
| [`/retro`](#retro) | **Eng Manager** | Team-aware weekly retro. Per-person breakdowns, shipping streaks, test health trends, growth opportunities. |
| [`/browse`](#browse) | **QA Engineer** | Give the agent eyes. Drives your Aside browser first — real sessions, real clicks, real screenshots — through deterministic `aside repl` scripts, and falls back to gstack's own Chromium (~100ms per command) when Aside isn't there. |
@@ -621,18 +621,30 @@ This is my **QA lead mode**.
`/browse` gives the agent eyes. `/qa` gives it a testing methodology.
The most common use case: you're on a feature branch, you just finished coding, and you want to verify everything works. Just say `/qa` — it reads your git diff, identifies which pages and routes your changes affect, opens them in tabs of your Aside browser, and tests each one. No URL required. No manual test plan.
The most common use case: you're on a feature branch, you just finished coding, and you want to verify everything works. Just say `/qa` — it uses your request, repository contracts, test plan and diff to select browser, functional (API, CLI, job, worker or webhook), or mixed surfaces. No URL or manual test plan is required. Browser targets still open affected pages in Aside tabs (or gstack's fallback browser); functional-only targets use documented native commands and isolated local fixtures without starting a browser.
Four modes:
Choose Full, Quick or Regression depth; diff-aware selects what to test:
- **Diff-aware** (automatic on feature branches) — reads `git diff main`, identifies affected pages, tests them specifically
- **Full** — systematic exploration of the entire app. 5-15 minutes. Documents 5-10 well-evidenced issues.
- **Quick** (`--quick`) — 30-second smoke test. Homepage + top 5 nav targets.
- **Regression** (`--regression baseline.json`) — run full mode, then diff against a previous baseline.
- **Diff-aware** (automatic on feature branches) — selects changed and adjacent behavior. Standalone `/qa` first resolves a dirty working tree through its commit/stash/abort question; it tests the resulting checkout. For browser targets it identifies affected pages and tests them specifically.
- **Full** — browser QA systematically explores the entire app (typically 5-15 minutes, documenting 5-10 well-evidenced issues); functional QA covers applicable documented contracts and reports blocked or untested ones separately.
- **Quick** (`--quick`) — browser QA keeps its 30-second homepage + top-five-navigation smoke; functional QA checks a successful operation and the highest-risk changed edge, marking other contracts not run.
- **Regression** (`--regression <previous-report-or-baseline>`) — browser QA runs full mode and diffs against a previous `baseline.json`; functional QA requires a readable prior functional report and replay evidence, repeats its failed probes against the intended contract, then checks changed adjacent behavior. A browser-only baseline is not a functional baseline.
Exploration retains a written trail: before each next discovery probe, QA saves an
`exploration-NNN.json` checkpoint in its owned report directory with the previous
command and result, the hypothesis and the next exact command. The final report
links those files. `/qa-only` and the bounded review/ship pass use the same evidence
contract without gaining permission to edit product code or tests.
Time limits include checkpoint and evidence work; unfinished probes remain untested.
New runs preserve prior reports and baselines, using a fresh owned run directory when
the selected output directory already contains artifacts. Mixed runs put browser and
functional results in separate sections of one report; browser scores never apply to
functional coverage. Conflicting Quick/Regression requests are resolved before probing.
### Automatic regression tests
When `/qa` fixes a bug and verifies it, it automatically generates a regression test that catches the exact scenario that broke. Tests include full attribution tracing back to the QA report.
For a reproduced defect, `/qa` writes a native regression test when infrastructure is available and proves it fails for that defect before the repair; CSS-only defects may use browser evidence instead. After the root-cause repair, it requires the original probe, adjacent happy path and native regression when available to pass before calling the fix verified. Tests trace back to the QA report. `/qa-only` can propose the case and retain replayable evidence but never changes product code or tests; missing native test infrastructure remains an explicit coverage limit, not permission to install a new framework for functional QA.
### Example
@@ -673,9 +685,11 @@ If your project doesn't have a test framework, `/ship` sets one up — detects y
Every `/ship` run builds a code path map from your diff, searches for corresponding tests, and produces an ASCII coverage diagram with quality stars. Gaps get tests auto-generated. Your PR body shows the coverage: `Tests: 42 → 47 (+5 new)`.
`/review` and `/ship` also run a bounded exploratory pass on changed behavior and nearby risks, even for a small diff without a plan or web server. Their existing approval and test rules govern any fixes or permanent tests; a blocked probe remains a coverage gap, not a passing QA result.
### Review gate
`/ship` checks the [Review Readiness Dashboard](#review-readiness-dashboard) before creating the PR. If the Eng Review is missing, it asks — but won't block you. Decisions are saved per-branch so you're never re-asked.
`/ship` displays historical review readiness in the [Review Readiness Dashboard](#review-readiness-dashboard) during preflight. A missing Eng Review is reported without an extra question; it does not replace or waive the current pre-landing review. Step 9 still runs the checklist, applicable specialists and bounded exploratory QA, with its existing approval and completion gates.
A lot of branches die when the interesting work is done and only the boring release work is left. Humans procrastinate that part. AI should not.
@@ -815,7 +829,7 @@ Claude: complete — assessed application routes, tenant authorization, secrets,
This is my **technical writer mode**.
After `/ship` creates the PR but before it merges, `/document-release` reads every documentation file in the project and cross-references it against the diff. It updates file paths, command lists, project structure trees, and anything else that drifted. Risky or subjective changes get surfaced as questions — everything else is handled automatically.
On every `/ship` run, including reruns and existing-PR updates, a ship-owned `/document-release` audit checks relevant authored docs against committed and selected uncommitted changes before the final commit, verification and publication. Clear factual corrections join the checked change; the ship parent owns versioning, Git and PR publication. A blocked or incomplete audit requires recovery or explicit acceptance of the named documentation risk before shipping, and never silently becomes current. You can still invoke `/document-release` standalone after a PR exists; that workflow retains its own approval, commit and PR-body steps.
```
You: /document-release
+1 -1
View File
@@ -138,5 +138,5 @@ Each one is short enough to maintain. Each one has a single job. The PR body sho
- **If you have gaps** /document-release flagged but didn't fill: run `/document-generate` again, scoped to those entities specifically.
- **If you want to understand why the four quadrants exist:** read [explanation-diataxis-in-gstack.md](./explanation-diataxis-in-gstack.md).
- **If you want to document one specific shipped feature** (not the whole project): read [howto-document-a-shipped-feature.md](./howto-document-a-shipped-feature.md).
- **If you want to document one specific feature before shipping** (not the whole project): read [howto-document-a-shipped-feature.md](./howto-document-a-shipped-feature.md).
- **Reference for the skill itself:** [`document-generate/SKILL.md`](../document-generate/SKILL.md).