v1.91.7.0 feat: add functional QA and pre-publication docs checks (#2983)

* feat: add surface-aware exploratory QA and ship documentation gates

* test: preserve delegated QA setup authority after main integration

* fix(qa): clarify exploration order and preserve report artifacts

* test(qa): follow the shared setup reference directly

* refactor(ship): make verification and recovery routes explicit

* test(ship): align evidence and review guards with explicit routes

* fix(workflows): clarify ship recovery and functional QA evidence

* fix(workflows): clarify approval recovery and full QA coverage

* refactor(workflows): order review transactions and clarify ship state

* fix(ship): clarify final verification and fail closed at publication

* fix(evals): attribute native atomic documentation writes

* fix(ship): clarify recovery and documentation lifecycle guidance

* fix(test): preserve observed native placeholder styling in CI

* fix(codex): report watchdog timeouts without a process-exit race

* Checkpoint functional QA implementation and workflow validation repairs

* Fix documentation and shared-review fixture contracts

* docs: clarify judge reuse and evaluation supervision

* test: align review evidence and selected case contracts

* test: verify append-only documentation checkpoints and recovery

* fix: qualify QA workflows and CI validation repairs

* fix: launch shared-libs fixture scripts on Windows

* fix: qualify QA deadlines, fixture isolation, and shard cleanup

* fix: preserve qualified QA and cancellation repairs

* fix: enforce functional fixture authority and share strict event decoding

* fix: retain free-test evidence and explain recovery

* fix: reject malformed native evidence after decoder consolidation

* test: use reliable capture for telemetry privacy filters

* test: refresh measured quick coverage and document validation costs

* Fix native fixture receipts and preserve VM validation evidence

* Align negative judge controls with upstream clarity policy

* Fix report-only QA preparation and public evidence handling

* Clarify QA-only preparation and current-report preservation

* Stream Ship quality judgments with an explicit 64k response contract

* Validate compact judge reasoning locally with supported wire schema

* Align functional QA fixture instructions with evidence acceptance

* Bind native browser diagnostics to execution evidence and align review verdicts

* Preserve native diagnostic line boundaries

* Serialize functional QA evidence from native captures

* Keep large QA evidence fixture payload out of Windows argv
This commit is contained in:
Garry Tan authored and GitHub committed 2026-09-29 06:07:35 -07:00
1 parent 65bfb0ce49
commit dcaea52800
333 files changed
+41755 -7357

No files matched your search

+104
View File
@@ -0,0 +1,104 @@
<!-- AUTO-GENERATED from exploratory.md.tmpl — do not edit directly -->
<!-- Regenerate: bun run gen:skill-docs -->
# Shared exploratory QA
The **caller** (/qa, /qa-only, /review or /ship) owns decisions, tests, fixes and publication. Discovery writes only reports/evidence
and owned fixture state; no workflows, framework installs or publication.
## 0. Preparation gate
Complete these Reads in order before writing charters or probing:
1. Read `sections/scope.md` relative to the installed `qa`/`gstack-qa` SKILL.md directory in full and select the surfaces.
2. Read the selected surface methods below in full.
Use this host's installed `qa`/`gstack-qa` SKILL.md directory for these reads:
**Functional surfaces:**
Read `sections/system-functional.md` in full.
**Browser surfaces only:**
Read `sections/qa-patterns.md` in full.
Await each successful Read result before continuing. A supplied target, isolation
description, section index or remembered method is not a completed instruction Read.
Do not repeat a Read already completed in this invocation; reuse only its acknowledged
full contents. If either required Read is missing, complete it now before Charter and preflight.
Missing or unreadable assets, prerequisites or permission block affected probes, not independent safe checks. Report QA setup blockers.
## 1. Charter and preflight
Reuse resolved REPORT_DIR; otherwise own a fresh `.gstack/qa-reports` subdirectory.
Write a **charter** per behavior: contract, risk, entrypoint, isolation, exit condition, source, commands and inputs. Save charters as Markdown in the report.
For /qa and /qa-only:
- Browser Quick: SECONDS=30. Browser Full/Regression: SECONDS=900.
- Functional Full, Quick and Regression have no default total timer.
Set SECONDS to the shorter mode/caller limit; an unlimited mode uses the caller's bound.
Without a total time limit, do not start D; announce finite command timeouts.
Stop when scoped contracts are tested or blocked.
Clocks/checkpoints use REPORT_DIR; mixed standalone runs use REPORT_DIR/browser and REPORT_DIR/functional, with one final report at REPORT_DIR. Caller paths win.
R = owned probe directory; D = R/deadline.json. Quote paths.
G = `$HOME/.claude/skills/gstack/bin/gstack-qa-deadline`; Q = `$HOME/.claude/skills/gstack/bin/gstack-qa-evidence`.
Start once before baseline: `bun G start D SECONDS [EARLIER_UTC]` if bounded.
EARLIER_UTC = caller's absolute deadline, if set.
Functional: `bun Q capture R NNN [--public] --deadline D -- COMMAND ARGS`.
Unbounded: use `--timeout-ms MS` instead. Use fresh three-digit IDs.
--public requires approved public/synthetic output; Q screens credentials. For complete private captures, await a safe Read of `R/.qa-evidence/NNN/observation.json`. Sensitive/incomplete captures cannot anchor checkpoints.
Bounded browsers: `bun G run D -- COMMAND ARGS`. No detached probes.
Never reset D/bypass G. Expiry or invalid/missing D stops probes; report unfinished coverage. QA_DEADLINE receipts are not observations.
## 2. Probe loop
Each probe is one native command/interaction plus checks, excluding bookkeeping.
Never batch probes.
1. First demonstrate success: output AND durable effects. Guard if bounded; await completion.
2. **Decide whether another probe is needed.** If bounded, run `bun G status D`.
If expired or no safe next probe remains, STOP exploration; write the report, not a checkpoint.
**Classify the last result before copying it.** For public or synthetic observations,
retain the entire result unchanged, including owned fixture paths, IDs, hashes and
existing credential placeholders. An absolute state path is not itself a secret.
For actual secrets/private payloads, withhold those values and disclose the redaction
and replay limits in the report. If no safe exact observation can be retained,
stop the affected probe chain; never invent a substitute path, identity or state.
**Publish before probing.** Create `exploration-NNN.json` in the probe directory, beside its deadline if bounded, with exactly four top-level fields:
observationCommand: last completed probe's full outer command, including guard.
observed: its exact decoded child JSON (no wrapper/extra keys), or its full non-JSON text.
For guarded text, copy the complete span between the guard's started and finished receipt lines.
Keep its whitespace and content fences verbatim. Do not summarize, relabel or add timing text.
The guard adds one newline before its finished receipt; that separator is not child text.
For unguarded text, copy the complete result instead.
If capture is incomplete, report that limit instead of reconstructing it.
hypothesis: why nextCommand. nextCommand: exact command/request, guarded if bounded.
Preserve every safe program-JSON key/value and identity hash unchanged.
Withhold unsafe values, disclose limits and stop that chain.
Check fields before publication. No drafts/placeholders or invented safe-path redactions; corrections cannot repair published notes.
Functional: `bun Q checkpoint R NNN CAPTURE_ID 'observationCommand' 'hypothesis' 'nextCommand'` with literal arguments. Q supplies observed; never transcribe it.
Browser checkpoints use Write.
Wait for successful checkpoint publication before dispatch.
Never backfill or overwrite notes.
3. Run that exact probe; G enforces the deadline when bounded.
Report refusals as not-run; retain initial state/inputs/results. Repeat from step 2.
4. Replay the exact failing command/request from the same initial fixture state via steps 2–3 (same native command, fresh capture ID)
to confirm it, then minimize via those gates. Expiry leaves confirmation/minimization incomplete.
Another input or a regression test is not that replay.
5. If the user or another process changes source, commands or fixtures, review the affected
contracts and return to step 2 for each affected revalidation. Do not make product changes yourself.
Keep the original limits/notes; update outcomes only from fresh evidence.
## 3. Parent handoff
Never change product code, tests, configuration, dependencies or Git through any tool,
including shell, rename, deletion, commit, stash or edit-then-restore. Return test_stub proposals
with their failing contract and expected assertion; never create tests or freeze buggy output.
## 4. Final report
Use the surface report template; link each checkpoint. Separate browser scores, functional outcomes and proposed/executed tests.
For evidence.json, Write R/annotations.json: {revision, runtime, cwd, evidence: [{capture, command, contract, expected, classification}], learning: [checkpoint IDs], limits}.
Run `bun Q materialize R annotations.json` before Markdown; Q fills observed/learning, not classifications. Retain all safe probes, including failures/replays; disclose withheld/incomplete evidence.
Evidence is invocation-local.
Missing prerequisites/expectations/observations, timeouts and refusal never pass.
Pass requires all required current-input contracts to pass with no required remainder.
Report blocked, inconclusive and not-run coverage without claiming success.
+1
View File
@@ -0,0 +1 @@
{{QA_EXPLORATORY}}
+19
View File
@@ -0,0 +1,19 @@
{
"$schema": "https://gstack.dev/schemas/section-manifest.json",
"skill": "qa-only",
"version": 1,
"sections": [
{
"id": "exploratory",
"file": "exploratory.md",
"title": "Report-only exploratory QA",
"trigger": "running selected report-only baseline and exploratory probes without product or test writes"
},
{
"id": "reporting",
"file": "reporting.md",
"title": "Evidence-grounded report finalization",
"trigger": "finalizing the report after probing stops"
}
]
}
+78
View File
@@ -0,0 +1,78 @@
<!-- AUTO-GENERATED from reporting.md.tmpl — do not edit directly -->
<!-- Regenerate: bun run gen:skill-docs -->
# Finalize a report from retained evidence
Complete these steps before the final report Write. They use retained results, not
new probes. Missing evidence stays unknown; an expired clock stays expired.
The caller's write boundary includes reports, learning notes and automatic memory.
Keep all of them in authorized destinations; a memory feature grants no extra path.
## 1. Establish each finding once
Ground the report and learnings in retained observations. For every finding, distinguish
the observed result, the expected contract and any untested causal hypothesis. Link the
supporting command/result or screenshot; unknown impact remains unknown. A console error
message does not establish an uncaught exception, failed payload or missing UI. Missing
text in a page-text extract does not establish an absent attribute or inaccessible element.
Verify those claims with an appropriate probe, or leave them unconfirmed when time expires.
Give each finding one ID and write its Observed, Expected, Evidence and Confirmation
fields first. A logged exception-shaped string proves a logged message, not that the
named operation executed. Keep possible causes in a separate Hypothesis field; omit
speculation that does not help the next investigation. Observed-once is not replay-confirmed.
## 2. Fill timing fields from their actual boundaries
Report **Probe budget** (configured limit) and **Guarded command time** (sum of measured
child spans). Measure guard start to child launch as pre-launch elapsed time, and child
start to finish as command duration, not a component's latency without its own measurement.
A deadline window is not total run time. Gaps between receipts do not measure
status/Write overhead or prove how many probes fit; if late, say only that this run
dispatched its follow-up after the deadline.
Use **Total session elapsed**: `unmeasured` for the invocation whose report is being written.
Its final report Write, acknowledgement and cleanup are not finished yet. Do not fetch
a clock merely to fill that field. An optional **Measured interval** must cite its actual
start/end receipts and name the work outside those boundaries, including later report
Writes and cleanup; it is never a completed-session measurement.
## 3. Assemble and check every repetition before writing
Use the caller's selected surface templates and assembly rules. Build headlines,
Top 3, summaries and completion text from each finding's Observed and Confirmation
fields, not its Hypothesis. Choose one conservative factual sentence per finding and
reuse it verbatim in those locations; do not introduce a new causal paraphrase.
A disclaimer in the detail does not qualify a stronger claim elsewhere.
Proposed regression assertions must detect the original observation on its actual
channel. For a logged console error, capture console errors; exception-only hooks
do not detect a console-only message. Additional causal tests remain separate proposals.
Apply these evidence limits to proposed tests and learnings too.
Before the final Write, check every mention of each finding against its evidence
fields, every proposed test against the observed channel, and each timing claim against
its named boundaries. Remove unsupported claims from all sections, not only the detail.
Keep refused/unstarted probes and untested categories explicit. Write the report only
after this consistency check; do not repair an evidence gap with invented facts.
Check claims about frequency and executed checks against the actual commands/results:
one observation proves neither recurrence nor an unexecuted check. Apply the same
evidence limits to the final response and any caller-authorized learning note.
## 4. Capture permitted notes, then write the report
Run the learning step below only if its destination is caller-authorized:
the user or invoking workflow explicitly permitted that learning-store path.
Invoking /qa-only alone does not grant this permission. Otherwise
keep notes in `REPORT_FILE`; do not write learning stores or automatic memory.
**No explicit permission:** skip learning-store writes and continue to the report.
**Explicit permission:** Read the named store first. Preserve its existing contents
and use the permitted write tool to append a verified note in that store's format.
If the format or write interface is unavailable, keep the note in the report instead.
Do not run logging helpers: they may write caches or enqueue synchronization outside
the permitted path. This branch never changes configuration or enables synchronization.
Write the checked report to the entrypoint's permitted destinations.
After the final Write, respond briefly with its path and verified coverage/limits;
do not append new findings or timing explanations.
+76
View File
@@ -0,0 +1,76 @@
# Finalize a report from retained evidence
Complete these steps before the final report Write. They use retained results, not
new probes. Missing evidence stays unknown; an expired clock stays expired.
The caller's write boundary includes reports, learning notes and automatic memory.
Keep all of them in authorized destinations; a memory feature grants no extra path.
## 1. Establish each finding once
Ground the report and learnings in retained observations. For every finding, distinguish
the observed result, the expected contract and any untested causal hypothesis. Link the
supporting command/result or screenshot; unknown impact remains unknown. A console error
message does not establish an uncaught exception, failed payload or missing UI. Missing
text in a page-text extract does not establish an absent attribute or inaccessible element.
Verify those claims with an appropriate probe, or leave them unconfirmed when time expires.
Give each finding one ID and write its Observed, Expected, Evidence and Confirmation
fields first. A logged exception-shaped string proves a logged message, not that the
named operation executed. Keep possible causes in a separate Hypothesis field; omit
speculation that does not help the next investigation. Observed-once is not replay-confirmed.
## 2. Fill timing fields from their actual boundaries
Report **Probe budget** (configured limit) and **Guarded command time** (sum of measured
child spans). Measure guard start to child launch as pre-launch elapsed time, and child
start to finish as command duration, not a component's latency without its own measurement.
A deadline window is not total run time. Gaps between receipts do not measure
status/Write overhead or prove how many probes fit; if late, say only that this run
dispatched its follow-up after the deadline.
Use **Total session elapsed**: `unmeasured` for the invocation whose report is being written.
Its final report Write, acknowledgement and cleanup are not finished yet. Do not fetch
a clock merely to fill that field. An optional **Measured interval** must cite its actual
start/end receipts and name the work outside those boundaries, including later report
Writes and cleanup; it is never a completed-session measurement.
## 3. Assemble and check every repetition before writing
Use the caller's selected surface templates and assembly rules. Build headlines,
Top 3, summaries and completion text from each finding's Observed and Confirmation
fields, not its Hypothesis. Choose one conservative factual sentence per finding and
reuse it verbatim in those locations; do not introduce a new causal paraphrase.
A disclaimer in the detail does not qualify a stronger claim elsewhere.
Proposed regression assertions must detect the original observation on its actual
channel. For a logged console error, capture console errors; exception-only hooks
do not detect a console-only message. Additional causal tests remain separate proposals.
Apply these evidence limits to proposed tests and learnings too.
Before the final Write, check every mention of each finding against its evidence
fields, every proposed test against the observed channel, and each timing claim against
its named boundaries. Remove unsupported claims from all sections, not only the detail.
Keep refused/unstarted probes and untested categories explicit. Write the report only
after this consistency check; do not repair an evidence gap with invented facts.
Check claims about frequency and executed checks against the actual commands/results:
one observation proves neither recurrence nor an unexecuted check. Apply the same
evidence limits to the final response and any caller-authorized learning note.
## 4. Capture permitted notes, then write the report
Run the learning step below only if its destination is caller-authorized:
the user or invoking workflow explicitly permitted that learning-store path.
Invoking /qa-only alone does not grant this permission. Otherwise
keep notes in `REPORT_FILE`; do not write learning stores or automatic memory.
**No explicit permission:** skip learning-store writes and continue to the report.
**Explicit permission:** Read the named store first. Preserve its existing contents
and use the permitted write tool to append a verified note in that store's format.
If the format or write interface is unavailable, keep the note in the report instead.
Do not run logging helpers: they may write caches or enqueue synchronization outside
the permitted path. This branch never changes configuration or enables synchronization.
Write the checked report to the entrypoint's permitted destinations.
After the final Write, respond briefly with its path and verified coverage/limits;
do not append new findings or timing explanations.