mirror of
https://github.com/garrytan/gstack.git
synced 2026-10-02 17:40:02 +02:00
v1.91.7.0 feat: add functional QA and pre-publication docs checks (#2983)
* feat: add surface-aware exploratory QA and ship documentation gates * test: preserve delegated QA setup authority after main integration * fix(qa): clarify exploration order and preserve report artifacts * test(qa): follow the shared setup reference directly * refactor(ship): make verification and recovery routes explicit * test(ship): align evidence and review guards with explicit routes * fix(workflows): clarify ship recovery and functional QA evidence * fix(workflows): clarify approval recovery and full QA coverage * refactor(workflows): order review transactions and clarify ship state * fix(ship): clarify final verification and fail closed at publication * fix(evals): attribute native atomic documentation writes * fix(ship): clarify recovery and documentation lifecycle guidance * fix(test): preserve observed native placeholder styling in CI * fix(codex): report watchdog timeouts without a process-exit race * Checkpoint functional QA implementation and workflow validation repairs * Fix documentation and shared-review fixture contracts * docs: clarify judge reuse and evaluation supervision * test: align review evidence and selected case contracts * test: verify append-only documentation checkpoints and recovery * fix: qualify QA workflows and CI validation repairs * fix: launch shared-libs fixture scripts on Windows * fix: qualify QA deadlines, fixture isolation, and shard cleanup * fix: preserve qualified QA and cancellation repairs * fix: enforce functional fixture authority and share strict event decoding * fix: retain free-test evidence and explain recovery * fix: reject malformed native evidence after decoder consolidation * test: use reliable capture for telemetry privacy filters * test: refresh measured quick coverage and document validation costs * Fix native fixture receipts and preserve VM validation evidence * Align negative judge controls with upstream clarity policy * Fix report-only QA preparation and public evidence handling * Clarify QA-only preparation and current-report preservation * Stream Ship quality judgments with an explicit 64k response contract * Validate compact judge reasoning locally with supported wire schema * Align functional QA fixture instructions with evidence acceptance * Bind native browser diagnostics to execution evidence and align review verdicts * Preserve native diagnostic line boundaries * Serialize functional QA evidence from native captures * Keep large QA evidence fixture payload out of Windows argv
This commit is contained in:
1 parent
65bfb0ce49
commit
dcaea52800
333 files changed
+41755
-7357
No files matched your search
@@ -0,0 +1,104 @@
|
||||
<!-- AUTO-GENERATED from exploratory.md.tmpl — do not edit directly -->
|
||||
<!-- Regenerate: bun run gen:skill-docs -->
|
||||
# Shared exploratory QA
|
||||
|
||||
The **caller** (/qa, /qa-only, /review or /ship) owns decisions, tests, fixes and publication. Discovery writes only reports/evidence
|
||||
and owned fixture state; no workflows, framework installs or publication.
|
||||
|
||||
## 0. Preparation gate
|
||||
|
||||
Complete these Reads in order before writing charters or probing:
|
||||
1. Read `sections/scope.md` relative to the installed `qa`/`gstack-qa` SKILL.md directory in full and select the surfaces.
|
||||
2. Read the selected surface methods below in full.
|
||||
|
||||
Use this host's installed `qa`/`gstack-qa` SKILL.md directory for these reads:
|
||||
|
||||
**Functional surfaces:**
|
||||
Read `sections/system-functional.md` in full.
|
||||
|
||||
**Browser surfaces only:**
|
||||
Read `sections/qa-patterns.md` in full.
|
||||
|
||||
Await each successful Read result before continuing. A supplied target, isolation
|
||||
description, section index or remembered method is not a completed instruction Read.
|
||||
Do not repeat a Read already completed in this invocation; reuse only its acknowledged
|
||||
full contents. If either required Read is missing, complete it now before Charter and preflight.
|
||||
Missing or unreadable assets, prerequisites or permission block affected probes, not independent safe checks. Report QA setup blockers.
|
||||
|
||||
## 1. Charter and preflight
|
||||
|
||||
Reuse resolved REPORT_DIR; otherwise own a fresh `.gstack/qa-reports` subdirectory.
|
||||
Write a **charter** per behavior: contract, risk, entrypoint, isolation, exit condition, source, commands and inputs. Save charters as Markdown in the report.
|
||||
|
||||
|
||||
For /qa and /qa-only:
|
||||
- Browser Quick: SECONDS=30. Browser Full/Regression: SECONDS=900.
|
||||
- Functional Full, Quick and Regression have no default total timer.
|
||||
Set SECONDS to the shorter mode/caller limit; an unlimited mode uses the caller's bound.
|
||||
Without a total time limit, do not start D; announce finite command timeouts.
|
||||
Stop when scoped contracts are tested or blocked.
|
||||
Clocks/checkpoints use REPORT_DIR; mixed standalone runs use REPORT_DIR/browser and REPORT_DIR/functional, with one final report at REPORT_DIR. Caller paths win.
|
||||
R = owned probe directory; D = R/deadline.json. Quote paths.
|
||||
G = `$HOME/.claude/skills/gstack/bin/gstack-qa-deadline`; Q = `$HOME/.claude/skills/gstack/bin/gstack-qa-evidence`.
|
||||
Start once before baseline: `bun G start D SECONDS [EARLIER_UTC]` if bounded.
|
||||
EARLIER_UTC = caller's absolute deadline, if set.
|
||||
Functional: `bun Q capture R NNN [--public] --deadline D -- COMMAND ARGS`.
|
||||
Unbounded: use `--timeout-ms MS` instead. Use fresh three-digit IDs.
|
||||
--public requires approved public/synthetic output; Q screens credentials. For complete private captures, await a safe Read of `R/.qa-evidence/NNN/observation.json`. Sensitive/incomplete captures cannot anchor checkpoints.
|
||||
Bounded browsers: `bun G run D -- COMMAND ARGS`. No detached probes.
|
||||
Never reset D/bypass G. Expiry or invalid/missing D stops probes; report unfinished coverage. QA_DEADLINE receipts are not observations.
|
||||
|
||||
## 2. Probe loop
|
||||
|
||||
Each probe is one native command/interaction plus checks, excluding bookkeeping.
|
||||
Never batch probes.
|
||||
|
||||
1. First demonstrate success: output AND durable effects. Guard if bounded; await completion.
|
||||
2. **Decide whether another probe is needed.** If bounded, run `bun G status D`.
|
||||
If expired or no safe next probe remains, STOP exploration; write the report, not a checkpoint.
|
||||
**Classify the last result before copying it.** For public or synthetic observations,
|
||||
retain the entire result unchanged, including owned fixture paths, IDs, hashes and
|
||||
existing credential placeholders. An absolute state path is not itself a secret.
|
||||
For actual secrets/private payloads, withhold those values and disclose the redaction
|
||||
and replay limits in the report. If no safe exact observation can be retained,
|
||||
stop the affected probe chain; never invent a substitute path, identity or state.
|
||||
**Publish before probing.** Create `exploration-NNN.json` in the probe directory, beside its deadline if bounded, with exactly four top-level fields:
|
||||
observationCommand: last completed probe's full outer command, including guard.
|
||||
observed: its exact decoded child JSON (no wrapper/extra keys), or its full non-JSON text.
|
||||
For guarded text, copy the complete span between the guard's started and finished receipt lines.
|
||||
Keep its whitespace and content fences verbatim. Do not summarize, relabel or add timing text.
|
||||
The guard adds one newline before its finished receipt; that separator is not child text.
|
||||
For unguarded text, copy the complete result instead.
|
||||
If capture is incomplete, report that limit instead of reconstructing it.
|
||||
hypothesis: why nextCommand. nextCommand: exact command/request, guarded if bounded.
|
||||
Preserve every safe program-JSON key/value and identity hash unchanged.
|
||||
Withhold unsafe values, disclose limits and stop that chain.
|
||||
Check fields before publication. No drafts/placeholders or invented safe-path redactions; corrections cannot repair published notes.
|
||||
Functional: `bun Q checkpoint R NNN CAPTURE_ID 'observationCommand' 'hypothesis' 'nextCommand'` with literal arguments. Q supplies observed; never transcribe it.
|
||||
Browser checkpoints use Write.
|
||||
Wait for successful checkpoint publication before dispatch.
|
||||
Never backfill or overwrite notes.
|
||||
3. Run that exact probe; G enforces the deadline when bounded.
|
||||
Report refusals as not-run; retain initial state/inputs/results. Repeat from step 2.
|
||||
4. Replay the exact failing command/request from the same initial fixture state via steps 2–3 (same native command, fresh capture ID)
|
||||
to confirm it, then minimize via those gates. Expiry leaves confirmation/minimization incomplete.
|
||||
Another input or a regression test is not that replay.
|
||||
5. If the user or another process changes source, commands or fixtures, review the affected
|
||||
contracts and return to step 2 for each affected revalidation. Do not make product changes yourself.
|
||||
Keep the original limits/notes; update outcomes only from fresh evidence.
|
||||
|
||||
## 3. Parent handoff
|
||||
|
||||
Never change product code, tests, configuration, dependencies or Git through any tool,
|
||||
including shell, rename, deletion, commit, stash or edit-then-restore. Return test_stub proposals
|
||||
with their failing contract and expected assertion; never create tests or freeze buggy output.
|
||||
|
||||
## 4. Final report
|
||||
|
||||
Use the surface report template; link each checkpoint. Separate browser scores, functional outcomes and proposed/executed tests.
|
||||
For evidence.json, Write R/annotations.json: {revision, runtime, cwd, evidence: [{capture, command, contract, expected, classification}], learning: [checkpoint IDs], limits}.
|
||||
Run `bun Q materialize R annotations.json` before Markdown; Q fills observed/learning, not classifications. Retain all safe probes, including failures/replays; disclose withheld/incomplete evidence.
|
||||
Evidence is invocation-local.
|
||||
Missing prerequisites/expectations/observations, timeouts and refusal never pass.
|
||||
Pass requires all required current-input contracts to pass with no required remainder.
|
||||
Report blocked, inconclusive and not-run coverage without claiming success.
|
||||
@@ -0,0 +1 @@
|
||||
{{QA_EXPLORATORY}}
|
||||
@@ -0,0 +1,19 @@
|
||||
{
|
||||
"$schema": "https://gstack.dev/schemas/section-manifest.json",
|
||||
"skill": "qa-only",
|
||||
"version": 1,
|
||||
"sections": [
|
||||
{
|
||||
"id": "exploratory",
|
||||
"file": "exploratory.md",
|
||||
"title": "Report-only exploratory QA",
|
||||
"trigger": "running selected report-only baseline and exploratory probes without product or test writes"
|
||||
},
|
||||
{
|
||||
"id": "reporting",
|
||||
"file": "reporting.md",
|
||||
"title": "Evidence-grounded report finalization",
|
||||
"trigger": "finalizing the report after probing stops"
|
||||
}
|
||||
]
|
||||
}
|
||||
@@ -0,0 +1,78 @@
|
||||
<!-- AUTO-GENERATED from reporting.md.tmpl — do not edit directly -->
|
||||
<!-- Regenerate: bun run gen:skill-docs -->
|
||||
# Finalize a report from retained evidence
|
||||
|
||||
Complete these steps before the final report Write. They use retained results, not
|
||||
new probes. Missing evidence stays unknown; an expired clock stays expired.
|
||||
The caller's write boundary includes reports, learning notes and automatic memory.
|
||||
Keep all of them in authorized destinations; a memory feature grants no extra path.
|
||||
|
||||
## 1. Establish each finding once
|
||||
|
||||
Ground the report and learnings in retained observations. For every finding, distinguish
|
||||
the observed result, the expected contract and any untested causal hypothesis. Link the
|
||||
supporting command/result or screenshot; unknown impact remains unknown. A console error
|
||||
message does not establish an uncaught exception, failed payload or missing UI. Missing
|
||||
text in a page-text extract does not establish an absent attribute or inaccessible element.
|
||||
Verify those claims with an appropriate probe, or leave them unconfirmed when time expires.
|
||||
|
||||
Give each finding one ID and write its Observed, Expected, Evidence and Confirmation
|
||||
fields first. A logged exception-shaped string proves a logged message, not that the
|
||||
named operation executed. Keep possible causes in a separate Hypothesis field; omit
|
||||
speculation that does not help the next investigation. Observed-once is not replay-confirmed.
|
||||
|
||||
## 2. Fill timing fields from their actual boundaries
|
||||
|
||||
Report **Probe budget** (configured limit) and **Guarded command time** (sum of measured
|
||||
child spans). Measure guard start to child launch as pre-launch elapsed time, and child
|
||||
start to finish as command duration, not a component's latency without its own measurement.
|
||||
A deadline window is not total run time. Gaps between receipts do not measure
|
||||
status/Write overhead or prove how many probes fit; if late, say only that this run
|
||||
dispatched its follow-up after the deadline.
|
||||
|
||||
Use **Total session elapsed**: `unmeasured` for the invocation whose report is being written.
|
||||
Its final report Write, acknowledgement and cleanup are not finished yet. Do not fetch
|
||||
a clock merely to fill that field. An optional **Measured interval** must cite its actual
|
||||
start/end receipts and name the work outside those boundaries, including later report
|
||||
Writes and cleanup; it is never a completed-session measurement.
|
||||
|
||||
## 3. Assemble and check every repetition before writing
|
||||
|
||||
Use the caller's selected surface templates and assembly rules. Build headlines,
|
||||
Top 3, summaries and completion text from each finding's Observed and Confirmation
|
||||
fields, not its Hypothesis. Choose one conservative factual sentence per finding and
|
||||
reuse it verbatim in those locations; do not introduce a new causal paraphrase.
|
||||
A disclaimer in the detail does not qualify a stronger claim elsewhere.
|
||||
|
||||
Proposed regression assertions must detect the original observation on its actual
|
||||
channel. For a logged console error, capture console errors; exception-only hooks
|
||||
do not detect a console-only message. Additional causal tests remain separate proposals.
|
||||
Apply these evidence limits to proposed tests and learnings too.
|
||||
|
||||
Before the final Write, check every mention of each finding against its evidence
|
||||
fields, every proposed test against the observed channel, and each timing claim against
|
||||
its named boundaries. Remove unsupported claims from all sections, not only the detail.
|
||||
Keep refused/unstarted probes and untested categories explicit. Write the report only
|
||||
after this consistency check; do not repair an evidence gap with invented facts.
|
||||
Check claims about frequency and executed checks against the actual commands/results:
|
||||
one observation proves neither recurrence nor an unexecuted check. Apply the same
|
||||
evidence limits to the final response and any caller-authorized learning note.
|
||||
|
||||
## 4. Capture permitted notes, then write the report
|
||||
|
||||
Run the learning step below only if its destination is caller-authorized:
|
||||
the user or invoking workflow explicitly permitted that learning-store path.
|
||||
Invoking /qa-only alone does not grant this permission. Otherwise
|
||||
keep notes in `REPORT_FILE`; do not write learning stores or automatic memory.
|
||||
|
||||
**No explicit permission:** skip learning-store writes and continue to the report.
|
||||
|
||||
**Explicit permission:** Read the named store first. Preserve its existing contents
|
||||
and use the permitted write tool to append a verified note in that store's format.
|
||||
If the format or write interface is unavailable, keep the note in the report instead.
|
||||
Do not run logging helpers: they may write caches or enqueue synchronization outside
|
||||
the permitted path. This branch never changes configuration or enables synchronization.
|
||||
|
||||
Write the checked report to the entrypoint's permitted destinations.
|
||||
After the final Write, respond briefly with its path and verified coverage/limits;
|
||||
do not append new findings or timing explanations.
|
||||
@@ -0,0 +1,76 @@
|
||||
# Finalize a report from retained evidence
|
||||
|
||||
Complete these steps before the final report Write. They use retained results, not
|
||||
new probes. Missing evidence stays unknown; an expired clock stays expired.
|
||||
The caller's write boundary includes reports, learning notes and automatic memory.
|
||||
Keep all of them in authorized destinations; a memory feature grants no extra path.
|
||||
|
||||
## 1. Establish each finding once
|
||||
|
||||
Ground the report and learnings in retained observations. For every finding, distinguish
|
||||
the observed result, the expected contract and any untested causal hypothesis. Link the
|
||||
supporting command/result or screenshot; unknown impact remains unknown. A console error
|
||||
message does not establish an uncaught exception, failed payload or missing UI. Missing
|
||||
text in a page-text extract does not establish an absent attribute or inaccessible element.
|
||||
Verify those claims with an appropriate probe, or leave them unconfirmed when time expires.
|
||||
|
||||
Give each finding one ID and write its Observed, Expected, Evidence and Confirmation
|
||||
fields first. A logged exception-shaped string proves a logged message, not that the
|
||||
named operation executed. Keep possible causes in a separate Hypothesis field; omit
|
||||
speculation that does not help the next investigation. Observed-once is not replay-confirmed.
|
||||
|
||||
## 2. Fill timing fields from their actual boundaries
|
||||
|
||||
Report **Probe budget** (configured limit) and **Guarded command time** (sum of measured
|
||||
child spans). Measure guard start to child launch as pre-launch elapsed time, and child
|
||||
start to finish as command duration, not a component's latency without its own measurement.
|
||||
A deadline window is not total run time. Gaps between receipts do not measure
|
||||
status/Write overhead or prove how many probes fit; if late, say only that this run
|
||||
dispatched its follow-up after the deadline.
|
||||
|
||||
Use **Total session elapsed**: `unmeasured` for the invocation whose report is being written.
|
||||
Its final report Write, acknowledgement and cleanup are not finished yet. Do not fetch
|
||||
a clock merely to fill that field. An optional **Measured interval** must cite its actual
|
||||
start/end receipts and name the work outside those boundaries, including later report
|
||||
Writes and cleanup; it is never a completed-session measurement.
|
||||
|
||||
## 3. Assemble and check every repetition before writing
|
||||
|
||||
Use the caller's selected surface templates and assembly rules. Build headlines,
|
||||
Top 3, summaries and completion text from each finding's Observed and Confirmation
|
||||
fields, not its Hypothesis. Choose one conservative factual sentence per finding and
|
||||
reuse it verbatim in those locations; do not introduce a new causal paraphrase.
|
||||
A disclaimer in the detail does not qualify a stronger claim elsewhere.
|
||||
|
||||
Proposed regression assertions must detect the original observation on its actual
|
||||
channel. For a logged console error, capture console errors; exception-only hooks
|
||||
do not detect a console-only message. Additional causal tests remain separate proposals.
|
||||
Apply these evidence limits to proposed tests and learnings too.
|
||||
|
||||
Before the final Write, check every mention of each finding against its evidence
|
||||
fields, every proposed test against the observed channel, and each timing claim against
|
||||
its named boundaries. Remove unsupported claims from all sections, not only the detail.
|
||||
Keep refused/unstarted probes and untested categories explicit. Write the report only
|
||||
after this consistency check; do not repair an evidence gap with invented facts.
|
||||
Check claims about frequency and executed checks against the actual commands/results:
|
||||
one observation proves neither recurrence nor an unexecuted check. Apply the same
|
||||
evidence limits to the final response and any caller-authorized learning note.
|
||||
|
||||
## 4. Capture permitted notes, then write the report
|
||||
|
||||
Run the learning step below only if its destination is caller-authorized:
|
||||
the user or invoking workflow explicitly permitted that learning-store path.
|
||||
Invoking /qa-only alone does not grant this permission. Otherwise
|
||||
keep notes in `REPORT_FILE`; do not write learning stores or automatic memory.
|
||||
|
||||
**No explicit permission:** skip learning-store writes and continue to the report.
|
||||
|
||||
**Explicit permission:** Read the named store first. Preserve its existing contents
|
||||
and use the permitted write tool to append a verified note in that store's format.
|
||||
If the format or write interface is unavailable, keep the note in the report instead.
|
||||
Do not run logging helpers: they may write caches or enqueue synchronization outside
|
||||
the permitted path. This branch never changes configuration or enables synchronization.
|
||||
|
||||
Write the checked report to the entrypoint's permitted destinations.
|
||||
After the final Write, respond briefly with its path and verified coverage/limits;
|
||||
do not append new findings or timing explanations.
|
||||
Reference in new issue
Block a user