Files
gstack/qa/SKILL.md.tmpl
T
Garry Tan dcaea52800 v1.91.7.0 feat: add functional QA and pre-publication docs checks (#2983)
* feat: add surface-aware exploratory QA and ship documentation gates

* test: preserve delegated QA setup authority after main integration

* fix(qa): clarify exploration order and preserve report artifacts

* test(qa): follow the shared setup reference directly

* refactor(ship): make verification and recovery routes explicit

* test(ship): align evidence and review guards with explicit routes

* fix(workflows): clarify ship recovery and functional QA evidence

* fix(workflows): clarify approval recovery and full QA coverage

* refactor(workflows): order review transactions and clarify ship state

* fix(ship): clarify final verification and fail closed at publication

* fix(evals): attribute native atomic documentation writes

* fix(ship): clarify recovery and documentation lifecycle guidance

* fix(test): preserve observed native placeholder styling in CI

* fix(codex): report watchdog timeouts without a process-exit race

* Checkpoint functional QA implementation and workflow validation repairs

* Fix documentation and shared-review fixture contracts

* docs: clarify judge reuse and evaluation supervision

* test: align review evidence and selected case contracts

* test: verify append-only documentation checkpoints and recovery

* fix: qualify QA workflows and CI validation repairs

* fix: launch shared-libs fixture scripts on Windows

* fix: qualify QA deadlines, fixture isolation, and shard cleanup

* fix: preserve qualified QA and cancellation repairs

* fix: enforce functional fixture authority and share strict event decoding

* fix: retain free-test evidence and explain recovery

* fix: reject malformed native evidence after decoder consolidation

* test: use reliable capture for telemetry privacy filters

* test: refresh measured quick coverage and document validation costs

* Fix native fixture receipts and preserve VM validation evidence

* Align negative judge controls with upstream clarity policy

* Fix report-only QA preparation and public evidence handling

* Clarify QA-only preparation and current-report preservation

* Stream Ship quality judgments with an explicit 64k response contract

* Validate compact judge reasoning locally with supported wire schema

* Align functional QA fixture instructions with evidence acceptance

* Bind native browser diagnostics to execution evidence and align review verdicts

* Preserve native diagnostic line boundaries

* Serialize functional QA evidence from native captures

* Keep large QA evidence fixture payload out of Windows argv
2026-09-29 06:07:35 -07:00

310 lines
11 KiB
Cheetah

---
name: qa
preamble-tier: 4
version: 2.0.0
description: |
Fix browser/API/CLI/job/worker/webhook bugs.
Commit verified fixes atomically. Use when asked to "qa", "QA", "test this site", "find bugs",
"test and fix", or "fix what's broken".
Proactively suggest when the user says a feature is ready for testing
or asks "does this work?". Three tiers: Quick (critical/high only),
Standard (+ medium), Exhaustive (+ cosmetic). Produces contract outcomes or browser health scores,
fix evidence, and a ship-readiness summary. For report-only mode, use /qa-only. (gstack)
voice-triggers:
- "quality check"
- "test the app"
- "run QA"
allowed-tools:
- Bash
- Read
- Write
- Edit
- Glob
- Grep
- AskUserQuestion
- WebSearch
triggers:
- qa test this
- find bugs on site
- test the site
---
{{PREAMBLE}}
{{BASE_BRANCH_DETECT}}
{{GBRAIN_CONTEXT_LOAD}}
# /qa: Test → Fix → Verify
---
{{SECTION_INDEX:qa}}
---
## Setup
{{SECTION:scope}}
**Parse the user's request for these parameters:**
| Parameter | Default | Override example |
|-----------|---------|-----------------:|
| Target | (infer from request/repository or ask) | Browser URL, API route, CLI command, job, worker or webhook |
| Tier | Standard | `--quick`, `--exhaustive` |
| Mode | full | `--quick`, `--regression <previous-report-or-baseline>` |
| Output dir | `.gstack/qa-reports/` | `Output to /tmp/qa` |
| Scope | Selected target (or diff-scoped) | `Focus on duplicate webhook delivery` |
| Auth | Isolated synthetic identity for functional probes | Browser session handling lives in browser setup; never request credentials in chat |
**Tiers determine which issues get fixed:**
- **Quick:** Fix critical + high severity only
- **Standard:** + medium severity (default)
- **Exhaustive:** + low/cosmetic severity
`--quick` also selects Quick exploration; `--exhaustive` changes only the fix tier.
Regression mode preserves the selected fix tier.
If both `--quick` and `--regression` are supplied, ask which exploration mode to use
before setup or probes. Keep the selected fix tier; this choice concerns exploration only.
**On a feature branch without an explicit scope:** Use diff-aware testing of changed
and adjacent behavior. Select the surface first; absence of a URL never forces a browser.
**Check for clean working tree:**
```bash
git status --porcelain
```
If dirty, **STOP** and use AskUserQuestion. Explain that a clean tree keeps QA fixes atomic:
- A) Commit all current changes with a descriptive message before QA (recommended).
- B) Stash changes, run QA, then pop the stash.
- C) Abort for manual cleanup.
Execute only the user's choice before continuing setup.
**Prepare report artifacts before browser setup.** Resolve any supplied prior report
and baseline paths before writing. Select the output override or `.gstack/qa-reports`.
Create that directory if absent. Use the directory as `REPORT_DIR`
only when it is empty; otherwise choose a fresh owned run subdirectory.
Use `run-YYYYMMDDTHHMMSSZ` in UTC, adding a suffix on collision. Keep all local evidence there.
Never overwrite previous reports, baselines, screenshots or exploration notes.
A caller's fixed artifact paths and permissions take precedence; if preserving them
safely is impossible, report the output blocker rather than expanding write authority.
**Browser surface only:** load its setup; functional-only runs skip this section.
{{SECTION:browser-setup}}
**Browser surface only:** check the test framework and use the existing bootstrap
offer if needed. Functional targets use supported native tests or report the gap;
they do not load this browser bootstrap or generate CI.
{{SECTION:test-bootstrap}}
---
{{LEARNINGS_SEARCH:query=qa testing bug regression flake fixture}}
## Test Plan Context
Prefer the richer of recent project test plans and plans in conversation over git diff:
1. **Project-scoped test plans:** Find the latest for this repo:
```bash
setopt +o nomatch 2>/dev/null || true # zsh compat
{{SLUG_EVAL}}
ls -t ~/.gstack/projects/$SLUG/*-test-plan-*.md 2>/dev/null | head -1
```
2. **Conversation context:** Prior `/plan-eng-review` or `/plan-ceo-review` test plans.
3. Fall back to git diff only if neither exists.
---
## Phases 1-6: QA Baseline
Follow the shared section's ordered preparation, then run its probe loop.
The numbered browser phases label techniques, not another workflow.
{{SECTION:exploratory}}
Report baseline findings before fixing. Keep browser scores and functional outcomes separate.
---
## Output Structure
Under `$REPORT_DIR`, write `qa-report-{target}-{YYYY-MM-DD}.md` and the browser's
`baseline.json`. Browser `{target}` is a safe hostname.
Browser evidence goes in `screenshots/`: `initial.jpg`,
`issue-NNN-step-N.jpg`, `issue-NNN-result.jpg`, annotated `issue-NNN.png` and
`issue-NNN-after.jpg` (Phase 5 is the before). Functional reports use a safe command/service
label and sanitized command/request/state evidence.
---
## Phase 7: Triage
Sort issues by severity and apply the selected fix tier. Mark lower-tier issues and
those not fixable from source (third-party widgets, infrastructure) as "deferred."
### Refresh learnings for the component/page where the bug lives
Before the fix loop, search again for the buggy component/page. Use ONE noun containing
only letters, digits or hyphens (e.g., `checkout-button`, `payment`), never a path,
quotes, whitespace or other punctuation; simplify to an alphanumeric stem if needed.
```bash
~/.claude/skills/gstack/bin/gstack-learnings-search --query "<your-keyword>" --limit 5 2>/dev/null || true
```
Name an applicable learning in one sentence, or continue if none applies.
---
## Phase 8: Fix Loop
For each fixable issue, in severity order:
### 8a. Diagnose and reproduce
Use the shared loop's causal hypothesis and minimized replay, recording actual versus
documented behavior before edits. Modify only responsible files. Environment failures
and unclear contracts never authorize repair.
### 8a.5. Regression test before repair
Match 2-3 nearby tests' naming, imports, assertions and fixtures. Reproduce the failure
in a new native test. Run its detected command before repair; prove the defect caused its
failure, not a bad fixture, import or service. Attribute it in the language's comment syntax:
```text
// Regression: ISSUE-NNN — short defect description
// Found by /qa on YYYY-MM-DD
// Report: .gstack/qa-reports/qa-report-{target}-{date}.md
```
A clear, healthy uncovered contract may gain a passing test without product edits.
Apply the shared exploratory section's native unit/integration/E2E rules.
CSS-only defects may use browser evidence. Missing infrastructure stays coverage debt.
Use the component's name and native extension in auto-incrementing `{name}.regression-N.test.{ext}`.
Set N to max number + 1, starting at 1; never replace an existing file.
Keep valid red regressions; narrowly correct a proved
fixture/test error or report the unresolved bug.
### 8b. Fix
Read the surrounding source and make the **minimal fix**. No unrelated refactors or features.
### 8c. Re-test
Re-run the regression, original failing probe and adjacent happy path. Inspect each
final state; acceptance alone cannot verify a worker repair. Failed/unavailable rechecks stay unresolved.
For browser defects only:
{{SECTION:browser-verify}}
### 8d. Commit verified work
```bash
git add <only-verified-source-and-regression-files>
git commit -m "fix(qa): ISSUE-NNN — short description"
```
Commit each verified fix with its regression, never unrelated fixes. Leave unresolved
repairs and valid red regressions/evidence uncommitted; tell the user what remains.
### 8e. Classify
- **verified**: passed 8c (native regression when available); disclose missing test coverage
- **best-effort**: fix applied but couldn't fully verify (e.g., needs auth state, external service)
- **reverted**: regression detected → undo only this run's repair (revert its commit if already committed), retain the valid regression/evidence, and mark the issue "deferred". Never discard user changes.
### 8e.5. Regression Test record
Record the test created before repair in 8a.5 and its re-test result from 8c:
file, command, attribution, tested boundary and red/green evidence, or why it is deferred.
This step records results; it does not create another test.
Healthy-contract commits use `test(qa): regression test for {contract}`.
**WTF-likelihood exclusion:** test-only commits do not count toward the heuristic.
### 8f. Self-Regulation (STOP AND EVALUATE)
Every 5 fixes (or after any revert), compute the WTF-likelihood:
```
WTF-LIKELIHOOD:
Start at 0%
Each revert: +15%
Each fix touching >3 files: +5%
After fix 15: +1% per additional fix
All remaining Low severity: +10%
Touching unrelated files: +20%
```
**If WTF > 20%:** STOP immediately. Show the user what you've done so far. Ask whether to continue.
**Hard cap: 50 fixes.** After 50 fixes, stop regardless of remaining issues.
---
## Phase 9: Final QA
Re-run affected contracts and adjacent happy paths on the final inputs.
Caller-required rechecks cannot be skipped as unaffected. For browser
surfaces, recheck affected pages and compute the final health score. Warn prominently
about a worse score or regressed contract; blocked/inconclusive rechecks never verify repairs.
---
## Phase 10: Report
Write the Output Structure report locally and copy the same content to project context:
**Project-scoped:** Write test outcome artifact for cross-session context:
```bash
{{SLUG_SETUP}}
```
Write to `~/.gstack/projects/{slug}/{user}-{branch}-test-outcome-{datetime}.md`
**Per-issue additions:**
- Fix Status: verified / best-effort / reverted / deferred
- Commit SHA (if fixed)
- Files Changed (if fixed)
- Before/After evidence: screenshots for browser, outputs/requests/durable state for functional
**Summary:** total issues, verified/best-effort/reverted fixes and deferred issues.
For browser coverage include the score delta. For functional coverage include
passing/failing/blocked/not-run contracts, permanent regressions and remaining risks,
never a score. Keep mixed results separate.
**PR Summary:** Include one line:
> "QA found N issues, fixed M, health score X → Y."
For functional targets, use those contract outcomes instead of a score in the PR summary.
---
## Phase 11: TODOS.md Update
If the repo has a `TODOS.md`:
1. **New deferred bugs** → add as TODOs with severity, category, and repro steps
2. **Fixed bugs that were in TODOS.md** → annotate with "Fixed by /qa on {branch}, {date}"
---
{{LEARNINGS_LOG}}
{{GBRAIN_SAVE_RESULTS}}
## Additional Rules (qa-specific)
**Outside an explicitly approved browser bootstrap:** Only create tests through authorized codification in Phase 8a.5. Never modify CI configuration or weaken existing tests; use new native test files.
When in doubt, stop and ask.