Files
gstack/qa/SKILL.md.tmpl
T

313 lines
11 KiB
Cheetah

---
name: qa
preamble-tier: 4
version: 2.0.0
description: |
Fix browser/API/CLI/job/worker/webhook bugs.
Commit verified fixes atomically. Use when asked to "qa", "QA", "test this site", "find bugs",
"test and fix", or "fix what's broken".
Proactively suggest when the user says a feature is ready for testing
or asks "does this work?". Three tiers: Quick (critical/high only),
Standard (+ medium), Exhaustive (+ cosmetic). Produces contract outcomes or browser health scores,
fix evidence, and a ship-readiness summary. For report-only mode, use /qa-only. (gstack)
voice-triggers:
- "quality check"
- "test the app"
- "run QA"
allowed-tools:
- Bash
- Read
- Write
- Edit
- Glob
- Grep
- AskUserQuestion
- WebSearch
triggers:
- qa test this
- find bugs on site
- test the site
---
{{PREAMBLE}}
{{BASE_BRANCH_DETECT}}
{{GBRAIN_CONTEXT_LOAD}}
# /qa: Test → Fix → Verify
---
{{SECTION_INDEX:qa}}
---
## Setup
{{SECTION:scope}}
**Parse the user's request for these parameters:**
| Parameter | Default | Override example |
|-----------|---------|-----------------:|
| Target | (infer from request/repository or ask) | Browser URL, API route, CLI command, job, worker or webhook |
| Tier | Standard | `--quick`, `--exhaustive` |
| Mode | full | `--quick`, `--regression <previous-report-or-baseline>` |
| Output dir | `.gstack/qa-reports/` | `Output to /tmp/qa` |
| Scope | Selected target (or diff-scoped) | `Focus on duplicate webhook delivery` |
| Auth | Isolated synthetic identity for functional probes | Browser session handling lives in browser setup; never request credentials in chat |
**Tiers determine which issues get fixed:**
- **Quick:** Fix critical + high severity only
- **Standard:** + medium severity (default)
- **Exhaustive:** + low/cosmetic severity
`--quick` also selects Quick exploration; `--exhaustive` changes only the fix tier.
Regression mode preserves the selected fix tier.
If both `--quick` and `--regression` are supplied, ask which exploration mode to use
before setup or probes. Keep the selected fix tier; this choice concerns exploration only.
**On a feature branch without an explicit scope:** Use diff-aware testing of changed
and adjacent behavior. Select the surface first; absence of a URL never forces a browser.
**Check for clean working tree:**
```bash
git status --porcelain
```
If dirty, **STOP** and use AskUserQuestion. Explain that a clean tree keeps QA fixes atomic:
- A) Commit all current changes with a descriptive message before QA (recommended).
- B) Stash changes, run QA, then pop the stash.
- C) Abort for manual cleanup.
Execute only the user's choice before continuing setup.
**Prepare report artifacts before browser setup.** Resolve any supplied prior report
and baseline paths before writing. Select the output override or `.gstack/qa-reports`.
Create that directory if absent. Use the directory as `REPORT_DIR`
only when it is empty; otherwise choose a fresh owned run subdirectory.
Use `run-YYYYMMDDTHHMMSSZ` in UTC, adding a suffix on collision. Keep all local evidence there.
Never overwrite previous reports, baselines, screenshots or exploration notes.
A caller's fixed artifact paths and permissions take precedence; if preserving them
safely is impossible, report the output blocker rather than expanding write authority.
**Browser surface only:** load its setup; functional-only runs skip this section.
{{SECTION:browser-setup}}
**Browser surface only:** check the test framework and use the existing bootstrap
offer if needed. Functional targets use supported native tests or report the gap;
they do not load this browser bootstrap or generate CI.
{{SECTION:test-bootstrap}}
---
{{LEARNINGS_SEARCH:query=qa testing bug regression flake fixture}}
## Test Plan Context
Prefer the richer of recent project test plans and plans in conversation over git diff:
1. **Project-scoped test plans:** Find the latest for this repo:
```bash
setopt +o nomatch 2>/dev/null || true # zsh compat
{{SLUG_EVAL}}
ls -t ~/.gstack/projects/$SLUG/*-test-plan-*.md 2>/dev/null | head -1
```
2. **Conversation context:** Prior `/plan-eng-review` or `/plan-ceo-review` test plans.
3. Fall back to git diff only if neither exists.
---
## Phases 1-6: QA Baseline
Follow the shared section's ordered preparation, then run its probe loop.
The numbered browser phases label techniques, not another workflow.
{{SECTION:exploratory}}
Report baseline findings before fixing. Keep browser scores and functional outcomes separate.
---
## Output Structure
Under `$REPORT_DIR`, write `qa-report-{target}-{YYYY-MM-DD}.md` and the browser's
`baseline.json`. Browser `{target}` is a safe hostname.
Browser evidence goes in `screenshots/`: `initial.jpg`,
`issue-NNN-step-N.jpg`, `issue-NNN-result.jpg`, annotated `issue-NNN.png` and
`issue-NNN-after.jpg` (Phase 5 is the before). Functional reports use a safe command/service
label and sanitized command/request/state evidence.
---
## Phase 7: Triage
Sort issues by severity and apply the selected fix tier. Mark lower-tier issues and
those not fixable from source (third-party widgets, infrastructure) as "deferred."
### Refresh learnings for the component/page where the bug lives
Before the fix loop, search again for the buggy component/page. Use ONE noun containing
only letters, digits or hyphens (e.g., `checkout-button`, `payment`), never a path,
quotes, whitespace or other punctuation; simplify to an alphanumeric stem if needed.
```bash
~/.claude/skills/gstack/bin/gstack-learnings-search --query "<your-keyword>" --limit 5 2>/dev/null || true
```
Name an applicable learning in one sentence, or continue if none applies.
---
## Phase 8: Fix Loop
For each fixable issue, in severity order:
### 8a. Diagnose and reproduce
Use the shared loop's causal hypothesis and minimized replay, recording actual versus
documented behavior before edits. Modify only responsible files. Environment failures
and unclear contracts never authorize repair.
### 8a.5. Regression test before repair
{{TEST_VALUE_BAR:qa}}
Extend an existing table or fixture when one covers the boundary; never add a production
seam for the test. Match 2-3 nearby tests' naming, imports, assertions and fixtures. Reproduce the failure
in a new native test. Run its detected command before repair; prove the defect caused its
failure, not a bad fixture, import or service. Attribute it in the language's comment syntax:
```text
// Regression: ISSUE-NNN — short defect description
// Found by /qa on YYYY-MM-DD
// Report: .gstack/qa-reports/qa-report-{target}-{date}.md
```
A clear, healthy uncovered contract may gain a passing test without product edits.
Apply the shared exploratory section's native unit/integration/E2E rules.
CSS-only defects may use browser evidence. Missing infrastructure stays coverage debt.
Use the component's name and native extension in auto-incrementing `{name}.regression-N.test.{ext}`.
Set N to max number + 1, starting at 1; never replace an existing file.
Keep valid red regressions; narrowly correct a proved
fixture/test error or report the unresolved bug.
### 8b. Fix
Read the surrounding source and make the **minimal fix**. No unrelated refactors or features.
### 8c. Re-test
Re-run the regression, original failing probe and adjacent happy path. Inspect each
final state; acceptance alone cannot verify a worker repair. Failed/unavailable rechecks stay unresolved.
For browser defects only:
{{SECTION:browser-verify}}
### 8d. Commit verified work
```bash
git add <only-verified-source-and-regression-files>
git commit -m "fix(qa): ISSUE-NNN — short description"
```
Commit each verified fix with its regression, never unrelated fixes. Leave unresolved
repairs and valid red regressions/evidence uncommitted; tell the user what remains.
### 8e. Classify
- **verified**: passed 8c (native regression when available); disclose missing test coverage
- **best-effort**: fix applied but couldn't fully verify (e.g., needs auth state, external service)
- **reverted**: regression detected → undo only this run's repair (revert its commit if already committed), retain the valid regression/evidence, and mark the issue "deferred". Never discard user changes.
### 8e.5. Regression Test record
Record the test created before repair in 8a.5 and its re-test result from 8c:
file, command, attribution, tested boundary, value card and red/green evidence, or why it is deferred.
This step records results; it does not create another test.
Healthy-contract commits use `test(qa): regression test for {contract}`.
**WTF-likelihood exclusion:** test-only commits do not count toward the heuristic.
### 8f. Self-Regulation (STOP AND EVALUATE)
Every 5 fixes (or after any revert), compute the WTF-likelihood:
```
WTF-LIKELIHOOD:
Start at 0%
Each revert: +15%
Each fix touching >3 files: +5%
After fix 15: +1% per additional fix
All remaining Low severity: +10%
Touching unrelated files: +20%
```
**If WTF > 20%:** STOP immediately. Show the user what you've done so far. Ask whether to continue.
**Hard cap: 50 fixes.** After 50 fixes, stop regardless of remaining issues.
---
## Phase 9: Final QA
Re-run affected contracts and adjacent happy paths on the final inputs.
Caller-required rechecks cannot be skipped as unaffected. For browser
surfaces, recheck affected pages and compute the final health score. Warn prominently
about a worse score or regressed contract; blocked/inconclusive rechecks never verify repairs.
---
## Phase 10: Report
Write the Output Structure report locally and copy the same content to project context:
**Project-scoped:** Write test outcome artifact for cross-session context:
```bash
{{SLUG_SETUP}}
```
Write to `~/.gstack/projects/{slug}/{user}-{branch}-test-outcome-{datetime}.md`
**Per-issue additions:**
- Fix Status: verified / best-effort / reverted / deferred
- Commit SHA (if fixed)
- Files Changed (if fixed)
- Before/After evidence: screenshots for browser, outputs/requests/durable state for functional
**Summary:** total issues, verified/best-effort/reverted fixes and deferred issues.
For browser coverage include the score delta. For functional coverage include
passing/failing/blocked/not-run contracts, permanent regressions and remaining risks,
never a score. Keep mixed results separate.
**PR Summary:** Include one line:
> "QA found N issues, fixed M, health score X → Y."
For functional targets, use those contract outcomes instead of a score in the PR summary.
---
## Phase 11: TODOS.md Update
If the repo has a `TODOS.md`:
1. **New deferred bugs** → add as TODOs with severity, category, and repro steps
2. **Fixed bugs that were in TODOS.md** → annotate with "Fixed by /qa on {branch}, {date}"
---
{{LEARNINGS_LOG}}
{{GBRAIN_SAVE_RESULTS}}
## Additional Rules (qa-specific)
**Outside an explicitly approved browser bootstrap:** Only create tests through authorized codification in Phase 8a.5. Never modify CI configuration or weaken existing tests; use new native test files.
When in doubt, stop and ask.