mirror of
https://github.com/garrytan/gstack.git
synced 2026-09-17 18:32:19 +02:00
feat(qa): carve QA patterns + health rubric into on-demand sections (68→48KB skeleton)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
This commit is contained in:
co-authored by
Claude Fable 5
parent
706091cbf9
commit
f1d21b65ca
+19
-462
@@ -459,6 +459,20 @@ branch name wherever the instructions say "the base branch" or `<default>`.
|
|||||||
|
|
||||||
You are a QA engineer AND a bug-fix engineer. Test web applications like a real user — click everything, fill every form, check every state. When you find bugs, fix them in source code with atomic commits, then re-verify. Produce a structured report with before/after evidence.
|
You are a QA engineer AND a bug-fix engineer. Test web applications like a real user — click everything, fill every form, check every state. When you find bugs, fix them in source code with atomic commits, then re-verify. Produce a structured report with before/after evidence.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Section index — Read each section when its situation applies
|
||||||
|
|
||||||
|
This skill is a decision-tree skeleton. The steps below point to on-demand
|
||||||
|
sections. Read a section in full before doing its step; do not work from memory.
|
||||||
|
|
||||||
|
| When | Read this section |
|
||||||
|
|------|-------------------|
|
||||||
|
| checking the project's test framework during Setup — ecosystem-marker detection, the bootstrap offer, framework install, CI pipeline generation, and first real tests (also needed at Phase 8e.5 if you skipped it and a regression test now requires a framework) | `sections/test-bootstrap.md` |
|
||||||
|
| running the QA baseline (Phases 1-6) — mode selection (Diff-aware/Full/Quick/Regression), the phase-by-phase browser workflow, the Health Score Rubric, framework-specific guidance, and the browser-testing Important Rules | `sections/qa-patterns.md` |
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
## Setup
|
## Setup
|
||||||
|
|
||||||
**Parse the user's request for these parameters:**
|
**Parse the user's request for these parameters:**
|
||||||
@@ -543,190 +557,8 @@ If `NEEDS_SETUP`:
|
|||||||
|
|
||||||
**Check test framework (bootstrap if needed):**
|
**Check test framework (bootstrap if needed):**
|
||||||
|
|
||||||
## Test Framework Bootstrap
|
> **STOP.** Before checking the project's test framework during Setup — ecosystem-marker detection, the bootstrap offer, framework install, CI pipeline generation, and first real tests (also needed at Phase 8e.5 if you skipped it and a regression test now requires a framework), Read `~/.claude/skills/gstack/qa/sections/test-bootstrap.md` and execute it
|
||||||
|
> in full. Do not work from memory — that section is the source of truth for this step.
|
||||||
**Read the project's CLAUDE.md (and TESTING.md if present) FIRST.** If it documents a test command, the project already told you: no detection, no bootstrap. Skip the rest of bootstrap and use that command in Step 5.
|
|
||||||
|
|
||||||
**Otherwise gather markers. Every marker below is EVIDENCE for the question you ask — never a command to run blind.** A marker tells you which ecosystem you're in and which command to OFFER. It does not tell you the command works. Do not execute a candidate test command to "check" it: a probe on a project that never had that runner fails loudly and teaches you nothing, and installing a second framework over a working one is worse.
|
|
||||||
|
|
||||||
```bash
|
|
||||||
setopt +o nomatch 2>/dev/null || true # zsh compat
|
|
||||||
# Definitive ecosystem markers (presence = ecosystem, NOT a command to run)
|
|
||||||
[ -f manage.py ] && echo "RUNTIME:python FRAMEWORK:django MARKER:manage.py"
|
|
||||||
{ [ -f pyproject.toml ] || [ -f pytest.ini ] || [ -f tox.ini ] || [ -f setup.cfg ] || [ -f requirements.txt ]; } && echo "RUNTIME:python"
|
|
||||||
[ -f Gemfile ] || [ -f Rakefile ] || [ -f .rspec ] && echo "RUNTIME:ruby"
|
|
||||||
[ -f package.json ] && echo "RUNTIME:node"
|
|
||||||
[ -f go.mod ] && echo "RUNTIME:go"
|
|
||||||
[ -f Cargo.toml ] && echo "RUNTIME:rust"
|
|
||||||
[ -f composer.json ] && echo "RUNTIME:php"
|
|
||||||
[ -f mix.exs ] && echo "RUNTIME:elixir"
|
|
||||||
[ -f pom.xml ] && echo "RUNTIME:jvm BUILD:maven"
|
|
||||||
{ [ -f build.gradle ] || [ -f build.gradle.kts ]; } && echo "RUNTIME:jvm BUILD:gradle"
|
|
||||||
# Detect sub-frameworks
|
|
||||||
[ -f Gemfile ] && grep -q "rails" Gemfile 2>/dev/null && echo "FRAMEWORK:rails"
|
|
||||||
[ -f package.json ] && grep -q '"next"' package.json 2>/dev/null && echo "FRAMEWORK:nextjs"
|
|
||||||
# Existing test path — config files, declared scripts, AND test FILES.
|
|
||||||
# A project with real tests and no config file is the common miss.
|
|
||||||
ls jest.config.* vitest.config.* playwright.config.* .rspec pytest.ini tox.ini phpunit.xml* 2>/dev/null
|
|
||||||
[ -f package.json ] && grep -q '"test"[[:space:]]*:' package.json && echo "SCRIPT:package.json test"
|
|
||||||
[ -f Makefile ] && grep -qE '^(test|check):' Makefile && echo "TARGET:make test"
|
|
||||||
[ -f pyproject.toml ] && grep -q "pytest" pyproject.toml && echo "CONFIG:pyproject pytest"
|
|
||||||
git ls-files | grep -cE '(^|/)(tests?|spec|__tests__)/|(^|/)tests?\.py$|(^|/)test_[^/]+\.py$|_test\.(go|py|rb|ts|js|exs)$|\.(test|spec)\.[jt]sx?$|_spec\.rb$|Test\.(java|kt)$' | sed 's/^/TESTFILES:/'
|
|
||||||
# Rust keeps unit tests inside src/, so file names alone miss them
|
|
||||||
[ -f Cargo.toml ] && git grep -lF '#[test]' -- 'src' >/dev/null 2>&1 && echo "TESTS:rust in-source"
|
|
||||||
# Check opt-out marker
|
|
||||||
[ -f .gstack/no-test-bootstrap ] && echo "BOOTSTRAP_DECLINED"
|
|
||||||
```
|
|
||||||
|
|
||||||
Map the markers to the command you will OFFER — never to one you run on a guess:
|
|
||||||
|
|
||||||
| Marker | Ecosystem | Candidate command to offer |
|
|
||||||
|--------|-----------|----------------------------|
|
|
||||||
| `manage.py` | Django | `python manage.py test` (or `pytest` when pytest-django is in the deps) |
|
|
||||||
| `pytest.ini` / `tox.ini` / pytest in `pyproject.toml` / `test_*.py` | Python | `pytest` |
|
|
||||||
| `go.mod` (+ any `*_test.go`) | Go | `go test ./...` |
|
|
||||||
| `Cargo.toml` | Rust | `cargo test` |
|
|
||||||
| `pom.xml` | JVM (Maven) | `mvn test` |
|
|
||||||
| `build.gradle` / `build.gradle.kts` | JVM (Gradle) | `./gradlew test` |
|
|
||||||
| `Gemfile` / `Rakefile` / `.rspec` | Ruby | `bundle exec rspec`, `bin/rails test`, or `rake test` |
|
|
||||||
| `mix.exs` | Elixir | `mix test` |
|
|
||||||
| `composer.json` | PHP | `composer test` or `./vendor/bin/phpunit` |
|
|
||||||
| `package.json` with a `test` script | Node | that script, run with the package manager the lockfile names |
|
|
||||||
| `Makefile` with a `test:` target | any | `make test` |
|
|
||||||
|
|
||||||
**If ANY existing-test evidence appears** (a config file, a declared test script or make target, a nonzero `TESTFILES:` count, or `TESTS:rust in-source`): the project has tests. **Do NOT bootstrap.** Print "Existing tests detected: {the evidence}." Then get the command the same way Step 5 does — CLAUDE.md/TESTING.md if documented, otherwise AskUserQuestion offering the candidates from the table above plus "Other", and persist the answer to CLAUDE.md's `## Testing` section so it is never asked again. When the ecosystem ships a runner (Django, Go, Rust, Elixir, Maven/Gradle), that runner is the candidate — never install a second framework beside a working one.
|
|
||||||
Read 2-3 existing test files to learn conventions (naming, imports, assertion style, setup patterns).
|
|
||||||
Store conventions as prose context for use in Phase 8e.5 or Step 7. **Skip the rest of bootstrap.**
|
|
||||||
|
|
||||||
Absent config files and absent `tests/` directories are NOT evidence of "no tests": Django keeps tests in `<app>/tests.py`, Go in `*_test.go` beside the source, Rust in `#[test]` blocks inside `src/`. A green `python manage.py test` with no `pytest.ini` is a tested project, not a bootstrap candidate.
|
|
||||||
|
|
||||||
**If BOOTSTRAP_DECLINED** appears: Print "Test bootstrap previously declined — skipping." **Skip the rest of bootstrap.**
|
|
||||||
|
|
||||||
**If NO ecosystem marker matched:** Use AskUserQuestion:
|
|
||||||
"I couldn't detect your project's language. What runtime are you using?"
|
|
||||||
Options: A) Node.js/TypeScript B) Ruby/Rails C) Python D) Go E) Rust F) PHP G) Elixir H) This project doesn't need tests.
|
|
||||||
If the runtime you need isn't listed, offer "Other" and take the runtime plus the test command as free text.
|
|
||||||
If user picks H → write `.gstack/no-test-bootstrap` and continue without tests.
|
|
||||||
|
|
||||||
**If an ecosystem matched but there is no existing-test evidence at all — bootstrap:**
|
|
||||||
|
|
||||||
### B2. Research best practices
|
|
||||||
|
|
||||||
Use WebSearch to find current best practices for the detected runtime:
|
|
||||||
- `"[runtime] best test framework 2025 2026"`
|
|
||||||
- `"[framework A] vs [framework B] comparison"`
|
|
||||||
|
|
||||||
If WebSearch is unavailable, use this built-in knowledge table:
|
|
||||||
|
|
||||||
| Runtime | Primary recommendation | Alternative |
|
|
||||||
|---------|----------------------|-------------|
|
|
||||||
| Ruby/Rails | minitest + fixtures + capybara | rspec + factory_bot + shoulda-matchers |
|
|
||||||
| Node.js | vitest + @testing-library | jest + @testing-library |
|
|
||||||
| Next.js | vitest + @testing-library/react + playwright | jest + cypress |
|
|
||||||
| Python | pytest + pytest-cov | unittest |
|
|
||||||
| Django | pytest + pytest-django | Django's built-in `manage.py test` (unittest) |
|
|
||||||
| Go | stdlib testing + testify | stdlib only |
|
|
||||||
| JVM (Maven/Gradle) | JUnit 5 + AssertJ | JUnit 5 only |
|
|
||||||
| Rust | cargo test (built-in) + mockall | — |
|
|
||||||
| PHP | phpunit + mockery | pest |
|
|
||||||
| Elixir | ExUnit (built-in) + ex_machina | — |
|
|
||||||
|
|
||||||
### B3. Framework selection
|
|
||||||
|
|
||||||
Use AskUserQuestion:
|
|
||||||
"I detected this is a [Runtime/Framework] project with no test framework. I researched current best practices. Here are the options:
|
|
||||||
A) [Primary] — [rationale]. Includes: [packages]. Supports: unit, integration, smoke, e2e
|
|
||||||
B) [Alternative] — [rationale]. Includes: [packages]
|
|
||||||
C) Skip — don't set up testing right now
|
|
||||||
RECOMMENDATION: Choose A because [reason based on project context]"
|
|
||||||
|
|
||||||
If user picks C → write `.gstack/no-test-bootstrap`. Tell user: "If you change your mind later, delete `.gstack/no-test-bootstrap` and re-run." Continue without tests.
|
|
||||||
|
|
||||||
If multiple runtimes detected (monorepo) → ask which runtime to set up first, with option to do both sequentially.
|
|
||||||
|
|
||||||
### B4. Install and configure
|
|
||||||
|
|
||||||
1. Install the chosen packages (npm/bun/gem/pip/etc.)
|
|
||||||
2. Create minimal config file
|
|
||||||
3. Create directory structure (test/, spec/, etc.)
|
|
||||||
4. Create one example test matching the project's code to verify setup works
|
|
||||||
|
|
||||||
If package installation fails → debug once. If still failing → revert with `git checkout -- package.json package-lock.json` (or equivalent for the runtime). Warn user and continue without tests.
|
|
||||||
|
|
||||||
### B4.5. First real tests
|
|
||||||
|
|
||||||
Generate 3-5 real tests for existing code:
|
|
||||||
|
|
||||||
1. **Find recently changed files:** `git log --since=30.days --name-only --format="" | sort | uniq -c | sort -rn | head -10`
|
|
||||||
2. **Prioritize by risk:** Error handlers > business logic with conditionals > API endpoints > pure functions
|
|
||||||
3. **For each file:** Write one test that tests real behavior with meaningful assertions. Never `expect(x).toBeDefined()` — test what the code DOES.
|
|
||||||
4. Run each test. Passes → keep. Fails → fix once. Still fails → delete silently.
|
|
||||||
5. Generate at least 1 test, cap at 5.
|
|
||||||
|
|
||||||
Never import secrets, API keys, or credentials in test files. Use environment variables or test fixtures.
|
|
||||||
|
|
||||||
### B5. Verify
|
|
||||||
|
|
||||||
```bash
|
|
||||||
# Run the full test suite to confirm everything works
|
|
||||||
{detected test command}
|
|
||||||
```
|
|
||||||
|
|
||||||
If tests fail → debug once. If still failing → revert all bootstrap changes and warn user.
|
|
||||||
|
|
||||||
### B5.5. CI/CD pipeline
|
|
||||||
|
|
||||||
```bash
|
|
||||||
# Check CI provider
|
|
||||||
ls -d .github/ 2>/dev/null && echo "CI:github"
|
|
||||||
ls .gitlab-ci.yml .circleci/ bitrise.yml 2>/dev/null
|
|
||||||
```
|
|
||||||
|
|
||||||
If `.github/` exists (or no CI detected — default to GitHub Actions):
|
|
||||||
Create `.github/workflows/test.yml` with:
|
|
||||||
- `runs-on: ubuntu-latest`
|
|
||||||
- Appropriate setup action for the runtime (setup-node, setup-ruby, setup-python, etc.)
|
|
||||||
- The same test command verified in B5
|
|
||||||
- Trigger: push + pull_request
|
|
||||||
|
|
||||||
If non-GitHub CI detected → skip CI generation with note: "Detected {provider} — CI pipeline generation supports GitHub Actions only. Add test step to your existing pipeline manually."
|
|
||||||
|
|
||||||
### B6. Create TESTING.md
|
|
||||||
|
|
||||||
First check: If TESTING.md already exists → read it and update/append rather than overwriting. Never destroy existing content.
|
|
||||||
|
|
||||||
Write TESTING.md with:
|
|
||||||
- Philosophy: "100% test coverage is the key to great vibe coding. Tests let you move fast, trust your instincts, and ship with confidence — without them, vibe coding is just yolo coding. With tests, it's a superpower."
|
|
||||||
- Framework name and version
|
|
||||||
- How to run tests (the verified command from B5)
|
|
||||||
- Test layers: Unit tests (what, where, when), Integration tests, Smoke tests, E2E tests
|
|
||||||
- Conventions: file naming, assertion style, setup/teardown patterns
|
|
||||||
|
|
||||||
### B7. Update CLAUDE.md
|
|
||||||
|
|
||||||
First check: If CLAUDE.md already has a `## Testing` section → skip. Don't duplicate.
|
|
||||||
|
|
||||||
Append a `## Testing` section:
|
|
||||||
- Run command and test directory
|
|
||||||
- Reference to TESTING.md
|
|
||||||
- Test expectations:
|
|
||||||
- 100% test coverage is the goal — tests make vibe coding safe
|
|
||||||
- When writing new functions, write a corresponding test
|
|
||||||
- When fixing a bug, write a regression test
|
|
||||||
- When adding error handling, write a test that triggers the error
|
|
||||||
- When adding a conditional (if/else, switch), write tests for BOTH paths
|
|
||||||
- Never commit code that makes existing tests fail
|
|
||||||
|
|
||||||
### B8. Commit
|
|
||||||
|
|
||||||
```bash
|
|
||||||
git status --porcelain
|
|
||||||
```
|
|
||||||
|
|
||||||
Only commit if there are changes. Stage all bootstrap files (config, test directory, TESTING.md, CLAUDE.md, .github/workflows/test.yml if created):
|
|
||||||
`git commit -m "chore: bootstrap test framework ({framework name})"`
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
**Create output directories:**
|
**Create output directories:**
|
||||||
|
|
||||||
@@ -791,285 +623,10 @@ Before falling back to git diff heuristics, check for richer test plan sources:
|
|||||||
|
|
||||||
## Phases 1-6: QA Baseline
|
## Phases 1-6: QA Baseline
|
||||||
|
|
||||||
## Modes
|
> **STOP.** Before running the QA baseline (Phases 1-6) — mode selection (Diff-aware/Full/Quick/Regression), the phase-by-phase browser workflow, the Health Score Rubric, framework-specific guidance, and the browser-testing Important Rules, Read `~/.claude/skills/gstack/qa/sections/qa-patterns.md` and execute it
|
||||||
|
> in full. Do not work from memory — that section is the source of truth for this step.
|
||||||
|
|
||||||
### Diff-aware (automatic when on a feature branch with no URL)
|
Record baseline health score at end of Phase 6 (per the Health Score Rubric in that section).
|
||||||
|
|
||||||
This is the **primary mode** for developers verifying their work. When the user says `/qa` without a URL and the repo is on a feature branch, automatically:
|
|
||||||
|
|
||||||
1. **Analyze the branch diff** to understand what changed:
|
|
||||||
```bash
|
|
||||||
git diff main...HEAD --name-only
|
|
||||||
git log main..HEAD --oneline
|
|
||||||
```
|
|
||||||
|
|
||||||
2. **Identify affected pages/routes** from the changed files:
|
|
||||||
- Controller/route files → which URL paths they serve
|
|
||||||
- View/template/component files → which pages render them
|
|
||||||
- Model/service files → which pages use those models (check controllers that reference them)
|
|
||||||
- CSS/style files → which pages include those stylesheets
|
|
||||||
- API endpoints → test them directly with `$B js "await fetch('/api/...')"`
|
|
||||||
- Static pages (markdown, HTML) → navigate to them directly
|
|
||||||
|
|
||||||
**If no obvious pages/routes are identified from the diff:** Do not skip browser testing. The user invoked /qa because they want browser-based verification. Fall back to Quick mode — navigate to the homepage, follow the top 5 navigation targets, check console for errors, and test any interactive elements found. Backend, config, and infrastructure changes affect app behavior — always verify the app still works.
|
|
||||||
|
|
||||||
3. **Detect the running app** — check common local dev ports:
|
|
||||||
```bash
|
|
||||||
$B goto http://localhost:3000 2>/dev/null && echo "Found app on :3000" || \
|
|
||||||
$B goto http://localhost:4000 2>/dev/null && echo "Found app on :4000" || \
|
|
||||||
$B goto http://localhost:8080 2>/dev/null && echo "Found app on :8080"
|
|
||||||
```
|
|
||||||
If no local app is found, check for a staging/preview URL in the PR or environment. If nothing works, ask the user for the URL.
|
|
||||||
|
|
||||||
4. **Test each affected page/route:**
|
|
||||||
- Navigate to the page
|
|
||||||
- Take a screenshot
|
|
||||||
- Check console for errors
|
|
||||||
- If the change was interactive (forms, buttons, flows), test the interaction end-to-end
|
|
||||||
- Use `snapshot -D` before and after actions to verify the change had the expected effect
|
|
||||||
|
|
||||||
5. **Cross-reference with commit messages and PR description** to understand *intent* — what should the change do? Verify it actually does that.
|
|
||||||
|
|
||||||
6. **Check TODOS.md** (if it exists) for known bugs or issues related to the changed files. If a TODO describes a bug that this branch should fix, add it to your test plan. If you find a new bug during QA that isn't in TODOS.md, note it in the report.
|
|
||||||
|
|
||||||
7. **Report findings** scoped to the branch changes:
|
|
||||||
- "Changes tested: N pages/routes affected by this branch"
|
|
||||||
- For each: does it work? Screenshot evidence.
|
|
||||||
- Any regressions on adjacent pages?
|
|
||||||
|
|
||||||
**If the user provides a URL with diff-aware mode:** Use that URL as the base but still scope testing to the changed files.
|
|
||||||
|
|
||||||
### Full (default when URL is provided)
|
|
||||||
Systematic exploration. Visit every reachable page. Document 5-10 well-evidenced issues. Produce health score. Takes 5-15 minutes depending on app size.
|
|
||||||
|
|
||||||
### Quick (`--quick`)
|
|
||||||
30-second smoke test. Visit homepage + top 5 navigation targets. Check: page loads? Console errors? Broken links? Produce health score. No detailed issue documentation.
|
|
||||||
|
|
||||||
### Regression (`--regression <baseline>`)
|
|
||||||
Run full mode, then load `baseline.json` from a previous run. Diff: which issues are fixed? Which are new? What's the score delta? Append regression section to report.
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
## Workflow
|
|
||||||
|
|
||||||
### Phase 1: Initialize
|
|
||||||
|
|
||||||
1. Find browse binary (see Setup above)
|
|
||||||
2. Create output directories
|
|
||||||
3. Copy report template from `qa/templates/qa-report-template.md` to output dir
|
|
||||||
4. Start timer for duration tracking
|
|
||||||
|
|
||||||
### Phase 2: Authenticate (if needed)
|
|
||||||
|
|
||||||
**If the user specified auth credentials:**
|
|
||||||
|
|
||||||
```bash
|
|
||||||
$B goto <login-url>
|
|
||||||
$B snapshot -i # find the login form
|
|
||||||
$B fill @e3 "user@example.com"
|
|
||||||
$B fill @e4 "[REDACTED]" # NEVER include real passwords in report
|
|
||||||
$B click @e5 # submit
|
|
||||||
$B snapshot -D # verify login succeeded
|
|
||||||
```
|
|
||||||
|
|
||||||
**If the user provided a cookie file:**
|
|
||||||
|
|
||||||
```bash
|
|
||||||
$B cookie-import cookies.json
|
|
||||||
$B goto <target-url>
|
|
||||||
```
|
|
||||||
|
|
||||||
**If 2FA/OTP is required:** Ask the user for the code and wait.
|
|
||||||
|
|
||||||
**If CAPTCHA blocks you:** Tell the user: "Please complete the CAPTCHA in the browser, then tell me to continue."
|
|
||||||
|
|
||||||
### Phase 3: Orient
|
|
||||||
|
|
||||||
Get a map of the application:
|
|
||||||
|
|
||||||
```bash
|
|
||||||
$B goto <target-url>
|
|
||||||
$B snapshot -i -a -o "$REPORT_DIR/screenshots/initial.png"
|
|
||||||
$B links # map navigation structure
|
|
||||||
$B console --errors # any errors on landing?
|
|
||||||
```
|
|
||||||
|
|
||||||
**Detect framework** (note in report metadata):
|
|
||||||
- `__next` in HTML or `_next/data` requests → Next.js
|
|
||||||
- `csrf-token` meta tag → Rails
|
|
||||||
- `wp-content` in URLs → WordPress
|
|
||||||
- Client-side routing with no page reloads → SPA
|
|
||||||
|
|
||||||
**For SPAs:** The `links` command may return few results because navigation is client-side. Use `snapshot -i` to find nav elements (buttons, menu items) instead.
|
|
||||||
|
|
||||||
### Phase 4: Explore
|
|
||||||
|
|
||||||
Visit pages systematically. At each page:
|
|
||||||
|
|
||||||
```bash
|
|
||||||
$B goto <page-url>
|
|
||||||
$B snapshot -i -a -o "$REPORT_DIR/screenshots/page-name.png"
|
|
||||||
$B console --errors
|
|
||||||
```
|
|
||||||
|
|
||||||
Then follow the **per-page exploration checklist** (see `qa/references/issue-taxonomy.md`):
|
|
||||||
|
|
||||||
1. **Visual scan** — Look at the annotated screenshot for layout issues
|
|
||||||
2. **Interactive elements** — Click buttons, links, controls. Do they work?
|
|
||||||
3. **Forms** — Fill and submit. Test empty, invalid, edge cases
|
|
||||||
4. **Navigation** — Check all paths in and out
|
|
||||||
5. **States** — Empty state, loading, error, overflow
|
|
||||||
6. **Console** — Any new JS errors after interactions?
|
|
||||||
7. **Responsiveness** — Check mobile viewport if relevant:
|
|
||||||
```bash
|
|
||||||
$B viewport 375x812
|
|
||||||
$B screenshot "$REPORT_DIR/screenshots/page-mobile.png"
|
|
||||||
$B viewport 1280x720
|
|
||||||
```
|
|
||||||
|
|
||||||
**Depth judgment:** Spend more time on core features (homepage, dashboard, checkout, search) and less on secondary pages (about, terms, privacy).
|
|
||||||
|
|
||||||
**Quick mode:** Only visit homepage + top 5 navigation targets from the Orient phase. Skip the per-page checklist — just check: loads? Console errors? Broken links visible?
|
|
||||||
|
|
||||||
### Phase 5: Document
|
|
||||||
|
|
||||||
Document each issue **immediately when found** — don't batch them.
|
|
||||||
|
|
||||||
**Two evidence tiers:**
|
|
||||||
|
|
||||||
**Interactive bugs** (broken flows, dead buttons, form failures):
|
|
||||||
1. Take a screenshot before the action
|
|
||||||
2. Perform the action
|
|
||||||
3. Take a screenshot showing the result
|
|
||||||
4. Use `snapshot -D` to show what changed
|
|
||||||
5. Write repro steps referencing screenshots
|
|
||||||
|
|
||||||
```bash
|
|
||||||
$B screenshot "$REPORT_DIR/screenshots/issue-001-step-1.png"
|
|
||||||
$B click @e5
|
|
||||||
$B screenshot "$REPORT_DIR/screenshots/issue-001-result.png"
|
|
||||||
$B snapshot -D
|
|
||||||
```
|
|
||||||
|
|
||||||
**Static bugs** (typos, layout issues, missing images):
|
|
||||||
1. Take a single annotated screenshot showing the problem
|
|
||||||
2. Describe what's wrong
|
|
||||||
|
|
||||||
```bash
|
|
||||||
$B snapshot -i -a -o "$REPORT_DIR/screenshots/issue-002.png"
|
|
||||||
```
|
|
||||||
|
|
||||||
**Write each issue to the report immediately** using the template format from `qa/templates/qa-report-template.md`.
|
|
||||||
|
|
||||||
### Phase 6: Wrap Up
|
|
||||||
|
|
||||||
1. **Compute health score** using the rubric below
|
|
||||||
2. **Write "Top 3 Things to Fix"** — the 3 highest-severity issues
|
|
||||||
3. **Write console health summary** — aggregate all console errors seen across pages
|
|
||||||
4. **Update severity counts** in the summary table
|
|
||||||
5. **Fill in report metadata** — date, duration, pages visited, screenshot count, framework
|
|
||||||
6. **Save baseline** — write `baseline.json` with:
|
|
||||||
```json
|
|
||||||
{
|
|
||||||
"date": "YYYY-MM-DD",
|
|
||||||
"url": "<target>",
|
|
||||||
"healthScore": N,
|
|
||||||
"issues": [{ "id": "ISSUE-001", "title": "...", "severity": "...", "category": "..." }],
|
|
||||||
"categoryScores": { "console": N, "links": N, ... }
|
|
||||||
}
|
|
||||||
```
|
|
||||||
|
|
||||||
**Regression mode:** After writing the report, load the baseline file. Compare:
|
|
||||||
- Health score delta
|
|
||||||
- Issues fixed (in baseline but not current)
|
|
||||||
- New issues (in current but not baseline)
|
|
||||||
- Append the regression section to the report
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
## Health Score Rubric
|
|
||||||
|
|
||||||
Compute each category score (0-100), then take the weighted average.
|
|
||||||
|
|
||||||
### Console (weight: 15%)
|
|
||||||
- 0 errors → 100
|
|
||||||
- 1-3 errors → 70
|
|
||||||
- 4-10 errors → 40
|
|
||||||
- 10+ errors → 10
|
|
||||||
|
|
||||||
### Links (weight: 10%)
|
|
||||||
- 0 broken → 100
|
|
||||||
- Each broken link → -15 (minimum 0)
|
|
||||||
|
|
||||||
### Per-Category Scoring (Visual, Functional, UX, Content, Performance, Accessibility)
|
|
||||||
Each category starts at 100. Deduct per finding:
|
|
||||||
- Critical issue → -25
|
|
||||||
- High issue → -15
|
|
||||||
- Medium issue → -8
|
|
||||||
- Low issue → -3
|
|
||||||
Minimum 0 per category.
|
|
||||||
|
|
||||||
### Weights
|
|
||||||
| Category | Weight |
|
|
||||||
|----------|--------|
|
|
||||||
| Console | 15% |
|
|
||||||
| Links | 10% |
|
|
||||||
| Visual | 10% |
|
|
||||||
| Functional | 20% |
|
|
||||||
| UX | 15% |
|
|
||||||
| Performance | 10% |
|
|
||||||
| Content | 5% |
|
|
||||||
| Accessibility | 15% |
|
|
||||||
|
|
||||||
### Final Score
|
|
||||||
`score = Σ (category_score × weight)`
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
## Framework-Specific Guidance
|
|
||||||
|
|
||||||
### Next.js
|
|
||||||
- Check console for hydration errors (`Hydration failed`, `Text content did not match`)
|
|
||||||
- Monitor `_next/data` requests in network — 404s indicate broken data fetching
|
|
||||||
- Test client-side navigation (click links, don't just `goto`) — catches routing issues
|
|
||||||
- Check for CLS (Cumulative Layout Shift) on pages with dynamic content
|
|
||||||
|
|
||||||
### Rails
|
|
||||||
- Check for N+1 query warnings in console (if development mode)
|
|
||||||
- Verify CSRF token presence in forms
|
|
||||||
- Test Turbo/Stimulus integration — do page transitions work smoothly?
|
|
||||||
- Check for flash messages appearing and dismissing correctly
|
|
||||||
|
|
||||||
### WordPress
|
|
||||||
- Check for plugin conflicts (JS errors from different plugins)
|
|
||||||
- Verify admin bar visibility for logged-in users
|
|
||||||
- Test REST API endpoints (`/wp-json/`)
|
|
||||||
- Check for mixed content warnings (common with WP)
|
|
||||||
|
|
||||||
### General SPA (React, Vue, Angular)
|
|
||||||
- Use `snapshot -i` for navigation — `links` command misses client-side routes
|
|
||||||
- Check for stale state (navigate away and back — does data refresh?)
|
|
||||||
- Test browser back/forward — does the app handle history correctly?
|
|
||||||
- Check for memory leaks (monitor console after extended use)
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
## Important Rules
|
|
||||||
|
|
||||||
1. **Repro is everything.** Every issue needs at least one screenshot. No exceptions.
|
|
||||||
2. **Verify before documenting.** Retry the issue once to confirm it's reproducible, not a fluke.
|
|
||||||
3. **Never include credentials.** Write `[REDACTED]` for passwords in repro steps.
|
|
||||||
4. **Write incrementally.** Append each issue to the report as you find it. Don't batch.
|
|
||||||
5. **Never read source code.** Test as a user, not a developer.
|
|
||||||
6. **Check console after every interaction.** JS errors that don't surface visually are still bugs.
|
|
||||||
7. **Test like a user.** Use realistic data. Walk through complete workflows end-to-end.
|
|
||||||
8. **Depth over breadth.** 5-10 well-documented issues with evidence > 20 vague descriptions.
|
|
||||||
9. **Never delete output files.** Screenshots and reports accumulate — that's intentional.
|
|
||||||
10. **Use `snapshot -C` for tricky UIs.** Finds clickable divs that the accessibility tree misses.
|
|
||||||
11. **Show screenshots to the user.** After every `$B screenshot`, `$B snapshot -a -o`, or `$B responsive` command, use the Read tool on the output file(s) so the user can see them inline. For `responsive` (3 files), Read all three. This is critical — without it, screenshots are invisible to the user.
|
|
||||||
12. **Never refuse to use the browser.** When the user invokes /qa or /qa-only, they are requesting browser-based testing. Never suggest evals, unit tests, or other alternatives as a substitute. Even if the diff appears to have no UI changes, backend changes affect app behavior — always open the browser and test.
|
|
||||||
|
|
||||||
Record baseline health score at end of Phase 6.
|
|
||||||
|
|
||||||
---
|
---
|
||||||
|
|
||||||
|
|||||||
+9
-3
@@ -40,6 +40,12 @@ triggers:
|
|||||||
|
|
||||||
You are a QA engineer AND a bug-fix engineer. Test web applications like a real user — click everything, fill every form, check every state. When you find bugs, fix them in source code with atomic commits, then re-verify. Produce a structured report with before/after evidence.
|
You are a QA engineer AND a bug-fix engineer. Test web applications like a real user — click everything, fill every form, check every state. When you find bugs, fix them in source code with atomic commits, then re-verify. Produce a structured report with before/after evidence.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
{{SECTION_INDEX:qa}}
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
## Setup
|
## Setup
|
||||||
|
|
||||||
**Parse the user's request for these parameters:**
|
**Parse the user's request for these parameters:**
|
||||||
@@ -90,7 +96,7 @@ After the user chooses, execute their choice (commit or stash), then continue wi
|
|||||||
|
|
||||||
**Check test framework (bootstrap if needed):**
|
**Check test framework (bootstrap if needed):**
|
||||||
|
|
||||||
{{TEST_BOOTSTRAP}}
|
{{SECTION:test-bootstrap}}
|
||||||
|
|
||||||
**Create output directories:**
|
**Create output directories:**
|
||||||
|
|
||||||
@@ -119,9 +125,9 @@ Before falling back to git diff heuristics, check for richer test plan sources:
|
|||||||
|
|
||||||
## Phases 1-6: QA Baseline
|
## Phases 1-6: QA Baseline
|
||||||
|
|
||||||
{{QA_METHODOLOGY}}
|
{{SECTION:qa-patterns}}
|
||||||
|
|
||||||
Record baseline health score at end of Phase 6.
|
Record baseline health score at end of Phase 6 (per the Health Score Rubric in that section).
|
||||||
|
|
||||||
---
|
---
|
||||||
|
|
||||||
|
|||||||
@@ -0,0 +1,20 @@
|
|||||||
|
{
|
||||||
|
"$schema": "https://gstack.dev/schemas/section-manifest.json",
|
||||||
|
"skill": "qa",
|
||||||
|
"version": 1,
|
||||||
|
"note": "PASSIVE registry (v2 plan T9 / CM2). Fields are IDs, file paths, human titles, and human-readable trigger text ONLY. The skeleton's decision-tree prose is the ONLY place that decides WHEN to read a section; required-reads live in the E2E fixtures. No machine predicate here — see docs/designs/v2_PLAN.md:663.",
|
||||||
|
"sections": [
|
||||||
|
{
|
||||||
|
"id": "test-bootstrap",
|
||||||
|
"file": "test-bootstrap.md",
|
||||||
|
"title": "Test Framework Bootstrap",
|
||||||
|
"trigger": "checking the project's test framework during Setup — ecosystem-marker detection, the bootstrap offer, framework install, CI pipeline generation, and first real tests (also needed at Phase 8e.5 if you skipped it and a regression test now requires a framework)"
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"id": "qa-patterns",
|
||||||
|
"file": "qa-patterns.md",
|
||||||
|
"title": "Core QA patterns — modes, Phases 1-6 workflow, health rubric, framework guidance, testing rules",
|
||||||
|
"trigger": "running the QA baseline (Phases 1-6) — mode selection (Diff-aware/Full/Quick/Regression), the phase-by-phase browser workflow, the Health Score Rubric, framework-specific guidance, and the browser-testing Important Rules"
|
||||||
|
}
|
||||||
|
]
|
||||||
|
}
|
||||||
@@ -0,0 +1,279 @@
|
|||||||
|
<!-- AUTO-GENERATED from qa-patterns.md.tmpl — do not edit directly -->
|
||||||
|
<!-- Regenerate: bun run gen:skill-docs -->
|
||||||
|
## Modes
|
||||||
|
|
||||||
|
### Diff-aware (automatic when on a feature branch with no URL)
|
||||||
|
|
||||||
|
This is the **primary mode** for developers verifying their work. When the user says `/qa` without a URL and the repo is on a feature branch, automatically:
|
||||||
|
|
||||||
|
1. **Analyze the branch diff** to understand what changed:
|
||||||
|
```bash
|
||||||
|
git diff main...HEAD --name-only
|
||||||
|
git log main..HEAD --oneline
|
||||||
|
```
|
||||||
|
|
||||||
|
2. **Identify affected pages/routes** from the changed files:
|
||||||
|
- Controller/route files → which URL paths they serve
|
||||||
|
- View/template/component files → which pages render them
|
||||||
|
- Model/service files → which pages use those models (check controllers that reference them)
|
||||||
|
- CSS/style files → which pages include those stylesheets
|
||||||
|
- API endpoints → test them directly with `$B js "await fetch('/api/...')"`
|
||||||
|
- Static pages (markdown, HTML) → navigate to them directly
|
||||||
|
|
||||||
|
**If no obvious pages/routes are identified from the diff:** Do not skip browser testing. The user invoked /qa because they want browser-based verification. Fall back to Quick mode — navigate to the homepage, follow the top 5 navigation targets, check console for errors, and test any interactive elements found. Backend, config, and infrastructure changes affect app behavior — always verify the app still works.
|
||||||
|
|
||||||
|
3. **Detect the running app** — check common local dev ports:
|
||||||
|
```bash
|
||||||
|
$B goto http://localhost:3000 2>/dev/null && echo "Found app on :3000" || \
|
||||||
|
$B goto http://localhost:4000 2>/dev/null && echo "Found app on :4000" || \
|
||||||
|
$B goto http://localhost:8080 2>/dev/null && echo "Found app on :8080"
|
||||||
|
```
|
||||||
|
If no local app is found, check for a staging/preview URL in the PR or environment. If nothing works, ask the user for the URL.
|
||||||
|
|
||||||
|
4. **Test each affected page/route:**
|
||||||
|
- Navigate to the page
|
||||||
|
- Take a screenshot
|
||||||
|
- Check console for errors
|
||||||
|
- If the change was interactive (forms, buttons, flows), test the interaction end-to-end
|
||||||
|
- Use `snapshot -D` before and after actions to verify the change had the expected effect
|
||||||
|
|
||||||
|
5. **Cross-reference with commit messages and PR description** to understand *intent* — what should the change do? Verify it actually does that.
|
||||||
|
|
||||||
|
6. **Check TODOS.md** (if it exists) for known bugs or issues related to the changed files. If a TODO describes a bug that this branch should fix, add it to your test plan. If you find a new bug during QA that isn't in TODOS.md, note it in the report.
|
||||||
|
|
||||||
|
7. **Report findings** scoped to the branch changes:
|
||||||
|
- "Changes tested: N pages/routes affected by this branch"
|
||||||
|
- For each: does it work? Screenshot evidence.
|
||||||
|
- Any regressions on adjacent pages?
|
||||||
|
|
||||||
|
**If the user provides a URL with diff-aware mode:** Use that URL as the base but still scope testing to the changed files.
|
||||||
|
|
||||||
|
### Full (default when URL is provided)
|
||||||
|
Systematic exploration. Visit every reachable page. Document 5-10 well-evidenced issues. Produce health score. Takes 5-15 minutes depending on app size.
|
||||||
|
|
||||||
|
### Quick (`--quick`)
|
||||||
|
30-second smoke test. Visit homepage + top 5 navigation targets. Check: page loads? Console errors? Broken links? Produce health score. No detailed issue documentation.
|
||||||
|
|
||||||
|
### Regression (`--regression <baseline>`)
|
||||||
|
Run full mode, then load `baseline.json` from a previous run. Diff: which issues are fixed? Which are new? What's the score delta? Append regression section to report.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Workflow
|
||||||
|
|
||||||
|
### Phase 1: Initialize
|
||||||
|
|
||||||
|
1. Find browse binary (see Setup above)
|
||||||
|
2. Create output directories
|
||||||
|
3. Copy report template from `qa/templates/qa-report-template.md` to output dir
|
||||||
|
4. Start timer for duration tracking
|
||||||
|
|
||||||
|
### Phase 2: Authenticate (if needed)
|
||||||
|
|
||||||
|
**If the user specified auth credentials:**
|
||||||
|
|
||||||
|
```bash
|
||||||
|
$B goto <login-url>
|
||||||
|
$B snapshot -i # find the login form
|
||||||
|
$B fill @e3 "user@example.com"
|
||||||
|
$B fill @e4 "[REDACTED]" # NEVER include real passwords in report
|
||||||
|
$B click @e5 # submit
|
||||||
|
$B snapshot -D # verify login succeeded
|
||||||
|
```
|
||||||
|
|
||||||
|
**If the user provided a cookie file:**
|
||||||
|
|
||||||
|
```bash
|
||||||
|
$B cookie-import cookies.json
|
||||||
|
$B goto <target-url>
|
||||||
|
```
|
||||||
|
|
||||||
|
**If 2FA/OTP is required:** Ask the user for the code and wait.
|
||||||
|
|
||||||
|
**If CAPTCHA blocks you:** Tell the user: "Please complete the CAPTCHA in the browser, then tell me to continue."
|
||||||
|
|
||||||
|
### Phase 3: Orient
|
||||||
|
|
||||||
|
Get a map of the application:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
$B goto <target-url>
|
||||||
|
$B snapshot -i -a -o "$REPORT_DIR/screenshots/initial.png"
|
||||||
|
$B links # map navigation structure
|
||||||
|
$B console --errors # any errors on landing?
|
||||||
|
```
|
||||||
|
|
||||||
|
**Detect framework** (note in report metadata):
|
||||||
|
- `__next` in HTML or `_next/data` requests → Next.js
|
||||||
|
- `csrf-token` meta tag → Rails
|
||||||
|
- `wp-content` in URLs → WordPress
|
||||||
|
- Client-side routing with no page reloads → SPA
|
||||||
|
|
||||||
|
**For SPAs:** The `links` command may return few results because navigation is client-side. Use `snapshot -i` to find nav elements (buttons, menu items) instead.
|
||||||
|
|
||||||
|
### Phase 4: Explore
|
||||||
|
|
||||||
|
Visit pages systematically. At each page:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
$B goto <page-url>
|
||||||
|
$B snapshot -i -a -o "$REPORT_DIR/screenshots/page-name.png"
|
||||||
|
$B console --errors
|
||||||
|
```
|
||||||
|
|
||||||
|
Then follow the **per-page exploration checklist** (see `qa/references/issue-taxonomy.md`):
|
||||||
|
|
||||||
|
1. **Visual scan** — Look at the annotated screenshot for layout issues
|
||||||
|
2. **Interactive elements** — Click buttons, links, controls. Do they work?
|
||||||
|
3. **Forms** — Fill and submit. Test empty, invalid, edge cases
|
||||||
|
4. **Navigation** — Check all paths in and out
|
||||||
|
5. **States** — Empty state, loading, error, overflow
|
||||||
|
6. **Console** — Any new JS errors after interactions?
|
||||||
|
7. **Responsiveness** — Check mobile viewport if relevant:
|
||||||
|
```bash
|
||||||
|
$B viewport 375x812
|
||||||
|
$B screenshot "$REPORT_DIR/screenshots/page-mobile.png"
|
||||||
|
$B viewport 1280x720
|
||||||
|
```
|
||||||
|
|
||||||
|
**Depth judgment:** Spend more time on core features (homepage, dashboard, checkout, search) and less on secondary pages (about, terms, privacy).
|
||||||
|
|
||||||
|
**Quick mode:** Only visit homepage + top 5 navigation targets from the Orient phase. Skip the per-page checklist — just check: loads? Console errors? Broken links visible?
|
||||||
|
|
||||||
|
### Phase 5: Document
|
||||||
|
|
||||||
|
Document each issue **immediately when found** — don't batch them.
|
||||||
|
|
||||||
|
**Two evidence tiers:**
|
||||||
|
|
||||||
|
**Interactive bugs** (broken flows, dead buttons, form failures):
|
||||||
|
1. Take a screenshot before the action
|
||||||
|
2. Perform the action
|
||||||
|
3. Take a screenshot showing the result
|
||||||
|
4. Use `snapshot -D` to show what changed
|
||||||
|
5. Write repro steps referencing screenshots
|
||||||
|
|
||||||
|
```bash
|
||||||
|
$B screenshot "$REPORT_DIR/screenshots/issue-001-step-1.png"
|
||||||
|
$B click @e5
|
||||||
|
$B screenshot "$REPORT_DIR/screenshots/issue-001-result.png"
|
||||||
|
$B snapshot -D
|
||||||
|
```
|
||||||
|
|
||||||
|
**Static bugs** (typos, layout issues, missing images):
|
||||||
|
1. Take a single annotated screenshot showing the problem
|
||||||
|
2. Describe what's wrong
|
||||||
|
|
||||||
|
```bash
|
||||||
|
$B snapshot -i -a -o "$REPORT_DIR/screenshots/issue-002.png"
|
||||||
|
```
|
||||||
|
|
||||||
|
**Write each issue to the report immediately** using the template format from `qa/templates/qa-report-template.md`.
|
||||||
|
|
||||||
|
### Phase 6: Wrap Up
|
||||||
|
|
||||||
|
1. **Compute health score** using the rubric below
|
||||||
|
2. **Write "Top 3 Things to Fix"** — the 3 highest-severity issues
|
||||||
|
3. **Write console health summary** — aggregate all console errors seen across pages
|
||||||
|
4. **Update severity counts** in the summary table
|
||||||
|
5. **Fill in report metadata** — date, duration, pages visited, screenshot count, framework
|
||||||
|
6. **Save baseline** — write `baseline.json` with:
|
||||||
|
```json
|
||||||
|
{
|
||||||
|
"date": "YYYY-MM-DD",
|
||||||
|
"url": "<target>",
|
||||||
|
"healthScore": N,
|
||||||
|
"issues": [{ "id": "ISSUE-001", "title": "...", "severity": "...", "category": "..." }],
|
||||||
|
"categoryScores": { "console": N, "links": N, ... }
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
**Regression mode:** After writing the report, load the baseline file. Compare:
|
||||||
|
- Health score delta
|
||||||
|
- Issues fixed (in baseline but not current)
|
||||||
|
- New issues (in current but not baseline)
|
||||||
|
- Append the regression section to the report
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Health Score Rubric
|
||||||
|
|
||||||
|
Compute each category score (0-100), then take the weighted average.
|
||||||
|
|
||||||
|
### Console (weight: 15%)
|
||||||
|
- 0 errors → 100
|
||||||
|
- 1-3 errors → 70
|
||||||
|
- 4-10 errors → 40
|
||||||
|
- 10+ errors → 10
|
||||||
|
|
||||||
|
### Links (weight: 10%)
|
||||||
|
- 0 broken → 100
|
||||||
|
- Each broken link → -15 (minimum 0)
|
||||||
|
|
||||||
|
### Per-Category Scoring (Visual, Functional, UX, Content, Performance, Accessibility)
|
||||||
|
Each category starts at 100. Deduct per finding:
|
||||||
|
- Critical issue → -25
|
||||||
|
- High issue → -15
|
||||||
|
- Medium issue → -8
|
||||||
|
- Low issue → -3
|
||||||
|
Minimum 0 per category.
|
||||||
|
|
||||||
|
### Weights
|
||||||
|
| Category | Weight |
|
||||||
|
|----------|--------|
|
||||||
|
| Console | 15% |
|
||||||
|
| Links | 10% |
|
||||||
|
| Visual | 10% |
|
||||||
|
| Functional | 20% |
|
||||||
|
| UX | 15% |
|
||||||
|
| Performance | 10% |
|
||||||
|
| Content | 5% |
|
||||||
|
| Accessibility | 15% |
|
||||||
|
|
||||||
|
### Final Score
|
||||||
|
`score = Σ (category_score × weight)`
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Framework-Specific Guidance
|
||||||
|
|
||||||
|
### Next.js
|
||||||
|
- Check console for hydration errors (`Hydration failed`, `Text content did not match`)
|
||||||
|
- Monitor `_next/data` requests in network — 404s indicate broken data fetching
|
||||||
|
- Test client-side navigation (click links, don't just `goto`) — catches routing issues
|
||||||
|
- Check for CLS (Cumulative Layout Shift) on pages with dynamic content
|
||||||
|
|
||||||
|
### Rails
|
||||||
|
- Check for N+1 query warnings in console (if development mode)
|
||||||
|
- Verify CSRF token presence in forms
|
||||||
|
- Test Turbo/Stimulus integration — do page transitions work smoothly?
|
||||||
|
- Check for flash messages appearing and dismissing correctly
|
||||||
|
|
||||||
|
### WordPress
|
||||||
|
- Check for plugin conflicts (JS errors from different plugins)
|
||||||
|
- Verify admin bar visibility for logged-in users
|
||||||
|
- Test REST API endpoints (`/wp-json/`)
|
||||||
|
- Check for mixed content warnings (common with WP)
|
||||||
|
|
||||||
|
### General SPA (React, Vue, Angular)
|
||||||
|
- Use `snapshot -i` for navigation — `links` command misses client-side routes
|
||||||
|
- Check for stale state (navigate away and back — does data refresh?)
|
||||||
|
- Test browser back/forward — does the app handle history correctly?
|
||||||
|
- Check for memory leaks (monitor console after extended use)
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Important Rules
|
||||||
|
|
||||||
|
1. **Repro is everything.** Every issue needs at least one screenshot. No exceptions.
|
||||||
|
2. **Verify before documenting.** Retry the issue once to confirm it's reproducible, not a fluke.
|
||||||
|
3. **Never include credentials.** Write `[REDACTED]` for passwords in repro steps.
|
||||||
|
4. **Write incrementally.** Append each issue to the report as you find it. Don't batch.
|
||||||
|
5. **Never read source code.** Test as a user, not a developer.
|
||||||
|
6. **Check console after every interaction.** JS errors that don't surface visually are still bugs.
|
||||||
|
7. **Test like a user.** Use realistic data. Walk through complete workflows end-to-end.
|
||||||
|
8. **Depth over breadth.** 5-10 well-documented issues with evidence > 20 vague descriptions.
|
||||||
|
9. **Never delete output files.** Screenshots and reports accumulate — that's intentional.
|
||||||
|
10. **Use `snapshot -C` for tricky UIs.** Finds clickable divs that the accessibility tree misses.
|
||||||
|
11. **Show screenshots to the user.** After every `$B screenshot`, `$B snapshot -a -o`, or `$B responsive` command, use the Read tool on the output file(s) so the user can see them inline. For `responsive` (3 files), Read all three. This is critical — without it, screenshots are invisible to the user.
|
||||||
|
12. **Never refuse to use the browser.** When the user invokes /qa or /qa-only, they are requesting browser-based testing. Never suggest evals, unit tests, or other alternatives as a substitute. Even if the diff appears to have no UI changes, backend changes affect app behavior — always open the browser and test.
|
||||||
@@ -0,0 +1 @@
|
|||||||
|
{{QA_METHODOLOGY}}
|
||||||
@@ -0,0 +1,186 @@
|
|||||||
|
<!-- AUTO-GENERATED from test-bootstrap.md.tmpl — do not edit directly -->
|
||||||
|
<!-- Regenerate: bun run gen:skill-docs -->
|
||||||
|
## Test Framework Bootstrap
|
||||||
|
|
||||||
|
**Read the project's CLAUDE.md (and TESTING.md if present) FIRST.** If it documents a test command, the project already told you: no detection, no bootstrap. Skip the rest of bootstrap and use that command in Step 5.
|
||||||
|
|
||||||
|
**Otherwise gather markers. Every marker below is EVIDENCE for the question you ask — never a command to run blind.** A marker tells you which ecosystem you're in and which command to OFFER. It does not tell you the command works. Do not execute a candidate test command to "check" it: a probe on a project that never had that runner fails loudly and teaches you nothing, and installing a second framework over a working one is worse.
|
||||||
|
|
||||||
|
```bash
|
||||||
|
setopt +o nomatch 2>/dev/null || true # zsh compat
|
||||||
|
# Definitive ecosystem markers (presence = ecosystem, NOT a command to run)
|
||||||
|
[ -f manage.py ] && echo "RUNTIME:python FRAMEWORK:django MARKER:manage.py"
|
||||||
|
{ [ -f pyproject.toml ] || [ -f pytest.ini ] || [ -f tox.ini ] || [ -f setup.cfg ] || [ -f requirements.txt ]; } && echo "RUNTIME:python"
|
||||||
|
[ -f Gemfile ] || [ -f Rakefile ] || [ -f .rspec ] && echo "RUNTIME:ruby"
|
||||||
|
[ -f package.json ] && echo "RUNTIME:node"
|
||||||
|
[ -f go.mod ] && echo "RUNTIME:go"
|
||||||
|
[ -f Cargo.toml ] && echo "RUNTIME:rust"
|
||||||
|
[ -f composer.json ] && echo "RUNTIME:php"
|
||||||
|
[ -f mix.exs ] && echo "RUNTIME:elixir"
|
||||||
|
[ -f pom.xml ] && echo "RUNTIME:jvm BUILD:maven"
|
||||||
|
{ [ -f build.gradle ] || [ -f build.gradle.kts ]; } && echo "RUNTIME:jvm BUILD:gradle"
|
||||||
|
# Detect sub-frameworks
|
||||||
|
[ -f Gemfile ] && grep -q "rails" Gemfile 2>/dev/null && echo "FRAMEWORK:rails"
|
||||||
|
[ -f package.json ] && grep -q '"next"' package.json 2>/dev/null && echo "FRAMEWORK:nextjs"
|
||||||
|
# Existing test path — config files, declared scripts, AND test FILES.
|
||||||
|
# A project with real tests and no config file is the common miss.
|
||||||
|
ls jest.config.* vitest.config.* playwright.config.* .rspec pytest.ini tox.ini phpunit.xml* 2>/dev/null
|
||||||
|
[ -f package.json ] && grep -q '"test"[[:space:]]*:' package.json && echo "SCRIPT:package.json test"
|
||||||
|
[ -f Makefile ] && grep -qE '^(test|check):' Makefile && echo "TARGET:make test"
|
||||||
|
[ -f pyproject.toml ] && grep -q "pytest" pyproject.toml && echo "CONFIG:pyproject pytest"
|
||||||
|
git ls-files | grep -cE '(^|/)(tests?|spec|__tests__)/|(^|/)tests?\.py$|(^|/)test_[^/]+\.py$|_test\.(go|py|rb|ts|js|exs)$|\.(test|spec)\.[jt]sx?$|_spec\.rb$|Test\.(java|kt)$' | sed 's/^/TESTFILES:/'
|
||||||
|
# Rust keeps unit tests inside src/, so file names alone miss them
|
||||||
|
[ -f Cargo.toml ] && git grep -lF '#[test]' -- 'src' >/dev/null 2>&1 && echo "TESTS:rust in-source"
|
||||||
|
# Check opt-out marker
|
||||||
|
[ -f .gstack/no-test-bootstrap ] && echo "BOOTSTRAP_DECLINED"
|
||||||
|
```
|
||||||
|
|
||||||
|
Map the markers to the command you will OFFER — never to one you run on a guess:
|
||||||
|
|
||||||
|
| Marker | Ecosystem | Candidate command to offer |
|
||||||
|
|--------|-----------|----------------------------|
|
||||||
|
| `manage.py` | Django | `python manage.py test` (or `pytest` when pytest-django is in the deps) |
|
||||||
|
| `pytest.ini` / `tox.ini` / pytest in `pyproject.toml` / `test_*.py` | Python | `pytest` |
|
||||||
|
| `go.mod` (+ any `*_test.go`) | Go | `go test ./...` |
|
||||||
|
| `Cargo.toml` | Rust | `cargo test` |
|
||||||
|
| `pom.xml` | JVM (Maven) | `mvn test` |
|
||||||
|
| `build.gradle` / `build.gradle.kts` | JVM (Gradle) | `./gradlew test` |
|
||||||
|
| `Gemfile` / `Rakefile` / `.rspec` | Ruby | `bundle exec rspec`, `bin/rails test`, or `rake test` |
|
||||||
|
| `mix.exs` | Elixir | `mix test` |
|
||||||
|
| `composer.json` | PHP | `composer test` or `./vendor/bin/phpunit` |
|
||||||
|
| `package.json` with a `test` script | Node | that script, run with the package manager the lockfile names |
|
||||||
|
| `Makefile` with a `test:` target | any | `make test` |
|
||||||
|
|
||||||
|
**If ANY existing-test evidence appears** (a config file, a declared test script or make target, a nonzero `TESTFILES:` count, or `TESTS:rust in-source`): the project has tests. **Do NOT bootstrap.** Print "Existing tests detected: {the evidence}." Then get the command the same way Step 5 does — CLAUDE.md/TESTING.md if documented, otherwise AskUserQuestion offering the candidates from the table above plus "Other", and persist the answer to CLAUDE.md's `## Testing` section so it is never asked again. When the ecosystem ships a runner (Django, Go, Rust, Elixir, Maven/Gradle), that runner is the candidate — never install a second framework beside a working one.
|
||||||
|
Read 2-3 existing test files to learn conventions (naming, imports, assertion style, setup patterns).
|
||||||
|
Store conventions as prose context for use in Phase 8e.5 or Step 7. **Skip the rest of bootstrap.**
|
||||||
|
|
||||||
|
Absent config files and absent `tests/` directories are NOT evidence of "no tests": Django keeps tests in `<app>/tests.py`, Go in `*_test.go` beside the source, Rust in `#[test]` blocks inside `src/`. A green `python manage.py test` with no `pytest.ini` is a tested project, not a bootstrap candidate.
|
||||||
|
|
||||||
|
**If BOOTSTRAP_DECLINED** appears: Print "Test bootstrap previously declined — skipping." **Skip the rest of bootstrap.**
|
||||||
|
|
||||||
|
**If NO ecosystem marker matched:** Use AskUserQuestion:
|
||||||
|
"I couldn't detect your project's language. What runtime are you using?"
|
||||||
|
Options: A) Node.js/TypeScript B) Ruby/Rails C) Python D) Go E) Rust F) PHP G) Elixir H) This project doesn't need tests.
|
||||||
|
If the runtime you need isn't listed, offer "Other" and take the runtime plus the test command as free text.
|
||||||
|
If user picks H → write `.gstack/no-test-bootstrap` and continue without tests.
|
||||||
|
|
||||||
|
**If an ecosystem matched but there is no existing-test evidence at all — bootstrap:**
|
||||||
|
|
||||||
|
### B2. Research best practices
|
||||||
|
|
||||||
|
Use WebSearch to find current best practices for the detected runtime:
|
||||||
|
- `"[runtime] best test framework 2025 2026"`
|
||||||
|
- `"[framework A] vs [framework B] comparison"`
|
||||||
|
|
||||||
|
If WebSearch is unavailable, use this built-in knowledge table:
|
||||||
|
|
||||||
|
| Runtime | Primary recommendation | Alternative |
|
||||||
|
|---------|----------------------|-------------|
|
||||||
|
| Ruby/Rails | minitest + fixtures + capybara | rspec + factory_bot + shoulda-matchers |
|
||||||
|
| Node.js | vitest + @testing-library | jest + @testing-library |
|
||||||
|
| Next.js | vitest + @testing-library/react + playwright | jest + cypress |
|
||||||
|
| Python | pytest + pytest-cov | unittest |
|
||||||
|
| Django | pytest + pytest-django | Django's built-in `manage.py test` (unittest) |
|
||||||
|
| Go | stdlib testing + testify | stdlib only |
|
||||||
|
| JVM (Maven/Gradle) | JUnit 5 + AssertJ | JUnit 5 only |
|
||||||
|
| Rust | cargo test (built-in) + mockall | — |
|
||||||
|
| PHP | phpunit + mockery | pest |
|
||||||
|
| Elixir | ExUnit (built-in) + ex_machina | — |
|
||||||
|
|
||||||
|
### B3. Framework selection
|
||||||
|
|
||||||
|
Use AskUserQuestion:
|
||||||
|
"I detected this is a [Runtime/Framework] project with no test framework. I researched current best practices. Here are the options:
|
||||||
|
A) [Primary] — [rationale]. Includes: [packages]. Supports: unit, integration, smoke, e2e
|
||||||
|
B) [Alternative] — [rationale]. Includes: [packages]
|
||||||
|
C) Skip — don't set up testing right now
|
||||||
|
RECOMMENDATION: Choose A because [reason based on project context]"
|
||||||
|
|
||||||
|
If user picks C → write `.gstack/no-test-bootstrap`. Tell user: "If you change your mind later, delete `.gstack/no-test-bootstrap` and re-run." Continue without tests.
|
||||||
|
|
||||||
|
If multiple runtimes detected (monorepo) → ask which runtime to set up first, with option to do both sequentially.
|
||||||
|
|
||||||
|
### B4. Install and configure
|
||||||
|
|
||||||
|
1. Install the chosen packages (npm/bun/gem/pip/etc.)
|
||||||
|
2. Create minimal config file
|
||||||
|
3. Create directory structure (test/, spec/, etc.)
|
||||||
|
4. Create one example test matching the project's code to verify setup works
|
||||||
|
|
||||||
|
If package installation fails → debug once. If still failing → revert with `git checkout -- package.json package-lock.json` (or equivalent for the runtime). Warn user and continue without tests.
|
||||||
|
|
||||||
|
### B4.5. First real tests
|
||||||
|
|
||||||
|
Generate 3-5 real tests for existing code:
|
||||||
|
|
||||||
|
1. **Find recently changed files:** `git log --since=30.days --name-only --format="" | sort | uniq -c | sort -rn | head -10`
|
||||||
|
2. **Prioritize by risk:** Error handlers > business logic with conditionals > API endpoints > pure functions
|
||||||
|
3. **For each file:** Write one test that tests real behavior with meaningful assertions. Never `expect(x).toBeDefined()` — test what the code DOES.
|
||||||
|
4. Run each test. Passes → keep. Fails → fix once. Still fails → delete silently.
|
||||||
|
5. Generate at least 1 test, cap at 5.
|
||||||
|
|
||||||
|
Never import secrets, API keys, or credentials in test files. Use environment variables or test fixtures.
|
||||||
|
|
||||||
|
### B5. Verify
|
||||||
|
|
||||||
|
```bash
|
||||||
|
# Run the full test suite to confirm everything works
|
||||||
|
{detected test command}
|
||||||
|
```
|
||||||
|
|
||||||
|
If tests fail → debug once. If still failing → revert all bootstrap changes and warn user.
|
||||||
|
|
||||||
|
### B5.5. CI/CD pipeline
|
||||||
|
|
||||||
|
```bash
|
||||||
|
# Check CI provider
|
||||||
|
ls -d .github/ 2>/dev/null && echo "CI:github"
|
||||||
|
ls .gitlab-ci.yml .circleci/ bitrise.yml 2>/dev/null
|
||||||
|
```
|
||||||
|
|
||||||
|
If `.github/` exists (or no CI detected — default to GitHub Actions):
|
||||||
|
Create `.github/workflows/test.yml` with:
|
||||||
|
- `runs-on: ubuntu-latest`
|
||||||
|
- Appropriate setup action for the runtime (setup-node, setup-ruby, setup-python, etc.)
|
||||||
|
- The same test command verified in B5
|
||||||
|
- Trigger: push + pull_request
|
||||||
|
|
||||||
|
If non-GitHub CI detected → skip CI generation with note: "Detected {provider} — CI pipeline generation supports GitHub Actions only. Add test step to your existing pipeline manually."
|
||||||
|
|
||||||
|
### B6. Create TESTING.md
|
||||||
|
|
||||||
|
First check: If TESTING.md already exists → read it and update/append rather than overwriting. Never destroy existing content.
|
||||||
|
|
||||||
|
Write TESTING.md with:
|
||||||
|
- Philosophy: "100% test coverage is the key to great vibe coding. Tests let you move fast, trust your instincts, and ship with confidence — without them, vibe coding is just yolo coding. With tests, it's a superpower."
|
||||||
|
- Framework name and version
|
||||||
|
- How to run tests (the verified command from B5)
|
||||||
|
- Test layers: Unit tests (what, where, when), Integration tests, Smoke tests, E2E tests
|
||||||
|
- Conventions: file naming, assertion style, setup/teardown patterns
|
||||||
|
|
||||||
|
### B7. Update CLAUDE.md
|
||||||
|
|
||||||
|
First check: If CLAUDE.md already has a `## Testing` section → skip. Don't duplicate.
|
||||||
|
|
||||||
|
Append a `## Testing` section:
|
||||||
|
- Run command and test directory
|
||||||
|
- Reference to TESTING.md
|
||||||
|
- Test expectations:
|
||||||
|
- 100% test coverage is the goal — tests make vibe coding safe
|
||||||
|
- When writing new functions, write a corresponding test
|
||||||
|
- When fixing a bug, write a regression test
|
||||||
|
- When adding error handling, write a test that triggers the error
|
||||||
|
- When adding a conditional (if/else, switch), write tests for BOTH paths
|
||||||
|
- Never commit code that makes existing tests fail
|
||||||
|
|
||||||
|
### B8. Commit
|
||||||
|
|
||||||
|
```bash
|
||||||
|
git status --porcelain
|
||||||
|
```
|
||||||
|
|
||||||
|
Only commit if there are changes. Stage all bootstrap files (config, test directory, TESTING.md, CLAUDE.md, .github/workflows/test.yml if created):
|
||||||
|
`git commit -m "chore: bootstrap test framework ({framework name})"`
|
||||||
|
|
||||||
|
---
|
||||||
@@ -0,0 +1 @@
|
|||||||
|
{{TEST_BOOTSTRAP}}
|
||||||
@@ -45,6 +45,7 @@ The test server is already running at: ${testServer.url}
|
|||||||
Target page: ${testServer.url}/basic.html
|
Target page: ${testServer.url}/basic.html
|
||||||
|
|
||||||
Read the file qa/SKILL.md for the QA workflow instructions.
|
Read the file qa/SKILL.md for the QA workflow instructions.
|
||||||
|
qa is a carved skill: when SKILL.md tells you to Read ~/.claude/skills/gstack/qa/sections/<file>, read qa/sections/<file> in this working directory instead (same content, local copy).
|
||||||
Skip the preamble bash block, lake intro, telemetry, and contributor mode sections — go straight to the QA workflow.
|
Skip the preamble bash block, lake intro, telemetry, and contributor mode sections — go straight to the QA workflow.
|
||||||
|
|
||||||
Run a Quick-depth QA test on ${testServer.url}/basic.html
|
Run a Quick-depth QA test on ${testServer.url}/basic.html
|
||||||
@@ -234,6 +235,7 @@ describeIfSelected('QA Fix Loop E2E', ['qa-fix-loop'], () => {
|
|||||||
prompt: `You have a browse binary at ${browseBin}. Assign it to B variable like: B="${browseBin}"
|
prompt: `You have a browse binary at ${browseBin}. Assign it to B variable like: B="${browseBin}"
|
||||||
|
|
||||||
Read the file qa/SKILL.md for the QA workflow instructions.
|
Read the file qa/SKILL.md for the QA workflow instructions.
|
||||||
|
qa is a carved skill: when SKILL.md tells you to Read ~/.claude/skills/gstack/qa/sections/<file>, read qa/sections/<file> in this working directory instead (same content, local copy).
|
||||||
Skip the preamble bash block, lake intro, telemetry, and contributor mode sections — go straight to the QA workflow.
|
Skip the preamble bash block, lake intro, telemetry, and contributor mode sections — go straight to the QA workflow.
|
||||||
|
|
||||||
Run a Quick-tier QA test on ${qaFixUrl}
|
Run a Quick-tier QA test on ${qaFixUrl}
|
||||||
|
|||||||
Reference in New Issue
Block a user