mirror of
https://github.com/garrytan/gstack.git
synced 2026-10-08 04:11:18 +02:00
v1.91.7.0 feat: add functional QA and pre-publication docs checks (#2983)
* feat: add surface-aware exploratory QA and ship documentation gates * test: preserve delegated QA setup authority after main integration * fix(qa): clarify exploration order and preserve report artifacts * test(qa): follow the shared setup reference directly * refactor(ship): make verification and recovery routes explicit * test(ship): align evidence and review guards with explicit routes * fix(workflows): clarify ship recovery and functional QA evidence * fix(workflows): clarify approval recovery and full QA coverage * refactor(workflows): order review transactions and clarify ship state * fix(ship): clarify final verification and fail closed at publication * fix(evals): attribute native atomic documentation writes * fix(ship): clarify recovery and documentation lifecycle guidance * fix(test): preserve observed native placeholder styling in CI * fix(codex): report watchdog timeouts without a process-exit race * Checkpoint functional QA implementation and workflow validation repairs * Fix documentation and shared-review fixture contracts * docs: clarify judge reuse and evaluation supervision * test: align review evidence and selected case contracts * test: verify append-only documentation checkpoints and recovery * fix: qualify QA workflows and CI validation repairs * fix: launch shared-libs fixture scripts on Windows * fix: qualify QA deadlines, fixture isolation, and shard cleanup * fix: preserve qualified QA and cancellation repairs * fix: enforce functional fixture authority and share strict event decoding * fix: retain free-test evidence and explain recovery * fix: reject malformed native evidence after decoder consolidation * test: use reliable capture for telemetry privacy filters * test: refresh measured quick coverage and document validation costs * Fix native fixture receipts and preserve VM validation evidence * Align negative judge controls with upstream clarity policy * Fix report-only QA preparation and public evidence handling * Clarify QA-only preparation and current-report preservation * Stream Ship quality judgments with an explicit 64k response contract * Validate compact judge reasoning locally with supported wire schema * Align functional QA fixture instructions with evidence acceptance * Bind native browser diagnostics to execution evidence and align review verdicts * Preserve native diagnostic line boundaries * Serialize functional QA evidence from native captures * Keep large QA evidence fixture payload out of Windows argv
This commit is contained in:
1 parent
65bfb0ce49
commit
dcaea52800
333 files changed
+41755
-7357
No files matched your search
+118
-185
@@ -3,13 +3,12 @@ name: qa
|
||||
preamble-tier: 4
|
||||
version: 2.0.0
|
||||
description: |
|
||||
Systematically QA test a web application and fix bugs found. Runs QA testing,
|
||||
then iteratively fixes bugs in source code, committing each fix atomically and
|
||||
re-verifying. Use when asked to "qa", "QA", "test this site", "find bugs",
|
||||
Fix browser/API/CLI/job/worker/webhook bugs.
|
||||
Commit verified fixes atomically. Use when asked to "qa", "QA", "test this site", "find bugs",
|
||||
"test and fix", or "fix what's broken".
|
||||
Proactively suggest when the user says a feature is ready for testing
|
||||
or asks "does this work?". Three tiers: Quick (critical/high only),
|
||||
Standard (+ medium), Exhaustive (+ cosmetic). Produces before/after health scores,
|
||||
Standard (+ medium), Exhaustive (+ cosmetic). Produces contract outcomes or browser health scores,
|
||||
fix evidence, and a ship-readiness summary. For report-only mode, use /qa-only. (gstack)
|
||||
voice-triggers:
|
||||
- "quality check"
|
||||
@@ -38,8 +37,6 @@ triggers:
|
||||
|
||||
# /qa: Test → Fix → Verify
|
||||
|
||||
You are a QA engineer AND a bug-fix engineer. Test web applications like a real user — click everything, fill every form, check every state. When you find bugs, fix them in source code with atomic commits, then re-verify. Produce a structured report with before/after evidence.
|
||||
|
||||
---
|
||||
|
||||
{{SECTION_INDEX:qa}}
|
||||
@@ -48,23 +45,31 @@ You are a QA engineer AND a bug-fix engineer. Test web applications like a real
|
||||
|
||||
## Setup
|
||||
|
||||
{{SECTION:scope}}
|
||||
|
||||
**Parse the user's request for these parameters:**
|
||||
|
||||
| Parameter | Default | Override example |
|
||||
|-----------|---------|-----------------:|
|
||||
| Target URL | (auto-detect or required) | `https://myapp.com`, `http://localhost:3000` |
|
||||
| Target | (infer from request/repository or ask) | Browser URL, API route, CLI command, job, worker or webhook |
|
||||
| Tier | Standard | `--quick`, `--exhaustive` |
|
||||
| Mode | full | `--regression .gstack/qa-reports/baseline.json` |
|
||||
| Mode | full | `--quick`, `--regression <previous-report-or-baseline>` |
|
||||
| Output dir | `.gstack/qa-reports/` | `Output to /tmp/qa` |
|
||||
| Scope | Full app (or diff-scoped) | `Focus on the billing page` |
|
||||
| Auth | Your Aside session (already signed in) | If a sign-in wall appears, you sign in yourself in Aside — no credentials in chat (see BROWSER SETUP). Fallback browser only: /setup-browser-cookies or `$B handoff` |
|
||||
| Scope | Selected target (or diff-scoped) | `Focus on duplicate webhook delivery` |
|
||||
| Auth | Isolated synthetic identity for functional probes | Browser session handling lives in browser setup; never request credentials in chat |
|
||||
|
||||
**Tiers determine which issues get fixed:**
|
||||
- **Quick:** Fix critical + high severity only
|
||||
- **Standard:** + medium severity (default)
|
||||
- **Exhaustive:** + low/cosmetic severity
|
||||
|
||||
**If no URL is given and you're on a feature branch:** Automatically enter **diff-aware mode** (see Modes below). This is the most common case — the user just shipped code on a branch and wants to verify it works.
|
||||
`--quick` also selects Quick exploration; `--exhaustive` changes only the fix tier.
|
||||
Regression mode preserves the selected fix tier.
|
||||
If both `--quick` and `--regression` are supplied, ask which exploration mode to use
|
||||
before setup or probes. Keep the selected fix tier; this choice concerns exploration only.
|
||||
|
||||
**On a feature branch without an explicit scope:** Use diff-aware testing of changed
|
||||
and adjacent behavior. Select the surface first; absence of a URL never forces a browser.
|
||||
|
||||
**Check for clean working tree:**
|
||||
|
||||
@@ -72,104 +77,89 @@ You are a QA engineer AND a bug-fix engineer. Test web applications like a real
|
||||
git status --porcelain
|
||||
```
|
||||
|
||||
If the output is non-empty (working tree is dirty), **STOP** and use AskUserQuestion:
|
||||
If dirty, **STOP** and use AskUserQuestion. Explain that a clean tree keeps QA fixes atomic:
|
||||
- A) Commit all current changes with a descriptive message before QA (recommended).
|
||||
- B) Stash changes, run QA, then pop the stash.
|
||||
- C) Abort for manual cleanup.
|
||||
|
||||
"Your working tree has uncommitted changes. /qa needs a clean tree so each bug fix gets its own atomic commit."
|
||||
Execute only the user's choice before continuing setup.
|
||||
|
||||
- A) Commit my changes — commit all current changes with a descriptive message, then start QA
|
||||
- B) Stash my changes — stash, run QA, pop the stash after
|
||||
- C) Abort — I'll clean up manually
|
||||
**Prepare report artifacts before browser setup.** Resolve any supplied prior report
|
||||
and baseline paths before writing. Select the output override or `.gstack/qa-reports`.
|
||||
Create that directory if absent. Use the directory as `REPORT_DIR`
|
||||
only when it is empty; otherwise choose a fresh owned run subdirectory.
|
||||
Use `run-YYYYMMDDTHHMMSSZ` in UTC, adding a suffix on collision. Keep all local evidence there.
|
||||
Never overwrite previous reports, baselines, screenshots or exploration notes.
|
||||
A caller's fixed artifact paths and permissions take precedence; if preserving them
|
||||
safely is impossible, report the output blocker rather than expanding write authority.
|
||||
|
||||
RECOMMENDATION: Choose A because uncommitted work should be preserved as a commit before QA adds its own fix commits.
|
||||
**Browser surface only:** load its setup; functional-only runs skip this section.
|
||||
|
||||
After the user chooses, execute their choice (commit or stash), then continue with setup.
|
||||
{{SECTION:browser-setup}}
|
||||
|
||||
**Browser: Aside**
|
||||
|
||||
{{ASIDE_SETUP}}
|
||||
|
||||
{{BROWSE_FALLBACK}}
|
||||
|
||||
**Check test framework (bootstrap if needed):**
|
||||
**Browser surface only:** check the test framework and use the existing bootstrap
|
||||
offer if needed. Functional targets use supported native tests or report the gap;
|
||||
they do not load this browser bootstrap or generate CI.
|
||||
|
||||
{{SECTION:test-bootstrap}}
|
||||
|
||||
**Create output directories:**
|
||||
|
||||
```bash
|
||||
REPORT_DIR=".gstack/qa-reports"
|
||||
mkdir -p "$REPORT_DIR/screenshots"
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
{{LEARNINGS_SEARCH:query=qa testing bug regression flake fixture}}
|
||||
|
||||
## Test Plan Context
|
||||
|
||||
Before falling back to git diff heuristics, check for richer test plan sources:
|
||||
Prefer the richer of recent project test plans and plans in conversation over git diff:
|
||||
|
||||
1. **Project-scoped test plans:** Check `~/.gstack/projects/` for recent `*-test-plan-*.md` files for this repo
|
||||
1. **Project-scoped test plans:** Find the latest for this repo:
|
||||
```bash
|
||||
setopt +o nomatch 2>/dev/null || true # zsh compat
|
||||
{{SLUG_EVAL}}
|
||||
ls -t ~/.gstack/projects/$SLUG/*-test-plan-*.md 2>/dev/null | head -1
|
||||
```
|
||||
2. **Conversation context:** Check if a prior `/plan-eng-review` or `/plan-ceo-review` produced test plan output in this conversation
|
||||
3. **Use whichever source is richer.** Fall back to git diff analysis only if neither is available.
|
||||
2. **Conversation context:** Prior `/plan-eng-review` or `/plan-ceo-review` test plans.
|
||||
3. Fall back to git diff only if neither exists.
|
||||
|
||||
---
|
||||
|
||||
## Phases 1-6: QA Baseline
|
||||
|
||||
{{SECTION:qa-patterns}}
|
||||
Follow the shared section's ordered preparation, then run its probe loop.
|
||||
The numbered browser phases label techniques, not another workflow.
|
||||
|
||||
Record baseline health score at end of Phase 6 (per the Health Score Rubric in that section).
|
||||
{{SECTION:exploratory}}
|
||||
|
||||
Report baseline findings before fixing. Keep browser scores and functional outcomes separate.
|
||||
|
||||
---
|
||||
|
||||
## Output Structure
|
||||
|
||||
```
|
||||
.gstack/qa-reports/
|
||||
├── qa-report-{domain}-{YYYY-MM-DD}.md # Structured report
|
||||
├── screenshots/
|
||||
│ ├── initial.jpg # Landing page screenshot
|
||||
│ ├── issue-001-step-1.jpg # Per-issue evidence
|
||||
│ ├── issue-001-result.jpg
|
||||
│ ├── issue-002.png # Annotated screenshot (static bugs)
|
||||
│ ├── issue-001-after.jpg # After fix (if fixed); the Phase 5 evidence is the before
|
||||
│ └── ...
|
||||
└── baseline.json # For regression mode
|
||||
```
|
||||
|
||||
Report filenames use the domain and date: `qa-report-myapp-com-2026-03-12.md`
|
||||
Under `$REPORT_DIR`, write `qa-report-{target}-{YYYY-MM-DD}.md` and the browser's
|
||||
`baseline.json`. Browser `{target}` is a safe hostname.
|
||||
Browser evidence goes in `screenshots/`: `initial.jpg`,
|
||||
`issue-NNN-step-N.jpg`, `issue-NNN-result.jpg`, annotated `issue-NNN.png` and
|
||||
`issue-NNN-after.jpg` (Phase 5 is the before). Functional reports use a safe command/service
|
||||
label and sanitized command/request/state evidence.
|
||||
|
||||
---
|
||||
|
||||
## Phase 7: Triage
|
||||
|
||||
Sort all discovered issues by severity, then decide which to fix based on the selected tier:
|
||||
|
||||
- **Quick:** Fix critical + high only. Mark medium/low as "deferred."
|
||||
- **Standard:** Fix critical + high + medium. Mark low as "deferred."
|
||||
- **Exhaustive:** Fix all, including cosmetic/low severity.
|
||||
|
||||
Mark issues that cannot be fixed from source code (e.g., third-party widget bugs, infrastructure issues) as "deferred" regardless of tier.
|
||||
Sort issues by severity and apply the selected fix tier. Mark lower-tier issues and
|
||||
those not fixable from source (third-party widgets, infrastructure) as "deferred."
|
||||
|
||||
### Refresh learnings for the component/page where the bug lives
|
||||
|
||||
The top-of-skill learnings pull was keyed to "qa testing" broadly. Before the fix loop, re-pull learnings keyed to the component or page where the bug you're about to fix lives so prior fixes for the same component-shape surface.
|
||||
|
||||
Pick ONE keyword that names the buggy component or page. The keyword should be a noun: the failing component name, the page route base, or the feature noun. The keyword MUST be alphanumeric or hyphen only — no quotes, slashes, dots, colons, or whitespace. If your candidate has any of those, simplify to just the alphanumeric stem.
|
||||
|
||||
Worked examples (qa-specific): good keywords are `checkout-button`, `signup-form`, `payment`. Bad: `tests are failing`, `<failing-test>`, `app/views/_checkout.html.erb`.
|
||||
Before the fix loop, search again for the buggy component/page. Use ONE noun containing
|
||||
only letters, digits or hyphens (e.g., `checkout-button`, `payment`), never a path,
|
||||
quotes, whitespace or other punctuation; simplify to an alphanumeric stem if needed.
|
||||
|
||||
```bash
|
||||
~/.claude/skills/gstack/bin/gstack-learnings-search --query "<your-keyword>" --limit 5 2>/dev/null || true
|
||||
```
|
||||
|
||||
If any learnings come back, name which one applies to the fix you're about to make in one sentence. If none come back, continue without reference — the absence is itself useful information.
|
||||
Name an applicable learning in one sentence, or continue if none applies.
|
||||
|
||||
---
|
||||
|
||||
@@ -177,123 +167,70 @@ If any learnings come back, name which one applies to the fix you're about to ma
|
||||
|
||||
For each fixable issue, in severity order:
|
||||
|
||||
### 8a. Locate source
|
||||
### 8a. Diagnose and reproduce
|
||||
|
||||
```bash
|
||||
# Grep for error messages, component names, route definitions
|
||||
# Glob for file patterns matching the affected page
|
||||
Use the shared loop's causal hypothesis and minimized replay, recording actual versus
|
||||
documented behavior before edits. Modify only responsible files. Environment failures
|
||||
and unclear contracts never authorize repair.
|
||||
|
||||
### 8a.5. Regression test before repair
|
||||
|
||||
Match 2-3 nearby tests' naming, imports, assertions and fixtures. Reproduce the failure
|
||||
in a new native test. Run its detected command before repair; prove the defect caused its
|
||||
failure, not a bad fixture, import or service. Attribute it in the language's comment syntax:
|
||||
|
||||
```text
|
||||
// Regression: ISSUE-NNN — short defect description
|
||||
// Found by /qa on YYYY-MM-DD
|
||||
// Report: .gstack/qa-reports/qa-report-{target}-{date}.md
|
||||
```
|
||||
|
||||
- Find the source file(s) responsible for the bug
|
||||
- ONLY modify files directly related to the issue
|
||||
A clear, healthy uncovered contract may gain a passing test without product edits.
|
||||
|
||||
Apply the shared exploratory section's native unit/integration/E2E rules.
|
||||
CSS-only defects may use browser evidence. Missing infrastructure stays coverage debt.
|
||||
|
||||
Use the component's name and native extension in auto-incrementing `{name}.regression-N.test.{ext}`.
|
||||
Set N to max number + 1, starting at 1; never replace an existing file.
|
||||
Keep valid red regressions; narrowly correct a proved
|
||||
fixture/test error or report the unresolved bug.
|
||||
|
||||
### 8b. Fix
|
||||
|
||||
- Read the source code, understand the context
|
||||
- Make the **minimal fix** — smallest change that resolves the issue
|
||||
- Do NOT refactor surrounding code, add features, or "improve" unrelated things
|
||||
Read the surrounding source and make the **minimal fix**. No unrelated refactors or features.
|
||||
|
||||
### 8c. Commit
|
||||
### 8c. Re-test
|
||||
|
||||
Re-run the regression, original failing probe and adjacent happy path. Inspect each
|
||||
final state; acceptance alone cannot verify a worker repair. Failed/unavailable rechecks stay unresolved.
|
||||
|
||||
For browser defects only:
|
||||
|
||||
{{SECTION:browser-verify}}
|
||||
|
||||
### 8d. Commit verified work
|
||||
|
||||
```bash
|
||||
git add <only-changed-files>
|
||||
git add <only-verified-source-and-regression-files>
|
||||
git commit -m "fix(qa): ISSUE-NNN — short description"
|
||||
```
|
||||
|
||||
- One commit per fix. Never bundle multiple fixes.
|
||||
- Message format: `fix(qa): ISSUE-NNN — short description`
|
||||
|
||||
### 8d. Re-test
|
||||
|
||||
- Navigate back to the affected page
|
||||
- Take **before/after screenshot pair** — the Phase 5 evidence is the before; capture the after now
|
||||
- Check console for errors
|
||||
- Compare the snapshot tree and `CONSOLE_ERRORS=` against the Phase 5 evidence to verify the change had the expected effect
|
||||
|
||||
One flow, one script (tabs close when the script ends, so re-navigate from the URL):
|
||||
|
||||
```bash
|
||||
aside repl '
|
||||
const HOOK = `(() => { window.__gstackErrs = window.__gstackErrs || []; const oe = console.error; console.error = (...a) => { window.__gstackErrs.push(a.map(String).join(" ")); oe.apply(console, a); }; window.addEventListener("error", e => window.__gstackErrs.push("uncaught: " + e.message)); })()`;
|
||||
const pg = await openTab("about:blank");
|
||||
await pg._sendToTarget("Page.addScriptToEvaluateOnNewDocument", { source: HOOK });
|
||||
await pg.goto("<affected-url>");
|
||||
const s = await snapshot(pg, { interactive: true });
|
||||
console.log(s.tree);
|
||||
console.log("CONSOLE_ERRORS=" + JSON.stringify(await pg.evaluate(() => window.__gstackErrs)));
|
||||
await pg.screenshot({ path: "issue-NNN-after.jpg", type: "jpeg", quality: 60, fullPage: true });
|
||||
console.log("ASIDE_DIR=" + pwd);
|
||||
await closeTab(pg);
|
||||
console.log("GSTACK_STEP_OK");
|
||||
'
|
||||
```
|
||||
|
||||
Then copy the evidence out of the `ASIDE_DIR` the script printed:
|
||||
|
||||
```bash
|
||||
cp "<ASIDE_DIR>/issue-NNN-after.jpg" "$REPORT_DIR/screenshots/issue-NNN-after.jpg"
|
||||
```
|
||||
|
||||
Read `$REPORT_DIR/screenshots/issue-NNN-after.jpg` so the user sees the after state inline. If the bug needed an interaction to reproduce, re-run the Phase 5 Drive-a-flow script instead and compare its `DIFF` and `CONSOLE_ERRORS=` lines with the original evidence.
|
||||
Commit each verified fix with its regression, never unrelated fixes. Leave unresolved
|
||||
repairs and valid red regressions/evidence uncommitted; tell the user what remains.
|
||||
|
||||
### 8e. Classify
|
||||
|
||||
- **verified**: re-test confirms the fix works, no new errors introduced
|
||||
- **verified**: passed 8c (native regression when available); disclose missing test coverage
|
||||
- **best-effort**: fix applied but couldn't fully verify (e.g., needs auth state, external service)
|
||||
- **reverted**: regression detected → `git revert HEAD` → mark issue as "deferred"
|
||||
- **reverted**: regression detected → undo only this run's repair (revert its commit if already committed), retain the valid regression/evidence, and mark the issue "deferred". Never discard user changes.
|
||||
|
||||
### 8e.5. Regression Test
|
||||
### 8e.5. Regression Test record
|
||||
|
||||
Skip if: classification is not "verified", OR the fix is purely visual/CSS with no JS behavior, OR no test framework was detected AND user declined bootstrap.
|
||||
|
||||
**1. Study the project's existing test patterns:**
|
||||
|
||||
Read 2-3 test files closest to the fix (same directory, same code type). Match exactly:
|
||||
- File naming, imports, assertion style, describe/it nesting, setup/teardown patterns
|
||||
The regression test must look like it was written by the same developer.
|
||||
|
||||
**2. Trace the bug's codepath, then write a regression test:**
|
||||
|
||||
Before writing the test, trace the data flow through the code you just fixed:
|
||||
- What input/state triggered the bug? (the exact precondition)
|
||||
- What codepath did it follow? (which branches, which function calls)
|
||||
- Where did it break? (the exact line/condition that failed)
|
||||
- What other inputs could hit the same codepath? (edge cases around the fix)
|
||||
|
||||
The test MUST:
|
||||
- Set up the precondition that triggered the bug (the exact state that made it break)
|
||||
- Perform the action that exposed the bug
|
||||
- Assert the correct behavior (NOT "it renders" or "it doesn't throw")
|
||||
- If you found adjacent edge cases while tracing, test those too (e.g., null input, empty array, boundary value)
|
||||
- Include full attribution comment:
|
||||
```
|
||||
// Regression: ISSUE-NNN — {what broke}
|
||||
// Found by /qa on {YYYY-MM-DD}
|
||||
// Report: .gstack/qa-reports/qa-report-{domain}-{date}.md
|
||||
```
|
||||
|
||||
Test type decision:
|
||||
- Console error / JS exception / logic bug → unit or integration test
|
||||
- Broken form / API failure / data flow bug → integration test with request/response
|
||||
- Visual bug with JS behavior (broken dropdown, animation) → component test
|
||||
- Pure CSS → skip (caught by QA reruns)
|
||||
|
||||
Generate unit tests. Mock all external dependencies (DB, API, Redis, file system).
|
||||
|
||||
Use auto-incrementing names to avoid collisions: check existing `{name}.regression-*.test.{ext}` files, take max number + 1.
|
||||
|
||||
**3. Run only the new test file:**
|
||||
|
||||
```bash
|
||||
{detected test command} {new-test-file}
|
||||
```
|
||||
|
||||
**4. Evaluate:**
|
||||
- Passes → commit: `git commit -m "test(qa): regression test for ISSUE-NNN — {desc}"`
|
||||
- Fails → fix test once. Still failing → delete test, defer.
|
||||
- Taking >2 min exploration → skip and defer.
|
||||
|
||||
**5. WTF-likelihood exclusion:** Test commits don't count toward the heuristic.
|
||||
Record the test created before repair in 8a.5 and its re-test result from 8c:
|
||||
file, command, attribution, tested boundary and red/green evidence, or why it is deferred.
|
||||
This step records results; it does not create another test.
|
||||
Healthy-contract commits use `test(qa): regression test for {contract}`.
|
||||
**WTF-likelihood exclusion:** test-only commits do not count toward the heuristic.
|
||||
|
||||
### 8f. Self-Regulation (STOP AND EVALUATE)
|
||||
|
||||
@@ -317,19 +254,16 @@ WTF-LIKELIHOOD:
|
||||
|
||||
## Phase 9: Final QA
|
||||
|
||||
After all fixes are applied:
|
||||
|
||||
1. Re-run QA on all affected pages
|
||||
2. Compute final health score
|
||||
3. **If final score is WORSE than baseline:** WARN prominently — something regressed
|
||||
Re-run affected contracts and adjacent happy paths on the final inputs.
|
||||
Caller-required rechecks cannot be skipped as unaffected. For browser
|
||||
surfaces, recheck affected pages and compute the final health score. Warn prominently
|
||||
about a worse score or regressed contract; blocked/inconclusive rechecks never verify repairs.
|
||||
|
||||
---
|
||||
|
||||
## Phase 10: Report
|
||||
|
||||
Write the report to both local and project-scoped locations:
|
||||
|
||||
**Local:** `.gstack/qa-reports/qa-report-{domain}-{YYYY-MM-DD}.md`
|
||||
Write the Output Structure report locally and copy the same content to project context:
|
||||
|
||||
**Project-scoped:** Write test outcome artifact for cross-session context:
|
||||
```bash
|
||||
@@ -337,21 +271,22 @@ Write the report to both local and project-scoped locations:
|
||||
```
|
||||
Write to `~/.gstack/projects/{slug}/{user}-{branch}-test-outcome-{datetime}.md`
|
||||
|
||||
**Per-issue additions** (beyond standard report template):
|
||||
**Per-issue additions:**
|
||||
- Fix Status: verified / best-effort / reverted / deferred
|
||||
- Commit SHA (if fixed)
|
||||
- Files Changed (if fixed)
|
||||
- Before/After screenshots (if fixed)
|
||||
- Before/After evidence: screenshots for browser, outputs/requests/durable state for functional
|
||||
|
||||
**Summary section:**
|
||||
- Total issues found
|
||||
- Fixes applied (verified: X, best-effort: Y, reverted: Z)
|
||||
- Deferred issues
|
||||
- Health score delta: baseline → final
|
||||
**Summary:** total issues, verified/best-effort/reverted fixes and deferred issues.
|
||||
For browser coverage include the score delta. For functional coverage include
|
||||
passing/failing/blocked/not-run contracts, permanent regressions and remaining risks,
|
||||
never a score. Keep mixed results separate.
|
||||
|
||||
**PR Summary:** Include a one-line summary suitable for PR descriptions:
|
||||
**PR Summary:** Include one line:
|
||||
> "QA found N issues, fixed M, health score X → Y."
|
||||
|
||||
For functional targets, use those contract outcomes instead of a score in the PR summary.
|
||||
|
||||
---
|
||||
|
||||
## Phase 11: TODOS.md Update
|
||||
@@ -369,8 +304,6 @@ If the repo has a `TODOS.md`:
|
||||
|
||||
## Additional Rules (qa-specific)
|
||||
|
||||
11. **Clean working tree required.** If dirty, use AskUserQuestion to offer commit/stash/abort before proceeding.
|
||||
12. **One commit per fix.** Never bundle multiple fixes into one commit.
|
||||
13. **Only modify tests when generating regression tests in Phase 8e.5.** Never modify CI configuration. Never modify existing tests — only create new test files.
|
||||
14. **Revert on regression.** If a fix makes things worse, `git revert HEAD` immediately.
|
||||
15. **Self-regulate.** Follow the WTF-likelihood heuristic. When in doubt, stop and ask.
|
||||
**Outside an explicitly approved browser bootstrap:** Only create tests through authorized codification in Phase 8a.5. Never modify CI configuration or weaken existing tests; use new native test files.
|
||||
|
||||
When in doubt, stop and ask.
|
||||
Reference in new issue
Block a user