mirror of
https://github.com/garrytan/gstack.git
synced 2026-10-02 17:40:02 +02:00
* feat: add surface-aware exploratory QA and ship documentation gates * test: preserve delegated QA setup authority after main integration * fix(qa): clarify exploration order and preserve report artifacts * test(qa): follow the shared setup reference directly * refactor(ship): make verification and recovery routes explicit * test(ship): align evidence and review guards with explicit routes * fix(workflows): clarify ship recovery and functional QA evidence * fix(workflows): clarify approval recovery and full QA coverage * refactor(workflows): order review transactions and clarify ship state * fix(ship): clarify final verification and fail closed at publication * fix(evals): attribute native atomic documentation writes * fix(ship): clarify recovery and documentation lifecycle guidance * fix(test): preserve observed native placeholder styling in CI * fix(codex): report watchdog timeouts without a process-exit race * Checkpoint functional QA implementation and workflow validation repairs * Fix documentation and shared-review fixture contracts * docs: clarify judge reuse and evaluation supervision * test: align review evidence and selected case contracts * test: verify append-only documentation checkpoints and recovery * fix: qualify QA workflows and CI validation repairs * fix: launch shared-libs fixture scripts on Windows * fix: qualify QA deadlines, fixture isolation, and shard cleanup * fix: preserve qualified QA and cancellation repairs * fix: enforce functional fixture authority and share strict event decoding * fix: retain free-test evidence and explain recovery * fix: reject malformed native evidence after decoder consolidation * test: use reliable capture for telemetry privacy filters * test: refresh measured quick coverage and document validation costs * Fix native fixture receipts and preserve VM validation evidence * Align negative judge controls with upstream clarity policy * Fix report-only QA preparation and public evidence handling * Clarify QA-only preparation and current-report preservation * Stream Ship quality judgments with an explicit 64k response contract * Validate compact judge reasoning locally with supported wire schema * Align functional QA fixture instructions with evidence acceptance * Bind native browser diagnostics to execution evidence and align review verdicts * Preserve native diagnostic line boundaries * Serialize functional QA evidence from native captures * Keep large QA evidence fixture payload out of Windows argv
83 lines
4.2 KiB
Markdown
83 lines
4.2 KiB
Markdown
# Testing Specialist Review Checklist
|
|
|
|
Scope: Always-on (every review)
|
|
Output: JSON objects, one finding per line. Schema:
|
|
{"severity":"CRITICAL|INFORMATIONAL","confidence":N,"path":"file","line":N,"category":"testing","summary":"...","fix":"...","fingerprint":"path:line:testing","specialist":"testing"}
|
|
Optional: line, fix, fingerprint, evidence, test_stub.
|
|
If no findings: output `NO FINDINGS` and nothing else.
|
|
|
|
If the caller explicitly asks for an ASCII coverage diagram, first read the
|
|
named source and test files directly in a dedicated tool call. Keep that tool
|
|
call limited to those two file reads, using either native Read calls or a simple
|
|
shell display such as `cat -n src/file && echo ---- && cat -n test/file`. Read
|
|
diffs, package files, configs, or other context in separate tool calls. Then
|
|
output the diagram before any JSON findings. Use one function root per public or
|
|
changed function, put the `[OK]` or `[GAP]` marker on the same branch row as the
|
|
tested or missing path, and keep the legend inside the same diagram block:
|
|
|
|
```text
|
|
src/billing.ts
|
|
processPayment(amount, currency)
|
|
├── valid USD happy path returns success [OK]
|
|
└── invalid amount / unsupported currency branches [GAP]
|
|
refundPayment(paymentId, reason)
|
|
└── refund success and guard branches not imported or untested [GAP]
|
|
Legend: [OK] tested [GAP] no test
|
|
```
|
|
|
|
Do not rely on a summary table, prose paragraph, or distant nested marker as the
|
|
only coverage evidence. Covered happy-path rows must name the successful, valid,
|
|
or concrete tested input path; gap rows must stay under the function that owns
|
|
the missing path.
|
|
|
|
---
|
|
|
|
## Categories
|
|
|
|
### Exploratory hypotheses
|
|
|
|
The parent runs one shared exploratory QA pass before Fix-First, including small diffs
|
|
that skip specialists. Supply high-risk changed contracts, adverse scenarios and native
|
|
test candidates to that pass; do not launch another explorer, edit product/tests or
|
|
commit. A test proposal uses `test_stub` and keeps the caller's approval requirements.
|
|
Check real request/queue/storage effects, retries, duplicates and interrupted recovery
|
|
where applicable. A diagram or test count is not executed proof. For discoveries promoted
|
|
to regressions, require failure for the reproduced bug before repair and green plus
|
|
original/adjacent probes afterward; never accept buggy-output goldens or discarded valid
|
|
red tests. Unit tests suit logic; real integration/E2E tests protect boundaries mocks hide.
|
|
|
|
### Missing Negative-Path Tests
|
|
- New code paths that handle errors, rejections, or invalid input with NO corresponding test
|
|
- Guard clauses and early returns that are untested
|
|
- Error branches in try/catch, rescue, or error boundaries with no failure-path test
|
|
- Permission/auth checks that are asserted in code but never tested for the "denied" case
|
|
|
|
### Missing Edge-Case Coverage
|
|
- Boundary values: zero, negative, max-int, empty string, empty array, nil/null/undefined
|
|
- Single-element collections (off-by-one on loops)
|
|
- Unicode and special characters in user-facing inputs
|
|
- Concurrent access patterns with no race-condition test
|
|
|
|
### Test Isolation Violations
|
|
- Tests sharing mutable state (class variables, global singletons, DB records not cleaned up)
|
|
- Order-dependent tests (pass in sequence, fail when randomized)
|
|
- Tests that depend on system clock, timezone, or locale
|
|
- Tests that make real network calls instead of using stubs/mocks
|
|
|
|
### Flaky Test Patterns
|
|
- Timing-dependent assertions (sleep, setTimeout, waitFor with tight timeouts)
|
|
- Assertions on ordering of unordered results (hash keys, Set iteration, async resolution order)
|
|
- Tests that depend on external services (APIs, databases) without fallback
|
|
- Randomized test data without seed control
|
|
|
|
### Security Enforcement Tests Missing
|
|
- Auth/authz checks in controllers with no test for the "unauthorized" case
|
|
- Rate limiting logic with no test proving it actually blocks
|
|
- Input sanitization with no test for malicious input
|
|
- CSRF/CORS configuration with no integration test
|
|
|
|
### Coverage Gaps
|
|
- New public methods/functions with zero test coverage
|
|
- Changed methods where existing tests only cover the old behavior, not the new branch
|
|
- Utility functions called from multiple places but tested only indirectly
|