mirror of
https://github.com/garrytan/gstack.git
synced 2026-10-02 17:40:02 +02:00
v1.91.7.0 feat: add functional QA and pre-publication docs checks (#2983)
* feat: add surface-aware exploratory QA and ship documentation gates * test: preserve delegated QA setup authority after main integration * fix(qa): clarify exploration order and preserve report artifacts * test(qa): follow the shared setup reference directly * refactor(ship): make verification and recovery routes explicit * test(ship): align evidence and review guards with explicit routes * fix(workflows): clarify ship recovery and functional QA evidence * fix(workflows): clarify approval recovery and full QA coverage * refactor(workflows): order review transactions and clarify ship state * fix(ship): clarify final verification and fail closed at publication * fix(evals): attribute native atomic documentation writes * fix(ship): clarify recovery and documentation lifecycle guidance * fix(test): preserve observed native placeholder styling in CI * fix(codex): report watchdog timeouts without a process-exit race * Checkpoint functional QA implementation and workflow validation repairs * Fix documentation and shared-review fixture contracts * docs: clarify judge reuse and evaluation supervision * test: align review evidence and selected case contracts * test: verify append-only documentation checkpoints and recovery * fix: qualify QA workflows and CI validation repairs * fix: launch shared-libs fixture scripts on Windows * fix: qualify QA deadlines, fixture isolation, and shard cleanup * fix: preserve qualified QA and cancellation repairs * fix: enforce functional fixture authority and share strict event decoding * fix: retain free-test evidence and explain recovery * fix: reject malformed native evidence after decoder consolidation * test: use reliable capture for telemetry privacy filters * test: refresh measured quick coverage and document validation costs * Fix native fixture receipts and preserve VM validation evidence * Align negative judge controls with upstream clarity policy * Fix report-only QA preparation and public evidence handling * Clarify QA-only preparation and current-report preservation * Stream Ship quality judgments with an explicit 64k response contract * Validate compact judge reasoning locally with supported wire schema * Align functional QA fixture instructions with evidence acceptance * Bind native browser diagnostics to execution evidence and align review verdicts * Preserve native diagnostic line boundaries * Serialize functional QA evidence from native captures * Keep large QA evidence fixture payload out of Windows argv
This commit is contained in:
1 parent
65bfb0ce49
commit
dcaea52800
333 files changed
+41755
-7357
No files matched your search
+31
-5
@@ -171,6 +171,22 @@ Bun auto-loads `.env` — no extra config. Conductor workspaces inherit `.env` f
|
||||
|
||||
### Test tiers
|
||||
|
||||
Functional QA changes need native fixture proof as well as prompt checks. Add declared
|
||||
CLI or loopback API/worker contracts in isolated temporary repositories, outside this
|
||||
checkout. Exercise success and adverse paths, durable effects, and setup failure.
|
||||
Report-only evaluations must leave mutation-capable tools available and independently
|
||||
detect forbidden writes, including an edit later restored; a clean final diff is not
|
||||
enough. Validate the observer with deliberately bad controls before a paid run.
|
||||
|
||||
For exploratory regressions, retain the actual pre-repair failure, post-repair pass,
|
||||
original probe and adjacent happy path. Automatic caller tests must enter through
|
||||
review/ship, not tell the agent to run the component being tested. Documentation tests
|
||||
must prove the real child completed and the parent used its result before publication;
|
||||
the existing dispatch-only test is narrower evidence. Register new cases and all
|
||||
consumed section/resolver inputs in touchfiles, tiers and the PR profile so they run.
|
||||
Share sanitized reproduction commands and fixture evidence when reporting a problem,
|
||||
never credentials, private payloads or an entire unreviewed agent transcript.
|
||||
|
||||
| Tier | Command | Cost | What it tests |
|
||||
|------|---------|------|---------------|
|
||||
| 1 — Static | `bun run test` | Free | Command validation, snapshot flags, Aside contract pins, render-wrapper option mapping, SKILL.md correctness, TODOS-format.md refs, observability unit tests |
|
||||
@@ -196,10 +212,10 @@ gate and periodic censuses run fresh weekly and on manual
|
||||
dispatch of `evals-periodic.yml`; `bun run eval:bg:release` runs both locally.
|
||||
Some broad behavioral failures will therefore be found after the PR gate.
|
||||
|
||||
CI enables verified first-attempt reuse for the 14 workflow quality judges for
|
||||
24 hours within the same PR. The other 11 quality cases and all dynamic agent
|
||||
cases stay fresh. Local runs stay fresh unless the complete scoped cache and
|
||||
runtime configuration is supplied. The key includes complete prompt bytes, generated inputs,
|
||||
CI enables verified first-attempt reuse for 16 workflow quality judges for
|
||||
24 hours within the same PR. The cookie workflow's custom input, the other 11
|
||||
quality cases and all dynamic agent cases stay fresh. Local runs stay fresh unless
|
||||
the complete scoped cache and runtime configuration is supplied. The key includes complete prompt bytes, generated inputs,
|
||||
fixtures, runner/rubric code, installed dependencies, model settings and runtime.
|
||||
The current assertions validate a reused score again. Records retain the original
|
||||
run, revision and time; reuse never renews that time. Failed, retried, partial or
|
||||
@@ -218,6 +234,11 @@ historical six-worker result below and the
|
||||
[four-CPU portfolio comparison](docs/TEST_PORTFOLIO.md#measurement-contract)
|
||||
are machine-specific measurements. CI setup, build and queue time are reported
|
||||
separately. Refresh measurements with `bun run test:ubicloud --record-durations`;
|
||||
before publication, classify new regressions for quick feedback using that seed
|
||||
and the existing `QUICK_CORE` list. Do not classify unknown files as fast or use
|
||||
quick results as release acceptance. The runner retains full logs in
|
||||
`.context/free-test-logs/` and explains the next repair step on failure; see
|
||||
[free-runner recovery](docs/TESTING_INTERNALS.md) for details. For full acceptance,
|
||||
the required free CI lane packs the complete inventory across isolated runners,
|
||||
then checks every shard's receipt before reporting success. Local worker counts
|
||||
remain bounded to avoid browser/process contention.
|
||||
@@ -282,7 +303,7 @@ Spawns `claude -p` as a subprocess with `--output-format stream-json --verbose`,
|
||||
|
||||
```bash
|
||||
# Must run from a plain terminal — can't nest inside Claude Code or Conductor
|
||||
EVALS=1 bun test test/skill-e2e-*.test.ts
|
||||
EVALS_RUN_ID="local-$(bun -e 'console.log(crypto.randomUUID())')" EVALS=1 bun test test/skill-e2e-*.test.ts
|
||||
```
|
||||
|
||||
- Gated by `EVALS=1` env var (prevents accidental expensive runs)
|
||||
@@ -292,6 +313,11 @@ EVALS=1 bun test test/skill-e2e-*.test.ts
|
||||
- Saves full NDJSON transcripts and failure JSON for debugging
|
||||
- Tests live in `test/skill-e2e-*.test.ts` (split by category), runner logic in `test/helpers/session-runner.ts`
|
||||
|
||||
Supply a fresh `EVALS_RUN_ID` for each invocation, including detached runs below.
|
||||
Functional QA and documentation cases refuse acceptance without it. CI supplies
|
||||
its own run/attempt/job/slice identity; see [Testing internals](docs/TESTING_INTERNALS.md)
|
||||
for the retained native-capture artifacts.
|
||||
|
||||
**Hermetic by default.** Every E2E runner (claude -p, the real-PTY plan-mode
|
||||
runner, the Agent SDK runner, plus the codex and gemini runners) spawns its child
|
||||
through `test/helpers/hermetic-env.ts`: an allowlist-scrubbed environment, a fresh
|
||||
|
||||
Reference in new issue
Block a user