v1.91.7.0 feat: add functional QA and pre-publication docs checks (#2983)

* feat: add surface-aware exploratory QA and ship documentation gates

* test: preserve delegated QA setup authority after main integration

* fix(qa): clarify exploration order and preserve report artifacts

* test(qa): follow the shared setup reference directly

* refactor(ship): make verification and recovery routes explicit

* test(ship): align evidence and review guards with explicit routes

* fix(workflows): clarify ship recovery and functional QA evidence

* fix(workflows): clarify approval recovery and full QA coverage

* refactor(workflows): order review transactions and clarify ship state

* fix(ship): clarify final verification and fail closed at publication

* fix(evals): attribute native atomic documentation writes

* fix(ship): clarify recovery and documentation lifecycle guidance

* fix(test): preserve observed native placeholder styling in CI

* fix(codex): report watchdog timeouts without a process-exit race

* Checkpoint functional QA implementation and workflow validation repairs

* Fix documentation and shared-review fixture contracts

* docs: clarify judge reuse and evaluation supervision

* test: align review evidence and selected case contracts

* test: verify append-only documentation checkpoints and recovery

* fix: qualify QA workflows and CI validation repairs

* fix: launch shared-libs fixture scripts on Windows

* fix: qualify QA deadlines, fixture isolation, and shard cleanup

* fix: preserve qualified QA and cancellation repairs

* fix: enforce functional fixture authority and share strict event decoding

* fix: retain free-test evidence and explain recovery

* fix: reject malformed native evidence after decoder consolidation

* test: use reliable capture for telemetry privacy filters

* test: refresh measured quick coverage and document validation costs

* Fix native fixture receipts and preserve VM validation evidence

* Align negative judge controls with upstream clarity policy

* Fix report-only QA preparation and public evidence handling

* Clarify QA-only preparation and current-report preservation

* Stream Ship quality judgments with an explicit 64k response contract

* Validate compact judge reasoning locally with supported wire schema

* Align functional QA fixture instructions with evidence acceptance

* Bind native browser diagnostics to execution evidence and align review verdicts

* Preserve native diagnostic line boundaries

* Serialize functional QA evidence from native captures

* Keep large QA evidence fixture payload out of Windows argv
This commit is contained in:
Garry Tan authored and GitHub committed 2026-09-29 06:07:35 -07:00
1 parent 65bfb0ce49
commit dcaea52800
333 files changed
+41755 -7357

No files matched your search

+31 -5
View File
@@ -171,6 +171,22 @@ Bun auto-loads `.env` — no extra config. Conductor workspaces inherit `.env` f
### Test tiers
Functional QA changes need native fixture proof as well as prompt checks. Add declared
CLI or loopback API/worker contracts in isolated temporary repositories, outside this
checkout. Exercise success and adverse paths, durable effects, and setup failure.
Report-only evaluations must leave mutation-capable tools available and independently
detect forbidden writes, including an edit later restored; a clean final diff is not
enough. Validate the observer with deliberately bad controls before a paid run.
For exploratory regressions, retain the actual pre-repair failure, post-repair pass,
original probe and adjacent happy path. Automatic caller tests must enter through
review/ship, not tell the agent to run the component being tested. Documentation tests
must prove the real child completed and the parent used its result before publication;
the existing dispatch-only test is narrower evidence. Register new cases and all
consumed section/resolver inputs in touchfiles, tiers and the PR profile so they run.
Share sanitized reproduction commands and fixture evidence when reporting a problem,
never credentials, private payloads or an entire unreviewed agent transcript.
| Tier | Command | Cost | What it tests |
|------|---------|------|---------------|
| 1 — Static | `bun run test` | Free | Command validation, snapshot flags, Aside contract pins, render-wrapper option mapping, SKILL.md correctness, TODOS-format.md refs, observability unit tests |
@@ -196,10 +212,10 @@ gate and periodic censuses run fresh weekly and on manual
dispatch of `evals-periodic.yml`; `bun run eval:bg:release` runs both locally.
Some broad behavioral failures will therefore be found after the PR gate.
CI enables verified first-attempt reuse for the 14 workflow quality judges for
24 hours within the same PR. The other 11 quality cases and all dynamic agent
cases stay fresh. Local runs stay fresh unless the complete scoped cache and
runtime configuration is supplied. The key includes complete prompt bytes, generated inputs,
CI enables verified first-attempt reuse for 16 workflow quality judges for
24 hours within the same PR. The cookie workflow's custom input, the other 11
quality cases and all dynamic agent cases stay fresh. Local runs stay fresh unless
the complete scoped cache and runtime configuration is supplied. The key includes complete prompt bytes, generated inputs,
fixtures, runner/rubric code, installed dependencies, model settings and runtime.
The current assertions validate a reused score again. Records retain the original
run, revision and time; reuse never renews that time. Failed, retried, partial or
@@ -218,6 +234,11 @@ historical six-worker result below and the
[four-CPU portfolio comparison](docs/TEST_PORTFOLIO.md#measurement-contract)
are machine-specific measurements. CI setup, build and queue time are reported
separately. Refresh measurements with `bun run test:ubicloud --record-durations`;
before publication, classify new regressions for quick feedback using that seed
and the existing `QUICK_CORE` list. Do not classify unknown files as fast or use
quick results as release acceptance. The runner retains full logs in
`.context/free-test-logs/` and explains the next repair step on failure; see
[free-runner recovery](docs/TESTING_INTERNALS.md) for details. For full acceptance,
the required free CI lane packs the complete inventory across isolated runners,
then checks every shard's receipt before reporting success. Local worker counts
remain bounded to avoid browser/process contention.
@@ -282,7 +303,7 @@ Spawns `claude -p` as a subprocess with `--output-format stream-json --verbose`,
```bash
# Must run from a plain terminal — can't nest inside Claude Code or Conductor
EVALS=1 bun test test/skill-e2e-*.test.ts
EVALS_RUN_ID="local-$(bun -e 'console.log(crypto.randomUUID())')" EVALS=1 bun test test/skill-e2e-*.test.ts
```
- Gated by `EVALS=1` env var (prevents accidental expensive runs)
@@ -292,6 +313,11 @@ EVALS=1 bun test test/skill-e2e-*.test.ts
- Saves full NDJSON transcripts and failure JSON for debugging
- Tests live in `test/skill-e2e-*.test.ts` (split by category), runner logic in `test/helpers/session-runner.ts`
Supply a fresh `EVALS_RUN_ID` for each invocation, including detached runs below.
Functional QA and documentation cases refuse acceptance without it. CI supplies
its own run/attempt/job/slice identity; see [Testing internals](docs/TESTING_INTERNALS.md)
for the retained native-capture artifacts.
**Hermetic by default.** Every E2E runner (claude -p, the real-PTY plan-mode
runner, the Agent SDK runner, plus the codex and gemini runners) spawns its child
through `test/helpers/hermetic-env.ts`: an allowlist-scrubbed environment, a fresh