Merge origin/main (v1.91.7.0) into test-audit-reduction

Keep both intents: v1.91.7.0's functional QA, docsync and exploratory
paid cases and their free owners stay; this branch's deletions stay
deleted. main's new paid keys follow the derived-closure touchfile rule
(free *.test.ts paths dropped, static helper/fixture closure added), its
new helper-only tests join the ratchet baseline, and its free selection
examples that named free test files now assert the derived selection.

Periodic CI keeps seven slices without the retired Autoplan slice; the
gate census keeps seven single-worker slices with --skip-judges. Wall
and census literals are recomputed from the merged planner, durations
are re-recorded on Ubicloud, and VERSION stays 1.91.8.0 above 1.91.7.0.
This commit is contained in:
garrytan committed 2026-09-29 13:48:02 +00:00
commit b421bba2c9
325 files changed
+42569 -8257

No files matched your search

+31 -5
View File
@@ -171,6 +171,22 @@ Bun auto-loads `.env` — no extra config. Conductor workspaces inherit `.env` f
### Test tiers
Functional QA changes need native fixture proof as well as prompt checks. Add declared
CLI or loopback API/worker contracts in isolated temporary repositories, outside this
checkout. Exercise success and adverse paths, durable effects, and setup failure.
Report-only evaluations must leave mutation-capable tools available and independently
detect forbidden writes, including an edit later restored; a clean final diff is not
enough. Validate the observer with deliberately bad controls before a paid run.
For exploratory regressions, retain the actual pre-repair failure, post-repair pass,
original probe and adjacent happy path. Automatic caller tests must enter through
review/ship, not tell the agent to run the component being tested. Documentation tests
must prove the real child completed and the parent used its result before publication;
the existing dispatch-only test is narrower evidence. Register new cases and all
consumed section/resolver inputs in touchfiles, tiers and the PR profile so they run.
Share sanitized reproduction commands and fixture evidence when reporting a problem,
never credentials, private payloads or an entire unreviewed agent transcript.
| Tier | Command | Cost | What it tests |
|------|---------|------|---------------|
| 1 — Static | `bun run test` | Free | Command validation, snapshot flags, Aside contract pins, render-wrapper option mapping, SKILL.md correctness, TODOS-format.md refs, observability unit tests |
@@ -196,10 +212,10 @@ gate and periodic censuses run fresh weekly and on manual
dispatch of `evals-periodic.yml`; `bun run eval:bg:release` runs both locally.
Some broad behavioral failures will therefore be found after the PR gate.
CI enables verified first-attempt reuse for the 14 workflow quality judges for
24 hours within the same PR. The other 11 quality cases and all dynamic agent
cases stay fresh. Local runs stay fresh unless the complete scoped cache and
runtime configuration is supplied. The key includes complete prompt bytes, generated inputs,
CI enables verified first-attempt reuse for 16 workflow quality judges for
24 hours within the same PR. The cookie workflow's custom input, the other 11
quality cases and all dynamic agent cases stay fresh. Local runs stay fresh unless
the complete scoped cache and runtime configuration is supplied. The key includes complete prompt bytes, generated inputs,
fixtures, runner/rubric code, installed dependencies, model settings and runtime.
The current assertions validate a reused score again. Records retain the original
run, revision and time; reuse never renews that time. Failed, retried, partial or
@@ -218,6 +234,11 @@ historical six-worker result below and the
[four-CPU portfolio comparison](docs/TEST_PORTFOLIO.md#measurement-contract)
are machine-specific measurements. CI setup, build and queue time are reported
separately. Refresh measurements with `bun run test:ubicloud --record-durations`;
before publication, classify new regressions for quick feedback using that seed
and the existing `QUICK_CORE` list. Do not classify unknown files as fast or use
quick results as release acceptance. The runner retains full logs in
`.context/free-test-logs/` and explains the next repair step on failure; see
[free-runner recovery](docs/TESTING_INTERNALS.md) for details. For full acceptance,
the required free CI lane packs the complete inventory across isolated runners,
then checks every shard's receipt before reporting success. Local worker counts
remain bounded to avoid browser/process contention.
@@ -297,7 +318,7 @@ Spawns `claude -p` as a subprocess with `--output-format stream-json --verbose`,
```bash
# Must run from a plain terminal — can't nest inside Claude Code or Conductor
EVALS=1 bun test test/skill-e2e-*.test.ts
EVALS_RUN_ID="local-$(bun -e 'console.log(crypto.randomUUID())')" EVALS=1 bun test test/skill-e2e-*.test.ts
```
- Gated by `EVALS=1` env var (prevents accidental expensive runs)
@@ -307,6 +328,11 @@ EVALS=1 bun test test/skill-e2e-*.test.ts
- Saves full NDJSON transcripts and failure JSON for debugging
- Tests live in `test/skill-e2e-*.test.ts` (split by category), runner logic in `test/helpers/session-runner.ts`
Supply a fresh `EVALS_RUN_ID` for each invocation, including detached runs below.
Functional QA and documentation cases refuse acceptance without it. CI supplies
its own run/attempt/job/slice identity; see [Testing internals](docs/TESTING_INTERNALS.md)
for the retained native-capture artifacts.
**Hermetic by default.** Every E2E runner (claude -p, the real-PTY plan-mode
runner, the Agent SDK runner, plus the codex and gemini runners) spawns its child
through `test/helpers/hermetic-env.ts`: an allowlist-scrubbed environment, a fresh