v1.91.7.0 feat: add functional QA and pre-publication docs checks (#2983)

* feat: add surface-aware exploratory QA and ship documentation gates

* test: preserve delegated QA setup authority after main integration

* fix(qa): clarify exploration order and preserve report artifacts

* test(qa): follow the shared setup reference directly

* refactor(ship): make verification and recovery routes explicit

* test(ship): align evidence and review guards with explicit routes

* fix(workflows): clarify ship recovery and functional QA evidence

* fix(workflows): clarify approval recovery and full QA coverage

* refactor(workflows): order review transactions and clarify ship state

* fix(ship): clarify final verification and fail closed at publication

* fix(evals): attribute native atomic documentation writes

* fix(ship): clarify recovery and documentation lifecycle guidance

* fix(test): preserve observed native placeholder styling in CI

* fix(codex): report watchdog timeouts without a process-exit race

* Checkpoint functional QA implementation and workflow validation repairs

* Fix documentation and shared-review fixture contracts

* docs: clarify judge reuse and evaluation supervision

* test: align review evidence and selected case contracts

* test: verify append-only documentation checkpoints and recovery

* fix: qualify QA workflows and CI validation repairs

* fix: launch shared-libs fixture scripts on Windows

* fix: qualify QA deadlines, fixture isolation, and shard cleanup

* fix: preserve qualified QA and cancellation repairs

* fix: enforce functional fixture authority and share strict event decoding

* fix: retain free-test evidence and explain recovery

* fix: reject malformed native evidence after decoder consolidation

* test: use reliable capture for telemetry privacy filters

* test: refresh measured quick coverage and document validation costs

* Fix native fixture receipts and preserve VM validation evidence

* Align negative judge controls with upstream clarity policy

* Fix report-only QA preparation and public evidence handling

* Clarify QA-only preparation and current-report preservation

* Stream Ship quality judgments with an explicit 64k response contract

* Validate compact judge reasoning locally with supported wire schema

* Align functional QA fixture instructions with evidence acceptance

* Bind native browser diagnostics to execution evidence and align review verdicts

* Preserve native diagnostic line boundaries

* Serialize functional QA evidence from native captures

* Keep large QA evidence fixture payload out of Windows argv
This commit is contained in:
Garry Tan authored and GitHub committed 2026-09-29 06:07:35 -07:00
1 parent 65bfb0ce49
commit dcaea52800
333 files changed
+41755 -7357

No files matched your search

+86 -17
View File
@@ -127,6 +127,21 @@ on a Mac they drive Aside and on Linux CI they drive the built browse binary,
skipping only when neither exists. The `$B`-driven E2E cases and `browse/test/`
run on every platform as before, so Linux CI proves the fallback engine live.
**Bootstrap dependency retention is opt-in qualification, not the behavior test.**
`qa-bootstrap` still runs its original unpinned Vitest installation and assertions
on macOS and Linux, including documented unsharded commands. Only a Linux paid
shard runner issues the owned retention scope: it binds each actual fixture and
native lifetime, retains locks, package manifests and the installed file/link
inventory, and acknowledges capture before deleting the fixture. The outer
runner also captures evidence when a callback is killed. Incomplete capture
fails qualification and preserves the source fixture as well as partial evidence.
Other platforms explicitly report retention as unavailable and still execute the
native behavior test. A run without the runner-issued scope earns no retained
dependency qualification credit; candidate acceptance requiring that evidence
must use the Linux sharded path and verify every attempt's complete capture,
acknowledgment and cleanup fallback. A passing unsharded or macOS behavior test
does not substitute for that evidence.
**The renderer picks the same way, so the render gates are engine-agnostic.**
`/make-pdf`, `/diagram`, and design previews print and screenshot their local
HTML through `lib/aside-render.ts` / `bin/gstack-render.ts`, which render in
@@ -177,13 +192,35 @@ fallback; unknown files get 75th-percentile pessimism, and both full-suite and
the long pole. Packed
shards get duration-aware walls (`max(base, predicted × 3, files × 5s)`). The
legacy `--shards N --shard i` path keeps stable hash indices. Required CI uses
one duration-packed `--ci-plan`, 20 isolated `--ci-run` machines, and a
`--ci-verify` aggregate. `TREE_MUTATING` is EMPTY:
`gen-skill-docs.ts` has a `main()` guard (imports never regenerate; pinned by
`test/gen-skill-docs-import-purity.test.ts`) and `--out-dir` renders every
host, so all former mutators render into mkdtemps and the trailing serial
shard is gone. The map remains a mechanism — a test that genuinely must write
shared artifacts in place earns a reasoned entry and is serialized again.
one duration-packed `--ci-plan`, 20 ordinary `--ci-run` shards plus a separate
exclusive-fixture shard, and a `--ci-verify` aggregate. CI shards still run on
independent machines without ordering unrelated jobs. The public
`TREE_MUTATING` map now classifies exclusive host-state fixtures; its sole
entry is `test/bootstrap-retention.test.ts`, whose same-UID nondumpable actors
affect host-wide procfs permission checks. Locally, this file runs only after
all parallel shards settle, and cancellation prevents that final phase from
starting. No selected files, retries, budgets, or receipt requirements are
removed. Former generator mutators still render into private output directories;
this serial phase protects process visibility, not in-place doc generation.
Full child output is retained in private files under `.context/free-test-logs/`,
outside each shard's temporary cleanup directory. The runner prints the path at
launch and completion. Losing the log fails the run even when the child exits
successfully. A redirected log directory is rejected before launching a child.
On failure, read that log first: the recovery message distinguishes incomplete
capture, unconfirmed cleanup, deadline expiry and a test/module failure. Fix the
demonstrated cause before rerunning. A focused `bun test` command is offered only
when every failure is attributable to existing selected files; it proves that
repair, not completion of the original selection. Preserve failed attempts when
sharing results, and inspect logs for private data before sharing them.
Before publication, classify new deterministic regressions for quick feedback.
Refresh the timing seed with the existing recorder on fixed inputs; do not edit
source while tests run. Critical boundary controls belong in `QUICK_CORE` when
their feedback cost is justified. Other measured files qualify at two seconds
or less; slow and unmeasured files remain outside quick, not outside full tests.
Report cold setup separately from warm execution, while retaining failed-attempt,
retry and cleanup time in the total cost.
**PTY fixture timing.** Plan-count sessions wake on terminal output or exit,
with at least 250ms between expensive observations and a 2s fallback for
@@ -230,6 +267,21 @@ key must name a living paid test (`test/touchfiles.test.ts`'s reverse
invariant), and `git show <sha>:path` fixtures are banned — vendor the bytes
instead (`test/git-ref-fixture-tripwire.test.ts`).
Functional QA and documentation acceptance require an explicit `EVALS_RUN_ID`;
`GSTACK_EVAL_DIR` alone does not satisfy their evidence-ownership guard. For each
local invocation, supply a fresh ID to the documented detached runner:
```bash
EVALS_RUN_ID="local-$(bun -e 'console.log(crypto.randomUUID())')" bun run eval:bg:pr
```
Neither package scripts nor `gstack-detach` invent this identity. CI's PR/manual
slices, periodic slices and weekly gate census supply an ID bound to the workflow
run, attempt, job and slice. Existing slice artifacts retain per-shard snapshots;
separate always-run `native-captures-<EVALS_RUN_ID>` artifacts retain project/legacy
`e2e-runs` and `evals/qa-callers` evidence for 90 days. These diagnostic artifacts
are not collector results and do not establish that an unfinished test passed.
**Fast PR profile and evidence reuse.** `test:pr` selects the changed cases in
`scripts/test-pr-profile.ts` plus every changed quality judge. `--profile full`
retains the broad census; no case IDs or tier assignments are removed. The plan
@@ -246,8 +298,9 @@ with matching before/after inputs. The audited workflow-judge adapter hashes the
actual expanded prompt, source/fixture/rubric/runner closure, installed SDK,
model parameters and runtime. Missing/unknown inputs force execution. Receipts
are scoped to the same repository and PR, expire after 24 hours, and contain
public scores and provenance rather than prompts or secrets. Only the 14 cases
using `runWorkflowJudge` are eligible; the other 11 quality cases remain fresh.
public scores and provenance rather than prompts or secrets. Of the 17 cases
using `runWorkflowJudge`, 16 are eligible; the cookie workflow's custom input
does not match the cache adapter and stays fresh, as do the other 11 quality cases.
CI supplies the scoped cache/runtime configuration; local runs are fresh by
default. Cached scores must
pass current assertions; reused records retain their original source and time
@@ -325,12 +378,24 @@ including two minutes for cleanup. No per-case budget grows. Overlay wrappers
have a 1,830-second minimum shard wall and run without Bun retries; see the
[overlay contract](OVERLAY_BENCHMARK_CONTRACT.md) for their unchanged work budget.
The quality file reserves 6,400 seconds for all 25 cases and their existing
retry, plus cleanup. Each still has 120 seconds of model work. Its 14 workflow
The quality file reserves 7,180 seconds for all 28 cases and their existing
retry, plus cleanup. Each still has 120 seconds of model work. Its 17 workflow
judges own their deadline and abort signal, with five seconds for terminal
recording inside a ten-second Bun grace; the other 11 retain their existing
120-second Bun timeout. Late responses cannot create records or cache passes.
The ship documentation file reserves 10,920 seconds for five 600-second cases and
eight 300-second fault cases, each with one retry, plus cleanup. The standalone
documentation child retains its 600-second case. The five review/ship explorer
cases reserve 3,270 seconds including their existing retry and finalization grace.
These are whole-file supervision limits, not additional model work per case.
The shared-library path file reserves 3,720 seconds for its three serial
600-second cases, each with one retry, plus 120 seconds for cleanup. Its
registered budget keeps the file in its own shard and binds the expected wall
to both the saved plan and the execution receipt; missing or stale budget
records fail reconciliation. Case deadlines, model budgets and retries do not grow.
`resolvePaidShardBudget(files, overrideMs?)` is the canonical per-job resolver.
Autoplan, each registered finding file, and each overlay wrapper require their
own shard, even with `--files-per-shard` above one. Mixed or multi-file overlay
@@ -341,20 +406,24 @@ Planner entries and execution results record the effective wall,
its source and policy identifier. Custom drivers must resolve each job instead
of passing their ordinary 1800-second default as an explicit Autoplan cap;
their outer controller/detach wall must also cover the allocated work and cleanup.
`eval:bg:pr` and `eval:bg:periodic` have 72000/66000-second outer caps; the PR
The current paid census has 122 files: 61 gate-tier and 103 periodic-tier.
`eval:bg:pr` and `eval:bg:periodic` have 92820/67380-second outer caps; the PR
wrapper covers a full-gate fallback at its default two workers. The broad gate
wrapper reserves 33600 seconds, and release reserves 100000 seconds for both
wrapper reserves 49320 seconds, and release reserves 116700 seconds for both
tiers. Legacy monolithic
`eval:bg`/`eval:bg:all` retain their shorter 5400/7200-second caps and do not
promise two complete Autoplan attempts; use the sharded periodic path for this policy.
Periodic CI plans `--slices 8 --autoplan-slice`: the eighth runs only Autoplan.
When overlays are selected, the seventh is reserved for their serial wrappers;
Periodic CI plans `--slices 9 --autoplan-slice`: the ninth runs only Autoplan.
When overlays are selected, the eighth is reserved for their serial wrappers;
registered finding files are distributed across the remaining ordinary slices
by their supervised walls. Each slice job has a 355-minute cap; Autoplan retains
by their supervised walls. Each slice job has a 360-minute cap; Autoplan retains
its 172-minute shard wall. Reconciliation rejects missing, duplicated or misplaced
registered work and absent budget records. The weekly gate census has a
350-minute cap and PR slices have a 220-minute cap. Free supervision tests
352-minute cap across eight single-worker slices with at most four running at
once. Its longest current work wall is 302 minutes. PR slices retain seven
two-worker slices with a 265-minute cap for their 242-minute work wall plus
setup. Free supervision tests
verify these bounds against the complete current census, configured retries,
and setup reserve. Ordinary paid tiers and the default 1800-second
shard wall remain unchanged; the registered and overlay policies above supply