Merge origin/main (v1.91.7.0) into test-audit-reduction

Keep both intents: v1.91.7.0's functional QA, docsync and exploratory
paid cases and their free owners stay; this branch's deletions stay
deleted. main's new paid keys follow the derived-closure touchfile rule
(free *.test.ts paths dropped, static helper/fixture closure added), its
new helper-only tests join the ratchet baseline, and its free selection
examples that named free test files now assert the derived selection.

Periodic CI keeps seven slices without the retired Autoplan slice; the
gate census keeps seven single-worker slices with --skip-judges. Wall
and census literals are recomputed from the merged planner, durations
are re-recorded on Ubicloud, and VERSION stays 1.91.8.0 above 1.91.7.0.
This commit is contained in:
garrytan committed 2026-09-29 13:48:02 +00:00
commit b421bba2c9
325 files changed
+42569 -8257

No files matched your search

+84 -15
View File
@@ -127,6 +127,21 @@ on a Mac they drive Aside and on Linux CI they drive the built browse binary,
skipping only when neither exists. The `$B`-driven E2E cases and `browse/test/`
run on every platform as before, so Linux CI proves the fallback engine live.
**Bootstrap dependency retention is opt-in qualification, not the behavior test.**
`qa-bootstrap` still runs its original unpinned Vitest installation and assertions
on macOS and Linux, including documented unsharded commands. Only a Linux paid
shard runner issues the owned retention scope: it binds each actual fixture and
native lifetime, retains locks, package manifests and the installed file/link
inventory, and acknowledges capture before deleting the fixture. The outer
runner also captures evidence when a callback is killed. Incomplete capture
fails qualification and preserves the source fixture as well as partial evidence.
Other platforms explicitly report retention as unavailable and still execute the
native behavior test. A run without the runner-issued scope earns no retained
dependency qualification credit; candidate acceptance requiring that evidence
must use the Linux sharded path and verify every attempt's complete capture,
acknowledgment and cleanup fallback. A passing unsharded or macOS behavior test
does not substitute for that evidence.
**The renderer picks the same way, so the render gates are engine-agnostic.**
`/make-pdf`, `/diagram`, and design previews print and screenshot their local
HTML through `lib/aside-render.ts` / `bin/gstack-render.ts`, which render in
@@ -177,13 +192,35 @@ fallback; unknown files get 75th-percentile pessimism, and both full-suite and
the long pole. Packed
shards get duration-aware walls (`max(base, predicted × 3, files × 5s)`). The
legacy `--shards N --shard i` path keeps stable hash indices. Required CI uses
one duration-packed `--ci-plan`, 20 isolated `--ci-run` machines, and a
`--ci-verify` aggregate. `TREE_MUTATING` is EMPTY:
`gen-skill-docs.ts` has a `main()` guard (imports never regenerate; pinned by
`test/gen-skill-docs-import-purity.test.ts`) and `--out-dir` renders every
host, so all former mutators render into mkdtemps and the trailing serial
shard is gone. The map remains a mechanism — a test that genuinely must write
shared artifacts in place earns a reasoned entry and is serialized again.
one duration-packed `--ci-plan`, 20 ordinary `--ci-run` shards plus a separate
exclusive-fixture shard, and a `--ci-verify` aggregate. CI shards still run on
independent machines without ordering unrelated jobs. The public
`TREE_MUTATING` map now classifies exclusive host-state fixtures; its sole
entry is `test/bootstrap-retention.test.ts`, whose same-UID nondumpable actors
affect host-wide procfs permission checks. Locally, this file runs only after
all parallel shards settle, and cancellation prevents that final phase from
starting. No selected files, retries, budgets, or receipt requirements are
removed. Former generator mutators still render into private output directories;
this serial phase protects process visibility, not in-place doc generation.
Full child output is retained in private files under `.context/free-test-logs/`,
outside each shard's temporary cleanup directory. The runner prints the path at
launch and completion. Losing the log fails the run even when the child exits
successfully. A redirected log directory is rejected before launching a child.
On failure, read that log first: the recovery message distinguishes incomplete
capture, unconfirmed cleanup, deadline expiry and a test/module failure. Fix the
demonstrated cause before rerunning. A focused `bun test` command is offered only
when every failure is attributable to existing selected files; it proves that
repair, not completion of the original selection. Preserve failed attempts when
sharing results, and inspect logs for private data before sharing them.
Before publication, classify new deterministic regressions for quick feedback.
Refresh the timing seed with the existing recorder on fixed inputs; do not edit
source while tests run. Critical boundary controls belong in `QUICK_CORE` when
their feedback cost is justified. Other measured files qualify at two seconds
or less; slow and unmeasured files remain outside quick, not outside full tests.
Report cold setup separately from warm execution, while retaining failed-attempt,
retry and cleanup time in the total cost.
**PTY fixture timing.** Plan-count sessions wake on terminal output or exit,
with at least 250ms between expensive observations and a 2s fallback for
@@ -230,6 +267,21 @@ key must name a living paid test (`test/touchfiles.test.ts`'s reverse
invariant), and `git show <sha>:path` fixtures are banned — vendor the bytes
instead (`test/git-ref-fixture-tripwire.test.ts`).
Functional QA and documentation acceptance require an explicit `EVALS_RUN_ID`;
`GSTACK_EVAL_DIR` alone does not satisfy their evidence-ownership guard. For each
local invocation, supply a fresh ID to the documented detached runner:
```bash
EVALS_RUN_ID="local-$(bun -e 'console.log(crypto.randomUUID())')" bun run eval:bg:pr
```
Neither package scripts nor `gstack-detach` invent this identity. CI's PR/manual
slices, periodic slices and weekly gate census supply an ID bound to the workflow
run, attempt, job and slice. Existing slice artifacts retain per-shard snapshots;
separate always-run `native-captures-<EVALS_RUN_ID>` artifacts retain project/legacy
`e2e-runs` and `evals/qa-callers` evidence for 90 days. These diagnostic artifacts
are not collector results and do not establish that an unfinished test passed.
**Fast PR profile and evidence reuse.** `test:pr` selects the changed cases in
`scripts/test-pr-profile.ts` plus every changed quality judge. `--profile full`
retains the broad census; no case IDs or tier assignments are removed. The plan
@@ -246,8 +298,9 @@ with matching before/after inputs. The audited workflow-judge adapter hashes the
actual expanded prompt, source/fixture/rubric/runner closure, installed SDK,
model parameters and runtime. Missing/unknown inputs force execution. Receipts
are scoped to the same repository and PR, expire after 24 hours, and contain
public scores and provenance rather than prompts or secrets. Only the 14 cases
using `runWorkflowJudge` are eligible; the other 11 quality cases remain fresh.
public scores and provenance rather than prompts or secrets. Of the 17 cases
using `runWorkflowJudge`, 16 are eligible; the cookie workflow's custom input
does not match the cache adapter and stays fresh, as do the other 11 quality cases.
CI supplies the scoped cache/runtime configuration; local runs are fresh by
default. Cached scores must
pass current assertions; reused records retain their original source and time
@@ -308,12 +361,24 @@ retry in a 52-minute shard wall, including two minutes for cleanup. No per-case
have a 1,830-second minimum shard wall and run without Bun retries; see the
[overlay contract](OVERLAY_BENCHMARK_CONTRACT.md) for their unchanged work budget.
The quality file reserves 6,400 seconds for all 25 cases and their existing
retry, plus cleanup. Each still has 120 seconds of model work. Its 14 workflow
The quality file reserves 7,180 seconds for all 28 cases and their existing
retry, plus cleanup. Each still has 120 seconds of model work. Its 17 workflow
judges own their deadline and abort signal, with five seconds for terminal
recording inside a ten-second Bun grace; the other 11 retain their existing
120-second Bun timeout. Late responses cannot create records or cache passes.
The ship documentation file reserves 10,920 seconds for five 600-second cases and
eight 300-second fault cases, each with one retry, plus cleanup. The standalone
documentation child retains its 600-second case. The five review/ship explorer
cases reserve 3,270 seconds including their existing retry and finalization grace.
These are whole-file supervision limits, not additional model work per case.
The shared-library path file reserves 3,720 seconds for its three serial
600-second cases, each with one retry, plus 120 seconds for cleanup. Its
registered budget keeps the file in its own shard and binds the expected wall
to both the saved plan and the execution receipt; missing or stale budget
records fail reconciliation. Case deadlines, model budgets and retries do not grow.
`resolvePaidShardBudget(files, overrideMs?)` is the canonical per-job resolver.
Each registered finding file and each overlay wrapper requires its
own shard, even with `--files-per-shard` above one. Mixed or multi-file overlay
@@ -324,9 +389,10 @@ Planner entries and execution results record the effective wall,
its source and policy identifier. Custom drivers must resolve each job instead
of passing their ordinary 1800-second default as an explicit cap;
their outer controller/detach wall must also cover the allocated work and cleanup.
`eval:bg:pr` and `eval:bg:periodic` have 72000/66000-second outer caps; the PR
The current paid census has 104 files: 46 gate-tier and 70 periodic-tier.
`eval:bg:pr` and `eval:bg:periodic` have 92820/67380-second outer caps; the PR
wrapper covers a full-gate fallback at its default two workers. The broad gate
wrapper reserves 33600 seconds, and release reserves 100000 seconds for both
wrapper reserves 49320 seconds, and release reserves 116700 seconds for both
tiers. Legacy monolithic
`eval:bg`/`eval:bg:all` retain their shorter 5400/7200-second caps and do not
promise every registered retry; use the sharded periodic path for this policy.
@@ -334,9 +400,12 @@ promise every registered retry; use the sharded periodic path for this policy.
Periodic CI plans `--slices 7`. When overlays are selected, the seventh is
reserved for their serial wrappers; registered finding files are distributed
across the remaining ordinary slices by their supervised walls. Each slice job
has a 358-minute cap. Reconciliation rejects missing, duplicated or misplaced
has a 360-minute cap. Reconciliation rejects missing, duplicated or misplaced
registered work and absent budget records. The weekly gate census has a
350-minute cap and PR slices have a 220-minute cap. Free supervision tests
352-minute cap across seven single-worker slices with at most four running at
once. Its longest current work wall is 272 minutes. PR slices retain seven
two-worker slices with a 265-minute cap for their 212-minute work wall plus
setup. Free supervision tests
verify these bounds against the complete current census, configured retries,
and setup reserve. Ordinary paid tiers and the default 1800-second
shard wall remain unchanged; the registered and overlay policies above supply