mirror of
https://github.com/garrytan/gstack.git
synced 2026-10-05 02:47:17 +02:00
v1.91.7.0 feat: add functional QA and pre-publication docs checks (#2983)
* feat: add surface-aware exploratory QA and ship documentation gates * test: preserve delegated QA setup authority after main integration * fix(qa): clarify exploration order and preserve report artifacts * test(qa): follow the shared setup reference directly * refactor(ship): make verification and recovery routes explicit * test(ship): align evidence and review guards with explicit routes * fix(workflows): clarify ship recovery and functional QA evidence * fix(workflows): clarify approval recovery and full QA coverage * refactor(workflows): order review transactions and clarify ship state * fix(ship): clarify final verification and fail closed at publication * fix(evals): attribute native atomic documentation writes * fix(ship): clarify recovery and documentation lifecycle guidance * fix(test): preserve observed native placeholder styling in CI * fix(codex): report watchdog timeouts without a process-exit race * Checkpoint functional QA implementation and workflow validation repairs * Fix documentation and shared-review fixture contracts * docs: clarify judge reuse and evaluation supervision * test: align review evidence and selected case contracts * test: verify append-only documentation checkpoints and recovery * fix: qualify QA workflows and CI validation repairs * fix: launch shared-libs fixture scripts on Windows * fix: qualify QA deadlines, fixture isolation, and shard cleanup * fix: preserve qualified QA and cancellation repairs * fix: enforce functional fixture authority and share strict event decoding * fix: retain free-test evidence and explain recovery * fix: reject malformed native evidence after decoder consolidation * test: use reliable capture for telemetry privacy filters * test: refresh measured quick coverage and document validation costs * Fix native fixture receipts and preserve VM validation evidence * Align negative judge controls with upstream clarity policy * Fix report-only QA preparation and public evidence handling * Clarify QA-only preparation and current-report preservation * Stream Ship quality judgments with an explicit 64k response contract * Validate compact judge reasoning locally with supported wire schema * Align functional QA fixture instructions with evidence acceptance * Bind native browser diagnostics to execution evidence and align review verdicts * Preserve native diagnostic line boundaries * Serialize functional QA evidence from native captures * Keep large QA evidence fixture payload out of Windows argv
This commit is contained in:
1 parent
65bfb0ce49
commit
dcaea52800
333 files changed
+41755
-7357
No files matched your search
+86
-17
@@ -127,6 +127,21 @@ on a Mac they drive Aside and on Linux CI they drive the built browse binary,
|
||||
skipping only when neither exists. The `$B`-driven E2E cases and `browse/test/`
|
||||
run on every platform as before, so Linux CI proves the fallback engine live.
|
||||
|
||||
**Bootstrap dependency retention is opt-in qualification, not the behavior test.**
|
||||
`qa-bootstrap` still runs its original unpinned Vitest installation and assertions
|
||||
on macOS and Linux, including documented unsharded commands. Only a Linux paid
|
||||
shard runner issues the owned retention scope: it binds each actual fixture and
|
||||
native lifetime, retains locks, package manifests and the installed file/link
|
||||
inventory, and acknowledges capture before deleting the fixture. The outer
|
||||
runner also captures evidence when a callback is killed. Incomplete capture
|
||||
fails qualification and preserves the source fixture as well as partial evidence.
|
||||
Other platforms explicitly report retention as unavailable and still execute the
|
||||
native behavior test. A run without the runner-issued scope earns no retained
|
||||
dependency qualification credit; candidate acceptance requiring that evidence
|
||||
must use the Linux sharded path and verify every attempt's complete capture,
|
||||
acknowledgment and cleanup fallback. A passing unsharded or macOS behavior test
|
||||
does not substitute for that evidence.
|
||||
|
||||
**The renderer picks the same way, so the render gates are engine-agnostic.**
|
||||
`/make-pdf`, `/diagram`, and design previews print and screenshot their local
|
||||
HTML through `lib/aside-render.ts` / `bin/gstack-render.ts`, which render in
|
||||
@@ -177,13 +192,35 @@ fallback; unknown files get 75th-percentile pessimism, and both full-suite and
|
||||
the long pole. Packed
|
||||
shards get duration-aware walls (`max(base, predicted × 3, files × 5s)`). The
|
||||
legacy `--shards N --shard i` path keeps stable hash indices. Required CI uses
|
||||
one duration-packed `--ci-plan`, 20 isolated `--ci-run` machines, and a
|
||||
`--ci-verify` aggregate. `TREE_MUTATING` is EMPTY:
|
||||
`gen-skill-docs.ts` has a `main()` guard (imports never regenerate; pinned by
|
||||
`test/gen-skill-docs-import-purity.test.ts`) and `--out-dir` renders every
|
||||
host, so all former mutators render into mkdtemps and the trailing serial
|
||||
shard is gone. The map remains a mechanism — a test that genuinely must write
|
||||
shared artifacts in place earns a reasoned entry and is serialized again.
|
||||
one duration-packed `--ci-plan`, 20 ordinary `--ci-run` shards plus a separate
|
||||
exclusive-fixture shard, and a `--ci-verify` aggregate. CI shards still run on
|
||||
independent machines without ordering unrelated jobs. The public
|
||||
`TREE_MUTATING` map now classifies exclusive host-state fixtures; its sole
|
||||
entry is `test/bootstrap-retention.test.ts`, whose same-UID nondumpable actors
|
||||
affect host-wide procfs permission checks. Locally, this file runs only after
|
||||
all parallel shards settle, and cancellation prevents that final phase from
|
||||
starting. No selected files, retries, budgets, or receipt requirements are
|
||||
removed. Former generator mutators still render into private output directories;
|
||||
this serial phase protects process visibility, not in-place doc generation.
|
||||
|
||||
Full child output is retained in private files under `.context/free-test-logs/`,
|
||||
outside each shard's temporary cleanup directory. The runner prints the path at
|
||||
launch and completion. Losing the log fails the run even when the child exits
|
||||
successfully. A redirected log directory is rejected before launching a child.
|
||||
On failure, read that log first: the recovery message distinguishes incomplete
|
||||
capture, unconfirmed cleanup, deadline expiry and a test/module failure. Fix the
|
||||
demonstrated cause before rerunning. A focused `bun test` command is offered only
|
||||
when every failure is attributable to existing selected files; it proves that
|
||||
repair, not completion of the original selection. Preserve failed attempts when
|
||||
sharing results, and inspect logs for private data before sharing them.
|
||||
|
||||
Before publication, classify new deterministic regressions for quick feedback.
|
||||
Refresh the timing seed with the existing recorder on fixed inputs; do not edit
|
||||
source while tests run. Critical boundary controls belong in `QUICK_CORE` when
|
||||
their feedback cost is justified. Other measured files qualify at two seconds
|
||||
or less; slow and unmeasured files remain outside quick, not outside full tests.
|
||||
Report cold setup separately from warm execution, while retaining failed-attempt,
|
||||
retry and cleanup time in the total cost.
|
||||
|
||||
**PTY fixture timing.** Plan-count sessions wake on terminal output or exit,
|
||||
with at least 250ms between expensive observations and a 2s fallback for
|
||||
@@ -230,6 +267,21 @@ key must name a living paid test (`test/touchfiles.test.ts`'s reverse
|
||||
invariant), and `git show <sha>:path` fixtures are banned — vendor the bytes
|
||||
instead (`test/git-ref-fixture-tripwire.test.ts`).
|
||||
|
||||
Functional QA and documentation acceptance require an explicit `EVALS_RUN_ID`;
|
||||
`GSTACK_EVAL_DIR` alone does not satisfy their evidence-ownership guard. For each
|
||||
local invocation, supply a fresh ID to the documented detached runner:
|
||||
|
||||
```bash
|
||||
EVALS_RUN_ID="local-$(bun -e 'console.log(crypto.randomUUID())')" bun run eval:bg:pr
|
||||
```
|
||||
|
||||
Neither package scripts nor `gstack-detach` invent this identity. CI's PR/manual
|
||||
slices, periodic slices and weekly gate census supply an ID bound to the workflow
|
||||
run, attempt, job and slice. Existing slice artifacts retain per-shard snapshots;
|
||||
separate always-run `native-captures-<EVALS_RUN_ID>` artifacts retain project/legacy
|
||||
`e2e-runs` and `evals/qa-callers` evidence for 90 days. These diagnostic artifacts
|
||||
are not collector results and do not establish that an unfinished test passed.
|
||||
|
||||
**Fast PR profile and evidence reuse.** `test:pr` selects the changed cases in
|
||||
`scripts/test-pr-profile.ts` plus every changed quality judge. `--profile full`
|
||||
retains the broad census; no case IDs or tier assignments are removed. The plan
|
||||
@@ -246,8 +298,9 @@ with matching before/after inputs. The audited workflow-judge adapter hashes the
|
||||
actual expanded prompt, source/fixture/rubric/runner closure, installed SDK,
|
||||
model parameters and runtime. Missing/unknown inputs force execution. Receipts
|
||||
are scoped to the same repository and PR, expire after 24 hours, and contain
|
||||
public scores and provenance rather than prompts or secrets. Only the 14 cases
|
||||
using `runWorkflowJudge` are eligible; the other 11 quality cases remain fresh.
|
||||
public scores and provenance rather than prompts or secrets. Of the 17 cases
|
||||
using `runWorkflowJudge`, 16 are eligible; the cookie workflow's custom input
|
||||
does not match the cache adapter and stays fresh, as do the other 11 quality cases.
|
||||
CI supplies the scoped cache/runtime configuration; local runs are fresh by
|
||||
default. Cached scores must
|
||||
pass current assertions; reused records retain their original source and time
|
||||
@@ -325,12 +378,24 @@ including two minutes for cleanup. No per-case budget grows. Overlay wrappers
|
||||
have a 1,830-second minimum shard wall and run without Bun retries; see the
|
||||
[overlay contract](OVERLAY_BENCHMARK_CONTRACT.md) for their unchanged work budget.
|
||||
|
||||
The quality file reserves 6,400 seconds for all 25 cases and their existing
|
||||
retry, plus cleanup. Each still has 120 seconds of model work. Its 14 workflow
|
||||
The quality file reserves 7,180 seconds for all 28 cases and their existing
|
||||
retry, plus cleanup. Each still has 120 seconds of model work. Its 17 workflow
|
||||
judges own their deadline and abort signal, with five seconds for terminal
|
||||
recording inside a ten-second Bun grace; the other 11 retain their existing
|
||||
120-second Bun timeout. Late responses cannot create records or cache passes.
|
||||
|
||||
The ship documentation file reserves 10,920 seconds for five 600-second cases and
|
||||
eight 300-second fault cases, each with one retry, plus cleanup. The standalone
|
||||
documentation child retains its 600-second case. The five review/ship explorer
|
||||
cases reserve 3,270 seconds including their existing retry and finalization grace.
|
||||
These are whole-file supervision limits, not additional model work per case.
|
||||
|
||||
The shared-library path file reserves 3,720 seconds for its three serial
|
||||
600-second cases, each with one retry, plus 120 seconds for cleanup. Its
|
||||
registered budget keeps the file in its own shard and binds the expected wall
|
||||
to both the saved plan and the execution receipt; missing or stale budget
|
||||
records fail reconciliation. Case deadlines, model budgets and retries do not grow.
|
||||
|
||||
`resolvePaidShardBudget(files, overrideMs?)` is the canonical per-job resolver.
|
||||
Autoplan, each registered finding file, and each overlay wrapper require their
|
||||
own shard, even with `--files-per-shard` above one. Mixed or multi-file overlay
|
||||
@@ -341,20 +406,24 @@ Planner entries and execution results record the effective wall,
|
||||
its source and policy identifier. Custom drivers must resolve each job instead
|
||||
of passing their ordinary 1800-second default as an explicit Autoplan cap;
|
||||
their outer controller/detach wall must also cover the allocated work and cleanup.
|
||||
`eval:bg:pr` and `eval:bg:periodic` have 72000/66000-second outer caps; the PR
|
||||
The current paid census has 122 files: 61 gate-tier and 103 periodic-tier.
|
||||
`eval:bg:pr` and `eval:bg:periodic` have 92820/67380-second outer caps; the PR
|
||||
wrapper covers a full-gate fallback at its default two workers. The broad gate
|
||||
wrapper reserves 33600 seconds, and release reserves 100000 seconds for both
|
||||
wrapper reserves 49320 seconds, and release reserves 116700 seconds for both
|
||||
tiers. Legacy monolithic
|
||||
`eval:bg`/`eval:bg:all` retain their shorter 5400/7200-second caps and do not
|
||||
promise two complete Autoplan attempts; use the sharded periodic path for this policy.
|
||||
|
||||
Periodic CI plans `--slices 8 --autoplan-slice`: the eighth runs only Autoplan.
|
||||
When overlays are selected, the seventh is reserved for their serial wrappers;
|
||||
Periodic CI plans `--slices 9 --autoplan-slice`: the ninth runs only Autoplan.
|
||||
When overlays are selected, the eighth is reserved for their serial wrappers;
|
||||
registered finding files are distributed across the remaining ordinary slices
|
||||
by their supervised walls. Each slice job has a 355-minute cap; Autoplan retains
|
||||
by their supervised walls. Each slice job has a 360-minute cap; Autoplan retains
|
||||
its 172-minute shard wall. Reconciliation rejects missing, duplicated or misplaced
|
||||
registered work and absent budget records. The weekly gate census has a
|
||||
350-minute cap and PR slices have a 220-minute cap. Free supervision tests
|
||||
352-minute cap across eight single-worker slices with at most four running at
|
||||
once. Its longest current work wall is 302 minutes. PR slices retain seven
|
||||
two-worker slices with a 265-minute cap for their 242-minute work wall plus
|
||||
setup. Free supervision tests
|
||||
verify these bounds against the complete current census, configured retries,
|
||||
and setup reserve. Ordinary paid tiers and the default 1800-second
|
||||
shard wall remain unchanged; the registered and overlay policies above supply
|
||||
|
||||
Reference in new issue
Block a user