Merge origin/main (v1.91.7.0) into test-audit-reduction

Keep both intents: v1.91.7.0's functional QA, docsync and exploratory
paid cases and their free owners stay; this branch's deletions stay
deleted. main's new paid keys follow the derived-closure touchfile rule
(free *.test.ts paths dropped, static helper/fixture closure added), its
new helper-only tests join the ratchet baseline, and its free selection
examples that named free test files now assert the derived selection.

Periodic CI keeps seven slices without the retired Autoplan slice; the
gate census keeps seven single-worker slices with --skip-judges. Wall
and census literals are recomputed from the merged planner, durations
are re-recorded on Ubicloud, and VERSION stays 1.91.8.0 above 1.91.7.0.
This commit is contained in:
garrytan committed 2026-09-29 13:48:02 +00:00
commit b421bba2c9
325 files changed
+42569 -8257

No files matched your search

+83
View File
@@ -48,6 +48,47 @@ never a new per-incident file; `test/test-of-test-ratchet.test.ts` enforces this
| Plan scope selection (`plan-scope-selection`) | `test/plan-scope-selection.test.ts` |
| `planCountPrerequisitePick` | `test/plan-count-prerequisite.test.ts` |
## Functional QA contract map
The deterministic owners below protect the failure boundary; their live partners
prove that an agent follows it. A shared fixture or captured event does not replace
an independent live trial. All free owners run in `bun run test`; quick eligibility
depends on measured duration or an explicit `QUICK_CORE` entry, not this table.
| Contract | Deterministic owner | Necessary live boundary | Host and lane |
| --- | --- | --- | --- |
| CLI/API/webhook QA without browser setup | `qa-functional-fixture`, `qa-functional-evidence`, `qa-lazy-sections` | `skill-e2e-qa-functional`: CLI and webhook report sessions | Linux/macOS free; selected PR gate; Windows only where curated |
| Report-only preserves local and remote authority | `qa-only-capability`, `qa-functional-observer`, `qa-functional-observer-atomic`, `qa-caller-authority` | Independent report-only sessions with synthetic owned endpoints/auth | Linux kernel observation; free callback controls plus selected PR gate |
| Repair reproduces the defect, adds a failing regression and rechecks adjacent behavior | `qa-fix-loop-fixture`, `qa-functional-evidence` | `skill-e2e-qa-functional-fix`: CLI and webhook repair sessions | Free controls plus selected PR gate |
| Review and Ship actually explore | `qa-exploratory-callers`, `qa-caller-report-observer`, `qa-checkpoint-evidence` | `skill-e2e-qa-callers`: actual Review/Ship callers | Free captures plus selected PR gate |
| Smoke expiry preserves required plan checks | `qa-deadline`, `qa-deadline-selection`, `qa-browser-deadline-evidence` | `ship-exploratory-plan-checks` | Free deadline/dispatch controls plus selected PR gate |
| Late changes invalidate affected results | `qa-caller-freshness-order`, `qa-deadline-publication-observer`, `shared-libs-revalidation-prompt` | `ship-exploratory-late-input` and the existing late-input documentation handoff | Free stale-input controls plus selected PR gate |
| Documentation completes before publication and respects protected files | `docsync-authority`, `docsync-atomic-writes`, `docsync-report-interface`, `docsync-lifecycle-interface` | `skill-e2e-ship-docsync`, `skill-e2e-docsync-spawned` | Free state/permission controls; registered gate/periodic scenarios retain their tiers |
| Cancellation drains owned work before another attempt | `shared-libs-cancellation`, `session-runner-stream-lifecycle`, `agent-sdk-runner`, `paid-shard-settlement` | Existing actual shared-library/SDK caller scenarios | Free real-callback/process controls; registered live gate/periodic trials remain independent |
| Missing tools or incomplete results never become verified coverage | `qa-probe-gates`, `qa-supervision-selection`, `test-free-shards`, `test-free-shards-capture`, `paid-shards` | `ship-exploratory-unavailable` and existing reporting-boundary sessions | Free negative controls plus selected PR gate; unsupported hosts remain unexecuted |
Names without a suffix refer to `test/<name>.test.ts`. Keep missing, stale,
duplicate, selected-but-unstarted, malformed/truncated and observer-overflow
controls distinct from legitimate empty selections. File restoration cannot
replace write observation, and a clean local tree cannot prove that an external
request made no mutation. Fixture endpoints and credentials must be synthetic
and owned; specifically authorized functional requests remain permitted.
Functional fixtures register their existing closed command policy as a native
PreToolUse hook, so an unsupported request is refused before execution. The
callback regression invokes the registered command with native hook input,
observes an isolated mutation target and permits the owned webhook positive
control. This is a command boundary, not a sandbox for arbitrary target code.
Its private CLI configuration is outside the observed product tree, and the
fixture's existing cleanup owns both directories.
Review/Ship observations now use the same strict native event decoder as QA
checkpoints and documentation. Caller-specific handoff/freshness interpretation
stays separate. Original missing, orphaned and duplicate-call controls were run
before replacing three incidental error-wording assertions with rejection checks;
the existing positive attribution case still runs, and a completed-ID reuse
negative control prevents incomplete evidence from becoming green.
## Complete inventory, not just the fast subset
At the audited revision, all 1,124 tracked Bun test files partition into 1,010 free
@@ -125,6 +166,41 @@ case/sample inventory and report skips and unavailable platforms separately.
Do not subtract failures from elapsed time or use a smaller selection as proof
that the complete suite got faster.
### Functional-QA cleanup measurement — September 28, 2026
On the same four-CPU Linux machine, using Bun 1.4.0, Node 22.20.0 and Claude
Code 2.1.251, the existing duration recorder measured all 1,113 free files.
The refreshed seed selects 931 files for quick feedback: 90 newly included and
20 newly excluded by measured cost, a net increase of 70. No files remain
unclassified. All 182 slow files remain in the complete suite. The functional
command observer, checkpoint decoder and log-capture controls are explicit
quick-core cases; each measured under two seconds.
| Existing command / attempt | Executed scope | Result | Wall time |
| --- | --- | --- | ---: |
| `bun run test:free --record-durations` | 1,113 files | 29,175 pass, 5 fail, 131 skip | 680.64s |
| `bun run test:quick`, first measured attempt | 931 files | 22,160 pass, 2 fail, 100 skip | 125.39s |
| `bun run test:quick`, repaired attempt | The same 931 files | 22,162 pass, 0 fail, 100 skip | 52.43s |
The profile's five failures came from the machine's Git identity wrapper
overwriting synthetic fixture authors. Running the two affected files with native
Git in the isolated test environment passed all 59 tests in 75.32s; normal checkout
commits retained the configured identity. Both quick attempts used that corrected
environment. Their two telemetry timeouts used Bun's synchronous piped-input
path; the repair reuses the existing file-backed command capture helper without
changing commands, assertions or deadlines. The seed retains observed costs,
including failed attempts; it is a scheduling hint, not a passing receipt.
Cold dependency installation took 0.477s and the integrated build took 3.84s,
separate from warm test execution; CLI installation was not independently timed.
An earlier 63.37s profile was cancelled for a decoder repair, with an additional
scoped browser cleanup, and earns no completion credit. Failed, cancelled and
repair runs are costs, not time removed from the workflow. The quick target of
one minute was met on this machine, but these measurements establish neither a
cross-environment speedup nor full release, live-model or Windows acceptance.
### Earlier component comparisons
Measured component comparisons:
| Workload | Before | After | Coverage retained |
@@ -190,6 +266,13 @@ not a fresh full-census runtime improvement.
## Evidence validity
After integrating main's September 28 Ubicloud improvements, the scheduling seed
uses upstream's CI-environment timings for shared files and preserves the 52
previously measured branch-only entries. These are scheduling hints from two
machines, not a matched performance comparison or acceptance result. Refresh
the whole seed with `bun run test:ubicloud --record-durations` when measuring a
new common baseline; do not infer a speedup by adding these measurements.
Check the executable actually used by each SDK, print-mode and terminal launcher.
A CLI version cached during preflight does not prove the version used by later
sessions if PATH contents change. Use native session-init versions, terminal