mirror of
https://github.com/garrytan/gstack.git
synced 2026-10-02 17:40:02 +02:00
v1.91.7.0 feat: add functional QA and pre-publication docs checks (#2983)
* feat: add surface-aware exploratory QA and ship documentation gates * test: preserve delegated QA setup authority after main integration * fix(qa): clarify exploration order and preserve report artifacts * test(qa): follow the shared setup reference directly * refactor(ship): make verification and recovery routes explicit * test(ship): align evidence and review guards with explicit routes * fix(workflows): clarify ship recovery and functional QA evidence * fix(workflows): clarify approval recovery and full QA coverage * refactor(workflows): order review transactions and clarify ship state * fix(ship): clarify final verification and fail closed at publication * fix(evals): attribute native atomic documentation writes * fix(ship): clarify recovery and documentation lifecycle guidance * fix(test): preserve observed native placeholder styling in CI * fix(codex): report watchdog timeouts without a process-exit race * Checkpoint functional QA implementation and workflow validation repairs * Fix documentation and shared-review fixture contracts * docs: clarify judge reuse and evaluation supervision * test: align review evidence and selected case contracts * test: verify append-only documentation checkpoints and recovery * fix: qualify QA workflows and CI validation repairs * fix: launch shared-libs fixture scripts on Windows * fix: qualify QA deadlines, fixture isolation, and shard cleanup * fix: preserve qualified QA and cancellation repairs * fix: enforce functional fixture authority and share strict event decoding * fix: retain free-test evidence and explain recovery * fix: reject malformed native evidence after decoder consolidation * test: use reliable capture for telemetry privacy filters * test: refresh measured quick coverage and document validation costs * Fix native fixture receipts and preserve VM validation evidence * Align negative judge controls with upstream clarity policy * Fix report-only QA preparation and public evidence handling * Clarify QA-only preparation and current-report preservation * Stream Ship quality judgments with an explicit 64k response contract * Validate compact judge reasoning locally with supported wire schema * Align functional QA fixture instructions with evidence acceptance * Bind native browser diagnostics to execution evidence and align review verdicts * Preserve native diagnostic line boundaries * Serialize functional QA evidence from native captures * Keep large QA evidence fixture payload out of Windows argv
This commit is contained in:
1 parent
65bfb0ce49
commit
dcaea52800
333 files changed
+41755
-7357
No files matched your search
@@ -27,6 +27,47 @@ Overlay efficacy experiments retain their full fixture/model/arm/trial matrix.
|
||||
Security cases retain their source, path, socket, process and lease identities.
|
||||
These are distinct scenario dimensions, not repeated work to delete.
|
||||
|
||||
## Functional QA contract map
|
||||
|
||||
The deterministic owners below protect the failure boundary; their live partners
|
||||
prove that an agent follows it. A shared fixture or captured event does not replace
|
||||
an independent live trial. All free owners run in `bun run test`; quick eligibility
|
||||
depends on measured duration or an explicit `QUICK_CORE` entry, not this table.
|
||||
|
||||
| Contract | Deterministic owner | Necessary live boundary | Host and lane |
|
||||
| --- | --- | --- | --- |
|
||||
| CLI/API/webhook QA without browser setup | `qa-functional-fixture`, `qa-functional-evidence`, `qa-lazy-sections` | `skill-e2e-qa-functional`: CLI and webhook report sessions | Linux/macOS free; selected PR gate; Windows only where curated |
|
||||
| Report-only preserves local and remote authority | `qa-only-capability`, `qa-functional-observer`, `qa-functional-observer-atomic`, `qa-caller-authority` | Independent report-only sessions with synthetic owned endpoints/auth | Linux kernel observation; free callback controls plus selected PR gate |
|
||||
| Repair reproduces the defect, adds a failing regression and rechecks adjacent behavior | `qa-fix-loop-fixture`, `qa-functional-evidence` | `skill-e2e-qa-functional-fix`: CLI and webhook repair sessions | Free controls plus selected PR gate |
|
||||
| Review and Ship actually explore | `qa-exploratory-callers`, `qa-caller-report-observer`, `qa-checkpoint-evidence` | `skill-e2e-qa-callers`: actual Review/Ship callers | Free captures plus selected PR gate |
|
||||
| Smoke expiry preserves required plan checks | `qa-deadline`, `qa-deadline-selection`, `qa-browser-deadline-evidence` | `ship-exploratory-plan-checks` | Free deadline/dispatch controls plus selected PR gate |
|
||||
| Late changes invalidate affected results | `qa-caller-freshness-order`, `qa-deadline-publication-observer`, `shared-libs-revalidation-prompt` | `ship-exploratory-late-input` and the existing late-input documentation handoff | Free stale-input controls plus selected PR gate |
|
||||
| Documentation completes before publication and respects protected files | `docsync-authority`, `docsync-atomic-writes`, `docsync-report-interface`, `docsync-lifecycle-interface` | `skill-e2e-ship-docsync`, `skill-e2e-docsync-spawned` | Free state/permission controls; registered gate/periodic scenarios retain their tiers |
|
||||
| Cancellation drains owned work before another attempt | `shared-libs-cancellation`, `session-runner-stream-lifecycle`, `agent-sdk-runner`, `paid-shard-settlement` | Existing actual shared-library/SDK caller scenarios | Free real-callback/process controls; registered live gate/periodic trials remain independent |
|
||||
| Missing tools or incomplete results never become verified coverage | `qa-probe-gates`, `qa-supervision-selection`, `test-free-shards`, `test-free-shards-capture`, `paid-shards` | `ship-exploratory-unavailable` and existing reporting-boundary sessions | Free negative controls plus selected PR gate; unsupported hosts remain unexecuted |
|
||||
|
||||
Names without a suffix refer to `test/<name>.test.ts`. Keep missing, stale,
|
||||
duplicate, selected-but-unstarted, malformed/truncated and observer-overflow
|
||||
controls distinct from legitimate empty selections. File restoration cannot
|
||||
replace write observation, and a clean local tree cannot prove that an external
|
||||
request made no mutation. Fixture endpoints and credentials must be synthetic
|
||||
and owned; specifically authorized functional requests remain permitted.
|
||||
|
||||
Functional fixtures register their existing closed command policy as a native
|
||||
PreToolUse hook, so an unsupported request is refused before execution. The
|
||||
callback regression invokes the registered command with native hook input,
|
||||
observes an isolated mutation target and permits the owned webhook positive
|
||||
control. This is a command boundary, not a sandbox for arbitrary target code.
|
||||
Its private CLI configuration is outside the observed product tree, and the
|
||||
fixture's existing cleanup owns both directories.
|
||||
|
||||
Review/Ship observations now use the same strict native event decoder as QA
|
||||
checkpoints and documentation. Caller-specific handoff/freshness interpretation
|
||||
stays separate. Original missing, orphaned and duplicate-call controls were run
|
||||
before replacing three incidental error-wording assertions with rejection checks;
|
||||
the existing positive attribution case still runs, and a completed-ID reuse
|
||||
negative control prevents incomplete evidence from becoming green.
|
||||
|
||||
## Complete inventory, not just the fast subset
|
||||
|
||||
At the audited revision, all 1,124 tracked Bun test files partition into 1,010 free
|
||||
@@ -104,6 +145,41 @@ case/sample inventory and report skips and unavailable platforms separately.
|
||||
Do not subtract failures from elapsed time or use a smaller selection as proof
|
||||
that the complete suite got faster.
|
||||
|
||||
### Functional-QA cleanup measurement — September 28, 2026
|
||||
|
||||
On the same four-CPU Linux machine, using Bun 1.4.0, Node 22.20.0 and Claude
|
||||
Code 2.1.251, the existing duration recorder measured all 1,113 free files.
|
||||
The refreshed seed selects 931 files for quick feedback: 90 newly included and
|
||||
20 newly excluded by measured cost, a net increase of 70. No files remain
|
||||
unclassified. All 182 slow files remain in the complete suite. The functional
|
||||
command observer, checkpoint decoder and log-capture controls are explicit
|
||||
quick-core cases; each measured under two seconds.
|
||||
|
||||
| Existing command / attempt | Executed scope | Result | Wall time |
|
||||
| --- | --- | --- | ---: |
|
||||
| `bun run test:free --record-durations` | 1,113 files | 29,175 pass, 5 fail, 131 skip | 680.64s |
|
||||
| `bun run test:quick`, first measured attempt | 931 files | 22,160 pass, 2 fail, 100 skip | 125.39s |
|
||||
| `bun run test:quick`, repaired attempt | The same 931 files | 22,162 pass, 0 fail, 100 skip | 52.43s |
|
||||
|
||||
The profile's five failures came from the machine's Git identity wrapper
|
||||
overwriting synthetic fixture authors. Running the two affected files with native
|
||||
Git in the isolated test environment passed all 59 tests in 75.32s; normal checkout
|
||||
commits retained the configured identity. Both quick attempts used that corrected
|
||||
environment. Their two telemetry timeouts used Bun's synchronous piped-input
|
||||
path; the repair reuses the existing file-backed command capture helper without
|
||||
changing commands, assertions or deadlines. The seed retains observed costs,
|
||||
including failed attempts; it is a scheduling hint, not a passing receipt.
|
||||
|
||||
Cold dependency installation took 0.477s and the integrated build took 3.84s,
|
||||
separate from warm test execution; CLI installation was not independently timed.
|
||||
An earlier 63.37s profile was cancelled for a decoder repair, with an additional
|
||||
scoped browser cleanup, and earns no completion credit. Failed, cancelled and
|
||||
repair runs are costs, not time removed from the workflow. The quick target of
|
||||
one minute was met on this machine, but these measurements establish neither a
|
||||
cross-environment speedup nor full release, live-model or Windows acceptance.
|
||||
|
||||
### Earlier component comparisons
|
||||
|
||||
Measured component comparisons:
|
||||
|
||||
| Workload | Before | After | Coverage retained |
|
||||
@@ -167,6 +243,13 @@ not a fresh full-census runtime improvement.
|
||||
|
||||
## Evidence validity
|
||||
|
||||
After integrating main's September 28 Ubicloud improvements, the scheduling seed
|
||||
uses upstream's CI-environment timings for shared files and preserves the 52
|
||||
previously measured branch-only entries. These are scheduling hints from two
|
||||
machines, not a matched performance comparison or acceptance result. Refresh
|
||||
the whole seed with `bun run test:ubicloud --record-durations` when measuring a
|
||||
new common baseline; do not infer a speedup by adding these measurements.
|
||||
|
||||
Check the executable actually used by each SDK, print-mode and terminal launcher.
|
||||
A CLI version cached during preflight does not prove the version used by later
|
||||
sessions if PATH contents change. Use native session-init versions, terminal
|
||||
|
||||
Reference in new issue
Block a user