v1.91.7.0 feat: add functional QA and pre-publication docs checks (#2983)

* feat: add surface-aware exploratory QA and ship documentation gates

* test: preserve delegated QA setup authority after main integration

* fix(qa): clarify exploration order and preserve report artifacts

* test(qa): follow the shared setup reference directly

* refactor(ship): make verification and recovery routes explicit

* test(ship): align evidence and review guards with explicit routes

* fix(workflows): clarify ship recovery and functional QA evidence

* fix(workflows): clarify approval recovery and full QA coverage

* refactor(workflows): order review transactions and clarify ship state

* fix(ship): clarify final verification and fail closed at publication

* fix(evals): attribute native atomic documentation writes

* fix(ship): clarify recovery and documentation lifecycle guidance

* fix(test): preserve observed native placeholder styling in CI

* fix(codex): report watchdog timeouts without a process-exit race

* Checkpoint functional QA implementation and workflow validation repairs

* Fix documentation and shared-review fixture contracts

* docs: clarify judge reuse and evaluation supervision

* test: align review evidence and selected case contracts

* test: verify append-only documentation checkpoints and recovery

* fix: qualify QA workflows and CI validation repairs

* fix: launch shared-libs fixture scripts on Windows

* fix: qualify QA deadlines, fixture isolation, and shard cleanup

* fix: preserve qualified QA and cancellation repairs

* fix: enforce functional fixture authority and share strict event decoding

* fix: retain free-test evidence and explain recovery

* fix: reject malformed native evidence after decoder consolidation

* test: use reliable capture for telemetry privacy filters

* test: refresh measured quick coverage and document validation costs

* Fix native fixture receipts and preserve VM validation evidence

* Align negative judge controls with upstream clarity policy

* Fix report-only QA preparation and public evidence handling

* Clarify QA-only preparation and current-report preservation

* Stream Ship quality judgments with an explicit 64k response contract

* Validate compact judge reasoning locally with supported wire schema

* Align functional QA fixture instructions with evidence acceptance

* Bind native browser diagnostics to execution evidence and align review verdicts

* Preserve native diagnostic line boundaries

* Serialize functional QA evidence from native captures

* Keep large QA evidence fixture payload out of Windows argv
This commit is contained in:
Garry Tan authored and GitHub committed 2026-09-29 06:07:35 -07:00
1 parent 65bfb0ce49
commit dcaea52800
333 files changed
+41755 -7357

No files matched your search

+26 -4
View File
@@ -1,14 +1,32 @@
## Step 7: Test Coverage Audit
**Dispatch this step as a subagent** using the Agent tool with `subagent_type: "general-purpose"`. The fresh-context subagent runs the audit; the parent only needs the conclusion.
### Shared subagent dispatch
{{FOREGROUND_DISPATCH_NOTE}} The parent needs this audit's LAST-line JSON before continuing.
For Steps 7, 8 and 10, use the Agent tool with `run_in_background: false`.
Omitting the flag runs the subagent in the background. The explicit flag waits
for a result while keeping a fresh context. Do not invoke the target as a Skill
or run it inline instead. Inline work is allowed only under that section's
documented fallback, after a failed subagent has stopped.
**Subagent prompt:** Pass the following instructions to the subagent, with `<base>` substituted with the base branch:
Dispatch the audit through Agent with `subagent_type: "general-purpose"` and
`run_in_background: false`, using the shared foreground-dispatch rule above.
Wait for its LAST-line JSON before applying the coverage gate.
**Generation allowance:** Maximum 2 generation passes total per invocation.
Count each generation-authorized attempt before dispatch/inline execution, including
the initial audit, failures and zero-test results. Re-entry never resets it.
Two passes already used means no further generation; read-only reassessment uses no pass.
**Subagent prompt:** Supply `<base>`, Step 4's framework/bootstrap decision,
permitted paths/commands, remaining gaps, passes used and generation allowance.
No allowance means audit only; missing permission is not approval. Preserve the
30-path/20-test/2-minute per-test caps.
````text
You are running a ship-workflow test coverage audit. Run `git diff origin/<base>` to include uncommitted tracked changes; also read relevant non-ignored untracked source/tests. Do not commit or push. Perform only this audit; return unresolved user decisions to the parent instead of asking or advancing to another workflow step.
Generation: <allowed|audit-only>; passes used: <N> of 2. Audit-only overrides every generation instruction below.
{{TEST_COVERAGE_AUDIT_SHIP}}
After your analysis, output a single JSON object on the LAST LINE of your response (no other text after it):
@@ -23,7 +41,11 @@ Use null for an undetermined or skipped coverage percentage, not zero. Include e
3. Embed `diagram` verbatim in the PR body's `## Test Coverage` section (Step 19).
4. Print a one-line summary: `Coverage: {coverage_pct}%, {gaps} gaps. {tests_added.length} tests added.`
**If the subagent fails, times out, returns invalid JSON, or never completes after ~10 minutes:** stop any live backgrounded task, then run the audit inline in the parent. Do not block /ship on subagent failure — partial results are better than none.
**Audit failure:** On failure, invalid JSON or no completion after ~10 minutes,
stop the child and confirm it stopped before running the same audit inline.
Fallback recovers the audit; it does not pass or bypass the coverage gate.
Apply that gate to the recovered results, including its undetermined-percentage
and test-only rules. Preserve partial results as incomplete, not passing coverage.
{{TEST_COVERAGE_GATE_SHIP}}