* feat: add surface-aware exploratory QA and ship documentation gates * test: preserve delegated QA setup authority after main integration * fix(qa): clarify exploration order and preserve report artifacts * test(qa): follow the shared setup reference directly * refactor(ship): make verification and recovery routes explicit * test(ship): align evidence and review guards with explicit routes * fix(workflows): clarify ship recovery and functional QA evidence * fix(workflows): clarify approval recovery and full QA coverage * refactor(workflows): order review transactions and clarify ship state * fix(ship): clarify final verification and fail closed at publication * fix(evals): attribute native atomic documentation writes * fix(ship): clarify recovery and documentation lifecycle guidance * fix(test): preserve observed native placeholder styling in CI * fix(codex): report watchdog timeouts without a process-exit race * Checkpoint functional QA implementation and workflow validation repairs * Fix documentation and shared-review fixture contracts * docs: clarify judge reuse and evaluation supervision * test: align review evidence and selected case contracts * test: verify append-only documentation checkpoints and recovery * fix: qualify QA workflows and CI validation repairs * fix: launch shared-libs fixture scripts on Windows * fix: qualify QA deadlines, fixture isolation, and shard cleanup * fix: preserve qualified QA and cancellation repairs * fix: enforce functional fixture authority and share strict event decoding * fix: retain free-test evidence and explain recovery * fix: reject malformed native evidence after decoder consolidation * test: use reliable capture for telemetry privacy filters * test: refresh measured quick coverage and document validation costs * Fix native fixture receipts and preserve VM validation evidence * Align negative judge controls with upstream clarity policy * Fix report-only QA preparation and public evidence handling * Clarify QA-only preparation and current-report preservation * Stream Ship quality judgments with an explicit 64k response contract * Validate compact judge reasoning locally with supported wire schema * Align functional QA fixture instructions with evidence acceptance * Bind native browser diagnostics to execution evidence and align review verdicts * Preserve native diagnostic line boundaries * Serialize functional QA evidence from native captures * Keep large QA evidence fixture payload out of Windows argv
4.8 KiB
Finalize a report from retained evidence
Complete these steps before the final report Write. They use retained results, not new probes. Missing evidence stays unknown; an expired clock stays expired. The caller's write boundary includes reports, learning notes and automatic memory. Keep all of them in authorized destinations; a memory feature grants no extra path.
1. Establish each finding once
Ground the report and learnings in retained observations. For every finding, distinguish the observed result, the expected contract and any untested causal hypothesis. Link the supporting command/result or screenshot; unknown impact remains unknown. A console error message does not establish an uncaught exception, failed payload or missing UI. Missing text in a page-text extract does not establish an absent attribute or inaccessible element. Verify those claims with an appropriate probe, or leave them unconfirmed when time expires.
Give each finding one ID and write its Observed, Expected, Evidence and Confirmation fields first. A logged exception-shaped string proves a logged message, not that the named operation executed. Keep possible causes in a separate Hypothesis field; omit speculation that does not help the next investigation. Observed-once is not replay-confirmed.
2. Fill timing fields from their actual boundaries
Report Probe budget (configured limit) and Guarded command time (sum of measured child spans). Measure guard start to child launch as pre-launch elapsed time, and child start to finish as command duration, not a component's latency without its own measurement. A deadline window is not total run time. Gaps between receipts do not measure status/Write overhead or prove how many probes fit; if late, say only that this run dispatched its follow-up after the deadline.
Use Total session elapsed: unmeasured for the invocation whose report is being written.
Its final report Write, acknowledgement and cleanup are not finished yet. Do not fetch
a clock merely to fill that field. An optional Measured interval must cite its actual
start/end receipts and name the work outside those boundaries, including later report
Writes and cleanup; it is never a completed-session measurement.
3. Assemble and check every repetition before writing
Use the caller's selected surface templates and assembly rules. Build headlines, Top 3, summaries and completion text from each finding's Observed and Confirmation fields, not its Hypothesis. Choose one conservative factual sentence per finding and reuse it verbatim in those locations; do not introduce a new causal paraphrase. A disclaimer in the detail does not qualify a stronger claim elsewhere.
Proposed regression assertions must detect the original observation on its actual channel. For a logged console error, capture console errors; exception-only hooks do not detect a console-only message. Additional causal tests remain separate proposals. Apply these evidence limits to proposed tests and learnings too.
Before the final Write, check every mention of each finding against its evidence fields, every proposed test against the observed channel, and each timing claim against its named boundaries. Remove unsupported claims from all sections, not only the detail. Keep refused/unstarted probes and untested categories explicit. Write the report only after this consistency check; do not repair an evidence gap with invented facts. Check claims about frequency and executed checks against the actual commands/results: one observation proves neither recurrence nor an unexecuted check. Apply the same evidence limits to the final response and any caller-authorized learning note.
4. Capture permitted notes, then write the report
Run the learning step below only if its destination is caller-authorized:
the user or invoking workflow explicitly permitted that learning-store path.
Invoking /qa-only alone does not grant this permission. Otherwise
keep notes in REPORT_FILE; do not write learning stores or automatic memory.
No explicit permission: skip learning-store writes and continue to the report.
Explicit permission: Read the named store first. Preserve its existing contents and use the permitted write tool to append a verified note in that store's format. If the format or write interface is unavailable, keep the note in the report instead. Do not run logging helpers: they may write caches or enqueue synchronization outside the permitted path. This branch never changes configuration or enables synchronization.
Write the checked report to the entrypoint's permitted destinations. After the final Write, respond briefly with its path and verified coverage/limits; do not append new findings or timing explanations.