v1.91.7.0 feat: add functional QA and pre-publication docs checks (#2983)

* feat: add surface-aware exploratory QA and ship documentation gates

* test: preserve delegated QA setup authority after main integration

* fix(qa): clarify exploration order and preserve report artifacts

* test(qa): follow the shared setup reference directly

* refactor(ship): make verification and recovery routes explicit

* test(ship): align evidence and review guards with explicit routes

* fix(workflows): clarify ship recovery and functional QA evidence

* fix(workflows): clarify approval recovery and full QA coverage

* refactor(workflows): order review transactions and clarify ship state

* fix(ship): clarify final verification and fail closed at publication

* fix(evals): attribute native atomic documentation writes

* fix(ship): clarify recovery and documentation lifecycle guidance

* fix(test): preserve observed native placeholder styling in CI

* fix(codex): report watchdog timeouts without a process-exit race

* Checkpoint functional QA implementation and workflow validation repairs

* Fix documentation and shared-review fixture contracts

* docs: clarify judge reuse and evaluation supervision

* test: align review evidence and selected case contracts

* test: verify append-only documentation checkpoints and recovery

* fix: qualify QA workflows and CI validation repairs

* fix: launch shared-libs fixture scripts on Windows

* fix: qualify QA deadlines, fixture isolation, and shard cleanup

* fix: preserve qualified QA and cancellation repairs

* fix: enforce functional fixture authority and share strict event decoding

* fix: retain free-test evidence and explain recovery

* fix: reject malformed native evidence after decoder consolidation

* test: use reliable capture for telemetry privacy filters

* test: refresh measured quick coverage and document validation costs

* Fix native fixture receipts and preserve VM validation evidence

* Align negative judge controls with upstream clarity policy

* Fix report-only QA preparation and public evidence handling

* Clarify QA-only preparation and current-report preservation

* Stream Ship quality judgments with an explicit 64k response contract

* Validate compact judge reasoning locally with supported wire schema

* Align functional QA fixture instructions with evidence acceptance

* Bind native browser diagnostics to execution evidence and align review verdicts

* Preserve native diagnostic line boundaries

* Serialize functional QA evidence from native captures

* Keep large QA evidence fixture payload out of Windows argv
This commit is contained in:
Garry Tan authored and GitHub committed 2026-09-29 06:07:35 -07:00
1 parent 65bfb0ce49
commit dcaea52800
333 files changed
+41755 -7357

No files matched your search

+34
View File
@@ -17,6 +17,40 @@
import type { TemplateContext } from './types';
export function generateConfidenceCalibration(_ctx: TemplateContext): string {
if (_ctx.skillName === 'review') return `## Confidence Calibration
Verify evidence first, then score every finding (1-10) and apply its display rule.
### Pre-emit verification gate
1. **Quote the specific code line:** file:line and verbatim text. For a missing field,
quote its class definition; for a nullable value, its initialization; for a race, both sides.
2. For framework-generated symbols, read and quote their generating metaclass,
descriptor, ORM Meta block, migration, decorator or schema. Missing literal
names in the class body or grep results do not prove absence.
3. **If you cannot quote the motivating line(s), the finding is unverified.**
Force its confidence to 4-5: use 4 for appendix-only reporting, or 5 only when
the finding belongs in the main report with the medium-confidence caveat below.
Never invent speculative confidence 7+.
| Score | Meaning | Display rule |
|-------|---------|-------------|
| 9-10 | Specific code verifies a concrete bug or exploit. | Show normally |
| 7-8 | High-confidence pattern match; very likely correct. | Show normally |
| 5-6 | Moderate; could be a false positive. | Show with caveat: "Medium confidence, verify this is actually an issue" |
| 3-4 | Suspicious but may be fine. | Suppress from main report. Include in appendix only. |
| 1-2 | Speculation. | Only report a suspected release-blocking catastrophe (widespread data loss, total outage or system-wide compromise); label it CRITICAL and explicitly speculative. |
**Finding format:**
\`[CRITICAL|INFORMATIONAL] (confidence: N/10) file:line — description\`
Example:
\`[CRITICAL] (confidence: 9/10) user.rb:42 — SQL injection via string interpolation\`
**Calibration learning:** If the user confirms a reported finding scored < 7 is
real, log the corrected pattern as a learning.`;
const result = `## Confidence Calibration
Every finding MUST include a confidence score (1-10):