mirror of
https://github.com/garrytan/gstack.git
synced 2026-10-04 18:36:54 +02:00
v1.91.7.0 feat: add functional QA and pre-publication docs checks (#2983)
* feat: add surface-aware exploratory QA and ship documentation gates * test: preserve delegated QA setup authority after main integration * fix(qa): clarify exploration order and preserve report artifacts * test(qa): follow the shared setup reference directly * refactor(ship): make verification and recovery routes explicit * test(ship): align evidence and review guards with explicit routes * fix(workflows): clarify ship recovery and functional QA evidence * fix(workflows): clarify approval recovery and full QA coverage * refactor(workflows): order review transactions and clarify ship state * fix(ship): clarify final verification and fail closed at publication * fix(evals): attribute native atomic documentation writes * fix(ship): clarify recovery and documentation lifecycle guidance * fix(test): preserve observed native placeholder styling in CI * fix(codex): report watchdog timeouts without a process-exit race * Checkpoint functional QA implementation and workflow validation repairs * Fix documentation and shared-review fixture contracts * docs: clarify judge reuse and evaluation supervision * test: align review evidence and selected case contracts * test: verify append-only documentation checkpoints and recovery * fix: qualify QA workflows and CI validation repairs * fix: launch shared-libs fixture scripts on Windows * fix: qualify QA deadlines, fixture isolation, and shard cleanup * fix: preserve qualified QA and cancellation repairs * fix: enforce functional fixture authority and share strict event decoding * fix: retain free-test evidence and explain recovery * fix: reject malformed native evidence after decoder consolidation * test: use reliable capture for telemetry privacy filters * test: refresh measured quick coverage and document validation costs * Fix native fixture receipts and preserve VM validation evidence * Align negative judge controls with upstream clarity policy * Fix report-only QA preparation and public evidence handling * Clarify QA-only preparation and current-report preservation * Stream Ship quality judgments with an explicit 64k response contract * Validate compact judge reasoning locally with supported wire schema * Align functional QA fixture instructions with evidence acceptance * Bind native browser diagnostics to execution evidence and align review verdicts * Preserve native diagnostic line boundaries * Serialize functional QA evidence from native captures * Keep large QA evidence fixture payload out of Windows argv
This commit is contained in:
1 parent
65bfb0ce49
commit
dcaea52800
333 files changed
+41755
-7357
No files matched your search
@@ -1,29 +1,50 @@
|
||||
## Step 8: Plan Completion Audit
|
||||
|
||||
**Dispatch this step as a subagent** using the Agent tool with `subagent_type: "general-purpose"`. The subagent reads the plan file and every referenced code file in its own fresh context. Parent gets only the conclusion.
|
||||
Complete this section in order:
|
||||
1. Dispatch the audit, validate its result and resolve its Gate Logic.
|
||||
2. Collect the plan's executable checks in Step 8.1; do not run them yet.
|
||||
3. Run Step 8.2 Scope Drift.
|
||||
4. Run Prior Learnings, including its setting question when offered, then proceed to Step 9 for review and QA.
|
||||
|
||||
{{FOREGROUND_DISPATCH_NOTE}} The Gate Logic below consumes this audit's LAST-line JSON before /ship can proceed.
|
||||
**Dispatch this step as a subagent** using Agent, `subagent_type: "general-purpose"`
|
||||
and `run_in_background: false`. Use Step 7's shared foreground-dispatch rule.
|
||||
The child reads the plan and every referenced
|
||||
code file; the parent validates its report and applies the gates below.
|
||||
|
||||
**Subagent prompt:** Pass these instructions to the subagent:
|
||||
**Subagent prompt:** Substitute `<base>` and supply the active plan's absolute path
|
||||
or complete text, including relevant user-approved scope changes. If none exists,
|
||||
say so explicitly and let the child use the fallback search below. The child does
|
||||
not inherit the parent's conversation.
|
||||
|
||||
````text
|
||||
You are running a ship-workflow plan completion audit. The base branch is `<base>`. Use `git diff origin/<base>` and inspect untracked files from `git status` to see the full proposed change. Do not commit or push. Report only: classify every item, but do not execute Gate Logic, ask the user, or advance the workflow. The parent applies those gates to your report.
|
||||
|
||||
{{PLAN_COMPLETION_AUDIT_SHIP}}
|
||||
|
||||
After your analysis, output a single JSON object on the LAST LINE of your response (no other text after it):
|
||||
After your analysis, output a single JSON object with exactly these seven fields on the LAST LINE of your response (no other text after it):
|
||||
{"total_items":N,"done":N,"changed":N,"partial":N,"not_done":N,"unverifiable":N,"summary":"<markdown checklist for PR body>"}
|
||||
Counts map one-to-one to the classifications above and sum to total_items. No plan or no actionable items means all counts are zero with the skip reason in summary. Do not classify work as deferred; only the parent can record a user-approved deferral.
|
||||
````
|
||||
|
||||
**Parent processing:**
|
||||
|
||||
1. Parse the LAST line as JSON. A non-null `error`, any missing count or count that is not a nonnegative integer, classification count sum unequal to `total_items`, or non-string `summary` takes the audit-failure fallback below. Validate every count field in the contract above. Valid no-plan/no-actionable-item reports retain zero counts and their summary.
|
||||
2. Store the counts for Step 20 metrics; use `summary` in PR body.
|
||||
3. Apply Gate Logic below to `not_done` and `unverifiable` before continuing. Carry approved deferrals, with item text and plan path, to Step 14; keep them separate from dropped scope. `partial` items receive a PR note, not the NOT DONE gate.
|
||||
4. Embed `summary` in PR body's `## Plan Completion` section (Step 19). For the UNVERIFIABLE gate, also embed `## Plan Completion — Manual Verifications` with each Y response's evidence and each D response's dropped item.
|
||||
1. Check the task's terminal status. Without successful completion and valid LAST-line
|
||||
JSON, use the audit-failure fallback below. Require exactly the seven declared
|
||||
fields: nonnegative integer counts whose classification sum equals `total_items`,
|
||||
and a string `summary`. Missing,
|
||||
extra or invalid fields fail. Valid no-plan/no-actionable reports retain zero counts
|
||||
and their summary.
|
||||
2. Store counts for Step 20 and `summary` for Step 19's `## Plan Completion`.
|
||||
3. Apply Gate Logic below before continuing. Carry approved deferrals, with item text
|
||||
and plan path, to Step 14; keep them separate from dropped scope. The gate supplies
|
||||
the required PR notes and per-item manual verification evidence.
|
||||
|
||||
**If the subagent fails, returns invalid JSON, or has no final output after ~10 minutes:** Stop any still-running background task before an inline fallback using the same extraction/classification logic; never race its late result. If fallback also fails, AskUserQuestion: "Audit failed ({reason}): A) Skip audit and ship anyway, recording the skip in PR body and Step 20 metrics; B) Stop and fix the audit (recommended/default)." Silent fail-open is the failure shape that VAS-449 surfaced.
|
||||
**Audit-failure fallback:** On failure, invalid JSON or no final output after ~10
|
||||
minutes, stop any live child and confirm it stopped before an inline audit with the same
|
||||
extraction/classification logic; never race a late result. If that also fails,
|
||||
AskUserQuestion: A) Skip audit and ship, recording the reason in the PR body and
|
||||
Step 20 metrics; B) Stop and fix the audit (recommended/default). Silent fail-open
|
||||
is the failure shape that VAS-449 surfaced.
|
||||
|
||||
---
|
||||
|
||||
@@ -33,9 +54,6 @@ Counts map one-to-one to the classifications above and sum to total_items. No pl
|
||||
|
||||
{{SCOPE_DRIFT}}
|
||||
|
||||
The parent now runs Prior Learnings and its cross-project setting question when
|
||||
offered, before Step 9, even when no plan file was found.
|
||||
|
||||
{{LEARNINGS_SEARCH:query=release ship version changelog merge pr}}
|
||||
|
||||
---
|
||||
Reference in new issue
Block a user