Merge origin/main (v1.91.7.0) into test-audit-reduction

Keep both intents: v1.91.7.0's functional QA, docsync and exploratory
paid cases and their free owners stay; this branch's deletions stay
deleted. main's new paid keys follow the derived-closure touchfile rule
(free *.test.ts paths dropped, static helper/fixture closure added), its
new helper-only tests join the ratchet baseline, and its free selection
examples that named free test files now assert the derived selection.

Periodic CI keeps seven slices without the retired Autoplan slice; the
gate census keeps seven single-worker slices with --skip-judges. Wall
and census literals are recomputed from the merged planner, durations
are re-recorded on Ubicloud, and VERSION stays 1.91.8.0 above 1.91.7.0.
This commit is contained in:
garrytan committed 2026-09-29 13:48:02 +00:00
commit b421bba2c9
325 files changed
+42569 -8257

No files matched your search

+107 -69
View File
@@ -31,28 +31,27 @@ Carry prior approvals into findings, tasks and the report. Routine auto-decide
cannot override user constraints or non-goals.
## CRITICAL RULE — How to ask questions
Follow the AskUserQuestion format from the Preamble above. Additional rules for plan reviews:
Use 0D's decision procedure and the preamble's AskUserQuestion format:
* **One decision unit = one AskUserQuestion call.** Use Step 0D boundaries, not topic labels.
* Describe the problem concretely, with file and line references.
* Present 2-3 options, including "do nothing" where reasonable.
* For each option: effort, risk, and maintenance burden in one line.
* Before calling AskUserQuestion, draft the recommended option as a complete remedy
for this one issue. Its offered description must state the rescue behavior,
verification, and failure visibility needed for that fix. Include those details
in the option itself. Omit irrelevant work, and keep independent findings and
new TODOs in their own questions.
* **Map the reasoning to my engineering preferences above.** One sentence connecting your recommendation to a specific preference.
* Use the preamble's `D<N>` question heading and A/B/C option labels. Cite the stable ledger ID separately so a reopened question keeps its earlier decision history.
* Describe the concrete problem with file/line references. Offer 2-3 options,
including "do nothing" when reasonable.
* Give each option one line covering effort, risk and maintenance.
* The recommended option's description must offer a complete remedy for this
issue: rescue behavior, verification and failure visibility. Exclude unrelated
work; ask about independent findings and new TODOs separately.
* Connect the recommendation to one engineering preference in a sentence.
* Use `D<N>` and A/B/C labels. Cite the stable ledger ID separately to retain
reopened decision history.
* An "obvious fix" still needs approval when it is not covered by an exact accepted choice.
## Formatting Rules
* Keep option labels short; use Step 0D's exact `currentDecision` fields for the question and option descriptions.
* Use short labels and 0D's exact `currentDecision` question and option descriptions.
* Use **CRITICAL GAP** / **WARNING** / **OK** for scannability.
## Mode Quick Reference
The mode changes which work is included, not review depth or section coverage.
Apply the review and outputs to the accepted work in every mode.
Mode controls included work, not depth or section coverage. Review and produce
outputs for accepted work in every mode.
| Step | SCOPE EXPANSION | SELECTIVE EXPANSION | HOLD SCOPE | SCOPE REDUCTION |
|------|-----------------|---------------------|------------|-----------------|
@@ -66,9 +65,9 @@ Apply the review and outputs to the accepted work in every mode.
| Future direction (Section 10) | Review accepted trajectory | Review accepted cherry-picks | Maintainability; no expansions | Maintainability of remaining scope |
| Design (Section 11) | Review if UI scope | Review if UI scope | Review if UI scope | Review if UI scope |
All modes produce the review content. Save it to the permitted working plan;
when no plan/report write is permitted, present it in chat as not persisted and
end with completion blocked. The CEO archive is additional expansion-mode output.
Save to the permitted working plan; with no permitted plan/report write, present
it in chat as not persisted and end with completion blocked. The CEO archive is
additional expansion-mode output.
### Working review decisions
@@ -82,14 +81,18 @@ and mitigations even if later text omits them. Flag approval conflicts. Unavaila
code proves neither failure nor safety; record unknown risks with their owners
and required verification.
**Resolve.** If this section needs a new decision or evidence warrants reopening
one, complete 0D through its post-answer save, then continue to Apply below.
Use the same row ID in the ledger, `currentDecision` and question; complete 0D's
pre-question checkpoint before each new or reopened question.
If all choices are settled, cite their exact answers and go straight to Apply.
Resolve critical risks now. Reference other pending rows in their owner sections;
do not decide them here. Keep independent safety fixes and throughput improvements
in separate rows, following 0D's test table.
**Resolve.** Take the first applicable path for each finding:
1. This section needs a new choice, or evidence warrants reopening its prior
answer: use 0D's Plan decision route through its post-answer save, then
return here to Apply.
Use the same row ID in the ledger, `currentDecision` and question; complete
the pre-question checkpoint before asking. Resolve critical risks now.
2. An exact prior answer covers it: cite that answer and go to Apply.
3. A non-blocking choice belongs to a later section: reference its pending row
and owner; leave it undecided here.
Keep independent safety fixes and throughput improvements in separate rows,
following 0D's test table. No path selects the mode again.
**Apply.** Check the saved plan against each answer's exact scope. Preserve existing
content, approved behavior, required implementation, tests and success/failure
@@ -104,7 +107,14 @@ review or no-UI skip, follow Closing sequence. Keep unresolved choices in the
ledger and report; an approval is not proof of implementation or verification.
### Section 1: Architecture Review
Publish **Current scope** in chat using the Step 0E mode-handoff format and the current ledger dispositions, including actual later scope-answer references. Retain mode, rationale and preference attribution. This updates scope after 0G; do not ask or log the mode again. Keep earlier answers as history, showing current accepted scope. Then say `Section 1: Architecture Review`.
Publish **Current scope** in chat before the architecture analysis:
- Retain 0E's selected mode, rationale and preference attribution.
- Show each governing row's ID, disposition and answer reference, including scope
decisions after 0E. Keep earlier answers as history.
- Distinguish accepted, deferred, rejected and pending work.
This is a scope update, not another mode handoff; do not ask or log the mode again.
Then say `Section 1: Architecture Review`.
Evaluate and diagram:
* System design and component boundaries. Draw the dependency graph.
@@ -353,11 +363,6 @@ _CODEX_CFG=$(~/.claude/skills/gstack/bin/gstack-config get codex_reviews 2>/dev/
source ~/.claude/skills/gstack/bin/gstack-codex-probe 2>/dev/null || true
if [ "$_CODEX_CFG" = "disabled" ]; then
_CODEX_MODE="disabled"
# Running-under-Codex presence probe (#2519): a live Codex session exports
# CODEX_THREAD_ID / CODEX_SANDBOX into every shell it spawns (verified
# against a live `codex exec 'env | grep -i codex'` capture, codex 0.147.0).
# Nested codex spawns from inside a Codex host multiply token burn
# (observed: one /review = 15M tokens). A stale own-harness artifact must stop.
elif { [ -n "${CODEX_THREAD_ID:-}" ] || [ -n "${CODEX_SANDBOX:-}" ] || [ "${GSTACK_ACTIVE_HOST:-}" = codex ]; }; then
_CODEX_MODE="under_codex"
elif ! command -v codex >/dev/null 2>&1; then
@@ -381,11 +386,11 @@ echo "CODEX_MODE: $_CODEX_MODE"
Branch on the echoed `CODEX_MODE`:
- **`disabled`** — the user turned Codex reviews off (`codex_reviews=disabled`). Skip the reviewer invocation; record disabled coverage as directed below; do NOT fall back to a Claude subagent — disabled means no extra review step. Print: "Codex review skipped (codex_reviews disabled). Re-enable: `gstack-config set codex_reviews enabled`."
- **`not_installed`** — Codex CLI absent. Print: "Codex not installed — falling back to a Claude subagent (fresh context, but the same harness; model identity is unknown). Install Codex for an actual outside-model read: `npm install -g @openai/codex`." Fall back to the Claude subagent path.
- **`not_installed`** — Codex CLI absent. Print: "Codex not installed; outside coverage unavailable. Install: `npm install -g @openai/codex`." Fall back to the Claude subagent path.
- **`under_codex`** — stale artifact selected its own harness. Print: "Codex outside review unavailable: harness mismatch; no outside process started. Missing coverage. Repair: setup --host codex." Skip the outside invocation and construct the prompt below, then follow **Native fallback**. Conflicting inherited harness markers are not grounds to guess another provider.
- **`not_authed`** — installed but no credentials. Print: "Codex installed but not authenticated — falling back to a Claude subagent (same harness; model identity is unknown). Run `codex login` or set `$CODEX_API_KEY`." Fall back to the Claude subagent path.
- **`broken_install`** — the CLI is on PATH but cannot execute (spawn ENOENT, non-executable binary, missing vendor payload). Print: "Codex is installed but its binary cannot run — Codex passes skipped. Reinstall: `npm install -g @openai/codex`." Relay the probe's HINT lines and fall back to the Claude subagent path. This state exists because a missing binary used to land in the model probe's fail-open bucket and report `ready`, so every Codex pass was skipped silently (#2742).
- **`model_unusable`** — authed but the account cannot use gstack's selected Codex model (#2477: HTTP 400 on every call). Relay the probe's HINT lines, tell the user the one-line fix (set `GSTACK_CODEX_MODEL=<supported-model>` or pass an explicit `-c model=...` override), and fall back to the Claude subagent path. The ~10s round trip is cached for 1h; timeouts fail open to `ready`.
- **`not_authed`** — installed but no credentials. Print: "Codex not authenticated; outside coverage unavailable. Run `codex login` or set `$CODEX_API_KEY`." Fall back to the Claude subagent path.
- **`broken_install`** — the CLI is on PATH but cannot execute (spawn ENOENT, non-executable binary, missing vendor payload). Print: "Codex is installed but its binary cannot run — Codex passes skipped. Reinstall: `npm install -g @openai/codex`." Relay the probe's HINT lines. Fall back to the Claude subagent path.
- **`model_unusable`** — authed but the account cannot use gstack's selected Codex model (#2477: HTTP 400 on every call). Relay the probe's HINT lines and tell the user the one-line fix (set `GSTACK_CODEX_MODEL=<supported-model>` or pass an explicit `-c model=...` override). Fall back to the Claude subagent path. The ~10s round trip is cached for 1h; timeouts fail open to `ready`.
- **`ready`** — run the Codex pass below.
**Outcome routing:** Follow the row for the current result. After an invocation, route its result
@@ -821,17 +826,18 @@ this run (an empty file means "ran, no findings" — distinct from "didn't run")
### Completion Summary
Fill this template from Review facts now, as part of the plan body. Artifact
outcomes remain pending until their writes are confirmed. Stage 3 publishes it
after report verification; forbidden writes stay labeled not persisted.
Fill this plan-body template from Review facts. Artifact outcomes stay pending
until writes are confirmed. Stage 3 publishes it after report verification;
forbidden writes stay labeled not persisted.
Use the full mode name from Step 0E; replace spaces with underscores only in the
review log's `MODE` field. "System Audit" summarizes repository findings from
Step 0 and the review sections. "Lake Score" counts complete options selected:
Y is the number of answered coverage questions offering a 10/10 option; X is
how many selected that option. Count a reopened choice only once, using its
latest answered option; superseded answers add nothing. Exclude kind-only and
unanswered questions; use `N/A` when Y is zero.
Step 0 and the review sections. Compute "Lake Score" (complete options selected):
1. Select answered questions scored for coverage under 0D that offered a 10/10
option. Exclude unscored mode/scope choices and unanswered questions.
2. Count a reopened choice only once, using its latest answered option.
3. Y is the number of eligible questions; X is how many selected the 10/10
option. Report X/Y, or `N/A` when Y is zero.
```
+====================================================================+
@@ -1049,17 +1055,69 @@ After completing the review, read the review log and config to display the dashb
~/.claude/skills/gstack/bin/gstack-review-read
```
Render each record using its recorded host, source, outside_provider, outside_status, and phase. Historical source "claude" means a native Claude subagent; source "claude-code" means the external CLI. Never infer a historical provider from the current harness. Unknown model identity remains unknown. Missing/disabled/skipped outside coverage is distinct from native completion.
**1. Choose the records to display.** Use the latest record for each row below.
Do not use a record older than 7 days to clear a row, and never substitute an older
success for a newer failure. Ship metrics are not review records.
Parse the output. Find the most recent entry for each skill (plan-ceo-review, plan-eng-review, review, plan-design-review, design-review-lite, adversarial-review, codex-review, codex-plan-review). Ignore entries with timestamps older than 7 days. For the Eng Review row, show whichever is more recent between `review` (diff-scoped pre-landing review) and `plan-eng-review` (plan-stage architecture review). Append "(DIFF)" or "(PLAN)" to the status to distinguish. For the Adversarial row, show whichever is more recent between `adversarial-review` (new auto-scaled) and `codex-review` (legacy). For Design Review, show whichever is more recent between `plan-design-review` (full visual audit) and `design-review-lite` (code-level check). Append "(FULL)" or "(LITE)" to the status to distinguish. For the Outside Voice row, show the most recent `codex-plan-review` entry — this captures outside voices from both /plan-ceo-review and /plan-eng-review.
| Row | Choose the latest of | Status suffix |
|---|---|---|
| Eng Review | `review` or `plan-eng-review` | (DIFF) or (PLAN) |
| CEO Review | `plan-ceo-review` | — |
| Design Review | `plan-design-review` or `design-review-lite` | (FULL) or (LITE) |
| Adversarial | `adversarial-review` or legacy `codex-review` | — |
| Outside Voice | `codex-plan-review` from CEO or Eng review | — |
**Source attribution:** If the most recent entry for a skill has a \`"via"\` field, append it to the status label in parentheses. Examples: `plan-eng-review` with `via:"autoplan"` shows as "CLEAR (PLAN via /autoplan)". `review` with `via:"ship"` shows as "CLEAR (DIFF via /ship)". Entries without a `via` field show as "CLEAR (PLAN)" or "CLEAR (DIFF)" as before.
Keep each record's host, source, outside_provider, outside_status and phase.
Historical source "claude" is a native subagent; "claude-code" is the external CLI.
Do not infer old providers or unknown models from today's harness. A native result
does not fill missing, disabled or skipped outside coverage.
From gstack-review-read output, use entries whose skill is `autoplan-voices` or `design-outside-voices` for the coverage detail below the dashboard. Group by workflow run and phase, not merely skill. Show each phase’s recorded provider and outside_status; partial coverage must remain partial. These records do not change the engineering gate.
**Source attribution:** Append a recorded `via` to the suffix, for example
"CLEAR (PLAN via /autoplan)" or "CLEAR (DIFF via /ship)". Without `via`, keep
"CLEAR (PLAN)" or "CLEAR (DIFF)". Below the dashboard, group `autoplan-voices`
and `design-outside-voices` by workflow run and phase. Show each phase's provider
and outside_status; retain partial coverage. These details do not clear Eng Review.
Display a fresh `clean` result as CLEAR and `issues_open` as ISSUES OPEN. Show missing, stale, disabled or unavailable results explicitly; none implies CLEAR. Keep the logged status unchanged.
**2. Check freshness before choosing a verdict.**
Display:
- **Content-first rule:** For `review`, `adversarial-review`, `codex-review`,
ship-stage reviews and `design-review-lite`, use `review_freshness.status`
and show its `reason`. CURRENT means a completed clean review whose start and
end content fingerprints equal the current `---WTREE---` fingerprint. This
fingerprint covers working-tree content, not just the commit.
STALE or UNVERIFIED cannot clear Eng Review. Missing `review_freshness`,
including legacy log-only records, means UNVERIFIED. Never fall back to HEAD
equality or commit distance for diff evidence, even at zero commits.
Show recorded cycles, completed/converged fields and missing source/phase
coverage. Unknown coverage is not a pass.
- **Plan records** (plan-ceo-review, plan-eng-review, plan-design-review and
codex-plan-review) use the 7-day window, not the working-tree fingerprint.
If `plan_sha256` is present, you may compare the plan file and report a mismatch.
For plan records only, compare the recorded commit with `---HEAD---`.
If different, run `git rev-list --count STORED_COMMIT..HEAD` and report
"Note: {skill} review from {date} may be stale — {N} commits since review".
A failed command means UNKNOWN, treated as stale. Without commit tracking,
retain the note to consider re-running. Omit staleness notes when all reviews
are current.
**3. Choose the historical verdict.** CLEARED requires the selected Eng Review
to be `clean`, within 7 days and fresh under step 2. Otherwise report NOT CLEARED
and its missing, stale or open-issue reason. If `skip_eng_review` is true, show
"SKIPPED (global)" for Eng Review and CLEARED for this dashboard.
Eng Review is required by default; `gstack-config set skip_eng_review true` disables that requirement.
Other rows provide context, not a substitute for Eng Review:
- Recommend CEO Review for product/business or scope decisions, not routine fixes or cleanup.
- Recommend Design Review for UI/UX work, not backend, infrastructure or prompt-only work.
- Adversarial review always includes a native pass. Available, enabled outside
challenges supplement it; diffs of 200+ lines also get the structured P1 gate.
- Outside Voice is the default-on plan review after CEO/Eng review. `codex_reviews`
disables that extra step. Provider failure uses native fallback and records
missing outside coverage; this dashboard row never gates shipping.
**4. Display the dashboard.** Show missing, stale, disabled or unavailable results
explicitly, never as CLEAR. Display a fresh `clean` result as CLEAR and
`issues_open` as ISSUES OPEN without changing the stored status.
```
+====================================================================+
@@ -1077,26 +1135,6 @@ Display:
+====================================================================+
```
**Review tiers:**
- **Eng Review (required by default):** The only review that gates shipping. Covers architecture, code quality, tests, performance. Can be disabled globally with \`gstack-config set skip_eng_review true\` (the "don't bother me" setting).
- **CEO Review (optional):** Use your judgment. Recommend it for big product/business changes, new user-facing features, or scope decisions. Skip for bug fixes, refactors, infra, and cleanup.
- **Design Review (optional):** Use your judgment. Recommend it for UI/UX changes. Skip for backend-only, infra, or prompt-only changes.
- **Adversarial Review (automatic):** Always-on for every review. Every diff gets a native adversarial pass and, when enabled and available, a host-selected outside challenge. Large diffs (200+ lines) additionally get a structured outside review with P1 gate.
- **Outside Voice (default-on):** Independent plan review through the host-selected provider after /plan-ceo-review and /plan-eng-review. The codex_reviews switch disables the entire extra step. Provider failure uses the existing native fallback and reports missing outside coverage. Never gates shipping.
**Verdict logic:**
- **CLEARED**: Eng Review has >= 1 entry within 7 days from either \`review\` or \`plan-eng-review\` with status "clean"; diff review must also grade CURRENT below (or \`skip_eng_review\` is \`true\`)
- **NOT CLEARED**: Eng Review missing, stale (>7 days), or has open issues
- CEO, Design, and outside reviews are shown for context but never block shipping
- If \`skip_eng_review\` config is \`true\`, Eng Review shows "SKIPPED (global)" and verdict is CLEARED
**Staleness detection:** Grade before deciding CLEARED:
- Ship telemetry reports metrics, not review coverage; it never satisfies a review row.
- **Content-first rule (diff-scoped rows only: `review`, `adversarial-review`, `codex-review`, ship-stage entries, `design-review-lite`).** Use the helper's computed `review_freshness.status` and show its `reason`. CURRENT requires a completed clean pass with captured start/end wtree equal to the current `---WTREE---`. STALE or UNVERIFIED never clears Eng Review. Missing `review_freshness` is UNVERIFIED, including legacy log-only rows. Never fall back to HEAD equality or commit distance for diff evidence, even at 0 commits. Show recorded cycles, completed/converged state, and missing per-source/phase coverage; unknown is not a pass.
- Plan-tier rows (plan-ceo-review, plan-eng-review, plan-design-review, codex-plan-review) grade a plan file, not the repo tree — never apply the wtree rule to them; they keep the 7-day freshness logic. If an entry carries `plan_sha256`, you MAY compare it with the plan file and note "plan changed since review" on mismatch.
- Plan-tier fallback only: parse `---HEAD---`. For entries with a different `commit`, count elapsed commits: `git rev-list --count STORED_COMMIT..HEAD`. If that command FAILS, grade UNKNOWN and treat as stale. Display: "Note: {skill} review from {date} may be stale — {N} commits since review". Missing commit tracking retains the legacy note to consider re-running.
- If all reviews grade CURRENT, do not display staleness notes
## Next Steps — Review Chaining
After displaying the Review Readiness Dashboard, recommend the next review(s) based on what this CEO review discovered. Read the dashboard output to see which reviews have already been run and whether they are stale.
@@ -29,28 +29,27 @@ Carry prior approvals into findings, tasks and the report. Routine auto-decide
cannot override user constraints or non-goals.
## CRITICAL RULE — How to ask questions
Follow the AskUserQuestion format from the Preamble above. Additional rules for plan reviews:
Use 0D's decision procedure and the preamble's AskUserQuestion format:
* **One decision unit = one AskUserQuestion call.** Use Step 0D boundaries, not topic labels.
* Describe the problem concretely, with file and line references.
* Present 2-3 options, including "do nothing" where reasonable.
* For each option: effort, risk, and maintenance burden in one line.
* Before calling AskUserQuestion, draft the recommended option as a complete remedy
for this one issue. Its offered description must state the rescue behavior,
verification, and failure visibility needed for that fix. Include those details
in the option itself. Omit irrelevant work, and keep independent findings and
new TODOs in their own questions.
* **Map the reasoning to my engineering preferences above.** One sentence connecting your recommendation to a specific preference.
* Use the preamble's `D<N>` question heading and A/B/C option labels. Cite the stable ledger ID separately so a reopened question keeps its earlier decision history.
* Describe the concrete problem with file/line references. Offer 2-3 options,
including "do nothing" when reasonable.
* Give each option one line covering effort, risk and maintenance.
* The recommended option's description must offer a complete remedy for this
issue: rescue behavior, verification and failure visibility. Exclude unrelated
work; ask about independent findings and new TODOs separately.
* Connect the recommendation to one engineering preference in a sentence.
* Use `D<N>` and A/B/C labels. Cite the stable ledger ID separately to retain
reopened decision history.
* An "obvious fix" still needs approval when it is not covered by an exact accepted choice.
## Formatting Rules
* Keep option labels short; use Step 0D's exact `currentDecision` fields for the question and option descriptions.
* Use short labels and 0D's exact `currentDecision` question and option descriptions.
* Use **CRITICAL GAP** / **WARNING** / **OK** for scannability.
## Mode Quick Reference
The mode changes which work is included, not review depth or section coverage.
Apply the review and outputs to the accepted work in every mode.
Mode controls included work, not depth or section coverage. Review and produce
outputs for accepted work in every mode.
| Step | SCOPE EXPANSION | SELECTIVE EXPANSION | HOLD SCOPE | SCOPE REDUCTION |
|------|-----------------|---------------------|------------|-----------------|
@@ -64,9 +63,9 @@ Apply the review and outputs to the accepted work in every mode.
| Future direction (Section 10) | Review accepted trajectory | Review accepted cherry-picks | Maintainability; no expansions | Maintainability of remaining scope |
| Design (Section 11) | Review if UI scope | Review if UI scope | Review if UI scope | Review if UI scope |
All modes produce the review content. Save it to the permitted working plan;
when no plan/report write is permitted, present it in chat as not persisted and
end with completion blocked. The CEO archive is additional expansion-mode output.
Save to the permitted working plan; with no permitted plan/report write, present
it in chat as not persisted and end with completion blocked. The CEO archive is
additional expansion-mode output.
### Working review decisions
@@ -80,14 +79,18 @@ and mitigations even if later text omits them. Flag approval conflicts. Unavaila
code proves neither failure nor safety; record unknown risks with their owners
and required verification.
**Resolve.** If this section needs a new decision or evidence warrants reopening
one, complete 0D through its post-answer save, then continue to Apply below.
Use the same row ID in the ledger, `currentDecision` and question; complete 0D's
pre-question checkpoint before each new or reopened question.
If all choices are settled, cite their exact answers and go straight to Apply.
Resolve critical risks now. Reference other pending rows in their owner sections;
do not decide them here. Keep independent safety fixes and throughput improvements
in separate rows, following 0D's test table.
**Resolve.** Take the first applicable path for each finding:
1. This section needs a new choice, or evidence warrants reopening its prior
answer: use 0D's Plan decision route through its post-answer save, then
return here to Apply.
Use the same row ID in the ledger, `currentDecision` and question; complete
the pre-question checkpoint before asking. Resolve critical risks now.
2. An exact prior answer covers it: cite that answer and go to Apply.
3. A non-blocking choice belongs to a later section: reference its pending row
and owner; leave it undecided here.
Keep independent safety fixes and throughput improvements in separate rows,
following 0D's test table. No path selects the mode again.
**Apply.** Check the saved plan against each answer's exact scope. Preserve existing
content, approved behavior, required implementation, tests and success/failure
@@ -102,7 +105,14 @@ review or no-UI skip, follow Closing sequence. Keep unresolved choices in the
ledger and report; an approval is not proof of implementation or verification.
### Section 1: Architecture Review
Publish **Current scope** in chat using the Step 0E mode-handoff format and the current ledger dispositions, including actual later scope-answer references. Retain mode, rationale and preference attribution. This updates scope after 0G; do not ask or log the mode again. Keep earlier answers as history, showing current accepted scope. Then say `Section 1: Architecture Review`.
Publish **Current scope** in chat before the architecture analysis:
- Retain 0E's selected mode, rationale and preference attribution.
- Show each governing row's ID, disposition and answer reference, including scope
decisions after 0E. Keep earlier answers as history.
- Distinguish accepted, deferred, rejected and pending work.
This is a scope update, not another mode handoff; do not ask or log the mode again.
Then say `Section 1: Architecture Review`.
Evaluate and diagram:
* System design and component boundaries. Draw the dependency graph.
@@ -443,17 +453,18 @@ List every ASCII diagram in files this plan touches. Still accurate?
{{TASKS_SECTION_EMIT:ceo-review}}
### Completion Summary
Fill this template from Review facts now, as part of the plan body. Artifact
outcomes remain pending until their writes are confirmed. Stage 3 publishes it
after report verification; forbidden writes stay labeled not persisted.
Fill this plan-body template from Review facts. Artifact outcomes stay pending
until writes are confirmed. Stage 3 publishes it after report verification;
forbidden writes stay labeled not persisted.
Use the full mode name from Step 0E; replace spaces with underscores only in the
review log's `MODE` field. "System Audit" summarizes repository findings from
Step 0 and the review sections. "Lake Score" counts complete options selected:
Y is the number of answered coverage questions offering a 10/10 option; X is
how many selected that option. Count a reopened choice only once, using its
latest answered option; superseded answers add nothing. Exclude kind-only and
unanswered questions; use `N/A` when Y is zero.
Step 0 and the review sections. Compute "Lake Score" (complete options selected):
1. Select answered questions scored for coverage under 0D that offered a 10/10
option. Exclude unscored mode/scope choices and unanswered questions.
2. Count a reopened choice only once, using its latest answered option.
3. Y is the number of eligible questions; X is how many selected the 10/10
option. Report X/Y, or `N/A` when Y is zero.
```
+====================================================================+