mirror of
https://github.com/garrytan/gstack.git
synced 2026-10-02 17:40:02 +02:00
Merge origin/main (v1.91.7.0) into test-audit-reduction
Keep both intents: v1.91.7.0's functional QA, docsync and exploratory paid cases and their free owners stay; this branch's deletions stay deleted. main's new paid keys follow the derived-closure touchfile rule (free *.test.ts paths dropped, static helper/fixture closure added), its new helper-only tests join the ratchet baseline, and its free selection examples that named free test files now assert the derived selection. Periodic CI keeps seven slices without the retired Autoplan slice; the gate census keeps seven single-worker slices with --skip-judges. Wall and census literals are recomputed from the merged planner, durations are re-recorded on Ubicloud, and VERSION stays 1.91.8.0 above 1.91.7.0.
This commit is contained in:
commit
b421bba2c9
325 files changed
+42569
-8257
No files matched your search
+69
-59
@@ -500,9 +500,9 @@ Never skip Step 0, system audit, error/rescue map or failure modes.
|
||||
|
||||
## Web research runs in Aside
|
||||
|
||||
For web research, do it through Aside's own agent first, using the user's signed-in browser. If Aside is not ready, fall back to the WebSearch tool when this host provides one.
|
||||
For research, do it through Aside's own agent first. If Aside is not ready, fall back to the WebSearch tool when this host provides one.
|
||||
|
||||
Check once (if this skill already ran this same probe, in BROWSER SETUP or Third-Party Web Actions, reuse its answer):
|
||||
Check once per run that Aside is ready (if this skill already ran this same probe, in BROWSER SETUP or Third-Party Web Actions, reuse its answer):
|
||||
|
||||
```bash
|
||||
_gs_d() { if command -v gtimeout >/dev/null; then gtimeout 30 "$@"; elif command -v timeout >/dev/null; then timeout 30 "$@"
|
||||
@@ -531,7 +531,7 @@ fi
|
||||
|
||||
- Any non-READY result: report only the safe status, never raw diagnostics. Run the same queries with the WebSearch tool if available, still read-only and untrusted. Otherwise say once: "Search unavailable — proceeding with in-distribution knowledge only." Never install Aside yourself; mention aside.com at most once per run. Continue the skill.
|
||||
|
||||
Sanitize every query before it leaves the machine: strip hostnames, IPs, file paths, SQL fragments, and anything that looks like a secret. Search for the error class and the library, not the user's data.
|
||||
Sanitize every query before it leaves the machine: strip hostnames, IPs, file paths, SQL and secrets. Search for the error class and library, never the user's data.
|
||||
|
||||
**Anti-shortcut clause:** Analyze → resolve → apply for each section before advancing. The plan file records the interactive review; it cannot replace it. Do not prewrite the remaining sections or their implementation tasks and then walk through a fixed question list. Proposed findings are not accepted plan changes: mark them pending until their actual decisions are made. Ask once per unresolved or reopened issue, wait for the answer, and apply only the exact accepted choice and scope to the working plan. An earlier approach selection does not authorize unrelated choices. Keep established contracts, accepted decisions, and their evidence available to later sections; new material risks or changed remedies still need approval. Cross-referencing settled decisions never replaces the full review and terminal report. Follow the working review decisions below; never invent a question merely because a new section starts.
|
||||
|
||||
@@ -822,15 +822,14 @@ single choice. To expand strategy-only into implementation design, use 0D with
|
||||
**A)** Keep this review strategy-only **B)** Add implementation design for the
|
||||
named capability. Recommend A unless a concrete blocker requires B; wait for the
|
||||
answer. B permits design detail for that capability only.
|
||||
Resolve a choice only when output would be wrong without it, a blocker would be
|
||||
hidden, or scope would change. Reuse prior answers only for the same scope.
|
||||
|
||||
Plain terms:
|
||||
- **Required choice:** a mode, scope, deferral, TODO, spec, outside-review or
|
||||
finding decision needed before the next step.
|
||||
- **Pending:** recorded in the ledger and waiting for approval.
|
||||
- **Settled:** answered by the user, directly instructed, or auto-authorized by
|
||||
the preamble.
|
||||
- **Required choice:** unanswered. Resolve a choice only when continuing would
|
||||
change scope, hide a blocker or produce the wrong output.
|
||||
- **Pending:** unapproved; keep in Proposed, not tasks or accepted work. Status is
|
||||
`unresolved` or `reopened`.
|
||||
- **Settled:** an answer, direct instruction or authorized auto-decision resolves
|
||||
this exact choice and scope; a recommendation does not.
|
||||
|
||||
Review depth controls the detail within each section. Review Sections 1–10 in every depth;
|
||||
run Section 11 only for UI. Strategy-only uses capability-level rows and
|
||||
@@ -838,11 +837,14 @@ run Section 11 only for UI. Strategy-only uses capability-level rows and
|
||||
Implementation-ready names interfaces, codepaths, rescue behavior and tests.
|
||||
For one narrow decision, apply every section to that choice and its dependencies.
|
||||
|
||||
**Keep the stated limits.** Record each measure, value, unit and prerequisite. Count all deliverables, including reused code, as scope; 0E estimates only files that will change. Changing a limit needs evidence and user approval.
|
||||
**Keep the stated limits.** Record each measure, value, unit and prerequisite.
|
||||
Count all deliverables, including reused code, against scope limits. Separately,
|
||||
0E counts changed files, excluding unchanged reuse, to recommend a mode.
|
||||
Neither count approves changes. Changing a limit needs evidence and user approval.
|
||||
|
||||
**Storage policy: choose before writing.** Honor user/host artifact and cleanup
|
||||
limits. One working plan: requested output, else reviewed plan, else host active
|
||||
plan. Use native Write for a missing file and scoped Edit for checkpoints;
|
||||
**Storage policy: choose before writing.** Honor user/host write and cleanup limits.
|
||||
Use one working plan: requested output, else reviewed plan, else host active plan.
|
||||
Use native Write for a missing file and scoped Edit for checkpoints;
|
||||
retain all current content, ledger rows and comparisons.
|
||||
|
||||
**Artifact outcomes:** Never claim an unconfirmed save, read-back or log.
|
||||
@@ -857,9 +859,9 @@ ExitPlanMode or next-skill handoff.
|
||||
| 0H spec-review metrics | Stop with the cause; reviewer availability does not waive this write. |
|
||||
| Review, decision and question history logs | Report cause and unsaved fields; continue. The plan's ledger is still required. |
|
||||
|
||||
Paths are per output: resolve the CEO archive as `CEO_PLANS` in 0H; tasks
|
||||
use `~/.gstack/projects/`, metrics use `~/.gstack/analytics/`, and log helpers
|
||||
choose their own paths. Do not substitute the CEO archive root for these paths.
|
||||
Paths differ: 0H resolves `CEO_PLANS`; tasks use `~/.gstack/projects/`, metrics
|
||||
use `~/.gstack/analytics/`, and log helpers choose their paths. Never substitute
|
||||
the CEO archive root for task, metric or log paths.
|
||||
|
||||
Keep one decision ledger through Step 0, Spec Review Loop and Outside Voice:
|
||||
|
||||
@@ -892,15 +894,17 @@ With no required choice, or after those choices settle, go to 0E.
|
||||
**Choose the question's route first:**
|
||||
- **Admin question:** mode, setup, navigation, document approval or promotion.
|
||||
Use its listed menu and the preamble question transport, then wait and record
|
||||
the answer. Skip steps 1–4; this approves no plan changes. For mode selection,
|
||||
0E defines the four-option menu and any authorized automatic preference;
|
||||
neither needs a plan-decision row, comparison grid or completeness score.
|
||||
the answer. Skip steps 1–4; this approves no plan changes. Resume that menu's
|
||||
next step. 0E owns mode selection; 0H owns document approval. Neither needs
|
||||
a plan-decision row or comparison grid.
|
||||
- **Plan decision:** review-depth expansion, scope additions/cuts, approach
|
||||
choices, TODOs, specs and review/outside findings. Start at step 1. Reuse exact
|
||||
prior approvals; run steps 2–4 only when a new answer is needed, even for one option.
|
||||
|
||||
If an admin answer requests a plan change, use the Plan decision route for that
|
||||
change. 0D never restarts mode selection.
|
||||
change before resuming. 0G proposals and section findings use this route even
|
||||
with prescribed menus. 0D returns to its caller, not to mode selection.
|
||||
For mode changes, follow 0E's **Mode change** instruction.
|
||||
|
||||
**1. Check sources and prior answers.**
|
||||
Compare input, source and answers; correct facts, flag conflicts and preserve unknowns.
|
||||
@@ -930,12 +934,11 @@ Build one `currentDecision` using these fields and the preamble format:
|
||||
| `header` and option labels | Final native text within host limits; exactly one label includes `(recommended)`. |
|
||||
| Each option's `description` | A 1–2 sentence summary; S/M/L/XL effort, low/medium/high risk, reuse, verification coverage, at least 2 ✅ pros and 1 ❌ con. Apply the preamble's minimum lengths and destructive-choice exception. |
|
||||
|
||||
For a plan decision without a prescribed menu, offer 2–3 options (prefer 3 for
|
||||
non-trivial plans). This default does not replace an admin or scope menu.
|
||||
Without a prescribed menu, offer 2–3 options (prefer 3 for non-trivial plans).
|
||||
For an option with no implementation, use effort S and state zero implementation
|
||||
work, never effort 0. Weigh diff size and long-term architecture equally,
|
||||
including rewrites: state the immediate changed-file cost and the future
|
||||
maintenance cost for each option, then explain both in the recommendation.
|
||||
including rewrites: compare immediate changed-file cost and future maintenance
|
||||
cost for each option; explain both in the recommendation.
|
||||
|
||||
In Proposed, compare every commitment in the labels, descriptions and pros/cons:
|
||||
|
||||
@@ -944,12 +947,16 @@ Commitment | Source/approval or pending | Current | A | B | C
|
||||
```
|
||||
|
||||
Include one column per option (add D for a four-option menu). Show unchanged,
|
||||
shared and pending values. Changes remain separate decisions even if they use the same framework.
|
||||
Keep other rows fixed or pending; preserve requirements, tests and fixes.
|
||||
shared and pending values. Keep independent changes separate even within one
|
||||
framework; other rows stay fixed or pending. Preserve requirements, tests and fixes.
|
||||
|
||||
Score this row's coverage differences: 10 = all edge cases, 7 = happy path,
|
||||
3 = shortcut. For different kinds of work, write:
|
||||
"Note: options differ in kind, not coverage — no completeness score."
|
||||
Choose scoring before saving:
|
||||
- **Same work, different coverage:** Score this row's coverage differences:
|
||||
10 = all edge cases, 7 = happy path, 3 = shortcut. Score each option.
|
||||
- **Different work:** For different kinds of work (including mode selection and
|
||||
Add/Defer/Skip or Defer/Keep), write:
|
||||
"Note: options differ in kind, not coverage — no completeness score."
|
||||
No score does not waive approval checkpoints.
|
||||
|
||||
**Pre-question checkpoint:** Validate every field above before saving.
|
||||
Find exactly one row by its assigned ID; verify owner, Current/Proposed, Status
|
||||
@@ -958,8 +965,8 @@ Effort/risk must each be one listed value, never a range. Correct missing or
|
||||
invalid fields and host-limit violations before saving.
|
||||
|
||||
- **Save.** Under the storage policy, save/present the complete current plan,
|
||||
pending rows and comparisons. Copy the grid and all exact fields below,
|
||||
without the illustrative fence delimiters:
|
||||
pending rows and comparisons. Copy the grid and all exact fields below
|
||||
(omit the fence delimiters):
|
||||
|
||||
```text
|
||||
## currentDecision (ROW-ID)
|
||||
@@ -973,13 +980,12 @@ invalid fields and host-limit violations before saving.
|
||||
<full second option description; repeat for all offered options>
|
||||
```
|
||||
|
||||
Replace the whole payload on revision.
|
||||
Keep answered decisions and their answers under separate headings.
|
||||
Replace the whole payload on revision; keep answered decisions under separate headings.
|
||||
- **Read-back.** After the latest successful Write/Edit, Read the ledger row and
|
||||
full payload through the last option's description; fetch continuations.
|
||||
Verify IDs and fields against `currentDecision`, citations against source.
|
||||
Read despite Edit's current-in-context hint. For chat, verify the complete text
|
||||
labeled **not persisted**. A grid, summary or pointer is insufficient.
|
||||
labeled **not persisted**, not a grid, summary or pointer.
|
||||
|
||||
A failed save stops the review. Correct mismatches, save and Read again before dispatch.
|
||||
|
||||
@@ -1000,8 +1006,8 @@ work. A recommendation is not approval; do not edit code.
|
||||
**Post-answer checkpoint:** Save or present the complete amended plan under the
|
||||
storage policy before taking another row.
|
||||
|
||||
If all options are declined, continue only with a viable current approach retained
|
||||
by the answer; otherwise leave the row unresolved and stop for direction.
|
||||
If all options are declined, continue only if the answer retains a viable current
|
||||
approach; otherwise leave the row unresolved and stop for direction.
|
||||
|
||||
Return to the calling step with the saved answer; do not ask it again.
|
||||
Record findings even after resolution; say "No issues, moving on." only with none.
|
||||
@@ -1010,12 +1016,15 @@ Record findings even after resolution; say "No issues, moving on." only with non
|
||||
Follow the preamble's session rules; `CONDUCTOR_SESSION: true` changes transport only.
|
||||
|
||||
1. An explicit choice skips steps 2–3. "Go big", "ambitious" or "cathedral" means SCOPE EXPANSION; "hold scope but tempt me", "show me options" or "cherry-pick" means SELECTIVE EXPANSION.
|
||||
2. Recommend without selecting. Count distinct planned file additions, edits and deletions, labeling estimates. For >15 planned changed files, recommend SCOPE REDUCTION. Otherwise: a new product/system (greenfield) → SCOPE EXPANSION; added capability → SELECTIVE EXPANSION; fix/refactor → HOLD SCOPE. If categories overlap or are unclear, explain why and recommend HOLD SCOPE.
|
||||
2. Recommend without selecting. Count distinct planned file additions, edits and
|
||||
deletions; mark estimated counts as estimates. Apply the first matching rule:
|
||||
- For >15 planned changed files, recommend SCOPE REDUCTION.
|
||||
- If categories overlap or are unclear, explain why and recommend HOLD SCOPE.
|
||||
- Otherwise: a new product/system (greenfield) → SCOPE EXPANSION;
|
||||
added capability → SELECTIVE EXPANSION; fix/refactor → HOLD SCOPE.
|
||||
In the Recommendation's `because` clause, connect a concrete plan fact or
|
||||
constraint to this mode's actual benefit or tradeoff. Count/category alone
|
||||
is not a reason.
|
||||
3. Resolve that recommendation. Mode selection is an admin choice, not a plan
|
||||
decision. When `QUESTION_TUNING: true`, first check `question_id=plan-ceo-review-mode` through the preamble.
|
||||
constraint to this mode's actual benefit or tradeoff, not just its count/category.
|
||||
3. Resolve that recommendation. When `QUESTION_TUNING: true`, first check `question_id=plan-ceo-review-mode` through the preamble.
|
||||
A check that exits 0 with `AUTO_DECIDE` selects the recommendation; go to the automatic handoff in
|
||||
step 4. When tuning is false, omit the lookup.
|
||||
Without that successful check, offer all four modes in one AskUserQuestion,
|
||||
@@ -1033,8 +1042,11 @@ Record mode provenance after the handoff:
|
||||
- **Actual question answer:** question, answer reference and mode; log `auto_decided: false`, including the question ID only when `QUESTION_TUNING: true`.
|
||||
|
||||
If 0D needed no approach choice, say "No new approach decision was needed" after
|
||||
the mode handoff. This records no plan decision, not automatic mode approval.
|
||||
Ask before changing a previously chosen mode.
|
||||
the mode handoff.
|
||||
**Mode change:** Pause and ask with the four-mode menu; keep the mode until
|
||||
answered. If changed, repeat the handoff/provenance record and complete newly
|
||||
applicable Step 0 work in route order, reusing completed work and scope answers.
|
||||
Then resume the paused step. If unchanged, resume directly.
|
||||
|
||||
Selecting a mode does not approve changes. Preserve 0D approvals and ask about
|
||||
each proposed addition or cut, including those prompted by file-count thresholds.
|
||||
@@ -1051,15 +1063,13 @@ Continue to Review Sections, outputs and report.
|
||||
|
||||
### 0F. Expansion Framing (shared by EXPANSION and SELECTIVE EXPANSION)
|
||||
|
||||
Prepare pending candidates for 0G: user experience, concrete addition, S/M/L/XL
|
||||
effort, risk and impact. Explain ambition enthusiastically in SCOPE EXPANSION;
|
||||
balance benefits and tradeoffs without unsupported promises in SELECTIVE
|
||||
EXPANSION. Mark one option `(recommended)` when presenting choices; this label
|
||||
does not approve scope. The user decides each proposal in 0G.
|
||||
Prepare 0G candidates: user experience, addition, S/M/L/XL effort, risk and impact.
|
||||
SCOPE EXPANSION is enthusiastic; SELECTIVE EXPANSION balances benefits and
|
||||
tradeoffs without unsupported promises. Mark one option `(recommended)`;
|
||||
the user still decides each proposal in 0G.
|
||||
|
||||
### 0G. Mode-Specific Analysis
|
||||
In expansion modes, extend 0F's pending list with this analysis, then resolve
|
||||
each proposal individually.
|
||||
In expansion modes, extend 0F's pending list.
|
||||
|
||||
**For SCOPE EXPANSION:**
|
||||
1. **10x check:** Describe 10x value for 2x effort.
|
||||
@@ -1086,24 +1096,24 @@ with the defer/keep menu below; retain the rest.
|
||||
separately per item: **A)** Defer this item to TODOS.md **B)** Keep it in scope.
|
||||
|
||||
Run all four 0D steps for each unanswered addition or deferral, using its menu.
|
||||
These scope choices differ in kind; do not score completeness. Keep other scope
|
||||
fixed or pending; wait for the answer before applying it.
|
||||
Omit completeness scores per 0D. Keep other scope fixed or pending; wait for the
|
||||
answer before applying it.
|
||||
A deferral changes only delivery scope: record its answer/reason beside the prior
|
||||
approval. Keep other approvals and limits unchanged. In later sections, review
|
||||
the retained work and accepted additions; list deferred or rejected work as excluded.
|
||||
approval. Keep other approvals and limits unchanged. Review retained work and
|
||||
accepted additions; exclude deferred or rejected work.
|
||||
|
||||
Save dispositions under the storage policy:
|
||||
- **Add / Keep:** accepted working-plan scope.
|
||||
- **Defer:** TODOS.md with context and NOT in scope with the deferral reason. This postpones work; it does not reject it.
|
||||
- **Skip / Cut:** NOT in scope with the rejection reason; no TODO.
|
||||
|
||||
Reuse answered scope decisions without another question or comparison. Inclusion
|
||||
does not settle pending implementation choices; keep those rows visible.
|
||||
Reuse answered scope decisions without another question or comparison.
|
||||
Implementation choices remain pending until answered.
|
||||
|
||||
### 0H. Persist CEO Plan (EXPANSION and SELECTIVE EXPANSION only)
|
||||
|
||||
Prepare the full amended working plan and a separate CEO scope summary. Keep
|
||||
behavior, requirements and scope consistent; the summary cannot serve as the plan.
|
||||
Prepare the full amended working plan and a separate, consistent CEO scope
|
||||
summary; the summary cannot serve as the plan.
|
||||
|
||||
**Save or present both inputs under the storage policy.** For permitted storage:
|
||||
|
||||
|
||||
@@ -205,15 +205,14 @@ single choice. To expand strategy-only into implementation design, use 0D with
|
||||
**A)** Keep this review strategy-only **B)** Add implementation design for the
|
||||
named capability. Recommend A unless a concrete blocker requires B; wait for the
|
||||
answer. B permits design detail for that capability only.
|
||||
Resolve a choice only when output would be wrong without it, a blocker would be
|
||||
hidden, or scope would change. Reuse prior answers only for the same scope.
|
||||
|
||||
Plain terms:
|
||||
- **Required choice:** a mode, scope, deferral, TODO, spec, outside-review or
|
||||
finding decision needed before the next step.
|
||||
- **Pending:** recorded in the ledger and waiting for approval.
|
||||
- **Settled:** answered by the user, directly instructed, or auto-authorized by
|
||||
the preamble.
|
||||
- **Required choice:** unanswered. Resolve a choice only when continuing would
|
||||
change scope, hide a blocker or produce the wrong output.
|
||||
- **Pending:** unapproved; keep in Proposed, not tasks or accepted work. Status is
|
||||
`unresolved` or `reopened`.
|
||||
- **Settled:** an answer, direct instruction or authorized auto-decision resolves
|
||||
this exact choice and scope; a recommendation does not.
|
||||
|
||||
Review depth controls the detail within each section. Review Sections 1–10 in every depth;
|
||||
run Section 11 only for UI. Strategy-only uses capability-level rows and
|
||||
@@ -221,11 +220,14 @@ run Section 11 only for UI. Strategy-only uses capability-level rows and
|
||||
Implementation-ready names interfaces, codepaths, rescue behavior and tests.
|
||||
For one narrow decision, apply every section to that choice and its dependencies.
|
||||
|
||||
**Keep the stated limits.** Record each measure, value, unit and prerequisite. Count all deliverables, including reused code, as scope; 0E estimates only files that will change. Changing a limit needs evidence and user approval.
|
||||
**Keep the stated limits.** Record each measure, value, unit and prerequisite.
|
||||
Count all deliverables, including reused code, against scope limits. Separately,
|
||||
0E counts changed files, excluding unchanged reuse, to recommend a mode.
|
||||
Neither count approves changes. Changing a limit needs evidence and user approval.
|
||||
|
||||
**Storage policy: choose before writing.** Honor user/host artifact and cleanup
|
||||
limits. One working plan: requested output, else reviewed plan, else host active
|
||||
plan. Use native Write for a missing file and scoped Edit for checkpoints;
|
||||
**Storage policy: choose before writing.** Honor user/host write and cleanup limits.
|
||||
Use one working plan: requested output, else reviewed plan, else host active plan.
|
||||
Use native Write for a missing file and scoped Edit for checkpoints;
|
||||
retain all current content, ledger rows and comparisons.
|
||||
|
||||
**Artifact outcomes:** Never claim an unconfirmed save, read-back or log.
|
||||
@@ -240,9 +242,9 @@ ExitPlanMode or next-skill handoff.
|
||||
| 0H spec-review metrics | Stop with the cause; reviewer availability does not waive this write. |
|
||||
| Review, decision and question history logs | Report cause and unsaved fields; continue. The plan's ledger is still required. |
|
||||
|
||||
Paths are per output: resolve the CEO archive as `CEO_PLANS` in 0H; tasks
|
||||
use `~/.gstack/projects/`, metrics use `~/.gstack/analytics/`, and log helpers
|
||||
choose their own paths. Do not substitute the CEO archive root for these paths.
|
||||
Paths differ: 0H resolves `CEO_PLANS`; tasks use `~/.gstack/projects/`, metrics
|
||||
use `~/.gstack/analytics/`, and log helpers choose their paths. Never substitute
|
||||
the CEO archive root for task, metric or log paths.
|
||||
|
||||
Keep one decision ledger through Step 0, Spec Review Loop and Outside Voice:
|
||||
|
||||
@@ -275,15 +277,17 @@ With no required choice, or after those choices settle, go to 0E.
|
||||
**Choose the question's route first:**
|
||||
- **Admin question:** mode, setup, navigation, document approval or promotion.
|
||||
Use its listed menu and the preamble question transport, then wait and record
|
||||
the answer. Skip steps 1–4; this approves no plan changes. For mode selection,
|
||||
0E defines the four-option menu and any authorized automatic preference;
|
||||
neither needs a plan-decision row, comparison grid or completeness score.
|
||||
the answer. Skip steps 1–4; this approves no plan changes. Resume that menu's
|
||||
next step. 0E owns mode selection; 0H owns document approval. Neither needs
|
||||
a plan-decision row or comparison grid.
|
||||
- **Plan decision:** review-depth expansion, scope additions/cuts, approach
|
||||
choices, TODOs, specs and review/outside findings. Start at step 1. Reuse exact
|
||||
prior approvals; run steps 2–4 only when a new answer is needed, even for one option.
|
||||
|
||||
If an admin answer requests a plan change, use the Plan decision route for that
|
||||
change. 0D never restarts mode selection.
|
||||
change before resuming. 0G proposals and section findings use this route even
|
||||
with prescribed menus. 0D returns to its caller, not to mode selection.
|
||||
For mode changes, follow 0E's **Mode change** instruction.
|
||||
|
||||
**1. Check sources and prior answers.**
|
||||
Compare input, source and answers; correct facts, flag conflicts and preserve unknowns.
|
||||
@@ -313,12 +317,11 @@ Build one `currentDecision` using these fields and the preamble format:
|
||||
| `header` and option labels | Final native text within host limits; exactly one label includes `(recommended)`. |
|
||||
| Each option's `description` | A 1–2 sentence summary; S/M/L/XL effort, low/medium/high risk, reuse, verification coverage, at least 2 ✅ pros and 1 ❌ con. Apply the preamble's minimum lengths and destructive-choice exception. |
|
||||
|
||||
For a plan decision without a prescribed menu, offer 2–3 options (prefer 3 for
|
||||
non-trivial plans). This default does not replace an admin or scope menu.
|
||||
Without a prescribed menu, offer 2–3 options (prefer 3 for non-trivial plans).
|
||||
For an option with no implementation, use effort S and state zero implementation
|
||||
work, never effort 0. Weigh diff size and long-term architecture equally,
|
||||
including rewrites: state the immediate changed-file cost and the future
|
||||
maintenance cost for each option, then explain both in the recommendation.
|
||||
including rewrites: compare immediate changed-file cost and future maintenance
|
||||
cost for each option; explain both in the recommendation.
|
||||
|
||||
In Proposed, compare every commitment in the labels, descriptions and pros/cons:
|
||||
|
||||
@@ -327,12 +330,16 @@ Commitment | Source/approval or pending | Current | A | B | C
|
||||
```
|
||||
|
||||
Include one column per option (add D for a four-option menu). Show unchanged,
|
||||
shared and pending values. Changes remain separate decisions even if they use the same framework.
|
||||
Keep other rows fixed or pending; preserve requirements, tests and fixes.
|
||||
shared and pending values. Keep independent changes separate even within one
|
||||
framework; other rows stay fixed or pending. Preserve requirements, tests and fixes.
|
||||
|
||||
Score this row's coverage differences: 10 = all edge cases, 7 = happy path,
|
||||
3 = shortcut. For different kinds of work, write:
|
||||
"Note: options differ in kind, not coverage — no completeness score."
|
||||
Choose scoring before saving:
|
||||
- **Same work, different coverage:** Score this row's coverage differences:
|
||||
10 = all edge cases, 7 = happy path, 3 = shortcut. Score each option.
|
||||
- **Different work:** For different kinds of work (including mode selection and
|
||||
Add/Defer/Skip or Defer/Keep), write:
|
||||
"Note: options differ in kind, not coverage — no completeness score."
|
||||
No score does not waive approval checkpoints.
|
||||
|
||||
**Pre-question checkpoint:** Validate every field above before saving.
|
||||
Find exactly one row by its assigned ID; verify owner, Current/Proposed, Status
|
||||
@@ -341,8 +348,8 @@ Effort/risk must each be one listed value, never a range. Correct missing or
|
||||
invalid fields and host-limit violations before saving.
|
||||
|
||||
- **Save.** Under the storage policy, save/present the complete current plan,
|
||||
pending rows and comparisons. Copy the grid and all exact fields below,
|
||||
without the illustrative fence delimiters:
|
||||
pending rows and comparisons. Copy the grid and all exact fields below
|
||||
(omit the fence delimiters):
|
||||
|
||||
```text
|
||||
## currentDecision (ROW-ID)
|
||||
@@ -356,13 +363,12 @@ invalid fields and host-limit violations before saving.
|
||||
<full second option description; repeat for all offered options>
|
||||
```
|
||||
|
||||
Replace the whole payload on revision.
|
||||
Keep answered decisions and their answers under separate headings.
|
||||
Replace the whole payload on revision; keep answered decisions under separate headings.
|
||||
- **Read-back.** After the latest successful Write/Edit, Read the ledger row and
|
||||
full payload through the last option's description; fetch continuations.
|
||||
Verify IDs and fields against `currentDecision`, citations against source.
|
||||
Read despite Edit's current-in-context hint. For chat, verify the complete text
|
||||
labeled **not persisted**. A grid, summary or pointer is insufficient.
|
||||
labeled **not persisted**, not a grid, summary or pointer.
|
||||
|
||||
A failed save stops the review. Correct mismatches, save and Read again before dispatch.
|
||||
|
||||
@@ -383,8 +389,8 @@ work. A recommendation is not approval; do not edit code.
|
||||
**Post-answer checkpoint:** Save or present the complete amended plan under the
|
||||
storage policy before taking another row.
|
||||
|
||||
If all options are declined, continue only with a viable current approach retained
|
||||
by the answer; otherwise leave the row unresolved and stop for direction.
|
||||
If all options are declined, continue only if the answer retains a viable current
|
||||
approach; otherwise leave the row unresolved and stop for direction.
|
||||
|
||||
Return to the calling step with the saved answer; do not ask it again.
|
||||
Record findings even after resolution; say "No issues, moving on." only with none.
|
||||
@@ -393,12 +399,15 @@ Record findings even after resolution; say "No issues, moving on." only with non
|
||||
Follow the preamble's session rules; `CONDUCTOR_SESSION: true` changes transport only.
|
||||
|
||||
1. An explicit choice skips steps 2–3. "Go big", "ambitious" or "cathedral" means SCOPE EXPANSION; "hold scope but tempt me", "show me options" or "cherry-pick" means SELECTIVE EXPANSION.
|
||||
2. Recommend without selecting. Count distinct planned file additions, edits and deletions, labeling estimates. For >15 planned changed files, recommend SCOPE REDUCTION. Otherwise: a new product/system (greenfield) → SCOPE EXPANSION; added capability → SELECTIVE EXPANSION; fix/refactor → HOLD SCOPE. If categories overlap or are unclear, explain why and recommend HOLD SCOPE.
|
||||
2. Recommend without selecting. Count distinct planned file additions, edits and
|
||||
deletions; mark estimated counts as estimates. Apply the first matching rule:
|
||||
- For >15 planned changed files, recommend SCOPE REDUCTION.
|
||||
- If categories overlap or are unclear, explain why and recommend HOLD SCOPE.
|
||||
- Otherwise: a new product/system (greenfield) → SCOPE EXPANSION;
|
||||
added capability → SELECTIVE EXPANSION; fix/refactor → HOLD SCOPE.
|
||||
In the Recommendation's `because` clause, connect a concrete plan fact or
|
||||
constraint to this mode's actual benefit or tradeoff. Count/category alone
|
||||
is not a reason.
|
||||
3. Resolve that recommendation. Mode selection is an admin choice, not a plan
|
||||
decision. When `QUESTION_TUNING: true`, first check `question_id=plan-ceo-review-mode` through the preamble.
|
||||
constraint to this mode's actual benefit or tradeoff, not just its count/category.
|
||||
3. Resolve that recommendation. When `QUESTION_TUNING: true`, first check `question_id=plan-ceo-review-mode` through the preamble.
|
||||
A check that exits 0 with `AUTO_DECIDE` selects the recommendation; go to the automatic handoff in
|
||||
step 4. When tuning is false, omit the lookup.
|
||||
Without that successful check, offer all four modes in one AskUserQuestion,
|
||||
@@ -416,8 +425,11 @@ Record mode provenance after the handoff:
|
||||
- **Actual question answer:** question, answer reference and mode; log `auto_decided: false`, including the question ID only when `QUESTION_TUNING: true`.
|
||||
|
||||
If 0D needed no approach choice, say "No new approach decision was needed" after
|
||||
the mode handoff. This records no plan decision, not automatic mode approval.
|
||||
Ask before changing a previously chosen mode.
|
||||
the mode handoff.
|
||||
**Mode change:** Pause and ask with the four-mode menu; keep the mode until
|
||||
answered. If changed, repeat the handoff/provenance record and complete newly
|
||||
applicable Step 0 work in route order, reusing completed work and scope answers.
|
||||
Then resume the paused step. If unchanged, resume directly.
|
||||
|
||||
Selecting a mode does not approve changes. Preserve 0D approvals and ask about
|
||||
each proposed addition or cut, including those prompted by file-count thresholds.
|
||||
@@ -434,15 +446,13 @@ Continue to Review Sections, outputs and report.
|
||||
|
||||
### 0F. Expansion Framing (shared by EXPANSION and SELECTIVE EXPANSION)
|
||||
|
||||
Prepare pending candidates for 0G: user experience, concrete addition, S/M/L/XL
|
||||
effort, risk and impact. Explain ambition enthusiastically in SCOPE EXPANSION;
|
||||
balance benefits and tradeoffs without unsupported promises in SELECTIVE
|
||||
EXPANSION. Mark one option `(recommended)` when presenting choices; this label
|
||||
does not approve scope. The user decides each proposal in 0G.
|
||||
Prepare 0G candidates: user experience, addition, S/M/L/XL effort, risk and impact.
|
||||
SCOPE EXPANSION is enthusiastic; SELECTIVE EXPANSION balances benefits and
|
||||
tradeoffs without unsupported promises. Mark one option `(recommended)`;
|
||||
the user still decides each proposal in 0G.
|
||||
|
||||
### 0G. Mode-Specific Analysis
|
||||
In expansion modes, extend 0F's pending list with this analysis, then resolve
|
||||
each proposal individually.
|
||||
In expansion modes, extend 0F's pending list.
|
||||
|
||||
**For SCOPE EXPANSION:**
|
||||
1. **10x check:** Describe 10x value for 2x effort.
|
||||
@@ -469,24 +479,24 @@ with the defer/keep menu below; retain the rest.
|
||||
separately per item: **A)** Defer this item to TODOS.md **B)** Keep it in scope.
|
||||
|
||||
Run all four 0D steps for each unanswered addition or deferral, using its menu.
|
||||
These scope choices differ in kind; do not score completeness. Keep other scope
|
||||
fixed or pending; wait for the answer before applying it.
|
||||
Omit completeness scores per 0D. Keep other scope fixed or pending; wait for the
|
||||
answer before applying it.
|
||||
A deferral changes only delivery scope: record its answer/reason beside the prior
|
||||
approval. Keep other approvals and limits unchanged. In later sections, review
|
||||
the retained work and accepted additions; list deferred or rejected work as excluded.
|
||||
approval. Keep other approvals and limits unchanged. Review retained work and
|
||||
accepted additions; exclude deferred or rejected work.
|
||||
|
||||
Save dispositions under the storage policy:
|
||||
- **Add / Keep:** accepted working-plan scope.
|
||||
- **Defer:** TODOS.md with context and NOT in scope with the deferral reason. This postpones work; it does not reject it.
|
||||
- **Skip / Cut:** NOT in scope with the rejection reason; no TODO.
|
||||
|
||||
Reuse answered scope decisions without another question or comparison. Inclusion
|
||||
does not settle pending implementation choices; keep those rows visible.
|
||||
Reuse answered scope decisions without another question or comparison.
|
||||
Implementation choices remain pending until answered.
|
||||
|
||||
### 0H. Persist CEO Plan (EXPANSION and SELECTIVE EXPANSION only)
|
||||
|
||||
Prepare the full amended working plan and a separate CEO scope summary. Keep
|
||||
behavior, requirements and scope consistent; the summary cannot serve as the plan.
|
||||
Prepare the full amended working plan and a separate, consistent CEO scope
|
||||
summary; the summary cannot serve as the plan.
|
||||
|
||||
**Save or present both inputs under the storage policy.** For permitted storage:
|
||||
|
||||
|
||||
@@ -31,28 +31,27 @@ Carry prior approvals into findings, tasks and the report. Routine auto-decide
|
||||
cannot override user constraints or non-goals.
|
||||
|
||||
## CRITICAL RULE — How to ask questions
|
||||
Follow the AskUserQuestion format from the Preamble above. Additional rules for plan reviews:
|
||||
Use 0D's decision procedure and the preamble's AskUserQuestion format:
|
||||
* **One decision unit = one AskUserQuestion call.** Use Step 0D boundaries, not topic labels.
|
||||
* Describe the problem concretely, with file and line references.
|
||||
* Present 2-3 options, including "do nothing" where reasonable.
|
||||
* For each option: effort, risk, and maintenance burden in one line.
|
||||
* Before calling AskUserQuestion, draft the recommended option as a complete remedy
|
||||
for this one issue. Its offered description must state the rescue behavior,
|
||||
verification, and failure visibility needed for that fix. Include those details
|
||||
in the option itself. Omit irrelevant work, and keep independent findings and
|
||||
new TODOs in their own questions.
|
||||
* **Map the reasoning to my engineering preferences above.** One sentence connecting your recommendation to a specific preference.
|
||||
* Use the preamble's `D<N>` question heading and A/B/C option labels. Cite the stable ledger ID separately so a reopened question keeps its earlier decision history.
|
||||
* Describe the concrete problem with file/line references. Offer 2-3 options,
|
||||
including "do nothing" when reasonable.
|
||||
* Give each option one line covering effort, risk and maintenance.
|
||||
* The recommended option's description must offer a complete remedy for this
|
||||
issue: rescue behavior, verification and failure visibility. Exclude unrelated
|
||||
work; ask about independent findings and new TODOs separately.
|
||||
* Connect the recommendation to one engineering preference in a sentence.
|
||||
* Use `D<N>` and A/B/C labels. Cite the stable ledger ID separately to retain
|
||||
reopened decision history.
|
||||
* An "obvious fix" still needs approval when it is not covered by an exact accepted choice.
|
||||
|
||||
## Formatting Rules
|
||||
* Keep option labels short; use Step 0D's exact `currentDecision` fields for the question and option descriptions.
|
||||
* Use short labels and 0D's exact `currentDecision` question and option descriptions.
|
||||
* Use **CRITICAL GAP** / **WARNING** / **OK** for scannability.
|
||||
|
||||
## Mode Quick Reference
|
||||
|
||||
The mode changes which work is included, not review depth or section coverage.
|
||||
Apply the review and outputs to the accepted work in every mode.
|
||||
Mode controls included work, not depth or section coverage. Review and produce
|
||||
outputs for accepted work in every mode.
|
||||
|
||||
| Step | SCOPE EXPANSION | SELECTIVE EXPANSION | HOLD SCOPE | SCOPE REDUCTION |
|
||||
|------|-----------------|---------------------|------------|-----------------|
|
||||
@@ -66,9 +65,9 @@ Apply the review and outputs to the accepted work in every mode.
|
||||
| Future direction (Section 10) | Review accepted trajectory | Review accepted cherry-picks | Maintainability; no expansions | Maintainability of remaining scope |
|
||||
| Design (Section 11) | Review if UI scope | Review if UI scope | Review if UI scope | Review if UI scope |
|
||||
|
||||
All modes produce the review content. Save it to the permitted working plan;
|
||||
when no plan/report write is permitted, present it in chat as not persisted and
|
||||
end with completion blocked. The CEO archive is additional expansion-mode output.
|
||||
Save to the permitted working plan; with no permitted plan/report write, present
|
||||
it in chat as not persisted and end with completion blocked. The CEO archive is
|
||||
additional expansion-mode output.
|
||||
|
||||
### Working review decisions
|
||||
|
||||
@@ -82,14 +81,18 @@ and mitigations even if later text omits them. Flag approval conflicts. Unavaila
|
||||
code proves neither failure nor safety; record unknown risks with their owners
|
||||
and required verification.
|
||||
|
||||
**Resolve.** If this section needs a new decision or evidence warrants reopening
|
||||
one, complete 0D through its post-answer save, then continue to Apply below.
|
||||
Use the same row ID in the ledger, `currentDecision` and question; complete 0D's
|
||||
pre-question checkpoint before each new or reopened question.
|
||||
If all choices are settled, cite their exact answers and go straight to Apply.
|
||||
Resolve critical risks now. Reference other pending rows in their owner sections;
|
||||
do not decide them here. Keep independent safety fixes and throughput improvements
|
||||
in separate rows, following 0D's test table.
|
||||
**Resolve.** Take the first applicable path for each finding:
|
||||
1. This section needs a new choice, or evidence warrants reopening its prior
|
||||
answer: use 0D's Plan decision route through its post-answer save, then
|
||||
return here to Apply.
|
||||
Use the same row ID in the ledger, `currentDecision` and question; complete
|
||||
the pre-question checkpoint before asking. Resolve critical risks now.
|
||||
2. An exact prior answer covers it: cite that answer and go to Apply.
|
||||
3. A non-blocking choice belongs to a later section: reference its pending row
|
||||
and owner; leave it undecided here.
|
||||
|
||||
Keep independent safety fixes and throughput improvements in separate rows,
|
||||
following 0D's test table. No path selects the mode again.
|
||||
|
||||
**Apply.** Check the saved plan against each answer's exact scope. Preserve existing
|
||||
content, approved behavior, required implementation, tests and success/failure
|
||||
@@ -104,7 +107,14 @@ review or no-UI skip, follow Closing sequence. Keep unresolved choices in the
|
||||
ledger and report; an approval is not proof of implementation or verification.
|
||||
|
||||
### Section 1: Architecture Review
|
||||
Publish **Current scope** in chat using the Step 0E mode-handoff format and the current ledger dispositions, including actual later scope-answer references. Retain mode, rationale and preference attribution. This updates scope after 0G; do not ask or log the mode again. Keep earlier answers as history, showing current accepted scope. Then say `Section 1: Architecture Review`.
|
||||
Publish **Current scope** in chat before the architecture analysis:
|
||||
- Retain 0E's selected mode, rationale and preference attribution.
|
||||
- Show each governing row's ID, disposition and answer reference, including scope
|
||||
decisions after 0E. Keep earlier answers as history.
|
||||
- Distinguish accepted, deferred, rejected and pending work.
|
||||
|
||||
This is a scope update, not another mode handoff; do not ask or log the mode again.
|
||||
Then say `Section 1: Architecture Review`.
|
||||
|
||||
Evaluate and diagram:
|
||||
* System design and component boundaries. Draw the dependency graph.
|
||||
@@ -353,11 +363,6 @@ _CODEX_CFG=$(~/.claude/skills/gstack/bin/gstack-config get codex_reviews 2>/dev/
|
||||
source ~/.claude/skills/gstack/bin/gstack-codex-probe 2>/dev/null || true
|
||||
if [ "$_CODEX_CFG" = "disabled" ]; then
|
||||
_CODEX_MODE="disabled"
|
||||
# Running-under-Codex presence probe (#2519): a live Codex session exports
|
||||
# CODEX_THREAD_ID / CODEX_SANDBOX into every shell it spawns (verified
|
||||
# against a live `codex exec 'env | grep -i codex'` capture, codex 0.147.0).
|
||||
# Nested codex spawns from inside a Codex host multiply token burn
|
||||
# (observed: one /review = 15M tokens). A stale own-harness artifact must stop.
|
||||
elif { [ -n "${CODEX_THREAD_ID:-}" ] || [ -n "${CODEX_SANDBOX:-}" ] || [ "${GSTACK_ACTIVE_HOST:-}" = codex ]; }; then
|
||||
_CODEX_MODE="under_codex"
|
||||
elif ! command -v codex >/dev/null 2>&1; then
|
||||
@@ -381,11 +386,11 @@ echo "CODEX_MODE: $_CODEX_MODE"
|
||||
|
||||
Branch on the echoed `CODEX_MODE`:
|
||||
- **`disabled`** — the user turned Codex reviews off (`codex_reviews=disabled`). Skip the reviewer invocation; record disabled coverage as directed below; do NOT fall back to a Claude subagent — disabled means no extra review step. Print: "Codex review skipped (codex_reviews disabled). Re-enable: `gstack-config set codex_reviews enabled`."
|
||||
- **`not_installed`** — Codex CLI absent. Print: "Codex not installed — falling back to a Claude subagent (fresh context, but the same harness; model identity is unknown). Install Codex for an actual outside-model read: `npm install -g @openai/codex`." Fall back to the Claude subagent path.
|
||||
- **`not_installed`** — Codex CLI absent. Print: "Codex not installed; outside coverage unavailable. Install: `npm install -g @openai/codex`." Fall back to the Claude subagent path.
|
||||
- **`under_codex`** — stale artifact selected its own harness. Print: "Codex outside review unavailable: harness mismatch; no outside process started. Missing coverage. Repair: setup --host codex." Skip the outside invocation and construct the prompt below, then follow **Native fallback**. Conflicting inherited harness markers are not grounds to guess another provider.
|
||||
- **`not_authed`** — installed but no credentials. Print: "Codex installed but not authenticated — falling back to a Claude subagent (same harness; model identity is unknown). Run `codex login` or set `$CODEX_API_KEY`." Fall back to the Claude subagent path.
|
||||
- **`broken_install`** — the CLI is on PATH but cannot execute (spawn ENOENT, non-executable binary, missing vendor payload). Print: "Codex is installed but its binary cannot run — Codex passes skipped. Reinstall: `npm install -g @openai/codex`." Relay the probe's HINT lines and fall back to the Claude subagent path. This state exists because a missing binary used to land in the model probe's fail-open bucket and report `ready`, so every Codex pass was skipped silently (#2742).
|
||||
- **`model_unusable`** — authed but the account cannot use gstack's selected Codex model (#2477: HTTP 400 on every call). Relay the probe's HINT lines, tell the user the one-line fix (set `GSTACK_CODEX_MODEL=<supported-model>` or pass an explicit `-c model=...` override), and fall back to the Claude subagent path. The ~10s round trip is cached for 1h; timeouts fail open to `ready`.
|
||||
- **`not_authed`** — installed but no credentials. Print: "Codex not authenticated; outside coverage unavailable. Run `codex login` or set `$CODEX_API_KEY`." Fall back to the Claude subagent path.
|
||||
- **`broken_install`** — the CLI is on PATH but cannot execute (spawn ENOENT, non-executable binary, missing vendor payload). Print: "Codex is installed but its binary cannot run — Codex passes skipped. Reinstall: `npm install -g @openai/codex`." Relay the probe's HINT lines. Fall back to the Claude subagent path.
|
||||
- **`model_unusable`** — authed but the account cannot use gstack's selected Codex model (#2477: HTTP 400 on every call). Relay the probe's HINT lines and tell the user the one-line fix (set `GSTACK_CODEX_MODEL=<supported-model>` or pass an explicit `-c model=...` override). Fall back to the Claude subagent path. The ~10s round trip is cached for 1h; timeouts fail open to `ready`.
|
||||
- **`ready`** — run the Codex pass below.
|
||||
|
||||
**Outcome routing:** Follow the row for the current result. After an invocation, route its result
|
||||
@@ -821,17 +826,18 @@ this run (an empty file means "ran, no findings" — distinct from "didn't run")
|
||||
|
||||
|
||||
### Completion Summary
|
||||
Fill this template from Review facts now, as part of the plan body. Artifact
|
||||
outcomes remain pending until their writes are confirmed. Stage 3 publishes it
|
||||
after report verification; forbidden writes stay labeled not persisted.
|
||||
Fill this plan-body template from Review facts. Artifact outcomes stay pending
|
||||
until writes are confirmed. Stage 3 publishes it after report verification;
|
||||
forbidden writes stay labeled not persisted.
|
||||
|
||||
Use the full mode name from Step 0E; replace spaces with underscores only in the
|
||||
review log's `MODE` field. "System Audit" summarizes repository findings from
|
||||
Step 0 and the review sections. "Lake Score" counts complete options selected:
|
||||
Y is the number of answered coverage questions offering a 10/10 option; X is
|
||||
how many selected that option. Count a reopened choice only once, using its
|
||||
latest answered option; superseded answers add nothing. Exclude kind-only and
|
||||
unanswered questions; use `N/A` when Y is zero.
|
||||
Step 0 and the review sections. Compute "Lake Score" (complete options selected):
|
||||
1. Select answered questions scored for coverage under 0D that offered a 10/10
|
||||
option. Exclude unscored mode/scope choices and unanswered questions.
|
||||
2. Count a reopened choice only once, using its latest answered option.
|
||||
3. Y is the number of eligible questions; X is how many selected the 10/10
|
||||
option. Report X/Y, or `N/A` when Y is zero.
|
||||
|
||||
```
|
||||
+====================================================================+
|
||||
@@ -1049,17 +1055,69 @@ After completing the review, read the review log and config to display the dashb
|
||||
~/.claude/skills/gstack/bin/gstack-review-read
|
||||
```
|
||||
|
||||
Render each record using its recorded host, source, outside_provider, outside_status, and phase. Historical source "claude" means a native Claude subagent; source "claude-code" means the external CLI. Never infer a historical provider from the current harness. Unknown model identity remains unknown. Missing/disabled/skipped outside coverage is distinct from native completion.
|
||||
**1. Choose the records to display.** Use the latest record for each row below.
|
||||
Do not use a record older than 7 days to clear a row, and never substitute an older
|
||||
success for a newer failure. Ship metrics are not review records.
|
||||
|
||||
Parse the output. Find the most recent entry for each skill (plan-ceo-review, plan-eng-review, review, plan-design-review, design-review-lite, adversarial-review, codex-review, codex-plan-review). Ignore entries with timestamps older than 7 days. For the Eng Review row, show whichever is more recent between `review` (diff-scoped pre-landing review) and `plan-eng-review` (plan-stage architecture review). Append "(DIFF)" or "(PLAN)" to the status to distinguish. For the Adversarial row, show whichever is more recent between `adversarial-review` (new auto-scaled) and `codex-review` (legacy). For Design Review, show whichever is more recent between `plan-design-review` (full visual audit) and `design-review-lite` (code-level check). Append "(FULL)" or "(LITE)" to the status to distinguish. For the Outside Voice row, show the most recent `codex-plan-review` entry — this captures outside voices from both /plan-ceo-review and /plan-eng-review.
|
||||
| Row | Choose the latest of | Status suffix |
|
||||
|---|---|---|
|
||||
| Eng Review | `review` or `plan-eng-review` | (DIFF) or (PLAN) |
|
||||
| CEO Review | `plan-ceo-review` | — |
|
||||
| Design Review | `plan-design-review` or `design-review-lite` | (FULL) or (LITE) |
|
||||
| Adversarial | `adversarial-review` or legacy `codex-review` | — |
|
||||
| Outside Voice | `codex-plan-review` from CEO or Eng review | — |
|
||||
|
||||
**Source attribution:** If the most recent entry for a skill has a \`"via"\` field, append it to the status label in parentheses. Examples: `plan-eng-review` with `via:"autoplan"` shows as "CLEAR (PLAN via /autoplan)". `review` with `via:"ship"` shows as "CLEAR (DIFF via /ship)". Entries without a `via` field show as "CLEAR (PLAN)" or "CLEAR (DIFF)" as before.
|
||||
Keep each record's host, source, outside_provider, outside_status and phase.
|
||||
Historical source "claude" is a native subagent; "claude-code" is the external CLI.
|
||||
Do not infer old providers or unknown models from today's harness. A native result
|
||||
does not fill missing, disabled or skipped outside coverage.
|
||||
|
||||
From gstack-review-read output, use entries whose skill is `autoplan-voices` or `design-outside-voices` for the coverage detail below the dashboard. Group by workflow run and phase, not merely skill. Show each phase’s recorded provider and outside_status; partial coverage must remain partial. These records do not change the engineering gate.
|
||||
**Source attribution:** Append a recorded `via` to the suffix, for example
|
||||
"CLEAR (PLAN via /autoplan)" or "CLEAR (DIFF via /ship)". Without `via`, keep
|
||||
"CLEAR (PLAN)" or "CLEAR (DIFF)". Below the dashboard, group `autoplan-voices`
|
||||
and `design-outside-voices` by workflow run and phase. Show each phase's provider
|
||||
and outside_status; retain partial coverage. These details do not clear Eng Review.
|
||||
|
||||
Display a fresh `clean` result as CLEAR and `issues_open` as ISSUES OPEN. Show missing, stale, disabled or unavailable results explicitly; none implies CLEAR. Keep the logged status unchanged.
|
||||
**2. Check freshness before choosing a verdict.**
|
||||
|
||||
Display:
|
||||
- **Content-first rule:** For `review`, `adversarial-review`, `codex-review`,
|
||||
ship-stage reviews and `design-review-lite`, use `review_freshness.status`
|
||||
and show its `reason`. CURRENT means a completed clean review whose start and
|
||||
end content fingerprints equal the current `---WTREE---` fingerprint. This
|
||||
fingerprint covers working-tree content, not just the commit.
|
||||
STALE or UNVERIFIED cannot clear Eng Review. Missing `review_freshness`,
|
||||
including legacy log-only records, means UNVERIFIED. Never fall back to HEAD
|
||||
equality or commit distance for diff evidence, even at zero commits.
|
||||
Show recorded cycles, completed/converged fields and missing source/phase
|
||||
coverage. Unknown coverage is not a pass.
|
||||
- **Plan records** (plan-ceo-review, plan-eng-review, plan-design-review and
|
||||
codex-plan-review) use the 7-day window, not the working-tree fingerprint.
|
||||
If `plan_sha256` is present, you may compare the plan file and report a mismatch.
|
||||
For plan records only, compare the recorded commit with `---HEAD---`.
|
||||
If different, run `git rev-list --count STORED_COMMIT..HEAD` and report
|
||||
"Note: {skill} review from {date} may be stale — {N} commits since review".
|
||||
A failed command means UNKNOWN, treated as stale. Without commit tracking,
|
||||
retain the note to consider re-running. Omit staleness notes when all reviews
|
||||
are current.
|
||||
|
||||
**3. Choose the historical verdict.** CLEARED requires the selected Eng Review
|
||||
to be `clean`, within 7 days and fresh under step 2. Otherwise report NOT CLEARED
|
||||
and its missing, stale or open-issue reason. If `skip_eng_review` is true, show
|
||||
"SKIPPED (global)" for Eng Review and CLEARED for this dashboard.
|
||||
Eng Review is required by default; `gstack-config set skip_eng_review true` disables that requirement.
|
||||
|
||||
Other rows provide context, not a substitute for Eng Review:
|
||||
- Recommend CEO Review for product/business or scope decisions, not routine fixes or cleanup.
|
||||
- Recommend Design Review for UI/UX work, not backend, infrastructure or prompt-only work.
|
||||
- Adversarial review always includes a native pass. Available, enabled outside
|
||||
challenges supplement it; diffs of 200+ lines also get the structured P1 gate.
|
||||
- Outside Voice is the default-on plan review after CEO/Eng review. `codex_reviews`
|
||||
disables that extra step. Provider failure uses native fallback and records
|
||||
missing outside coverage; this dashboard row never gates shipping.
|
||||
|
||||
**4. Display the dashboard.** Show missing, stale, disabled or unavailable results
|
||||
explicitly, never as CLEAR. Display a fresh `clean` result as CLEAR and
|
||||
`issues_open` as ISSUES OPEN without changing the stored status.
|
||||
|
||||
```
|
||||
+====================================================================+
|
||||
@@ -1077,26 +1135,6 @@ Display:
|
||||
+====================================================================+
|
||||
```
|
||||
|
||||
**Review tiers:**
|
||||
- **Eng Review (required by default):** The only review that gates shipping. Covers architecture, code quality, tests, performance. Can be disabled globally with \`gstack-config set skip_eng_review true\` (the "don't bother me" setting).
|
||||
- **CEO Review (optional):** Use your judgment. Recommend it for big product/business changes, new user-facing features, or scope decisions. Skip for bug fixes, refactors, infra, and cleanup.
|
||||
- **Design Review (optional):** Use your judgment. Recommend it for UI/UX changes. Skip for backend-only, infra, or prompt-only changes.
|
||||
- **Adversarial Review (automatic):** Always-on for every review. Every diff gets a native adversarial pass and, when enabled and available, a host-selected outside challenge. Large diffs (200+ lines) additionally get a structured outside review with P1 gate.
|
||||
- **Outside Voice (default-on):** Independent plan review through the host-selected provider after /plan-ceo-review and /plan-eng-review. The codex_reviews switch disables the entire extra step. Provider failure uses the existing native fallback and reports missing outside coverage. Never gates shipping.
|
||||
|
||||
**Verdict logic:**
|
||||
- **CLEARED**: Eng Review has >= 1 entry within 7 days from either \`review\` or \`plan-eng-review\` with status "clean"; diff review must also grade CURRENT below (or \`skip_eng_review\` is \`true\`)
|
||||
- **NOT CLEARED**: Eng Review missing, stale (>7 days), or has open issues
|
||||
- CEO, Design, and outside reviews are shown for context but never block shipping
|
||||
- If \`skip_eng_review\` config is \`true\`, Eng Review shows "SKIPPED (global)" and verdict is CLEARED
|
||||
|
||||
**Staleness detection:** Grade before deciding CLEARED:
|
||||
- Ship telemetry reports metrics, not review coverage; it never satisfies a review row.
|
||||
- **Content-first rule (diff-scoped rows only: `review`, `adversarial-review`, `codex-review`, ship-stage entries, `design-review-lite`).** Use the helper's computed `review_freshness.status` and show its `reason`. CURRENT requires a completed clean pass with captured start/end wtree equal to the current `---WTREE---`. STALE or UNVERIFIED never clears Eng Review. Missing `review_freshness` is UNVERIFIED, including legacy log-only rows. Never fall back to HEAD equality or commit distance for diff evidence, even at 0 commits. Show recorded cycles, completed/converged state, and missing per-source/phase coverage; unknown is not a pass.
|
||||
- Plan-tier rows (plan-ceo-review, plan-eng-review, plan-design-review, codex-plan-review) grade a plan file, not the repo tree — never apply the wtree rule to them; they keep the 7-day freshness logic. If an entry carries `plan_sha256`, you MAY compare it with the plan file and note "plan changed since review" on mismatch.
|
||||
- Plan-tier fallback only: parse `---HEAD---`. For entries with a different `commit`, count elapsed commits: `git rev-list --count STORED_COMMIT..HEAD`. If that command FAILS, grade UNKNOWN and treat as stale. Display: "Note: {skill} review from {date} may be stale — {N} commits since review". Missing commit tracking retains the legacy note to consider re-running.
|
||||
- If all reviews grade CURRENT, do not display staleness notes
|
||||
|
||||
## Next Steps — Review Chaining
|
||||
|
||||
After displaying the Review Readiness Dashboard, recommend the next review(s) based on what this CEO review discovered. Read the dashboard output to see which reviews have already been run and whether they are stale.
|
||||
|
||||
@@ -29,28 +29,27 @@ Carry prior approvals into findings, tasks and the report. Routine auto-decide
|
||||
cannot override user constraints or non-goals.
|
||||
|
||||
## CRITICAL RULE — How to ask questions
|
||||
Follow the AskUserQuestion format from the Preamble above. Additional rules for plan reviews:
|
||||
Use 0D's decision procedure and the preamble's AskUserQuestion format:
|
||||
* **One decision unit = one AskUserQuestion call.** Use Step 0D boundaries, not topic labels.
|
||||
* Describe the problem concretely, with file and line references.
|
||||
* Present 2-3 options, including "do nothing" where reasonable.
|
||||
* For each option: effort, risk, and maintenance burden in one line.
|
||||
* Before calling AskUserQuestion, draft the recommended option as a complete remedy
|
||||
for this one issue. Its offered description must state the rescue behavior,
|
||||
verification, and failure visibility needed for that fix. Include those details
|
||||
in the option itself. Omit irrelevant work, and keep independent findings and
|
||||
new TODOs in their own questions.
|
||||
* **Map the reasoning to my engineering preferences above.** One sentence connecting your recommendation to a specific preference.
|
||||
* Use the preamble's `D<N>` question heading and A/B/C option labels. Cite the stable ledger ID separately so a reopened question keeps its earlier decision history.
|
||||
* Describe the concrete problem with file/line references. Offer 2-3 options,
|
||||
including "do nothing" when reasonable.
|
||||
* Give each option one line covering effort, risk and maintenance.
|
||||
* The recommended option's description must offer a complete remedy for this
|
||||
issue: rescue behavior, verification and failure visibility. Exclude unrelated
|
||||
work; ask about independent findings and new TODOs separately.
|
||||
* Connect the recommendation to one engineering preference in a sentence.
|
||||
* Use `D<N>` and A/B/C labels. Cite the stable ledger ID separately to retain
|
||||
reopened decision history.
|
||||
* An "obvious fix" still needs approval when it is not covered by an exact accepted choice.
|
||||
|
||||
## Formatting Rules
|
||||
* Keep option labels short; use Step 0D's exact `currentDecision` fields for the question and option descriptions.
|
||||
* Use short labels and 0D's exact `currentDecision` question and option descriptions.
|
||||
* Use **CRITICAL GAP** / **WARNING** / **OK** for scannability.
|
||||
|
||||
## Mode Quick Reference
|
||||
|
||||
The mode changes which work is included, not review depth or section coverage.
|
||||
Apply the review and outputs to the accepted work in every mode.
|
||||
Mode controls included work, not depth or section coverage. Review and produce
|
||||
outputs for accepted work in every mode.
|
||||
|
||||
| Step | SCOPE EXPANSION | SELECTIVE EXPANSION | HOLD SCOPE | SCOPE REDUCTION |
|
||||
|------|-----------------|---------------------|------------|-----------------|
|
||||
@@ -64,9 +63,9 @@ Apply the review and outputs to the accepted work in every mode.
|
||||
| Future direction (Section 10) | Review accepted trajectory | Review accepted cherry-picks | Maintainability; no expansions | Maintainability of remaining scope |
|
||||
| Design (Section 11) | Review if UI scope | Review if UI scope | Review if UI scope | Review if UI scope |
|
||||
|
||||
All modes produce the review content. Save it to the permitted working plan;
|
||||
when no plan/report write is permitted, present it in chat as not persisted and
|
||||
end with completion blocked. The CEO archive is additional expansion-mode output.
|
||||
Save to the permitted working plan; with no permitted plan/report write, present
|
||||
it in chat as not persisted and end with completion blocked. The CEO archive is
|
||||
additional expansion-mode output.
|
||||
|
||||
### Working review decisions
|
||||
|
||||
@@ -80,14 +79,18 @@ and mitigations even if later text omits them. Flag approval conflicts. Unavaila
|
||||
code proves neither failure nor safety; record unknown risks with their owners
|
||||
and required verification.
|
||||
|
||||
**Resolve.** If this section needs a new decision or evidence warrants reopening
|
||||
one, complete 0D through its post-answer save, then continue to Apply below.
|
||||
Use the same row ID in the ledger, `currentDecision` and question; complete 0D's
|
||||
pre-question checkpoint before each new or reopened question.
|
||||
If all choices are settled, cite their exact answers and go straight to Apply.
|
||||
Resolve critical risks now. Reference other pending rows in their owner sections;
|
||||
do not decide them here. Keep independent safety fixes and throughput improvements
|
||||
in separate rows, following 0D's test table.
|
||||
**Resolve.** Take the first applicable path for each finding:
|
||||
1. This section needs a new choice, or evidence warrants reopening its prior
|
||||
answer: use 0D's Plan decision route through its post-answer save, then
|
||||
return here to Apply.
|
||||
Use the same row ID in the ledger, `currentDecision` and question; complete
|
||||
the pre-question checkpoint before asking. Resolve critical risks now.
|
||||
2. An exact prior answer covers it: cite that answer and go to Apply.
|
||||
3. A non-blocking choice belongs to a later section: reference its pending row
|
||||
and owner; leave it undecided here.
|
||||
|
||||
Keep independent safety fixes and throughput improvements in separate rows,
|
||||
following 0D's test table. No path selects the mode again.
|
||||
|
||||
**Apply.** Check the saved plan against each answer's exact scope. Preserve existing
|
||||
content, approved behavior, required implementation, tests and success/failure
|
||||
@@ -102,7 +105,14 @@ review or no-UI skip, follow Closing sequence. Keep unresolved choices in the
|
||||
ledger and report; an approval is not proof of implementation or verification.
|
||||
|
||||
### Section 1: Architecture Review
|
||||
Publish **Current scope** in chat using the Step 0E mode-handoff format and the current ledger dispositions, including actual later scope-answer references. Retain mode, rationale and preference attribution. This updates scope after 0G; do not ask or log the mode again. Keep earlier answers as history, showing current accepted scope. Then say `Section 1: Architecture Review`.
|
||||
Publish **Current scope** in chat before the architecture analysis:
|
||||
- Retain 0E's selected mode, rationale and preference attribution.
|
||||
- Show each governing row's ID, disposition and answer reference, including scope
|
||||
decisions after 0E. Keep earlier answers as history.
|
||||
- Distinguish accepted, deferred, rejected and pending work.
|
||||
|
||||
This is a scope update, not another mode handoff; do not ask or log the mode again.
|
||||
Then say `Section 1: Architecture Review`.
|
||||
|
||||
Evaluate and diagram:
|
||||
* System design and component boundaries. Draw the dependency graph.
|
||||
@@ -443,17 +453,18 @@ List every ASCII diagram in files this plan touches. Still accurate?
|
||||
{{TASKS_SECTION_EMIT:ceo-review}}
|
||||
|
||||
### Completion Summary
|
||||
Fill this template from Review facts now, as part of the plan body. Artifact
|
||||
outcomes remain pending until their writes are confirmed. Stage 3 publishes it
|
||||
after report verification; forbidden writes stay labeled not persisted.
|
||||
Fill this plan-body template from Review facts. Artifact outcomes stay pending
|
||||
until writes are confirmed. Stage 3 publishes it after report verification;
|
||||
forbidden writes stay labeled not persisted.
|
||||
|
||||
Use the full mode name from Step 0E; replace spaces with underscores only in the
|
||||
review log's `MODE` field. "System Audit" summarizes repository findings from
|
||||
Step 0 and the review sections. "Lake Score" counts complete options selected:
|
||||
Y is the number of answered coverage questions offering a 10/10 option; X is
|
||||
how many selected that option. Count a reopened choice only once, using its
|
||||
latest answered option; superseded answers add nothing. Exclude kind-only and
|
||||
unanswered questions; use `N/A` when Y is zero.
|
||||
Step 0 and the review sections. Compute "Lake Score" (complete options selected):
|
||||
1. Select answered questions scored for coverage under 0D that offered a 10/10
|
||||
option. Exclude unscored mode/scope choices and unanswered questions.
|
||||
2. Count a reopened choice only once, using its latest answered option.
|
||||
3. Y is the number of eligible questions; X is how many selected the 10/10
|
||||
option. Report X/Y, or `N/A` when Y is zero.
|
||||
|
||||
```
|
||||
+====================================================================+
|
||||
|
||||
Reference in new issue
Block a user