Files
gstack/test/fixtures/autoplan/u-ceo-original-loss.json
Garry TanandOpenAI Codex 9f81911136 v1.86.0.0 feat: route outside reviews by harness (#2850)
* feat: add a restricted and supervised Claude Code runner

Preserve configured authentication and models while enforcing tool access, strict completion JSON, bounded output and process cleanup. Cover argv, failure handling, session metadata and Windows process containment.

* feat: route outside reviews by harness and migrate wrapper installs

Use Claude Code from Codex and Codex from other supported hosts, with shared invocation rendering, positive gate validation and per-phase provenance. Rename /claude to /claude-code, repair managed shared and copied installations safely, and generate native Kiro skills. Add installed-workflow, failure-injection and live cross-harness regression coverage.

* test: recognize CEO mode labels without terminal spacing

The paid workflow rendered SCOPEEXPANSION at option 4, but its driver required a literal space. Match the leading mode title without cursor-spacing artifacts and ignore adjacent preview text. Preserve missing-target failures and downstream posture assertions.

* test: isolate plan-count fixtures before starting review workflows

Seed the complete test plan in a private git repository before launching Claude, so a bare slash command cannot review the live workspace while a delayed fixture message remains queued. Preserve count thresholds, parsers and budgets. Add initial-context and installed-discovery tests, and retain startup/terminal diagnostics on failed evaluations.

* test: stabilize review fixtures and Claude eval startup

Preserve source boundaries in workflow judge inputs, isolate CEO mode plans, and wait for interactive trust input readiness. Keep startup failure evidence and retain existing models, budgets, and assertions.

Co-Authored-By: OpenAI Codex <noreply@openai.com>

* test: classify collapsed review modes and isolate seeded findings

Keep review questions out of the setup count when terminal cursor positioning removes spaces. State existing webhook safeguards so the five-finding control measures its seeded defects without accidental extra security and concurrency gaps. Preserve question bands and the paired control.

Co-Authored-By: OpenAI Codex <noreply@openai.com>

* test: isolate browser daemon state across free shards

Co-Authored-By: OpenAI Codex <noreply@openai.com>

* test: stabilize native review counting and interactive navigation

Co-Authored-By: OpenAI Codex <noreply@openai.com>

* chore: prepare v1.82.0.0 release

Co-Authored-By: OpenAI Codex <noreply@openai.com>

* fix: eliminate browser and process-cleanup test flakes

Pin every CI surface to Bun 1.4.0 to avoid extra-stdio finalizers closing
reused live sockets. Add an isolated GC/listener regression that fails on
Bun 1.3.13, and prevent coordinated rollback to an affected CI runtime.

Check renderer cleanup against the render's own staging directory so
concurrent renders cannot invalidate the assertion. Make the no-pgrep
process-tree walk tolerate disappearing /proc entries, and synchronize
its test fixture through child readiness and pipe EOF instead of sleeps.

Validation: 9,157 passed, 31 skipped, zero failures across 556 files with
retries disabled. Build, all-host generation freshness, and skill checks
passed. All three races have failing-before/passing-after regressions.

* fix: count completed native review questions in evals

* fix: drive review navigation from confirmed native choices

* fix: require complete section-loading eval reports

* test: isolate telemetry HTTP transport from local assertions

* fix: keep review input on the active native question

* test: let tunnel revocation daemon choose an available port

* test: allocate available ports for pairing and watchdog fixtures

* fix: stabilize planning eval navigation and phase reporting

* test: isolate installed runtime paths in planning evals

* test: stabilize review evidence and concurrent refresh fixtures

* fix: resolve design findings before editing the plan

* fix: honor and persist disabled outside plan reviews

* fix: preserve planning decisions and terminal evidence

Load installed host reviews at autoplan phase entry and wait for completed
reviewers and saved artifacts. Reuse approved remedies while preserving
individual finding decisions.

Drive interactive evals from the current terminal viewport, bind native
questions across scrolling, and require complete native report evidence.
Cover captured stale menus, permission lifecycles, setup classification,
and disabled-review tool availability with deterministic regressions.

Advance release metadata and the upgrade migration to the unclaimed
1.83.0.0 slot.

* fix: drive native review questions and preserve current plans

Use the native single-choice keyboard protocol and current terminal viewport,
with per-question navigation inside packets and completed-call coverage.
Keep permissions, multi-select menus, and Submit controls distinct.

Send Autoplan reviewers the amended implementation plan, keep its review record
separate, and supply retained application contracts in the chain fixture.
Clarify individual DevEx decisions and complete CEO fix options; use one active
plan destination for the section-loading report.

* fix: preserve complete plan-review decisions

* fix: recognize native plan dialogs and reviewer controls

* fix: preserve review decisions and phase completion

* fix: recognize completed reviews without losing findings

* fix: preserve review continuity and native eval completion

* test: fix native review completion and eval retry isolation

* test: handle native review menus and complete eval fixtures

* test: fix native review setup, completion, and isolation failures

* test: limit native skill discovery to runtime assets

* fix: bind Autoplan reviews to full ordered phase inputs

* test: fix planning eval routing, counting, and timeout handling

* chore: advance queued release to v1.84.0.0

* fix: preserve complete review inputs and planning decisions

* fix: reconcile review approvals and preserve phase obligations

* fix: preserve review obligations and unblock eval permissions

Carry recorded Autoplan requirements into blind phase inputs, require Eng
review approvals before exit, and exercise combined asynchronous flows in
CEO reviews. Correct native finding and handoff classification and unblock
repeated report edits using scoped request identities.

* fix: retain plan requirements and complete native review dialogs

* fix: complete native review prompts and retain plan references

* fix: preserve review inputs and classify native eval evidence

* fix: check competing completion orders in CEO reviews

* fix: recognize review decisions and require phase methodology

Require the current phase methodology before Autoplan snapshots. Correct
substantive decision, closed handoff, and cache-finding classification, and
honor the recommended implementation approach in native review dialogs.

Add captured-transcript regressions without changing review thresholds,
provider models, retries, or deadlines.

* test: bind native review decisions and close completed handoffs

* fix: complete review dialogs and verify methodology delivery

* fix: preserve review evidence and unblock native eval prompts

* fix: handle native review question completions

* fix: recognize native review narration and controls

* fix: count native review decisions and isolate eval fixtures

* test: verify seeded review coverage and current artifact permissions

* test: isolate model and brain-aware skill renders

* fix: repair native workflow evaluation and clarify review steps

* fix: stabilize workflow eval evidence and review guidance

* test: repair native workflow observation and fixture isolation

* fix: recognize completed workflow evidence and owned skill reads

* test: repair seeded workflow delivery and completion evidence

* test: recognize current review evidence across native forms

* test: handle native review variants and permission redraws

* fix: honor review preferences and recognize native eval evidence

* test: recognize completed review decisions and queued permissions

* test: match current review contracts and partial-line edits

* test: recognize completed workflow evidence and bounded human waits

* fix: preserve review entry gates and native eval interactions

* fix: recognize native workflow evidence and preserve review gates

* test: recognize current review evidence and preconfigure workflow fixtures

* test: recognize completed review findings and scoped artifact permissions

* fix: stabilize native workflow review and permission evidence

* fix: recognize current review evidence and scoped edit confirmations

Clarify Design and engineering review entry instructions and Design scoring.
Recognize required legacy coverage and public Autoplan completion recaps.
Bind the pending Edit confirmation to its exact file, ordered digest, and
one-request approval when a preceding command display remains visible.
Keep reviews within their existing size limits and preserve scope gates
when extracting workflow fixtures from either supported preamble header.

Keep failure outcomes, review thresholds, provider choices, and eval budgets.

* fix: recover review workflow progress and eval evidence

* fix: recognize valid review evidence and scope selection

* test: fix review evidence parsing and repeated artifact prompts

* test: recognize valid review decisions and pending native cards

* fix(plan-eng-review): keep final navigation consistent with approved tasks

* test: recognize valid review evidence and bind legacy diff requests

* fix: stabilize review eval evidence and harness repair guidance

* docs: update project documentation for v1.85.0.0

Co-Authored-By: OpenAI Codex <noreply@openai.com>

* test: fix Windows CI fixtures and credential scan

Rebase captured JSON values and filesystem evidence using the appropriate
path convention. Compile native fake CLIs on Windows and synchronize pipe
holder readiness, with cleanup retained when assertions fail.

Assemble synthetic credential fixtures at runtime so the added-line scan
keeps enforcing the same gate without flagging its own rejection controls.

Discover generated skills directly for the empty-find regression check,
avoiding a recursive scan through saved evaluation artifacts and dependencies.

* fix: preserve source renders on Windows

Compare canonical generator paths using native separators so an output
sidecar pointing at the source cannot overwrite its skill or metadata.
Keep the regression fixture isolated from the real checkout and expose
freshness diagnostics before asserting subprocess status.

Detach Windows drain-test pipe holders from the fake provider's automatic
child cleanup while preserving the enclosing runner job and its assertions.

* fix: clarify outside review fallback and CEO decisions

Render one applicable own-harness fallback path and retain native review,
disabled policy, and missing-coverage semantics. Align report field names
and mode labels, and make the existing per-cut scope approval explicit.

Regenerate skill outputs and keep the workflow judge's model, thresholds,
and retry policy unchanged.

* chore: move release to free version slot (v1.86.0.0)

PR #2852 now claims v1.85.0.0. Align the release metadata and
rename migration so upgrades from that version still receive it.

Co-Authored-By: OpenAI Codex <noreply@openai.com>

* fix: include engineering review prerequisites and restore branch context

* fix: recognize coverage diagrams and clarify design review instructions

* fix: preserve file identities and join Windows test processes

---------

Co-authored-by: OpenAI Codex <noreply@openai.com>
2026-09-14 14:32:45 -07:00

7 lines
31 KiB
JSON
Raw Permalink Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
{
"sourceRun": "ship-source-u-full-paid-20260909-0740",
"initialImplementation": "# Plan: User Dashboard Page\n\n## Context\nWe're shipping a new user dashboard at `/dashboard` showing recent activity,\nnotifications panel, and quick-action buttons. Users land here after login.\n\n## UI Scope\n- New React page component `UserDashboard.tsx` at `src/pages/`\n- Three new sub-components: `ActivityFeed`, `NotificationsPanel`, `QuickActions`\n- Tailwind CSS for layout, mobile-first responsive (breakpoints: sm/md/lg)\n- Empty state, loading skeleton, error state for each panel\n- Hover states + focus-visible outlines on every interactive element\n- Modal dialog for \"Mark all as read\" on notifications panel\n- Toast notification system for action feedback\n\n## Backend\n- New REST endpoint `GET /api/dashboard` returns `{ activity, notifications, quickActions }`\n- Backed by existing PostgreSQL tables; no schema changes\n\n## Out of scope\n- Dark mode (separate plan)\n- Personalization / customization (separate plan)\n\n## Existing product and application contracts\n\nThis is the existing single-role member workspace, not a new product or a new\nonboarding flow. Members currently visit three separate pages after login to\nresume work, check alerts, and inspect recent changes. In the team's last task\nwalkthrough, finding the next item took a median 75 seconds. The dashboard's\nsuccess measure is login-to-first-completed-task time, targeting 45 seconds,\nwith completed-task rate and permission-error rate as guardrails. Existing\nanalytics records login, action start, action completion, and permission errors;\nthe new page still needs its own exposure and interaction instrumentation.\n\nActivity is the immutable audit history of workspace changes. Notifications are\nmember-specific alerts with persistent read state; acknowledging an alert does\nnot alter audit history. The existing action registry supplies three actions\n(create an item, resume assigned work, invite a member), with stable IDs, labels,\nroute targets, and server-side eligibility predicates. These are links into\nexisting workflows; action ranking and a new configuration service do not exist.\n\nThe application already uses cookie sessions and workspace membership middleware.\nIts request context supplies the authenticated member and workspace IDs. Existing\nrepository methods apply both IDs where appropriate; callers do not accept a\nworkspace ID from query parameters. Mutations already require CSRF tokens. The\nnew dashboard endpoint must compose these methods and follow the same boundaries;\nits handler, authorization integration, and failure paths have not been written.\n\nExisting list methods return the latest 20 records plus a cursor and have indexed\nworkspace/member and created-at access paths. The existing full activity and\nnotification pages own older-page navigation. The member-scoped bulk-read API is\nidempotent and marks only notifications at or before the supplied snapshot time,\nso later arrivals remain unread. Existing HTTP clients expose typed unauthenticated,\nforbidden, validation, retryable-service, and network errors. Each dashboard panel\nstill needs to map these results to its loading, empty, error, retry, and success\nstates; the aggregate endpoint's response composition and partial-failure behavior\nremain new implementation work. No schema migration or new mutation API is needed.\n\nThe app already has Tailwind spacing/color/type tokens, a responsive page shell,\nbuttons, links, and a dialog primitive with focus trapping, Escape dismissal, and\nfocus return. These primitives do not implement any dashboard panel, confirmation\nflow, or toast system. The new modal and toast feedback must also work with keyboard\nand screen readers; existing accessibility policy requires named controls, a live\nregion for nonblocking feedback, sufficient contrast, and reduced-motion support.\nThe dashboard still needs its own layout, content hierarchy, mobile behavior, and\nstate-specific copy at sm/md/lg breakpoints.\n\nVitest, React Testing Library, and Playwright already run in CI. Existing fixtures\ncover authenticated members, another workspace, empty lists, and service failures;\nthere are no dashboard-specific tests yet. Existing staging feature flags and\nrequest/error metrics support a member-cohort rollout and rollback to the current\nlanding page. The dashboard's rollout criteria, endpoint performance checks,\ninteraction tests, and accessibility verification must be specified and added.\n\nAll dashboard screen, panel, aggregate-endpoint, modal, and toast work listed above\nis new. The existing contracts describe dependencies to reuse, not completed work\nor prior approval of an implementation approach.\n",
"activeAfterAmend": "# Plan: User Dashboard Page\n\n## Implementation plan\n\n### Original plan\n\n**Context:** Shipping a new user dashboard at `/dashboard` — activity feed, notifications panel, quick-action buttons. Users land here after login. Success metric: login-to-first-completed-task median from 75 seconds to 45 seconds.\n\n**UI Scope:**\n- `UserDashboard.tsx` at `src/pages/`\n- Sub-components: `ActivityFeed`, `NotificationsPanel`, `QuickActions`\n- Tailwind CSS, mobile-first (sm/md/lg breakpoints)\n- Loading skeleton, empty state, error state per panel\n- Hover + focus-visible on interactive elements\n- Modal: \"Mark all as read\" on notifications panel\n- Toast notification system for action feedback\n\n**Backend:**\n- `GET /api/dashboard` → `{ activity, notifications, quickActions }`\n- Existing PostgreSQL tables; no schema changes\n\n**Out of scope:** dark mode, personalization.\n\n### Amendments from CEO review (A1-A14)\n\n**A1 — Partial-failure contract (REQUIRED before implementation):**\nResponse schema:\n```\n{\n activity: ActivityData | null,\n notifications: NotificationData | null,\n quickActions: QuickActionData | null,\n panelErrors: {\n activity?: { code: string, message: string },\n notifications?: { code: string, message: string },\n quickActions?: { code: string, message: string }\n }\n}\n```\nEach panel renders based on its own data/error field independently.\n\n**A2 — Parallel backend queries:** Handler uses Promise.all (or equivalent) for all three data sources. Serial execution prohibited.\n\n**A3 — Rollout criteria (pre-specified, required before feature flag enables):**\n1. p95 GET /api/dashboard latency < 300ms (staging verification)\n2. Zero WCAG 2.1 AA violations (keyboard + screen reader paths)\n3. Permission-error rate at or below current baseline\n4. Login-to-first-completed-task median < 50s in first cohort (abort if > 65s)\n\n**A4 — Double-click guard:** \"Mark all as read\" confirm button disables immediately on click; re-enables on API response.\n\n**A5 — Toast system:**\n- Use existing installed toast library if present; otherwise install `sonner` (not hand-rolled)\n- Toast container at AppShell level (not inside UserDashboard) — `role=\"status\"` / `aria-live=\"polite\"`\n- Max 3 visible toasts; FIFO replacement; 4-second auto-dismiss\n- Suppress animations via `@media (prefers-reduced-motion)`\n\n**A6 — PanelWrapper:** Shared component handling loading skeleton, empty state, error state + retry button, success slot. All three panels use it.\n\n**A7 — QuickActions empty state:** When all 3 actions ineligible, render defined empty state (e.g., \"No actions available\").\n\n**A8 — Panel order:** QuickActions → NotificationsPanel → ActivityFeed (mobile and lg breakpoint).\n\n**A9 — Content rendering:** All user-generated content in ActivityFeed and NotificationsPanel rendered as plaintext (React's default `{value}`). No `dangerouslySetInnerHTML`.\n\n**A10 — QuickActions service layer:** Data fetching via `fetchQuickActions()` service function, not direct action registry import. Enables future ranking service substitution.\n\n**A11 — Analytics events (required):**\n- `dashboard_viewed` on page mount\n- `panel_loaded {panel, duration_ms, status}` per panel\n- `quick_action_clicked {action_id}`\n- `notification_marked_read {count}` on bulk-read success\n- `panel_error {panel, error_code}`\n\n**A12 — Alert:** Dashboard endpoint p95 > 500ms for 5 minutes → existing alert channel.\n\n**A13 — Runbook:** 1-page operational runbook: what to check when endpoint is slow/erroring, how to disable feature flag, escalation path.\n\n**A14 — Required tests:**\n- Playwright: login → all panels render (happy path)\n- Playwright: one panel API fails → other two panels still render\n- RTL: \"Mark all as read\" modal — confirm / cancel / error states\n- RTL: double-click guard on confirm button\n- RTL: all panel states (loading / empty / error)\n- Playwright: keyboard navigation through all panels (Tab, Enter, Escape)\n- Playwright: reduced-motion mode (no skeleton animations)\n- Vitest: GET /api/dashboard partial failure response structure\n- Playwright staging: p95 latency verification check\n\n**Deferred to TODOS.md:**\n- State-aware redirect for members with in-progress work (P2)\n- Notification badge in page `<title>` (P3)\n- Keyboard shortcut for bulk read (P3)\n- Action ranking service integration (P2)\n- Validate action registry coverage vs. post-login analytics frequency (P1 — do before implementation if analytics accessible)\n\n---\n\n\n<!-- autoplan-accepted:ceo -->\n- A1: Partial-failure contract for GET /api/dashboard — per-panel status fields in response; each panel renders independently\n- A2: Parallel backend query execution required (Promise.all or equivalent)\n- A3: Rollout criteria pre-specified: p95 < 300ms staging, WCAG 2.1 AA clean, error rate baseline, median < 50s first cohort\n- A4: Double-click guard on Mark all as read confirm button\n- A5: Toast system: existing library or sonner; AppShell-level; max 3; FIFO; 4s dismiss; reduced-motion; ARIA live region\n- A6: PanelWrapper shared component for panel states\n- A7: QuickActions empty state when all actions ineligible\n- A8: Panel order: QuickActions → NotificationsPanel → ActivityFeed\n- A9: Plaintext rendering for user-generated content (no dangerouslySetInnerHTML)\n- A10: QuickActions data via service layer for future ranking decoupling\n- A11: 5 analytics events specified\n- A12: Dashboard p95 alert (> 500ms / 5min → existing alert channel)\n- A13: 1-page operational runbook for dashboard endpoint failures\n- A14: 9 required test specifications (Playwright + RTL + Vitest)\n<!-- /autoplan-accepted:ceo -->\n## Review record\n\n### CEO Review (Phase 1) — SELECTIVE EXPANSION\n\n**Mode:** SELECTIVE EXPANSION (auto-selected by /autoplan)\n**Dual voices:** Claude subagent completed; Codex disabled (codex_reviews=disabled)\n\n#### Pre-Review System Audit\nFixture repo — single plan file, one commit, no application code. Plan evaluated on specification quality only. No design doc, TODOS.md, or CLAUDE.md.\n\n#### 0A: Premise Challenge\n\n| Premise | Status |\n|---|---|\n| Navigation fragmentation causes 75s median | ASSUMED — root cause (load time vs. cognitive load) not validated |\n| Three quick actions cover top post-login behaviors | ASSUMED — action registry designed for eligibility, not frequency |\n| No schema changes needed | VALID |\n| Existing dialog primitive handles modal | VALID |\n| Feature flag supports rollout | VALID |\n| Partial-failure behavior can be decided during implementation | PROBLEMATIC — now required upfront (A1) |\n| Rollout criteria can be specified later | WRONG — queued as User Challenge for Final Gate |\n\n**User Challenge (queued for Final Gate):** Both CEO subagent and primary review flagged that root cause of 75-second median is unvalidated. If bottleneck is task ambiguity (not page count), dashboard may not move the needle. Does not block implementation; surfaced for user judgment at Final Gate.\n\n#### 0B: Existing Code Leverage Map\n\n| Sub-problem | Existing code |\n|---|---|\n| Auth / session | Cookie sessions + workspace membership middleware |\n| Activity data | Existing activityRepo.list() (20 records + cursor, indexed) |\n| Notification data | Existing notificationRepo.list() (20 records + cursor, indexed) |\n| Quick actions | Existing action registry (3 actions, eligibility predicates) |\n| HTTP error types | Typed: unauthenticated, forbidden, validation, retryable, network |\n| UI primitives | Tailwind tokens, responsive page shell, buttons, links, dialog |\n| Bulk-read | Existing idempotent mark-read API (snapshot-time scoped) |\n| Tests | Vitest + RTL + Playwright + existing fixtures |\n| Rollout | Existing feature flag + request/error metrics |\n\nNew (no existing equivalent): dashboard layout, aggregate endpoint, PanelWrapper, Toast system, dashboard analytics, rollout criteria.\n\n#### 0C: Dream State\n\n```\nCURRENT STATE THIS PLAN 12-MONTH IDEAL\n3 pages after login --> /dashboard (all 3 --> State-aware redirect\n(activity, notifs, panels). Static + personalized actions\nwork). 75s median. quick actions. + < 30s median\n 45s target.\n```\n\n#### 0C-bis: Implementation Alternatives\n\n```\nAPPROACH A: Aggregate Endpoint (the plan)\n Effort: M | Risk: Med | Single request; one loading state\n Con: Single point of failure; partial-failure contract needed (A1 fixes)\n\nAPPROACH B: Per-Panel Independent Fetch\n Effort: M | Risk: Low | Better fault isolation; natural progressive load\n Con: 3 requests vs 1; waterfall risk if not parallelized\n \nRECOMMENDATION: Both valid. Approach A retained (plan spec). Partial-failure\ncontract now required by A1. Logging as Taste Decision.\n```\n\n#### 0D: SELECTIVE EXPANSION Analysis\n\nComplexity: 8-10 files — at threshold but justified.\n\nState-aware redirect (10x reframe from subagent): ~1d human / ~30min CC. DEFERRED — outside blast radius, requires new routing logic; validate after first cohort data.\n\nDelight opportunities deferred: badge in `<title>`, optimistic bulk-read, keyboard shortcuts, \"last updated\" timestamps, skeleton-to-content transitions. All to TODOS.md.\n\n#### 0E: Temporal Interrogation\n\n```\nHOUR 1: Partial-failure schema + Toast placement (AppShell) — decide before any code\nHOUR 2-3: Parallel queries; Mark-all-read optimistic vs. confirmed; query parallelism\nHOUR 4-5: CSRF token for bulk-read; toast copy strings needed\nHOUR 6+: Rollout criteria verification; Playwright reduced-motion test\n```\n\n#### Section 1: Architecture\n\n```\n Browser\n │ GET /dashboard (feature-flagged)\n ▼\n UserDashboard.tsx\n ├── QuickActions.tsx ──────────────────┐\n ├── NotificationsPanel.tsx ────────── │── GET /api/dashboard\n │ └── [Mark all as read modal] │ ├── activityRepo.list() ─┐\n └── ActivityFeed.tsx ────────────────┘ ├── notifRepo.list() ─ │─ Promise.all\n └── actionRegistry() ─┘\n ▼\n PostgreSQL (existing)\n AppShell\n └── ToastProvider (new, root-level)\n```\n\n**Findings (auto-decided):**\n- Single-point-of-failure: → A1 (partial-failure contract). IN SCOPE.\n- Serial queries: → A2 (parallel). IN SCOPE.\n- Rollback: feature flag flip < 1 minute. GOOD.\n\n#### Section 2: Error & Rescue Map\n\n```\nMETHOD | WHAT CAN GO WRONG | EXCEPTION\n────────────────────────|─────────────────────────────|──────────────────\nGET /api/dashboard | One panel query timeout | QueryTimeoutError\n | All panels fail | (cascading)\n | Unhandled exception | UnhandledError ← GAP\nPOST /bulk-read | CSRF stale (403) | CSRFError\n | Auth expired (401) | AuthError\n | Network error | NetworkError\n\nEXCEPTION | RESCUED? | ACTION | USER SEES\n───────────────────|──────────|──────────────────────────────|───────────────────\nQueryTimeoutError | YES* | panelErrors field in response| Panel error + retry\nUnhandledError | NO ← GAP | — | 500 page ← BAD\nCSRFError | YES | Toast + modal close | \"Session expired\"\nAuthError | YES | Redirect /login | Login redirect\nNetworkError | YES | Error toast | \"Failed, try again\"\n```\n*Requires partial-failure contract (A1). GAP fixed by adding error middleware requirement.\n\n#### Section 3: Security\n\n| Threat | Mitigated? |\n|---|---|\n| Workspace ID injection via query param | YES — middleware provides IDs; contracts prohibit query param |\n| Direct object reference (cross-member data) | YES — existing repo methods scope to workspace+member |\n| CSRF on bulk-read | YES — CSRF tokens required for mutations |\n| XSS via activity/notification content | NO ← GAP → A9 (plaintext rendering) |\n\n#### Section 4: Interaction Edge Cases\n\n| Interaction | Edge Case | Status |\n|---|---|---|\n| Mark all read | Double-click confirm | GAP → A4 |\n| Mark all read | Navigation during request | OK (idempotent, silent complete) |\n| Notifications | New notif after snapshot | HANDLED (snapshot-time scoped API) |\n| Toast | Multiple simultaneous | GAP → A5 (max 3, FIFO) |\n| QuickActions | All actions ineligible | GAP → A7 |\n| Dashboard | Session expires mid-load | HANDLED (AuthError → redirect) |\n\n#### Section 5: Code Quality\n\n- DRY gap: 3 panels × 4 states without shared component → A6 (PanelWrapper)\n- Toast: don't hand-roll → A5 (use existing library or sonner)\n- QuickActions: direct registry import → A10 (service layer for future decoupling)\n\n#### Section 6: Tests\n\nNew things requiring tests:\n```\nNEW UX FLOWS: dashboard load, panel states (loading/empty/error/success),\n mark-all-read modal, quick action navigation, toast\nNEW DATA FLOWS: GET /api/dashboard (full + partial failure),\n POST /bulk-read\nNEW CODEPATHS: feature flag enabled/disabled, per-panel error/success/empty,\n double-click guard, auth expired during load\n```\n\nTest gaps → A14 (9 test specifications required).\n\n**2am Friday test:** Playwright full load + one panel failure + mark-all-read + keyboard nav + toast.\n\n#### Section 7: Performance\n\n| Concern | Fix |\n|---|---|\n| 3 serial DB queries = sum of latencies | A2: parallel execution required |\n| p95 latency target unspecified | A3: < 300ms in staging |\n| Existing indexes on workspace+member+created_at | No new indexes needed |\n\n#### Section 8: Observability\n\n5 analytics events missing → A11. Alert missing → A12. Runbook missing → A13.\n\nStructured logging requirement: each panel error logs `{panel, error_code, workspace_id, member_id, timestamp}`.\n\n#### Section 9: Deployment\n\n- No DB migration. GOOD.\n- Feature flag: existing system supports member-cohort rollout.\n- Rollback: flag flip < 1 minute.\n- Rollout criteria: CRITICAL GAP → A3.\n- Post-deploy: monitor p95 + panel error rate in first 15 minutes.\n\n#### Section 10: Long-Term Trajectory\n\n- Toast system: use library (A5) prevents tech debt.\n- QuickActions decoupling (A10): enables ranking service without UI changes.\n- PanelWrapper (A6): reusable by future dashboard-like features.\n- Reversibility: 4/5 (feature flag makes rollback trivial; no DB changes).\n\n#### Section 11: Design & UX\n\nPanel order: QuickActions first → A8 (information hierarchy for success metric).\n\n```\nLogin ──► QuickActions (action intent) ──► Resume work ──► Goal achieved\n │\n ├── NotificationsPanel (3 unread) ──► [optional: mark all read]\n │\n └── ActivityFeed (ambient context)\n```\n\nAccessibility gaps → A5 (live region for toast), A8 (tab order), A5 (reduced-motion).\n\nState coverage gaps: QuickActions empty → A7; partial dashboard failure → A1.\n\n---\n\n### CEO Dual Voices\n\n**CLAUDE SUBAGENT (independent CEO review):**\n1. Critical: Rollout criteria not specified\n2. High: Root cause of 75s median not validated\n3. High: Partial-failure contract deferred\n4. Medium: Action registry not validated vs. behavior frequency\n5. Medium: State-aware redirect not evaluated\n6. Medium: Alternatives section absent\n7. Medium: Toast + modal not connected to success metric\n\n**Codex:** disabled (codex_reviews=disabled)\n\n```\nCEO DUAL VOICES — CONSENSUS TABLE:\n═══════════════════════════════════════════════════════════════\n Dimension Claude Codex Consensus\n ──────────────────────────────────── ─────── ─────── ─────────\n 1. Premises valid? PARTIAL N/A N/A\n 2. Right problem to solve? MOSTLY N/A N/A\n 3. Scope calibration correct? YES N/A N/A\n 4. Alternatives sufficiently explored? NO N/A N/A\n 5. Competitive/market risks covered? YES N/A N/A\n 6. 6-month trajectory sound? MOSTLY N/A N/A\n═══════════════════════════════════════════════════════════════\nOutside disabled — 6 Consensus cells N/A, never CONFIRMED.\n```\n\n---\n\n### Required Outputs\n\n#### NOT in scope\n- Dark mode (existing out-of-scope)\n- Personalization (existing out-of-scope)\n- State-aware post-login redirect (DEFER — outside blast radius)\n- Notification badge in page `<title>` (DEFER)\n- Optimistic Mark-all-read (optional — implementer choice)\n- Action ranking service (DEFER — A10 decouples for future)\n- Keyboard shortcut for bulk read (DEFER)\n- Multi-page notification navigation (existing full notification page owns)\n\n#### What already exists\nSee Section 0B. Reusable: auth middleware, list methods, HTTP error types, UI primitives, dialog primitive, bulk-read API, test infrastructure, feature flags.\n\n#### Dream State Delta\nThis plan consolidates 3 pages → 1. Leaves open: state-aware routing, personalized actions, sub-30s median. Correct scope for v1.\n\n#### Error & Rescue Registry\nSee Section 2. 5 paths mapped, 1 critical gap (unhandled exception → 500) addressed by adding error middleware requirement.\n\n#### Failure Modes Registry\n```\nCODEPATH | FAILURE MODE | RESCUED? | TEST? | USER SEES? | LOGGED?\n──────────────────────|─────────────────────|──────────|──────-|─────────────────|────────\nGET /api/dashboard | One panel timeout | YES (A1) | YES(A14)| Panel error | YES\nGET /api/dashboard | All panels fail | YES (A1) | YES(A14)| All errors | YES\nGET /api/dashboard | Unhandled exception | YES* | NO→A14 | 500 page | YES*\nPOST /bulk-read | CSRF 403 | YES | NO→A14 | Toast+close | YES\nPOST /bulk-read | Auth 401 | YES | NO→A14 | Login redirect| YES\nActivityFeed | XSS via content | YES (A9) | NO→A14 | Plaintext | N/A\nNotifPanel | XSS via content | YES (A9) | NO→A14 | Plaintext | N/A\n```\n*Requires error middleware addition (addressed).\n\n**CRITICAL GAPS identified: 3 (all addressed by amendments A1, A9, error middleware)**\n\n---\n\n### CEO Completion Summary\n\n```\n+====================================================================+\n| MEGA PLAN REVIEW (CEO) — COMPLETION SUMMARY |\n+====================================================================+\n| Mode selected | SELECTIVE EXPANSION |\n| System Audit | Fixture only — plan specification evaluated |\n| Step 0 | SELECTIVE EXPANSION; premise challenge; |\n| | approach A retained + partial-failure added |\n| Section 1 (Arch) | 2 findings → A1, A2 |\n| Section 2 (Errors) | 5 paths mapped, 1 gap → error middleware |\n| Section 3 (Security) | 2 XSS gaps → A9 |\n| Section 4 (Data/UX) | 3 unhandled edge cases → A4, A5, A7 |\n| Section 5 (Quality) | 2 issues → A5, A6, A10 |\n| Section 6 (Tests) | 9 test gaps → A14 |\n| Section 7 (Perf) | 1 issue → A2; p95 target → A3 |\n| Section 8 (Observ) | 5 events missing → A11; alert → A12 |\n| Section 9 (Deploy) | Rollout criteria gap → A3; runbook → A13 |\n| Section 10 (Future) | Reversibility 4/5; decoupling → A10 |\n| Section 11 (Design) | Panel order → A8; a11y gaps → A5 |\n+--------------------------------------------------------------------+\n| NOT in scope | 8 items written |\n| What already exists | Written |\n| Dream state delta | Written |\n| Error/rescue registry | 5 paths, 1 gap addressed |\n| Failure modes | 7 rows, 3 critical gaps addressed |\n| TODOS.md updates | 5 items deferred |\n| Scope proposals | 0 accepted cherry-picks |\n| CEO plan | Not written (no accepted expansions) |\n| Outside voice | Claude subagent: 7 findings; Codex disabled |\n| Lake Score | 14/14 chose complete option |\n| Amendments | A1-A14 (14 requirements added) |\n| Unresolved decisions | 1 (root cause of 75s median) |\n+====================================================================+\n```\n\n<!-- autoplan-accepted:ceo -->\n- A1: Partial-failure contract for GET /api/dashboard — per-panel status fields in response; each panel renders independently\n- A2: Parallel backend query execution required (Promise.all or equivalent)\n- A3: Rollout criteria pre-specified: p95 < 300ms staging, WCAG 2.1 AA clean, error rate baseline, median < 50s first cohort\n- A4: Double-click guard on Mark all as read confirm button\n- A5: Toast system: existing library or sonner; AppShell-level; max 3; FIFO; 4s dismiss; reduced-motion; ARIA live region\n- A6: PanelWrapper shared component for panel states\n- A7: QuickActions empty state when all actions ineligible\n- A8: Panel order: QuickActions → NotificationsPanel → ActivityFeed\n- A9: Plaintext rendering for user-generated content (no dangerouslySetInnerHTML)\n- A10: QuickActions data via service layer for future ranking decoupling\n- A11: 5 analytics events specified\n- A12: Dashboard p95 alert (> 500ms / 5min → existing alert channel)\n- A13: 1-page operational runbook for dashboard endpoint failures\n- A14: 9 required test specifications (Playwright + RTL + Vitest)\n<!-- /autoplan-accepted:ceo -->\n\n---\n\n## Decision Audit Trail\n\n<!-- AUTONOMOUS DECISION LOG -->\n\n| # | Phase | Decision | Classification | Principle | Rationale | Rejected |\n|---|-------|----------|----------------|-----------|-----------|----------|\n| 1 | CEO | Mode: SELECTIVE EXPANSION | Mechanical | P6 | Feature iteration on existing workspace; autoplan pre-selects | EXPANSION, HOLD, REDUCTION |\n| 2 | CEO | Approach A retained + partial-failure contract required | Taste | P1+P5 | Both A and B valid; A is plan-specified; A1 addresses key concern | Approach B, Approach C |\n| 3 | CEO | A1: Partial-failure contract → IN SCOPE | Mechanical | P1+P2 | Critical gap; in blast radius; < 1d CC | Defer to implementation |\n| 4 | CEO | A2: Parallel queries → IN SCOPE | Mechanical | P1+P7 | Serial adds 200-400ms; correctness not taste | Accept serial |\n| 5 | CEO | A3: Rollout criteria → IN SCOPE | Mechanical | P1 | Critical governance gap; cannot enable flag safely without | Defer |\n| 6 | CEO | A4: Double-click guard → IN SCOPE | Mechanical | P1+P5 | Interaction correctness; idempotent API | Ignore |\n| 7 | CEO | A5: Toast at AppShell, use library | Mechanical | P4+P5 | DRY; existing library > hand-rolled; placement is correctness | Hand-roll |\n| 8 | CEO | A6: PanelWrapper → IN SCOPE | Mechanical | P4+P5 | 3 panels × 4 states = DRY without shared component | Separate per-panel |\n| 9 | CEO | A7: QuickActions empty state → IN SCOPE | Mechanical | P1 | Edge case; in blast radius | Omit |\n| 10 | CEO | A8: Panel order (QuickActions first) | Mechanical | P5 | Information hierarchy for success metric | Keep original order |\n| 11 | CEO | A9: Plaintext rendering (XSS) | Mechanical | P1+P3 | Security correctness; React default | Allow HTML |\n| 12 | CEO | A10: QuickActions service layer | Mechanical | P5 | Enables future ranking; explicit boundary | Direct import |\n| 13 | CEO | A11-A13: Observability + alert + runbook | Mechanical | P1+P8 | Observability is scope | Post-launch |\n| 14 | CEO | A14: Test specifications → IN SCOPE | Mechanical | P1+P6 | Concrete specs reduce ambiguity | Vague \"add tests\" |\n| 15 | CEO | State-aware redirect → DEFER | Mechanical | P3 | Outside blast radius; new routing logic | Include in plan |\n| 16 | CEO | Root cause of 75s median → USER CHALLENGE | User Challenge | — | Both models flagged; queued for Final Gate | Auto-decide |\n\n---\n\n## GSTACK REVIEW REPORT\n\n| Review | Trigger | Why | Runs | Status | Findings |\n|--------|---------|-----|------|--------|----------|\n| CEO Review | `/plan-ceo-review` | Scope & strategy | 1 | issues_open | 14 amendments added; 1 User Challenge queued for gate |\n| Outside Review | Claude subagent (Codex disabled) | Independent 2nd opinion | 1 | issues_found [subagent-only] | 7 findings; Codex: disabled |\n| Eng Review | `/plan-eng-review` | Architecture & tests (required) | 0 | — | Not yet run |\n| Design Review | `/plan-design-review` | UI/UX gaps | 0 | — | Not yet run |\n| DX Review | `/plan-devex-review` | Developer experience gaps | 0 | — | Not yet run |\n\n**OUTSIDE COVERAGE:** CEO phase — Claude subagent completed (source: in-host); Codex: disabled. Consensus N/A (outside disabled). Design phase: not yet run. DX phase: not yet run. Eng phase: not yet run.\n\n**VERDICT:** CEO review complete with amendments. Eng review required before implementation.\n\n**UNRESOLVED DECISIONS:**\n- User Challenge: Root cause of 75-second median not validated — both models flagged that the dashboard may solve the wrong bottleneck if cognitive load (task ambiguity) rather than navigation fragmentation is the real cause. Your original direction (build the dashboard) stands unless you explicitly change it at the Final Gate.\n",
"note": "Actual U closed CEO after replacing the original with a summary; raw operator-cancelled/incomplete outcome is unchanged."
}