Files
gstack/test/fixtures/autoplan/t-ceo-omitted-obligations.json
Garry TanandOpenAI Codex 9f81911136 v1.86.0.0 feat: route outside reviews by harness (#2850)
* feat: add a restricted and supervised Claude Code runner

Preserve configured authentication and models while enforcing tool access, strict completion JSON, bounded output and process cleanup. Cover argv, failure handling, session metadata and Windows process containment.

* feat: route outside reviews by harness and migrate wrapper installs

Use Claude Code from Codex and Codex from other supported hosts, with shared invocation rendering, positive gate validation and per-phase provenance. Rename /claude to /claude-code, repair managed shared and copied installations safely, and generate native Kiro skills. Add installed-workflow, failure-injection and live cross-harness regression coverage.

* test: recognize CEO mode labels without terminal spacing

The paid workflow rendered SCOPEEXPANSION at option 4, but its driver required a literal space. Match the leading mode title without cursor-spacing artifacts and ignore adjacent preview text. Preserve missing-target failures and downstream posture assertions.

* test: isolate plan-count fixtures before starting review workflows

Seed the complete test plan in a private git repository before launching Claude, so a bare slash command cannot review the live workspace while a delayed fixture message remains queued. Preserve count thresholds, parsers and budgets. Add initial-context and installed-discovery tests, and retain startup/terminal diagnostics on failed evaluations.

* test: stabilize review fixtures and Claude eval startup

Preserve source boundaries in workflow judge inputs, isolate CEO mode plans, and wait for interactive trust input readiness. Keep startup failure evidence and retain existing models, budgets, and assertions.

Co-Authored-By: OpenAI Codex <noreply@openai.com>

* test: classify collapsed review modes and isolate seeded findings

Keep review questions out of the setup count when terminal cursor positioning removes spaces. State existing webhook safeguards so the five-finding control measures its seeded defects without accidental extra security and concurrency gaps. Preserve question bands and the paired control.

Co-Authored-By: OpenAI Codex <noreply@openai.com>

* test: isolate browser daemon state across free shards

Co-Authored-By: OpenAI Codex <noreply@openai.com>

* test: stabilize native review counting and interactive navigation

Co-Authored-By: OpenAI Codex <noreply@openai.com>

* chore: prepare v1.82.0.0 release

Co-Authored-By: OpenAI Codex <noreply@openai.com>

* fix: eliminate browser and process-cleanup test flakes

Pin every CI surface to Bun 1.4.0 to avoid extra-stdio finalizers closing
reused live sockets. Add an isolated GC/listener regression that fails on
Bun 1.3.13, and prevent coordinated rollback to an affected CI runtime.

Check renderer cleanup against the render's own staging directory so
concurrent renders cannot invalidate the assertion. Make the no-pgrep
process-tree walk tolerate disappearing /proc entries, and synchronize
its test fixture through child readiness and pipe EOF instead of sleeps.

Validation: 9,157 passed, 31 skipped, zero failures across 556 files with
retries disabled. Build, all-host generation freshness, and skill checks
passed. All three races have failing-before/passing-after regressions.

* fix: count completed native review questions in evals

* fix: drive review navigation from confirmed native choices

* fix: require complete section-loading eval reports

* test: isolate telemetry HTTP transport from local assertions

* fix: keep review input on the active native question

* test: let tunnel revocation daemon choose an available port

* test: allocate available ports for pairing and watchdog fixtures

* fix: stabilize planning eval navigation and phase reporting

* test: isolate installed runtime paths in planning evals

* test: stabilize review evidence and concurrent refresh fixtures

* fix: resolve design findings before editing the plan

* fix: honor and persist disabled outside plan reviews

* fix: preserve planning decisions and terminal evidence

Load installed host reviews at autoplan phase entry and wait for completed
reviewers and saved artifacts. Reuse approved remedies while preserving
individual finding decisions.

Drive interactive evals from the current terminal viewport, bind native
questions across scrolling, and require complete native report evidence.
Cover captured stale menus, permission lifecycles, setup classification,
and disabled-review tool availability with deterministic regressions.

Advance release metadata and the upgrade migration to the unclaimed
1.83.0.0 slot.

* fix: drive native review questions and preserve current plans

Use the native single-choice keyboard protocol and current terminal viewport,
with per-question navigation inside packets and completed-call coverage.
Keep permissions, multi-select menus, and Submit controls distinct.

Send Autoplan reviewers the amended implementation plan, keep its review record
separate, and supply retained application contracts in the chain fixture.
Clarify individual DevEx decisions and complete CEO fix options; use one active
plan destination for the section-loading report.

* fix: preserve complete plan-review decisions

* fix: recognize native plan dialogs and reviewer controls

* fix: preserve review decisions and phase completion

* fix: recognize completed reviews without losing findings

* fix: preserve review continuity and native eval completion

* test: fix native review completion and eval retry isolation

* test: handle native review menus and complete eval fixtures

* test: fix native review setup, completion, and isolation failures

* test: limit native skill discovery to runtime assets

* fix: bind Autoplan reviews to full ordered phase inputs

* test: fix planning eval routing, counting, and timeout handling

* chore: advance queued release to v1.84.0.0

* fix: preserve complete review inputs and planning decisions

* fix: reconcile review approvals and preserve phase obligations

* fix: preserve review obligations and unblock eval permissions

Carry recorded Autoplan requirements into blind phase inputs, require Eng
review approvals before exit, and exercise combined asynchronous flows in
CEO reviews. Correct native finding and handoff classification and unblock
repeated report edits using scoped request identities.

* fix: retain plan requirements and complete native review dialogs

* fix: complete native review prompts and retain plan references

* fix: preserve review inputs and classify native eval evidence

* fix: check competing completion orders in CEO reviews

* fix: recognize review decisions and require phase methodology

Require the current phase methodology before Autoplan snapshots. Correct
substantive decision, closed handoff, and cache-finding classification, and
honor the recommended implementation approach in native review dialogs.

Add captured-transcript regressions without changing review thresholds,
provider models, retries, or deadlines.

* test: bind native review decisions and close completed handoffs

* fix: complete review dialogs and verify methodology delivery

* fix: preserve review evidence and unblock native eval prompts

* fix: handle native review question completions

* fix: recognize native review narration and controls

* fix: count native review decisions and isolate eval fixtures

* test: verify seeded review coverage and current artifact permissions

* test: isolate model and brain-aware skill renders

* fix: repair native workflow evaluation and clarify review steps

* fix: stabilize workflow eval evidence and review guidance

* test: repair native workflow observation and fixture isolation

* fix: recognize completed workflow evidence and owned skill reads

* test: repair seeded workflow delivery and completion evidence

* test: recognize current review evidence across native forms

* test: handle native review variants and permission redraws

* fix: honor review preferences and recognize native eval evidence

* test: recognize completed review decisions and queued permissions

* test: match current review contracts and partial-line edits

* test: recognize completed workflow evidence and bounded human waits

* fix: preserve review entry gates and native eval interactions

* fix: recognize native workflow evidence and preserve review gates

* test: recognize current review evidence and preconfigure workflow fixtures

* test: recognize completed review findings and scoped artifact permissions

* fix: stabilize native workflow review and permission evidence

* fix: recognize current review evidence and scoped edit confirmations

Clarify Design and engineering review entry instructions and Design scoring.
Recognize required legacy coverage and public Autoplan completion recaps.
Bind the pending Edit confirmation to its exact file, ordered digest, and
one-request approval when a preceding command display remains visible.
Keep reviews within their existing size limits and preserve scope gates
when extracting workflow fixtures from either supported preamble header.

Keep failure outcomes, review thresholds, provider choices, and eval budgets.

* fix: recover review workflow progress and eval evidence

* fix: recognize valid review evidence and scope selection

* test: fix review evidence parsing and repeated artifact prompts

* test: recognize valid review decisions and pending native cards

* fix(plan-eng-review): keep final navigation consistent with approved tasks

* test: recognize valid review evidence and bind legacy diff requests

* fix: stabilize review eval evidence and harness repair guidance

* docs: update project documentation for v1.85.0.0

Co-Authored-By: OpenAI Codex <noreply@openai.com>

* test: fix Windows CI fixtures and credential scan

Rebase captured JSON values and filesystem evidence using the appropriate
path convention. Compile native fake CLIs on Windows and synchronize pipe
holder readiness, with cleanup retained when assertions fail.

Assemble synthetic credential fixtures at runtime so the added-line scan
keeps enforcing the same gate without flagging its own rejection controls.

Discover generated skills directly for the empty-find regression check,
avoiding a recursive scan through saved evaluation artifacts and dependencies.

* fix: preserve source renders on Windows

Compare canonical generator paths using native separators so an output
sidecar pointing at the source cannot overwrite its skill or metadata.
Keep the regression fixture isolated from the real checkout and expose
freshness diagnostics before asserting subprocess status.

Detach Windows drain-test pipe holders from the fake provider's automatic
child cleanup while preserving the enclosing runner job and its assertions.

* fix: clarify outside review fallback and CEO decisions

Render one applicable own-harness fallback path and retain native review,
disabled policy, and missing-coverage semantics. Align report field names
and mode labels, and make the existing per-cut scope approval explicit.

Regenerate skill outputs and keep the workflow judge's model, thresholds,
and retry policy unchanged.

* chore: move release to free version slot (v1.86.0.0)

PR #2852 now claims v1.85.0.0. Align the release metadata and
rename migration so upgrades from that version still receive it.

Co-Authored-By: OpenAI Codex <noreply@openai.com>

* fix: include engineering review prerequisites and restore branch context

* fix: recognize coverage diagrams and clarify design review instructions

* fix: preserve file identities and join Windows test processes

---------

Co-authored-by: OpenAI Codex <noreply@openai.com>
2026-09-14 14:32:45 -07:00

9 lines
51 KiB
JSON
Raw Permalink Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
{
"sourceRun": "ship-source-t-full-paid-20260909-0639",
"sourcePlanSha256": "052033a02ef68e995bb1f432e73f2d9bdb11065781afa1e2d0d729466b3b7277",
"initialImplementation": "\nWe're shipping a new user dashboard at `/dashboard` showing recent activity,\nnotifications panel, and quick-action buttons. Users land here after login.\n\n### UI Scope\n- New React page component `UserDashboard.tsx` at `src/pages/`\n- Three new sub-components: `ActivityFeed`, `NotificationsPanel`, `QuickActions`\n- Tailwind CSS for layout, mobile-first responsive (breakpoints: sm/md/lg)\n- Empty state, loading skeleton, error state for each panel\n- Hover states + focus-visible outlines on every interactive element\n- Modal dialog for \"Mark all as read\" on notifications panel\n- Toast notification system for action feedback\n\n### Backend\n- New REST endpoint `GET /api/dashboard` returns `{ activity, notifications, quickActions }`\n- Backed by existing PostgreSQL tables; no schema changes\n\n### Out of scope\n- Dark mode (separate plan)\n- Personalization / customization (separate plan)\n\n### Existing product and application contracts\n\nThis is the existing single-role member workspace, not a new product or a new\nonboarding flow. Members currently visit three separate pages after login to\nresume work, check alerts, and inspect recent changes. In the team's last task\nwalkthrough, finding the next item took a median 75 seconds. The dashboard's\nsuccess measure is login-to-first-completed-task time, targeting 45 seconds,\nwith completed-task rate and permission-error rate as guardrails. Existing\nanalytics records login, action start, action completion, and permission errors;\nthe new page still needs its own exposure and interaction instrumentation.\n\nActivity is the immutable audit history of workspace changes. Notifications are\nmember-specific alerts with persistent read state; acknowledging an alert does\nnot alter audit history. The existing action registry supplies three actions\n(create an item, resume assigned work, invite a member), with stable IDs, labels,\nroute targets, and server-side eligibility predicates. These are links into\nexisting workflows; action ranking and a new configuration service do not exist.\n\nThe application already uses cookie sessions and workspace membership middleware.\nIts request context supplies the authenticated member and workspace IDs. Existing\nrepository methods apply both IDs where appropriate; callers do not accept a\nworkspace ID from query parameters. Mutations already require CSRF tokens. The\nnew dashboard endpoint must compose these methods and follow the same boundaries;\nits handler, authorization integration, and failure paths have not been written.\n\nExisting list methods return the latest 20 records plus a cursor and have indexed\nworkspace/member and created-at access paths. The existing full activity and\nnotification pages own older-page navigation. The member-scoped bulk-read API is\nidempotent and marks only notifications at or before the supplied snapshot time,\nso later arrivals remain unread. Existing HTTP clients expose typed unauthenticated,\nforbidden, validation, retryable-service, and network errors. Each dashboard panel\nstill needs to map these results to its loading, empty, error, retry, and success\nstates; the aggregate endpoint's response composition and partial-failure behavior\nremain new implementation work. No schema migration or new mutation API is needed.\n\nThe app already has Tailwind spacing/color/type tokens, a responsive page shell,\nbuttons, links, and a dialog primitive with focus trapping, Escape dismissal, and\nfocus return. These primitives do not implement any dashboard panel, confirmation\nflow, or toast system. The new modal and toast feedback must also work with keyboard\nand screen readers; existing accessibility policy requires named controls, a live\nregion for nonblocking feedback, sufficient contrast, and reduced-motion support.\nThe dashboard still needs its own layout, content hierarchy, mobile behavior, and\nstate-specific copy at sm/md/lg breakpoints.\n\nVitest, React Testing Library, and Playwright already run in CI. Existing fixtures\ncover authenticated members, another workspace, empty lists, and service failures;\nthere are no dashboard-specific tests yet. Existing staging feature flags and\nrequest/error metrics support a member-cohort rollout and rollback to the current\nlanding page. The dashboard's rollout criteria, endpoint performance checks,\ninteraction tests, and accessibility verification must be specified and added.\n\nAll dashboard screen, panel, aggregate-endpoint, modal, and toast work listed above\nis new. The existing contracts describe dependencies to reuse, not completed work\nor prior approval of an implementation approach.\n\n",
"activeAtBoundary": "# Plan: User Dashboard Page\n\n## Implementation plan\n\nWe're shipping a new user dashboard at `/dashboard` showing recent activity,\nnotifications panel, and quick-action buttons. Users land here after login.\n\n**Success metric:** Login-to-first-completed-task time < 45 seconds (baseline: 75s median).\n**Guardrails:** completed-task rate, permission-error rate.\n\n### UI Scope\n- New React page component `UserDashboard.tsx` at `src/pages/`\n- New shared `PanelWrapper` component (loading/empty/error states, retry callback) — DRY across panels [CEO cherry-pick]\n- Three panel components: `ActivityFeed`, `NotificationsPanel`, `QuickActions`\n- Layout order: **Quick Actions first**, then Notifications, then Activity [action-first for 45s metric]\n- Tailwind CSS for layout, mobile-first responsive (breakpoints: sm/md/lg); Quick Actions stacks first on mobile\n- Empty state, loading skeleton, error state for each panel (via PanelWrapper)\n- Hover states + focus-visible outlines on every interactive element\n- \"Mark all as read\" — direct button (no confirmation modal); undo toast on success [CEO §11 UX]\n- Toast notification system for action feedback; toast queue/stack spec: newer toasts append below; max 3 visible; oldest auto-dismiss at 4s\n- Post-login redirect to `/dashboard` for authenticated members [CEO cherry-pick]\n\n### Backend\n- New REST endpoint `GET /api/dashboard`:\n - Runs ActivityRepository.list(), NotificationRepository.list(), ActionRegistry.getActions() in parallel (Promise.all)\n - **Partial-failure policy:** Returns 200 with `{ activity: { status, data?, error? }, notifications: { status, data?, error? }, quickActions: { status, data?, error? } }`. Each panel degraded independently; failed panel → error status + panel shows error state.\n - **Error handling:** ConnectionPoolExhausted → 503 with Retry-After header; RegistryError → QuickActions returns `{ status: 'error' }`; QueryTimeoutError → that panel returns `{ status: 'error' }`.\n - **Logging:** Structured log on entry (memberId, workspaceId, requestId) and on each panel failure (error class, duration).\n- New mutation `POST /api/notifications/mark-all-read`:\n - Validates snapshotTime (must be ≤ now; rejects future timestamps with 400)\n - Requires CSRF token (existing pattern)\n - Creates server-side audit log entry (memberId, workspaceId, snapshotTime, timestamp)\n- Backed by existing PostgreSQL tables; no schema changes\n\n### Analytics instrumentation (new)\n- `dashboard.exposure` — fired on page mount (memberId, workspaceId, panelStatuses)\n- `dashboard.first_action_click_ms` — latency from exposure to first QuickAction click\n- `dashboard.mark_all_read` — fired on successful bulk-read\n- `api.dashboard.latency` histogram by overall status (all-ok / partial / error)\n\n### Rollout criteria (must be verified before expanding beyond 5% cohort)\n- `api.dashboard.latency` p95 < 400ms at 10% rollout\n- Login-to-first-task median < 60s at 50% rollout (measured via `dashboard.first_action_click_ms`)\n- Zero critical accessibility failures in Playwright a11y suite\n- `api.dashboard.error_rate` < 0.5% at any rollout level triggers rollback\n\n### Out of scope\n- Dark mode (separate plan)\n- Personalization / customization (separate plan)\n- Real-time updates (separate plan)\n- Dashboard response caching (revisit at scale)\n- Keyboard shortcuts for quick actions (P3 follow-up)\n\n### Existing product and application contracts\n\nThis is the existing single-role member workspace, not a new product or a new\nonboarding flow. Members currently visit three separate pages after login to\nresume work, check alerts, and inspect recent changes. In the team's last task\nwalkthrough, finding the next item took a median 75 seconds. The dashboard's\nsuccess measure is login-to-first-completed-task time, targeting 45 seconds,\nwith completed-task rate and permission-error rate as guardrails. Existing\nanalytics records login, action start, action completion, and permission errors;\nthe new page still needs its own exposure and interaction instrumentation.\n\nActivity is the immutable audit history of workspace changes. Notifications are\nmember-specific alerts with persistent read state; acknowledging an alert does\nnot alter audit history. The existing action registry supplies three actions\n(create an item, resume assigned work, invite a member), with stable IDs, labels,\nroute targets, and server-side eligibility predicates. These are links into\nexisting workflows; action ranking and a new configuration service do not exist.\n\nThe application already uses cookie sessions and workspace membership middleware.\nIts request context supplies the authenticated member and workspace IDs. Existing\nrepository methods apply both IDs where appropriate; callers do not accept a\nworkspace ID from query parameters. Mutations already require CSRF tokens. The\nnew dashboard endpoint must compose these methods and follow the same boundaries;\nits handler, authorization integration, and failure paths have not been written.\n\nExisting list methods return the latest 20 records plus a cursor and have indexed\nworkspace/member and created-at access paths. The existing full activity and\nnotification pages own older-page navigation. The member-scoped bulk-read API is\nidempotent and marks only notifications at or before the supplied snapshot time,\nso later arrivals remain unread. Existing HTTP clients expose typed unauthenticated,\nforbidden, validation, retryable-service, and network errors. Each dashboard panel\nstill needs to map these results to its loading, empty, error, retry, and success\nstates; the aggregate endpoint's response composition and partial-failure behavior\nremain new implementation work. No schema migration or new mutation API is needed.\n\nThe app already has Tailwind spacing/color/type tokens, a responsive page shell,\nbuttons, links, and a dialog primitive with focus trapping, Escape dismissal, and\nfocus return. These primitives do not implement any dashboard panel, confirmation\nflow, or toast system. The new modal and toast feedback must also work with keyboard\nand screen readers; existing accessibility policy requires named controls, a live\nregion for nonblocking feedback, sufficient contrast, and reduced-motion support.\nThe dashboard still needs its own layout, content hierarchy, mobile behavior, and\nstate-specific copy at sm/md/lg breakpoints.\n\nVitest, React Testing Library, and Playwright already run in CI. Existing fixtures\ncover authenticated members, another workspace, empty lists, and service failures;\nthere are no dashboard-specific tests yet. Existing staging feature flags and\nrequest/error metrics support a member-cohort rollout and rollback to the current\nlanding page. The dashboard's rollout criteria, endpoint performance checks,\ninteraction tests, and accessibility verification must be specified and added.\n\nAll dashboard screen, panel, aggregate-endpoint, modal, and toast work listed above\nis new. The existing contracts describe dependencies to reuse, not completed work\nor prior approval of an implementation approach.\n\n## Review record\n\n**Autoplan run:** 2026-09-09 · Branch: main\n**UI scope:** yes (dashboard, modal, component, screen, form, button)\n**DX scope:** yes (dxRequired=true)\n**Codex:** disabled — Claude-native passes only\n**Phases:** CEO → Design → DX → Eng\n\n## Decision Audit Trail\n\n| # | Phase | Decision | Classification | Principle | Rationale | Rejected |\n|---|-------|----------|----------------|-----------|-----------|----------|\n| 1 | Phase 0 | Mode: SELECTIVE EXPANSION | Mechanical | P3/P6 | New greenfield feature on existing system; default SELECTIVE EXPANSION | SCOPE EXPANSION, HOLD SCOPE, SCOPE REDUCTION |\n| 2 | CEO §0C-bis | Implementation approach B (aggregate endpoint + parallel queries + per-panel failure) | Mechanical | P1+P3 | Plan already describes B; most resilient without over-engineering separate endpoints | A (sequential, single-failure all-or-nothing), C (3 separate endpoints) |\n| 3 | CEO §0D cherry | Post-login redirect to /dashboard | Mechanical | P2 | Trivial effort (1 line), directly enables 45s success metric | SKIP |\n| 4 | CEO §0D cherry | PanelWrapper shared loading/empty/error component | Mechanical | P4 | DRY — all 3 panels share identical state UI | SKIP |\n| 5 | CEO §0D cherry | Real-time SSE/WebSocket updates | Mechanical | P3 | Significant infra, separate plan warranted, not in blast radius | ACCEPT |\n| 6 | CEO §0D cherry | Dashboard response caching 30s TTL | Mechanical | P3 | Not blocking for v1; revisit at scale | ACCEPT |\n| 7 | CEO §1 | ConnectionPoolExhausted → P1 task, add explicit pool exhaustion handling | Mechanical | P1 | Silent 500 is a critical gap per zero-silent-failures prime directive | Skip |\n| 8 | CEO §2 | RegistryError → P1 task, add registry error guard | Mechanical | P1 | Unhandled crash of QuickActions panel | Skip |\n| 9 | CEO §4 | Double-click guard for Mark All Read → P1 task | Mechanical | P1 | Race condition, optimistic update or debounce needed | Skip |\n| 10 | CEO §4 | Toast queue spec → P2 task | Mechanical | P1 | Multiple simultaneous actions produce undefined toast behavior | Skip |\n| 11 | CEO §6 | Rollout criteria → P1 task (specify advancement gates) | Mechanical | P1 | Cannot roll back without agreed threshold; Codex + native both flag | Skip |\n| 12 | CEO §8 | Analytics event spec → P1 task (exposure + interaction events) | Mechanical | P1 | Plan says \"still needs instrumentation\" but events not defined | Skip |\n\n---\n\n## Phase 1 — CEO Review\n\n**Mode:** SELECTIVE EXPANSION \n**Snapshot hash:** ec5c6e27628a06ea9262d68200082b610a69ed0de018c0b8204624db6738cb15\n\n### Step 0A: Premise Challenge\n\nPremises evaluated:\n1. **75s median is a navigation problem** — Partially validated (team task walkthrough shows 3-page navigation is the workflow). The measurement exists and is concrete. However, the subagent flags that the bottleneck could be cognitive (deciding what to do) rather than navigational. Risk: medium. Decision: accept premise but add a TASTE DECISION for the Final Gate re: whether session recordings should be done first.\n2. **45s target is achievable** — No evidence against it; the three-panel consolidated view is a direct path to reducing navigation overhead. Accepted.\n3. **Existing action registry covers the right quick actions** — Stated constraint, not a premise. Three actions with stable IDs. Accepted.\n4. **Aggregate endpoint is the right backend shape** — This is the implementation premise with the highest risk (see Section 1). Partial failure semantics unspecified. FLAGGED.\n\n### Step 0B: Existing Code Leverage\n\n| Sub-problem | Existing code | Reused? |\n|---|---|---|\n| Auth / session | Cookie sessions + workspace membership middleware | Yes — compose at handler level |\n| Activity data | ActivityRepository.list(workspaceId, memberId, cursor) | Yes — index on workspace/member + created-at |\n| Notification data | NotificationRepository.list(memberId, cursor) | Yes |\n| Mark all read | NotificationRepository.markAllRead(memberId, snapshotTime) | Yes — idempotent |\n| Quick actions | ActionRegistry.getActions(memberId) | Yes — stable IDs, labels, eligibility predicates |\n| Modal | Existing dialog primitive (focus trap, Escape, focus return) | Yes |\n| Layout / tokens | Tailwind spacing/color/type tokens, page shell, buttons, links | Yes |\n| Analytics | Existing events (login, action start/complete, permission error) | Extend — new exposure + interaction events needed |\n| Feature flag | Existing staging flags + request/error metrics | Yes — member-cohort rollout |\n| Test fixtures | Authenticated member, empty lists, service failures | Extend — no dashboard-specific fixtures yet |\n| Test runners | Vitest, RTL, Playwright | Yes — already in CI |\n\nNo parallel re-implementations detected.\n\n### Step 0C: Dream State\n\n```\nCURRENT STATE THIS PLAN 12-MONTH IDEAL\n─────────────────────── ───────────────────────── ────────────────────────────\n3 separate pages Unified /dashboard Personalized dashboard\n75s median to first task Target 45s login-to-task <20s with smart prioritization\nNo cross-page visibility Activity + Notifications + Real-time updates, mobile app,\n Quick Actions in one view push notifications, digest email\nManual navigation Default post-login dest. Proactive \"what's urgent now\"\n```\n\n**Dream state delta:** This plan gets us to the \"unified view\" waypoint. It leaves real-time updates, personalization, mobile, and smart prioritization for later plans — all of which the architecture supports cleanly (no path dependency).\n\n### Step 0C-bis: Implementation Alternatives\n\n```\nAPPROACH A: Sequential Endpoint (all-or-nothing)\n Summary: Single GET /api/dashboard with sequential DB queries; one failure = 500.\n Effort: S | Risk: High\n Pros: Simpler handler; easy to test; no partial-failure spec\n Cons: One slow panel stalls everything; worst p99 = sum of 3 queries\n\nAPPROACH B: Aggregate Endpoint + Parallel Queries + Per-Panel Failure (RECOMMENDED)\n Summary: GET /api/dashboard runs queries in parallel; each panel can fail independently.\n Effort: M | Risk: Medium\n Pros: Resilient to individual panel failures; lower latency; independently retryable\n Cons: More complex handler; partial-failure semantics need explicit spec\n Reuses: Existing repository methods, existing HTTP error types\n\nAPPROACH C: Separate Endpoints Per Panel\n Summary: Three endpoints loaded in parallel client-side.\n Effort: L | Risk: Low\n Pros: Maximum isolation; each evolves independently; easy per-panel caching\n Cons: 3x roundtrips (though parallel); harder to test combined view\n```\n\n**RECOMMENDATION: Approach B.** Plan already describes B. Right balance of resilience and complexity. Auto-decided (Mechanical).\n\n### Step 0D: Cherry-Pick Ceremony (SELECTIVE EXPANSION)\n\n- ✅ ACCEPTED: Post-login redirect to /dashboard (1 line, enables success metric)\n- ✅ ACCEPTED: PanelWrapper shared component (DRY, avoids repeating loading/empty/error in 3 panels)\n- 📋 DEFERRED: Keyboard shortcuts for quick actions (P3 follow-up)\n- 📋 DEFERRED: Dashboard response caching 30s TTL (revisit at scale)\n- ❌ SKIPPED: Real-time SSE/WebSocket (separate plan)\n\n### Step 0E: Temporal Interrogation\n\n```\nHOUR 1 (foundations):\n - What is the partial-failure response shape? (200+partial? 207?) — decide before writing handler\n - Does the aggregate endpoint run queries in parallel (Promise.all) or sequential? — parallel preferred\n - What is the default redirect for unauthenticated GET /dashboard? — existing login redirect middleware should handle\n\nHOUR 2-3 (core logic):\n - How does each panel distinguish \"loading\" vs \"empty\" when the endpoint returns [] ?\n - Toast system: is this a new component or does the app have one? — plan says new\n - PanelWrapper props shape: onRetry callback, data, status enum?\n\nHOUR 4-5 (integration):\n - Mark all as read: optimistic update or wait for server confirmation?\n - Modal: what focus target receives focus when modal opens? (first focusable element or heading)\n - CSRF token source: from meta tag or request header? — follow existing app pattern\n\nHOUR 6+ (polish/tests):\n - Test fixture for partial-failure response (one panel missing)\n - Playwright: session expiry mid-page-load (401 on in-flight request)\n - Reduced-motion: CSS media query for skeleton animation\n```\n\n*Human implementation: ~6 hours. With CC+gstack: ~30-45 minutes.*\n\n### CEO Dual Voices\n\n**CODEX:** Disabled (`codex_reviews disabled`). Skipped.\n\n**CLAUDE SUBAGENT:** Completed. 6 findings:\n1. (Critical) Wrong-problem risk: bottleneck may be cognitive, not navigational\n2. (High) Aggregate endpoint partial failure undefined — **CONFIRMED** vs native\n3. (High) No rollout criteria — **CONFIRMED** vs native\n4. (Medium) Three equal quick actions may add noise (eligibility + ranking)\n5. (Medium) Activity feed prominence may be wrong for post-login intent\n6. (Medium) Modal adds friction; recommend direct button + undo toast\n\n```\nCEO DUAL VOICES — CONSENSUS TABLE:\n═══════════════════════════════════════════════════════════════════════\n Dimension Native Subagent Consensus\n ────────────────────────────────────── ────── ──────── ──────────\n 1. Premises valid? MOSTLY CHALLENGE DISAGREE → TASTE\n 2. Right problem to solve? YES UNCERTAIN DISAGREE → TASTE\n 3. Scope calibration correct? YES YES CONFIRMED\n 4. Alternatives sufficiently explored? MOSTLY NO DISAGREE → TASTE\n 5. Competitive/market risks covered? N/A N/A N/A (internal workspace)\n 6. 6-month trajectory sound? YES YES CONFIRMED\n═══════════════════════════════════════════════════════════════════════\nCONFIRMED = native + subagent agree. DISAGREE → surfaced at Final Gate.\nCodex: disabled. Subagent: in-host (model identity unknown).\n```\n\n**Taste decisions queued for Final Gate:**\n- T1: Premise validation first (session recordings) vs proceed-on-existing-data\n- T2: Layout order — Quick Actions first vs Activity-first per current plan\n- T3: Modal vs undo-toast for \"Mark all as read\"\n\n---\n\n### Section 1: Architecture\n\n```\nSYSTEM ARCHITECTURE:\n─────────────────────────────────────────────────────────────────────\nBrowser\n └─ /dashboard (UserDashboard.tsx)\n ├─ PanelWrapper → ActivityFeed ─┐\n ├─ PanelWrapper → NotificationsPanel ─┼─ all → GET /api/dashboard\n │ └─ Modal (existing dialog) │\n └─ PanelWrapper → QuickActions ─┘\n └─ Toast (new)\n\nGET /api/dashboard [new handler]\n ├─ Auth: existing cookie-session + workspace-membership middleware\n ├─ Promise.all([\n │ ActivityRepository.list(workspaceId, memberId) [existing]\n │ NotificationRepository.list(memberId) [existing]\n │ ActionRegistry.getActions(memberId) [existing]\n │ ])\n └─ Partial-failure composition [new — semantics unspecified ← GAP]\n\nPOST /api/notifications/mark-all-read [mutation — inferred]\n ├─ Auth + CSRF validation (existing)\n └─ NotificationRepository.markAllRead(memberId, snapshotTime) [existing]\n─────────────────────────────────────────────────────────────────────\n```\n\n**Data flow (four paths):**\n- Happy: parallel queries complete → { activity, notifications, quickActions } → rendered panels\n- Nil: no session cookie → 401 → middleware redirects to login\n- Empty: valid session, no items → { activity: [], notifications: [], quickActions: [action1, action2, action3] } → empty states (plan specifies ✓)\n- Error: one query throws → **PARTIAL FAILURE UNDEFINED ← CRITICAL GAP**\n\n**Panel state machine:**\n```\nLOADING ──→ SUCCESS\n ──→ ERROR ──→ LOADING (retry)\n ──→ EMPTY\n```\n\n**Auth boundary:** GET /api/dashboard is read-only, cookie session only. No workspace ID from query params (plan correct). CSRF required for mark-all-read mutation.\n\n**Rollback:** No schema changes. Feature flag flip = instant rollback. Git revert = 5 minutes.\n\n**CRITICAL GAP:** Partial failure semantics unspecified. → Implementation Task T1.\n\n### Section 2: Error & Rescue Map\n\n```\nMETHOD/CODEPATH WHAT CAN GO WRONG EXCEPTION CLASS\n──────────────────────────── ────────────────────────── ─────────────────────\nGET /api/dashboard handler Unauthenticated request UnauthenticatedError\n Forbidden workspace ForbiddenError\n DB connection pool exhaust ConnectionPoolExhausted\nActivityRepository.list() DB query timeout QueryTimeoutError\n Invalid cursor ValidationError\nNotificationRepository.list() DB query timeout QueryTimeoutError\nNotificationRepository DB timeout QueryTimeoutError\n .markAllRead() Invalid snapshot time ValidationError\n Stale/missing CSRF CSRFError\nActionRegistry.getActions() Registry misconfiguration RegistryError\n\nEXCEPTION CLASS RESCUED? RESCUE ACTION USER SEES\n──────────────────────── ───────── ───────────────────────── ─────────────────\nUnauthenticatedError Y 401 → middleware redirect Login page\nForbiddenError Y 403 response Error page\nQueryTimeoutError PARTIAL? Return partial response? Panel error state?\nConnectionPoolExhausted N ← GAP — 500 (silent) ← BAD\nValidationError Y 400 response Panel error toast\nCSRFError Y 403 Toast: \"Try again\"\nRegistryError N ← GAP — QuickActions crash ← BAD\n```\n\n**CRITICAL GAPS:** ConnectionPoolExhausted → silent 500; RegistryError → QuickActions panel crash. → Tasks T2, T3.\n\n### Section 3: Security\n\n- **Auth:** Existing middleware handles session + workspace. No new auth surface introduced.\n- **IDOR:** No workspace ID from query params (plan explicit). Repository methods use request-context IDs. No IDOR risk.\n- **Mutations:** \"Mark all as read\" requires CSRF (plan states). Idempotent. Time-snapshotted (later arrivals unread).\n- **Input validation:** Snapshot time for markAllRead needs validation (reject future timestamps). Cursor format validation if passed as query param.\n- **New secrets:** None.\n- **PII:** Notifications and activity data are existing PII surfaces. No new exposure.\n- **Audit logging:** **WARNING** — \"Mark all as read\" is a state-modifying action with no audit trail specified. Medium concern. → Task T4.\n\nNo high-severity security issues found. 1 warning.\n\n### Section 4: Data Flow & Interaction Edge Cases\n\n**Dashboard load flow:**\n```\nREQUEST → Auth middleware → handler → Promise.all(queries) → compose response → JSON\n │ │ │ │\n ▼ ▼ ▼ ▼\n[no cookie?] [invalid?] [one fails?] [partial result?]\n → 401 → 403 → UNDEFINED GAP → UNDEFINED GAP\n```\n\n**Interaction edge cases:**\n```\nINTERACTION EDGE CASE HANDLED? FIX\n─────────────────────────────────────────────────────────────────────\nMark all as read Double-click NO ← GAP Debounce/disabled state\n Network failure NO ← GAP Toast error + retry\n Session expired mid-click ? CSRF failure → 403 toast\nDashboard load Session expires mid-load ? In-flight 401 → redirect\nToast Multiple simultaneous NO ← GAP Toast queue/stack spec\nQuick action click Action ineligible at click OK Server-side gate handles\nNotification arrival New item during page view OK No real-time (v1 OK)\n```\n\n**GAPS:** Double-click guard (T5), network failure handling (T6), toast queue (T7). → Tasks.\n\n### Section 5: Code Quality\n\n- Component naming: `UserDashboard`, `ActivityFeed`, `NotificationsPanel`, `QuickActions`, `PanelWrapper` — clear, domain-matched.\n- DRY: PanelWrapper accepted in cherry-pick (avoids repeating loading/empty/error across 3 panels).\n- Aggregate endpoint partial-failure logic will be complex — needs internal comments explaining the policy once defined.\n- No over-engineering detected. No premature abstractions.\n- Risk: If partial-failure handler uses nested conditionals for each panel state, complexity could spike. Recommend a `composePanelResult(settled)` helper function to keep the handler flat.\n\nNo issues beyond what's tracked in Sections 1-4.\n\n### Section 6: Tests\n\n```\nNEW UX FLOWS:\n - Dashboard load (loading → success, loading → error, loading → empty per panel)\n - ActivityFeed render (20 items, empty state, error state)\n - NotificationsPanel render (items, empty, error, unread count)\n - Mark all as read (open modal → confirm, cancel, keyboard Escape)\n - Toast: success + error messages for mark-all-read\n - QuickActions render (eligible actions, ineligible hidden/disabled)\n - Post-login redirect to /dashboard (accepted cherry-pick)\n\nNEW DATA FLOWS:\n - GET /api/dashboard: happy path, partial failure (one panel fails), 401, 403, timeout\n\nNEW CODEPATHS:\n - UserDashboard.tsx page component\n - ActivityFeed, NotificationsPanel, QuickActions, PanelWrapper components\n - Modal confirmation flow\n - Toast notification system\n - Aggregate endpoint handler + parallel query composition\n - Partial failure response composition (once spec is defined)\n - Post-login redirect logic\n\nNEW INTEGRATIONS:\n - GET /api/dashboard (new endpoint, backed by existing repo methods)\n - POST /api/notifications/mark-all-read\n\nMISSING TESTS (gaps):\n - Panel-level retry: error state → LOADING → SUCCESS (RTL)\n - Toast queue: two simultaneous mark-all-read attempts (edge case)\n - Keyboard navigation Tab order through dashboard panels (RTL)\n - Screen reader: live region announces toast (axe-core + RTL)\n - Reduced-motion: skeleton animation disabled (@media prefers-reduced-motion)\n - Double-click guard: button disabled after first click (RTL)\n - Playwright E2E: full dashboard flow post-login\n - Playwright E2E: mark-all-read confirmation keyboard flow\n - Playwright E2E: session expiry mid-page-load (401 redirect)\n - Rollout criteria verification (p95 latency, interaction test pass rate)\n```\n\nTest coverage gaps → Tasks T8T12.\n\n### Section 7: Performance\n\n- **Parallel queries:** GET /api/dashboard runs ActivityRepository, NotificationRepository, ActionRegistry in parallel. p99 ≈ max(single slowest query). Good.\n- **Indexed access paths:** Plan confirms workspace/member and created-at indexes exist. No N+1 risk (20 flat records per panel, no relationship traversals).\n- **Connection pool:** 3 concurrent queries per dashboard request. Under 100 concurrent users = 300 simultaneous queries. Acceptable for moderate load; revisit at 1000 concurrent.\n- **ActionRegistry:** Memory-resident. Zero DB cost.\n- **Caching:** Deferred to TODOS.md (30s TTL would reduce DB load at scale; not blocking v1).\n- **Slow paths:** p99 = max(activity_query, notifications_query). Estimate: <100ms each on indexed tables with 20-record limit. Acceptable.\n\nNo blocking performance issues found for v1 scale.\n\n### Section 8: Observability\n\n**Gaps:**\n- **Analytics events:** Plan says \"still needs its own exposure and interaction instrumentation\" — but specific events not named. → Task T12.\n- **Endpoint metrics:** `api.dashboard.latency` histogram + `api.dashboard.error_rate` + status code breakdown needed on day 1.\n- **Partial failure logging:** When a panel fails, the handler must log: which repository failed, the error class, request context (member ID, workspace ID, request ID). Without this, debugging a production issue requires guessing which panel is slow.\n- **Toast/interaction logging:** mark-all-read click events need client-side analytics.\n\n**Day 1 observability wish list:**\n- `api.dashboard.latency` by panel status (all-success / partial / error)\n- `dashboard.exposure` event (member saw dashboard)\n- `dashboard.time_to_first_action_ms` (login → first quick action click)\n- `notifications.mark_all_read` count\n- Error rate alert: `api.dashboard.error_rate` > 1% → page on-call\n\n**Runbook:** dashboard 500s → check `api.dashboard.error_rate` → roll back via feature flag.\n\nLogging gaps → Task T13.\n\n### Section 9: Deployment & Rollout\n\n- **Schema migration:** None. Zero-downtime deploy.\n- **Feature flag:** Member-cohort rollout via existing staging flags. ✓\n- **Rollout order:** Deploy → enable flag for 5% cohort → monitor → expand.\n- **Rollback:** Feature flag flip (instant). Git revert (5 min).\n- **CRITICAL GAP: Rollout criteria not specified.** Plan says \"must be specified and added\" but they are not. What threshold triggers rollback? What p95 latency is acceptable? → Task T14.\n- **Post-deploy verification:**\n 1. Check GET /api/dashboard returns 200 for test account\n 2. Check partial-failure case (disable one repo method): panels degrade independently\n 3. Check mark-all-read works end-to-end\n 4. Check analytics events fire in staging\n 5. Check toast/modal keyboard accessibility\n 6. Check login redirect for existing users\n\nRollout criteria gap is the only critical deployment finding.\n\n### Section 10: Long-Term Trajectory\n\n- **Technical debt:** Toast system (new component) could be replaced by design-system toast later. PanelWrapper is reusable for future panels. Low debt.\n- **Path dependency:** Dashboard at /dashboard is a new route, no existing routes displaced. Easy to extend (dark mode, personalization add panels; real-time adds WebSocket to existing handler).\n- **Reversibility:** 5/5 — pure addition, no schema changes, feature flag.\n- **Platform potential:** ActionRegistry pattern is reusable. PanelWrapper is reusable. Aggregate endpoint becomes the foundation for personalization phase.\n- **1-year question:** Architecture will be clear to a new engineer. The only opacity is the partial-failure handler — document the policy in a comment once decided.\n- **Phase 2 trajectory:** dark mode (feature flag on existing Tailwind tokens), personalization (add to ActionRegistry), real-time (add SSE to existing endpoint). Clean trajectory.\n\nNo blocking long-term issues.\n\n### Section 11: Design & UX Review\n\n**Information architecture:**\n- Current plan order: Activity Feed, Notifications Panel, Quick Actions\n- Recommended order: **Quick Actions first**, then Notifications, then Activity (action-first since success metric is \"login-to-first-completed-task\")\n- Rationale: Activity is audit history (low urgency post-login); Quick Actions are the primary goal\n\n→ TASTE DECISION queued for Final Gate.\n\n**Interaction state coverage:**\n```\nFEATURE LOADING EMPTY ERROR SUCCESS PARTIAL-FAILURE\n─────────────────────────────────────────────────────────────────\nActivityFeed Y Y Y Y N/A (handled by PanelWrapper)\nNotifPanel Y Y Y Y N/A\nQuickActions Y N/A Y Y N/A\nDashboard (agg) Y N/A Y Y GAP — spec needed\nModal N/A N/A Y Y N/A\nToast N/A N/A Y Y N/A\n```\n\n**GAP: Dashboard partial state** — what renders when Activity loads but Notifications errors? Needs explicit spec (accepted panels render normally; failed panels show panel-level error state). → Task T15.\n\n**User journey (storyboard):**\n```\nLogin → /dashboard (skeleton loading, ~100ms)\n → Dashboard renders:\n [Quick Actions: resume assigned work | create item | invite member]\n [Notifications: 3 unread | Mark all as read button]\n [Activity: last 20 workspace changes]\n → User clicks \"resume assigned work\" → exits to existing workflow\n (success: <45s from login) ✓\n OR user clicks \"Mark all as read\" → [modal? or direct button?] → toast \"All caught up\"\n```\n\n**Accessibility:** Focus trapping in modal, Escape dismissal, live region for toast, contrast, reduced-motion — all explicitly specified. ✓\n\n**Mobile:** sm/md/lg breakpoints specified. Stacking order for mobile not specified — Quick Actions should stack first (same action-first rationale).\n\n**Delight opportunity:** After marking all read, a subtle empty-state transition (\"You're all caught up ✓\") before the panel goes quiet. 15min CC. → DEFER to TODOS.md.\n\n**AI slop risk:** Plan describes specific states per panel (empty state, loading skeleton, error state) with named components. Not generic. ✓\n\nRecommend running `/plan-design-review` after CEO review for deep UI/UX audit.\n\n---\n\n### CEO Required Outputs\n\n#### NOT in scope\n- Dark mode (deferred to separate plan)\n- Personalization / customization (deferred to separate plan)\n- Real-time updates via SSE/WebSocket (deferred — significant infra)\n- Dashboard response caching (deferred — revisit at scale)\n- Keyboard shortcuts (deferred P3)\n- Push notifications / digest email (separate system)\n- Admin tooling for notification debugging (deferred P3)\n- Delight animation for empty state after mark-all-read (deferred P3)\n\n#### What already exists\n(See Step 0B table above — all repository methods, dialog primitive, Tailwind tokens, page shell, analytics, feature flags, test fixtures, CI runners are reused.)\n\n#### Dream state delta\nThis plan moves from 3-page navigation to a consolidated dashboard, targeting 45s login-to-first-task. It leaves the 12-month ideal (real-time, personalization, mobile, smart prioritization) as future phases. The accepted cherry-picks (PanelWrapper, post-login redirect) set clean foundations for those phases.\n\n#### Error & Rescue Registry\n(See Section 2 table above. Critical gaps: ConnectionPoolExhausted, RegistryError.)\n\n#### Failure Modes Registry\n\n```\nCODEPATH FAILURE MODE RESCUED? TEST? USER SEES? LOGGED?\n────────────────────────── ────────────────── ──────── ───── ────────────── ───────\nGET /api/dashboard ConnectionPoolExh. N ← GAP N 500 (silent) N ← BAD\nGET /api/dashboard One panel timeout PARTIAL? N UNDEFINED N ← BAD\nGET /api/dashboard All panels fail ? N 500? UI error? N ← BAD\nmark-all-read Network failure N ← GAP N Nothing N ← BAD\nActionRegistry.getActions RegistryError N ← GAP N Panel crash N ← BAD\nmark-all-read Double-click N ← GAP N Duplicate req. N\n```\n\n**CRITICAL GAPS:** Rows with RESCUED=N, TEST=N, USER SEES=Silent → 5 critical gaps. → Tasks T1T6.\n\n#### Scope Expansion Decisions\n- Accepted: Post-login redirect, PanelWrapper\n- Deferred: Keyboard shortcuts, caching\n- Skipped: Real-time, delight animation\n\n---\n\n## CEO Implementation Tasks\n\n- [ ] **T1 (P1, human: ~2h / CC: ~15min)** — Aggregate endpoint — Specify and implement partial-failure policy\n - Surfaced by: Section 1 architecture + Section 2 error map + Subagent Finding 2\n - Spec: 200 response with `{ activity: Result, notifications: Result, quickActions: Result }` where each Result has `{ status: 'ok'|'error', data?: [...], error?: string }`. Failed panel → error status, panel renders error state.\n - Files: `src/api/dashboard.ts`, `src/pages/UserDashboard.tsx`\n - Verify: RTL test with one repository mock throwing → panel shows error, others render normally\n\n- [ ] **T2 (P1, human: ~30min / CC: ~5min)** — GET /api/dashboard — Handle ConnectionPoolExhausted explicitly\n - Surfaced by: Section 2 — silent 500 critical gap\n - Files: `src/api/dashboard.ts`\n - Verify: Test with pool-exhausted mock → returns 503 with retry-after header\n\n- [ ] **T3 (P1, human: ~30min / CC: ~5min)** — QuickActions — Handle RegistryError explicitly\n - Surfaced by: Section 2 — RegistryError unhandled, panel crashes\n - Files: `src/components/QuickActions.tsx`, `src/api/dashboard.ts`\n - Verify: Test with registry mock throwing → panel shows error state, not crash\n\n- [ ] **T4 (P2, human: ~1h / CC: ~10min)** — mark-all-read — Add server-side audit log entry\n - Surfaced by: Section 3 — mutation with no audit trail\n - Files: `src/api/notifications.ts`\n - Verify: Test confirms audit log entry created with memberId + snapshotTime\n\n- [ ] **T5 (P1, human: ~30min / CC: ~5min)** — NotificationsPanel — Double-click guard on mark-all-read\n - Surfaced by: Section 4 — race condition on double-click\n - Files: `src/components/NotificationsPanel.tsx`\n - Verify: RTL test simulates two rapid clicks → only one request fired\n\n- [ ] **T6 (P1, human: ~30min / CC: ~5min)** — NotificationsPanel — Network failure handling for mark-all-read\n - Surfaced by: Section 4 — no error handling for failed mutation\n - Files: `src/components/NotificationsPanel.tsx`\n - Verify: RTL test with fetch mock rejecting → toast \"Failed to mark as read. Try again.\"\n\n- [ ] **T7 (P2, human: ~1h / CC: ~10min)** — Toast — Define queue/stack behavior for simultaneous toasts\n - Surfaced by: Section 4 — multiple simultaneous toasts undefined\n - Files: `src/components/Toast.tsx`\n - Verify: RTL test fires two toasts simultaneously → both visible or stacked correctly\n\n- [ ] **T8T12 (P1/P2)** — Tests — Panel retry flow, keyboard nav, screen reader live region, reduced-motion, Playwright E2E\n - Surfaced by: Section 6 test coverage gaps\n - Files: `src/components/*.test.tsx`, `e2e/dashboard.spec.ts`\n - Verify: All tests pass in CI\n\n- [ ] **T13 (P1, human: ~1h / CC: ~10min)** — Observability — Structured logging for partial failures + analytics event spec\n - Surfaced by: Section 8 — no logging spec, analytics events unnamed\n - Files: `src/api/dashboard.ts`, `src/pages/UserDashboard.tsx`\n - Verify: Dashboard load fires `dashboard.exposure` event; failed panel logs member/workspace/error class\n\n- [ ] **T14 (P1, human: ~1h / CC: ~10min)** — Rollout — Specify and document advancement gates\n - Surfaced by: Section 9 + Subagent Finding 3 — no rollout criteria\n - Spec: p95 < 400ms at 10% rollout; login-to-first-task median < 60s at 50%; zero critical a11y failures in Playwright\n - Files: deployment runbook / feature flag config\n - Verify: Gates measurable in existing metrics dashboard\n\n- [ ] **T15 (P1, human: ~30min / CC: ~5min)** — UserDashboard — Spec partial-success render behavior\n - Surfaced by: Section 11 — dashboard partial state (some panels success, some error) not specified\n - Spec: Each PanelWrapper renders its own success/error state independently; no full-page error on partial failure\n - Files: `src/pages/UserDashboard.tsx`, `src/components/PanelWrapper.tsx`\n - Verify: RTL test with mixed panel results → partial render correct\n\n---\n\n### CEO Completion Summary\n\n```\n+====================================================================+\n| MEGA PLAN REVIEW — CEO COMPLETION SUMMARY |\n+====================================================================+\n| Mode selected | SELECTIVE EXPANSION |\n| System Audit | Clean repo; 1 plan file; no prior reviews |\n| Step 0 | SELECTIVE EXPANSION; Approach B; 2 accepted |\n| | cherry-picks (redirect + PanelWrapper) |\n| Section 1 (Arch) | 1 critical gap (partial-failure semantics) |\n| Section 2 (Errors) | 8 error paths mapped, 2 CRITICAL GAPS |\n| Section 3 (Security)| 0 high severity; 1 warning (audit log) |\n| Section 4 (Data/UX) | 3 edge cases unhandled |\n| Section 5 (Quality) | 0 blocking issues |\n| Section 6 (Tests) | Diagram produced, 8 gaps |\n| Section 7 (Perf) | 0 blocking issues |\n| Section 8 (Observ) | 2 gaps (logging, analytics event spec) |\n| Section 9 (Deploy) | 1 critical gap (rollout criteria) |\n| Section 10 (Future) | Reversibility: 5/5; debt items: 1 (toast) |\n| Section 11 (Design) | 1 taste (layout order); 1 gap (partial UI) |\n+--------------------------------------------------------------------+\n| NOT in scope | written (8 items) |\n| What already exists | written |\n| Dream state delta | written |\n| Error/rescue registry| 8 methods, 2 CRITICAL GAPS |\n| Failure modes | 6 total, 5 CRITICAL GAPS |\n| TODOS.md updates | 4 items (keyboard shortcuts, caching, |\n| | delight animation, admin tooling) |\n| Scope proposals | 5 proposed, 2 accepted, 2 deferred, 1 skip |\n| CEO plan | written |\n| Outside voice | Subagent: 6 findings; Codex: disabled |\n| Diagrams produced | 4 (system arch, data flow, state machine, |\n| | error & rescue, failure modes) |\n| Stale diagrams | 0 (no prior diagrams exist) |\n| Unresolved decisions | 3 TASTE DECISIONS queued for Final Gate |\n+====================================================================+\n```\n\n**Phase 1 complete.**\nCodex: disabled. Claude subagent: completed — 6 findings, 3 confirmed cross-model.\nConsensus: 2/6 confirmed, 3 disagreements → surfaced at Final Gate as Taste Decisions.\nPassing to Phase 2 (Design Review).\n",
"acceptedObligations": "- Use aggregate endpoint, parallel queries and independent per-panel failures.\n- Redirect authenticated members to /dashboard after login.\n- Add shared PanelWrapper for loading, empty and error states.\n- **T1 (P1, human: ~2h / CC: ~15min)** — Aggregate endpoint — Specify and implement partial-failure policy\n - Spec: 200 response with `{ activity: Result, notifications: Result, quickActions: Result }` where each Result has `{ status: 'ok'|'error', data?: [...], error?: string }`. Failed panel → error status, panel renders error state.\n - Files: `src/api/dashboard.ts`, `src/pages/UserDashboard.tsx`\n - Verify: RTL test with one repository mock throwing → panel shows error, others render normally\n\n- **T2 (P1, human: ~30min / CC: ~5min)** — GET /api/dashboard — Handle ConnectionPoolExhausted explicitly\n - Files: `src/api/dashboard.ts`\n - Verify: Test with pool-exhausted mock → returns 503 with retry-after header\n\n- **T3 (P1, human: ~30min / CC: ~5min)** — QuickActions — Handle RegistryError explicitly\n - Files: `src/components/QuickActions.tsx`, `src/api/dashboard.ts`\n - Verify: Test with registry mock throwing → panel shows error state, not crash\n\n- **T4 (P2, human: ~1h / CC: ~10min)** — mark-all-read — Add server-side audit log entry\n - Files: `src/api/notifications.ts`\n - Verify: Test confirms audit log entry created with memberId + snapshotTime\n\n- **T5 (P1, human: ~30min / CC: ~5min)** — NotificationsPanel — Double-click guard on mark-all-read\n - Files: `src/components/NotificationsPanel.tsx`\n - Verify: RTL test simulates two rapid clicks → only one request fired\n\n- **T6 (P1, human: ~30min / CC: ~5min)** — NotificationsPanel — Network failure handling for mark-all-read\n - Files: `src/components/NotificationsPanel.tsx`\n - Verify: RTL test with fetch mock rejecting → toast \"Failed to mark as read. Try again.\"\n\n- **T7 (P2, human: ~1h / CC: ~10min)** — Toast — Define queue/stack behavior for simultaneous toasts\n - Files: `src/components/Toast.tsx`\n - Verify: RTL test fires two toasts simultaneously → both visible or stacked correctly\n\n- **T8T12 (P1/P2)** — Tests — Panel retry flow, keyboard nav, screen reader live region, reduced-motion, Playwright E2E\n - Files: `src/components/*.test.tsx`, `e2e/dashboard.spec.ts`\n - Verify: All tests pass in CI\n\n- **T13 (P1, human: ~1h / CC: ~10min)** — Observability — Structured logging for partial failures + analytics event spec\n - Files: `src/api/dashboard.ts`, `src/pages/UserDashboard.tsx`\n - Verify: Dashboard load fires `dashboard.exposure` event; failed panel logs member/workspace/error class\n\n- **T14 (P1, human: ~1h / CC: ~10min)** — Rollout — Specify and document advancement gates\n - Spec: p95 < 400ms at 10% rollout; login-to-first-task median < 60s at 50%; zero critical a11y failures in Playwright\n - Files: deployment runbook / feature flag config\n - Verify: Gates measurable in existing metrics dashboard\n\n- **T15 (P1, human: ~30min / CC: ~5min)** — UserDashboard — Spec partial-success render behavior\n - Spec: Each PanelWrapper renders its own success/error state independently; no full-page error on partial failure\n - Files: `src/pages/UserDashboard.tsx`, `src/components/PanelWrapper.tsx`\n - Verify: RTL test with mixed panel results → partial render correct\n- Preserve the accepted test details:\n - Panel-level retry: error state → LOADING → SUCCESS (RTL)\n - Toast queue: two simultaneous mark-all-read attempts (edge case)\n - Keyboard navigation Tab order through dashboard panels (RTL)\n - Screen reader: live region announces toast (axe-core + RTL)\n - Reduced-motion: skeleton animation disabled (@media prefers-reduced-motion)\n - Double-click guard: button disabled after first click (RTL)\n - Playwright E2E: full dashboard flow post-login\n - Playwright E2E: mark-all-read confirmation keyboard flow\n - Playwright E2E: session expiry mid-page-load (401 redirect)\n - Rollout criteria verification (p95 latency, interaction test pass rate)\n",
"note": "Free regression derived from exact T record. Recorded implementation tasks retain all conditions/verification; finding-attribution lines are omitted as analysis, and accepted test detail is retained. No historical rescore."
}