mirror of
https://github.com/garrytan/gstack.git
synced 2026-09-16 09:55:29 +02:00
* feat: add a restricted and supervised Claude Code runner Preserve configured authentication and models while enforcing tool access, strict completion JSON, bounded output and process cleanup. Cover argv, failure handling, session metadata and Windows process containment. * feat: route outside reviews by harness and migrate wrapper installs Use Claude Code from Codex and Codex from other supported hosts, with shared invocation rendering, positive gate validation and per-phase provenance. Rename /claude to /claude-code, repair managed shared and copied installations safely, and generate native Kiro skills. Add installed-workflow, failure-injection and live cross-harness regression coverage. * test: recognize CEO mode labels without terminal spacing The paid workflow rendered SCOPEEXPANSION at option 4, but its driver required a literal space. Match the leading mode title without cursor-spacing artifacts and ignore adjacent preview text. Preserve missing-target failures and downstream posture assertions. * test: isolate plan-count fixtures before starting review workflows Seed the complete test plan in a private git repository before launching Claude, so a bare slash command cannot review the live workspace while a delayed fixture message remains queued. Preserve count thresholds, parsers and budgets. Add initial-context and installed-discovery tests, and retain startup/terminal diagnostics on failed evaluations. * test: stabilize review fixtures and Claude eval startup Preserve source boundaries in workflow judge inputs, isolate CEO mode plans, and wait for interactive trust input readiness. Keep startup failure evidence and retain existing models, budgets, and assertions. Co-Authored-By: OpenAI Codex <noreply@openai.com> * test: classify collapsed review modes and isolate seeded findings Keep review questions out of the setup count when terminal cursor positioning removes spaces. State existing webhook safeguards so the five-finding control measures its seeded defects without accidental extra security and concurrency gaps. Preserve question bands and the paired control. Co-Authored-By: OpenAI Codex <noreply@openai.com> * test: isolate browser daemon state across free shards Co-Authored-By: OpenAI Codex <noreply@openai.com> * test: stabilize native review counting and interactive navigation Co-Authored-By: OpenAI Codex <noreply@openai.com> * chore: prepare v1.82.0.0 release Co-Authored-By: OpenAI Codex <noreply@openai.com> * fix: eliminate browser and process-cleanup test flakes Pin every CI surface to Bun 1.4.0 to avoid extra-stdio finalizers closing reused live sockets. Add an isolated GC/listener regression that fails on Bun 1.3.13, and prevent coordinated rollback to an affected CI runtime. Check renderer cleanup against the render's own staging directory so concurrent renders cannot invalidate the assertion. Make the no-pgrep process-tree walk tolerate disappearing /proc entries, and synchronize its test fixture through child readiness and pipe EOF instead of sleeps. Validation: 9,157 passed, 31 skipped, zero failures across 556 files with retries disabled. Build, all-host generation freshness, and skill checks passed. All three races have failing-before/passing-after regressions. * fix: count completed native review questions in evals * fix: drive review navigation from confirmed native choices * fix: require complete section-loading eval reports * test: isolate telemetry HTTP transport from local assertions * fix: keep review input on the active native question * test: let tunnel revocation daemon choose an available port * test: allocate available ports for pairing and watchdog fixtures * fix: stabilize planning eval navigation and phase reporting * test: isolate installed runtime paths in planning evals * test: stabilize review evidence and concurrent refresh fixtures * fix: resolve design findings before editing the plan * fix: honor and persist disabled outside plan reviews * fix: preserve planning decisions and terminal evidence Load installed host reviews at autoplan phase entry and wait for completed reviewers and saved artifacts. Reuse approved remedies while preserving individual finding decisions. Drive interactive evals from the current terminal viewport, bind native questions across scrolling, and require complete native report evidence. Cover captured stale menus, permission lifecycles, setup classification, and disabled-review tool availability with deterministic regressions. Advance release metadata and the upgrade migration to the unclaimed 1.83.0.0 slot. * fix: drive native review questions and preserve current plans Use the native single-choice keyboard protocol and current terminal viewport, with per-question navigation inside packets and completed-call coverage. Keep permissions, multi-select menus, and Submit controls distinct. Send Autoplan reviewers the amended implementation plan, keep its review record separate, and supply retained application contracts in the chain fixture. Clarify individual DevEx decisions and complete CEO fix options; use one active plan destination for the section-loading report. * fix: preserve complete plan-review decisions * fix: recognize native plan dialogs and reviewer controls * fix: preserve review decisions and phase completion * fix: recognize completed reviews without losing findings * fix: preserve review continuity and native eval completion * test: fix native review completion and eval retry isolation * test: handle native review menus and complete eval fixtures * test: fix native review setup, completion, and isolation failures * test: limit native skill discovery to runtime assets * fix: bind Autoplan reviews to full ordered phase inputs * test: fix planning eval routing, counting, and timeout handling * chore: advance queued release to v1.84.0.0 * fix: preserve complete review inputs and planning decisions * fix: reconcile review approvals and preserve phase obligations * fix: preserve review obligations and unblock eval permissions Carry recorded Autoplan requirements into blind phase inputs, require Eng review approvals before exit, and exercise combined asynchronous flows in CEO reviews. Correct native finding and handoff classification and unblock repeated report edits using scoped request identities. * fix: retain plan requirements and complete native review dialogs * fix: complete native review prompts and retain plan references * fix: preserve review inputs and classify native eval evidence * fix: check competing completion orders in CEO reviews * fix: recognize review decisions and require phase methodology Require the current phase methodology before Autoplan snapshots. Correct substantive decision, closed handoff, and cache-finding classification, and honor the recommended implementation approach in native review dialogs. Add captured-transcript regressions without changing review thresholds, provider models, retries, or deadlines. * test: bind native review decisions and close completed handoffs * fix: complete review dialogs and verify methodology delivery * fix: preserve review evidence and unblock native eval prompts * fix: handle native review question completions * fix: recognize native review narration and controls * fix: count native review decisions and isolate eval fixtures * test: verify seeded review coverage and current artifact permissions * test: isolate model and brain-aware skill renders * fix: repair native workflow evaluation and clarify review steps * fix: stabilize workflow eval evidence and review guidance * test: repair native workflow observation and fixture isolation * fix: recognize completed workflow evidence and owned skill reads * test: repair seeded workflow delivery and completion evidence * test: recognize current review evidence across native forms * test: handle native review variants and permission redraws * fix: honor review preferences and recognize native eval evidence * test: recognize completed review decisions and queued permissions * test: match current review contracts and partial-line edits * test: recognize completed workflow evidence and bounded human waits * fix: preserve review entry gates and native eval interactions * fix: recognize native workflow evidence and preserve review gates * test: recognize current review evidence and preconfigure workflow fixtures * test: recognize completed review findings and scoped artifact permissions * fix: stabilize native workflow review and permission evidence * fix: recognize current review evidence and scoped edit confirmations Clarify Design and engineering review entry instructions and Design scoring. Recognize required legacy coverage and public Autoplan completion recaps. Bind the pending Edit confirmation to its exact file, ordered digest, and one-request approval when a preceding command display remains visible. Keep reviews within their existing size limits and preserve scope gates when extracting workflow fixtures from either supported preamble header. Keep failure outcomes, review thresholds, provider choices, and eval budgets. * fix: recover review workflow progress and eval evidence * fix: recognize valid review evidence and scope selection * test: fix review evidence parsing and repeated artifact prompts * test: recognize valid review decisions and pending native cards * fix(plan-eng-review): keep final navigation consistent with approved tasks * test: recognize valid review evidence and bind legacy diff requests * fix: stabilize review eval evidence and harness repair guidance * docs: update project documentation for v1.85.0.0 Co-Authored-By: OpenAI Codex <noreply@openai.com> * test: fix Windows CI fixtures and credential scan Rebase captured JSON values and filesystem evidence using the appropriate path convention. Compile native fake CLIs on Windows and synchronize pipe holder readiness, with cleanup retained when assertions fail. Assemble synthetic credential fixtures at runtime so the added-line scan keeps enforcing the same gate without flagging its own rejection controls. Discover generated skills directly for the empty-find regression check, avoiding a recursive scan through saved evaluation artifacts and dependencies. * fix: preserve source renders on Windows Compare canonical generator paths using native separators so an output sidecar pointing at the source cannot overwrite its skill or metadata. Keep the regression fixture isolated from the real checkout and expose freshness diagnostics before asserting subprocess status. Detach Windows drain-test pipe holders from the fake provider's automatic child cleanup while preserving the enclosing runner job and its assertions. * fix: clarify outside review fallback and CEO decisions Render one applicable own-harness fallback path and retain native review, disabled policy, and missing-coverage semantics. Align report field names and mode labels, and make the existing per-cut scope approval explicit. Regenerate skill outputs and keep the workflow judge's model, thresholds, and retry policy unchanged. * chore: move release to free version slot (v1.86.0.0) PR #2852 now claims v1.85.0.0. Align the release metadata and rename migration so upgrades from that version still receive it. Co-Authored-By: OpenAI Codex <noreply@openai.com> * fix: include engineering review prerequisites and restore branch context * fix: recognize coverage diagrams and clarify design review instructions * fix: preserve file identities and join Windows test processes --------- Co-authored-by: OpenAI Codex <noreply@openai.com>
94 lines
49 KiB
JSON
94 lines
49 KiB
JSON
{
|
||
"cwd": "/tmp/gstack-paid-shard-MPtf0T/tmp/gstack-autoplan-chain-LukJXD",
|
||
"config": "/tmp/gstack-paid-shard-MPtf0T/tmp/gstack-hermetic-3587976-PXLrEr/with-skills/.claude",
|
||
"stateRoot": "/tmp/gstack-paid-shard-MPtf0T/tmp/gstack-hermetic-3587976-PXLrEr/skill-home-cMZrbW/.gstack",
|
||
"hook": {
|
||
"version": 1,
|
||
"cwd": "/tmp/gstack-paid-shard-MPtf0T/tmp/gstack-autoplan-chain-LukJXD",
|
||
"config": "/tmp/gstack-paid-shard-MPtf0T/tmp/gstack-hermetic-3587976-PXLrEr/with-skills/.claude",
|
||
"stateRoot": "/tmp/gstack-paid-shard-MPtf0T/tmp/gstack-hermetic-3587976-PXLrEr/skill-home-cMZrbW/.gstack",
|
||
"seenIds": [
|
||
"toolu_019ngUUTCGyDufyLQaN9Tvgm"
|
||
],
|
||
"pending": {
|
||
"source": "pre_tool_use",
|
||
"sessionId": "037c9f5e-7035-45e2-935d-0b950c435930",
|
||
"toolUseId": "toolu_019ngUUTCGyDufyLQaN9Tvgm",
|
||
"tool": "Edit",
|
||
"file": "/tmp/gstack-paid-shard-MPtf0T/tmp/gstack-hermetic-3587976-PXLrEr/skill-home-cMZrbW/.gstack/projects/gstack-autoplan-chain-LukJXD/ceo-plans/2026-09-10-user-dashboard.md",
|
||
"timestamp": "2026-09-10T14:35:35.363Z",
|
||
"transcriptPath": "/tmp/gstack-paid-shard-MPtf0T/tmp/gstack-hermetic-3587976-PXLrEr/with-skills/.claude/projects/-tmp-gstack-paid-shard-MPtf0T-tmp-gstack-autoplan-chain-LukJXD/037c9f5e-7035-45e2-935d-0b950c435930.jsonl",
|
||
"editDigest": {
|
||
"version": 1,
|
||
"beforeSHA256": "6771404ca60011824af336f0a7011e3aba465b8728c011eef935a4b847e72bdd",
|
||
"requestSHA256": "7ea64ad22d29d9c779b6e945119ea238e2f40862a368ca518100626a1e7978de",
|
||
"oldLineHashes": [
|
||
"95b34a72ab3b36711dc4f7cffc93cfe9e4ac45adaff3cbe06170576bb2b7787d"
|
||
],
|
||
"newLineHashes": [
|
||
"b69061a320491c4d23ea3b0e11040b1edfdf54c90d67287236c5b49695b74a6d",
|
||
"f01cbe18eb0051492cb3e8740542a491038881c5f23dac05003dfa84a4a59087",
|
||
"2a23528279bef0c9dcde730b41c9cb25768c8d6de1d40195181e48efa8e6a295",
|
||
"8110f2d8d27cb7e397562be0be38bc995e33cdf5d570d3d18aae23fe83db5116",
|
||
"4fef7666a0b82562e292a3d53f26f0721c09a74ceb0a30a970351f213468b07e",
|
||
"2bbfa283b38391e7d66c205292efb1e8565120fa337a098863f4758576fe1239",
|
||
"be4ed8d0778c7ba2036e19657ac496f2b24ee5d3ceb4d1240c3f7993567f4f89",
|
||
"861a718d78d38dd744ab59540c3139c5283d0e67a3deed5b41138f9581084d34",
|
||
"418ec10dc288464be75940c80f56faa734f5380ed8c2a957dcc52ed75f60f942",
|
||
"a59e48fe92bfcea3ec1cad39996f1a01039861faaa75daa251be8d64b102abdc",
|
||
"c701beb15c91ba566b32cd004f4ba46eb8f317061549c0f699041c151065a387",
|
||
"44aa4bd0d2a4cc4cbce51a9ea6bc485d5fb2cfbebfc01a1fe275d8e165f51c92",
|
||
"e8fd6eaf078d6ccc6028e6213e12acfa82511902136e4aec48e1398990582884",
|
||
"e3b0c44298fc1c149afbf4c8996fb92427ae41e4649b934ca495991b7852b855",
|
||
"cfc849f68e95805d8b84b46121a94a5ee70daf8fe374e20224d88a4290553089",
|
||
"1e619e0fe087d96a85514e4f9f741b2aae28e3c65a73f3833b6c1869e282ed5b",
|
||
"0ef460f920e57467c135c9b9605f343e189fc4397565d14b12d0fc9ff5f31e4f",
|
||
"d9134cd12c765f0e72b00438c6f70fa3af52d4a4602dfb27c8ae4fe29a7ee4cc",
|
||
"e3b0c44298fc1c149afbf4c8996fb92427ae41e4649b934ca495991b7852b855",
|
||
"95b34a72ab3b36711dc4f7cffc93cfe9e4ac45adaff3cbe06170576bb2b7787d"
|
||
]
|
||
}
|
||
},
|
||
"sessionId": "037c9f5e-7035-45e2-935d-0b950c435930"
|
||
},
|
||
"before": "---\nstatus: ACTIVE\n---\n# CEO Plan: User Dashboard Page\nGenerated by /plan-ceo-review (via /autoplan) on 2026-09-10\nBranch: main | Mode: SELECTIVE EXPANSION\nRepo: gstack-autoplan-chain-LukJXD (local, no remote)\n\n## Problem and goal\n\nMembers of the existing single-role workspace visit three separate pages after login to resume work, check alerts, and inspect recent changes. One team walkthrough measured a median of 75 seconds to find the next item. The dashboard at `/dashboard` is meant to cut the median login-to-first-completed-task time from ~75 s to 45 s (a 40 % reduction), with completed-task rate and permission-error rate as guardrails.\n\n**Primary ship metric:** median login-to-first-completed-task time, treatment vs control. **Ship threshold:** ≥ 40 % reduction (equivalent to 75 → 45 s) sustained over the ≥ 2-week measurement window with no guardrail regression. **Diagnostic split (not a ship gate):** login→first-action-start (the leg the dashboard moves) and action-start→completion (unchanged by this plan), both derivable from the existing login / action start / action completion events, so the read-out attributes the delta correctly. The production baseline for all three is computed from existing events before the cohort opens; the 75 s walkthrough figure is not load-bearing.\n\n**Experiment design.** Assignment is a deterministic hash of member id, fixed for the experiment's duration. Default split 10 % treatment / 90 % control. Minimum sample: ≥ 300 logins per arm within the window; if the workspace's active-member count makes that unreachable at 10 % (record the count at approval), switch to a 50/50 split, and if still unreachable, fall back to a before/after comparison over the same window and say so in the read-out. `/dashboard` is flag-gated (below), so control members cannot reach it and contaminate the read-out.\n\n## Vision\n\n### 10x Check\nThe dashboard becomes the member's assistant rather than a summary page: a ranked \"next item\" (\"You have 2 items due today\"), alerts grouped by source, small tasks completed inline, live updates as teammates act, and a home that other features publish panels into. Concrete shape: ranking service over the action registry, a panel registry, a push channel. Human ~4-6 weeks / CC ~2-3 days. Out of this plan's blast radius; the personalisation plan owns it. This plan lays the seam (per-panel envelopes + a shared panel shell) without building the platform.\n\n### Platonic Ideal\nSkipped (SELECTIVE EXPANSION).\n\n## Baseline held (plan scope)\n- `/dashboard` page with Quick actions, Notifications, Recent activity panels; Tailwind, mobile-first (sm/md/lg).\n- Loading skeleton, empty, error, success per panel; hover and focus-visible on every control.\n- \"Mark all as read\" confirmation modal on the existing dialog primitive.\n- Toast feedback for actions.\n- `GET /api/dashboard` aggregate over existing PostgreSQL tables; no schema change.\n- Out of scope by the user: dark mode, personalisation.\n\n## Implementation approach chosen\nApproach A: aggregate endpoint composed server-side from the existing list methods and action registry, returning per-panel result envelopes so one failing panel never sinks the page. Alternatives: B client-side composition of three existing endpoints (7/10, close second, surfaced as taste decision T1); C smart post-login redirect + shell badge (4/10 against the stated goal, rejected).\n\n**If the Final Gate selects B instead of A:** the \"Endpoint\", \"Response shape\", and \"Retry\" paragraphs below are replaced by three per-panel fetch hooks wrapping the existing list endpoints and a small eligibility endpoint over the registry; each hook produces the same `Envelope<T>` client-side, so panel states, empty CTAs, mark-all-read, latest-wins, redirect precedence, caching headers, toast, accessibility, analytics, and rollout are unchanged. Composer tests become hook tests. **If the gate selects dashboard-local for T3:** the toast code moves under `src/components/dashboard/` with the same API and the same live-region rule; nothing else changes.\n\n## Verified preconditions (confirm before implementation; each has a fallback)\n- **Bulk-read API accepts a snapshot time.** The plan's contracts section states it \"marks only notifications at or before the supplied snapshot time\". Confirm the parameter name and type. If it turns out not to exist, mark-all-read cannot be made safe against later arrivals and the modal flow is **blocked** until the API gains the parameter (this becomes the same API work E8 needs; raise at the gate, do not ship a clock-unsafe version).\n- **Bulk-read API returns the affected count.** Confirm; otherwise the success toast reads \"Marked all as read\" without a number.\n- **Feature-flag system supports stable per-member assignment** (hash of member id). Confirm; otherwise implement the hash in the redirect module and store only the flag's on/off and percentage in the flag system.\n- **Analytics pipeline already emits login, action start, action completion, permission error** with member id and timestamps queryable for the baseline. Confirm; otherwise the baseline query is the first task and the cohort does not open until it runs.\n- **Metrics and alerting stack exists** (the plan names existing request/error metrics). Confirm the alert-rule mechanism; otherwise alerts are the first observability task.\n- **Activity list method preloads actor display names.** Confirm; otherwise the composer batches one lookup (never per row).\n- **Notification list can return a DB-clock timestamp** (`SELECT now()` in the same statement or transaction). Confirm; otherwise add a read-only wrapper that does.\n\n## Contract details (assume Approach A; see the B fallback above)\n\n**Route gating.** `/dashboard` is served only to members whose `dashboard_landing` assignment is on; others are redirected to the current landing page (302). No nav link is added in v1; treatment members reach the page through the post-login redirect and by URL. `dashboard_viewed{cohort}` records the member's assignment.\n\n**Endpoint.** `GET /api/dashboard`. Authentication and workspace membership are checked by the existing middleware *before* the handler runs: no session → 401, non-member → 403, both short-circuit the whole request (the client redirects to login or shows the access message). Flag-off members receive 404 from this route. Per-panel envelopes cover upstream failures only.\n\n**Response shape.**\n```\nDashboardResponse = {\n activity: Envelope<{ rows: ActivityRow[] }>,\n notifications: Envelope<{ rows: NotificationRow[], unreadTotal: number }>,\n quickActions: Envelope<{ rows: QuickAction[] }>\n}\nEnvelope<T> = { ok: true, fetchedAt: string, data: T }\n | { ok: false, fetchedAt: string, error: { code: 'timeout' | 'upstream_error', retryable: true, requestId: string } }\n\nDashboardUnavailable = { error: { code: 'dashboard_unavailable', retryable: true, requestId: string } } // 503 body only, not an envelope\n```\n`fetchedAt` is per envelope. For notifications it is the **database clock** (`now()` read in the same statement or transaction as the list and count), so a row whose `created_at` is later than `fetchedAt` by the same clock is never covered by a snapshot taken from it. For activity and quick actions it is the app-server time before the call (informational only). The composer runs the three calls with `Promise.allSettled` and a 2 s budget per call enforced **both** at the composer (race) and at the query (`statement_timeout` of 2 s or the driver's `AbortSignal`), so a timed-out query releases its pool connection instead of leaking it. The notifications count runs inside the notifications call and shares its timeout; if the count fails the notifications panel is `ok: false`. If all three fail the response is 503 with `DashboardUnavailable`; otherwise 200. A registry predicate that throws is caught: that action is omitted, one error log line with `actionId` is written, `dashboard_action_predicate_error_total{actionId}` is incremented, and the panel still succeeds. `unreadTotal` comes from a member-scoped read-state count over the existing indexed path; if no repository count method exists, a read-only count method is added (a new read query, not a schema or mutation change).\n\n**Limits and ordering.** Activity and notifications return at most 5 rows. Notifications: unread first, then `created_at` desc. Activity: `created_at` desc. Quick actions: all eligible actions in registry order (the registry holds 3; hard cap 5 for safety); Quick actions has **no** \"See all\". Activity and Notifications each have a \"See all\" link to the existing full page, which owns older-page navigation.\n\n**Retry.** A panel's Retry button refetches the whole aggregate (one request). Only panels that were in ERROR show a skeleton; panels that already have data enter `refreshing`.\n\n**Panel states.** `LOADING` (skeleton, `aria-busy`), `EMPTY` (copy + CTA), `ERROR` (message + Retry), `SUCCESS`, and `refreshing` = previous data stays rendered with a subtle indicator; used for post-mutation refetch and for Retry when the panel already has data. State is derived only from the last settled envelope, so SUCCESS→EMPTY or ERROR→SUCCESS cannot happen without a fetch.\n\n**Empty-state CTAs.** Activity empty → \"Create your first item\" linking to the create action's route when that action is present in `quickActions.data.rows`; otherwise link to the existing full activity page. Notifications empty → \"All caught up\" with a link to the full notifications page. Quick actions empty → \"Nothing to do right now\" (no CTA).\n\n**Mark all as read.** The trigger is rendered only when the notifications panel is `SUCCESS` and `unreadTotal > 0`; in every other state (LOADING, EMPTY, ERROR, refreshing) it is absent. Confirmation dialog on the existing dialog primitive (focus trap, Escape, focus return). Confirm button disabled while the request is pending; the server API is idempotent regardless. The request calls the existing member-scoped bulk-read API with the notifications envelope's `fetchedAt` as the snapshot, never the client clock. No new mutation API. On success: toast \"Marked N notifications as read\" where N is the count returned by the bulk-read response (or \"Marked all as read\" if the API returns no count), then a refetch. On failure: 401 → login; 403 (including CSRF) → the existing client refreshes the token once and retries, then an error toast; validation / retryable / network → error toast with a Retry action.\n\n**Latest-wins fetching.** One `AbortController` per request. Starting the mutation aborts any in-flight GET. Each request carries a sequence number; only the newest may commit state. This excludes the schedule where a slow pre-mutation GET lands after the post-mutation GET and repaints stale unread rows.\n\n**Post-login redirect precedence.** (1) A `next` value that is a same-origin relative path → go to `next`. An absolute, protocol-relative, or foreign-origin `next` is treated as absent and logged at warn. (2) Otherwise, if `dashboard_landing` is on for this member → `/dashboard`. (3) Otherwise → the current landing page. If the flag lookup itself fails, treat the flag as off (step 3).\n\n**Caching and limits.** `Cache-Control: private, no-store` on the GET. The route joins the same rate-limit tier as the existing authenticated list endpoints. Error envelopes carry the request id so a report can be joined to server logs.\n\n**Toast primitive.** `ToastProvider` + `useToast` under `src/components/toast/`, mounted at the app root so the `aria-live=\"polite\"` region exists before the first toast. Queue holds at most 3; a 4th evicts the oldest. Informational toasts auto-dismiss after 5 s; toasts carrying an action (Retry) persist until dismissed or acted on. Respects `prefers-reduced-motion`. In development, `useToast()` outside the provider throws, and a toast with an empty message throws; in production both are no-ops with `console.error`.\n\n**Accessibility.** Named controls, focus-visible outlines, 44 px touch targets on action buttons, reduced motion on skeleton shimmer and toast motion, axe passes on every panel state and with the dialog open, keyboard-only mark-all-read passes in Playwright.\n\n## Scope Decisions\n\n| # | Proposal | Effort (human / CC) | Decision | Reasoning |\n|---|----------|---------------------|----------|-----------|\n| E1 | Per-panel result envelopes with Retry | S (~4h / ~20m) | ACCEPTED | closes the plan's own open partial-failure work |\n| E2 | Preserve intended destination; `/dashboard` only when none | S (~1h / ~10m) | ACCEPTED | protects deep links and bookmarks; one file |\n| E3 | Unread-first ordering + `unreadTotal` in header | S (~2h / ~10m) | ACCEPTED | one sort, one count, one span |\n| E4 | \"See all\" link (activity, notifications) to the existing full page | S (~1h / ~5m) | ACCEPTED | full pages already own pagination |\n| E5 | Empty-state copy pointing at the relevant quick action (with fallback) | S (~1h / ~5m) | ACCEPTED | empty states are features |\n| E6 | Relative timestamps with absolute time for assistive tech | S (~1h / ~5m) | ACCEPTED | small, on the row component |\n| E7 | Toast built as a shared app primitive, not dashboard-local | S (+1h / +10m) | ACCEPTED (taste T3) | second consumer is inevitable; a11y policy requires a live region |\n| R1 | Required-to-ship measurement block (flag gating, events, baseline, logs, metrics, alerts, staging load check, test matrix) | M (~3 days / ~1.5h) | ACCEPTED | without it the ship rule cannot be evaluated |\n| E8 | Undo after Mark all as read | M | DEFERRED | needs a new \"unmark since snapshot\" mutation; the existing bulk-read API only marks read |\n| E9 | Live / polling notification updates | L | DEFERRED | new infra, outside blast radius |\n| E10 | Keyboard shortcuts (`g d`, `j/k`) | S | DEFERRED | no existing shortcut map to extend |\n| E11 | Stale-while-revalidate client cache | S | DEFERRED | private data; `no-store` in v1 |\n| E12 | Ranking / next-item recommendation | XL | SKIPPED | user's stated non-goal (personalisation plan) |\n\n## Accepted Scope (added to this plan)\nEverything under \"Contract details\" and \"Verified preconditions\" above, plus:\n\n**Required to ship (R1; gates the cohort):**\n- `dashboard_landing` flag with deterministic per-member assignment, route gating, and flag-flip rollback; no migration; endpoint additive.\n- Analytics events: `dashboard_viewed{cohort}` once per page view; `dashboard_panel_state{panel,state}` once per page view per panel at its first terminal state (EMPTY / ERROR / SUCCESS) only; `dashboard_action_clicked{actionId}`; `dashboard_see_all_clicked{panel}`; `dashboard_mark_all_read{result}`.\n- Production baseline for login-to-first-completed-task, login→action-start, action-start→completion computed from existing events before the cohort opens; active-member count recorded and the split chosen per the experiment design.\n- The pre-registered ship rule above, with a named owner for the read-out recorded at approval.\n- Structured request log per GET (member id, workspace id, request id, per-panel status and ms, total ms) and per mark-all-read.\n- Metrics `dashboard_request_total{panel,status}`, `dashboard_panel_duration_ms{panel}`, `dashboard_mark_all_read_total{result}`, `dashboard_action_predicate_error_total{actionId}`, DB pool wait on the route; production alerts: 5xx or `dashboard_unavailable` > 1 % for 5 min, p95 > 800 ms for 10 min, mark-all-read error rate > 5 %.\n- Tests: composer settle matrix + all-fail 503 + predicate-throw + query-timeout-releases-connection; envelope→state mapper; panel, dialog, toast, latest-wins (controlled promise release), redirect precedence, route gating unit/RTL suites; integration for member 200 / other-workspace 403 / no-session 401 / flag-off 404 / service failure / `no-store` header; Playwright login→dashboard→action, 503→Retry, keyboard-only mark-all-read, axe on every state.\n- Rollout order: dark deploy → internal members on staging (Playwright + axe suite + a staging load check p95 < 300 ms at 50 rps on staging fixtures) → treatment cohort for the ≥ 2-week measurement window → decision per the ship rule.\n\n**Operability polish (same PR, does not gate the cohort):**\n- `Server-Timing` per panel on the GET.\n- JSDoc on `composeDashboard`, `DashboardPanel`, and `ToastProvider` carrying the architecture / panel-state / envelope-rationale diagrams. No separate README.\n\nThe p95 latency target is enforced on staging before cohort expansion and in production alerting, not as a per-PR CI assertion (CI runners are too noisy for a latency SLO).\n\n## Deferred to TODOS.md\n- Undo after Mark all as read — blocked on a new \"unmark since snapshot\" mutation API (P2).\n- Live notification updates — needs a push/polling infra decision (P3).\n- Dashboard keyboard shortcuts — needs an app-wide shortcut map (P3).\n- SWR client cache for the dashboard payload — needs a privacy review of client caching (P3).\n\n## Taste decisions carried to the Final Gate\n- T1 aggregate endpoint (A) vs client-side composition (B) — open; the B fallback paragraph above states exactly what changes.\n- T2 confirmation modal for Mark all as read — recorded decision, surfaced for information: the independent reviewer preferred direct action + undo, but undo (E8) needs a new mutation API and is deferred, so the modal stays unless E8 is pulled into scope.\n- T3 toast as a shared primitive vs dashboard-local — open; the T3 fallback sentence above states what changes.\n- Premise note: the 75 s baseline is one team walkthrough; the ship rule is relative (≥ 40 %) and the analytics/rollout task measures the true baseline, so the launch does not depend on the figure.\n",
|
||
"viewport": " +t. \n 139 +- **Snapshot consistency.** Notifications rows, `unreadTotal`, and `now()` come from a single statement (CTE) so t\n +hey share one READ COMMITTED snapshot. Query budget via `SET LOCAL statement_timeout = '2s'` inside that statement\n +'s transaction (or the driver's per-query timeout option); the \"timed-out query releases its connection\" test asse\n +rts on pool metrics, not on the promise alone. \n 140 +- **Route gating in a client-rendered app.** The page route performs a client-side redirect to the current landing\n + after the flag lookup (no server 302 unless the app already server-renders routes); `GET /api/dashboard` returnin\n +g 404 for flag-off members is the authoritative gate. \n 141 +- **Forcing panel states in tests.** Playwright uses `page.route()` interception on `/api/dashboard` and the bulk-\n +read route to inject 503 / per-panel `ok:false` / empty payloads; no test-only headers in production builds. \n 142 +- **Identifiers.** The create action is `actionId = 'create_item'` (from the registry's stable ids); `fetchedAt` i\n +s ISO 8601 UTC with milliseconds; the 44 px target rule applies to every interactive control inside the panels (ac\n +tions, Retry, See all, Mark all as read). \n 143 + \n 144 +## Reviewer Concerns (unresolved after 3 iterations) \n 145 +- **Snapshot window (r3 #5).** A notification inserted in a transaction whose `now()` predates `fetchedAt` but whi\n +ch commits after the dashboard read has `created_at <= fetchedAt`, is invisible to the list, and is still marked r\n +ead. Accepted and documented: the window is the duration of an in-flight insert transaction (milliseconds), the ex\n +isting bulk-read contract is clock-based, and changing it to id-based is the E8-adjacent API work already deferred\n +. Record as a known limitation in the JSDoc on the mark-all-read handler. \n 146 +- **Provisional approach vs open taste items (r2 #4, r3 #6).** The reviewer wants T1/T3 closed or A/E7 marked prov\n +isional. They are provisional by design: /autoplan surfaces taste decisions at one Final Gate and the fallback par\n +agraphs define what changes. Not a contradiction to resolve in this document. \n 147 +- **Population size (r3 #14).** The single-workspace population may be too small for the default design; the fallb\n +acks exist, and the member count is now recorded before R1 starts. \n 148 + \n 149 ## Taste decisions carried to the Final Gate\n 150 - T1 aggregate endpoint (A) vs client-side composition (B) — open; the B fallback paragraph above states exactly w\n hat changes.\n 151 - T2 confirmation modal for Mark all as read — recorded decision, surfaced for information: the independent review\n er preferred direct action + undo, but undo (E8) needs a new mutation API and is deferred, so the modal stays unle\n ss E8 is pulled into scope.\n╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌\n Do you want to make this edit to 2026-09-10-user-dashboard.md?\n ❯ 1. Yes\n 2. Yes, and switch to accept edits (auto-approve file edits and common file commands) for this session (shift+tab)\n 3. No\n\n Esc to cancel · Tab to amend\n",
|
||
"request": {
|
||
"old_string": "## Taste decisions carried to the Final Gate",
|
||
"new_string": "## Iteration-3 amendments (override earlier wording where they conflict)\n- **Guardrail bounds.** Completed-task rate in treatment must not be lower than control by more than 2 percentage points; permission-error rate must not be higher than control by more than 1 percentage point, both over the same window.\n- **Control-arm assignment recorded.** Emit `dashboard_assignment{cohort, hashVersion}` on the existing login event for every member (both arms); record the hash salt/version at approval. `dashboard_viewed` remains treatment-only.\n- **Ship-rule estimator.** Unit of analysis is the member (first login-to-first-completed-task per member per day, then per-member median). Estimator: bootstrap 95 % CI on the difference of medians; ship when the CI lower bound shows ≥ 40 % reduction. \"≥ 300 logins per arm\" becomes \"≥ 300 member-days per arm\".\n- **Record the active-member count before starting R1**, not after; choose the split then.\n- **Whole-response failures on the GET.** 503 `DashboardUnavailable` or a network/parse failure → all three panels ERROR with one shared page-level Retry; 404 (assignment flipped off mid-session) → client redirects to the current landing; 401 → login. Envelope codes remain per-panel only.\n- **Panel state model.** `{ status: LOADING | EMPTY | ERROR | SUCCESS, refreshing: boolean }`. `status` derives only from the last settled envelope; `refreshing` is set at request start when the panel already has a settled status other than LOADING, and cleared on settle. EMPTY panels refresh with their empty copy still rendered. The Mark-all-read trigger stays rendered but disabled while `refreshing`; it is rendered only when `status === SUCCESS && unreadTotal > 0`.\n- **CSRF vs membership 403.** Retry-once applies only when the existing client's typed error indicates a CSRF failure (its existing discriminator, body code or header); a membership 403 shows the access message with no retry.\n- **Preconditions heading** is to be read as \"Preconditions to verify before implementation\". The bulk-read semantic requirement is: it marks read only notifications with `created_at <= snapshot` by the DB clock and is idempotent.\n- **Snapshot consistency.** Notifications rows, `unreadTotal`, and `now()` come from a single statement (CTE) so they share one READ COMMITTED snapshot. Query budget via `SET LOCAL statement_timeout = '2s'` inside that statement's transaction (or the driver's per-query timeout option); the \"timed-out query releases its connection\" test asserts on pool metrics, not on the promise alone.\n- **Route gating in a client-rendered app.** The page route performs a client-side redirect to the current landing after the flag lookup (no server 302 unless the app already server-renders routes); `GET /api/dashboard` returning 404 for flag-off members is the authoritative gate.\n- **Forcing panel states in tests.** Playwright uses `page.route()` interception on `/api/dashboard` and the bulk-read route to inject 503 / per-panel `ok:false` / empty payloads; no test-only headers in production builds.\n- **Identifiers.** The create action is `actionId = 'create_item'` (from the registry's stable ids); `fetchedAt` is ISO 8601 UTC with milliseconds; the 44 px target rule applies to every interactive control inside the panels (actions, Retry, See all, Mark all as read).\n\n## Reviewer Concerns (unresolved after 3 iterations)\n- **Snapshot window (r3 #5).** A notification inserted in a transaction whose `now()` predates `fetchedAt` but which commits after the dashboard read has `created_at <= fetchedAt`, is invisible to the list, and is still marked read. Accepted and documented: the window is the duration of an in-flight insert transaction (milliseconds), the existing bulk-read contract is clock-based, and changing it to id-based is the E8-adjacent API work already deferred. Record as a known limitation in the JSDoc on the mark-all-read handler.\n- **Provisional approach vs open taste items (r2 #4, r3 #6).** The reviewer wants T1/T3 closed or A/E7 marked provisional. They are provisional by design: /autoplan surfaces taste decisions at one Final Gate and the fallback paragraphs define what changes. Not a contradiction to resolve in this document.\n- **Population size (r3 #14).** The single-workspace population may be too small for the default design; the fallbacks exist, and the member count is now recorded before R1 starts.\n\n## Taste decisions carried to the Final Gate"
|
||
},
|
||
"publicTools": [
|
||
{
|
||
"sessionId": "037c9f5e-7035-45e2-935d-0b950c435930",
|
||
"timestamp": "2026-09-10T14:32:00.204Z",
|
||
"toolUseId": "toolu_01A78r3KVhp1i9RAkz57ECSC",
|
||
"kind": "use",
|
||
"name": "Write",
|
||
"input": {
|
||
"file_path": "/tmp/gstack-paid-shard-MPtf0T/tmp/gstack-hermetic-3587976-PXLrEr/skill-home-cMZrbW/.gstack/projects/gstack-autoplan-chain-LukJXD/ceo-plans/2026-09-10-user-dashboard.md",
|
||
"content": "---\nstatus: ACTIVE\n---\n# CEO Plan: User Dashboard Page\nGenerated by /plan-ceo-review (via /autoplan) on 2026-09-10\nBranch: main | Mode: SELECTIVE EXPANSION\nRepo: gstack-autoplan-chain-LukJXD (local, no remote)\n\n## Problem and goal\n\nMembers of the existing single-role workspace visit three separate pages after login to resume work, check alerts, and inspect recent changes. One team walkthrough measured a median of 75 seconds to find the next item. The dashboard at `/dashboard` is meant to cut the median login-to-first-completed-task time from ~75 s to 45 s (a 40 % reduction), with completed-task rate and permission-error rate as guardrails.\n\n**Primary ship metric:** median login-to-first-completed-task time, treatment vs control. **Ship threshold:** ≥ 40 % reduction (equivalent to 75 → 45 s) sustained over the ≥ 2-week measurement window with no guardrail regression. **Diagnostic split (not a ship gate):** login→first-action-start (the leg the dashboard moves) and action-start→completion (unchanged by this plan), both derivable from the existing login / action start / action completion events, so the read-out attributes the delta correctly. The production baseline for all three is computed from existing events before the cohort opens; the 75 s walkthrough figure is not load-bearing.\n\n**Experiment design.** Assignment is a deterministic hash of member id, fixed for the experiment's duration. Default split 10 % treatment / 90 % control. Minimum sample: ≥ 300 logins per arm within the window; if the workspace's active-member count makes that unreachable at 10 % (record the count at approval), switch to a 50/50 split, and if still unreachable, fall back to a before/after comparison over the same window and say so in the read-out. `/dashboard` is flag-gated (below), so control members cannot reach it and contaminate the read-out.\n\n## Vision\n\n### 10x Check\nThe dashboard becomes the member's assistant rather than a summary page: a ranked \"next item\" (\"You have 2 items due today\"), alerts grouped by source, small tasks completed inline, live updates as teammates act, and a home that other features publish panels into. Concrete shape: ranking service over the action registry, a panel registry, a push channel. Human ~4-6 weeks / CC ~2-3 days. Out of this plan's blast radius; the personalisation plan owns it. This plan lays the seam (per-panel envelopes + a shared panel shell) without building the platform.\n\n### Platonic Ideal\nSkipped (SELECTIVE EXPANSION).\n\n## Baseline held (plan scope)\n- `/dashboard` page with Quick actions, Notifications, Recent activity panels; Tailwind, mobile-first (sm/md/lg).\n- Loading skeleton, empty, error, success per panel; hover and focus-visible on every control.\n- \"Mark all as read\" confirmation modal on the existing dialog primitive.\n- Toast feedback for actions.\n- `GET /api/dashboard` aggregate over existing PostgreSQL tables; no schema change.\n- Out of scope by the user: dark mode, personalisation.\n\n## Implementation approach chosen\nApproach A: aggregate endpoint composed server-side from the existing list methods and action registry, returning per-panel result envelopes so one failing panel never sinks the page. Alternatives: B client-side composition of three existing endpoints (7/10, close second, surfaced as taste decision T1); C smart post-login redirect + shell badge (4/10 against the stated goal, rejected).\n\n**If the Final Gate selects B instead of A:** the \"Endpoint\", \"Response shape\", and \"Retry\" paragraphs below are replaced by three per-panel fetch hooks wrapping the existing list endpoints and a small eligibility endpoint over the registry; each hook produces the same `Envelope<T>` client-side, so panel states, empty CTAs, mark-all-read, latest-wins, redirect precedence, caching headers, toast, accessibility, analytics, and rollout are unchanged. Composer tests become hook tests. **If the gate selects dashboard-local for T3:** the toast code moves under `src/components/dashboard/` with the same API and the same live-region rule; nothing else changes.\n\n## Verified preconditions (confirm before implementation; each has a fallback)\n- **Bulk-read API accepts a snapshot time.** The plan's contracts section states it \"marks only notifications at or before the supplied snapshot time\". Confirm the parameter name and type. If it turns out not to exist, mark-all-read cannot be made safe against later arrivals and the modal flow is **blocked** until the API gains the parameter (this becomes the same API work E8 needs; raise at the gate, do not ship a clock-unsafe version).\n- **Bulk-read API returns the affected count.** Confirm; otherwise the success toast reads \"Marked all as read\" without a number.\n- **Feature-flag system supports stable per-member assignment** (hash of member id). Confirm; otherwise implement the hash in the redirect module and store only the flag's on/off and percentage in the flag system.\n- **Analytics pipeline already emits login, action start, action completion, permission error** with member id and timestamps queryable for the baseline. Confirm; otherwise the baseline query is the first task and the cohort does not open until it runs.\n- **Metrics and alerting stack exists** (the plan names existing request/error metrics). Confirm the alert-rule mechanism; otherwise alerts are the first observability task.\n- **Activity list method preloads actor display names.** Confirm; otherwise the composer batches one lookup (never per row).\n- **Notification list can return a DB-clock timestamp** (`SELECT now()` in the same statement or transaction). Confirm; otherwise add a read-only wrapper that does.\n\n## Contract details (assume Approach A; see the B fallback above)\n\n**Route gating.** `/dashboard` is served only to members whose `dashboard_landing` assignment is on; others are redirected to the current landing page (302). No nav link is added in v1; treatment members reach the page through the post-login redirect and by URL. `dashboard_viewed{cohort}` records the member's assignment.\n\n**Endpoint.** `GET /api/dashboard`. Authentication and workspace membership are checked by the existing middleware *before* the handler runs: no session → 401, non-member → 403, both short-circuit the whole request (the client redirects to login or shows the access message). Flag-off members receive 404 from this route. Per-panel envelopes cover upstream failures only.\n\n**Response shape.**\n```\nDashboardResponse = {\n activity: Envelope<{ rows: ActivityRow[] }>,\n notifications: Envelope<{ rows: NotificationRow[], unreadTotal: number }>,\n quickActions: Envelope<{ rows: QuickAction[] }>\n}\nEnvelope<T> = { ok: true, fetchedAt: string, data: T }\n | { ok: false, fetchedAt: string, error: { code: 'timeout' | 'upstream_error', retryable: true, requestId: string } }\n\nDashboardUnavailable = { error: { code: 'dashboard_unavailable', retryable: true, requestId: string } } // 503 body only, not an envelope\n```\n`fetchedAt` is per envelope. For notifications it is the **database clock** (`now()` read in the same statement or transaction as the list and count), so a row whose `created_at` is later than `fetchedAt` by the same clock is never covered by a snapshot taken from it. For activity and quick actions it is the app-server time before the call (informational only). The composer runs the three calls with `Promise.allSettled` and a 2 s budget per call enforced **both** at the composer (race) and at the query (`statement_timeout` of 2 s or the driver's `AbortSignal`), so a timed-out query releases its pool connection instead of leaking it. The notifications count runs inside the notifications call and shares its timeout; if the count fails the notifications panel is `ok: false`. If all three fail the response is 503 with `DashboardUnavailable`; otherwise 200. A registry predicate that throws is caught: that action is omitted, one error log line with `actionId` is written, `dashboard_action_predicate_error_total{actionId}` is incremented, and the panel still succeeds. `unreadTotal` comes from a member-scoped read-state count over the existing indexed path; if no repository count method exists, a read-only count method is added (a new read query, not a schema or mutation change).\n\n**Limits and ordering.** Activity and notifications return at most 5 rows. Notifications: unread first, then `created_at` desc. Activity: `created_at` desc. Quick actions: all eligible actions in registry order (the registry holds 3; hard cap 5 for safety); Quick actions has **no** \"See all\". Activity and Notifications each have a \"See all\" link to the existing full page, which owns older-page navigation.\n\n**Retry.** A panel's Retry button refetches the whole aggregate (one request). Only panels that were in ERROR show a skeleton; panels that already have data enter `refreshing`.\n\n**Panel states.** `LOADING` (skeleton, `aria-busy`), `EMPTY` (copy + CTA), `ERROR` (message + Retry), `SUCCESS`, and `refreshing` = previous data stays rendered with a subtle indicator; used for post-mutation refetch and for Retry when the panel already has data. State is derived only from the last settled envelope, so SUCCESS→EMPTY or ERROR→SUCCESS cannot happen without a fetch.\n\n**Empty-state CTAs.** Activity empty → \"Create your first item\" linking to the create action's route when that action is present in `quickActions.data.rows`; otherwise link to the existing full activity page. Notifications empty → \"All caught up\" with a link to the full notifications page. Quick actions empty → \"Nothing to do right now\" (no CTA).\n\n**Mark all as read.** The trigger is rendered only when the notifications panel is `SUCCESS` and `unreadTotal > 0`; in every other state (LOADING, EMPTY, ERROR, refreshing) it is absent. Confirmation dialog on the existing dialog primitive (focus trap, Escape, focus return). Confirm button disabled while the request is pending; the server API is idempotent regardless. The request calls the existing member-scoped bulk-read API with the notifications envelope's `fetchedAt` as the snapshot, never the client clock. No new mutation API. On success: toast \"Marked N notifications as read\" where N is the count returned by the bulk-read response (or \"Marked all as read\" if the API returns no count), then a refetch. On failure: 401 → login; 403 (including CSRF) → the existing client refreshes the token once and retries, then an error toast; validation / retryable / network → error toast with a Retry action.\n\n**Latest-wins fetching.** One `AbortController` per request. Starting the mutation aborts any in-flight GET. Each request carries a sequence number; only the newest may commit state. This excludes the schedule where a slow pre-mutation GET lands after the post-mutation GET and repaints stale unread rows.\n\n**Post-login redirect precedence.** (1) A `next` value that is a same-origin relative path → go to `next`. An absolute, protocol-relative, or foreign-origin `next` is treated as absent and logged at warn. (2) Otherwise, if `dashboard_landing` is on for this member → `/dashboard`. (3) Otherwise → the current landing page. If the flag lookup itself fails, treat the flag as off (step 3).\n\n**Caching and limits.** `Cache-Control: private, no-store` on the GET. The route joins the same rate-limit tier as the existing authenticated list endpoints. Error envelopes carry the request id so a report can be joined to server logs.\n\n**Toast primitive.** `ToastProvider` + `useToast` under `src/components/toast/`, mounted at the app root so the `aria-live=\"polite\"` region exists before the first toast. Queue holds at most 3; a 4th evicts the oldest. Informational toasts auto-dismiss after 5 s; toasts carrying an action (Retry) persist until dismissed or acted on. Respects `prefers-reduced-motion`. In development, `useToast()` outside the provider throws, and a toast with an empty message throws; in production both are no-ops with `console.error`.\n\n**Accessibility.** Named controls, focus-visible outlines, 44 px touch targets on action buttons, reduced motion on skeleton shimmer and toast motion, axe passes on every panel state and with the dialog open, keyboard-only mark-all-read passes in Playwright.\n\n## Scope Decisions\n\n| # | Proposal | Effort (human / CC) | Decision | Reasoning |\n|---|----------|---------------------|----------|-----------|\n| E1 | Per-panel result envelopes with Retry | S (~4h / ~20m) | ACCEPTED | closes the plan's own open partial-failure work |\n| E2 | Preserve intended destination; `/dashboard` only when none | S (~1h / ~10m) | ACCEPTED | protects deep links and bookmarks; one file |\n| E3 | Unread-first ordering + `unreadTotal` in header | S (~2h / ~10m) | ACCEPTED | one sort, one count, one span |\n| E4 | \"See all\" link (activity, notifications) to the existing full page | S (~1h / ~5m) | ACCEPTED | full pages already own pagination |\n| E5 | Empty-state copy pointing at the relevant quick action (with fallback) | S (~1h / ~5m) | ACCEPTED | empty states are features |\n| E6 | Relative timestamps with absolute time for assistive tech | S (~1h / ~5m) | ACCEPTED | small, on the row component |\n| E7 | Toast built as a shared app primitive, not dashboard-local | S (+1h / +10m) | ACCEPTED (taste T3) | second consumer is inevitable; a11y policy requires a live region |\n| R1 | Required-to-ship measurement block (flag gating, events, baseline, logs, metrics, alerts, staging load check, test matrix) | M (~3 days / ~1.5h) | ACCEPTED | without it the ship rule cannot be evaluated |\n| E8 | Undo after Mark all as read | M | DEFERRED | needs a new \"unmark since snapshot\" mutation; the existing bulk-read API only marks read |\n| E9 | Live / polling notification updates | L | DEFERRED | new infra, outside blast radius |\n| E10 | Keyboard shortcuts (`g d`, `j/k`) | S | DEFERRED | no existing shortcut map to extend |\n| E11 | Stale-while-revalidate client cache | S | DEFERRED | private data; `no-store` in v1 |\n| E12 | Ranking / next-item recommendation | XL | SKIPPED | user's stated non-goal (personalisation plan) |\n\n## Accepted Scope (added to this plan)\nEverything under \"Contract details\" and \"Verified preconditions\" above, plus:\n\n**Required to ship (R1; gates the cohort):**\n- `dashboard_landing` flag with deterministic per-member assignment, route gating, and flag-flip rollback; no migration; endpoint additive.\n- Analytics events: `dashboard_viewed{cohort}` once per page view; `dashboard_panel_state{panel,state}` once per page view per panel at its first terminal state (EMPTY / ERROR / SUCCESS) only; `dashboard_action_clicked{actionId}`; `dashboard_see_all_clicked{panel}`; `dashboard_mark_all_read{result}`.\n- Production baseline for login-to-first-completed-task, login→action-start, action-start→completion computed from existing events before the cohort opens; active-member count recorded and the split chosen per the experiment design.\n- The pre-registered ship rule above, with a named owner for the read-out recorded at approval.\n- Structured request log per GET (member id, workspace id, request id, per-panel status and ms, total ms) and per mark-all-read.\n- Metrics `dashboard_request_total{panel,status}`, `dashboard_panel_duration_ms{panel}`, `dashboard_mark_all_read_total{result}`, `dashboard_action_predicate_error_total{actionId}`, DB pool wait on the route; production alerts: 5xx or `dashboard_unavailable` > 1 % for 5 min, p95 > 800 ms for 10 min, mark-all-read error rate > 5 %.\n- Tests: composer settle matrix + all-fail 503 + predicate-throw + query-timeout-releases-connection; envelope→state mapper; panel, dialog, toast, latest-wins (controlled promise release), redirect precedence, route gating unit/RTL suites; integration for member 200 / other-workspace 403 / no-session 401 / flag-off 404 / service failure / `no-store` header; Playwright login→dashboard→action, 503→Retry, keyboard-only mark-all-read, axe on every state.\n- Rollout order: dark deploy → internal members on staging (Playwright + axe suite + a staging load check p95 < 300 ms at 50 rps on staging fixtures) → treatment cohort for the ≥ 2-week measurement window → decision per the ship rule.\n\n**Operability polish (same PR, does not gate the cohort):**\n- `Server-Timing` per panel on the GET.\n- JSDoc on `composeDashboard`, `DashboardPanel`, and `ToastProvider` carrying the architecture / panel-state / envelope-rationale diagrams. No separate README.\n\nThe p95 latency target is enforced on staging before cohort expansion and in production alerting, not as a per-PR CI assertion (CI runners are too noisy for a latency SLO).\n\n## Deferred to TODOS.md\n- Undo after Mark all as read — blocked on a new \"unmark since snapshot\" mutation API (P2).\n- Live notification updates — needs a push/polling infra decision (P3).\n- Dashboard keyboard shortcuts — needs an app-wide shortcut map (P3).\n- SWR client cache for the dashboard payload — needs a privacy review of client caching (P3).\n\n## Taste decisions carried to the Final Gate\n- T1 aggregate endpoint (A) vs client-side composition (B) — open; the B fallback paragraph above states exactly what changes.\n- T2 confirmation modal for Mark all as read — recorded decision, surfaced for information: the independent reviewer preferred direct action + undo, but undo (E8) needs a new mutation API and is deferred, so the modal stays unless E8 is pulled into scope.\n- T3 toast as a shared primitive vs dashboard-local — open; the T3 fallback sentence above states what changes.\n- Premise note: the 75 s baseline is one team walkthrough; the ship rule is relative (≥ 40 %) and the analytics/rollout task measures the true baseline, so the launch does not depend on the figure.\n"
|
||
},
|
||
"messageId": "msg_011CeuvH91RGhGXA5XHMPN7J",
|
||
"requestId": "req_011CeuvH6MgNDWkEkYtuWxs9"
|
||
},
|
||
{
|
||
"sessionId": "037c9f5e-7035-45e2-935d-0b950c435930",
|
||
"timestamp": "2026-09-10T14:32:00.969Z",
|
||
"toolUseId": "toolu_01A78r3KVhp1i9RAkz57ECSC",
|
||
"kind": "result",
|
||
"content": "The file /tmp/gstack-paid-shard-MPtf0T/tmp/gstack-hermetic-3587976-PXLrEr/skill-home-cMZrbW/.gstack/projects/gstack-autoplan-chain-LukJXD/ceo-plans/2026-09-10-user-dashboard.md has been updated successfully. (file state is current in your context — no need to Read it back)",
|
||
"isError": false
|
||
}
|
||
],
|
||
"provenance": {
|
||
"status": "Exact original request reconstruction independently equals original requestSHA; no new digest existed in the paid attempt.",
|
||
"requestSHA256": "7ea64ad22d29d9c779b6e945119ea238e2f40862a368ca518100626a1e7978de",
|
||
"beforeSHA256": "6771404ca60011824af336f0a7011e3aba465b8728c011eef935a4b847e72bdd",
|
||
"paneSHA256": "c5662a9c0d67ba04de8720c25d51378b4662266273cb186f48f7c3109abd2ee2",
|
||
"sourceProof": "ap-cropped-addition-plan-v1/proof.json",
|
||
"reconstruction": "ap-cropped-addition-request-reconstruction-plan-v1.json",
|
||
"publicTools": "Exact acknowledged same-file predecessor pair from full94 public events.",
|
||
"actualCoverage": "operator-cancelled-incomplete; no phase credit"
|
||
}
|
||
}
|