mirror of
https://github.com/garrytan/gstack.git
synced 2026-09-16 09:55:29 +02:00
* feat: add a restricted and supervised Claude Code runner Preserve configured authentication and models while enforcing tool access, strict completion JSON, bounded output and process cleanup. Cover argv, failure handling, session metadata and Windows process containment. * feat: route outside reviews by harness and migrate wrapper installs Use Claude Code from Codex and Codex from other supported hosts, with shared invocation rendering, positive gate validation and per-phase provenance. Rename /claude to /claude-code, repair managed shared and copied installations safely, and generate native Kiro skills. Add installed-workflow, failure-injection and live cross-harness regression coverage. * test: recognize CEO mode labels without terminal spacing The paid workflow rendered SCOPEEXPANSION at option 4, but its driver required a literal space. Match the leading mode title without cursor-spacing artifacts and ignore adjacent preview text. Preserve missing-target failures and downstream posture assertions. * test: isolate plan-count fixtures before starting review workflows Seed the complete test plan in a private git repository before launching Claude, so a bare slash command cannot review the live workspace while a delayed fixture message remains queued. Preserve count thresholds, parsers and budgets. Add initial-context and installed-discovery tests, and retain startup/terminal diagnostics on failed evaluations. * test: stabilize review fixtures and Claude eval startup Preserve source boundaries in workflow judge inputs, isolate CEO mode plans, and wait for interactive trust input readiness. Keep startup failure evidence and retain existing models, budgets, and assertions. Co-Authored-By: OpenAI Codex <noreply@openai.com> * test: classify collapsed review modes and isolate seeded findings Keep review questions out of the setup count when terminal cursor positioning removes spaces. State existing webhook safeguards so the five-finding control measures its seeded defects without accidental extra security and concurrency gaps. Preserve question bands and the paired control. Co-Authored-By: OpenAI Codex <noreply@openai.com> * test: isolate browser daemon state across free shards Co-Authored-By: OpenAI Codex <noreply@openai.com> * test: stabilize native review counting and interactive navigation Co-Authored-By: OpenAI Codex <noreply@openai.com> * chore: prepare v1.82.0.0 release Co-Authored-By: OpenAI Codex <noreply@openai.com> * fix: eliminate browser and process-cleanup test flakes Pin every CI surface to Bun 1.4.0 to avoid extra-stdio finalizers closing reused live sockets. Add an isolated GC/listener regression that fails on Bun 1.3.13, and prevent coordinated rollback to an affected CI runtime. Check renderer cleanup against the render's own staging directory so concurrent renders cannot invalidate the assertion. Make the no-pgrep process-tree walk tolerate disappearing /proc entries, and synchronize its test fixture through child readiness and pipe EOF instead of sleeps. Validation: 9,157 passed, 31 skipped, zero failures across 556 files with retries disabled. Build, all-host generation freshness, and skill checks passed. All three races have failing-before/passing-after regressions. * fix: count completed native review questions in evals * fix: drive review navigation from confirmed native choices * fix: require complete section-loading eval reports * test: isolate telemetry HTTP transport from local assertions * fix: keep review input on the active native question * test: let tunnel revocation daemon choose an available port * test: allocate available ports for pairing and watchdog fixtures * fix: stabilize planning eval navigation and phase reporting * test: isolate installed runtime paths in planning evals * test: stabilize review evidence and concurrent refresh fixtures * fix: resolve design findings before editing the plan * fix: honor and persist disabled outside plan reviews * fix: preserve planning decisions and terminal evidence Load installed host reviews at autoplan phase entry and wait for completed reviewers and saved artifacts. Reuse approved remedies while preserving individual finding decisions. Drive interactive evals from the current terminal viewport, bind native questions across scrolling, and require complete native report evidence. Cover captured stale menus, permission lifecycles, setup classification, and disabled-review tool availability with deterministic regressions. Advance release metadata and the upgrade migration to the unclaimed 1.83.0.0 slot. * fix: drive native review questions and preserve current plans Use the native single-choice keyboard protocol and current terminal viewport, with per-question navigation inside packets and completed-call coverage. Keep permissions, multi-select menus, and Submit controls distinct. Send Autoplan reviewers the amended implementation plan, keep its review record separate, and supply retained application contracts in the chain fixture. Clarify individual DevEx decisions and complete CEO fix options; use one active plan destination for the section-loading report. * fix: preserve complete plan-review decisions * fix: recognize native plan dialogs and reviewer controls * fix: preserve review decisions and phase completion * fix: recognize completed reviews without losing findings * fix: preserve review continuity and native eval completion * test: fix native review completion and eval retry isolation * test: handle native review menus and complete eval fixtures * test: fix native review setup, completion, and isolation failures * test: limit native skill discovery to runtime assets * fix: bind Autoplan reviews to full ordered phase inputs * test: fix planning eval routing, counting, and timeout handling * chore: advance queued release to v1.84.0.0 * fix: preserve complete review inputs and planning decisions * fix: reconcile review approvals and preserve phase obligations * fix: preserve review obligations and unblock eval permissions Carry recorded Autoplan requirements into blind phase inputs, require Eng review approvals before exit, and exercise combined asynchronous flows in CEO reviews. Correct native finding and handoff classification and unblock repeated report edits using scoped request identities. * fix: retain plan requirements and complete native review dialogs * fix: complete native review prompts and retain plan references * fix: preserve review inputs and classify native eval evidence * fix: check competing completion orders in CEO reviews * fix: recognize review decisions and require phase methodology Require the current phase methodology before Autoplan snapshots. Correct substantive decision, closed handoff, and cache-finding classification, and honor the recommended implementation approach in native review dialogs. Add captured-transcript regressions without changing review thresholds, provider models, retries, or deadlines. * test: bind native review decisions and close completed handoffs * fix: complete review dialogs and verify methodology delivery * fix: preserve review evidence and unblock native eval prompts * fix: handle native review question completions * fix: recognize native review narration and controls * fix: count native review decisions and isolate eval fixtures * test: verify seeded review coverage and current artifact permissions * test: isolate model and brain-aware skill renders * fix: repair native workflow evaluation and clarify review steps * fix: stabilize workflow eval evidence and review guidance * test: repair native workflow observation and fixture isolation * fix: recognize completed workflow evidence and owned skill reads * test: repair seeded workflow delivery and completion evidence * test: recognize current review evidence across native forms * test: handle native review variants and permission redraws * fix: honor review preferences and recognize native eval evidence * test: recognize completed review decisions and queued permissions * test: match current review contracts and partial-line edits * test: recognize completed workflow evidence and bounded human waits * fix: preserve review entry gates and native eval interactions * fix: recognize native workflow evidence and preserve review gates * test: recognize current review evidence and preconfigure workflow fixtures * test: recognize completed review findings and scoped artifact permissions * fix: stabilize native workflow review and permission evidence * fix: recognize current review evidence and scoped edit confirmations Clarify Design and engineering review entry instructions and Design scoring. Recognize required legacy coverage and public Autoplan completion recaps. Bind the pending Edit confirmation to its exact file, ordered digest, and one-request approval when a preceding command display remains visible. Keep reviews within their existing size limits and preserve scope gates when extracting workflow fixtures from either supported preamble header. Keep failure outcomes, review thresholds, provider choices, and eval budgets. * fix: recover review workflow progress and eval evidence * fix: recognize valid review evidence and scope selection * test: fix review evidence parsing and repeated artifact prompts * test: recognize valid review decisions and pending native cards * fix(plan-eng-review): keep final navigation consistent with approved tasks * test: recognize valid review evidence and bind legacy diff requests * fix: stabilize review eval evidence and harness repair guidance * docs: update project documentation for v1.85.0.0 Co-Authored-By: OpenAI Codex <noreply@openai.com> * test: fix Windows CI fixtures and credential scan Rebase captured JSON values and filesystem evidence using the appropriate path convention. Compile native fake CLIs on Windows and synchronize pipe holder readiness, with cleanup retained when assertions fail. Assemble synthetic credential fixtures at runtime so the added-line scan keeps enforcing the same gate without flagging its own rejection controls. Discover generated skills directly for the empty-find regression check, avoiding a recursive scan through saved evaluation artifacts and dependencies. * fix: preserve source renders on Windows Compare canonical generator paths using native separators so an output sidecar pointing at the source cannot overwrite its skill or metadata. Keep the regression fixture isolated from the real checkout and expose freshness diagnostics before asserting subprocess status. Detach Windows drain-test pipe holders from the fake provider's automatic child cleanup while preserving the enclosing runner job and its assertions. * fix: clarify outside review fallback and CEO decisions Render one applicable own-harness fallback path and retain native review, disabled policy, and missing-coverage semantics. Align report field names and mode labels, and make the existing per-cut scope approval explicit. Regenerate skill outputs and keep the workflow judge's model, thresholds, and retry policy unchanged. * chore: move release to free version slot (v1.86.0.0) PR #2852 now claims v1.85.0.0. Align the release metadata and rename migration so upgrades from that version still receive it. Co-Authored-By: OpenAI Codex <noreply@openai.com> * fix: include engineering review prerequisites and restore branch context * fix: recognize coverage diagrams and clarify design review instructions * fix: preserve file identities and join Windows test processes --------- Co-authored-by: OpenAI Codex <noreply@openai.com>
22 lines
78 KiB
JSON
22 lines
78 KiB
JSON
{
|
|
"description": "Exact public completed PLAN.md artifacts for the AG SDK semantic assertion failures; not a live pass.",
|
|
"first": {
|
|
"sourceProof": ".context/ship-source-ag-delta-paid-20260910-v1/sdk-first-report-native-v1/proof.json",
|
|
"sourceProofSha256": "e2e5849b7ee71acce22dc2e3a56df20e9a5c94caf1d9fe203cb564c6fcb542a0",
|
|
"reportSha256": "60d0f6a690c6fb2e25c145ded7d19e7c370f26cfc00e557c1e6cbc5614b9c756",
|
|
"report": "# Plan: cache profile summaries in one process\n\nReviewed by /plan-ceo-review on 2026-09-10 (mode: HOLD SCOPE). Original plan content is\npreserved below; accepted amendments are marked `[Amended: D<n>]` and cross-reference the\ndecision and finding registries in the review record.\n\n## Measured problem and accepted scope\nThe existing profile-summary service has one active process. A one-week trace\nshows repeated reads of about 900 hot keys: DB CPU is 70%, with read p95 120 ms.\nAdd a process-local LRU wrapper to the existing repository. Acceptance targets\nare at least 60% cache hits, DB CPU below 50%, and read p95 below 60 ms, with the\nexisting error-rate and correctness SLOs unchanged. This is an internal backend\nchange with no UI, API, schema, pricing, or developer onboarding change.\n\n## Existing contracts retained\n- All reads and writes use this repository in the same process; there are no\n external DB writers. Multi-process operation remains unsupported and startup\n rejects that configuration while caching is enabled.\n- Authentication and authorization run before repository access. Keys encode\n the authenticated tenant ID and validated profile ID without ambiguity.\n Values are immutable profile-summary DTOs; secrets and cache keys are never\n logged. Cached results cannot bypass authorization.\n- The existing LRU adapter supports 1000 entries, a 16 MiB byte cap, and a\n 30-second TTL. Recorded hot data fits those limits. Absent records use a\n distinct sentinel with a 10-second TTL; undefined means a cache miss.\n- Cache operations are synchronous and atomic in the single JS event loop.\n On any cache failure the existing adapter bypasses the cache until an empty\n cache is reinitialized; repository errors keep the current typed API error\n mapping. The existing per-key\n single-flight wrapper coalesces simultaneous misses and releases on failure.\n- A read already in progress when a write commits may return its earlier DB\n snapshot to that caller. Every read begun after that write completes must\n observe the committed version. TTL expiry is not a substitute for this rule.\n\n## Proposed wrapper integration\nKeep the current read-through repository interface and shared adapters.\n\n`[Amended: D3, D4, D5, D7]` The original sketch stated that no coordination between a\ncache fill and a write was needed. Review finding F1 shows that claim violates the\nretained read-after-write contract (see Async schedule in the review record). The\ncomplete read/write ordering rules are now:\n\n1. A miss registers an in-flight fill for the key in the existing single-flight\n wrapper. The fill handle carries an `invalidated` flag.\n2. A write marks and detaches the key's in-flight fill before awaiting the DB, and\n again after the write settles (success, typed failure, or indeterminate failure),\n then deletes the key. A detached fill still resolves for its own waiters (allowed by\n the contract) but never populates the cache; it increments `profile_cache.fill_discarded`.\n3. Absent results are stored as the sentinel with the 10-second TTL and mapped back to\n the repository's existing absent result on a hit, so callers see identical behavior\n with and without the cache. Values use the 30-second TTL.\n4. The feature flag is evaluated once per operation with the same deterministic key\n hash for reads and writes. A flag-read failure evaluates as disabled. Any change to\n the flag value (off/on or percentage) reinitializes an empty cache and invalidates\n all in-flight fills.\n\n```javascript\n// Fill/write ordering (F1, F3): a fill that overlaps any write never populates.\n// miss \u2500\u25b6 flight{invalidated:false} \u2500\u25b6 await repository.read \u2500\u25b6 invalidated ? discard : set\n// write \u2500\u25b6 invalidate(key) \u2500\u25b6 await repository.write \u2500\u25b6 finally: invalidate(key); cache.delete(key)\nasync function readProfile(key) {\n if (!flag.enabledFor(key)) return repository.read(key); // flag error => disabled (F5)\n const cached = cache.get(key);\n if (cached !== undefined) return fromCacheValue(cached); // sentinel -> existing absent result (F2)\n return singleFlight.run(key, async (flight) => {\n const value = await repository.read(key); // typed errors propagate; nothing cached\n if (flight.invalidated) metrics.increment('profile_cache.fill_discarded');\n else cache.set(key, toCacheValue(value), ttlFor(value)); // 30 s value / 10 s sentinel (F2)\n return value;\n });\n}\n\nasync function writeProfile(key, update) {\n if (!flag.enabledFor(key)) return repository.write(key, update);\n singleFlight.invalidate(key); // fills started before the write cannot populate (F1)\n try {\n return await repository.write(key, update);\n } finally {\n singleFlight.invalidate(key); // fills issued during the write are pre-commit snapshots (F1)\n cache.delete(key); // every outcome, including indeterminate failures (F3)\n }\n}\n```\n\n## Verification and rollout\nExisting repository contract tests cover tenant isolation, key validation,\nabsence, DB failures, and authorization. New wrapper tests cover hit/miss,\neviction and byte limits, TTL, adapter-failure fallback, successful-write\ninvalidation, and concurrent-miss coalescing.\n\n`[Amended: D3, D4, D5, D7]` Additional required tests (see Section 6 and tasks T1 to T4):\ncontrolled-schedule read-after-write tests for both completion orders and for a fill\nissued during a write; sentinel round-trip and negative-cache tests; write-failure\ninvalidation (replaces the original \"failed-write preservation\" test, which asserted the\nprohibited stale result after an ambiguous-outcome commit); flag-change flush and\nin-flight discard; same-hash flag evaluation for reads and writes; TTL tests use fake timers.\n\nThe rollout uses the existing runtime feature flag: enable for 10% of keys,\nthen 50%, then all keys after one healthy hour at each stage. Monitor hit/miss,\neviction, cache bytes, fallback errors, DB CPU, and read p95 without raw IDs.\nOn error-rate or latency regression, disable the flag immediately; both reads\nand writes bypass the cache while disabled, and enabling creates an empty cache.\nCold starts remain within the existing DB capacity. The service owner monitors\nthe rollout and records the results against the acceptance targets.\n\n`[Amended: D6]` \"Healthy hour\" is defined and alerted (task T5): error rate within the\nexisting SLO, read p95 not above baseline, zero adapter fallback errors, cache in ACTIVE\nstate, `fill_discarded` below 1% of fills, and at the 100% stage hit ratio at least 60%\nafter a 5-minute warmup. A day-1 dashboard, alerts, and a runbook are deliverables of\nthis change, not follow-ups.\n\n## Out of scope\nDistributed caching, cross-process coherence, prewarming, changing consistency\nsemantics, or adding new product surfaces. The repository interface preserves a\nfuture replacement path without introducing a general cache framework now.\n\n---\n\n# CEO Review Record (HOLD SCOPE)\n\n## Run conditions and skipped steps\n- Automated run, no human present. Every decision point auto-selected the recommended option; each is recorded in the Decision Registry.\n- Skipped per run rules: preamble/system audit, environment setup, telemetry, codebase exploration, git commands, review-log and decision-log persistence, tasks JSONL artifact, handoff cleanup, brain context. File paths in tasks are therefore descriptive, not verified.\n- Outside voice: `.gstack-section-state-xMGd4Z/config.yaml` sets `codex_reviews: disabled`. The entire extra outside-review step was skipped, including the native subagent fallback. Outside coverage: disabled. The persistence command for the disabled record was not run (no shell in this run).\n- No design doc, handoff note, TODOS.md, or CLAUDE.md is present in the repo. Prerequisite /office-hours offer: auto-declined (standard review).\n- Frontend/UI scope detection: none. Section 11 skipped with justification.\n\n## Step 0: Scope challenge and mode\n\n**0A Premise.** Real, measured pain: 70% DB CPU and 120 ms p95 from about 900 hot keys in a one-process service. A process-local read-through cache is the most direct path. Doing nothing leaves the DB as the scaling ceiling. No proxy problem detected: the targets (hit ratio, CPU, p95) are the user-facing outcome.\n\n**0B Existing leverage.** The plan reuses the LRU adapter (limits, TTL, sentinel, failure bypass), the per-key single-flight wrapper, the repository's typed error mapping, contract tests, and the runtime feature flag. Nothing is rebuilt. The one amendment to shared code is adding `invalidate(key)` to the single-flight wrapper (F1).\n\n**0C Dream state.**\n```\n CURRENT STATE THIS PLAN 12-MONTH IDEAL\n every read hits DB; ---> process-local LRU behind ---> same repository interface;\n 70% DB CPU, p95 120ms the repository, flag-gated, cache tier swappable (local\n invariant-safe fills or shared) if multi-process\n ever arrives; DB CPU headroom\n```\nMoves toward the ideal: the interface is unchanged and the wrapper is replaceable.\n\n**0C-bis Implementation alternatives** (decision D1).\n```\nAPPROACH A: Process-local LRU wrapper around the repository (plan as written, corrected)\n Effort: S Risk: Low Completeness: 10/10 once F1 to F5 remedies land\n Pros: reuses every existing adapter; flag rollback in seconds; smallest diff meeting targets\n Cons: single-process only; needs fill/write ordering guard the sketch omitted\n Reuses: LRU adapter, single-flight, flag, typed errors, contract tests\nAPPROACH B: Minimal viable: sketch as written, no fill coordination\n Effort: XS Risk: High Completeness: 4/10\n Pros: fewest lines\n Cons: violates the retained read-after-write contract (F1); silent negative-cache defect (F2)\n Reuses: same as A\nAPPROACH C: Ideal architecture: shared cache tier (Redis) with pub/sub invalidation\n Effort: L Risk: Med Completeness: 10/10 for multi-process, but out of the accepted scope\n Pros: survives a future multi-process move; prewarm-friendly\n Cons: new dependency, network hop, cross-process coherence is explicitly out of scope\n Reuses: repository interface only\nRECOMMENDATION: A. It is the complete version of the accepted scope; B fails a stated invariant and C is declared out of scope.\n```\n\n**0F Mode.** HOLD SCOPE, set by the user (decision D2). Backend performance change on an existing system; no expansions surfaced.\n\n**0D HOLD SCOPE analysis.** Complexity: one wrapper module, one small addition to the single-flight wrapper, tests, and dashboards. Under the 8-file smell threshold. Minimum change meeting the stated invariants is Approach A with F1 to F5; nothing further can be deferred without breaking a stated contract or Prime Directive 5.\n\n**0E Temporal interrogation.** Decisions resolved now rather than during implementation: fill/write ordering mechanism (D3), sentinel representation and TTL per type (D4), write-failure cache behavior (D5), health definition and alerts (D6), flag semantics at transitions (D7). Human effort about 1 day total; CC about 1 hour.\n\n## Decision Registry\n| ID | Decision | Choice | Basis |\n|----|----------|--------|-------|\n| D1 | Implementation approach (0C-bis) | A: process-local LRU wrapper, corrected | Recommended; auto-selected |\n| D2 | Review mode (0F) | HOLD SCOPE | Set by user |\n| D3 | F1 remedy | A: per-key in-flight fill invalidation in single-flight (10/10) over B: global write epoch (8/10, hit-rate loss under write bursts) or C: keep sketch (rejected, violates contract) | Recommended; auto-selected |\n| D4 | F2 remedy | A: explicit `toCacheValue`/`fromCacheValue` mapping with per-type TTL (10/10) over B: store absent as `null` only (7/10, breaks if absent is a typed error) | Recommended; auto-selected |\n| D5 | F3 remedy | A: invalidate on every write outcome via `finally` (10/10) over B: classify definite vs indeterminate failures (8/10, needs driver error taxonomy) or C: keep preservation (rejected, serves stale data) | Recommended; auto-selected |\n| D6 | F4 remedy | A: define health, alerts, dashboard, runbook as deliverables (10/10) over B: monitor manually (5/10) | Recommended; auto-selected |\n| D7 | F5 remedy | A: same-hash evaluation, flush plus in-flight invalidation on any flag change, fail-closed flag read (10/10) over B: flush only on off-to-on (6/10) | Recommended; auto-selected |\n\nLake Score: 5/5 remedy recommendations chose the complete option. No shortcuts accepted; no `gstack-shortcut` markers required.\n\n## Findings Registry\n| ID | Section | Severity | Evidence (from PLAN.md as submitted) | Selected remedy | Residual risk | Verification |\n|----|---------|----------|--------------------------------------|-----------------|---------------|--------------|\n| F1 | 1, 4 | CRITICAL GAP (fixed by D3) | Contract: \"Every read begun after that write completes must observe the committed version.\" Sketch: `cache.set(key, value)` after `await repository.read` with \"no additional version checks or coordination\". Schedule below shows a post-write reader receiving the pre-write snapshot via a stale fill or via single-flight coalescing. | D3: single-flight `invalidate(key)` before and after the write; invalidated fills never `set`; `fill_discarded` metric | Fills overlapping a write are not cached (bounded hit-rate cost, visible in metric) | T1 tests: both completion orders plus during-write fill; assert post-write reader sees v2 and cache never holds v1 |\n| F2 | 2, 5 | WARNING (fixed by D4) | Contract: sentinel for absent, \"undefined means a cache miss\". Sketch returns `cached` raw (leaks sentinel to callers) and calls `cache.set(key, value)` with no mapping or TTL: if absent is `undefined`, the set is a silent no-op and negative caching never works; if absent is a typed error, the throw skips `set`. | D4: explicit mapping both ways, TTL per type | None beyond adapter behavior | T2 tests: second absent read within 10 s does not hit DB; caller result identical to uncached path; sentinel expires at 10 s, value at 30 s |\n| F3 | 2 | WARNING (fixed by D5) | Sketch runs `cache.delete` only after a successful write; tests specify \"failed-write preservation\". A commit-ack timeout or connection drop after commit leaves v2 in the DB and v1 cached for up to 30 s, which the uncached system never did and the contract forbids. | D5: `finally { invalidate; delete }` on every write outcome | One extra DB read after a failed write | T3 test: write rejects after simulated commit; next read hits DB and returns v2 |\n| F4 | 8, 9 | GAP (fixed by D6) | Plan says \"Monitor ...\" and \"one healthy hour\" with no thresholds, alerts, dashboard, or runbook. Adapter bypass after a cache failure is silent to operators without an alert. | D6: health definition, alerts, day-1 dashboard, runbook | Threshold tuning after first stage | T5: alerts fire in staging fault injection (forced adapter failure, forced hit-ratio drop) |\n| F5 | 9 | WARNING (fixed by D7) | \"enable for 10% of keys\" and \"enabling creates an empty cache\" leave open: whether reads and writes evaluate the same bucket, whether 10 to 50 or 50 to 10 changes flush, what a flag-read failure does, and whether in-flight fills spanning a flush can repopulate. A fill started before a flush that lands after it, with a write in the disabled window, would violate the contract. | D7: same deterministic hash; any value change flushes and invalidates in-flight fills; flag-read error means disabled | None | T4 tests: read and write buckets agree for sampled keys; flag change flushes and discards in-flight fill; flag error bypasses cache |\n\n## Section outcomes (1 to 11)\n1. **Architecture.** One finding, F1. Dependency graph, data flow paths, and rollback posture in Diagrams. Coupling added: wrapper depends on single-flight handle and flag; justified and small. Scaling: at 10x load the entry cap (1000) and coalescing hold; at 100x the DB is still the ceiling for misses, unchanged from today. Single point of failure: the one process, unchanged. Security architecture unchanged (Section 3).\n2. **Error and rescue map.** 10 error paths mapped (registry below). Two gaps, F2 and F3, both closed. No catch-all handlers introduced; all repository errors keep typed mapping.\n3. **Security and threat model.** No issues. Threats assessed: cross-tenant read via key collision (mitigated by retained unambiguous key contract and validated IDs; likelihood Low, impact High); authorization bypass via cached value (mitigated: authz runs before repository; Low/High); memory exhaustion via many distinct keys (mitigated by 1000-entry and 16 MiB caps; Low/Med); key or PII leakage in logs and metrics (mitigated: no raw IDs, `fill_discarded` has no labels; Low/Med); cross-tenant eviction thrash (availability only, visible in eviction metric; Med/Low). No new dependencies, endpoints, secrets, or inputs.\n4. **Data flow and interaction edge cases.** Async schedule proved F1 and its remedy (see Diagrams). Shadow paths traced: nil input rejected by retained key validation before the wrapper; empty result maps to sentinel (F2); upstream error propagates uncached with single-flight release. No user-visible interactions; interaction table not applicable. No additional findings.\n5. **Code quality.** One finding, F2. Naming: `toCacheValue`/`fromCacheValue`/`ttlFor` describe intent. No DRY violation: invalidation logic lives once in the single-flight wrapper. Branching in `readProfile` stays under 5. No over-engineering: no general cache framework.\n6. **Test review.** Diagram produced (below). Five test gaps identified, all mapped to accepted remedies (T1 to T5); no unresolved behavioral choice remains. 2 a.m. test: T1 controlled-schedule suite. Hostile QA test: T3 commit-then-timeout. Chaos test: forced adapter failure mid-fill during the 50% stage (T5). Pyramid: mostly unit with fake timers and paused promises; one integration test against the real repository for sentinel mapping; staging fault injection for alerts. Flakiness: TTL and schedule tests must use fake timers and explicit pause/release points, never sleeps.\n7. **Performance.** No issues. Memory bounded by the 16 MiB cap plus the in-flight map (bounded by concurrency). Residual risk, no plan change: with about 900 hot keys and a 1000-entry cap, long-tail keys can evict hot entries and hold hit ratio under 60%; the eviction metric and the 100% stage acceptance check detect this.\n8. **Observability.** One gap, F4. Metrics: hits, misses, evictions, bytes, fallback errors, adapter state, `fill_discarded`, DB CPU, read p95. Logs: structured line on adapter ACTIVE to BYPASS and on reinit, on flag value change, never containing keys. Debuggability: a stale-read report three weeks later can be reconstructed from flag-change and `fill_discarded` timelines.\n9. **Deployment and rollout.** One risk, F5. No migration. Single process, so no old/new code overlap. Rollback: flag off, seconds. Deployment sequence and rollback flowchart in Diagrams. Post-deploy: first 5 minutes check fallback errors zero and hit ratio climbing; first hour check p95, DB CPU, error rate against the health definition.\n10. **Long-term trajectory.** Reversibility 5/5 (flag off, then delete the wrapper). Debt: none beyond a code-comment diagram of fill/write ordering, included in T1. Path dependency: repository interface unchanged, so a shared cache tier later swaps the adapter. A new engineer in 12 months reads one wrapper and one diagram.\n11. **Design and UX.** SKIPPED, justified: internal backend change with no UI, API, or user-visible interaction. No /plan-design-review needed.\n\n## NOT in scope\n- Distributed or shared cache tier: explicitly out of scope; Approach C recorded for the future path.\n- Cross-process coherence and prewarming: excluded by the plan; startup rejects multi-process with caching on.\n- Error-class taxonomy for write failures (D5 option B): unnecessary once every write outcome invalidates.\n- Trace-replay load harness for pre-production hit-ratio prediction: the staged rollout with the 100% stage acceptance check verifies the target; not added.\n\n## What already exists\n- LRU adapter with limits, TTL, sentinel, and failure bypass: reused unchanged.\n- Per-key single-flight wrapper: reused, extended with `invalidate(key)` and a fill handle (F1).\n- Repository typed error mapping and contract tests: reused unchanged.\n- Runtime feature flag with per-key percentage: reused; semantics clarified (F5).\n\n## Dream state delta\nThis plan lands the first cache tier behind an unchanged repository interface and with invariant-safe fill ordering. The 12-month ideal (swappable local or shared tier) remains reachable without rework; only the adapter would change.\n\n## Error and Rescue Registry\n| Method/codepath | What can go wrong | Error class | Rescued? | Action | User sees |\n|---|---|---|---|---|---|\n| readProfile: flag.enabledFor | Flag store unavailable | Flag read error | Y | Evaluate as disabled, bypass cache, log once (F5) | Nothing; DB path |\n| readProfile: cache.get | Adapter failure | Adapter error | Y (existing) | Adapter enters BYPASS, metric, log | Nothing; DB path |\n| readProfile: repository.read | Timeout, connection error, pool exhausted | Existing typed API errors | Y (existing) | Propagate typed error; single-flight releases; nothing cached | Existing error response |\n| readProfile: repository.read | Record absent | Existing absent result (value or typed NotFound) | Y | Store sentinel 10 s; hit maps back to same result (F2) | Same as uncached |\n| readProfile: cache.set | Adapter failure, byte cap exceeded | Adapter error | Y (existing) | BYPASS or skip entry; value still returned | Nothing |\n| readProfile: fill resolves after write | Stale snapshot | none (logic) | Y | Discard, `fill_discarded` (F1) | Nothing |\n| writeProfile: repository.write | Validation or constraint failure | Existing typed errors | Y | Propagate; finally invalidates and deletes (F3) | Existing error response |\n| writeProfile: repository.write | Commit-ack timeout (indeterminate) | Existing timeout error | Y | Propagate; finally invalidates and deletes (F3) | Existing error response |\n| writeProfile: cache.delete | Adapter failure | Adapter error | Y (existing) | BYPASS, metric, log | Nothing |\n| Flag value change | In-flight fill lands after flush | none (logic) | Y | Flush also invalidates in-flight fills (F5) | Nothing |\n\nNo catch-all handlers. No CRITICAL GAPS remaining.\n\n## Failure Modes Registry\n| Codepath | Failure mode | Rescued? | Test? | User sees? | Logged? |\n|---|---|---|---|---|---|\n| Fill vs write ordering | Stale fill cached after write (F1) | Y (D3) | Y (T1) | Nothing | Metric |\n| Single-flight coalescing | Post-write reader joins pre-write flight (F1) | Y (D3) | Y (T1) | Nothing | Metric |\n| Negative caching | Sentinel leaked or never stored (F2) | Y (D4) | Y (T2) | Nothing | n/a |\n| Write failure | Stale value kept after indeterminate commit (F3) | Y (D5) | Y (T3) | Existing error | Existing |\n| Adapter failure | Silent bypass, hit ratio to zero (F4) | Y (existing) | Y (T5) | Nothing | Alert + log |\n| Flag transition | Bucket mismatch or unflushed fill (F5) | Y (D7) | Y (T4) | Nothing | Log |\n| DB error on miss | Error cached | Y (existing) | Y (existing) | Existing error | Existing |\n| Capacity | Hot set evicted by tail keys | n/a | Rollout check | Slower reads | Metric |\n\n0 rows with RESCUED=N, TEST=N, USER SEES=Silent. 0 CRITICAL GAPS remaining (1 found, closed by D3).\n\n## TODOS.md updates\nNo TODOs proposed. HOLD SCOPE permits only evidenced gaps in accepted scope; every evidenced gap (F1 to F5) is remedied in this plan, and the remaining candidates (Approach C, replay harness, error taxonomy) are expansions or alternatives to adequate approved remedies. TODOS-format.md was not read (outside the permitted path set); no TODOS.md exists in the repo.\n\n## Diagrams\n\n**1. System architecture (before and after)**\n```\n BEFORE: caller \u2500\u25b6 authn/authz \u2500\u25b6 repository \u2500\u25b6 DB\n\n AFTER: caller \u2500\u25b6 authn/authz \u2500\u25b6 readProfile/writeProfile (wrapper)\n \u2502 \u251c\u2500\u25b6 flag.enabledFor(key) (same hash, reads & writes)\n \u2502 \u251c\u2500\u25b6 LRU adapter get/set/delete (1000 / 16 MiB / 30 s, sentinel 10 s)\n \u2502 \u251c\u2500\u25b6 single-flight run/invalidate (in-flight fills)\n \u2502 \u2514\u2500\u25b6 repository \u2500\u25b6 DB (typed errors unchanged)\n \u2514\u2500\u25b6 metrics: hit, miss, evict, bytes, fallback, fill_discarded\n```\n\n**2. Data flow with shadow paths (read)**\n```\n key \u2500\u25b6 [validated upstream] \u2500\u25b6 flag? \u2500\u25b6 cache.get \u2500\u25b6 miss \u2500\u25b6 single-flight \u2500\u25b6 repository.read \u2500\u25b6 set? \u2500\u25b6 return\n \u2502 \u2502 \u2502 \u2502 \u2502 \u2502\n nil/invalid: rejected off/error: hit: fromCacheValue coalesce: join error: throw invalidated:\n before wrapper DB path (sentinel -> absent) same flight typed, no set discard+metric\n absent: sentinel 10 s\n```\n\n**3. Async schedule (F1) with one column per operation and shared state**\n```\n t | R1 readProfile(k) | W writeProfile(k, v2) | R2 readProfile(k) | cache[k] | inflight[k]\n 1 | miss; flight f1; await read | | | - | f1\n 2 | | await write ... commit v2 | | - | f1\n 3 | | delete(k) no-op; return | | - | f1\n 4 | | | begins; miss; joins f1 | - | f1\n 5 | read resolves v1; set(k, v1) | | receives v1 VIOLATION | v1 BAD | -\n Order B: R2 begins after t5 -> cache hit v1 VIOLATION (until TTL or next write)\n Contract: R2 began after W completed, so R2 must observe v2. Sketch has no preventing mechanism.\n\n AMENDED (D3):\n 2 | (f1.invalidated = true; inflight.delete(k) at write start)\n 4 | | | begins; miss; new flight f2 (post-commit read -> v2)\n 5 | f1 resolves v1 -> returned to R1 only (R1 began before W: allowed); not cached; fill_discarded+1\n During-write fill: f2 issued between write start and commit reads v1 pre-commit; finally-invalidate marks f2; discarded.\n```\n\n**4. State machines**\n```\n cache[key]: ABSENT \u2500\u2500fill(value)\u2500\u2500\u25b6 PRESENT(value, 30 s) \u2500\u2500TTL/evict/delete\u2500\u2500\u25b6 ABSENT\n ABSENT \u2500\u2500fill(absent)\u2500\u25b6 PRESENT(sentinel, 10 s) \u2500\u2500TTL/evict/delete\u2500\u2500\u25b6 ABSENT\n Impossible: PRESENT \u2500\u2500\u25b6 PRESENT with older version (prevented by invalidate-before-set, D3)\n fill: FILLING \u2500\u2500resolve\u2500\u2500\u25b6 SET FILLING \u2500\u2500invalidate\u2500\u2500\u25b6 DETACHED \u2500\u2500resolve\u2500\u2500\u25b6 DISCARDED (metric)\n adapter: ACTIVE \u2500\u2500any failure\u2500\u2500\u25b6 BYPASS \u2500\u2500reinit (flag change or operator)\u2500\u2500\u25b6 ACTIVE (empty)\n flag: OFF \u2500\u2500\u25b6 10% \u2500\u2500\u25b6 50% \u2500\u2500\u25b6 100% ; any transition (either direction) \u2500\u2500\u25b6 flush + invalidate in-flight\n```\n\n**5. Error flow**\n```\n repository error \u2500\u25b6 typed API error (unchanged) \u2500\u25b6 caller; single-flight releases; nothing cached\n adapter error \u2500\u25b6 BYPASS + metric + log \u2500\u25b6 reads/writes go to DB \u2500\u25b6 alert if BYPASS > 1 min (T5)\n write error \u2500\u25b6 finally: invalidate(key); delete(key) \u2500\u25b6 typed error to caller (D5)\n flag error \u2500\u25b6 treated as disabled \u2500\u25b6 DB path \u2500\u25b6 log once\n```\n\n**6. Deployment sequence**\n```\n deploy code (flag OFF) \u2500\u25b6 smoke: reads/writes bypass, metrics emit zeros\n \u2500\u25b6 flag 10% \u2500\u25b6 1 healthy hour (T5 definition) \u2500\u25b6 50% \u2500\u25b6 1 healthy hour \u2500\u25b6 100%\n \u2500\u25b6 5 min warmup \u2500\u25b6 hit ratio >= 60%, DB CPU < 50%, p95 < 60 ms \u2500\u25b6 record results\n```\n\n**7. Rollback flowchart**\n```\n regression (error rate, p95, fallback alert, hit ratio) \u2500\u25b6 flag OFF (seconds)\n \u2500\u25b6 verify: hits -> 0, DB CPU within capacity, error rate baseline\n \u2500\u25b6 persistent defect? \u2500\u25b6 revert wrapper commit \u2500\u25b6 redeploy (repository interface unchanged)\n```\n\n## Stale diagram audit\nNo ASCII diagrams exist in the submitted plan or (per plan description) in the touched adapters. New code-comment diagram of fill/write ordering is added by T1 and must be kept current with the wrapper.\n\n## Implementation Tasks\nSynthesized from this review's findings. Each task derives from a specific\nfinding above. Run with Claude Code or Codex; checkbox as you ship.\n\n- [ ] **T1 (P1, human: ~4h / CC: ~20min)** \u2014 single-flight wrapper + profile cache wrapper \u2014 Add `invalidate(key)` and fill handle; invalidate before and after writes; discard invalidated fills with `profile_cache.fill_discarded`; add ordering comment diagram\n - Surfaced by: Section 1 / Section 4 \u2014 F1 stale fill and coalescing violate read-after-write contract\n - Files: single-flight wrapper module; new profile cache wrapper module; wrapper tests (paths unverified this run)\n - Verify: controlled-schedule tests for order A (fill resolves before R2), order B (R2 joins in-flight), and during-write fill; assert R2 gets v2 and cache never holds v1\n- [ ] **T2 (P1, human: ~2h / CC: ~10min)** \u2014 profile cache wrapper \u2014 Implement `toCacheValue`/`fromCacheValue`/`ttlFor`; sentinel 10 s, value 30 s\n - Surfaced by: Section 2 / Section 5 \u2014 F2 sketch leaks sentinel or silently never negative-caches\n - Files: profile cache wrapper module; wrapper tests; one integration test against the repository's absent result\n - Verify: second absent read within 10 s does not call repository; caller result identical to uncached; fake-timer TTL expiry at 10 s and 30 s\n- [ ] **T3 (P1, human: ~1h / CC: ~5min)** \u2014 profile cache wrapper \u2014 Move invalidate and delete into `finally` for every write outcome; replace \"failed-write preservation\" test with \"failed-write invalidation\"\n - Surfaced by: Section 2 \u2014 F3 ambiguous-outcome write leaves stale value for up to 30 s\n - Files: profile cache wrapper module; wrapper tests\n - Verify: write rejects after simulated commit; next read calls repository and returns v2\n- [ ] **T4 (P1, human: ~2h / CC: ~10min)** \u2014 profile cache wrapper + flag integration \u2014 Same deterministic hash for reads and writes; any flag value change flushes and invalidates in-flight fills; flag-read error evaluates as disabled\n - Surfaced by: Section 9 \u2014 F5 flag transition semantics unspecified\n - Files: profile cache wrapper module; flag integration; wrapper tests\n - Verify: property test that read and write buckets agree; flag 10 to 50 and 50 to 10 both flush; in-flight fill across flush is discarded; flag error bypasses cache\n- [ ] **T5 (P2, human: ~4h / CC: ~30min)** \u2014 observability \u2014 Dashboard panels (hit ratio, evictions, bytes, adapter state, fallback errors, fill_discarded, DB CPU, p95), alerts (BYPASS > 1 min; hit ratio < 60% at 100% stage after warmup; fallback errors > 0; fill_discarded > 1% of fills), runbook (disable flag, verify, reinit), health definition for staged rollout\n - Surfaced by: Section 8 / Section 9 \u2014 F4 monitoring named without thresholds, alerts, or runbook\n - Files: metrics emission in wrapper; dashboard and alert config; runbook doc\n - Verify: staging fault injection (forced adapter failure, forced miss storm) fires each alert; runbook steps restore ACTIVE state\n\n_No new tasks from Section 3, Section 7, Section 10, Section 11._\n\nTasks JSONL artifact: not written this run (no shell; ~/.gstack writes excluded by run rules).\n\n## Completion Summary\n```\n +====================================================================+\n | MEGA PLAN REVIEW \u2014 COMPLETION SUMMARY |\n +====================================================================+\n | Mode selected | HOLD SCOPE |\n | System Audit | SKIPPED per run rules (no repo exploration) |\n | Step 0 | Approach A (D1), HOLD SCOPE (D2), 0E: D3-D7 |\n | Section 1 (Arch) | 1 issue found (F1) |\n | Section 2 (Errors) | 10 error paths mapped, 2 GAPS (F2, F3) |\n | Section 3 (Security)| 0 issues found, 0 High severity |\n | Section 4 (Data/UX) | 8 edge cases mapped, 0 unhandled (F1 fixed) |\n | Section 5 (Quality) | 1 issue found (F2) |\n | Section 6 (Tests) | Diagram produced, 5 gaps (all in T1-T5) |\n | Section 7 (Perf) | 0 issues found (1 residual risk noted) |\n | Section 8 (Observ) | 1 gap found (F4) |\n | Section 9 (Deploy) | 1 risk flagged (F5) |\n | Section 10 (Future) | Reversibility: 5/5, debt items: 0 |\n | Section 11 (Design) | SKIPPED (no UI scope) |\n +--------------------------------------------------------------------+\n | NOT in scope | written (4 items) |\n | What already exists | written |\n | Dream state delta | written |\n | Error/rescue registry| 10 methods, 0 CRITICAL GAPS |\n | Failure modes | 8 total, 0 CRITICAL GAPS (1 found, closed) |\n | TODOS.md updates | 0 items proposed |\n | Scope proposals | 0 proposed, 0 accepted (HOLD) |\n | CEO plan | skipped (HOLD) |\n | Outside voice | codex: disabled (config); no fallback run |\n | Lake Score | 5/5 recommendations chose complete option |\n | Diagrams produced | 7 (arch, data flow, async schedule, state, |\n | | error flow, deployment, rollback) |\n | Stale diagrams found | 0 |\n | Unresolved decisions | 0 |\n +====================================================================+\n```\n\n## Unresolved Decisions\nNone. All seven decision points were auto-selected to the recommended option per run rules and are recorded in the Decision Registry.\n\n## Next steps\nRun `/plan-eng-review` next (required shipping gate). No design review needed. Review-log, decision-log, telemetry, and dashboard persistence were skipped per run rules; the report below reflects this run only.\n\n## GSTACK REVIEW REPORT\n\n| Review | Trigger | Why | Runs | Status | Findings |\n|--------|---------|-----|------|--------|----------|\n| CEO Review | `/plan-ceo-review` | Scope & strategy | 1 | clean | mode: HOLD_SCOPE, 5 findings, 5 remedies accepted, 0 critical gaps open |\n| Outside Review | codex (`codex_reviews: disabled`) | Independent 2nd opinion | 0 | disabled | none (step skipped, no fallback) |\n| Eng Review | `/plan-eng-review` | Architecture & tests (required) | 0 | \u2014 | \u2014 |\n| Design Review | `/plan-design-review` | UI/UX gaps | 0 | \u2014 | \u2014 |\n| DX Review | `/plan-devex-review` | Developer experience gaps | 0 | \u2014 | \u2014 |\n\n- **OUTSIDE COVERAGE:** provider codex, phase plan-review, status disabled by `.gstack-section-state-xMGd4Z/config.yaml`; no outside or native fallback pass ran; no outside findings.\n- **VERDICT:** CEO CLEARED (HOLD SCOPE, 0 unresolved, 0 critical gaps) \u2014 eng review required\n\nNO UNRESOLVED DECISIONS\n",
|
|
"finding": "| F1 | 1, 4 | CRITICAL GAP (fixed by D3) | Contract: \"Every read begun after that write completes must observe the committed version.\" Sketch: `cache.set(key, value)` after `await repository.read` with \"no additional version checks or coordination\". Schedule below shows a post-write reader receiving the pre-write snapshot via a stale fill or via single-flight coalescing. | D3: single-flight `invalidate(key)` before and after the write; invalidated fills never `set`; `fill_discarded` metric | Fills overlapping a write are not cached (bounded hit-rate cost, visible in metric) | T1 tests: both completion orders plus during-write fill; assert post-write reader sees v2 and cache never holds v1 |",
|
|
"heading": "**3. Async schedule (F1) with one column per operation and shared state**",
|
|
"trace": " t | R1 readProfile(k) | W writeProfile(k, v2) | R2 readProfile(k) | cache[k] | inflight[k]\n 1 | miss; flight f1; await read | | | - | f1\n 2 | | await write ... commit v2 | | - | f1\n 3 | | delete(k) no-op; return | | - | f1\n 4 | | | begins; miss; joins f1 | - | f1\n 5 | read resolves v1; set(k, v1) | | receives v1 VIOLATION | v1 BAD | -\n Order B: R2 begins after t5 -> cache hit v1 VIOLATION (until TTL or next write)\n Contract: R2 began after W completed, so R2 must observe v2. Sketch has no preventing mechanism.\n\n AMENDED (D3):\n 2 | (f1.invalidated = true; inflight.delete(k) at write start)\n 4 | | | begins; miss; new flight f2 (post-commit read -> v2)\n 5 | f1 resolves v1 -> returned to R1 only (R1 began before W: allowed); not cached; fill_discarded+1\n During-write fill: f2 issued between write start and commit reads v1 pre-commit; finally-invalidate marks f2; discarded."
|
|
},
|
|
"retry": {
|
|
"sourceProof": ".context/ship-source-ag-delta-paid-20260910-v1/sdk-retry-report-native-v1/proof.json",
|
|
"sourceProofSha256": "abc769713669f6759b934c7ae23306d411f2bfc1fb2f8ba1ff022e19c9e24bce",
|
|
"reportSha256": "02bcd206f50606a61b5f6240d5d455f1330f3df2608149d2a402294ea697456e",
|
|
"report": "# Plan: cache profile summaries in one process\n\nReviewed by /plan-ceo-review on 2026-09-10 (HOLD SCOPE, spawned run: every decision\nauto-selected the recommended option; see Decision Register). Original plan content is\npreserved below; amendments are marked `[Amended: F#]` and point at the finding record.\n\n## Measured problem and accepted scope\nThe existing profile-summary service has one active process. A one-week trace\nshows repeated reads of about 900 hot keys: DB CPU is 70%, with read p95 120 ms.\nAdd a process-local LRU wrapper to the existing repository. Acceptance targets\nare at least 60% cache hits, DB CPU below 50%, and read p95 below 60 ms, with the\nexisting error-rate and correctness SLOs unchanged. This is an internal backend\nchange with no UI, API, schema, pricing, or developer onboarding change.\n\n`[Amended: F7]` Targets are unchanged. Hit rate is measured over flag-enabled keys\nonly, and read latency is recorded separately for hits and misses so the owner can\nattribute any p95 result (see Metric Schema).\n\n## Existing contracts retained\n- All reads and writes use this repository in the same process; there are no\n external DB writers. Multi-process operation remains unsupported and startup\n rejects that configuration while caching is enabled.\n- Authentication and authorization run before repository access. Keys encode\n the authenticated tenant ID and validated profile ID without ambiguity.\n Values are immutable profile-summary DTOs; secrets and cache keys are never\n logged. Cached results cannot bypass authorization.\n- The existing LRU adapter supports 1000 entries, a 16 MiB byte cap, and a\n 30-second TTL. Recorded hot data fits those limits. Absent records use a\n distinct sentinel with a 10-second TTL; undefined means a cache miss.\n- Cache operations are synchronous and atomic in the single JS event loop.\n On any cache failure the existing adapter bypasses the cache until an empty\n cache is reinitialized; repository errors keep the current typed API error\n mapping. The existing per-key\n single-flight wrapper coalesces simultaneous misses and releases on failure.\n- A read already in progress when a write commits may return its earlier DB\n snapshot to that caller. Every read begun after that write completes must\n observe the committed version. TTL expiry is not a substitute for this rule.\n\n`[Amended: F1]` Boundary definition used by this plan: a write \"completes\" when\n`writeProfile` resolves to its caller. `invalidate()` runs synchronously inside that\nsame continuation, before resolution, so any read that begins after the caller\nobserves completion cannot hit a pre-write entry or join a pre-write DB read.\n\n## Proposed wrapper integration\nKeep the current read-through repository interface and shared adapters.\n\n`[Amended: F1, F2, F4, F10]` The original sketch stated that no coordination between a\ncache fill and a write was proposed. The review showed that omission violates the\nretained read-after-write contract (F1). The rules below replace it and are the\ncomplete read/write ordering rules.\n\n```javascript\n// Composition (read): flag \u2500\u25b6 cache.get \u2500\u25b6 pending-fill token \u2500\u25b6 single-flight \u2500\u25b6 repository.read \u2500\u25b6 guarded fill\n// Composition (write): repository.write \u2500\u25b6 invalidate(key) [always, flag-independent]\nconst pendingFills = new Map(); // key -> Set<{cancelled:boolean}>; entries exist only while a miss is in flight\n\nasync function readProfile(key) {\n if (!cacheFlag.enabledFor(key)) return repository.read(key); // F10: flag gates reads only\n const cached = cache.get(key); // undefined = miss or adapter bypass\n if (cached !== undefined) return cached === ABSENT ? null : cached; // F4: unwrap sentinel (null = existing absent result)\n const token = { cancelled: false };\n track(pendingFills, key, token);\n let value;\n try {\n value = await singleFlight.run(key, () => repository.read(key)); // existing wrapper; releases on failure\n } finally {\n untrack(pendingFills, key, token);\n }\n if (token.cancelled) { // F1: a write invalidated this key mid-read\n metrics.inc('profile_cache_fill_cancelled_total');\n } else {\n cache.set(key, value === null ? ABSENT : value); // adapter applies 10 s / 30 s TTL, byte cap, oversize reject (F6)\n }\n return value; // caller keeps its own snapshot (allowed by contract)\n}\n\nasync function writeProfile(key, update) {\n let saved;\n try {\n saved = await repository.write(key, update);\n } catch (err) {\n if (isIndeterminateWriteError(err)) invalidate(key, 'write_indeterminate'); // F2: DB state unknown\n throw err; // existing typed mapping unchanged\n }\n invalidate(key, 'write_committed');\n return saved;\n}\n\nfunction invalidate(key, reason) {\n for (const t of pendingFills.get(key) ?? []) t.cancelled = true; // (a) in-flight fills must not land\n singleFlight.detach(key); // (b) later readers start a fresh DB read\n cache.delete(key); // (c) drop PRESENT or ABSENT entry\n metrics.inc('profile_cache_invalidate_total', { reason });\n}\n```\n\n`isIndeterminateWriteError` maps the repository's existing typed errors: validation,\nauthorization, not-found and conflict rejections are deterministic pre-commit\nfailures and preserve the cache; timeout, connection-lost, and any unrecognised\nclass are indeterminate and invalidate. Unknown defaults to invalidate.\n`singleFlight.detach(key)` is added to the existing wrapper if it lacks one: the\nin-flight promise still resolves to its existing waiters (they began before the\ncommit); the slot is simply no longer joinable.\n\n## Verification and rollout\nExisting repository contract tests cover tenant isolation, key validation,\nabsence, DB failures, and authorization. New wrapper tests cover hit/miss,\neviction and byte limits, TTL, adapter-failure fallback, successful-write\ninvalidation, failed-write preservation, and concurrent-miss coalescing.\n\n`[Amended: F1, F2, F6, F10]` Additional wrapper tests: the four async schedules\nS1\u2013S4 (Section 4) with controlled pause/release points; failed-write behaviour split\ninto deterministic-rejection preservation and indeterminate-error invalidation;\noversize value rejected without evicting others; flag percentage change yields an\nempty cache; write on an unflagged key still invalidates. Assertions are listed in\nSection 6.\n\nThe rollout uses the existing runtime feature flag: enable for 10% of keys,\nthen 50%, then all keys after one healthy hour at each stage. Monitor hit/miss,\neviction, cache bytes, fallback errors, DB CPU, and read p95 without raw IDs.\nOn error-rate or latency regression, disable the flag immediately; both reads\nand writes bypass the cache while disabled, and enabling creates an empty cache.\nCold starts remain within the existing DB capacity. The service owner monitors\nthe rollout and records the results against the acceptance targets.\n\n`[Amended: F10, F11]` Any change of the flag value, including percentage steps in\neither direction, reinitializes an empty cache. Writes call `invalidate()` regardless\nof flag state (a no-op on an empty cache). Deploy ships with the flag off; the owner\nenables 10% only after confirming exactly one process is serving.\n\n## Out of scope\nDistributed caching, cross-process coherence, prewarming, changing consistency\nsemantics, or adding new product surfaces. The repository interface preserves a\nfuture replacement path without introducing a general cache framework now.\n\n---\n\n# CEO REVIEW RECORD (HOLD SCOPE)\n\n## Run conditions\n- Mode: HOLD SCOPE, selected by the user up front. No expansions surfaced.\n- Skipped by instruction: system audit, environment setup, telemetry, codebase exploration, design-doc and handoff checks, prior-learnings and brain context.\n- No shell tool in this run: `gstack-review-log`, `gstack-decision-log`, tasks JSONL, learnings log, handoff cleanup and `gstack-review-read` were not executed. Recorded here instead.\n- Outside voice: `codex_reviews: disabled` in `.gstack-section-state-0O8BAs/config.yaml`. Entire step skipped, no native fallback. `outside_status: disabled` (record not persisted, no shell).\n- Codex review skipped (codex_reviews disabled). Re-enable: `gstack-config set codex_reviews enabled`.\n\n## Step 0\n**0A Premise.** Real, measured pain (DB CPU 70%, p95 120 ms, 900 hot keys over one week). Doing nothing leaves no headroom for growth and no lever short of DB scaling. A process-local read-through cache is the most direct path; the single-process, single-writer topology makes it correct without distributed coordination. Not a proxy problem.\n\n**0B Existing leverage.** Reused: LRU adapter (limits, TTL, sentinel, failure bypass), per-key single-flight wrapper, runtime feature flag, typed repository error mapping, repository contract tests. Nothing is rebuilt. One addition to an existing component: `singleFlight.detach(key)`.\n\n**0C Dream state.**\n```\nCURRENT STATE THIS PLAN 12-MONTH IDEAL\nsingle process, every read read-through LRU in front of same repository interface; cache\nhits DB; CPU 70%, p95 120 ms --> repository; explicit invalidate --> backend swappable (local or shared)\nno invalidation protocol protocol; metrics per outcome if multi-process ever arrives\n```\nMoves toward the ideal: invalidation lives in one function behind the existing interface.\n\n**0C-bis Approaches.**\n| | A: Sketch as written | B: Wrapper + in-process invalidation protocol | C: DB row-version compare on fill |\n|---|---|---|---|\n| Summary | get/set/delete, no coordination | pending-fill tokens, single-flight detach, flag-independent invalidate | fill only if row version unchanged |\n| Effort | S | S (human ~1 day / CC ~30 min) | M |\n| Risk | High: violates read-after-write contract | Low | Med |\n| Completeness | 4/10 | 10/10 | 9/10 |\n| Reuses | adapter, flag | adapter, flag, single-flight, typed errors | needs version column: schema change, out of scope |\n\nD1 auto-selected **B** (recommended): only approach meeting the stated invariant with no schema change; smallest diff that is still correct (engineering preference: explicit over clever, right-sized diff).\n\n**0F Mode.** D2: HOLD SCOPE, pre-selected by user. Committed.\n\n**0D HOLD analysis.** Touches about five files (wrapper module, single-flight detach, flag hook, metrics, tests), one new module. No complexity smell. Minimum set equals accepted scope; repairs F1\u2013F12 are required to meet stated invariants and are in scope. Nothing deferrable without breaking a stated target.\n\n**0E Temporal interrogation.** Resolved now: write-completion boundary (F1), error classification (F2), sentinel unwrap and composition order (F4), flag semantics on change (F10), metric names (Metric Schema), deploy ordering (F11). Human ~6 h / CC ~45 min.\n\n## Findings Register\nOne row per finding. Sections cross-reference by F#.\n\n| F# | Sec | Sev | Evidence (plan text) | Remedy (auto-selected, recommended) | Residual risk | Verification |\n|---|---|---|---|---|---|---|\n| F1 | 1, 4 | CRITICAL | \"no additional version checks or coordination between a cache fill and a write\" vs contract \"every read begun after that write completes must observe the committed version\". Schedule S1: R1 misses, W commits and deletes, R1 fills stale v1, R2 hits v1. S2: R2 joins R1's in-flight single-flight read after commit. | `invalidate()` cancels pending fills, detaches single-flight slot, deletes entry; runs before `writeProfile` resolves. Counter `fill_cancelled_total`. | Contract boundary defined as caller-visible completion; owner may confirm | Tests S1\u2013S4 with pause/release |\n| F2 | 2 | HIGH | \"failed-write preservation\" with no distinction for errors where DB state is unknown (timeout after commit). Stale entry up to 30 s despite committed write. | Classify typed errors; deterministic rejections preserve, indeterminate or unknown invalidate. | Classification list must match repository's error classes | Two tests: preserve on validation error; invalidate on timeout |\n| F3 | 3 | LOW | Sentinel entries occupy LRU slots; an authenticated user requesting many absent IDs evicts hot entries. | Visibility: `evictions_total{kind}` and `entries{kind}` gauges. Impact already bounded by 10 s TTL and per-tenant keys. | Hit rate can dip under abuse; visible, not silent | Metric present in tests |\n| F4 | 5 | MED | Sketch returns raw `cached` (would hand the ABSENT sentinel to callers); omits single-flight and flag composition. | Amended sketch with unwrap and explicit composition order. | None | Test: absent record read returns existing absent result twice (miss then hit) |\n| F6 | 2, 7 | MED | Byte cap stated; behaviour when one value exceeds cap unspecified. Shadow path of `cache.set`. | Adapter rejects oversize set, returns without caching, `oversize_total` counter. | None | Test: oversize value not cached, other entries intact |\n| F7 | 7 | MED | Targets: 60% hits and p95 < 60 ms. At 60% hits p95 is a miss (~120 ms) unless DB latency itself falls. | Record hit and miss latency separately; acceptance write-up states which mechanism delivered p95. Targets unchanged. | p95 target may be missed at 60% hits; now attributable | Dashboard panel by outcome |\n| F8 | 8 | HIGH | \"adapter bypasses the cache until an empty cache is reinitialized\": no trigger, alert or runbook named. Silent return to 70% DB CPU. | Structured error log (class, operation, no key), `fallback_total`, gauge `bypassed`, alert if bypassed > 5 min, runbook: flag off then on. | None | Test: adapter throw sets gauge and logs |\n| F9 | 8 | MED | Hit rate at 10% stage would read ~6% overall; \"healthy hour\" uninterpretable. | `lookup_total{result=hit,miss,bypass_unflagged,bypass_fallback}`; hit rate = hit/(hit+miss). | None | Metric label test |\n| F10 | 9 | HIGH | Percentage ramp-down leaves entries for now-unflagged keys; writes to them bypass cache; ramp-up serves stale. \"enabling creates an empty cache\" covers off-to-on only. | Any flag value change reinitializes empty cache; writes invalidate regardless of flag. | None | Tests: flag change empties cache; unflagged write invalidates |\n| F11 | 9 | MED | \"one active process\" is a startup check; a deploy overlap window has two processes, one with a cache. | Deploy flag-off; enable 10% after single process confirmed; disable before any deploy. | Relies on owner discipline; add to runbook | Rollout checklist |\n| F12 | 10 | LOW | Invalidation protocol is non-obvious to a reader in 12 months. | ASCII schedule (Section 4) in wrapper module header comment; keep it in sync. | Stale diagram if protocol changes | Review checklist item |\n\nF5 was merged into F4 during review (same failure mode: sketch omissions). Numbering retained for traceability.\n\n## Decision Register\n| D# | Subject | Choice | Basis |\n|---|---|---|---|\n| D1 | Approach | B | Only approach that meets the invariant without schema change |\n| D2 | Mode | HOLD SCOPE | User-selected |\n| D3\u2013D13 | F1, F2, F3, F4, F6, F7, F8, F9, F10, F11, F12 | Recommended remedy, each individually | Spawned run: auto-selected; each is the complete option |\n| D14 | TODOS.md | None proposed | No evidenced gap left out of scope under HOLD |\n\nLake Score: 12/12 decisions with a completeness axis chose the complete option.\n\n## Section outcomes\n1. **Architecture.** F1 found. Diagram: System Architecture, Dependency Before/After. Coupling added: wrapper now depends on single-flight detach and flag; justified, all in-process. 10x load: 900 hot keys fit; misses scale with key cardinality, bounded by LRU. 100x: byte cap hit first, evictions rise, visible via F3 metrics. SPOF: the single process (pre-existing). Rollback: flag off (seconds), then revert.\n2. **Error and rescue.** F2, F6 found. Registry below. No catch-all handlers introduced; `writeProfile` rethrows after classification.\n3. **Security.** F3 found (Low/Med). No new endpoints, params, secrets or dependencies. Keys tenant-scoped and validated upstream. Added assertion, not a finding: repository read result is caller-independent (field filtering happens above the repository), so a cached DTO cannot widen authorization. Injection: keys built from validated IDs; no string interpolation into queries added. No PII added to logs.\n4. **Data flow and edge cases.** F1 (schedules) found. Diagram: Data Flow with shadow paths, Async Schedule. No UI interactions.\n5. **Code quality.** F4 found. One new module; `invalidate()` is the single place for invalidation (DRY). No method branches more than four times.\n6. **Tests.** Gaps map to F1, F2, F4, F6, F10; test diagram and assertions below. Pyramid: many unit, two integration (adapter failure, flag change), no E2E needed. Flakiness: TTL tests use fake timers; schedule tests use explicit pause/release, not sleeps.\n7. **Performance.** F6, F7 found. `cache.get` O(1); size estimate per fill acceptable at hot-key volume. Miss storm at enable equals today's baseline (single-flight coalesces). No N+1, no new indexes, no new connections.\n8. **Observability.** F8, F9 found. Metric Schema below. Debuggability residual: keys are never logged by contract, so post-hoc investigation uses request IDs plus per-outcome metrics.\n9. **Deployment.** F10, F11 found. No migration. Diagrams: Deployment Sequence, Rollback Flowchart.\n10. **Long-term.** F12 found. Debt: none structural; `pendingFills` is bounded by concurrent misses (no leak). Reversibility 5/5. Interface preserved for a future backend swap (Out of scope respected).\n11. **Design and UX.** SKIPPED, justified: no UI, API or user-visible surface.\n\n## Error and Rescue Registry\n```\nMETHOD/CODEPATH | WHAT CAN GO WRONG | EXCEPTION / SIGNAL | RESCUED? | ACTION | USER SEES\n---------------------------|------------------------------------------|-----------------------------|----------|-----------------------------------------------|---------------------------\nreadProfile \u2192 cache.get | adapter failure | adapter internal error | Y | adapter bypass, log class, gauge bypassed=1 (F8)| nothing; DB latency\nreadProfile \u2192 repo.read | DB timeout / connection / pool exhausted | existing typed DB errors | Y (exist)| single-flight releases, no fill, rethrow | existing typed API error\nreadProfile \u2192 repo.read | record absent | existing absent result | Y | cache ABSENT 10 s, unwrap on hit (F4) | existing not-found\nreadProfile \u2192 cache.set | value exceeds byte cap | adapter oversize reject | Y | not cached, oversize_total (F6) | nothing\nreadProfile \u2192 fill | write committed mid-read | token.cancelled | Y | skip fill, fill_cancelled_total (F1) | its own snapshot (allowed)\nwriteProfile \u2192 repo.write | deterministic rejection (validation/authz/not-found/conflict) | typed errors | Y | preserve cache, rethrow (F2) | existing typed API error\nwriteProfile \u2192 repo.write | indeterminate (timeout, connection lost, unknown) | typed/unknown errors | Y | invalidate, rethrow (F2) | existing typed API error\ninvalidate \u2192 cache.delete | adapter failure | adapter internal error | Y | adapter bypass (empty cache on reinit), F8 signals | nothing\ncacheFlag change | stale entries for re-included keys | none (state hazard) | Y | reinit empty cache (F10) | nothing\nstartup | more than one process configured | existing startup rejection | Y (exist)| process refuses to start | deploy failure\n```\nNo unrescued rows. CRITICAL GAPS: 0 after remedies (F1 was CRITICAL before remedy).\n\n## Failure Modes Registry\n```\nCODEPATH | FAILURE MODE | RESCUED? | TEST? | USER SEES? | LOGGED?\n--------------------|----------------------------------|----------|-------|-----------------------|-------------------------\nread fill | stale fill after commit (F1) | Y | Y S1 | correct data | metric fill_cancelled\nread single-flight | join pre-commit read (F1) | Y | Y S2 | correct data | metric invalidate\nwrite | indeterminate error (F2) | Y | Y | typed error | metric invalidate{reason}\nwrite | deterministic rejection (F2) | Y | Y | typed error | existing\ncache adapter | any adapter failure (F8) | Y | Y | slower reads | error log + gauge + alert\ncache set | oversize value (F6) | Y | Y | nothing | metric oversize\nflag | percentage change (F10) | Y | Y | nothing | metric cache_reset\ndeploy | two processes overlap (F11) | Y (proc) | N* | none (flag off) | runbook step\nLRU | sentinel pollution (F3) | Y | Y | lower hit rate | metric evictions{kind}\n```\n*F11 is a procedural control verified by the rollout checklist, not an automated test. CRITICAL GAPS: 0.\n\n## Metric Schema (F3, F6, F7, F8, F9, F1, F10)\n| Metric | Labels | Purpose |\n|---|---|---|\n| profile_cache_lookup_total | result=hit,miss,bypass_unflagged,bypass_fallback | hit rate over flagged keys (F9) |\n| profile_read_duration_seconds | outcome=hit,miss,bypass | attribute p95 (F7) |\n| profile_cache_invalidate_total | reason=write_committed,write_indeterminate | F1, F2 |\n| profile_cache_fill_cancelled_total | none | F1 guard firing |\n| profile_cache_evictions_total | kind=present,absent | F3 |\n| profile_cache_entries | kind=present,absent | F3 |\n| profile_cache_bytes | none | existing requirement |\n| profile_cache_oversize_total | none | F6 |\n| profile_cache_fallback_total, profile_cache_bypassed (gauge) | none | F8; alert bypassed=1 for 5 min |\n| profile_cache_reset_total | reason=flag_change,enable | F10 |\nNo raw IDs or keys in any label or log line.\n\n## Diagrams\n\n### 1. System architecture and dependency graph (before \u2192 after)\n```\nBEFORE: handler \u2500\u25b6 authn/authz \u2500\u25b6 repository \u2500\u25b6 DB\nAFTER: handler \u2500\u25b6 authn/authz \u2500\u25b6 profileCacheWrapper \u2500\u252c\u2500\u25b6 cacheFlag\n \u2502 \u251c\u2500\u25b6 LRU adapter (limits, TTL, sentinel, bypass)\n \u2502 \u251c\u2500\u25b6 singleFlight (run, detach*) *new method\n \u2502 \u251c\u2500\u25b6 metrics\n \u2514\u2500\u2500\u2500\u2500\u2500\u2500\u2500\u2500\u2500\u2500\u2500\u2500\u2500\u2500\u2500\u2500\u2534\u2500\u25b6 repository \u2500\u25b6 DB\n```\n\n### 2. Data flow with shadow paths\n```\nkey \u2500\u2500\u25b6 flag check \u2500\u2500\u25b6 cache.get \u2500\u2500\u25b6 singleFlight(repo.read) \u2500\u2500\u25b6 guarded fill \u2500\u2500\u25b6 value\n \u2502 \u2502 \u2502 \u2502 \u2502\n nil/empty: rejected undefined: miss DB error: release, cancelled: skip fill\n upstream (existing) ABSENT: unwrap rethrow typed (existing) oversize: reject (F6)\n unflagged: repo.read adapter fail: absent: null \u2192 ABSENT adapter fail: bypass (F8)\n bypass (F8)\n```\n\n### 3. State machine (cache entry per key)\n```\n fill(present) fill(absent)\nMISS \u2500\u2500\u2500\u2500\u2500\u2500\u2500\u2500\u2500\u2500\u2500\u2500\u2500\u2500\u2500\u25b6 PRESENT MISS \u2500\u2500\u2500\u2500\u2500\u2500\u2500\u2500\u2500\u2500\u25b6 ABSENT\n \u25b2 TTL 30 s / evict \u2502 \u25b2 TTL 10 s / evict \u2502\n \u2514\u2500\u2500\u2500\u2500\u2500\u2500\u2500\u2500\u2500\u2500\u2500\u2500\u2500\u2500\u2500\u2500\u2500\u2500\u2500\u2500\u2500\u2518 \u2514\u2500\u2500\u2500\u2500\u2500\u2500\u2500\u2500\u2500\u2500\u2500\u2500\u2500\u2500\u2500\u2500\u2500\u2500\u2500\u2500\u2518\n \u25b2 invalidate(key) from PRESENT or ABSENT \u2500\u2500\u25b6 MISS\nInvalid: PRESENT \u2500\u25b6 ABSENT directly (prevented: only fill or invalidate mutate; fill requires MISS)\nInvalid: fill while a newer write has landed (prevented: token.cancelled, F1)\nFlag: OFF \u2500enable\u2500\u25b6 ON(empty) ; ON(p%) \u2500change\u2500\u25b6 ON(empty) ; ON \u2500disable\u2500\u25b6 OFF(dropped) (F10)\n```\n\n### 4. Async schedule (F1)\n```\nstep | R1 read (began before commit) | W write | R2 read (began after W resolves) | cache / pending\n1 | get\u2192miss; token t1 | | | \u2205 / {t1}\n2 | await singleFlight(repo.read) | | |\n3 | | await repo.write \u2192 v2 committed | |\n4 | | invalidate: cancel t1, detach, delete | | \u2205 / {t1\u2717}\n4' | | writeProfile resolves (write \"complete\") | |\n5 | | | get\u2192miss; token t2; fresh DB read | \u2205 / {t1\u2717,t2}\n6 | read resolves v1; t1\u2717 \u2192 no fill | | | \u2205 / {t2}\n7 | | | resolves v2; fill; return v2 | v2 / \u2205\nAlternate order (6 before 4): R1 fills v1, then step 4 deletes it; R2 misses and reads v2. Safe.\nOriginal sketch: step 6 fills v1 after step 4 \u2192 R2 hits v1 \u2192 VIOLATION (S1).\nWithout detach: R2 at step 5 joins R1's promise \u2192 v1 \u2192 VIOLATION (S2).\nS3: W commits before R1 issues read \u2192 R1 reads v2, fill v2 (t1 not cancelled since invalidate preceded track): correct.\nS4: indeterminate write error after commit \u2192 invalidate on error path \u2192 R2 reads v2 (F2).\n```\n\n### 5. Error flow\n```\nrepo.write rejects \u2500\u2500\u25b6 isIndeterminate? \u2500\u2500yes\u2500\u2500\u25b6 invalidate(reason=write_indeterminate) \u2500\u2500\u25b6 rethrow\n \u2514\u2500\u2500no\u2500\u2500\u2500\u2500\u2500\u2500\u2500\u2500\u2500\u2500\u2500\u2500\u2500\u2500\u2500\u2500\u2500\u2500\u2500\u2500\u2500\u2500\u2500\u2500\u2500\u2500\u2500\u2500\u2500\u2500\u2500\u2500\u2500\u2500\u2500\u2500\u2500\u2500\u2500\u2500\u2500\u2500\u2500\u2500\u2500\u2500\u2500\u2500\u2500\u2500\u2500\u2500\u2500\u2500\u25b6 rethrow (cache preserved)\ncache adapter throws \u2500\u2500\u25b6 adapter bypass \u2500\u2500\u25b6 error log (class, op) \u2500\u2500\u25b6 bypassed=1 \u2500\u2500\u25b6 alert after 5 min \u2500\u2500\u25b6 runbook: flag off\u2192on\n```\n\n### 6. Deployment sequence\n```\ndeploy code (flag OFF) \u2500\u25b6 confirm 1 process \u2500\u25b6 flag 10% (empty cache) \u2500\u25b6 1 healthy hour \u2500\u25b6 50% (empty) \u2500\u25b6 1 h \u2500\u25b6 100% (empty)\nhealth = error rate flat, p95 not worse, bypassed=0, hit rate over flagged keys rising\n```\n\n### 7. Rollback flowchart\n```\nregression? \u2500\u2500yes\u2500\u2500\u25b6 flag OFF (cache dropped, reads+writes bypass) \u2500\u2500\u25b6 resolved? \u2500\u2500yes\u2500\u2500\u25b6 investigate with metrics\n \u2502 \u2514\u2500\u2500no\u2500\u2500\u25b6 git revert wrapper commit \u2500\u25b6 redeploy\n no \u2500\u2500\u25b6 continue ramp\n```\n\n### Test diagram (Section 6)\n```\nNEW UX FLOWS: none\nNEW DATA FLOWS: read-through fill; sentinel fill/unwrap; guarded fill\nNEW CODEPATHS: flag gate; token track/untrack; invalidate(a,b,c); isIndeterminate classification; flag-change reset\nNEW ASYNC WORK: none (in-process)\nNEW INTEGRATIONS: singleFlight.detach\nNEW ERROR PATHS: indeterminate write invalidate; oversize reject; adapter bypass signals\n```\nAssertions (unit unless noted):\n- hit/miss: second read of same key makes zero repository calls; wrong result rejected: repository called twice.\n- S1: after W resolves, a fresh read returns v2; rejects v1.\n- S2: fresh read after W does not share R1's promise (repository called again); rejects single call.\n- S3, S4 as in schedule diagram.\n- Deterministic write rejection: cache entry still present, exactly one repository read afterwards; indeterminate: entry gone, one repository read.\n- Absent: read returns existing absent result on miss and on hit; sentinel never escapes.\n- Oversize: `set` returns without storing; unrelated entry still present; counter incremented.\n- TTL (fake timers): entry absent at 30 s + 1 ms; sentinel absent at 10 s + 1 ms.\n- Adapter failure (integration): reads succeed via DB; `bypassed` gauge 1; error log has class and op, no key.\n- Flag change (integration): entries gone after any percentage change; unflagged write still calls invalidate.\n- Authorization: same key yields identical DTO for two callers of different roles (repository is caller-independent).\n- Hostile QA: 1000 concurrent misses on one key produce one repository read; chaos: adapter throws on every third `set`, reads stay correct and gauge flips.\n\n### Stale diagram audit\nCodebase exploration was skipped; no existing diagrams were enumerated. T9 adds the schedule diagram to the module header and the same task checks any existing diagram in the touched files.\n\n## NOT in scope\n- Distributed or shared cache, cross-process coherence: stated out of scope; single process.\n- Prewarming: cold miss burst equals today's baseline load.\n- Hashed-key debug logging: conflicts with the explicit \"keys never logged\" constraint.\n- Auto-rollback on metric thresholds: manual flag-off is seconds; owner monitors each stage.\n- DB row-version compare (Approach C): needs schema change.\n\n## What already exists\nLRU adapter (limits, TTL, sentinel, bypass) reused; single-flight wrapper reused plus `detach`; runtime flag reused with reset-on-change; typed error mapping reused for classification; contract tests reused unchanged.\n\n## Dream state delta\nInvalidation is one function behind the unchanged repository interface. A future shared backend replaces the adapter and `invalidate()` fan-out without touching callers. No step away from the ideal.\n\n## TODOS.md updates\nNone proposed (D14). Every evidenced gap was remedied in scope under HOLD.\n\n## Implementation Tasks\nSynthesized from this review's findings. Each task derives from a specific\nfinding above. Run with Claude Code or Codex; checkbox as you ship.\n\n- [ ] **T1 (P1, human: ~4h / CC: ~20min)** \u2014 wrapper \u2014 Implement pending-fill tokens and `invalidate()` (cancel, detach, delete); add `singleFlight.detach(key)`\n - Surfaced by: Section 1/4 \u2014 F1 stale fill and single-flight join after commit\n - Files: profile cache wrapper module, single-flight wrapper\n - Verify: schedule tests S1\u2013S4 with pause/release points\n- [ ] **T2 (P1, human: ~2h / CC: ~10min)** \u2014 wrapper \u2014 Classify write errors; invalidate on indeterminate, preserve on deterministic rejection\n - Surfaced by: Section 2 \u2014 F2\n - Files: wrapper module, error classification helper\n - Verify: preservation and invalidation tests; unknown class invalidates\n- [ ] **T3 (P1, human: ~1h / CC: ~5min)** \u2014 wrapper \u2014 Sentinel unwrap and explicit composition order (flag, cache, token, single-flight)\n - Surfaced by: Section 5 \u2014 F4\n - Files: wrapper module\n - Verify: absent-record miss-then-hit test; sentinel never returned\n- [ ] **T4 (P1, human: ~2h / CC: ~10min)** \u2014 flag hook \u2014 Reset cache on any flag value change; writes invalidate regardless of flag\n - Surfaced by: Section 9 \u2014 F10\n - Files: flag subscription, wrapper module\n - Verify: flag-change and unflagged-write tests\n- [ ] **T5 (P2, human: ~3h / CC: ~15min)** \u2014 metrics \u2014 Emit Metric Schema (lookup result, latency by outcome, invalidate reason, fill_cancelled, evictions and entries by kind, oversize, fallback gauge, reset)\n - Surfaced by: Sections 3, 7, 8 \u2014 F3, F7, F9, F8, F1\n - Files: wrapper module, adapter hooks, dashboard config\n - Verify: label tests; dashboard panels present\n- [ ] **T6 (P2, human: ~2h / CC: ~10min)** \u2014 ops \u2014 Alert on `bypassed=1` for 5 min; runbook: flag off then on to reinitialize; include deploy ordering\n - Surfaced by: Sections 8, 9 \u2014 F8, F11\n - Files: alert rules, runbook doc\n - Verify: alert rule test; runbook reviewed by owner\n- [ ] **T7 (P2, human: ~1h / CC: ~5min)** \u2014 adapter \u2014 Reject oversize value without evicting others, counter\n - Surfaced by: Section 2/7 \u2014 F6\n - Files: LRU adapter, tests\n - Verify: oversize test\n- [ ] **T8 (P2, human: ~30min / CC: ~5min)** \u2014 rollout \u2014 Deploy with flag off; single-process confirmation step before 10%\n - Surfaced by: Section 9 \u2014 F11\n - Files: rollout checklist\n - Verify: checklist item present\n- [ ] **T9 (P2, human: ~30min / CC: ~5min)** \u2014 docs \u2014 Add async schedule ASCII diagram to wrapper module header; audit existing diagrams in touched files\n - Surfaced by: Section 10 \u2014 F12; Stale diagram audit\n - Files: wrapper module header\n - Verify: diagram matches Section 4 schedule\n\nTasks JSONL artifact: not written (no shell in this run). Re-run `/plan-ceo-review` with a shell to emit it for `/autoplan`.\n\n## Completion Summary\n```\n +====================================================================+\n | MEGA PLAN REVIEW \u2014 COMPLETION SUMMARY |\n +====================================================================+\n | Mode selected | HOLD |\n | System Audit | skipped by instruction |\n | Step 0 | HOLD; approach B (invalidation protocol) |\n | Section 1 (Arch) | 1 issue found (F1 CRITICAL) |\n | Section 2 (Errors) | 10 error paths mapped, 0 GAPS after F2, F6 |\n | Section 3 (Security)| 1 issue found, 0 High severity |\n | Section 4 (Data/UX) | 4 schedules + 5 shadow paths, 0 unhandled |\n | Section 5 (Quality) | 1 issue found (F4) |\n | Section 6 (Tests) | Diagram produced, 5 gaps (mapped, remedied) |\n | Section 7 (Perf) | 2 issues found (F6, F7) |\n | Section 8 (Observ) | 2 gaps found (F8, F9) |\n | Section 9 (Deploy) | 2 risks flagged (F10, F11) |\n | Section 10 (Future) | Reversibility: 5/5, debt items: 1 (F12) |\n | Section 11 (Design) | SKIPPED (no UI scope) |\n +--------------------------------------------------------------------+\n | NOT in scope | written (5 items) |\n | What already exists | written |\n | Dream state delta | written |\n | Error/rescue registry| 10 rows, 0 CRITICAL GAPS |\n | Failure modes | 9 total, 0 CRITICAL GAPS |\n | TODOS.md updates | 0 items proposed |\n | Scope proposals | 0 proposed, 0 accepted (HOLD) |\n | CEO plan | skipped (HOLD) |\n | Outside voice | codex, disabled (config) |\n | Lake Score | 12/12 recommendations chose complete option |\n | Diagrams produced | 8 (arch, data flow, state, schedule, error, |\n | | deploy, rollback, test) |\n | Stale diagrams found | 0 enumerated (exploration skipped; T9) |\n | Unresolved decisions | 0 |\n +====================================================================+\n```\nReview log, decision log, learnings log, handoff cleanup, telemetry: not executed (no shell tool; mutating commands prohibited). No durable learnings this session.\n\n## GSTACK REVIEW REPORT\n\n| Review | Trigger | Why | Runs | Status | Findings |\n|--------|---------|-----|------|--------|----------|\n| CEO Review | `/plan-ceo-review` | Scope & strategy | 1 | clean | mode: HOLD_SCOPE, 0 critical gaps (11 findings remedied) |\n| Outside Review | codex via `/plan-ceo-review` | Independent 2nd opinion | 0 | disabled | codex_reviews disabled; no fallback run |\n| Eng Review | `/plan-eng-review` | Architecture & tests (required) | 0 | \u2014 | \u2014 |\n| Design Review | `/plan-design-review` | UI/UX gaps | 0 | \u2014 | no UI scope |\n| DX Review | `/plan-devex-review` | Developer experience gaps | 0 | \u2014 | \u2014 |\n\n- **OUTSIDE COVERAGE:** provider codex, phase plan-review, status disabled (config `codex_reviews: disabled`); no native fallback dispatched; record not persisted (no shell).\n- **VERDICT:** CEO CLEARED \u2014 eng review required\n\nNO UNRESOLVED DECISIONS\n",
|
|
"finding": "| F1 | 1, 4 | CRITICAL | \"no additional version checks or coordination between a cache fill and a write\" vs contract \"every read begun after that write completes must observe the committed version\". Schedule S1: R1 misses, W commits and deletes, R1 fills stale v1, R2 hits v1. S2: R2 joins R1's in-flight single-flight read after commit. | `invalidate()` cancels pending fills, detaches single-flight slot, deletes entry; runs before `writeProfile` resolves. Counter `fill_cancelled_total`. | Contract boundary defined as caller-visible completion; owner may confirm | Tests S1\u2013S4 with pause/release |",
|
|
"heading": "### 4. Async schedule (F1)",
|
|
"trace": "step | R1 read (began before commit) | W write | R2 read (began after W resolves) | cache / pending\n1 | get\u2192miss; token t1 | | | \u2205 / {t1}\n2 | await singleFlight(repo.read) | | |\n3 | | await repo.write \u2192 v2 committed | |\n4 | | invalidate: cancel t1, detach, delete | | \u2205 / {t1\u2717}\n4' | | writeProfile resolves (write \"complete\") | |\n5 | | | get\u2192miss; token t2; fresh DB read | \u2205 / {t1\u2717,t2}\n6 | read resolves v1; t1\u2717 \u2192 no fill | | | \u2205 / {t2}\n7 | | | resolves v2; fill; return v2 | v2 / \u2205\nAlternate order (6 before 4): R1 fills v1, then step 4 deletes it; R2 misses and reads v2. Safe.\nOriginal sketch: step 6 fills v1 after step 4 \u2192 R2 hits v1 \u2192 VIOLATION (S1).\nWithout detach: R2 at step 5 joins R1's promise \u2192 v1 \u2192 VIOLATION (S2).\nS3: W commits before R1 issues read \u2192 R1 reads v2, fill v2 (t1 not cancelled since invalidate preceded track): correct.\nS4: indeterminate write error after commit \u2192 invalidate on error path \u2192 R2 reads v2 (F2)."
|
|
}
|
|
}
|