{ "source": "b176520c966d347fb8da05f631da3a9bb9fca3be", "originalFirstAttempt": "FAIL", "configuredRetry": "CANCELLED after demonstrated driver defect; no pass credit", "frames": [ { "fingerprint": { "signature": "8c2809a2-42f4-479b-a5e4-a6f2e5016e61:toolu_01AZB1veZuBaHDZ76jTVXHGX", "promptSnippet": "Routing D1 \u2014 Add gstack skill routing rules to CLAUDE.md? Project/branch/task: main branch of the plan-review fixture repo; one-time gstack onboarding step before the plan review. ELI10: gstack works best when your project's CLAUDE.md inclu", "options": [ { "index": 1, "label": "Add routing rules (recommended)" }, { "index": 2, "label": "No thanks, invoke manually" } ], "observedAtMs": 38146, "preReview": true, "nativeCall": { "sessionId": "8c2809a2-42f4-479b-a5e4-a6f2e5016e61", "toolUseId": "toolu_01AZB1veZuBaHDZ76jTVXHGX", "questions": [ { "question": "D1 \u2014 Add gstack skill routing rules to CLAUDE.md?\nProject/branch/task: main branch of the plan-review fixture repo; one-time gstack onboarding step before the plan review.\nELI10: gstack works best when your project's CLAUDE.md includes skill routing rules. These tell Claude which /skill to reach for when you say things like \"review the architecture\" or \"ship this\", so you get the right workflow without typing slash commands. Without them, skills only run when you invoke them by name.\nStakes if we pick wrong: minor either way; you either get automatic routing or keep invoking skills manually.\nRecommendation: A because routing rules make every future session cheaper to steer, and they are a small appended section you can delete anytime.\nNote: options differ in kind, not coverage \u2014 no completeness score.\nNet: convenience of auto-routing vs keeping CLAUDE.md untouched. (Note: plan mode is active, so the CLAUDE.md write and commit would happen after plan mode ends.)", "header": "Routing", "multiSelect": false, "options": [ { "label": "Add routing rules (recommended)", "description": "\u2705 Natural-language requests auto-route to the matching gstack skill in future sessions\n\u2705 Small appended section, easy to remove or edit later; committed so teammates get it too\n\u274c Adds ~20 lines to CLAUDE.md and a commit on main; write deferred until plan mode exits" }, { "label": "No thanks, invoke manually", "description": "\u2705 CLAUDE.md stays exactly as-is, no extra commit on the branch\n\u2705 You keep full control over when a skill runs by typing /skill-name\n\u274c Skills never fire from natural language; you must remember the slash commands" } ] } ], "answered": true, "failed": false, "answers": { "D1 \u2014 Add gstack skill routing rules to CLAUDE.md?\nProject/branch/task: main branch of the plan-review fixture repo; one-time gstack onboarding step before the plan review.\nELI10: gstack works best when your project's CLAUDE.md includes skill routing rules. These tell Claude which /skill to reach for when you say things like \"review the architecture\" or \"ship this\", so you get the right workflow without typing slash commands. Without them, skills only run when you invoke them by name.\nStakes if we pick wrong: minor either way; you either get automatic routing or keep invoking skills manually.\nRecommendation: A because routing rules make every future session cheaper to steer, and they are a small appended section you can delete anytime.\nNote: options differ in kind, not coverage \u2014 no completeness score.\nNet: convenience of auto-routing vs keeping CLAUDE.md untouched. (Note: plan mode is active, so the CLAUDE.md write and commit would happen after plan mode ends.)": "Add routing rules (recommended)" }, "unansweredQuestionIndices": [], "answeredAt": "2026-09-16T05:53:50.574Z" } }, "questionAt": "2026-09-16T05:53:48.625Z", "answeredAt": "2026-09-16T05:53:50.574Z", "preAskPlan": "", "snapshot": null }, { "fingerprint": { "signature": "8c2809a2-42f4-479b-a5e4-a6f2e5016e61:toolu_01U1oW29UJZAy6kmrvtxY72p", "promptSnippet": "Prereq D2 \u2014 Run /office-hours first, or go straight to the engineering review? Project/branch/task: main branch; reviewing PLAN.md \"Add background job retry framework\". ELI10: No design doc exists for this change. /office-hours is a ~10 min", "options": [ { "index": 1, "label": "Skip, standard review (recommended)" }, { "index": 2, "label": "Run /office-hours now" } ], "observedAtMs": 70198, "preReview": true, "nativeCall": { "sessionId": "8c2809a2-42f4-479b-a5e4-a6f2e5016e61", "toolUseId": "toolu_01U1oW29UJZAy6kmrvtxY72p", "questions": [ { "question": "D2 \u2014 Run /office-hours first, or go straight to the engineering review?\nProject/branch/task: main branch; reviewing PLAN.md \"Add background job retry framework\".\nELI10: No design doc exists for this change. /office-hours is a ~10 minute structured session that produces a problem statement, challenges the premise, and lists alternatives considered. It gives this review sharper input, because right now the plan says what it will build but not why retries are needed, what the failure modes are, or what \"at-most-once\" currently protects.\nStakes if we pick wrong: skipping means I review the plan's mechanics without a stated problem; running it costs ~10 minutes before any review output.\nRecommendation: B because the plan is short and its four sections already expose the key architecture and test risks; I can flag the missing problem statement inside the review instead.\nNote: options differ in kind, not coverage \u2014 no completeness score.\nNet: sharper problem framing now vs faster feedback on a plan whose issues are already visible.", "header": "Prereq", "multiSelect": false, "options": [ { "label": "Skip, standard review (recommended)", "description": "\u2705 Review starts immediately; the plan is small enough to evaluate on its own terms\n\u2705 Missing problem statement gets flagged as a finding rather than blocking progress\n\u274c Alternatives (library retry hooks vs custom) get evaluated without a documented rationale" }, { "label": "Run /office-hours now", "description": "\u2705 Produces a design doc with problem statement and explored alternatives (human: ~1 hr / CC: ~10 min)\n\u2705 Forces the \"why custom scheduler?\" question before the review debates it\n\u274c Adds ~10 minutes before any review findings land" } ] } ], "answered": true, "failed": false, "answers": { "D2 \u2014 Run /office-hours first, or go straight to the engineering review?\nProject/branch/task: main branch; reviewing PLAN.md \"Add background job retry framework\".\nELI10: No design doc exists for this change. /office-hours is a ~10 minute structured session that produces a problem statement, challenges the premise, and lists alternatives considered. It gives this review sharper input, because right now the plan says what it will build but not why retries are needed, what the failure modes are, or what \"at-most-once\" currently protects.\nStakes if we pick wrong: skipping means I review the plan's mechanics without a stated problem; running it costs ~10 minutes before any review output.\nRecommendation: B because the plan is short and its four sections already expose the key architecture and test risks; I can flag the missing problem statement inside the review instead.\nNote: options differ in kind, not coverage \u2014 no completeness score.\nNet: sharper problem framing now vs faster feedback on a plan whose issues are already visible.": "Skip, standard review (recommended)" }, "unansweredQuestionIndices": [], "answeredAt": "2026-09-16T05:54:22.625Z" } }, "questionAt": "2026-09-16T05:54:20.648Z", "answeredAt": "2026-09-16T05:54:22.625Z", "preAskPlan": "", "snapshot": null }, { "fingerprint": { "signature": "8c2809a2-42f4-479b-a5e4-a6f2e5016e61:toolu_016jUdzFxz252h24Qbw1tZNe", "promptSnippet": "Learnings D3 \u2014 Enable cross-project learnings search? Project/branch/task: main branch; one-time gstack setting before loading prior learnings for the PLAN.md review. ELI10: gstack keeps a local log of lessons from past review sessions (pit", "options": [ { "index": 1, "label": "Enable cross-project (recommended)" }, { "index": 2, "label": "Keep project-scoped only" } ], "observedAtMs": 114285, "preReview": true, "nativeCall": { "sessionId": "8c2809a2-42f4-479b-a5e4-a6f2e5016e61", "toolUseId": "toolu_016jUdzFxz252h24Qbw1tZNe", "questions": [ { "question": "D3 \u2014 Enable cross-project learnings search?\nProject/branch/task: main branch; one-time gstack setting before loading prior learnings for the PLAN.md review.\nELI10: gstack keeps a local log of lessons from past review sessions (pitfalls, patterns, project quirks). By default it only searches this project's lessons. Cross-project mode also searches lessons from your other repos on this machine, so a retry-scheduler pitfall you hit elsewhere would surface here. Nothing leaves your machine either way.\nStakes if we pick wrong: enabling on a machine with multiple client codebases could mix one client's context into another's review; keeping it scoped just means fewer hints.\nRecommendation: A because you appear to be a solo developer on this machine and more local hints make reviews sharper at zero privacy cost.\nNote: options differ in kind, not coverage \u2014 no completeness score.\nNet: broader local recall vs strict per-project isolation.", "header": "Learnings", "multiSelect": false, "options": [ { "label": "Enable cross-project (recommended)", "description": "\u2705 Pitfalls and patterns from your other repos surface during this and future reviews\n\u2705 Stays fully local; no data leaves the machine, and you can flip it off with one config command\n\u274c If you consult for multiple clients, one client's project quirks could appear in another's review" }, { "label": "Keep project-scoped only", "description": "\u2705 Strict isolation: each repo's learnings never appear outside that repo\n\u2705 Safe default for multi-client or shared-machine setups\n\u274c This brand-new project has zero learnings, so the search returns nothing today" } ] } ], "answered": true, "failed": false, "answers": { "D3 \u2014 Enable cross-project learnings search?\nProject/branch/task: main branch; one-time gstack setting before loading prior learnings for the PLAN.md review.\nELI10: gstack keeps a local log of lessons from past review sessions (pitfalls, patterns, project quirks). By default it only searches this project's lessons. Cross-project mode also searches lessons from your other repos on this machine, so a retry-scheduler pitfall you hit elsewhere would surface here. Nothing leaves your machine either way.\nStakes if we pick wrong: enabling on a machine with multiple client codebases could mix one client's context into another's review; keeping it scoped just means fewer hints.\nRecommendation: A because you appear to be a solo developer on this machine and more local hints make reviews sharper at zero privacy cost.\nNote: options differ in kind, not coverage \u2014 no completeness score.\nNet: broader local recall vs strict per-project isolation.": "Enable cross-project (recommended)" }, "unansweredQuestionIndices": [], "answeredAt": "2026-09-16T05:55:06.711Z" } }, "questionAt": "2026-09-16T05:55:04.959Z", "answeredAt": "2026-09-16T05:55:06.711Z", "preAskPlan": "", "snapshot": null }, { "fingerprint": { "signature": "8c2809a2-42f4-479b-a5e4-a6f2e5016e61:toolu_012MvjGqGe6XZZcBmGbujY7p", "promptSnippet": "Retry mech D4 \u2014 Retry mechanism: use the job library's built-in retry hooks, or roll a custom inline scheduler? Project/branch/task: main branch; PLAN.md \"Add background job retry framework\", Architecture section. ELI10: The plan rebuilds r", "options": [ { "index": 1, "label": "Library retry hooks + one custom backoff function (recommended)" }, { "index": 2, "label": "Custom inline scheduler in each worker (as proposed)" } ], "observedAtMs": 250132, "preReview": true, "nativeCall": { "sessionId": "8c2809a2-42f4-479b-a5e4-a6f2e5016e61", "toolUseId": "toolu_012MvjGqGe6XZZcBmGbujY7p", "questions": [ { "question": "D4 \u2014 Retry mechanism: use the job library's built-in retry hooks, or roll a custom inline scheduler?\nProject/branch/task: main branch; PLAN.md \"Add background job retry framework\", Architecture section.\nELI10: The plan rebuilds retry scheduling inside each worker even though the job library already has retry hooks. The stated reason is \"full control over the curve\", but these libraries let you plug in your own backoff function, so you get the exact curve you want while the library keeps doing the hard parts: persisting attempt counts across worker crashes, scheduling the delayed re-run, and exposing retries to its dashboard and metrics. A hand-rolled inline scheduler has to re-solve all of that in five places.\nStakes if we pick wrong: a worker that crashes mid-sleep in a custom scheduler loses the job silently; a library-managed retry survives the crash and is visible in the queue UI.\nRecommendation: A because the library hook gives the same curve control with crash-safe attempt state and zero new scheduling code; the custom path only wins if the library has no backoff hook at all, which the plan does not claim.\nCompleteness: A=10/10, B=6/10\nNet: same curve either way; the trade is crash-safety and observability for free vs. rebuilding the library's job.", "header": "Retry mech", "multiSelect": false, "options": [ { "label": "Library retry hooks + one custom backoff function (recommended)", "description": "\u2705 Attempt count and next-run time persist in the queue, so a worker crash mid-backoff cannot lose the job\n\u2705 Retries show up in the library's dashboard, metrics and dead-letter handling for free\n\u2705 One backoff function to test instead of five inline schedulers (human: ~1 day / CC: ~20 min)\n\u274c Curve is bounded by whatever the hook API allows (delay per attempt); exotic behavior like per-tenant curves needs a wrapper" }, { "label": "Custom inline scheduler in each worker (as proposed)", "description": "\u2705 Unconstrained control over every retry decision inside worker code\n\u2705 No dependency on the library's hook signature or its upgrade path\n\u274c Attempt state lives in process memory unless you build persistence; a crash or deploy mid-backoff drops the retry\n\u274c Five schedulers to keep in sync, invisible to the queue UI and metrics (human: ~3 days / CC: ~45 min)" } ] } ], "answered": true, "failed": false, "answers": { "D4 \u2014 Retry mechanism: use the job library's built-in retry hooks, or roll a custom inline scheduler?\nProject/branch/task: main branch; PLAN.md \"Add background job retry framework\", Architecture section.\nELI10: The plan rebuilds retry scheduling inside each worker even though the job library already has retry hooks. The stated reason is \"full control over the curve\", but these libraries let you plug in your own backoff function, so you get the exact curve you want while the library keeps doing the hard parts: persisting attempt counts across worker crashes, scheduling the delayed re-run, and exposing retries to its dashboard and metrics. A hand-rolled inline scheduler has to re-solve all of that in five places.\nStakes if we pick wrong: a worker that crashes mid-sleep in a custom scheduler loses the job silently; a library-managed retry survives the crash and is visible in the queue UI.\nRecommendation: A because the library hook gives the same curve control with crash-safe attempt state and zero new scheduling code; the custom path only wins if the library has no backoff hook at all, which the plan does not claim.\nCompleteness: A=10/10, B=6/10\nNet: same curve either way; the trade is crash-safety and observability for free vs. rebuilding the library's job.": "Library retry hooks + one custom backoff function (recommended)" }, "unansweredQuestionIndices": [], "answeredAt": "2026-09-16T05:57:22.558Z" } }, "questionAt": "2026-09-16T05:57:21.694Z", "answeredAt": "2026-09-16T05:57:22.558Z", "preAskPlan": "# Eng Review: Add background job retry framework\n\nReviewed target: `PLAN.md` (repo root, commit 7ce6818, branch main).\nReview: `/plan-eng-review`, 2026-09-16. Report file requested by the user.\n\n## Original plan (as reviewed)\n\n> ## Architecture\n> We'll roll a custom exponential-backoff scheduler inline in each worker\n> rather than use the existing job library's built-in retry hooks. Same\n> shape as the library version, but we want full control over the curve.\n>\n> ## Code quality\n> The retry envelope (compute delay, log attempt, dispatch) is duplicated\n> across 5 worker files with copy-pasted bodies. We will leave the\n> duplication for now and refactor \"later.\"\n>\n> ## Tests\n> The existing `processWebhookJob()` flow gets rewritten as part of this\n> change. No regression test for the prior at-most-once delivery guarantee\n> is planned.\n>\n> ## Performance\n> On every retry we re-fetch the full job payload from the database, then\n> iterate the payload to recompute the dependency graph. Could cache the\n> graph on the first attempt; not planned.\n\n## Step 0: Scope Challenge findings\n\n1. [P1] (confidence: 8/10) PLAN.md:6-8 \u2014 custom inline scheduler rebuilds the job library's built-in retry hooks (Layer 1). Custom backoff curves are a supported extension point. Disposition: pending (R1 / D4).\n2. [P1] (confidence: 9/10) PLAN.md:11-13 \u2014 retry envelope copy-pasted across 5 worker files; refactor deferred. Disposition: pending (Code Quality).\n3. [P0] (confidence: 9/10) PLAN.md:16-18 \u2014 `processWebhookJob()` rewritten with no regression test; retries change the guarantee from at-most-once to at-least-once with no idempotency key. Disposition: pending (Architecture + Tests).\n4. [P2] (confidence: 7/10) PLAN.md:21-23 \u2014 payload re-fetched and dependency graph recomputed on every retry. Disposition: pending (Performance).\n5. [P1] (confidence: 8/10) PLAN.md:6-8 \u2014 \"full control over the curve\" but no curve specified: no max attempts, delay cap, jitter, dead-letter, or retryable-error classification. Disposition: pending (Architecture).\n6. Factual gap: no design doc / problem statement (which failure classes retries target). Recorded; folded into finding 5.\n\nComplexity gate: not triggered (5 worker files + 1 rewrite, 0 new classes). TODOS.md: absent. Distribution: no new artifact.\n\n## Decision ledger\n\n### R1: Retry mechanism \u2014 library built-in retry hooks vs custom inline scheduler\nFinding: #1, P1, confidence 8/10, PLAN.md:6-8, reviewer: Claude (plan-eng-review)\nPlan baseline: original proposal \u2014 custom exponential-backoff scheduler inline in each worker; library hooks unused.\nRuntime evidence: unknown \u2014 the repo under review contains only PLAN.md and CLAUDE.md; the job library and worker files are not present. Web search (2026) shows mainstream job libraries expose a custom backoff function / retry_in hook.\nState: pending\nComparison grid:\n\n| Choice | Current | A | B |\n|---|---|---|---|\n| R1 retry mechanism | custom inline scheduler per worker (proposed) | library built-in retry hooks with one shared custom backoff function | custom inline scheduler per worker (as proposed) |\n| R2 envelope duplication | 5 copies, deferred (pending) | pending | pending |\n| R3 webhook delivery semantics | unspecified (pending) | pending | pending |\n| R4 retry policy bounds | unspecified (pending) | pending | pending |\n| R5 webhook regression contract | none planned (pending) | pending | pending |\n| R7 graph recompute per retry | recompute every retry (pending) | pending | pending |\n\nQuestion D4:\nD4 \u2014 Retry mechanism: use the job library's built-in retry hooks, or roll a custom inline scheduler?\nProject/branch/task: main branch; PLAN.md \"Add background job retry framework\", Architecture section.\nELI10: The plan rebuilds retry scheduling inside each worker even though the job library already has retry hooks. The stated reason is \"full control over the curve\", but these libraries let you plug in your own backoff function, so you get the exact curve you want while the library keeps doing the hard parts: persisting attempt counts across worker crashes, scheduling the delayed re-run, and exposing retries to its dashboard and metrics. A hand-rolled inline scheduler has to re-solve all of that in five places.\nStakes if we pick wrong: a worker that crashes mid-sleep in a custom scheduler loses the job silently; a library-managed retry survives the crash and is visible in the queue UI.\nRecommendation: A because the library hook gives the same curve control with crash-safe attempt state and zero new scheduling code; the custom path only wins if the library has no backoff hook at all, which the plan does not claim.\nCompleteness: A=10/10, B=6/10\nPros / cons:\nA) Library retry hooks + one custom backoff function (recommended)\n \u2705 Attempt count and next-run time persist in the queue, so a worker crash mid-backoff cannot lose the job\n \u2705 Retries show up in the library's dashboard, metrics and dead-letter handling for free\n \u2705 One backoff function to test instead of five inline schedulers (human: ~1 day / CC: ~20 min)\n \u274c Curve is bounded by whatever the hook API allows (delay per attempt); exotic behavior like per-tenant curves needs a wrapper\nB) Custom inline scheduler in each worker (as proposed)\n \u2705 Unconstrained control over every retry decision inside worker code\n \u2705 No dependency on the library's hook signature or its upgrade path\n \u274c Attempt state lives in process memory unless you build persistence; a crash or deploy mid-backoff drops the retry\n \u274c Five schedulers to keep in sync, invisible to the queue UI and metrics (human: ~3 days / CC: ~45 min)\nNet: same curve either way; the trade is crash-safety and observability for free vs. rebuilding the library's job.\nHeader: Retry mech\nOptions:\nA) Library retry hooks + one custom backoff function (recommended)\n\u2705 Attempt count and next-run time persist in the queue, so a worker crash mid-backoff cannot lose the job\n\u2705 Retries show up in the library's dashboard, metrics and dead-letter handling for free\n\u2705 One backoff function to test instead of five inline schedulers (human: ~1 day / CC: ~20 min)\n\u274c Curve is bounded by whatever the hook API allows (delay per attempt); exotic behavior like per-tenant curves needs a wrapper\nB) Custom inline scheduler in each worker (as proposed)\n\u2705 Unconstrained control over every retry decision inside worker code\n\u2705 No dependency on the library's hook signature or its upgrade path\n\u274c Attempt state lives in process memory unless you build persistence; a crash or deploy mid-backoff drops the retry\n\u274c Five schedulers to keep in sync, invisible to the queue UI and metrics (human: ~3 days / CC: ~45 min)\nActual answer: unanswered\nAccepted scope: none\nHistory: none\n", "snapshot": { "at": "2026-09-16T05:57:07.228Z", "kind": "owned-plan-or-report", "source": "/tmp/g-pyj7vm4y/gstack-paid-shard-qDJLTL/tmp/gstack-e2e-plan-eng-batching-IpzCNz/gstack-test-plan-eng-batching.md", "artifact": "objects/96b1bb4f2a1ba9b5d88118b6639dcb196335b5789fc26319cf8c5f86ef06befa.md", "sha256": "96b1bb4f2a1ba9b5d88118b6639dcb196335b5789fc26319cf8c5f86ef06befa", "bytes": 6751, "mtimeMs": 1789538227044.3108, "provenance": "Exact observed file bytes; never reconstructed from tool text." } }, { "fingerprint": { "signature": "8c2809a2-42f4-479b-a5e4-a6f2e5016e61:toolu_01TYqesfX2vGSEQMM3KxrJ6q", "promptSnippet": "Webhook sem D5 \u2014 When processWebhookJob() gains retries, what delivery guarantee does it make? Project/branch/task: main branch; PLAN.md \"Add background job retry framework\", Tests/Architecture sections. ELI10: Today the webhook job runs at", "options": [ { "index": 1, "label": "At-least-once + idempotency key (recommended)" }, { "index": 2, "label": "Keep at-most-once: no retries for the webhook job" }, { "index": 3, "label": "Retry with no idempotency key (as proposed)" } ], "observedAtMs": 307813, "preReview": true, "nativeCall": { "sessionId": "8c2809a2-42f4-479b-a5e4-a6f2e5016e61", "toolUseId": "toolu_01TYqesfX2vGSEQMM3KxrJ6q", "questions": [ { "question": "D5 \u2014 When processWebhookJob() gains retries, what delivery guarantee does it make?\nProject/branch/task: main branch; PLAN.md \"Add background job retry framework\", Tests/Architecture sections.\nELI10: Today the webhook job runs at most once: if it fails, nothing is redelivered, so it can never double-fire. The moment you add retries, a job that timed out after the receiver already processed it will run again, and the receiver sees the same webhook twice. That is a real behavior change, not just a missing test. The standard fix is an idempotency key: record the event ID before doing the side effect, and if the key is already there, treat the retry as already done.\nStakes if we pick wrong: duplicate webhooks mean double charges, double emails, or double state changes for whoever receives them, and you find out from a customer, not a test.\nRecommendation: A because retrying is the point of this plan, and an idempotency key is a few lines plus one unique index; it turns duplicates from an incident into a no-op.\nCompleteness: A=10/10, B=7/10, C=3/10\nNet: you are choosing between fixing lost webhooks safely, not fixing them, or fixing them in a way that creates duplicates.", "header": "Webhook sem", "multiSelect": false, "options": [ { "label": "At-least-once + idempotency key (recommended)", "description": "\u2705 Transient failures get retried and duplicates are absorbed, so receivers never see a double effect\n\u2705 Small change: atomic INSERT ... ON CONFLICT DO NOTHING (or SET NX) on the event ID before dispatch (human: ~1 day / CC: ~15 min)\n\u274c Needs a dedup store with a TTL at least as long as the retry window, and a decision on what the key is when the payload has no natural ID" }, { "label": "Keep at-most-once: no retries for the webhook job", "description": "\u2705 Zero semantic change for webhook receivers; regression risk drops to the rewrite itself\n\u2705 Simplest to prove: max attempts = 1 for this job, existing behavior stays\n\u274c Webhook deliveries that fail on a transient blip are still lost, which is the pain this plan exists to fix" }, { "label": "Retry with no idempotency key (as proposed)", "description": "\u2705 No extra storage or key design work up front\n\u274c Duplicate side effects on every timeout-after-success; the at-most-once guarantee silently disappears" } ] } ], "answered": true, "failed": false, "answers": { "D5 \u2014 When processWebhookJob() gains retries, what delivery guarantee does it make?\nProject/branch/task: main branch; PLAN.md \"Add background job retry framework\", Tests/Architecture sections.\nELI10: Today the webhook job runs at most once: if it fails, nothing is redelivered, so it can never double-fire. The moment you add retries, a job that timed out after the receiver already processed it will run again, and the receiver sees the same webhook twice. That is a real behavior change, not just a missing test. The standard fix is an idempotency key: record the event ID before doing the side effect, and if the key is already there, treat the retry as already done.\nStakes if we pick wrong: duplicate webhooks mean double charges, double emails, or double state changes for whoever receives them, and you find out from a customer, not a test.\nRecommendation: A because retrying is the point of this plan, and an idempotency key is a few lines plus one unique index; it turns duplicates from an incident into a no-op.\nCompleteness: A=10/10, B=7/10, C=3/10\nNet: you are choosing between fixing lost webhooks safely, not fixing them, or fixing them in a way that creates duplicates.": "At-least-once + idempotency key (recommended)" }, "unansweredQuestionIndices": [], "answeredAt": "2026-09-16T05:58:20.235Z" } }, "questionAt": "2026-09-16T05:58:19.602Z", "answeredAt": "2026-09-16T05:58:20.235Z", "preAskPlan": "# Eng Review: Add background job retry framework\n\nReviewed target: `PLAN.md` (repo root, commit 7ce6818, branch main).\nReview: `/plan-eng-review`, 2026-09-16. Report file requested by the user.\n\n## Original plan (as reviewed)\n\n> ## Architecture\n> We'll roll a custom exponential-backoff scheduler inline in each worker\n> rather than use the existing job library's built-in retry hooks. Same\n> shape as the library version, but we want full control over the curve.\n>\n> ## Code quality\n> The retry envelope (compute delay, log attempt, dispatch) is duplicated\n> across 5 worker files with copy-pasted bodies. We will leave the\n> duplication for now and refactor \"later.\"\n>\n> ## Tests\n> The existing `processWebhookJob()` flow gets rewritten as part of this\n> change. No regression test for the prior at-most-once delivery guarantee\n> is planned.\n>\n> ## Performance\n> On every retry we re-fetch the full job payload from the database, then\n> iterate the payload to recompute the dependency graph. Could cache the\n> graph on the first attempt; not planned.\n\n## Step 0: Scope Challenge findings\n\n1. [P1] (confidence: 8/10) PLAN.md:6-8 \u2014 custom inline scheduler rebuilds the job library's built-in retry hooks (Layer 1). Custom backoff curves are a supported extension point. Disposition: pending (R1 / D4).\n2. [P1] (confidence: 9/10) PLAN.md:11-13 \u2014 retry envelope copy-pasted across 5 worker files; refactor deferred. Disposition: pending (Code Quality).\n3. [P0] (confidence: 9/10) PLAN.md:16-18 \u2014 `processWebhookJob()` rewritten with no regression test; retries change the guarantee from at-most-once to at-least-once with no idempotency key. Disposition: pending (Architecture + Tests).\n4. [P2] (confidence: 7/10) PLAN.md:21-23 \u2014 payload re-fetched and dependency graph recomputed on every retry. Disposition: pending (Performance).\n5. [P1] (confidence: 8/10) PLAN.md:6-8 \u2014 \"full control over the curve\" but no curve specified: no max attempts, delay cap, jitter, dead-letter, or retryable-error classification. Disposition: pending (Architecture).\n6. Factual gap: no design doc / problem statement (which failure classes retries target). Recorded; folded into finding 5.\n\nComplexity gate: not triggered (5 worker files + 1 rewrite, 0 new classes). TODOS.md: absent. Distribution: no new artifact.\n\n## Decision ledger\n\n### R1: Retry mechanism \u2014 library built-in retry hooks vs custom inline scheduler\nFinding: #1, P1, confidence 8/10, PLAN.md:6-8, reviewer: Claude (plan-eng-review)\nPlan baseline: original proposal \u2014 custom exponential-backoff scheduler inline in each worker; library hooks unused.\nRuntime evidence: unknown \u2014 the repo under review contains only PLAN.md and CLAUDE.md; the job library and worker files are not present. Web search (2026) shows mainstream job libraries expose a custom backoff function / retry_in hook.\nState: pending\nComparison grid:\n\n| Choice | Current | A | B |\n|---|---|---|---|\n| R1 retry mechanism | custom inline scheduler per worker (proposed) | library built-in retry hooks with one shared custom backoff function | custom inline scheduler per worker (as proposed) |\n| R2 envelope duplication | 5 copies, deferred (pending) | pending | pending |\n| R3 webhook delivery semantics | unspecified (pending) | pending | pending |\n| R4 retry policy bounds | unspecified (pending) | pending | pending |\n| R5 webhook regression contract | none planned (pending) | pending | pending |\n| R7 graph recompute per retry | recompute every retry (pending) | pending | pending |\n\nQuestion D4:\nD4 \u2014 Retry mechanism: use the job library's built-in retry hooks, or roll a custom inline scheduler?\nProject/branch/task: main branch; PLAN.md \"Add background job retry framework\", Architecture section.\nELI10: The plan rebuilds retry scheduling inside each worker even though the job library already has retry hooks. The stated reason is \"full control over the curve\", but these libraries let you plug in your own backoff function, so you get the exact curve you want while the library keeps doing the hard parts: persisting attempt counts across worker crashes, scheduling the delayed re-run, and exposing retries to its dashboard and metrics. A hand-rolled inline scheduler has to re-solve all of that in five places.\nStakes if we pick wrong: a worker that crashes mid-sleep in a custom scheduler loses the job silently; a library-managed retry survives the crash and is visible in the queue UI.\nRecommendation: A because the library hook gives the same curve control with crash-safe attempt state and zero new scheduling code; the custom path only wins if the library has no backoff hook at all, which the plan does not claim.\nCompleteness: A=10/10, B=6/10\nPros / cons:\nA) Library retry hooks + one custom backoff function (recommended)\n \u2705 Attempt count and next-run time persist in the queue, so a worker crash mid-backoff cannot lose the job\n \u2705 Retries show up in the library's dashboard, metrics and dead-letter handling for free\n \u2705 One backoff function to test instead of five inline schedulers (human: ~1 day / CC: ~20 min)\n \u274c Curve is bounded by whatever the hook API allows (delay per attempt); exotic behavior like per-tenant curves needs a wrapper\nB) Custom inline scheduler in each worker (as proposed)\n \u2705 Unconstrained control over every retry decision inside worker code\n \u2705 No dependency on the library's hook signature or its upgrade path\n \u274c Attempt state lives in process memory unless you build persistence; a crash or deploy mid-backoff drops the retry\n \u274c Five schedulers to keep in sync, invisible to the queue UI and metrics (human: ~3 days / CC: ~45 min)\nNet: same curve either way; the trade is crash-safety and observability for free vs. rebuilding the library's job.\nHeader: Retry mech\nOptions:\nA) Library retry hooks + one custom backoff function (recommended)\n\u2705 Attempt count and next-run time persist in the queue, so a worker crash mid-backoff cannot lose the job\n\u2705 Retries show up in the library's dashboard, metrics and dead-letter handling for free\n\u2705 One backoff function to test instead of five inline schedulers (human: ~1 day / CC: ~20 min)\n\u274c Curve is bounded by whatever the hook API allows (delay per attempt); exotic behavior like per-tenant curves needs a wrapper\nB) Custom inline scheduler in each worker (as proposed)\n\u2705 Unconstrained control over every retry decision inside worker code\n\u2705 No dependency on the library's hook signature or its upgrade path\n\u274c Attempt state lives in process memory unless you build persistence; a crash or deploy mid-backoff drops the retry\n\u274c Five schedulers to keep in sync, invisible to the queue UI and metrics (human: ~3 days / CC: ~45 min)\nActual answer: A) Library retry hooks + one custom backoff function \u2014 user answer to D4\nAccepted scope: Replace the custom inline scheduler with the job library's built-in retry hooks; implement one shared custom backoff function that the hooks call; delete the plan's inline per-worker scheduler. Scope Challenge result: scope reduced per recommendation.\nHistory: original proposal (custom inline scheduler) superseded by D4.\nState: approved\n\n### R3: Delivery semantics for `processWebhookJob()` under retries\nFinding: #3, P0, confidence 9/10, PLAN.md:16-18, reviewer: Claude (plan-eng-review)\nPlan baseline: original proposal \u2014 processWebhookJob() is rewritten to retry; the prior at-most-once guarantee is neither preserved nor replaced with an explicit alternative.\nRuntime evidence: unknown \u2014 worker source not present in the reviewed repo. The plan itself states the current guarantee is at-most-once (PLAN.md:17). Web sources (Svix, Hookdeck, 2026) agree that retries imply at-least-once and require an idempotency key to be safe.\nState: pending\nComparison grid:\n\n| Choice | Current | A | B | C |\n|---|---|---|---|---|\n| R1 retry mechanism | approved: library hooks + shared backoff fn (D4) | fixed | fixed | fixed |\n| R3 webhook delivery semantics | at-most-once today; plan retries without stating new guarantee | at-least-once with idempotency key (provider event ID, or hash of raw body) claimed atomically before side effects; duplicates return success | keep at-most-once: exclude processWebhookJob from retries (max attempts = 1); other 4 workers retry | retry with no idempotency key (as proposed) |\n| R2 envelope duplication | pending | pending | pending | pending |\n| R4 retry policy bounds | pending | pending | pending | pending |\n| R5 webhook regression contract | pending | pending | pending | pending |\n| R7 graph recompute per retry | pending | pending | pending | pending |\n\nQuestion D5:\nD5 \u2014 When processWebhookJob() gains retries, what delivery guarantee does it make?\nProject/branch/task: main branch; PLAN.md \"Add background job retry framework\", Tests/Architecture sections.\nELI10: Today the webhook job runs at most once: if it fails, nothing is redelivered, so it can never double-fire. The moment you add retries, a job that timed out after the receiver already processed it will run again, and the receiver sees the same webhook twice. That is a real behavior change, not just a missing test. The standard fix is an idempotency key: record the event ID before doing the side effect, and if the key is already there, treat the retry as already done.\nStakes if we pick wrong: duplicate webhooks mean double charges, double emails, or double state changes for whoever receives them, and you find out from a customer, not a test.\nRecommendation: A because retrying is the point of this plan, and an idempotency key is a few lines plus one unique index; it turns duplicates from an incident into a no-op.\nCompleteness: A=10/10, B=7/10, C=3/10\nPros / cons:\nA) At-least-once + idempotency key (recommended)\n \u2705 Transient failures get retried and duplicates are absorbed, so receivers never see a double effect\n \u2705 Small change: atomic INSERT ... ON CONFLICT DO NOTHING (or SET NX) on the event ID before dispatch (human: ~1 day / CC: ~15 min)\n \u274c Needs a dedup store with a TTL at least as long as the retry window, and a decision on what the key is when the payload has no natural ID\nB) Keep at-most-once: no retries for the webhook job\n \u2705 Zero semantic change for webhook receivers; regression risk drops to the rewrite itself\n \u2705 Simplest to prove: max attempts = 1 for this job, existing behavior stays\n \u274c Webhook deliveries that fail on a transient blip are still lost, which is the pain this plan exists to fix\nC) Retry with no idempotency key (as proposed)\n \u2705 No extra storage or key design work up front\n \u274c Duplicate side effects on every timeout-after-success; the at-most-once guarantee silently disappears\nNet: you are choosing between fixing lost webhooks safely, not fixing them, or fixing them in a way that creates duplicates.\nHeader: Webhook sem\nOptions:\nA) At-least-once + idempotency key (recommended)\n\u2705 Transient failures get retried and duplicates are absorbed, so receivers never see a double effect\n\u2705 Small change: atomic INSERT ... ON CONFLICT DO NOTHING (or SET NX) on the event ID before dispatch (human: ~1 day / CC: ~15 min)\n\u274c Needs a dedup store with a TTL at least as long as the retry window, and a decision on what the key is when the payload has no natural ID\nB) Keep at-most-once: no retries for the webhook job\n\u2705 Zero semantic change for webhook receivers; regression risk drops to the rewrite itself\n\u2705 Simplest to prove: max attempts = 1 for this job, existing behavior stays\n\u274c Webhook deliveries that fail on a transient blip are still lost, which is the pain this plan exists to fix\nC) Retry with no idempotency key (as proposed)\n\u2705 No extra storage or key design work up front\n\u274c Duplicate side effects on every timeout-after-success; the at-most-once guarantee silently disappears\nActual answer: unanswered\nAccepted scope: none\nHistory: none\n", "snapshot": { "at": "2026-09-16T05:57:57.008Z", "kind": "owned-plan-or-report", "source": "/tmp/g-pyj7vm4y/gstack-paid-shard-qDJLTL/tmp/gstack-e2e-plan-eng-batching-IpzCNz/gstack-test-plan-eng-batching.md", "artifact": "objects/822bf8d2cbaa5160c46fde3e8b93ac06053ae3088eca9e0580160ddf23fbedbf.md", "sha256": "822bf8d2cbaa5160c46fde3e8b93ac06053ae3088eca9e0580160ddf23fbedbf", "bytes": 11839, "mtimeMs": 1789538276688.31, "provenance": "Exact observed file bytes; never reconstructed from tool text." } }, { "fingerprint": { "signature": "8c2809a2-42f4-479b-a5e4-a6f2e5016e61:toolu_01JFaHGcCQrx5TzBYHogy15g", "promptSnippet": "Retry policy D6 \u2014 How completely should the plan specify the retry policy (attempts, cap, jitter, exhaustion, error classes)? Project/branch/task: main branch; PLAN.md \"Add background job retry framework\", Architecture section. ELI10: Expon", "options": [ { "index": 1, "label": "Full policy: max 5 attempts, 10-min cap, full jitter, dead-letter + alert on exhaustion, retry only transient errors (recommended)" }, { "index": 2, "label": "Max attempts only (5), everything else unspecified" }, { "index": 3, "label": "Curve only, as proposed" } ], "observedAtMs": 386578, "preReview": true, "nativeCall": { "sessionId": "8c2809a2-42f4-479b-a5e4-a6f2e5016e61", "toolUseId": "toolu_01JFaHGcCQrx5TzBYHogy15g", "questions": [ { "question": "D6 \u2014 How completely should the plan specify the retry policy (attempts, cap, jitter, exhaustion, error classes)?\nProject/branch/task: main branch; PLAN.md \"Add background job retry framework\", Architecture section.\nELI10: Exponential backoff is only half a policy. Without a cap, attempt 20 waits about 12 days. Without jitter, every job that failed at the same moment retries at the same moment and hammers the downstream again (a retry storm). Without a max, a permanently broken payload retries forever. And retrying a \"400 invalid payload\" is pointless: only transient failures (timeouts, 5xx, dropped connections) deserve a retry. The plan says \"full control over the curve\" but writes none of these down, so each of the five workers will pick its own answer.\nStakes if we pick wrong: a poisoned job retrying forever, or a downstream outage made worse by synchronized retries; both look like mystery load at 3am.\nRecommendation: A because every one of these is a one-line constant or a small predicate in the single shared backoff function you already approved, and the defaults below are the standard ones (adjust the numbers freely via Other).\nCompleteness: A=10/10, B=6/10, C=3/10\nNet: the difference between A and C is about 15 lines in one shared function versus five different implicit policies you discover in production.", "header": "Retry policy", "multiSelect": false, "options": [ { "label": "Full policy: max 5 attempts, 10-min cap, full jitter, dead-letter + alert on exhaustion, retry only transient errors (recommended)", "description": "\u2705 Bounded worst case: a broken job costs at most 5 attempts and ~20 minutes, then lands in a dead-letter set someone gets paged for\n\u2705 Full jitter spreads synchronized failures so a downstream blip does not become a retry storm\n\u2705 Terminal errors fail fast, so bad payloads surface immediately instead of after 5 retries (human: ~1 day / CC: ~15 min)\n\u274c Five constants and one error predicate to agree on now; per-job overrides add a little config surface" }, { "label": "Max attempts only (5), everything else unspecified", "description": "\u2705 Prevents infinite retries with one constant\n\u2705 Smallest possible addition to the plan text\n\u274c Uncapped delay and no jitter still allow retry storms and multi-hour waits within those 5 attempts; terminal errors waste retries" }, { "label": "Curve only, as proposed", "description": "\u2705 No policy decisions needed before coding starts\n\u274c Each worker author invents their own bounds (or none); unbounded delay and infinite retries are possible" } ] } ], "answered": true, "failed": false, "answers": { "D6 \u2014 How completely should the plan specify the retry policy (attempts, cap, jitter, exhaustion, error classes)?\nProject/branch/task: main branch; PLAN.md \"Add background job retry framework\", Architecture section.\nELI10: Exponential backoff is only half a policy. Without a cap, attempt 20 waits about 12 days. Without jitter, every job that failed at the same moment retries at the same moment and hammers the downstream again (a retry storm). Without a max, a permanently broken payload retries forever. And retrying a \"400 invalid payload\" is pointless: only transient failures (timeouts, 5xx, dropped connections) deserve a retry. The plan says \"full control over the curve\" but writes none of these down, so each of the five workers will pick its own answer.\nStakes if we pick wrong: a poisoned job retrying forever, or a downstream outage made worse by synchronized retries; both look like mystery load at 3am.\nRecommendation: A because every one of these is a one-line constant or a small predicate in the single shared backoff function you already approved, and the defaults below are the standard ones (adjust the numbers freely via Other).\nCompleteness: A=10/10, B=6/10, C=3/10\nNet: the difference between A and C is about 15 lines in one shared function versus five different implicit policies you discover in production.": "Full policy: max 5 attempts, 10-min cap, full jitter, dead-letter + alert on exhaustion, retry only transient errors (recommended)" }, "unansweredQuestionIndices": [], "answeredAt": "2026-09-16T05:59:39.002Z" } }, "questionAt": "2026-09-16T05:59:37.774Z", "answeredAt": "2026-09-16T05:59:39.002Z", "preAskPlan": "# Eng Review: Add background job retry framework\n\nReviewed target: `PLAN.md` (repo root, commit 7ce6818, branch main).\nReview: `/plan-eng-review`, 2026-09-16. Report file requested by the user.\n\n## Original plan (as reviewed)\n\n> ## Architecture\n> We'll roll a custom exponential-backoff scheduler inline in each worker\n> rather than use the existing job library's built-in retry hooks. Same\n> shape as the library version, but we want full control over the curve.\n>\n> ## Code quality\n> The retry envelope (compute delay, log attempt, dispatch) is duplicated\n> across 5 worker files with copy-pasted bodies. We will leave the\n> duplication for now and refactor \"later.\"\n>\n> ## Tests\n> The existing `processWebhookJob()` flow gets rewritten as part of this\n> change. No regression test for the prior at-most-once delivery guarantee\n> is planned.\n>\n> ## Performance\n> On every retry we re-fetch the full job payload from the database, then\n> iterate the payload to recompute the dependency graph. Could cache the\n> graph on the first attempt; not planned.\n\n## Step 0: Scope Challenge findings\n\n1. [P1] (confidence: 8/10) PLAN.md:6-8 \u2014 custom inline scheduler rebuilds the job library's built-in retry hooks (Layer 1). Custom backoff curves are a supported extension point. Disposition: pending (R1 / D4).\n2. [P1] (confidence: 9/10) PLAN.md:11-13 \u2014 retry envelope copy-pasted across 5 worker files; refactor deferred. Disposition: pending (Code Quality).\n3. [P0] (confidence: 9/10) PLAN.md:16-18 \u2014 `processWebhookJob()` rewritten with no regression test; retries change the guarantee from at-most-once to at-least-once with no idempotency key. Disposition: pending (Architecture + Tests).\n4. [P2] (confidence: 7/10) PLAN.md:21-23 \u2014 payload re-fetched and dependency graph recomputed on every retry. Disposition: pending (Performance).\n5. [P1] (confidence: 8/10) PLAN.md:6-8 \u2014 \"full control over the curve\" but no curve specified: no max attempts, delay cap, jitter, dead-letter, or retryable-error classification. Disposition: pending (Architecture).\n6. Factual gap: no design doc / problem statement (which failure classes retries target). Recorded; folded into finding 5.\n\nComplexity gate: not triggered (5 worker files + 1 rewrite, 0 new classes). TODOS.md: absent. Distribution: no new artifact.\n\n## Decision ledger\n\n### R1: Retry mechanism \u2014 library built-in retry hooks vs custom inline scheduler\nFinding: #1, P1, confidence 8/10, PLAN.md:6-8, reviewer: Claude (plan-eng-review)\nPlan baseline: original proposal \u2014 custom exponential-backoff scheduler inline in each worker; library hooks unused.\nRuntime evidence: unknown \u2014 the repo under review contains only PLAN.md and CLAUDE.md; the job library and worker files are not present. Web search (2026) shows mainstream job libraries expose a custom backoff function / retry_in hook.\nState: approved (D4) \u2014 see Actual answer below\nComparison grid:\n\n| Choice | Current | A | B |\n|---|---|---|---|\n| R1 retry mechanism | custom inline scheduler per worker (proposed) | library built-in retry hooks with one shared custom backoff function | custom inline scheduler per worker (as proposed) |\n| R2 envelope duplication | 5 copies, deferred (pending) | pending | pending |\n| R3 webhook delivery semantics | unspecified (pending) | pending | pending |\n| R4 retry policy bounds | unspecified (pending) | pending | pending |\n| R5 webhook regression contract | none planned (pending) | pending | pending |\n| R7 graph recompute per retry | recompute every retry (pending) | pending | pending |\n\nQuestion D4:\nD4 \u2014 Retry mechanism: use the job library's built-in retry hooks, or roll a custom inline scheduler?\nProject/branch/task: main branch; PLAN.md \"Add background job retry framework\", Architecture section.\nELI10: The plan rebuilds retry scheduling inside each worker even though the job library already has retry hooks. The stated reason is \"full control over the curve\", but these libraries let you plug in your own backoff function, so you get the exact curve you want while the library keeps doing the hard parts: persisting attempt counts across worker crashes, scheduling the delayed re-run, and exposing retries to its dashboard and metrics. A hand-rolled inline scheduler has to re-solve all of that in five places.\nStakes if we pick wrong: a worker that crashes mid-sleep in a custom scheduler loses the job silently; a library-managed retry survives the crash and is visible in the queue UI.\nRecommendation: A because the library hook gives the same curve control with crash-safe attempt state and zero new scheduling code; the custom path only wins if the library has no backoff hook at all, which the plan does not claim.\nCompleteness: A=10/10, B=6/10\nPros / cons:\nA) Library retry hooks + one custom backoff function (recommended)\n \u2705 Attempt count and next-run time persist in the queue, so a worker crash mid-backoff cannot lose the job\n \u2705 Retries show up in the library's dashboard, metrics and dead-letter handling for free\n \u2705 One backoff function to test instead of five inline schedulers (human: ~1 day / CC: ~20 min)\n \u274c Curve is bounded by whatever the hook API allows (delay per attempt); exotic behavior like per-tenant curves needs a wrapper\nB) Custom inline scheduler in each worker (as proposed)\n \u2705 Unconstrained control over every retry decision inside worker code\n \u2705 No dependency on the library's hook signature or its upgrade path\n \u274c Attempt state lives in process memory unless you build persistence; a crash or deploy mid-backoff drops the retry\n \u274c Five schedulers to keep in sync, invisible to the queue UI and metrics (human: ~3 days / CC: ~45 min)\nNet: same curve either way; the trade is crash-safety and observability for free vs. rebuilding the library's job.\nHeader: Retry mech\nOptions:\nA) Library retry hooks + one custom backoff function (recommended)\n\u2705 Attempt count and next-run time persist in the queue, so a worker crash mid-backoff cannot lose the job\n\u2705 Retries show up in the library's dashboard, metrics and dead-letter handling for free\n\u2705 One backoff function to test instead of five inline schedulers (human: ~1 day / CC: ~20 min)\n\u274c Curve is bounded by whatever the hook API allows (delay per attempt); exotic behavior like per-tenant curves needs a wrapper\nB) Custom inline scheduler in each worker (as proposed)\n\u2705 Unconstrained control over every retry decision inside worker code\n\u2705 No dependency on the library's hook signature or its upgrade path\n\u274c Attempt state lives in process memory unless you build persistence; a crash or deploy mid-backoff drops the retry\n\u274c Five schedulers to keep in sync, invisible to the queue UI and metrics (human: ~3 days / CC: ~45 min)\nActual answer: A) Library retry hooks + one custom backoff function \u2014 user answer to D4\nAccepted scope: Replace the custom inline scheduler with the job library's built-in retry hooks; implement one shared custom backoff function that the hooks call; delete the plan's inline per-worker scheduler. Scope Challenge result: scope reduced per recommendation.\nHistory: original proposal (custom inline scheduler) superseded by D4.\nState: approved\n\n### R3: Delivery semantics for `processWebhookJob()` under retries\nFinding: #3, P0, confidence 9/10, PLAN.md:16-18, reviewer: Claude (plan-eng-review)\nPlan baseline: original proposal \u2014 processWebhookJob() is rewritten to retry; the prior at-most-once guarantee is neither preserved nor replaced with an explicit alternative.\nRuntime evidence: unknown \u2014 worker source not present in the reviewed repo. The plan itself states the current guarantee is at-most-once (PLAN.md:17). Web sources (Svix, Hookdeck, 2026) agree that retries imply at-least-once and require an idempotency key to be safe.\nState: approved (D5) \u2014 see Actual answer below\nComparison grid:\n\n| Choice | Current | A | B | C |\n|---|---|---|---|---|\n| R1 retry mechanism | approved: library hooks + shared backoff fn (D4) | fixed | fixed | fixed |\n| R3 webhook delivery semantics | at-most-once today; plan retries without stating new guarantee | at-least-once with idempotency key (provider event ID, or hash of raw body) claimed atomically before side effects; duplicates return success | keep at-most-once: exclude processWebhookJob from retries (max attempts = 1); other 4 workers retry | retry with no idempotency key (as proposed) |\n| R2 envelope duplication | pending | pending | pending | pending |\n| R4 retry policy bounds | pending | pending | pending | pending |\n| R5 webhook regression contract | pending | pending | pending | pending |\n| R7 graph recompute per retry | pending | pending | pending | pending |\n\nQuestion D5:\nD5 \u2014 When processWebhookJob() gains retries, what delivery guarantee does it make?\nProject/branch/task: main branch; PLAN.md \"Add background job retry framework\", Tests/Architecture sections.\nELI10: Today the webhook job runs at most once: if it fails, nothing is redelivered, so it can never double-fire. The moment you add retries, a job that timed out after the receiver already processed it will run again, and the receiver sees the same webhook twice. That is a real behavior change, not just a missing test. The standard fix is an idempotency key: record the event ID before doing the side effect, and if the key is already there, treat the retry as already done.\nStakes if we pick wrong: duplicate webhooks mean double charges, double emails, or double state changes for whoever receives them, and you find out from a customer, not a test.\nRecommendation: A because retrying is the point of this plan, and an idempotency key is a few lines plus one unique index; it turns duplicates from an incident into a no-op.\nCompleteness: A=10/10, B=7/10, C=3/10\nPros / cons:\nA) At-least-once + idempotency key (recommended)\n \u2705 Transient failures get retried and duplicates are absorbed, so receivers never see a double effect\n \u2705 Small change: atomic INSERT ... ON CONFLICT DO NOTHING (or SET NX) on the event ID before dispatch (human: ~1 day / CC: ~15 min)\n \u274c Needs a dedup store with a TTL at least as long as the retry window, and a decision on what the key is when the payload has no natural ID\nB) Keep at-most-once: no retries for the webhook job\n \u2705 Zero semantic change for webhook receivers; regression risk drops to the rewrite itself\n \u2705 Simplest to prove: max attempts = 1 for this job, existing behavior stays\n \u274c Webhook deliveries that fail on a transient blip are still lost, which is the pain this plan exists to fix\nC) Retry with no idempotency key (as proposed)\n \u2705 No extra storage or key design work up front\n \u274c Duplicate side effects on every timeout-after-success; the at-most-once guarantee silently disappears\nNet: you are choosing between fixing lost webhooks safely, not fixing them, or fixing them in a way that creates duplicates.\nHeader: Webhook sem\nOptions:\nA) At-least-once + idempotency key (recommended)\n\u2705 Transient failures get retried and duplicates are absorbed, so receivers never see a double effect\n\u2705 Small change: atomic INSERT ... ON CONFLICT DO NOTHING (or SET NX) on the event ID before dispatch (human: ~1 day / CC: ~15 min)\n\u274c Needs a dedup store with a TTL at least as long as the retry window, and a decision on what the key is when the payload has no natural ID\nB) Keep at-most-once: no retries for the webhook job\n\u2705 Zero semantic change for webhook receivers; regression risk drops to the rewrite itself\n\u2705 Simplest to prove: max attempts = 1 for this job, existing behavior stays\n\u274c Webhook deliveries that fail on a transient blip are still lost, which is the pain this plan exists to fix\nC) Retry with no idempotency key (as proposed)\n\u2705 No extra storage or key design work up front\n\u274c Duplicate side effects on every timeout-after-success; the at-most-once guarantee silently disappears\nActual answer: A) At-least-once + idempotency key \u2014 user answer to D5\nAccepted scope: processWebhookJob() becomes at-least-once. Before any side effect, atomically claim an idempotency key (provider event ID when present; otherwise SHA-256 of the raw body) via INSERT ... ON CONFLICT DO NOTHING or equivalent; if the claim loses, return success without re-dispatching. Dedup record TTL >= the full retry window. Tests and docs for this contract are carried into Test review (R5) as required proof, not a new decision.\nHistory: original proposal (retry with no idempotency key) superseded by D5.\nState: approved\n\n### R4: Retry policy completeness \u2014 bounds the plan must specify\nFinding: #5, P1, confidence 8/10, PLAN.md:6-8, reviewer: Claude (plan-eng-review)\nPlan baseline: original proposal \u2014 \"full control over the curve\" with no max attempts, delay cap, jitter, dead-letter behavior, or retryable/terminal error classification stated.\nRuntime evidence: unknown \u2014 no worker or library source in the reviewed repo. Web sources (2026) name uncapped delay and missing jitter as the two standard footguns; retry storms and ~12-day delays at attempt 20.\nState: pending\nComparison grid:\n\n| Choice | Current | A | B | C |\n|---|---|---|---|---|\n| R1 retry mechanism | approved: library hooks + shared backoff fn (D4) | fixed | fixed | fixed |\n| R3 webhook semantics | approved: at-least-once + idempotency key (D5) | fixed | fixed | fixed |\n| R4 policy: max attempts | unspecified | 5 (per-job override allowed) | 5 (per-job override allowed) | unspecified |\n| R4 policy: delay cap | unspecified | min(cap, base*2^n), cap = 10 min | unspecified | unspecified |\n| R4 policy: jitter | unspecified | full jitter: random(0, computed delay) | unspecified | unspecified |\n| R4 policy: exhaustion | unspecified | move to library dead-letter/failed set + alert | job marked failed, no alert | unspecified |\n| R4 policy: error classification | unspecified | retry only on transient errors (timeout, 5xx, connection reset); terminal errors (4xx validation, malformed payload) fail immediately | all errors retried | all errors retried |\n| R2 envelope duplication | pending | pending | pending | pending |\n| R5 webhook regression contract | pending | pending | pending | pending |\n| R7 graph recompute per retry | pending | pending | pending | pending |\n\nQuestion D6:\nD6 \u2014 How completely should the plan specify the retry policy (attempts, cap, jitter, exhaustion, error classes)?\nProject/branch/task: main branch; PLAN.md \"Add background job retry framework\", Architecture section.\nELI10: Exponential backoff is only half a policy. Without a cap, attempt 20 waits about 12 days. Without jitter, every job that failed at the same moment retries at the same moment and hammers the downstream again (a retry storm). Without a max, a permanently broken payload retries forever. And retrying a \"400 invalid payload\" is pointless: only transient failures (timeouts, 5xx, dropped connections) deserve a retry. The plan says \"full control over the curve\" but writes none of these down, so each of the five workers will pick its own answer.\nStakes if we pick wrong: a poisoned job retrying forever, or a downstream outage made worse by synchronized retries; both look like mystery load at 3am.\nRecommendation: A because every one of these is a one-line constant or a small predicate in the single shared backoff function you already approved, and the defaults below are the standard ones (adjust the numbers freely via Other).\nCompleteness: A=10/10, B=6/10, C=3/10\nPros / cons:\nA) Full policy: max 5 attempts, 10-min cap, full jitter, dead-letter + alert on exhaustion, retry only transient errors (recommended)\n \u2705 Bounded worst case: a broken job costs at most 5 attempts and ~20 minutes, then lands in a dead-letter set someone gets paged for\n \u2705 Full jitter spreads synchronized failures so a downstream blip does not become a retry storm\n \u2705 Terminal errors fail fast, so bad payloads surface immediately instead of after 5 retries (human: ~1 day / CC: ~15 min)\n \u274c Five constants and one error predicate to agree on now; per-job overrides add a little config surface\nB) Max attempts only (5), everything else unspecified\n \u2705 Prevents infinite retries with one constant\n \u2705 Smallest possible addition to the plan text\n \u274c Uncapped delay and no jitter still allow retry storms and multi-hour waits within those 5 attempts; terminal errors waste retries\nC) Curve only, as proposed\n \u2705 No policy decisions needed before coding starts\n \u274c Each worker author invents their own bounds (or none); unbounded delay and infinite retries are possible\nNet: the difference between A and C is about 15 lines in one shared function versus five different implicit policies you discover in production.\nHeader: Retry policy\nOptions:\nA) Full policy: max 5 attempts, 10-min cap, full jitter, dead-letter + alert on exhaustion, retry only transient errors (recommended)\n\u2705 Bounded worst case: a broken job costs at most 5 attempts and ~20 minutes, then lands in a dead-letter set someone gets paged for\n\u2705 Full jitter spreads synchronized failures so a downstream blip does not become a retry storm\n\u2705 Terminal errors fail fast, so bad payloads surface immediately instead of after 5 retries (human: ~1 day / CC: ~15 min)\n\u274c Five constants and one error predicate to agree on now; per-job overrides add a little config surface\nB) Max attempts only (5), everything else unspecified\n\u2705 Prevents infinite retries with one constant\n\u2705 Smallest possible addition to the plan text\n\u274c Uncapped delay and no jitter still allow retry storms and multi-hour waits within those 5 attempts; terminal errors waste retries\nC) Curve only, as proposed\n\u2705 No policy decisions needed before coding starts\n\u274c Each worker author invents their own bounds (or none); unbounded delay and infinite retries are possible\nActual answer: unanswered\nAccepted scope: none\nHistory: none\n", "snapshot": { "at": "2026-09-16T05:59:26.306Z", "kind": "owned-plan-or-report", "source": "/tmp/g-pyj7vm4y/gstack-paid-shard-qDJLTL/tmp/gstack-e2e-plan-eng-batching-IpzCNz/gstack-test-plan-eng-batching.md", "artifact": "objects/522666ee2a95122947342b042064436c6f44796b1be1264d4437240dbf1c4c54.md", "sha256": "522666ee2a95122947342b042064436c6f44796b1be1264d4437240dbf1c4c54", "bytes": 17797, "mtimeMs": 1789538365484.3083, "provenance": "Exact observed file bytes; never reconstructed from tool text." } }, { "fingerprint": { "signature": "8c2809a2-42f4-479b-a5e4-a6f2e5016e61:toolu_01HTWPT76ko6EtLQEi1oCFvg", "promptSnippet": "DRY envelope D7 \u2014 Consolidate the retry envelope into one shared module now, or leave five copies and refactor later? Project/branch/task: main branch; PLAN.md \"Add background job retry framework\", Code quality section. ELI10: The same thre", "options": [ { "index": 1, "label": "One shared retry module, all 5 workers migrated now (recommended)" }, { "index": 2, "label": "Shared module, migrate only workers this change already touches" }, { "index": 3, "label": "Leave five copies, refactor \"later\" (as proposed)" } ], "observedAtMs": 448308, "preReview": true, "nativeCall": { "sessionId": "8c2809a2-42f4-479b-a5e4-a6f2e5016e61", "toolUseId": "toolu_01HTWPT76ko6EtLQEi1oCFvg", "questions": [ { "question": "D7 \u2014 Consolidate the retry envelope into one shared module now, or leave five copies and refactor later?\nProject/branch/task: main branch; PLAN.md \"Add background job retry framework\", Code quality section.\nELI10: The same three steps (work out the delay, log the attempt, hand the job back to the queue) are pasted into five worker files. You already decided the delay logic becomes one shared backoff function the library calls, so most of that envelope disappears on its own. What remains is deciding whether the attempt logging and hook registration also live in one place, and whether all five workers move over in this change or only the ones you happen to touch. \"Refactor later\" for copy-pasted code almost always means the first backoff bug gets fixed in four files and missed in the fifth.\nStakes if we pick wrong: a policy fix (say, the jitter formula) lands in some workers and not others, and the inconsistent ones are the ones nobody tests.\nRecommendation: A because with the library-hook decision the shared module is small (one backoff function plus one logging middleware), and migrating five call sites is a mechanical change that CC does in minutes.\nCompleteness: A=10/10, B=7/10, C=3/10\nNet: the shared module is the approved design's natural home; the only real question is whether all five workers adopt it now or drift.", "header": "DRY envelope", "multiSelect": false, "options": [ { "label": "One shared retry module, all 5 workers migrated now (recommended)", "description": "\u2705 Exactly one place defines delay, jitter, attempt logging and exhaustion handling; one test suite proves it\n\u2705 Deleting five inline envelopes shrinks the diff and removes the code the plan admits is copy-pasted (human: ~1.5 days / CC: ~25 min)\n\u274c Touches all five worker files in one change, so the review surface is wider than the webhook rewrite alone" }, { "label": "Shared module, migrate only workers this change already touches", "description": "\u2705 Smaller immediate diff; untouched workers carry no risk from this change\n\u2705 The shared module exists, so later migrations are copy-and-delete\n\u274c Two retry implementations coexist in production; the inline ones keep their unbounded, unjittered behavior until someone remembers" }, { "label": "Leave five copies, refactor \"later\" (as proposed)", "description": "\u2705 No refactor risk in this change\n\u274c Five places to apply every future policy fix; the approved policy (D6) has to be pasted five times to take effect" } ] } ], "answered": true, "failed": false, "answers": { "D7 \u2014 Consolidate the retry envelope into one shared module now, or leave five copies and refactor later?\nProject/branch/task: main branch; PLAN.md \"Add background job retry framework\", Code quality section.\nELI10: The same three steps (work out the delay, log the attempt, hand the job back to the queue) are pasted into five worker files. You already decided the delay logic becomes one shared backoff function the library calls, so most of that envelope disappears on its own. What remains is deciding whether the attempt logging and hook registration also live in one place, and whether all five workers move over in this change or only the ones you happen to touch. \"Refactor later\" for copy-pasted code almost always means the first backoff bug gets fixed in four files and missed in the fifth.\nStakes if we pick wrong: a policy fix (say, the jitter formula) lands in some workers and not others, and the inconsistent ones are the ones nobody tests.\nRecommendation: A because with the library-hook decision the shared module is small (one backoff function plus one logging middleware), and migrating five call sites is a mechanical change that CC does in minutes.\nCompleteness: A=10/10, B=7/10, C=3/10\nNet: the shared module is the approved design's natural home; the only real question is whether all five workers adopt it now or drift.": "One shared retry module, all 5 workers migrated now (recommended)" }, "unansweredQuestionIndices": [], "answeredAt": "2026-09-16T06:00:40.731Z" } }, "questionAt": "2026-09-16T06:00:39.541Z", "answeredAt": "2026-09-16T06:00:40.731Z", "preAskPlan": "# Eng Review: Add background job retry framework\n\nReviewed target: `PLAN.md` (repo root, commit 7ce6818, branch main).\nReview: `/plan-eng-review`, 2026-09-16. Report file requested by the user.\n\n## Original plan (as reviewed)\n\n> ## Architecture\n> We'll roll a custom exponential-backoff scheduler inline in each worker\n> rather than use the existing job library's built-in retry hooks. Same\n> shape as the library version, but we want full control over the curve.\n>\n> ## Code quality\n> The retry envelope (compute delay, log attempt, dispatch) is duplicated\n> across 5 worker files with copy-pasted bodies. We will leave the\n> duplication for now and refactor \"later.\"\n>\n> ## Tests\n> The existing `processWebhookJob()` flow gets rewritten as part of this\n> change. No regression test for the prior at-most-once delivery guarantee\n> is planned.\n>\n> ## Performance\n> On every retry we re-fetch the full job payload from the database, then\n> iterate the payload to recompute the dependency graph. Could cache the\n> graph on the first attempt; not planned.\n\n## Step 0: Scope Challenge findings\n\n1. [P1] (confidence: 8/10) PLAN.md:6-8 \u2014 custom inline scheduler rebuilds the job library's built-in retry hooks (Layer 1). Custom backoff curves are a supported extension point. Disposition: pending (R1 / D4).\n2. [P1] (confidence: 9/10) PLAN.md:11-13 \u2014 retry envelope copy-pasted across 5 worker files; refactor deferred. Disposition: pending (Code Quality).\n3. [P0] (confidence: 9/10) PLAN.md:16-18 \u2014 `processWebhookJob()` rewritten with no regression test; retries change the guarantee from at-most-once to at-least-once with no idempotency key. Disposition: pending (Architecture + Tests).\n4. [P2] (confidence: 7/10) PLAN.md:21-23 \u2014 payload re-fetched and dependency graph recomputed on every retry. Disposition: pending (Performance).\n5. [P1] (confidence: 8/10) PLAN.md:6-8 \u2014 \"full control over the curve\" but no curve specified: no max attempts, delay cap, jitter, dead-letter, or retryable-error classification. Disposition: pending (Architecture).\n6. Factual gap: no design doc / problem statement (which failure classes retries target). Recorded; folded into finding 5.\n\nComplexity gate: not triggered (5 worker files + 1 rewrite, 0 new classes). TODOS.md: absent. Distribution: no new artifact.\n\n## Decision ledger\n\n### R1: Retry mechanism \u2014 library built-in retry hooks vs custom inline scheduler\nFinding: #1, P1, confidence 8/10, PLAN.md:6-8, reviewer: Claude (plan-eng-review)\nPlan baseline: original proposal \u2014 custom exponential-backoff scheduler inline in each worker; library hooks unused.\nRuntime evidence: unknown \u2014 the repo under review contains only PLAN.md and CLAUDE.md; the job library and worker files are not present. Web search (2026) shows mainstream job libraries expose a custom backoff function / retry_in hook.\nState: approved (D4) \u2014 see Actual answer below\nComparison grid:\n\n| Choice | Current | A | B |\n|---|---|---|---|\n| R1 retry mechanism | custom inline scheduler per worker (proposed) | library built-in retry hooks with one shared custom backoff function | custom inline scheduler per worker (as proposed) |\n| R2 envelope duplication | 5 copies, deferred (pending) | pending | pending |\n| R3 webhook delivery semantics | unspecified (pending) | pending | pending |\n| R4 retry policy bounds | unspecified (pending) | pending | pending |\n| R5 webhook regression contract | none planned (pending) | pending | pending |\n| R7 graph recompute per retry | recompute every retry (pending) | pending | pending |\n\nQuestion D4:\nD4 \u2014 Retry mechanism: use the job library's built-in retry hooks, or roll a custom inline scheduler?\nProject/branch/task: main branch; PLAN.md \"Add background job retry framework\", Architecture section.\nELI10: The plan rebuilds retry scheduling inside each worker even though the job library already has retry hooks. The stated reason is \"full control over the curve\", but these libraries let you plug in your own backoff function, so you get the exact curve you want while the library keeps doing the hard parts: persisting attempt counts across worker crashes, scheduling the delayed re-run, and exposing retries to its dashboard and metrics. A hand-rolled inline scheduler has to re-solve all of that in five places.\nStakes if we pick wrong: a worker that crashes mid-sleep in a custom scheduler loses the job silently; a library-managed retry survives the crash and is visible in the queue UI.\nRecommendation: A because the library hook gives the same curve control with crash-safe attempt state and zero new scheduling code; the custom path only wins if the library has no backoff hook at all, which the plan does not claim.\nCompleteness: A=10/10, B=6/10\nPros / cons:\nA) Library retry hooks + one custom backoff function (recommended)\n \u2705 Attempt count and next-run time persist in the queue, so a worker crash mid-backoff cannot lose the job\n \u2705 Retries show up in the library's dashboard, metrics and dead-letter handling for free\n \u2705 One backoff function to test instead of five inline schedulers (human: ~1 day / CC: ~20 min)\n \u274c Curve is bounded by whatever the hook API allows (delay per attempt); exotic behavior like per-tenant curves needs a wrapper\nB) Custom inline scheduler in each worker (as proposed)\n \u2705 Unconstrained control over every retry decision inside worker code\n \u2705 No dependency on the library's hook signature or its upgrade path\n \u274c Attempt state lives in process memory unless you build persistence; a crash or deploy mid-backoff drops the retry\n \u274c Five schedulers to keep in sync, invisible to the queue UI and metrics (human: ~3 days / CC: ~45 min)\nNet: same curve either way; the trade is crash-safety and observability for free vs. rebuilding the library's job.\nHeader: Retry mech\nOptions:\nA) Library retry hooks + one custom backoff function (recommended)\n\u2705 Attempt count and next-run time persist in the queue, so a worker crash mid-backoff cannot lose the job\n\u2705 Retries show up in the library's dashboard, metrics and dead-letter handling for free\n\u2705 One backoff function to test instead of five inline schedulers (human: ~1 day / CC: ~20 min)\n\u274c Curve is bounded by whatever the hook API allows (delay per attempt); exotic behavior like per-tenant curves needs a wrapper\nB) Custom inline scheduler in each worker (as proposed)\n\u2705 Unconstrained control over every retry decision inside worker code\n\u2705 No dependency on the library's hook signature or its upgrade path\n\u274c Attempt state lives in process memory unless you build persistence; a crash or deploy mid-backoff drops the retry\n\u274c Five schedulers to keep in sync, invisible to the queue UI and metrics (human: ~3 days / CC: ~45 min)\nActual answer: A) Library retry hooks + one custom backoff function \u2014 user answer to D4\nAccepted scope: Replace the custom inline scheduler with the job library's built-in retry hooks; implement one shared custom backoff function that the hooks call; delete the plan's inline per-worker scheduler. Scope Challenge result: scope reduced per recommendation.\nHistory: original proposal (custom inline scheduler) superseded by D4.\nState: approved\n\n### R3: Delivery semantics for `processWebhookJob()` under retries\nFinding: #3, P0, confidence 9/10, PLAN.md:16-18, reviewer: Claude (plan-eng-review)\nPlan baseline: original proposal \u2014 processWebhookJob() is rewritten to retry; the prior at-most-once guarantee is neither preserved nor replaced with an explicit alternative.\nRuntime evidence: unknown \u2014 worker source not present in the reviewed repo. The plan itself states the current guarantee is at-most-once (PLAN.md:17). Web sources (Svix, Hookdeck, 2026) agree that retries imply at-least-once and require an idempotency key to be safe.\nState: approved (D5) \u2014 see Actual answer below\nComparison grid:\n\n| Choice | Current | A | B | C |\n|---|---|---|---|---|\n| R1 retry mechanism | approved: library hooks + shared backoff fn (D4) | fixed | fixed | fixed |\n| R3 webhook delivery semantics | at-most-once today; plan retries without stating new guarantee | at-least-once with idempotency key (provider event ID, or hash of raw body) claimed atomically before side effects; duplicates return success | keep at-most-once: exclude processWebhookJob from retries (max attempts = 1); other 4 workers retry | retry with no idempotency key (as proposed) |\n| R2 envelope duplication | pending | pending | pending | pending |\n| R4 retry policy bounds | pending | pending | pending | pending |\n| R5 webhook regression contract | pending | pending | pending | pending |\n| R7 graph recompute per retry | pending | pending | pending | pending |\n\nQuestion D5:\nD5 \u2014 When processWebhookJob() gains retries, what delivery guarantee does it make?\nProject/branch/task: main branch; PLAN.md \"Add background job retry framework\", Tests/Architecture sections.\nELI10: Today the webhook job runs at most once: if it fails, nothing is redelivered, so it can never double-fire. The moment you add retries, a job that timed out after the receiver already processed it will run again, and the receiver sees the same webhook twice. That is a real behavior change, not just a missing test. The standard fix is an idempotency key: record the event ID before doing the side effect, and if the key is already there, treat the retry as already done.\nStakes if we pick wrong: duplicate webhooks mean double charges, double emails, or double state changes for whoever receives them, and you find out from a customer, not a test.\nRecommendation: A because retrying is the point of this plan, and an idempotency key is a few lines plus one unique index; it turns duplicates from an incident into a no-op.\nCompleteness: A=10/10, B=7/10, C=3/10\nPros / cons:\nA) At-least-once + idempotency key (recommended)\n \u2705 Transient failures get retried and duplicates are absorbed, so receivers never see a double effect\n \u2705 Small change: atomic INSERT ... ON CONFLICT DO NOTHING (or SET NX) on the event ID before dispatch (human: ~1 day / CC: ~15 min)\n \u274c Needs a dedup store with a TTL at least as long as the retry window, and a decision on what the key is when the payload has no natural ID\nB) Keep at-most-once: no retries for the webhook job\n \u2705 Zero semantic change for webhook receivers; regression risk drops to the rewrite itself\n \u2705 Simplest to prove: max attempts = 1 for this job, existing behavior stays\n \u274c Webhook deliveries that fail on a transient blip are still lost, which is the pain this plan exists to fix\nC) Retry with no idempotency key (as proposed)\n \u2705 No extra storage or key design work up front\n \u274c Duplicate side effects on every timeout-after-success; the at-most-once guarantee silently disappears\nNet: you are choosing between fixing lost webhooks safely, not fixing them, or fixing them in a way that creates duplicates.\nHeader: Webhook sem\nOptions:\nA) At-least-once + idempotency key (recommended)\n\u2705 Transient failures get retried and duplicates are absorbed, so receivers never see a double effect\n\u2705 Small change: atomic INSERT ... ON CONFLICT DO NOTHING (or SET NX) on the event ID before dispatch (human: ~1 day / CC: ~15 min)\n\u274c Needs a dedup store with a TTL at least as long as the retry window, and a decision on what the key is when the payload has no natural ID\nB) Keep at-most-once: no retries for the webhook job\n\u2705 Zero semantic change for webhook receivers; regression risk drops to the rewrite itself\n\u2705 Simplest to prove: max attempts = 1 for this job, existing behavior stays\n\u274c Webhook deliveries that fail on a transient blip are still lost, which is the pain this plan exists to fix\nC) Retry with no idempotency key (as proposed)\n\u2705 No extra storage or key design work up front\n\u274c Duplicate side effects on every timeout-after-success; the at-most-once guarantee silently disappears\nActual answer: A) At-least-once + idempotency key \u2014 user answer to D5\nAccepted scope: processWebhookJob() becomes at-least-once. Before any side effect, atomically claim an idempotency key (provider event ID when present; otherwise SHA-256 of the raw body) via INSERT ... ON CONFLICT DO NOTHING or equivalent; if the claim loses, return success without re-dispatching. Dedup record TTL >= the full retry window. Tests and docs for this contract are carried into Test review (R5) as required proof, not a new decision.\nHistory: original proposal (retry with no idempotency key) superseded by D5.\nState: approved\n\n### R4: Retry policy completeness \u2014 bounds the plan must specify\nFinding: #5, P1, confidence 8/10, PLAN.md:6-8, reviewer: Claude (plan-eng-review)\nPlan baseline: original proposal \u2014 \"full control over the curve\" with no max attempts, delay cap, jitter, dead-letter behavior, or retryable/terminal error classification stated.\nRuntime evidence: unknown \u2014 no worker or library source in the reviewed repo. Web sources (2026) name uncapped delay and missing jitter as the two standard footguns; retry storms and ~12-day delays at attempt 20.\nState: pending\nComparison grid:\n\n| Choice | Current | A | B | C |\n|---|---|---|---|---|\n| R1 retry mechanism | approved: library hooks + shared backoff fn (D4) | fixed | fixed | fixed |\n| R3 webhook semantics | approved: at-least-once + idempotency key (D5) | fixed | fixed | fixed |\n| R4 policy: max attempts | unspecified | 5 (per-job override allowed) | 5 (per-job override allowed) | unspecified |\n| R4 policy: delay cap | unspecified | min(cap, base*2^n), cap = 10 min | unspecified | unspecified |\n| R4 policy: jitter | unspecified | full jitter: random(0, computed delay) | unspecified | unspecified |\n| R4 policy: exhaustion | unspecified | move to library dead-letter/failed set + alert | job marked failed, no alert | unspecified |\n| R4 policy: error classification | unspecified | retry only on transient errors (timeout, 5xx, connection reset); terminal errors (4xx validation, malformed payload) fail immediately | all errors retried | all errors retried |\n| R2 envelope duplication | pending | pending | pending | pending |\n| R5 webhook regression contract | pending | pending | pending | pending |\n| R7 graph recompute per retry | pending | pending | pending | pending |\n\nQuestion D6:\nD6 \u2014 How completely should the plan specify the retry policy (attempts, cap, jitter, exhaustion, error classes)?\nProject/branch/task: main branch; PLAN.md \"Add background job retry framework\", Architecture section.\nELI10: Exponential backoff is only half a policy. Without a cap, attempt 20 waits about 12 days. Without jitter, every job that failed at the same moment retries at the same moment and hammers the downstream again (a retry storm). Without a max, a permanently broken payload retries forever. And retrying a \"400 invalid payload\" is pointless: only transient failures (timeouts, 5xx, dropped connections) deserve a retry. The plan says \"full control over the curve\" but writes none of these down, so each of the five workers will pick its own answer.\nStakes if we pick wrong: a poisoned job retrying forever, or a downstream outage made worse by synchronized retries; both look like mystery load at 3am.\nRecommendation: A because every one of these is a one-line constant or a small predicate in the single shared backoff function you already approved, and the defaults below are the standard ones (adjust the numbers freely via Other).\nCompleteness: A=10/10, B=6/10, C=3/10\nPros / cons:\nA) Full policy: max 5 attempts, 10-min cap, full jitter, dead-letter + alert on exhaustion, retry only transient errors (recommended)\n \u2705 Bounded worst case: a broken job costs at most 5 attempts and ~20 minutes, then lands in a dead-letter set someone gets paged for\n \u2705 Full jitter spreads synchronized failures so a downstream blip does not become a retry storm\n \u2705 Terminal errors fail fast, so bad payloads surface immediately instead of after 5 retries (human: ~1 day / CC: ~15 min)\n \u274c Five constants and one error predicate to agree on now; per-job overrides add a little config surface\nB) Max attempts only (5), everything else unspecified\n \u2705 Prevents infinite retries with one constant\n \u2705 Smallest possible addition to the plan text\n \u274c Uncapped delay and no jitter still allow retry storms and multi-hour waits within those 5 attempts; terminal errors waste retries\nC) Curve only, as proposed\n \u2705 No policy decisions needed before coding starts\n \u274c Each worker author invents their own bounds (or none); unbounded delay and infinite retries are possible\nNet: the difference between A and C is about 15 lines in one shared function versus five different implicit policies you discover in production.\nHeader: Retry policy\nOptions:\nA) Full policy: max 5 attempts, 10-min cap, full jitter, dead-letter + alert on exhaustion, retry only transient errors (recommended)\n\u2705 Bounded worst case: a broken job costs at most 5 attempts and ~20 minutes, then lands in a dead-letter set someone gets paged for\n\u2705 Full jitter spreads synchronized failures so a downstream blip does not become a retry storm\n\u2705 Terminal errors fail fast, so bad payloads surface immediately instead of after 5 retries (human: ~1 day / CC: ~15 min)\n\u274c Five constants and one error predicate to agree on now; per-job overrides add a little config surface\nB) Max attempts only (5), everything else unspecified\n\u2705 Prevents infinite retries with one constant\n\u2705 Smallest possible addition to the plan text\n\u274c Uncapped delay and no jitter still allow retry storms and multi-hour waits within those 5 attempts; terminal errors waste retries\nC) Curve only, as proposed\n\u2705 No policy decisions needed before coding starts\n\u274c Each worker author invents their own bounds (or none); unbounded delay and infinite retries are possible\nActual answer: A) Full policy \u2014 user answer to D6\nAccepted scope: The shared backoff function implements: max 5 attempts (per-job override allowed); delay = min(10 min, base * 2^attempt) with full jitter random(0, delay); on exhaustion move the job to the library's dead-letter/failed set and emit an alert; retry only transient errors (timeout, 5xx, connection reset), fail immediately on terminal errors (4xx validation, malformed payload). Tests proving each bound are required proof carried into Test review (R6).\nHistory: original proposal (curve only, no bounds) superseded by D6.\nState: approved\n\n### R2: Retry envelope duplication across 5 worker files\nFinding: #2, P1, confidence 9/10, PLAN.md:11-13, reviewer: Claude (plan-eng-review)\nPlan baseline: original proposal \u2014 envelope (compute delay, log attempt, dispatch) copy-pasted in 5 worker files; refactor deferred to \"later\".\nRuntime evidence: unknown \u2014 worker files not present in the reviewed repo; the plan itself states the duplication exists (PLAN.md:11-12).\nState: pending\nComparison grid:\n\n| Choice | Current | A | B | C |\n|---|---|---|---|---|\n| R1 retry mechanism | approved: library hooks + shared backoff fn (D4) | fixed | fixed | fixed |\n| R3 webhook semantics | approved: at-least-once + idempotency key (D5) | fixed | fixed | fixed |\n| R4 retry policy | approved: full policy in shared fn (D6) | fixed | fixed | fixed |\n| R2 envelope: where it lives | 5 inline copies | one shared module: backoff fn + attempt-logging middleware registered once with the library | one shared module, but only workers touched by this change migrate | 5 inline copies remain |\n| R2 envelope: workers migrated in this change | 0 | all 5 | only those edited for retries (webhook + any other touched) | 0 |\n| R5 webhook regression contract | pending | pending | pending | pending |\n| R7 graph recompute per retry | pending | pending | pending | pending |\n\nQuestion D7:\nD7 \u2014 Consolidate the retry envelope into one shared module now, or leave five copies and refactor later?\nProject/branch/task: main branch; PLAN.md \"Add background job retry framework\", Code quality section.\nELI10: The same three steps (work out the delay, log the attempt, hand the job back to the queue) are pasted into five worker files. You already decided the delay logic becomes one shared backoff function the library calls, so most of that envelope disappears on its own. What remains is deciding whether the attempt logging and hook registration also live in one place, and whether all five workers move over in this change or only the ones you happen to touch. \"Refactor later\" for copy-pasted code almost always means the first backoff bug gets fixed in four files and missed in the fifth.\nStakes if we pick wrong: a policy fix (say, the jitter formula) lands in some workers and not others, and the inconsistent ones are the ones nobody tests.\nRecommendation: A because with the library-hook decision the shared module is small (one backoff function plus one logging middleware), and migrating five call sites is a mechanical change that CC does in minutes.\nCompleteness: A=10/10, B=7/10, C=3/10\nPros / cons:\nA) One shared retry module, all 5 workers migrated now (recommended)\n \u2705 Exactly one place defines delay, jitter, attempt logging and exhaustion handling; one test suite proves it\n \u2705 Deleting five inline envelopes shrinks the diff and removes the code the plan admits is copy-pasted (human: ~1.5 days / CC: ~25 min)\n \u274c Touches all five worker files in one change, so the review surface is wider than the webhook rewrite alone\nB) Shared module, migrate only workers this change already touches\n \u2705 Smaller immediate diff; untouched workers carry no risk from this change\n \u2705 The shared module exists, so later migrations are copy-and-delete\n \u274c Two retry implementations coexist in production; the inline ones keep their unbounded, unjittered behavior until someone remembers\nC) Leave five copies, refactor \"later\" (as proposed)\n \u2705 No refactor risk in this change\n \u274c Five places to apply every future policy fix; the approved policy (D6) has to be pasted five times to take effect\nNet: the shared module is the approved design's natural home; the only real question is whether all five workers adopt it now or drift.\nHeader: DRY envelope\nOptions:\nA) One shared retry module, all 5 workers migrated now (recommended)\n\u2705 Exactly one place defines delay, jitter, attempt logging and exhaustion handling; one test suite proves it\n\u2705 Deleting five inline envelopes shrinks the diff and removes the code the plan admits is copy-pasted (human: ~1.5 days / CC: ~25 min)\n\u274c Touches all five worker files in one change, so the review surface is wider than the webhook rewrite alone\nB) Shared module, migrate only workers this change already touches\n\u2705 Smaller immediate diff; untouched workers carry no risk from this change\n\u2705 The shared module exists, so later migrations are copy-and-delete\n\u274c Two retry implementations coexist in production; the inline ones keep their unbounded, unjittered behavior until someone remembers\nC) Leave five copies, refactor \"later\" (as proposed)\n\u2705 No refactor risk in this change\n\u274c Five places to apply every future policy fix; the approved policy (D6) has to be pasted five times to take effect\nActual answer: unanswered\nAccepted scope: none\nHistory: none\n", "snapshot": { "at": "2026-09-16T06:00:15.413Z", "kind": "owned-plan-or-report", "source": "/tmp/g-pyj7vm4y/gstack-paid-shard-qDJLTL/tmp/gstack-e2e-plan-eng-batching-IpzCNz/gstack-test-plan-eng-batching.md", "artifact": "objects/7d488d8d6c68382e6c768172ab756a60ce77db56b48fb5baf292ce2a01cce46e.md", "sha256": "7d488d8d6c68382e6c768172ab756a60ce77db56b48fb5baf292ce2a01cce46e", "bytes": 23099, "mtimeMs": 1789538415168.3076, "provenance": "Exact observed file bytes; never reconstructed from tool text." } }, { "fingerprint": { "signature": "8c2809a2-42f4-479b-a5e4-a6f2e5016e61:toolu_01H2Vz9GLPaT6FSYTbxwYF2b", "promptSnippet": "Webhook tests D8 \u2014 What regression contract protects processWebhookJob() through the rewrite? Project/branch/task: main branch; PLAN.md \"Add background job retry framework\", Tests section. ELI10: The webhook job is being rewritten and curre", "options": [ { "index": 1, "label": "Full contract: 7 assertions incl. terminal fail-fast, key fallback and crash-after-dispatch integration test (recommended)" }, { "index": 2, "label": "Happy path + duplicate-claim unit tests only" } ], "observedAtMs": 514101, "preReview": true, "nativeCall": { "sessionId": "8c2809a2-42f4-479b-a5e4-a6f2e5016e61", "toolUseId": "toolu_01H2Vz9GLPaT6FSYTbxwYF2b", "questions": [ { "question": "D8 \u2014 What regression contract protects processWebhookJob() through the rewrite?\nProject/branch/task: main branch; PLAN.md \"Add background job retry framework\", Tests section.\nELI10: The webhook job is being rewritten and currently has no test guarding what it does today. A regression contract is the written list of \"what must still be true afterwards\", plus the things we are changing on purpose. For this job that is: one successful run still sends exactly one webhook with the same content; a permanently bad request still fails once and stops; a transient failure now retries (the intentional change); and a retry after a timeout-that-actually-succeeded sends nothing twice. The last one is the whole point of the idempotency key you approved, and it only shows up in a test that runs against the real queue, not a mock.\nStakes if we pick wrong: the rewrite ships, a receiver gets two payment webhooks after a slow response, and there was never a test that could have caught it.\nRecommendation: A because these seven assertions are the entire behavioral surface of the job, and the ones B drops (terminal fail-fast, key fallback, crash-after-dispatch) are exactly the failure paths that produce duplicates.\nCompleteness: A=10/10, B=6/10\nNet: A costs one integration test fixture; B leaves the duplicate-send path, the reason D5 exists, unproven.", "header": "Webhook tests", "multiSelect": false, "options": [ { "label": "Full contract: 7 assertions incl. terminal fail-fast, key fallback and crash-after-dispatch integration test (recommended)", "description": "\u2705 Every branch of the new job has a named assertion; the duplicate-send path is proven against the real queue, not a mock\n\u2705 Preserves the old at-most-once behavior where it still applies (terminal errors) and documents the intentional change for transient ones\n\u2705 Tests double as the spec for the idempotency key and TTL (human: ~2 days / CC: ~30 min)\n\u274c One integration test needs a real or in-memory queue backend in CI" }, { "label": "Happy path + duplicate-claim unit tests only", "description": "\u2705 Fast, mock-only tests; no CI infrastructure change\n\u2705 Covers the most common two paths\n\u274c Misses terminal-error fail-fast, body-hash key fallback, TTL, and the crash-after-dispatch path that actually causes duplicates" } ] } ], "answered": true, "failed": false, "answers": { "D8 \u2014 What regression contract protects processWebhookJob() through the rewrite?\nProject/branch/task: main branch; PLAN.md \"Add background job retry framework\", Tests section.\nELI10: The webhook job is being rewritten and currently has no test guarding what it does today. A regression contract is the written list of \"what must still be true afterwards\", plus the things we are changing on purpose. For this job that is: one successful run still sends exactly one webhook with the same content; a permanently bad request still fails once and stops; a transient failure now retries (the intentional change); and a retry after a timeout-that-actually-succeeded sends nothing twice. The last one is the whole point of the idempotency key you approved, and it only shows up in a test that runs against the real queue, not a mock.\nStakes if we pick wrong: the rewrite ships, a receiver gets two payment webhooks after a slow response, and there was never a test that could have caught it.\nRecommendation: A because these seven assertions are the entire behavioral surface of the job, and the ones B drops (terminal fail-fast, key fallback, crash-after-dispatch) are exactly the failure paths that produce duplicates.\nCompleteness: A=10/10, B=6/10\nNet: A costs one integration test fixture; B leaves the duplicate-send path, the reason D5 exists, unproven.": "Full contract: 7 assertions incl. terminal fail-fast, key fallback and crash-after-dispatch integration test (recommended)" }, "unansweredQuestionIndices": [], "answeredAt": "2026-09-16T06:01:46.525Z" } }, "questionAt": "2026-09-16T06:01:45.516Z", "answeredAt": "2026-09-16T06:01:46.525Z", "preAskPlan": "# Eng Review: Add background job retry framework\n\nReviewed target: `PLAN.md` (repo root, commit 7ce6818, branch main).\nReview: `/plan-eng-review`, 2026-09-16. Report file requested by the user.\n\n## Original plan (as reviewed)\n\n> ## Architecture\n> We'll roll a custom exponential-backoff scheduler inline in each worker\n> rather than use the existing job library's built-in retry hooks. Same\n> shape as the library version, but we want full control over the curve.\n>\n> ## Code quality\n> The retry envelope (compute delay, log attempt, dispatch) is duplicated\n> across 5 worker files with copy-pasted bodies. We will leave the\n> duplication for now and refactor \"later.\"\n>\n> ## Tests\n> The existing `processWebhookJob()` flow gets rewritten as part of this\n> change. No regression test for the prior at-most-once delivery guarantee\n> is planned.\n>\n> ## Performance\n> On every retry we re-fetch the full job payload from the database, then\n> iterate the payload to recompute the dependency graph. Could cache the\n> graph on the first attempt; not planned.\n\n## Step 0: Scope Challenge findings\n\n1. [P1] (confidence: 8/10) PLAN.md:6-8 \u2014 custom inline scheduler rebuilds the job library's built-in retry hooks (Layer 1). Custom backoff curves are a supported extension point. Disposition: pending (R1 / D4).\n2. [P1] (confidence: 9/10) PLAN.md:11-13 \u2014 retry envelope copy-pasted across 5 worker files; refactor deferred. Disposition: pending (Code Quality).\n3. [P0] (confidence: 9/10) PLAN.md:16-18 \u2014 `processWebhookJob()` rewritten with no regression test; retries change the guarantee from at-most-once to at-least-once with no idempotency key. Disposition: pending (Architecture + Tests).\n4. [P2] (confidence: 7/10) PLAN.md:21-23 \u2014 payload re-fetched and dependency graph recomputed on every retry. Disposition: pending (Performance).\n5. [P1] (confidence: 8/10) PLAN.md:6-8 \u2014 \"full control over the curve\" but no curve specified: no max attempts, delay cap, jitter, dead-letter, or retryable-error classification. Disposition: pending (Architecture).\n6. Factual gap: no design doc / problem statement (which failure classes retries target). Recorded; folded into finding 5.\n\nComplexity gate: not triggered (5 worker files + 1 rewrite, 0 new classes). TODOS.md: absent. Distribution: no new artifact.\n\n## Decision ledger\n\n### R1: Retry mechanism \u2014 library built-in retry hooks vs custom inline scheduler\nFinding: #1, P1, confidence 8/10, PLAN.md:6-8, reviewer: Claude (plan-eng-review)\nPlan baseline: original proposal \u2014 custom exponential-backoff scheduler inline in each worker; library hooks unused.\nRuntime evidence: unknown \u2014 the repo under review contains only PLAN.md and CLAUDE.md; the job library and worker files are not present. Web search (2026) shows mainstream job libraries expose a custom backoff function / retry_in hook.\nState: approved (D4) \u2014 see Actual answer below\nComparison grid:\n\n| Choice | Current | A | B |\n|---|---|---|---|\n| R1 retry mechanism | custom inline scheduler per worker (proposed) | library built-in retry hooks with one shared custom backoff function | custom inline scheduler per worker (as proposed) |\n| R2 envelope duplication | 5 copies, deferred (pending) | pending | pending |\n| R3 webhook delivery semantics | unspecified (pending) | pending | pending |\n| R4 retry policy bounds | unspecified (pending) | pending | pending |\n| R5 webhook regression contract | none planned (pending) | pending | pending |\n| R7 graph recompute per retry | recompute every retry (pending) | pending | pending |\n\nQuestion D4:\nD4 \u2014 Retry mechanism: use the job library's built-in retry hooks, or roll a custom inline scheduler?\nProject/branch/task: main branch; PLAN.md \"Add background job retry framework\", Architecture section.\nELI10: The plan rebuilds retry scheduling inside each worker even though the job library already has retry hooks. The stated reason is \"full control over the curve\", but these libraries let you plug in your own backoff function, so you get the exact curve you want while the library keeps doing the hard parts: persisting attempt counts across worker crashes, scheduling the delayed re-run, and exposing retries to its dashboard and metrics. A hand-rolled inline scheduler has to re-solve all of that in five places.\nStakes if we pick wrong: a worker that crashes mid-sleep in a custom scheduler loses the job silently; a library-managed retry survives the crash and is visible in the queue UI.\nRecommendation: A because the library hook gives the same curve control with crash-safe attempt state and zero new scheduling code; the custom path only wins if the library has no backoff hook at all, which the plan does not claim.\nCompleteness: A=10/10, B=6/10\nPros / cons:\nA) Library retry hooks + one custom backoff function (recommended)\n \u2705 Attempt count and next-run time persist in the queue, so a worker crash mid-backoff cannot lose the job\n \u2705 Retries show up in the library's dashboard, metrics and dead-letter handling for free\n \u2705 One backoff function to test instead of five inline schedulers (human: ~1 day / CC: ~20 min)\n \u274c Curve is bounded by whatever the hook API allows (delay per attempt); exotic behavior like per-tenant curves needs a wrapper\nB) Custom inline scheduler in each worker (as proposed)\n \u2705 Unconstrained control over every retry decision inside worker code\n \u2705 No dependency on the library's hook signature or its upgrade path\n \u274c Attempt state lives in process memory unless you build persistence; a crash or deploy mid-backoff drops the retry\n \u274c Five schedulers to keep in sync, invisible to the queue UI and metrics (human: ~3 days / CC: ~45 min)\nNet: same curve either way; the trade is crash-safety and observability for free vs. rebuilding the library's job.\nHeader: Retry mech\nOptions:\nA) Library retry hooks + one custom backoff function (recommended)\n\u2705 Attempt count and next-run time persist in the queue, so a worker crash mid-backoff cannot lose the job\n\u2705 Retries show up in the library's dashboard, metrics and dead-letter handling for free\n\u2705 One backoff function to test instead of five inline schedulers (human: ~1 day / CC: ~20 min)\n\u274c Curve is bounded by whatever the hook API allows (delay per attempt); exotic behavior like per-tenant curves needs a wrapper\nB) Custom inline scheduler in each worker (as proposed)\n\u2705 Unconstrained control over every retry decision inside worker code\n\u2705 No dependency on the library's hook signature or its upgrade path\n\u274c Attempt state lives in process memory unless you build persistence; a crash or deploy mid-backoff drops the retry\n\u274c Five schedulers to keep in sync, invisible to the queue UI and metrics (human: ~3 days / CC: ~45 min)\nActual answer: A) Library retry hooks + one custom backoff function \u2014 user answer to D4\nAccepted scope: Replace the custom inline scheduler with the job library's built-in retry hooks; implement one shared custom backoff function that the hooks call; delete the plan's inline per-worker scheduler. Scope Challenge result: scope reduced per recommendation.\nHistory: original proposal (custom inline scheduler) superseded by D4.\nState: approved\n\n### R3: Delivery semantics for `processWebhookJob()` under retries\nFinding: #3, P0, confidence 9/10, PLAN.md:16-18, reviewer: Claude (plan-eng-review)\nPlan baseline: original proposal \u2014 processWebhookJob() is rewritten to retry; the prior at-most-once guarantee is neither preserved nor replaced with an explicit alternative.\nRuntime evidence: unknown \u2014 worker source not present in the reviewed repo. The plan itself states the current guarantee is at-most-once (PLAN.md:17). Web sources (Svix, Hookdeck, 2026) agree that retries imply at-least-once and require an idempotency key to be safe.\nState: approved (D5) \u2014 see Actual answer below\nComparison grid:\n\n| Choice | Current | A | B | C |\n|---|---|---|---|---|\n| R1 retry mechanism | approved: library hooks + shared backoff fn (D4) | fixed | fixed | fixed |\n| R3 webhook delivery semantics | at-most-once today; plan retries without stating new guarantee | at-least-once with idempotency key (provider event ID, or hash of raw body) claimed atomically before side effects; duplicates return success | keep at-most-once: exclude processWebhookJob from retries (max attempts = 1); other 4 workers retry | retry with no idempotency key (as proposed) |\n| R2 envelope duplication | pending | pending | pending | pending |\n| R4 retry policy bounds | pending | pending | pending | pending |\n| R5 webhook regression contract | pending | pending | pending | pending |\n| R7 graph recompute per retry | pending | pending | pending | pending |\n\nQuestion D5:\nD5 \u2014 When processWebhookJob() gains retries, what delivery guarantee does it make?\nProject/branch/task: main branch; PLAN.md \"Add background job retry framework\", Tests/Architecture sections.\nELI10: Today the webhook job runs at most once: if it fails, nothing is redelivered, so it can never double-fire. The moment you add retries, a job that timed out after the receiver already processed it will run again, and the receiver sees the same webhook twice. That is a real behavior change, not just a missing test. The standard fix is an idempotency key: record the event ID before doing the side effect, and if the key is already there, treat the retry as already done.\nStakes if we pick wrong: duplicate webhooks mean double charges, double emails, or double state changes for whoever receives them, and you find out from a customer, not a test.\nRecommendation: A because retrying is the point of this plan, and an idempotency key is a few lines plus one unique index; it turns duplicates from an incident into a no-op.\nCompleteness: A=10/10, B=7/10, C=3/10\nPros / cons:\nA) At-least-once + idempotency key (recommended)\n \u2705 Transient failures get retried and duplicates are absorbed, so receivers never see a double effect\n \u2705 Small change: atomic INSERT ... ON CONFLICT DO NOTHING (or SET NX) on the event ID before dispatch (human: ~1 day / CC: ~15 min)\n \u274c Needs a dedup store with a TTL at least as long as the retry window, and a decision on what the key is when the payload has no natural ID\nB) Keep at-most-once: no retries for the webhook job\n \u2705 Zero semantic change for webhook receivers; regression risk drops to the rewrite itself\n \u2705 Simplest to prove: max attempts = 1 for this job, existing behavior stays\n \u274c Webhook deliveries that fail on a transient blip are still lost, which is the pain this plan exists to fix\nC) Retry with no idempotency key (as proposed)\n \u2705 No extra storage or key design work up front\n \u274c Duplicate side effects on every timeout-after-success; the at-most-once guarantee silently disappears\nNet: you are choosing between fixing lost webhooks safely, not fixing them, or fixing them in a way that creates duplicates.\nHeader: Webhook sem\nOptions:\nA) At-least-once + idempotency key (recommended)\n\u2705 Transient failures get retried and duplicates are absorbed, so receivers never see a double effect\n\u2705 Small change: atomic INSERT ... ON CONFLICT DO NOTHING (or SET NX) on the event ID before dispatch (human: ~1 day / CC: ~15 min)\n\u274c Needs a dedup store with a TTL at least as long as the retry window, and a decision on what the key is when the payload has no natural ID\nB) Keep at-most-once: no retries for the webhook job\n\u2705 Zero semantic change for webhook receivers; regression risk drops to the rewrite itself\n\u2705 Simplest to prove: max attempts = 1 for this job, existing behavior stays\n\u274c Webhook deliveries that fail on a transient blip are still lost, which is the pain this plan exists to fix\nC) Retry with no idempotency key (as proposed)\n\u2705 No extra storage or key design work up front\n\u274c Duplicate side effects on every timeout-after-success; the at-most-once guarantee silently disappears\nActual answer: A) At-least-once + idempotency key \u2014 user answer to D5\nAccepted scope: processWebhookJob() becomes at-least-once. Before any side effect, atomically claim an idempotency key (provider event ID when present; otherwise SHA-256 of the raw body) via INSERT ... ON CONFLICT DO NOTHING or equivalent; if the claim loses, return success without re-dispatching. Dedup record TTL >= the full retry window. Tests and docs for this contract are carried into Test review (R5) as required proof, not a new decision.\nHistory: original proposal (retry with no idempotency key) superseded by D5.\nState: approved\n\n### R4: Retry policy completeness \u2014 bounds the plan must specify\nFinding: #5, P1, confidence 8/10, PLAN.md:6-8, reviewer: Claude (plan-eng-review)\nPlan baseline: original proposal \u2014 \"full control over the curve\" with no max attempts, delay cap, jitter, dead-letter behavior, or retryable/terminal error classification stated.\nRuntime evidence: unknown \u2014 no worker or library source in the reviewed repo. Web sources (2026) name uncapped delay and missing jitter as the two standard footguns; retry storms and ~12-day delays at attempt 20.\nState: pending\nComparison grid:\n\n| Choice | Current | A | B | C |\n|---|---|---|---|---|\n| R1 retry mechanism | approved: library hooks + shared backoff fn (D4) | fixed | fixed | fixed |\n| R3 webhook semantics | approved: at-least-once + idempotency key (D5) | fixed | fixed | fixed |\n| R4 policy: max attempts | unspecified | 5 (per-job override allowed) | 5 (per-job override allowed) | unspecified |\n| R4 policy: delay cap | unspecified | min(cap, base*2^n), cap = 10 min | unspecified | unspecified |\n| R4 policy: jitter | unspecified | full jitter: random(0, computed delay) | unspecified | unspecified |\n| R4 policy: exhaustion | unspecified | move to library dead-letter/failed set + alert | job marked failed, no alert | unspecified |\n| R4 policy: error classification | unspecified | retry only on transient errors (timeout, 5xx, connection reset); terminal errors (4xx validation, malformed payload) fail immediately | all errors retried | all errors retried |\n| R2 envelope duplication | pending | pending | pending | pending |\n| R5 webhook regression contract | pending | pending | pending | pending |\n| R7 graph recompute per retry | pending | pending | pending | pending |\n\nQuestion D6:\nD6 \u2014 How completely should the plan specify the retry policy (attempts, cap, jitter, exhaustion, error classes)?\nProject/branch/task: main branch; PLAN.md \"Add background job retry framework\", Architecture section.\nELI10: Exponential backoff is only half a policy. Without a cap, attempt 20 waits about 12 days. Without jitter, every job that failed at the same moment retries at the same moment and hammers the downstream again (a retry storm). Without a max, a permanently broken payload retries forever. And retrying a \"400 invalid payload\" is pointless: only transient failures (timeouts, 5xx, dropped connections) deserve a retry. The plan says \"full control over the curve\" but writes none of these down, so each of the five workers will pick its own answer.\nStakes if we pick wrong: a poisoned job retrying forever, or a downstream outage made worse by synchronized retries; both look like mystery load at 3am.\nRecommendation: A because every one of these is a one-line constant or a small predicate in the single shared backoff function you already approved, and the defaults below are the standard ones (adjust the numbers freely via Other).\nCompleteness: A=10/10, B=6/10, C=3/10\nPros / cons:\nA) Full policy: max 5 attempts, 10-min cap, full jitter, dead-letter + alert on exhaustion, retry only transient errors (recommended)\n \u2705 Bounded worst case: a broken job costs at most 5 attempts and ~20 minutes, then lands in a dead-letter set someone gets paged for\n \u2705 Full jitter spreads synchronized failures so a downstream blip does not become a retry storm\n \u2705 Terminal errors fail fast, so bad payloads surface immediately instead of after 5 retries (human: ~1 day / CC: ~15 min)\n \u274c Five constants and one error predicate to agree on now; per-job overrides add a little config surface\nB) Max attempts only (5), everything else unspecified\n \u2705 Prevents infinite retries with one constant\n \u2705 Smallest possible addition to the plan text\n \u274c Uncapped delay and no jitter still allow retry storms and multi-hour waits within those 5 attempts; terminal errors waste retries\nC) Curve only, as proposed\n \u2705 No policy decisions needed before coding starts\n \u274c Each worker author invents their own bounds (or none); unbounded delay and infinite retries are possible\nNet: the difference between A and C is about 15 lines in one shared function versus five different implicit policies you discover in production.\nHeader: Retry policy\nOptions:\nA) Full policy: max 5 attempts, 10-min cap, full jitter, dead-letter + alert on exhaustion, retry only transient errors (recommended)\n\u2705 Bounded worst case: a broken job costs at most 5 attempts and ~20 minutes, then lands in a dead-letter set someone gets paged for\n\u2705 Full jitter spreads synchronized failures so a downstream blip does not become a retry storm\n\u2705 Terminal errors fail fast, so bad payloads surface immediately instead of after 5 retries (human: ~1 day / CC: ~15 min)\n\u274c Five constants and one error predicate to agree on now; per-job overrides add a little config surface\nB) Max attempts only (5), everything else unspecified\n\u2705 Prevents infinite retries with one constant\n\u2705 Smallest possible addition to the plan text\n\u274c Uncapped delay and no jitter still allow retry storms and multi-hour waits within those 5 attempts; terminal errors waste retries\nC) Curve only, as proposed\n\u2705 No policy decisions needed before coding starts\n\u274c Each worker author invents their own bounds (or none); unbounded delay and infinite retries are possible\nActual answer: A) Full policy \u2014 user answer to D6\nAccepted scope: The shared backoff function implements: max 5 attempts (per-job override allowed); delay = min(10 min, base * 2^attempt) with full jitter random(0, delay); on exhaustion move the job to the library's dead-letter/failed set and emit an alert; retry only transient errors (timeout, 5xx, connection reset), fail immediately on terminal errors (4xx validation, malformed payload). Tests proving each bound are required proof carried into Test review (R6).\nHistory: original proposal (curve only, no bounds) superseded by D6.\nState: approved\n\n### R2: Retry envelope duplication across 5 worker files\nFinding: #2, P1, confidence 9/10, PLAN.md:11-13, reviewer: Claude (plan-eng-review)\nPlan baseline: original proposal \u2014 envelope (compute delay, log attempt, dispatch) copy-pasted in 5 worker files; refactor deferred to \"later\".\nRuntime evidence: unknown \u2014 worker files not present in the reviewed repo; the plan itself states the duplication exists (PLAN.md:11-12).\nState: pending\nComparison grid:\n\n| Choice | Current | A | B | C |\n|---|---|---|---|---|\n| R1 retry mechanism | approved: library hooks + shared backoff fn (D4) | fixed | fixed | fixed |\n| R3 webhook semantics | approved: at-least-once + idempotency key (D5) | fixed | fixed | fixed |\n| R4 retry policy | approved: full policy in shared fn (D6) | fixed | fixed | fixed |\n| R2 envelope: where it lives | 5 inline copies | one shared module: backoff fn + attempt-logging middleware registered once with the library | one shared module, but only workers touched by this change migrate | 5 inline copies remain |\n| R2 envelope: workers migrated in this change | 0 | all 5 | only those edited for retries (webhook + any other touched) | 0 |\n| R5 webhook regression contract | pending | pending | pending | pending |\n| R7 graph recompute per retry | pending | pending | pending | pending |\n\nQuestion D7:\nD7 \u2014 Consolidate the retry envelope into one shared module now, or leave five copies and refactor later?\nProject/branch/task: main branch; PLAN.md \"Add background job retry framework\", Code quality section.\nELI10: The same three steps (work out the delay, log the attempt, hand the job back to the queue) are pasted into five worker files. You already decided the delay logic becomes one shared backoff function the library calls, so most of that envelope disappears on its own. What remains is deciding whether the attempt logging and hook registration also live in one place, and whether all five workers move over in this change or only the ones you happen to touch. \"Refactor later\" for copy-pasted code almost always means the first backoff bug gets fixed in four files and missed in the fifth.\nStakes if we pick wrong: a policy fix (say, the jitter formula) lands in some workers and not others, and the inconsistent ones are the ones nobody tests.\nRecommendation: A because with the library-hook decision the shared module is small (one backoff function plus one logging middleware), and migrating five call sites is a mechanical change that CC does in minutes.\nCompleteness: A=10/10, B=7/10, C=3/10\nPros / cons:\nA) One shared retry module, all 5 workers migrated now (recommended)\n \u2705 Exactly one place defines delay, jitter, attempt logging and exhaustion handling; one test suite proves it\n \u2705 Deleting five inline envelopes shrinks the diff and removes the code the plan admits is copy-pasted (human: ~1.5 days / CC: ~25 min)\n \u274c Touches all five worker files in one change, so the review surface is wider than the webhook rewrite alone\nB) Shared module, migrate only workers this change already touches\n \u2705 Smaller immediate diff; untouched workers carry no risk from this change\n \u2705 The shared module exists, so later migrations are copy-and-delete\n \u274c Two retry implementations coexist in production; the inline ones keep their unbounded, unjittered behavior until someone remembers\nC) Leave five copies, refactor \"later\" (as proposed)\n \u2705 No refactor risk in this change\n \u274c Five places to apply every future policy fix; the approved policy (D6) has to be pasted five times to take effect\nNet: the shared module is the approved design's natural home; the only real question is whether all five workers adopt it now or drift.\nHeader: DRY envelope\nOptions:\nA) One shared retry module, all 5 workers migrated now (recommended)\n\u2705 Exactly one place defines delay, jitter, attempt logging and exhaustion handling; one test suite proves it\n\u2705 Deleting five inline envelopes shrinks the diff and removes the code the plan admits is copy-pasted (human: ~1.5 days / CC: ~25 min)\n\u274c Touches all five worker files in one change, so the review surface is wider than the webhook rewrite alone\nB) Shared module, migrate only workers this change already touches\n\u2705 Smaller immediate diff; untouched workers carry no risk from this change\n\u2705 The shared module exists, so later migrations are copy-and-delete\n\u274c Two retry implementations coexist in production; the inline ones keep their unbounded, unjittered behavior until someone remembers\nC) Leave five copies, refactor \"later\" (as proposed)\n\u2705 No refactor risk in this change\n\u274c Five places to apply every future policy fix; the approved policy (D6) has to be pasted five times to take effect\nActual answer: A) One shared retry module, all 5 workers migrated now \u2014 user answer to D7\nAccepted scope: Create one shared retry module (e.g. `jobs/retry/` or the project's equivalent): the backoff/policy function from D6 plus an attempt-logging middleware (job id, job type, attempt n/max, error class, next delay ms) registered once with the job library. Delete the inline envelope from all 5 worker files and register them with the shared module. Existing worker behavior other than retry scheduling is unchanged.\nHistory: original proposal (leave 5 copies, refactor later) superseded by D7.\nState: approved\n\n### R5: Regression contract for the `processWebhookJob()` rewrite\nFinding: #3, P0, confidence 9/10, PLAN.md:16-18, reviewer: Claude (plan-eng-review)\nPlan baseline: original proposal \u2014 \"No regression test for the prior at-most-once delivery guarantee is planned.\" Approved behavior to protect comes from D5 (at-least-once + idempotency key) and D6 (retry only transient errors).\nRuntime evidence: unknown \u2014 no test framework or test files exist in the reviewed repo (git ls-files: CLAUDE.md, PLAN.md only). Existing callers of processWebhookJob() cannot be traced here; treat every current behavior as at risk.\nState: pending\nComparison grid:\n\n| Choice | Current | A | B |\n|---|---|---|---|\n| R1, R3, R4, R2 | approved (D4-D7) | fixed | fixed |\n| R5 behavior preserved: success dispatches exactly once with unchanged payload/headers | untested | asserted | asserted |\n| R5 behavior preserved: terminal error (4xx/malformed) \u2192 exactly one attempt, job failed, no retry (old at-most-once holds for this class) | untested | asserted | not asserted |\n| R5 intentional change: transient error \u2192 retried per D6 policy, attempt count visible | untested | asserted | asserted |\n| R5 new contract: second run with an already-claimed key \u2192 no dispatch, returns success | untested | asserted | asserted |\n| R5 new contract: key derivation \u2014 provider event ID when present, body hash when absent | untested | asserted (both branches) | not asserted |\n| R5 new contract: crash/timeout after dispatch, before ack \u2192 retry finds claimed key \u2192 zero duplicate sends [\u2192E2E] | untested | asserted via integration test against the real queue | not asserted |\n| R5: dedup TTL >= retry window | untested | asserted | not asserted |\n| R6 policy proof depth | pending | pending | pending |\n| R7 graph recompute | pending | pending | pending |\n\nQuestion D8:\nD8 \u2014 What regression contract protects processWebhookJob() through the rewrite?\nProject/branch/task: main branch; PLAN.md \"Add background job retry framework\", Tests section.\nELI10: The webhook job is being rewritten and currently has no test guarding what it does today. A regression contract is the written list of \"what must still be true afterwards\", plus the things we are changing on purpose. For this job that is: one successful run still sends exactly one webhook with the same content; a permanently bad request still fails once and stops; a transient failure now retries (the intentional change); and a retry after a timeout-that-actually-succeeded sends nothing twice. The last one is the whole point of the idempotency key you approved, and it only shows up in a test that runs against the real queue, not a mock.\nStakes if we pick wrong: the rewrite ships, a receiver gets two payment webhooks after a slow response, and there was never a test that could have caught it.\nRecommendation: A because these seven assertions are the entire behavioral surface of the job, and the ones B drops (terminal fail-fast, key fallback, crash-after-dispatch) are exactly the failure paths that produce duplicates.\nCompleteness: A=10/10, B=6/10\nPros / cons:\nA) Full contract: 7 assertions incl. terminal fail-fast, key fallback and crash-after-dispatch integration test (recommended)\n \u2705 Every branch of the new job has a named assertion; the duplicate-send path is proven against the real queue, not a mock\n \u2705 Preserves the old at-most-once behavior where it still applies (terminal errors) and documents the intentional change for transient ones\n \u2705 Tests double as the spec for the idempotency key and TTL (human: ~2 days / CC: ~30 min)\n \u274c One integration test needs a real or in-memory queue backend in CI\nB) Happy path + duplicate-claim unit tests only\n \u2705 Fast, mock-only tests; no CI infrastructure change\n \u2705 Covers the most common two paths\n \u274c Misses terminal-error fail-fast, body-hash key fallback, TTL, and the crash-after-dispatch path that actually causes duplicates\nNet: A costs one integration test fixture; B leaves the duplicate-send path, the reason D5 exists, unproven.\nHeader: Webhook tests\nOptions:\nA) Full contract: 7 assertions incl. terminal fail-fast, key fallback and crash-after-dispatch integration test (recommended)\n\u2705 Every branch of the new job has a named assertion; the duplicate-send path is proven against the real queue, not a mock\n\u2705 Preserves the old at-most-once behavior where it still applies (terminal errors) and documents the intentional change for transient ones\n\u2705 Tests double as the spec for the idempotency key and TTL (human: ~2 days / CC: ~30 min)\n\u274c One integration test needs a real or in-memory queue backend in CI\nB) Happy path + duplicate-claim unit tests only\n\u2705 Fast, mock-only tests; no CI infrastructure change\n\u2705 Covers the most common two paths\n\u274c Misses terminal-error fail-fast, body-hash key fallback, TTL, and the crash-after-dispatch path that actually causes duplicates\nActual answer: unanswered\nAccepted scope: none\nHistory: none\n", "snapshot": { "at": "2026-09-16T06:01:25.858Z", "kind": "owned-plan-or-report", "source": "/tmp/g-pyj7vm4y/gstack-paid-shard-qDJLTL/tmp/gstack-e2e-plan-eng-batching-IpzCNz/gstack-test-plan-eng-batching.md", "artifact": "objects/2ed153e59fb97fe1c38f20e927851887957aa7d7357cbdd8b056830e42e20061.md", "sha256": "2ed153e59fb97fe1c38f20e927851887957aa7d7357cbdd8b056830e42e20061", "bytes": 28627, "mtimeMs": 1789538484952.3064, "provenance": "Exact observed file bytes; never reconstructed from tool text." } }, { "fingerprint": { "signature": "8c2809a2-42f4-479b-a5e4-a6f2e5016e61:toolu_014uRBqjs6e1n19DrHNp1TQX", "promptSnippet": "Policy tests D9 \u2014 How deeply do we prove the shared retry policy and the five worker migrations? Project/branch/task: main branch; PLAN.md \"Add background job retry framework\", Tests section. ELI10: You already approved unit tests for each ", "options": [ { "index": 1, "label": "Per-bound units + property test + 5 per-worker integration tests + exhaustion E2E (recommended)" }, { "index": 2, "label": "Per-bound unit tests only (as carried from D6)" } ], "observedAtMs": 567840, "preReview": true, "nativeCall": { "sessionId": "8c2809a2-42f4-479b-a5e4-a6f2e5016e61", "toolUseId": "toolu_014uRBqjs6e1n19DrHNp1TQX", "questions": [ { "question": "D9 \u2014 How deeply do we prove the shared retry policy and the five worker migrations?\nProject/branch/task: main branch; PLAN.md \"Add background job retry framework\", Tests section.\nELI10: You already approved unit tests for each policy bound. Two extra layers are on the table. First, a property test: instead of checking attempt 1, 2, 3 by hand, run the delay function for every attempt from 0 to 50 and assert it never exceeds the cap, never goes negative, and never overflows (2^50 milliseconds is a real number that bites). Second, integration tests: for each of the five workers, force one transient failure and confirm the queue library actually reschedules it through the shared policy and logs the attempt. That last one is what proves the migration did anything; a unit test on the policy function cannot tell you a worker forgot to register it.\nStakes if we pick wrong: a worker silently keeps no-retry behavior after \"migration\", or the delay math overflows at high attempt counts, and nothing in CI notices.\nRecommendation: A because the property test is ~10 lines and the five integration tests share one fixture with the D8 crash-after-dispatch test you already approved, so the marginal CC cost is minutes.\nCompleteness: A=10/10, B=7/10\nNet: B proves the math; A proves the math and that production uses it.", "header": "Policy tests", "multiSelect": false, "options": [ { "label": "Per-bound units + property test + 5 per-worker integration tests + exhaustion E2E (recommended)", "description": "\u2705 Proves each worker is actually wired to the shared policy, not just that the policy function is correct\n\u2705 Property test catches cap/overflow/negative-delay bugs across the whole attempt range in one assertion\n\u2705 Exhaustion E2E proves the dead-letter + alert path fires exactly once (human: ~2 days / CC: ~30 min)\n\u274c Six more integration tests in CI on top of D8's fixture; slower suite by seconds" }, { "label": "Per-bound unit tests only (as carried from D6)", "description": "\u2705 Fast, no queue backend needed for this module\n\u2705 Already covers every named bound at specific attempts\n\u274c Cannot detect an unmigrated worker, a broken hook registration, or overflow at attempts you did not hand-pick" } ] } ], "answered": true, "failed": false, "answers": { "D9 \u2014 How deeply do we prove the shared retry policy and the five worker migrations?\nProject/branch/task: main branch; PLAN.md \"Add background job retry framework\", Tests section.\nELI10: You already approved unit tests for each policy bound. Two extra layers are on the table. First, a property test: instead of checking attempt 1, 2, 3 by hand, run the delay function for every attempt from 0 to 50 and assert it never exceeds the cap, never goes negative, and never overflows (2^50 milliseconds is a real number that bites). Second, integration tests: for each of the five workers, force one transient failure and confirm the queue library actually reschedules it through the shared policy and logs the attempt. That last one is what proves the migration did anything; a unit test on the policy function cannot tell you a worker forgot to register it.\nStakes if we pick wrong: a worker silently keeps no-retry behavior after \"migration\", or the delay math overflows at high attempt counts, and nothing in CI notices.\nRecommendation: A because the property test is ~10 lines and the five integration tests share one fixture with the D8 crash-after-dispatch test you already approved, so the marginal CC cost is minutes.\nCompleteness: A=10/10, B=7/10\nNet: B proves the math; A proves the math and that production uses it.": "Per-bound units + property test + 5 per-worker integration tests + exhaustion E2E (recommended)" }, "unansweredQuestionIndices": [], "answeredAt": "2026-09-16T06:02:40.263Z" } }, "questionAt": "2026-09-16T06:02:38.731Z", "answeredAt": "2026-09-16T06:02:40.263Z", "preAskPlan": "# Eng Review: Add background job retry framework\n\nReviewed target: `PLAN.md` (repo root, commit 7ce6818, branch main).\nReview: `/plan-eng-review`, 2026-09-16. Report file requested by the user.\n\n## Original plan (as reviewed)\n\n> ## Architecture\n> We'll roll a custom exponential-backoff scheduler inline in each worker\n> rather than use the existing job library's built-in retry hooks. Same\n> shape as the library version, but we want full control over the curve.\n>\n> ## Code quality\n> The retry envelope (compute delay, log attempt, dispatch) is duplicated\n> across 5 worker files with copy-pasted bodies. We will leave the\n> duplication for now and refactor \"later.\"\n>\n> ## Tests\n> The existing `processWebhookJob()` flow gets rewritten as part of this\n> change. No regression test for the prior at-most-once delivery guarantee\n> is planned.\n>\n> ## Performance\n> On every retry we re-fetch the full job payload from the database, then\n> iterate the payload to recompute the dependency graph. Could cache the\n> graph on the first attempt; not planned.\n\n## Step 0: Scope Challenge findings\n\n1. [P1] (confidence: 8/10) PLAN.md:6-8 \u2014 custom inline scheduler rebuilds the job library's built-in retry hooks (Layer 1). Custom backoff curves are a supported extension point. Disposition: pending (R1 / D4).\n2. [P1] (confidence: 9/10) PLAN.md:11-13 \u2014 retry envelope copy-pasted across 5 worker files; refactor deferred. Disposition: pending (Code Quality).\n3. [P0] (confidence: 9/10) PLAN.md:16-18 \u2014 `processWebhookJob()` rewritten with no regression test; retries change the guarantee from at-most-once to at-least-once with no idempotency key. Disposition: pending (Architecture + Tests).\n4. [P2] (confidence: 7/10) PLAN.md:21-23 \u2014 payload re-fetched and dependency graph recomputed on every retry. Disposition: pending (Performance).\n5. [P1] (confidence: 8/10) PLAN.md:6-8 \u2014 \"full control over the curve\" but no curve specified: no max attempts, delay cap, jitter, dead-letter, or retryable-error classification. Disposition: pending (Architecture).\n6. Factual gap: no design doc / problem statement (which failure classes retries target). Recorded; folded into finding 5.\n\nComplexity gate: not triggered (5 worker files + 1 rewrite, 0 new classes). TODOS.md: absent. Distribution: no new artifact.\n\n## Decision ledger\n\n### R1: Retry mechanism \u2014 library built-in retry hooks vs custom inline scheduler\nFinding: #1, P1, confidence 8/10, PLAN.md:6-8, reviewer: Claude (plan-eng-review)\nPlan baseline: original proposal \u2014 custom exponential-backoff scheduler inline in each worker; library hooks unused.\nRuntime evidence: unknown \u2014 the repo under review contains only PLAN.md and CLAUDE.md; the job library and worker files are not present. Web search (2026) shows mainstream job libraries expose a custom backoff function / retry_in hook.\nState: approved (D4) \u2014 see Actual answer below\nComparison grid:\n\n| Choice | Current | A | B |\n|---|---|---|---|\n| R1 retry mechanism | custom inline scheduler per worker (proposed) | library built-in retry hooks with one shared custom backoff function | custom inline scheduler per worker (as proposed) |\n| R2 envelope duplication | 5 copies, deferred (pending) | pending | pending |\n| R3 webhook delivery semantics | unspecified (pending) | pending | pending |\n| R4 retry policy bounds | unspecified (pending) | pending | pending |\n| R5 webhook regression contract | none planned (pending) | pending | pending |\n| R7 graph recompute per retry | recompute every retry (pending) | pending | pending |\n\nQuestion D4:\nD4 \u2014 Retry mechanism: use the job library's built-in retry hooks, or roll a custom inline scheduler?\nProject/branch/task: main branch; PLAN.md \"Add background job retry framework\", Architecture section.\nELI10: The plan rebuilds retry scheduling inside each worker even though the job library already has retry hooks. The stated reason is \"full control over the curve\", but these libraries let you plug in your own backoff function, so you get the exact curve you want while the library keeps doing the hard parts: persisting attempt counts across worker crashes, scheduling the delayed re-run, and exposing retries to its dashboard and metrics. A hand-rolled inline scheduler has to re-solve all of that in five places.\nStakes if we pick wrong: a worker that crashes mid-sleep in a custom scheduler loses the job silently; a library-managed retry survives the crash and is visible in the queue UI.\nRecommendation: A because the library hook gives the same curve control with crash-safe attempt state and zero new scheduling code; the custom path only wins if the library has no backoff hook at all, which the plan does not claim.\nCompleteness: A=10/10, B=6/10\nPros / cons:\nA) Library retry hooks + one custom backoff function (recommended)\n \u2705 Attempt count and next-run time persist in the queue, so a worker crash mid-backoff cannot lose the job\n \u2705 Retries show up in the library's dashboard, metrics and dead-letter handling for free\n \u2705 One backoff function to test instead of five inline schedulers (human: ~1 day / CC: ~20 min)\n \u274c Curve is bounded by whatever the hook API allows (delay per attempt); exotic behavior like per-tenant curves needs a wrapper\nB) Custom inline scheduler in each worker (as proposed)\n \u2705 Unconstrained control over every retry decision inside worker code\n \u2705 No dependency on the library's hook signature or its upgrade path\n \u274c Attempt state lives in process memory unless you build persistence; a crash or deploy mid-backoff drops the retry\n \u274c Five schedulers to keep in sync, invisible to the queue UI and metrics (human: ~3 days / CC: ~45 min)\nNet: same curve either way; the trade is crash-safety and observability for free vs. rebuilding the library's job.\nHeader: Retry mech\nOptions:\nA) Library retry hooks + one custom backoff function (recommended)\n\u2705 Attempt count and next-run time persist in the queue, so a worker crash mid-backoff cannot lose the job\n\u2705 Retries show up in the library's dashboard, metrics and dead-letter handling for free\n\u2705 One backoff function to test instead of five inline schedulers (human: ~1 day / CC: ~20 min)\n\u274c Curve is bounded by whatever the hook API allows (delay per attempt); exotic behavior like per-tenant curves needs a wrapper\nB) Custom inline scheduler in each worker (as proposed)\n\u2705 Unconstrained control over every retry decision inside worker code\n\u2705 No dependency on the library's hook signature or its upgrade path\n\u274c Attempt state lives in process memory unless you build persistence; a crash or deploy mid-backoff drops the retry\n\u274c Five schedulers to keep in sync, invisible to the queue UI and metrics (human: ~3 days / CC: ~45 min)\nActual answer: A) Library retry hooks + one custom backoff function \u2014 user answer to D4\nAccepted scope: Replace the custom inline scheduler with the job library's built-in retry hooks; implement one shared custom backoff function that the hooks call; delete the plan's inline per-worker scheduler. Scope Challenge result: scope reduced per recommendation.\nHistory: original proposal (custom inline scheduler) superseded by D4.\nState: approved\n\n### R3: Delivery semantics for `processWebhookJob()` under retries\nFinding: #3, P0, confidence 9/10, PLAN.md:16-18, reviewer: Claude (plan-eng-review)\nPlan baseline: original proposal \u2014 processWebhookJob() is rewritten to retry; the prior at-most-once guarantee is neither preserved nor replaced with an explicit alternative.\nRuntime evidence: unknown \u2014 worker source not present in the reviewed repo. The plan itself states the current guarantee is at-most-once (PLAN.md:17). Web sources (Svix, Hookdeck, 2026) agree that retries imply at-least-once and require an idempotency key to be safe.\nState: approved (D5) \u2014 see Actual answer below\nComparison grid:\n\n| Choice | Current | A | B | C |\n|---|---|---|---|---|\n| R1 retry mechanism | approved: library hooks + shared backoff fn (D4) | fixed | fixed | fixed |\n| R3 webhook delivery semantics | at-most-once today; plan retries without stating new guarantee | at-least-once with idempotency key (provider event ID, or hash of raw body) claimed atomically before side effects; duplicates return success | keep at-most-once: exclude processWebhookJob from retries (max attempts = 1); other 4 workers retry | retry with no idempotency key (as proposed) |\n| R2 envelope duplication | pending | pending | pending | pending |\n| R4 retry policy bounds | pending | pending | pending | pending |\n| R5 webhook regression contract | pending | pending | pending | pending |\n| R7 graph recompute per retry | pending | pending | pending | pending |\n\nQuestion D5:\nD5 \u2014 When processWebhookJob() gains retries, what delivery guarantee does it make?\nProject/branch/task: main branch; PLAN.md \"Add background job retry framework\", Tests/Architecture sections.\nELI10: Today the webhook job runs at most once: if it fails, nothing is redelivered, so it can never double-fire. The moment you add retries, a job that timed out after the receiver already processed it will run again, and the receiver sees the same webhook twice. That is a real behavior change, not just a missing test. The standard fix is an idempotency key: record the event ID before doing the side effect, and if the key is already there, treat the retry as already done.\nStakes if we pick wrong: duplicate webhooks mean double charges, double emails, or double state changes for whoever receives them, and you find out from a customer, not a test.\nRecommendation: A because retrying is the point of this plan, and an idempotency key is a few lines plus one unique index; it turns duplicates from an incident into a no-op.\nCompleteness: A=10/10, B=7/10, C=3/10\nPros / cons:\nA) At-least-once + idempotency key (recommended)\n \u2705 Transient failures get retried and duplicates are absorbed, so receivers never see a double effect\n \u2705 Small change: atomic INSERT ... ON CONFLICT DO NOTHING (or SET NX) on the event ID before dispatch (human: ~1 day / CC: ~15 min)\n \u274c Needs a dedup store with a TTL at least as long as the retry window, and a decision on what the key is when the payload has no natural ID\nB) Keep at-most-once: no retries for the webhook job\n \u2705 Zero semantic change for webhook receivers; regression risk drops to the rewrite itself\n \u2705 Simplest to prove: max attempts = 1 for this job, existing behavior stays\n \u274c Webhook deliveries that fail on a transient blip are still lost, which is the pain this plan exists to fix\nC) Retry with no idempotency key (as proposed)\n \u2705 No extra storage or key design work up front\n \u274c Duplicate side effects on every timeout-after-success; the at-most-once guarantee silently disappears\nNet: you are choosing between fixing lost webhooks safely, not fixing them, or fixing them in a way that creates duplicates.\nHeader: Webhook sem\nOptions:\nA) At-least-once + idempotency key (recommended)\n\u2705 Transient failures get retried and duplicates are absorbed, so receivers never see a double effect\n\u2705 Small change: atomic INSERT ... ON CONFLICT DO NOTHING (or SET NX) on the event ID before dispatch (human: ~1 day / CC: ~15 min)\n\u274c Needs a dedup store with a TTL at least as long as the retry window, and a decision on what the key is when the payload has no natural ID\nB) Keep at-most-once: no retries for the webhook job\n\u2705 Zero semantic change for webhook receivers; regression risk drops to the rewrite itself\n\u2705 Simplest to prove: max attempts = 1 for this job, existing behavior stays\n\u274c Webhook deliveries that fail on a transient blip are still lost, which is the pain this plan exists to fix\nC) Retry with no idempotency key (as proposed)\n\u2705 No extra storage or key design work up front\n\u274c Duplicate side effects on every timeout-after-success; the at-most-once guarantee silently disappears\nActual answer: A) At-least-once + idempotency key \u2014 user answer to D5\nAccepted scope: processWebhookJob() becomes at-least-once. Before any side effect, atomically claim an idempotency key (provider event ID when present; otherwise SHA-256 of the raw body) via INSERT ... ON CONFLICT DO NOTHING or equivalent; if the claim loses, return success without re-dispatching. Dedup record TTL >= the full retry window. Tests and docs for this contract are carried into Test review (R5) as required proof, not a new decision.\nHistory: original proposal (retry with no idempotency key) superseded by D5.\nState: approved\n\n### R4: Retry policy completeness \u2014 bounds the plan must specify\nFinding: #5, P1, confidence 8/10, PLAN.md:6-8, reviewer: Claude (plan-eng-review)\nPlan baseline: original proposal \u2014 \"full control over the curve\" with no max attempts, delay cap, jitter, dead-letter behavior, or retryable/terminal error classification stated.\nRuntime evidence: unknown \u2014 no worker or library source in the reviewed repo. Web sources (2026) name uncapped delay and missing jitter as the two standard footguns; retry storms and ~12-day delays at attempt 20.\nState: pending\nComparison grid:\n\n| Choice | Current | A | B | C |\n|---|---|---|---|---|\n| R1 retry mechanism | approved: library hooks + shared backoff fn (D4) | fixed | fixed | fixed |\n| R3 webhook semantics | approved: at-least-once + idempotency key (D5) | fixed | fixed | fixed |\n| R4 policy: max attempts | unspecified | 5 (per-job override allowed) | 5 (per-job override allowed) | unspecified |\n| R4 policy: delay cap | unspecified | min(cap, base*2^n), cap = 10 min | unspecified | unspecified |\n| R4 policy: jitter | unspecified | full jitter: random(0, computed delay) | unspecified | unspecified |\n| R4 policy: exhaustion | unspecified | move to library dead-letter/failed set + alert | job marked failed, no alert | unspecified |\n| R4 policy: error classification | unspecified | retry only on transient errors (timeout, 5xx, connection reset); terminal errors (4xx validation, malformed payload) fail immediately | all errors retried | all errors retried |\n| R2 envelope duplication | pending | pending | pending | pending |\n| R5 webhook regression contract | pending | pending | pending | pending |\n| R7 graph recompute per retry | pending | pending | pending | pending |\n\nQuestion D6:\nD6 \u2014 How completely should the plan specify the retry policy (attempts, cap, jitter, exhaustion, error classes)?\nProject/branch/task: main branch; PLAN.md \"Add background job retry framework\", Architecture section.\nELI10: Exponential backoff is only half a policy. Without a cap, attempt 20 waits about 12 days. Without jitter, every job that failed at the same moment retries at the same moment and hammers the downstream again (a retry storm). Without a max, a permanently broken payload retries forever. And retrying a \"400 invalid payload\" is pointless: only transient failures (timeouts, 5xx, dropped connections) deserve a retry. The plan says \"full control over the curve\" but writes none of these down, so each of the five workers will pick its own answer.\nStakes if we pick wrong: a poisoned job retrying forever, or a downstream outage made worse by synchronized retries; both look like mystery load at 3am.\nRecommendation: A because every one of these is a one-line constant or a small predicate in the single shared backoff function you already approved, and the defaults below are the standard ones (adjust the numbers freely via Other).\nCompleteness: A=10/10, B=6/10, C=3/10\nPros / cons:\nA) Full policy: max 5 attempts, 10-min cap, full jitter, dead-letter + alert on exhaustion, retry only transient errors (recommended)\n \u2705 Bounded worst case: a broken job costs at most 5 attempts and ~20 minutes, then lands in a dead-letter set someone gets paged for\n \u2705 Full jitter spreads synchronized failures so a downstream blip does not become a retry storm\n \u2705 Terminal errors fail fast, so bad payloads surface immediately instead of after 5 retries (human: ~1 day / CC: ~15 min)\n \u274c Five constants and one error predicate to agree on now; per-job overrides add a little config surface\nB) Max attempts only (5), everything else unspecified\n \u2705 Prevents infinite retries with one constant\n \u2705 Smallest possible addition to the plan text\n \u274c Uncapped delay and no jitter still allow retry storms and multi-hour waits within those 5 attempts; terminal errors waste retries\nC) Curve only, as proposed\n \u2705 No policy decisions needed before coding starts\n \u274c Each worker author invents their own bounds (or none); unbounded delay and infinite retries are possible\nNet: the difference between A and C is about 15 lines in one shared function versus five different implicit policies you discover in production.\nHeader: Retry policy\nOptions:\nA) Full policy: max 5 attempts, 10-min cap, full jitter, dead-letter + alert on exhaustion, retry only transient errors (recommended)\n\u2705 Bounded worst case: a broken job costs at most 5 attempts and ~20 minutes, then lands in a dead-letter set someone gets paged for\n\u2705 Full jitter spreads synchronized failures so a downstream blip does not become a retry storm\n\u2705 Terminal errors fail fast, so bad payloads surface immediately instead of after 5 retries (human: ~1 day / CC: ~15 min)\n\u274c Five constants and one error predicate to agree on now; per-job overrides add a little config surface\nB) Max attempts only (5), everything else unspecified\n\u2705 Prevents infinite retries with one constant\n\u2705 Smallest possible addition to the plan text\n\u274c Uncapped delay and no jitter still allow retry storms and multi-hour waits within those 5 attempts; terminal errors waste retries\nC) Curve only, as proposed\n\u2705 No policy decisions needed before coding starts\n\u274c Each worker author invents their own bounds (or none); unbounded delay and infinite retries are possible\nActual answer: A) Full policy \u2014 user answer to D6\nAccepted scope: The shared backoff function implements: max 5 attempts (per-job override allowed); delay = min(10 min, base * 2^attempt) with full jitter random(0, delay); on exhaustion move the job to the library's dead-letter/failed set and emit an alert; retry only transient errors (timeout, 5xx, connection reset), fail immediately on terminal errors (4xx validation, malformed payload). Tests proving each bound are required proof carried into Test review (R6).\nHistory: original proposal (curve only, no bounds) superseded by D6.\nState: approved\n\n### R2: Retry envelope duplication across 5 worker files\nFinding: #2, P1, confidence 9/10, PLAN.md:11-13, reviewer: Claude (plan-eng-review)\nPlan baseline: original proposal \u2014 envelope (compute delay, log attempt, dispatch) copy-pasted in 5 worker files; refactor deferred to \"later\".\nRuntime evidence: unknown \u2014 worker files not present in the reviewed repo; the plan itself states the duplication exists (PLAN.md:11-12).\nState: pending\nComparison grid:\n\n| Choice | Current | A | B | C |\n|---|---|---|---|---|\n| R1 retry mechanism | approved: library hooks + shared backoff fn (D4) | fixed | fixed | fixed |\n| R3 webhook semantics | approved: at-least-once + idempotency key (D5) | fixed | fixed | fixed |\n| R4 retry policy | approved: full policy in shared fn (D6) | fixed | fixed | fixed |\n| R2 envelope: where it lives | 5 inline copies | one shared module: backoff fn + attempt-logging middleware registered once with the library | one shared module, but only workers touched by this change migrate | 5 inline copies remain |\n| R2 envelope: workers migrated in this change | 0 | all 5 | only those edited for retries (webhook + any other touched) | 0 |\n| R5 webhook regression contract | pending | pending | pending | pending |\n| R7 graph recompute per retry | pending | pending | pending | pending |\n\nQuestion D7:\nD7 \u2014 Consolidate the retry envelope into one shared module now, or leave five copies and refactor later?\nProject/branch/task: main branch; PLAN.md \"Add background job retry framework\", Code quality section.\nELI10: The same three steps (work out the delay, log the attempt, hand the job back to the queue) are pasted into five worker files. You already decided the delay logic becomes one shared backoff function the library calls, so most of that envelope disappears on its own. What remains is deciding whether the attempt logging and hook registration also live in one place, and whether all five workers move over in this change or only the ones you happen to touch. \"Refactor later\" for copy-pasted code almost always means the first backoff bug gets fixed in four files and missed in the fifth.\nStakes if we pick wrong: a policy fix (say, the jitter formula) lands in some workers and not others, and the inconsistent ones are the ones nobody tests.\nRecommendation: A because with the library-hook decision the shared module is small (one backoff function plus one logging middleware), and migrating five call sites is a mechanical change that CC does in minutes.\nCompleteness: A=10/10, B=7/10, C=3/10\nPros / cons:\nA) One shared retry module, all 5 workers migrated now (recommended)\n \u2705 Exactly one place defines delay, jitter, attempt logging and exhaustion handling; one test suite proves it\n \u2705 Deleting five inline envelopes shrinks the diff and removes the code the plan admits is copy-pasted (human: ~1.5 days / CC: ~25 min)\n \u274c Touches all five worker files in one change, so the review surface is wider than the webhook rewrite alone\nB) Shared module, migrate only workers this change already touches\n \u2705 Smaller immediate diff; untouched workers carry no risk from this change\n \u2705 The shared module exists, so later migrations are copy-and-delete\n \u274c Two retry implementations coexist in production; the inline ones keep their unbounded, unjittered behavior until someone remembers\nC) Leave five copies, refactor \"later\" (as proposed)\n \u2705 No refactor risk in this change\n \u274c Five places to apply every future policy fix; the approved policy (D6) has to be pasted five times to take effect\nNet: the shared module is the approved design's natural home; the only real question is whether all five workers adopt it now or drift.\nHeader: DRY envelope\nOptions:\nA) One shared retry module, all 5 workers migrated now (recommended)\n\u2705 Exactly one place defines delay, jitter, attempt logging and exhaustion handling; one test suite proves it\n\u2705 Deleting five inline envelopes shrinks the diff and removes the code the plan admits is copy-pasted (human: ~1.5 days / CC: ~25 min)\n\u274c Touches all five worker files in one change, so the review surface is wider than the webhook rewrite alone\nB) Shared module, migrate only workers this change already touches\n\u2705 Smaller immediate diff; untouched workers carry no risk from this change\n\u2705 The shared module exists, so later migrations are copy-and-delete\n\u274c Two retry implementations coexist in production; the inline ones keep their unbounded, unjittered behavior until someone remembers\nC) Leave five copies, refactor \"later\" (as proposed)\n\u2705 No refactor risk in this change\n\u274c Five places to apply every future policy fix; the approved policy (D6) has to be pasted five times to take effect\nActual answer: A) One shared retry module, all 5 workers migrated now \u2014 user answer to D7\nAccepted scope: Create one shared retry module (e.g. `jobs/retry/` or the project's equivalent): the backoff/policy function from D6 plus an attempt-logging middleware (job id, job type, attempt n/max, error class, next delay ms) registered once with the job library. Delete the inline envelope from all 5 worker files and register them with the shared module. Existing worker behavior other than retry scheduling is unchanged.\nHistory: original proposal (leave 5 copies, refactor later) superseded by D7.\nState: approved\n\n### R5: Regression contract for the `processWebhookJob()` rewrite\nFinding: #3, P0, confidence 9/10, PLAN.md:16-18, reviewer: Claude (plan-eng-review)\nPlan baseline: original proposal \u2014 \"No regression test for the prior at-most-once delivery guarantee is planned.\" Approved behavior to protect comes from D5 (at-least-once + idempotency key) and D6 (retry only transient errors).\nRuntime evidence: unknown \u2014 no test framework or test files exist in the reviewed repo (git ls-files: CLAUDE.md, PLAN.md only). Existing callers of processWebhookJob() cannot be traced here; treat every current behavior as at risk.\nState: pending\nComparison grid:\n\n| Choice | Current | A | B |\n|---|---|---|---|\n| R1, R3, R4, R2 | approved (D4-D7) | fixed | fixed |\n| R5 behavior preserved: success dispatches exactly once with unchanged payload/headers | untested | asserted | asserted |\n| R5 behavior preserved: terminal error (4xx/malformed) \u2192 exactly one attempt, job failed, no retry (old at-most-once holds for this class) | untested | asserted | not asserted |\n| R5 intentional change: transient error \u2192 retried per D6 policy, attempt count visible | untested | asserted | asserted |\n| R5 new contract: second run with an already-claimed key \u2192 no dispatch, returns success | untested | asserted | asserted |\n| R5 new contract: key derivation \u2014 provider event ID when present, body hash when absent | untested | asserted (both branches) | not asserted |\n| R5 new contract: crash/timeout after dispatch, before ack \u2192 retry finds claimed key \u2192 zero duplicate sends [\u2192E2E] | untested | asserted via integration test against the real queue | not asserted |\n| R5: dedup TTL >= retry window | untested | asserted | not asserted |\n| R6 policy proof depth | pending | pending | pending |\n| R7 graph recompute | pending | pending | pending |\n\nQuestion D8:\nD8 \u2014 What regression contract protects processWebhookJob() through the rewrite?\nProject/branch/task: main branch; PLAN.md \"Add background job retry framework\", Tests section.\nELI10: The webhook job is being rewritten and currently has no test guarding what it does today. A regression contract is the written list of \"what must still be true afterwards\", plus the things we are changing on purpose. For this job that is: one successful run still sends exactly one webhook with the same content; a permanently bad request still fails once and stops; a transient failure now retries (the intentional change); and a retry after a timeout-that-actually-succeeded sends nothing twice. The last one is the whole point of the idempotency key you approved, and it only shows up in a test that runs against the real queue, not a mock.\nStakes if we pick wrong: the rewrite ships, a receiver gets two payment webhooks after a slow response, and there was never a test that could have caught it.\nRecommendation: A because these seven assertions are the entire behavioral surface of the job, and the ones B drops (terminal fail-fast, key fallback, crash-after-dispatch) are exactly the failure paths that produce duplicates.\nCompleteness: A=10/10, B=6/10\nPros / cons:\nA) Full contract: 7 assertions incl. terminal fail-fast, key fallback and crash-after-dispatch integration test (recommended)\n \u2705 Every branch of the new job has a named assertion; the duplicate-send path is proven against the real queue, not a mock\n \u2705 Preserves the old at-most-once behavior where it still applies (terminal errors) and documents the intentional change for transient ones\n \u2705 Tests double as the spec for the idempotency key and TTL (human: ~2 days / CC: ~30 min)\n \u274c One integration test needs a real or in-memory queue backend in CI\nB) Happy path + duplicate-claim unit tests only\n \u2705 Fast, mock-only tests; no CI infrastructure change\n \u2705 Covers the most common two paths\n \u274c Misses terminal-error fail-fast, body-hash key fallback, TTL, and the crash-after-dispatch path that actually causes duplicates\nNet: A costs one integration test fixture; B leaves the duplicate-send path, the reason D5 exists, unproven.\nHeader: Webhook tests\nOptions:\nA) Full contract: 7 assertions incl. terminal fail-fast, key fallback and crash-after-dispatch integration test (recommended)\n\u2705 Every branch of the new job has a named assertion; the duplicate-send path is proven against the real queue, not a mock\n\u2705 Preserves the old at-most-once behavior where it still applies (terminal errors) and documents the intentional change for transient ones\n\u2705 Tests double as the spec for the idempotency key and TTL (human: ~2 days / CC: ~30 min)\n\u274c One integration test needs a real or in-memory queue backend in CI\nB) Happy path + duplicate-claim unit tests only\n\u2705 Fast, mock-only tests; no CI infrastructure change\n\u2705 Covers the most common two paths\n\u274c Misses terminal-error fail-fast, body-hash key fallback, TTL, and the crash-after-dispatch path that actually causes duplicates\nActual answer: A) Full contract \u2014 user answer to D8\nAccepted scope: CRITICAL regression contract for processWebhookJob(), 7 assertions: (1) success \u2192 exactly one dispatch, payload/headers byte-identical to pre-rewrite; (2) terminal error (4xx/malformed) \u2192 exactly one attempt, job failed, no retry scheduled; (3) transient error \u2192 retry scheduled per D6 policy, attempt count incremented; (4) already-claimed key \u2192 no dispatch, job returns success; (5) key = provider event ID when present, SHA-256(raw body) when absent; (6) crash/timeout after dispatch before ack \u2192 retry finds claimed key \u2192 zero duplicate sends, integration test against a real/in-memory queue [\u2192E2E]; (7) dedup record TTL >= max total retry window (5 attempts at 10-min cap \u2248 50 min + base; assert TTL \u2265 1h or the computed window).\nHistory: original proposal (no regression test) superseded by D8.\nState: approved\n\n### R6: Proof depth for the shared retry policy module and worker migration\nFinding: #5 and #2 follow-on, P1, confidence 8/10, PLAN.md:6-13, reviewer: Claude (plan-eng-review)\nPlan baseline: D6 approved the policy and its per-bound tests as required proof; D7 approved migrating all 5 workers. Not yet decided: whether the jitter/cap curve gets a property-based test across the attempt range and whether each migrated worker gets an integration test proving the library actually calls the shared policy for it.\nRuntime evidence: unknown \u2014 no test framework detected in the reviewed repo.\nState: pending\nComparison grid:\n\n| Choice | Current | A | B |\n|---|---|---|---|\n| R1-R5 | approved (D4-D8) | fixed | fixed |\n| R6 per-bound unit tests (max attempts, cap, jitter range, exhaustion \u2192 DLQ + alert, transient vs terminal predicate) | required by D6 | included (carried) | included (carried) |\n| R6 property test: for attempt in 0..50, 0 \u2264 delay \u2264 min(cap, base*2^n), monotone cap, no overflow | not planned | included | not included |\n| R6 per-worker integration test: each of 5 workers, forced transient failure \u2192 library reschedules via shared policy, attempt log line emitted [\u2192E2E] | not planned | included (5 tests, shared fixture) | not included |\n| R6 exhaustion E2E: 5 forced failures \u2192 job in dead-letter set, alert emitted once [\u2192E2E] | not planned | included | not included |\n| R7 graph recompute | pending | pending | pending |\n\nQuestion D9:\nD9 \u2014 How deeply do we prove the shared retry policy and the five worker migrations?\nProject/branch/task: main branch; PLAN.md \"Add background job retry framework\", Tests section.\nELI10: You already approved unit tests for each policy bound. Two extra layers are on the table. First, a property test: instead of checking attempt 1, 2, 3 by hand, run the delay function for every attempt from 0 to 50 and assert it never exceeds the cap, never goes negative, and never overflows (2^50 milliseconds is a real number that bites). Second, integration tests: for each of the five workers, force one transient failure and confirm the queue library actually reschedules it through the shared policy and logs the attempt. That last one is what proves the migration did anything; a unit test on the policy function cannot tell you a worker forgot to register it.\nStakes if we pick wrong: a worker silently keeps no-retry behavior after \"migration\", or the delay math overflows at high attempt counts, and nothing in CI notices.\nRecommendation: A because the property test is ~10 lines and the five integration tests share one fixture with the D8 crash-after-dispatch test you already approved, so the marginal CC cost is minutes.\nCompleteness: A=10/10, B=7/10\nPros / cons:\nA) Per-bound units + property test + 5 per-worker integration tests + exhaustion E2E (recommended)\n \u2705 Proves each worker is actually wired to the shared policy, not just that the policy function is correct\n \u2705 Property test catches cap/overflow/negative-delay bugs across the whole attempt range in one assertion\n \u2705 Exhaustion E2E proves the dead-letter + alert path fires exactly once (human: ~2 days / CC: ~30 min)\n \u274c Six more integration tests in CI on top of D8's fixture; slower suite by seconds\nB) Per-bound unit tests only (as carried from D6)\n \u2705 Fast, no queue backend needed for this module\n \u2705 Already covers every named bound at specific attempts\n \u274c Cannot detect an unmigrated worker, a broken hook registration, or overflow at attempts you did not hand-pick\nNet: B proves the math; A proves the math and that production uses it.\nHeader: Policy tests\nOptions:\nA) Per-bound units + property test + 5 per-worker integration tests + exhaustion E2E (recommended)\n\u2705 Proves each worker is actually wired to the shared policy, not just that the policy function is correct\n\u2705 Property test catches cap/overflow/negative-delay bugs across the whole attempt range in one assertion\n\u2705 Exhaustion E2E proves the dead-letter + alert path fires exactly once (human: ~2 days / CC: ~30 min)\n\u274c Six more integration tests in CI on top of D8's fixture; slower suite by seconds\nB) Per-bound unit tests only (as carried from D6)\n\u2705 Fast, no queue backend needed for this module\n\u2705 Already covers every named bound at specific attempts\n\u274c Cannot detect an unmigrated worker, a broken hook registration, or overflow at attempts you did not hand-pick\nActual answer: unanswered\nAccepted scope: none\nHistory: none\n", "snapshot": { "at": "2026-09-16T06:02:22.822Z", "kind": "owned-plan-or-report", "source": "/tmp/g-pyj7vm4y/gstack-paid-shard-qDJLTL/tmp/gstack-e2e-plan-eng-batching-IpzCNz/gstack-test-plan-eng-batching.md", "artifact": "objects/52e76dbf87c4b0ff3f41e1d9b0d4256fa22b6e14b9bf6192309bf2e98d211fd5.md", "sha256": "52e76dbf87c4b0ff3f41e1d9b0d4256fa22b6e14b9bf6192309bf2e98d211fd5", "bytes": 33935, "mtimeMs": 1789538542736.3054, "provenance": "Exact observed file bytes; never reconstructed from tool text." } }, { "fingerprint": { "signature": "8c2809a2-42f4-479b-a5e4-a6f2e5016e61:toolu_01FYULvw5ZgYovtTydZYWig3", "promptSnippet": "Graph cache D10 \u2014 Cache the dependency graph across retries, keep recomputing, or measure first? Project/branch/task: main branch; PLAN.md \"Add background job retry framework\", Performance section. ELI10: Every retry currently reloads the j", "options": [ { "index": 1, "label": "Cache graph on first attempt, keyed by payload version; recompute on mismatch (recommended)" }, { "index": 2, "label": "Recompute every retry (as proposed)" }, { "index": 3, "label": "Instrument only: add timing metrics for fetch + graph build, decide after data" } ], "observedAtMs": 669869, "preReview": true, "nativeCall": { "sessionId": "8c2809a2-42f4-479b-a5e4-a6f2e5016e61", "toolUseId": "toolu_01FYULvw5ZgYovtTydZYWig3", "questions": [ { "question": "D10 \u2014 Cache the dependency graph across retries, keep recomputing, or measure first?\nProject/branch/task: main branch; PLAN.md \"Add background job retry framework\", Performance section.\nELI10: Every retry currently reloads the job's payload from the database and rebuilds a dependency graph by walking that payload. With the approved policy that is at most five rebuilds per job, so this is not an outage risk on its own. It is waste, and for large payloads it is waste on the exact retry path that is already under stress. Caching the graph is safe only if a stale graph cannot be used: keying the cache on a payload version (or hash) means a changed payload forces a rebuild, and re-fetching the small payload row each time stays cheap.\nStakes if we pick wrong: caching without a version check serves a stale graph after a payload edit; not caching means retries of big jobs pay full price while the downstream is already struggling.\nRecommendation: A because a version-keyed cache is a handful of lines with three tests, keeps re-fetch correctness, and removes the O(payload) walk from attempts 2-5 where it matters most.\nCompleteness: A=10/10, B=5/10, C=6/10\nNet: A is small and self-protecting; B is fine only if payloads are known to be tiny, which the plan does not say.", "header": "Graph cache", "multiSelect": false, "options": [ { "label": "Cache graph on first attempt, keyed by payload version; recompute on mismatch (recommended)", "description": "\u2705 Attempts 2-5 skip the O(payload) walk; the version key makes a stale graph impossible by construction\n\u2705 Payload re-fetch is unchanged, so any correctness that depends on fresh payload data is preserved\n\u2705 Three tests (miss, hit, mismatch) pin the behavior (human: ~1 day / CC: ~15 min)\n\u274c One more piece of per-job state to store and evict when the job finishes or dead-letters" }, { "label": "Recompute every retry (as proposed)", "description": "\u2705 Zero new state; simplest possible code path\n\u274c Retries of large payloads do 5x the graph work exactly when the system is already failing; no upper bound on payload size stated" }, { "label": "Instrument only: add timing metrics for fetch + graph build, decide after data", "description": "\u2705 Decision grounded in real numbers instead of guesses about payload size\n\u2705 Metrics are useful regardless of the eventual choice\n\u274c Leaves the 5x work in place for now and requires a second change later; the cache is cheap enough that measuring first buys little" } ] } ], "answered": true, "failed": false, "answers": { "D10 \u2014 Cache the dependency graph across retries, keep recomputing, or measure first?\nProject/branch/task: main branch; PLAN.md \"Add background job retry framework\", Performance section.\nELI10: Every retry currently reloads the job's payload from the database and rebuilds a dependency graph by walking that payload. With the approved policy that is at most five rebuilds per job, so this is not an outage risk on its own. It is waste, and for large payloads it is waste on the exact retry path that is already under stress. Caching the graph is safe only if a stale graph cannot be used: keying the cache on a payload version (or hash) means a changed payload forces a rebuild, and re-fetching the small payload row each time stays cheap.\nStakes if we pick wrong: caching without a version check serves a stale graph after a payload edit; not caching means retries of big jobs pay full price while the downstream is already struggling.\nRecommendation: A because a version-keyed cache is a handful of lines with three tests, keeps re-fetch correctness, and removes the O(payload) walk from attempts 2-5 where it matters most.\nCompleteness: A=10/10, B=5/10, C=6/10\nNet: A is small and self-protecting; B is fine only if payloads are known to be tiny, which the plan does not say.": "Cache graph on first attempt, keyed by payload version; recompute on mismatch (recommended)" }, "unansweredQuestionIndices": [], "answeredAt": "2026-09-16T06:04:22.292Z" } }, "questionAt": "2026-09-16T06:04:20.498Z", "answeredAt": "2026-09-16T06:04:22.292Z", "preAskPlan": "# Eng Review: Add background job retry framework\n\nReviewed target: `PLAN.md` (repo root, commit 7ce6818, branch main).\nReview: `/plan-eng-review`, 2026-09-16. Report file requested by the user.\n\n## Original plan (as reviewed)\n\n> ## Architecture\n> We'll roll a custom exponential-backoff scheduler inline in each worker\n> rather than use the existing job library's built-in retry hooks. Same\n> shape as the library version, but we want full control over the curve.\n>\n> ## Code quality\n> The retry envelope (compute delay, log attempt, dispatch) is duplicated\n> across 5 worker files with copy-pasted bodies. We will leave the\n> duplication for now and refactor \"later.\"\n>\n> ## Tests\n> The existing `processWebhookJob()` flow gets rewritten as part of this\n> change. No regression test for the prior at-most-once delivery guarantee\n> is planned.\n>\n> ## Performance\n> On every retry we re-fetch the full job payload from the database, then\n> iterate the payload to recompute the dependency graph. Could cache the\n> graph on the first attempt; not planned.\n\n## Step 0: Scope Challenge findings\n\n1. [P1] (confidence: 8/10) PLAN.md:6-8 \u2014 custom inline scheduler rebuilds the job library's built-in retry hooks (Layer 1). Custom backoff curves are a supported extension point. Disposition: pending (R1 / D4).\n2. [P1] (confidence: 9/10) PLAN.md:11-13 \u2014 retry envelope copy-pasted across 5 worker files; refactor deferred. Disposition: pending (Code Quality).\n3. [P0] (confidence: 9/10) PLAN.md:16-18 \u2014 `processWebhookJob()` rewritten with no regression test; retries change the guarantee from at-most-once to at-least-once with no idempotency key. Disposition: pending (Architecture + Tests).\n4. [P2] (confidence: 7/10) PLAN.md:21-23 \u2014 payload re-fetched and dependency graph recomputed on every retry. Disposition: pending (Performance).\n5. [P1] (confidence: 8/10) PLAN.md:6-8 \u2014 \"full control over the curve\" but no curve specified: no max attempts, delay cap, jitter, dead-letter, or retryable-error classification. Disposition: pending (Architecture).\n6. Factual gap: no design doc / problem statement (which failure classes retries target). Recorded; folded into finding 5.\n\nComplexity gate: not triggered (5 worker files + 1 rewrite, 0 new classes). TODOS.md: absent. Distribution: no new artifact.\n\n## Decision ledger\n\n### R1: Retry mechanism \u2014 library built-in retry hooks vs custom inline scheduler\nFinding: #1, P1, confidence 8/10, PLAN.md:6-8, reviewer: Claude (plan-eng-review)\nPlan baseline: original proposal \u2014 custom exponential-backoff scheduler inline in each worker; library hooks unused.\nRuntime evidence: unknown \u2014 the repo under review contains only PLAN.md and CLAUDE.md; the job library and worker files are not present. Web search (2026) shows mainstream job libraries expose a custom backoff function / retry_in hook.\nState: approved (D4) \u2014 see Actual answer below\nComparison grid:\n\n| Choice | Current | A | B |\n|---|---|---|---|\n| R1 retry mechanism | custom inline scheduler per worker (proposed) | library built-in retry hooks with one shared custom backoff function | custom inline scheduler per worker (as proposed) |\n| R2 envelope duplication | 5 copies, deferred (pending) | pending | pending |\n| R3 webhook delivery semantics | unspecified (pending) | pending | pending |\n| R4 retry policy bounds | unspecified (pending) | pending | pending |\n| R5 webhook regression contract | none planned (pending) | pending | pending |\n| R7 graph recompute per retry | recompute every retry (pending) | pending | pending |\n\nQuestion D4:\nD4 \u2014 Retry mechanism: use the job library's built-in retry hooks, or roll a custom inline scheduler?\nProject/branch/task: main branch; PLAN.md \"Add background job retry framework\", Architecture section.\nELI10: The plan rebuilds retry scheduling inside each worker even though the job library already has retry hooks. The stated reason is \"full control over the curve\", but these libraries let you plug in your own backoff function, so you get the exact curve you want while the library keeps doing the hard parts: persisting attempt counts across worker crashes, scheduling the delayed re-run, and exposing retries to its dashboard and metrics. A hand-rolled inline scheduler has to re-solve all of that in five places.\nStakes if we pick wrong: a worker that crashes mid-sleep in a custom scheduler loses the job silently; a library-managed retry survives the crash and is visible in the queue UI.\nRecommendation: A because the library hook gives the same curve control with crash-safe attempt state and zero new scheduling code; the custom path only wins if the library has no backoff hook at all, which the plan does not claim.\nCompleteness: A=10/10, B=6/10\nPros / cons:\nA) Library retry hooks + one custom backoff function (recommended)\n \u2705 Attempt count and next-run time persist in the queue, so a worker crash mid-backoff cannot lose the job\n \u2705 Retries show up in the library's dashboard, metrics and dead-letter handling for free\n \u2705 One backoff function to test instead of five inline schedulers (human: ~1 day / CC: ~20 min)\n \u274c Curve is bounded by whatever the hook API allows (delay per attempt); exotic behavior like per-tenant curves needs a wrapper\nB) Custom inline scheduler in each worker (as proposed)\n \u2705 Unconstrained control over every retry decision inside worker code\n \u2705 No dependency on the library's hook signature or its upgrade path\n \u274c Attempt state lives in process memory unless you build persistence; a crash or deploy mid-backoff drops the retry\n \u274c Five schedulers to keep in sync, invisible to the queue UI and metrics (human: ~3 days / CC: ~45 min)\nNet: same curve either way; the trade is crash-safety and observability for free vs. rebuilding the library's job.\nHeader: Retry mech\nOptions:\nA) Library retry hooks + one custom backoff function (recommended)\n\u2705 Attempt count and next-run time persist in the queue, so a worker crash mid-backoff cannot lose the job\n\u2705 Retries show up in the library's dashboard, metrics and dead-letter handling for free\n\u2705 One backoff function to test instead of five inline schedulers (human: ~1 day / CC: ~20 min)\n\u274c Curve is bounded by whatever the hook API allows (delay per attempt); exotic behavior like per-tenant curves needs a wrapper\nB) Custom inline scheduler in each worker (as proposed)\n\u2705 Unconstrained control over every retry decision inside worker code\n\u2705 No dependency on the library's hook signature or its upgrade path\n\u274c Attempt state lives in process memory unless you build persistence; a crash or deploy mid-backoff drops the retry\n\u274c Five schedulers to keep in sync, invisible to the queue UI and metrics (human: ~3 days / CC: ~45 min)\nActual answer: A) Library retry hooks + one custom backoff function \u2014 user answer to D4\nAccepted scope: Replace the custom inline scheduler with the job library's built-in retry hooks; implement one shared custom backoff function that the hooks call; delete the plan's inline per-worker scheduler. Scope Challenge result: scope reduced per recommendation.\nHistory: original proposal (custom inline scheduler) superseded by D4.\nState: approved\n\n### R3: Delivery semantics for `processWebhookJob()` under retries\nFinding: #3, P0, confidence 9/10, PLAN.md:16-18, reviewer: Claude (plan-eng-review)\nPlan baseline: original proposal \u2014 processWebhookJob() is rewritten to retry; the prior at-most-once guarantee is neither preserved nor replaced with an explicit alternative.\nRuntime evidence: unknown \u2014 worker source not present in the reviewed repo. The plan itself states the current guarantee is at-most-once (PLAN.md:17). Web sources (Svix, Hookdeck, 2026) agree that retries imply at-least-once and require an idempotency key to be safe.\nState: approved (D5) \u2014 see Actual answer below\nComparison grid:\n\n| Choice | Current | A | B | C |\n|---|---|---|---|---|\n| R1 retry mechanism | approved: library hooks + shared backoff fn (D4) | fixed | fixed | fixed |\n| R3 webhook delivery semantics | at-most-once today; plan retries without stating new guarantee | at-least-once with idempotency key (provider event ID, or hash of raw body) claimed atomically before side effects; duplicates return success | keep at-most-once: exclude processWebhookJob from retries (max attempts = 1); other 4 workers retry | retry with no idempotency key (as proposed) |\n| R2 envelope duplication | pending | pending | pending | pending |\n| R4 retry policy bounds | pending | pending | pending | pending |\n| R5 webhook regression contract | pending | pending | pending | pending |\n| R7 graph recompute per retry | pending | pending | pending | pending |\n\nQuestion D5:\nD5 \u2014 When processWebhookJob() gains retries, what delivery guarantee does it make?\nProject/branch/task: main branch; PLAN.md \"Add background job retry framework\", Tests/Architecture sections.\nELI10: Today the webhook job runs at most once: if it fails, nothing is redelivered, so it can never double-fire. The moment you add retries, a job that timed out after the receiver already processed it will run again, and the receiver sees the same webhook twice. That is a real behavior change, not just a missing test. The standard fix is an idempotency key: record the event ID before doing the side effect, and if the key is already there, treat the retry as already done.\nStakes if we pick wrong: duplicate webhooks mean double charges, double emails, or double state changes for whoever receives them, and you find out from a customer, not a test.\nRecommendation: A because retrying is the point of this plan, and an idempotency key is a few lines plus one unique index; it turns duplicates from an incident into a no-op.\nCompleteness: A=10/10, B=7/10, C=3/10\nPros / cons:\nA) At-least-once + idempotency key (recommended)\n \u2705 Transient failures get retried and duplicates are absorbed, so receivers never see a double effect\n \u2705 Small change: atomic INSERT ... ON CONFLICT DO NOTHING (or SET NX) on the event ID before dispatch (human: ~1 day / CC: ~15 min)\n \u274c Needs a dedup store with a TTL at least as long as the retry window, and a decision on what the key is when the payload has no natural ID\nB) Keep at-most-once: no retries for the webhook job\n \u2705 Zero semantic change for webhook receivers; regression risk drops to the rewrite itself\n \u2705 Simplest to prove: max attempts = 1 for this job, existing behavior stays\n \u274c Webhook deliveries that fail on a transient blip are still lost, which is the pain this plan exists to fix\nC) Retry with no idempotency key (as proposed)\n \u2705 No extra storage or key design work up front\n \u274c Duplicate side effects on every timeout-after-success; the at-most-once guarantee silently disappears\nNet: you are choosing between fixing lost webhooks safely, not fixing them, or fixing them in a way that creates duplicates.\nHeader: Webhook sem\nOptions:\nA) At-least-once + idempotency key (recommended)\n\u2705 Transient failures get retried and duplicates are absorbed, so receivers never see a double effect\n\u2705 Small change: atomic INSERT ... ON CONFLICT DO NOTHING (or SET NX) on the event ID before dispatch (human: ~1 day / CC: ~15 min)\n\u274c Needs a dedup store with a TTL at least as long as the retry window, and a decision on what the key is when the payload has no natural ID\nB) Keep at-most-once: no retries for the webhook job\n\u2705 Zero semantic change for webhook receivers; regression risk drops to the rewrite itself\n\u2705 Simplest to prove: max attempts = 1 for this job, existing behavior stays\n\u274c Webhook deliveries that fail on a transient blip are still lost, which is the pain this plan exists to fix\nC) Retry with no idempotency key (as proposed)\n\u2705 No extra storage or key design work up front\n\u274c Duplicate side effects on every timeout-after-success; the at-most-once guarantee silently disappears\nActual answer: A) At-least-once + idempotency key \u2014 user answer to D5\nAccepted scope: processWebhookJob() becomes at-least-once. Before any side effect, atomically claim an idempotency key (provider event ID when present; otherwise SHA-256 of the raw body) via INSERT ... ON CONFLICT DO NOTHING or equivalent; if the claim loses, return success without re-dispatching. Dedup record TTL >= the full retry window. Tests and docs for this contract are carried into Test review (R5) as required proof, not a new decision.\nHistory: original proposal (retry with no idempotency key) superseded by D5.\nState: approved\n\n### R4: Retry policy completeness \u2014 bounds the plan must specify\nFinding: #5, P1, confidence 8/10, PLAN.md:6-8, reviewer: Claude (plan-eng-review)\nPlan baseline: original proposal \u2014 \"full control over the curve\" with no max attempts, delay cap, jitter, dead-letter behavior, or retryable/terminal error classification stated.\nRuntime evidence: unknown \u2014 no worker or library source in the reviewed repo. Web sources (2026) name uncapped delay and missing jitter as the two standard footguns; retry storms and ~12-day delays at attempt 20.\nState: pending\nComparison grid:\n\n| Choice | Current | A | B | C |\n|---|---|---|---|---|\n| R1 retry mechanism | approved: library hooks + shared backoff fn (D4) | fixed | fixed | fixed |\n| R3 webhook semantics | approved: at-least-once + idempotency key (D5) | fixed | fixed | fixed |\n| R4 policy: max attempts | unspecified | 5 (per-job override allowed) | 5 (per-job override allowed) | unspecified |\n| R4 policy: delay cap | unspecified | min(cap, base*2^n), cap = 10 min | unspecified | unspecified |\n| R4 policy: jitter | unspecified | full jitter: random(0, computed delay) | unspecified | unspecified |\n| R4 policy: exhaustion | unspecified | move to library dead-letter/failed set + alert | job marked failed, no alert | unspecified |\n| R4 policy: error classification | unspecified | retry only on transient errors (timeout, 5xx, connection reset); terminal errors (4xx validation, malformed payload) fail immediately | all errors retried | all errors retried |\n| R2 envelope duplication | pending | pending | pending | pending |\n| R5 webhook regression contract | pending | pending | pending | pending |\n| R7 graph recompute per retry | pending | pending | pending | pending |\n\nQuestion D6:\nD6 \u2014 How completely should the plan specify the retry policy (attempts, cap, jitter, exhaustion, error classes)?\nProject/branch/task: main branch; PLAN.md \"Add background job retry framework\", Architecture section.\nELI10: Exponential backoff is only half a policy. Without a cap, attempt 20 waits about 12 days. Without jitter, every job that failed at the same moment retries at the same moment and hammers the downstream again (a retry storm). Without a max, a permanently broken payload retries forever. And retrying a \"400 invalid payload\" is pointless: only transient failures (timeouts, 5xx, dropped connections) deserve a retry. The plan says \"full control over the curve\" but writes none of these down, so each of the five workers will pick its own answer.\nStakes if we pick wrong: a poisoned job retrying forever, or a downstream outage made worse by synchronized retries; both look like mystery load at 3am.\nRecommendation: A because every one of these is a one-line constant or a small predicate in the single shared backoff function you already approved, and the defaults below are the standard ones (adjust the numbers freely via Other).\nCompleteness: A=10/10, B=6/10, C=3/10\nPros / cons:\nA) Full policy: max 5 attempts, 10-min cap, full jitter, dead-letter + alert on exhaustion, retry only transient errors (recommended)\n \u2705 Bounded worst case: a broken job costs at most 5 attempts and ~20 minutes, then lands in a dead-letter set someone gets paged for\n \u2705 Full jitter spreads synchronized failures so a downstream blip does not become a retry storm\n \u2705 Terminal errors fail fast, so bad payloads surface immediately instead of after 5 retries (human: ~1 day / CC: ~15 min)\n \u274c Five constants and one error predicate to agree on now; per-job overrides add a little config surface\nB) Max attempts only (5), everything else unspecified\n \u2705 Prevents infinite retries with one constant\n \u2705 Smallest possible addition to the plan text\n \u274c Uncapped delay and no jitter still allow retry storms and multi-hour waits within those 5 attempts; terminal errors waste retries\nC) Curve only, as proposed\n \u2705 No policy decisions needed before coding starts\n \u274c Each worker author invents their own bounds (or none); unbounded delay and infinite retries are possible\nNet: the difference between A and C is about 15 lines in one shared function versus five different implicit policies you discover in production.\nHeader: Retry policy\nOptions:\nA) Full policy: max 5 attempts, 10-min cap, full jitter, dead-letter + alert on exhaustion, retry only transient errors (recommended)\n\u2705 Bounded worst case: a broken job costs at most 5 attempts and ~20 minutes, then lands in a dead-letter set someone gets paged for\n\u2705 Full jitter spreads synchronized failures so a downstream blip does not become a retry storm\n\u2705 Terminal errors fail fast, so bad payloads surface immediately instead of after 5 retries (human: ~1 day / CC: ~15 min)\n\u274c Five constants and one error predicate to agree on now; per-job overrides add a little config surface\nB) Max attempts only (5), everything else unspecified\n\u2705 Prevents infinite retries with one constant\n\u2705 Smallest possible addition to the plan text\n\u274c Uncapped delay and no jitter still allow retry storms and multi-hour waits within those 5 attempts; terminal errors waste retries\nC) Curve only, as proposed\n\u2705 No policy decisions needed before coding starts\n\u274c Each worker author invents their own bounds (or none); unbounded delay and infinite retries are possible\nActual answer: A) Full policy \u2014 user answer to D6\nAccepted scope: The shared backoff function implements: max 5 attempts (per-job override allowed); delay = min(10 min, base * 2^attempt) with full jitter random(0, delay); on exhaustion move the job to the library's dead-letter/failed set and emit an alert; retry only transient errors (timeout, 5xx, connection reset), fail immediately on terminal errors (4xx validation, malformed payload). Tests proving each bound are required proof carried into Test review (R6).\nHistory: original proposal (curve only, no bounds) superseded by D6.\nState: approved\n\n### R2: Retry envelope duplication across 5 worker files\nFinding: #2, P1, confidence 9/10, PLAN.md:11-13, reviewer: Claude (plan-eng-review)\nPlan baseline: original proposal \u2014 envelope (compute delay, log attempt, dispatch) copy-pasted in 5 worker files; refactor deferred to \"later\".\nRuntime evidence: unknown \u2014 worker files not present in the reviewed repo; the plan itself states the duplication exists (PLAN.md:11-12).\nState: pending\nComparison grid:\n\n| Choice | Current | A | B | C |\n|---|---|---|---|---|\n| R1 retry mechanism | approved: library hooks + shared backoff fn (D4) | fixed | fixed | fixed |\n| R3 webhook semantics | approved: at-least-once + idempotency key (D5) | fixed | fixed | fixed |\n| R4 retry policy | approved: full policy in shared fn (D6) | fixed | fixed | fixed |\n| R2 envelope: where it lives | 5 inline copies | one shared module: backoff fn + attempt-logging middleware registered once with the library | one shared module, but only workers touched by this change migrate | 5 inline copies remain |\n| R2 envelope: workers migrated in this change | 0 | all 5 | only those edited for retries (webhook + any other touched) | 0 |\n| R5 webhook regression contract | pending | pending | pending | pending |\n| R7 graph recompute per retry | pending | pending | pending | pending |\n\nQuestion D7:\nD7 \u2014 Consolidate the retry envelope into one shared module now, or leave five copies and refactor later?\nProject/branch/task: main branch; PLAN.md \"Add background job retry framework\", Code quality section.\nELI10: The same three steps (work out the delay, log the attempt, hand the job back to the queue) are pasted into five worker files. You already decided the delay logic becomes one shared backoff function the library calls, so most of that envelope disappears on its own. What remains is deciding whether the attempt logging and hook registration also live in one place, and whether all five workers move over in this change or only the ones you happen to touch. \"Refactor later\" for copy-pasted code almost always means the first backoff bug gets fixed in four files and missed in the fifth.\nStakes if we pick wrong: a policy fix (say, the jitter formula) lands in some workers and not others, and the inconsistent ones are the ones nobody tests.\nRecommendation: A because with the library-hook decision the shared module is small (one backoff function plus one logging middleware), and migrating five call sites is a mechanical change that CC does in minutes.\nCompleteness: A=10/10, B=7/10, C=3/10\nPros / cons:\nA) One shared retry module, all 5 workers migrated now (recommended)\n \u2705 Exactly one place defines delay, jitter, attempt logging and exhaustion handling; one test suite proves it\n \u2705 Deleting five inline envelopes shrinks the diff and removes the code the plan admits is copy-pasted (human: ~1.5 days / CC: ~25 min)\n \u274c Touches all five worker files in one change, so the review surface is wider than the webhook rewrite alone\nB) Shared module, migrate only workers this change already touches\n \u2705 Smaller immediate diff; untouched workers carry no risk from this change\n \u2705 The shared module exists, so later migrations are copy-and-delete\n \u274c Two retry implementations coexist in production; the inline ones keep their unbounded, unjittered behavior until someone remembers\nC) Leave five copies, refactor \"later\" (as proposed)\n \u2705 No refactor risk in this change\n \u274c Five places to apply every future policy fix; the approved policy (D6) has to be pasted five times to take effect\nNet: the shared module is the approved design's natural home; the only real question is whether all five workers adopt it now or drift.\nHeader: DRY envelope\nOptions:\nA) One shared retry module, all 5 workers migrated now (recommended)\n\u2705 Exactly one place defines delay, jitter, attempt logging and exhaustion handling; one test suite proves it\n\u2705 Deleting five inline envelopes shrinks the diff and removes the code the plan admits is copy-pasted (human: ~1.5 days / CC: ~25 min)\n\u274c Touches all five worker files in one change, so the review surface is wider than the webhook rewrite alone\nB) Shared module, migrate only workers this change already touches\n\u2705 Smaller immediate diff; untouched workers carry no risk from this change\n\u2705 The shared module exists, so later migrations are copy-and-delete\n\u274c Two retry implementations coexist in production; the inline ones keep their unbounded, unjittered behavior until someone remembers\nC) Leave five copies, refactor \"later\" (as proposed)\n\u2705 No refactor risk in this change\n\u274c Five places to apply every future policy fix; the approved policy (D6) has to be pasted five times to take effect\nActual answer: A) One shared retry module, all 5 workers migrated now \u2014 user answer to D7\nAccepted scope: Create one shared retry module (e.g. `jobs/retry/` or the project's equivalent): the backoff/policy function from D6 plus an attempt-logging middleware (job id, job type, attempt n/max, error class, next delay ms) registered once with the job library. Delete the inline envelope from all 5 worker files and register them with the shared module. Existing worker behavior other than retry scheduling is unchanged.\nHistory: original proposal (leave 5 copies, refactor later) superseded by D7.\nState: approved\n\n### R5: Regression contract for the `processWebhookJob()` rewrite\nFinding: #3, P0, confidence 9/10, PLAN.md:16-18, reviewer: Claude (plan-eng-review)\nPlan baseline: original proposal \u2014 \"No regression test for the prior at-most-once delivery guarantee is planned.\" Approved behavior to protect comes from D5 (at-least-once + idempotency key) and D6 (retry only transient errors).\nRuntime evidence: unknown \u2014 no test framework or test files exist in the reviewed repo (git ls-files: CLAUDE.md, PLAN.md only). Existing callers of processWebhookJob() cannot be traced here; treat every current behavior as at risk.\nState: pending\nComparison grid:\n\n| Choice | Current | A | B |\n|---|---|---|---|\n| R1, R3, R4, R2 | approved (D4-D7) | fixed | fixed |\n| R5 behavior preserved: success dispatches exactly once with unchanged payload/headers | untested | asserted | asserted |\n| R5 behavior preserved: terminal error (4xx/malformed) \u2192 exactly one attempt, job failed, no retry (old at-most-once holds for this class) | untested | asserted | not asserted |\n| R5 intentional change: transient error \u2192 retried per D6 policy, attempt count visible | untested | asserted | asserted |\n| R5 new contract: second run with an already-claimed key \u2192 no dispatch, returns success | untested | asserted | asserted |\n| R5 new contract: key derivation \u2014 provider event ID when present, body hash when absent | untested | asserted (both branches) | not asserted |\n| R5 new contract: crash/timeout after dispatch, before ack \u2192 retry finds claimed key \u2192 zero duplicate sends [\u2192E2E] | untested | asserted via integration test against the real queue | not asserted |\n| R5: dedup TTL >= retry window | untested | asserted | not asserted |\n| R6 policy proof depth | pending | pending | pending |\n| R7 graph recompute | pending | pending | pending |\n\nQuestion D8:\nD8 \u2014 What regression contract protects processWebhookJob() through the rewrite?\nProject/branch/task: main branch; PLAN.md \"Add background job retry framework\", Tests section.\nELI10: The webhook job is being rewritten and currently has no test guarding what it does today. A regression contract is the written list of \"what must still be true afterwards\", plus the things we are changing on purpose. For this job that is: one successful run still sends exactly one webhook with the same content; a permanently bad request still fails once and stops; a transient failure now retries (the intentional change); and a retry after a timeout-that-actually-succeeded sends nothing twice. The last one is the whole point of the idempotency key you approved, and it only shows up in a test that runs against the real queue, not a mock.\nStakes if we pick wrong: the rewrite ships, a receiver gets two payment webhooks after a slow response, and there was never a test that could have caught it.\nRecommendation: A because these seven assertions are the entire behavioral surface of the job, and the ones B drops (terminal fail-fast, key fallback, crash-after-dispatch) are exactly the failure paths that produce duplicates.\nCompleteness: A=10/10, B=6/10\nPros / cons:\nA) Full contract: 7 assertions incl. terminal fail-fast, key fallback and crash-after-dispatch integration test (recommended)\n \u2705 Every branch of the new job has a named assertion; the duplicate-send path is proven against the real queue, not a mock\n \u2705 Preserves the old at-most-once behavior where it still applies (terminal errors) and documents the intentional change for transient ones\n \u2705 Tests double as the spec for the idempotency key and TTL (human: ~2 days / CC: ~30 min)\n \u274c One integration test needs a real or in-memory queue backend in CI\nB) Happy path + duplicate-claim unit tests only\n \u2705 Fast, mock-only tests; no CI infrastructure change\n \u2705 Covers the most common two paths\n \u274c Misses terminal-error fail-fast, body-hash key fallback, TTL, and the crash-after-dispatch path that actually causes duplicates\nNet: A costs one integration test fixture; B leaves the duplicate-send path, the reason D5 exists, unproven.\nHeader: Webhook tests\nOptions:\nA) Full contract: 7 assertions incl. terminal fail-fast, key fallback and crash-after-dispatch integration test (recommended)\n\u2705 Every branch of the new job has a named assertion; the duplicate-send path is proven against the real queue, not a mock\n\u2705 Preserves the old at-most-once behavior where it still applies (terminal errors) and documents the intentional change for transient ones\n\u2705 Tests double as the spec for the idempotency key and TTL (human: ~2 days / CC: ~30 min)\n\u274c One integration test needs a real or in-memory queue backend in CI\nB) Happy path + duplicate-claim unit tests only\n\u2705 Fast, mock-only tests; no CI infrastructure change\n\u2705 Covers the most common two paths\n\u274c Misses terminal-error fail-fast, body-hash key fallback, TTL, and the crash-after-dispatch path that actually causes duplicates\nActual answer: A) Full contract \u2014 user answer to D8\nAccepted scope: CRITICAL regression contract for processWebhookJob(), 7 assertions: (1) success \u2192 exactly one dispatch, payload/headers byte-identical to pre-rewrite; (2) terminal error (4xx/malformed) \u2192 exactly one attempt, job failed, no retry scheduled; (3) transient error \u2192 retry scheduled per D6 policy, attempt count incremented; (4) already-claimed key \u2192 no dispatch, job returns success; (5) key = provider event ID when present, SHA-256(raw body) when absent; (6) crash/timeout after dispatch before ack \u2192 retry finds claimed key \u2192 zero duplicate sends, integration test against a real/in-memory queue [\u2192E2E]; (7) dedup record TTL >= max total retry window (5 attempts at 10-min cap \u2248 50 min + base; assert TTL \u2265 1h or the computed window).\nHistory: original proposal (no regression test) superseded by D8.\nState: approved\n\n### R6: Proof depth for the shared retry policy module and worker migration\nFinding: #5 and #2 follow-on, P1, confidence 8/10, PLAN.md:6-13, reviewer: Claude (plan-eng-review)\nPlan baseline: D6 approved the policy and its per-bound tests as required proof; D7 approved migrating all 5 workers. Not yet decided: whether the jitter/cap curve gets a property-based test across the attempt range and whether each migrated worker gets an integration test proving the library actually calls the shared policy for it.\nRuntime evidence: unknown \u2014 no test framework detected in the reviewed repo.\nState: pending\nComparison grid:\n\n| Choice | Current | A | B |\n|---|---|---|---|\n| R1-R5 | approved (D4-D8) | fixed | fixed |\n| R6 per-bound unit tests (max attempts, cap, jitter range, exhaustion \u2192 DLQ + alert, transient vs terminal predicate) | required by D6 | included (carried) | included (carried) |\n| R6 property test: for attempt in 0..50, 0 \u2264 delay \u2264 min(cap, base*2^n), monotone cap, no overflow | not planned | included | not included |\n| R6 per-worker integration test: each of 5 workers, forced transient failure \u2192 library reschedules via shared policy, attempt log line emitted [\u2192E2E] | not planned | included (5 tests, shared fixture) | not included |\n| R6 exhaustion E2E: 5 forced failures \u2192 job in dead-letter set, alert emitted once [\u2192E2E] | not planned | included | not included |\n| R7 graph recompute | pending | pending | pending |\n\nQuestion D9:\nD9 \u2014 How deeply do we prove the shared retry policy and the five worker migrations?\nProject/branch/task: main branch; PLAN.md \"Add background job retry framework\", Tests section.\nELI10: You already approved unit tests for each policy bound. Two extra layers are on the table. First, a property test: instead of checking attempt 1, 2, 3 by hand, run the delay function for every attempt from 0 to 50 and assert it never exceeds the cap, never goes negative, and never overflows (2^50 milliseconds is a real number that bites). Second, integration tests: for each of the five workers, force one transient failure and confirm the queue library actually reschedules it through the shared policy and logs the attempt. That last one is what proves the migration did anything; a unit test on the policy function cannot tell you a worker forgot to register it.\nStakes if we pick wrong: a worker silently keeps no-retry behavior after \"migration\", or the delay math overflows at high attempt counts, and nothing in CI notices.\nRecommendation: A because the property test is ~10 lines and the five integration tests share one fixture with the D8 crash-after-dispatch test you already approved, so the marginal CC cost is minutes.\nCompleteness: A=10/10, B=7/10\nPros / cons:\nA) Per-bound units + property test + 5 per-worker integration tests + exhaustion E2E (recommended)\n \u2705 Proves each worker is actually wired to the shared policy, not just that the policy function is correct\n \u2705 Property test catches cap/overflow/negative-delay bugs across the whole attempt range in one assertion\n \u2705 Exhaustion E2E proves the dead-letter + alert path fires exactly once (human: ~2 days / CC: ~30 min)\n \u274c Six more integration tests in CI on top of D8's fixture; slower suite by seconds\nB) Per-bound unit tests only (as carried from D6)\n \u2705 Fast, no queue backend needed for this module\n \u2705 Already covers every named bound at specific attempts\n \u274c Cannot detect an unmigrated worker, a broken hook registration, or overflow at attempts you did not hand-pick\nNet: B proves the math; A proves the math and that production uses it.\nHeader: Policy tests\nOptions:\nA) Per-bound units + property test + 5 per-worker integration tests + exhaustion E2E (recommended)\n\u2705 Proves each worker is actually wired to the shared policy, not just that the policy function is correct\n\u2705 Property test catches cap/overflow/negative-delay bugs across the whole attempt range in one assertion\n\u2705 Exhaustion E2E proves the dead-letter + alert path fires exactly once (human: ~2 days / CC: ~30 min)\n\u274c Six more integration tests in CI on top of D8's fixture; slower suite by seconds\nB) Per-bound unit tests only (as carried from D6)\n\u2705 Fast, no queue backend needed for this module\n\u2705 Already covers every named bound at specific attempts\n\u274c Cannot detect an unmigrated worker, a broken hook registration, or overflow at attempts you did not hand-pick\nActual answer: A) Per-bound units + property test + 5 per-worker integration tests + exhaustion E2E \u2014 user answer to D9\nAccepted scope: Unit tests for each D6 bound; property test over attempts 0..50 (0 \u2264 delay \u2264 min(cap, base*2^n), no negative/overflow); 5 per-worker integration tests (forced transient failure \u2192 library reschedules via shared policy, attempt log emitted); one exhaustion E2E (5 failures \u2192 dead-letter set, alert exactly once). Shares the queue fixture with D8 assertion (6).\nHistory: not planned in original proposal; added by D9.\nState: approved\n\n### Test coverage diagram (Section 3)\n\nFramework: unknown \u2014 the reviewed repo contains no source, test config or test files (`git ls-files`: CLAUDE.md, PLAN.md). Every path below is a GAP today; the approved contracts (D8, D9) define the tests to write. Test Plan Artifact: `~/.gstack/projects/gstack-plan-count-vjhjkU/vercel-sandbox-main-eng-review-test-plan-20260916-060250.md`.\n\n```\nCODE PATHS USER / OPERATOR FLOWS\n[+] jobs/retry/policy (shared, D4/D6/D7) [+] Transient webhook failure\n \u251c\u2500\u2500 backoff(attempt, err) \u251c\u2500\u2500 [GAP] [\u2192E2E] attempt 1 times out \u2192 retried \u2192 attempt 2 succeeds\n \u2502 \u251c\u2500\u2500 [GAP] attempt < max \u2192 delay = min(cap, base*2^n) \u2502 \u2192 receiver sees exactly ONE webhook\n \u2502 \u251c\u2500\u2500 [GAP] full jitter: 0 \u2264 delay \u2264 computed \u2514\u2500\u2500 [GAP] attempt log lines visible per attempt\n \u2502 \u251c\u2500\u2500 [GAP] property: attempts 0..50 in range, no overflow\n \u2502 \u251c\u2500\u2500 [GAP] attempt >= max \u2192 exhaustion [+] Poisoned job\n \u2502 \u2502 \u251c\u2500\u2500 [GAP] [\u2192E2E] job \u2192 dead-letter/failed set \u251c\u2500\u2500 [GAP] [\u2192E2E] 5 attempts \u2248 20 min \u2192 dead-letter\n \u2502 \u2502 \u2514\u2500\u2500 [GAP] [\u2192E2E] alert emitted exactly once \u2514\u2500\u2500 [GAP] operator sees it in queue UI + one alert\n \u2502 \u2514\u2500\u2500 isTransient(err)\n \u2502 \u251c\u2500\u2500 [GAP] timeout / 5xx / conn reset \u2192 true [+] Bad payload\n \u2502 \u2514\u2500\u2500 [GAP] 4xx validation / malformed \u2192 false \u2514\u2500\u2500 [GAP] fails once, no retries, error visible immediately\n \u2514\u2500\u2500 attemptLogger middleware\n \u2514\u2500\u2500 [GAP] logs job id, type, n/max, error class, delay [+] Worker crash mid-retry (D4 rationale)\n \u2514\u2500\u2500 [GAP] [\u2192E2E] queue-persisted attempt survives restart\n[+] workers/* (5 files, D7 migration)\n \u2514\u2500\u2500 [GAP] [\u2192E2E] x5: forced transient failure \u2192 library [+] Duplicate delivery (D5)\n reschedules via shared policy (one per worker) \u251c\u2500\u2500 [GAP] [\u2192E2E] timeout-after-success retry \u2192 zero duplicates\n \u2514\u2500\u2500 [GAP] same key twice \u2192 second is a no-op success\n[+] processWebhookJob() (rewrite, D5/D8)\n \u251c\u2500\u2500 deriveKey(payload)\n \u2502 \u251c\u2500\u2500 [GAP] provider event ID present \u2192 key = ID\n \u2502 \u2514\u2500\u2500 [GAP] absent \u2192 key = sha256(raw body)\n \u251c\u2500\u2500 claimKey(key) INSERT ... ON CONFLICT DO NOTHING\n \u2502 \u251c\u2500\u2500 [GAP] claim won \u2192 dispatch\n \u2502 \u2514\u2500\u2500 [GAP] claim lost \u2192 return success, no dispatch\n \u251c\u2500\u2500 dispatch()\n \u2502 \u251c\u2500\u2500 [GAP] success \u2192 exactly one send, payload/headers byte-identical\n \u2502 \u251c\u2500\u2500 [GAP] transient error \u2192 retry scheduled (D6)\n \u2502 \u2514\u2500\u2500 [GAP] terminal error \u2192 fail, one attempt (old at-most-once preserved for this class)\n \u251c\u2500\u2500 [GAP] [\u2192E2E] crash after dispatch before ack \u2192 retry finds claimed key \u2192 0 duplicates CRITICAL\n \u2514\u2500\u2500 [GAP] dedup TTL \u2265 retry window (\u2265 ~1h)\n\n[+] dependency graph per retry (R7 \u2014 pending D10)\n \u251c\u2500\u2500 [GAP] cache miss (first attempt) \u2192 compute + store\n \u251c\u2500\u2500 [GAP] cache hit, payload version unchanged \u2192 reuse\n \u2514\u2500\u2500 [GAP] payload version changed \u2192 recompute\n\nCOVERAGE: 0/32 paths tested (0%) | Code paths: 0/24 (0%) | User flows: 0/8 (0%)\nQUALITY: \u2605\u2605\u2605:0 \u2605\u2605:0 \u2605:0 | GAPS: 32 (11 E2E, 0 eval) | Approved as required tests: 29 (R7's 3 pending D10)\n```\n\nLegend: \u2605\u2605\u2605 behavior + edge + error | \u2605\u2605 happy path | \u2605 smoke | [\u2192E2E] = integration test against a real/in-memory queue. LLM/eval scope: none (no prompt or LLM code in this plan).\n\n### R7: Dependency graph recomputed on every retry\nFinding: #4, P2, confidence 7/10, PLAN.md:21-23, reviewer: Claude (plan-eng-review)\nPlan baseline: original proposal \u2014 \"On every retry we re-fetch the full job payload from the database, then iterate the payload to recompute the dependency graph. Could cache the graph on the first attempt; not planned.\"\nRuntime evidence: unknown \u2014 payload size, mutability between attempts, and graph compute cost are not measurable here; no source in the reviewed repo. With D6 approved, worst case is 5 fetch+recompute cycles per job, spread over ~20 minutes.\nState: pending\nComparison grid:\n\n| Choice | Current | A | B | C |\n|---|---|---|---|---|\n| R1-R6 | approved (D4-D9) | fixed | fixed | fixed |\n| R7 payload fetch per retry | full re-fetch every attempt | re-fetch (one row; keeps correctness if payload can change) | re-fetch | re-fetch |\n| R7 graph compute per retry | recompute every attempt | compute on attempt 1, store keyed by (job id, payload version/hash); reuse while version matches, recompute on mismatch | recompute every attempt (as proposed) | add timing instrumentation only; decide after one week of data |\n| R7 cache storage | none | alongside job (job metadata field or sidecar table), evicted when job completes/dead-letters | none | none |\n| R7 tests | none | 3 (miss, hit, version mismatch) | none | none |\n\nQuestion D10:\nD10 \u2014 Cache the dependency graph across retries, keep recomputing, or measure first?\nProject/branch/task: main branch; PLAN.md \"Add background job retry framework\", Performance section.\nELI10: Every retry currently reloads the job's payload from the database and rebuilds a dependency graph by walking that payload. With the approved policy that is at most five rebuilds per job, so this is not an outage risk on its own. It is waste, and for large payloads it is waste on the exact retry path that is already under stress. Caching the graph is safe only if a stale graph cannot be used: keying the cache on a payload version (or hash) means a changed payload forces a rebuild, and re-fetching the small payload row each time stays cheap.\nStakes if we pick wrong: caching without a version check serves a stale graph after a payload edit; not caching means retries of big jobs pay full price while the downstream is already struggling.\nRecommendation: A because a version-keyed cache is a handful of lines with three tests, keeps re-fetch correctness, and removes the O(payload) walk from attempts 2-5 where it matters most.\nCompleteness: A=10/10, B=5/10, C=6/10\nPros / cons:\nA) Cache graph on first attempt, keyed by payload version; recompute on mismatch (recommended)\n \u2705 Attempts 2-5 skip the O(payload) walk; the version key makes a stale graph impossible by construction\n \u2705 Payload re-fetch is unchanged, so any correctness that depends on fresh payload data is preserved\n \u2705 Three tests (miss, hit, mismatch) pin the behavior (human: ~1 day / CC: ~15 min)\n \u274c One more piece of per-job state to store and evict when the job finishes or dead-letters\nB) Recompute every retry (as proposed)\n \u2705 Zero new state; simplest possible code path\n \u274c Retries of large payloads do 5x the graph work exactly when the system is already failing; no upper bound on payload size stated\nC) Instrument only: add timing metrics for fetch + graph build, decide after data\n \u2705 Decision grounded in real numbers instead of guesses about payload size\n \u2705 Metrics are useful regardless of the eventual choice\n \u274c Leaves the 5x work in place for now and requires a second change later; the cache is cheap enough that measuring first buys little\nNet: A is small and self-protecting; B is fine only if payloads are known to be tiny, which the plan does not say.\nHeader: Graph cache\nOptions:\nA) Cache graph on first attempt, keyed by payload version; recompute on mismatch (recommended)\n\u2705 Attempts 2-5 skip the O(payload) walk; the version key makes a stale graph impossible by construction\n\u2705 Payload re-fetch is unchanged, so any correctness that depends on fresh payload data is preserved\n\u2705 Three tests (miss, hit, mismatch) pin the behavior (human: ~1 day / CC: ~15 min)\n\u274c One more piece of per-job state to store and evict when the job finishes or dead-letters\nB) Recompute every retry (as proposed)\n\u2705 Zero new state; simplest possible code path\n\u274c Retries of large payloads do 5x the graph work exactly when the system is already failing; no upper bound on payload size stated\nC) Instrument only: add timing metrics for fetch + graph build, decide after data\n\u2705 Decision grounded in real numbers instead of guesses about payload size\n\u2705 Metrics are useful regardless of the eventual choice\n\u274c Leaves the 5x work in place for now and requires a second change later; the cache is cheap enough that measuring first buys little\nActual answer: unanswered\nAccepted scope: none\nHistory: none\n", "snapshot": { "at": "2026-09-16T06:03:59.029Z", "kind": "owned-plan-or-report", "source": "/tmp/g-pyj7vm4y/gstack-paid-shard-qDJLTL/tmp/gstack-e2e-plan-eng-batching-IpzCNz/gstack-test-plan-eng-batching.md", "artifact": "objects/85d76350e65834a60bce86fefeb9fc041b06c44f187ddf725c0de7d4dd45cd04.md", "sha256": "85d76350e65834a60bce86fefeb9fc041b06c44f187ddf725c0de7d4dd45cd04", "bytes": 43219, "mtimeMs": 1789538638728.3037, "provenance": "Exact observed file bytes; never reconstructed from tool text." } }, { "fingerprint": { "signature": "8c2809a2-42f4-479b-a5e4-a6f2e5016e61:toolu_01EgPWFprwF6HMxMLywZr4CN", "promptSnippet": "TODO obs D11 \u2014 Capture \"retry observability dashboard + rate alerts\" as a TODO, skip it, or build it now? Project/branch/task: main branch; PLAN.md \"Add background job retry framework\", follow-up work. ELI10: You approved one alert when a j", "options": [ { "index": 1, "label": "Add to TODOS.md (recommended)" }, { "index": 2, "label": "Skip, not valuable enough" }, { "index": 3, "label": "Build it now in this change" } ], "observedAtMs": 743315, "preReview": true, "nativeCall": { "sessionId": "8c2809a2-42f4-479b-a5e4-a6f2e5016e61", "toolUseId": "toolu_01EgPWFprwF6HMxMLywZr4CN", "questions": [ { "question": "D11 \u2014 Capture \"retry observability dashboard + rate alerts\" as a TODO, skip it, or build it now?\nProject/branch/task: main branch; PLAN.md \"Add background job retry framework\", follow-up work.\nELI10: You approved one alert when a job exhausts its retries. That tells you about individual poisoned jobs, not about trends: a downstream that starts failing 30% of webhook attempts will show up as lots of log lines and, later, a pile of dead-letters, but nothing pages you on the retry rate itself. A small dashboard (retry rate per job type, dead-letter depth, p95 attempts to success) plus one rate alert would catch that early. It is real work with its own design questions (which metrics stack, what thresholds), so it does not belong in this change.\nStakes if we pick wrong: skipping it means the first sign of a degraded downstream is a full dead-letter queue; building it now widens a change that already touches five workers.\nRecommendation: A because the work is valuable, clearly separate, and a TODO with context keeps the reasoning from being lost.\nNote: options differ in kind, not coverage \u2014 no completeness score.\nNet: capture now, build after the framework has a week of real attempt data to set thresholds from.", "header": "TODO obs", "multiSelect": false, "options": [ { "label": "Add to TODOS.md (recommended)", "description": "\u2705 Captures the why (rate-level blindness after this change) and where to start (attempt log lines already carry the fields)\n\u2705 Keeps this change focused on the retry framework and its tests\n\u274c TODOS.md does not exist yet in this repo; creating it is a small extra file" }, { "label": "Skip, not valuable enough", "description": "\u2705 Nothing to maintain; the exhaustion alert may be enough for current volume\n\u274c The idea and its reasoning are lost until the first slow-burn downstream incident" }, { "label": "Build it now in this change", "description": "\u2705 Ships observability with the feature it observes\n\u274c Adds metrics-stack and threshold decisions to a change that is already five workers wide" } ] } ], "answered": true, "failed": false, "answers": { "D11 \u2014 Capture \"retry observability dashboard + rate alerts\" as a TODO, skip it, or build it now?\nProject/branch/task: main branch; PLAN.md \"Add background job retry framework\", follow-up work.\nELI10: You approved one alert when a job exhausts its retries. That tells you about individual poisoned jobs, not about trends: a downstream that starts failing 30% of webhook attempts will show up as lots of log lines and, later, a pile of dead-letters, but nothing pages you on the retry rate itself. A small dashboard (retry rate per job type, dead-letter depth, p95 attempts to success) plus one rate alert would catch that early. It is real work with its own design questions (which metrics stack, what thresholds), so it does not belong in this change.\nStakes if we pick wrong: skipping it means the first sign of a degraded downstream is a full dead-letter queue; building it now widens a change that already touches five workers.\nRecommendation: A because the work is valuable, clearly separate, and a TODO with context keeps the reasoning from being lost.\nNote: options differ in kind, not coverage \u2014 no completeness score.\nNet: capture now, build after the framework has a week of real attempt data to set thresholds from.": "Add to TODOS.md (recommended)" }, "unansweredQuestionIndices": [], "answeredAt": "2026-09-16T06:05:35.735Z" } }, "questionAt": "2026-09-16T06:05:35.483Z", "answeredAt": "2026-09-16T06:05:35.735Z", "preAskPlan": "# Eng Review: Add background job retry framework\n\nReviewed target: `PLAN.md` (repo root, commit 7ce6818, branch main).\nReview: `/plan-eng-review`, 2026-09-16. Report file requested by the user.\n\n## Original plan (as reviewed)\n\n> ## Architecture\n> We'll roll a custom exponential-backoff scheduler inline in each worker\n> rather than use the existing job library's built-in retry hooks. Same\n> shape as the library version, but we want full control over the curve.\n>\n> ## Code quality\n> The retry envelope (compute delay, log attempt, dispatch) is duplicated\n> across 5 worker files with copy-pasted bodies. We will leave the\n> duplication for now and refactor \"later.\"\n>\n> ## Tests\n> The existing `processWebhookJob()` flow gets rewritten as part of this\n> change. No regression test for the prior at-most-once delivery guarantee\n> is planned.\n>\n> ## Performance\n> On every retry we re-fetch the full job payload from the database, then\n> iterate the payload to recompute the dependency graph. Could cache the\n> graph on the first attempt; not planned.\n\n## Step 0: Scope Challenge findings\n\n1. [P1] (confidence: 8/10) PLAN.md:6-8 \u2014 custom inline scheduler rebuilds the job library's built-in retry hooks (Layer 1). Custom backoff curves are a supported extension point. Disposition: pending (R1 / D4).\n2. [P1] (confidence: 9/10) PLAN.md:11-13 \u2014 retry envelope copy-pasted across 5 worker files; refactor deferred. Disposition: pending (Code Quality).\n3. [P0] (confidence: 9/10) PLAN.md:16-18 \u2014 `processWebhookJob()` rewritten with no regression test; retries change the guarantee from at-most-once to at-least-once with no idempotency key. Disposition: pending (Architecture + Tests).\n4. [P2] (confidence: 7/10) PLAN.md:21-23 \u2014 payload re-fetched and dependency graph recomputed on every retry. Disposition: pending (Performance).\n5. [P1] (confidence: 8/10) PLAN.md:6-8 \u2014 \"full control over the curve\" but no curve specified: no max attempts, delay cap, jitter, dead-letter, or retryable-error classification. Disposition: pending (Architecture).\n6. Factual gap: no design doc / problem statement (which failure classes retries target). Recorded; folded into finding 5.\n\nComplexity gate: not triggered (5 worker files + 1 rewrite, 0 new classes). TODOS.md: absent. Distribution: no new artifact.\n\n## Decision ledger\n\n### R1: Retry mechanism \u2014 library built-in retry hooks vs custom inline scheduler\nFinding: #1, P1, confidence 8/10, PLAN.md:6-8, reviewer: Claude (plan-eng-review)\nPlan baseline: original proposal \u2014 custom exponential-backoff scheduler inline in each worker; library hooks unused.\nRuntime evidence: unknown \u2014 the repo under review contains only PLAN.md and CLAUDE.md; the job library and worker files are not present. Web search (2026) shows mainstream job libraries expose a custom backoff function / retry_in hook.\nState: approved (D4) \u2014 see Actual answer below\nComparison grid:\n\n| Choice | Current | A | B |\n|---|---|---|---|\n| R1 retry mechanism | custom inline scheduler per worker (proposed) | library built-in retry hooks with one shared custom backoff function | custom inline scheduler per worker (as proposed) |\n| R2 envelope duplication | 5 copies, deferred (pending) | pending | pending |\n| R3 webhook delivery semantics | unspecified (pending) | pending | pending |\n| R4 retry policy bounds | unspecified (pending) | pending | pending |\n| R5 webhook regression contract | none planned (pending) | pending | pending |\n| R7 graph recompute per retry | recompute every retry (pending) | pending | pending |\n\nQuestion D4:\nD4 \u2014 Retry mechanism: use the job library's built-in retry hooks, or roll a custom inline scheduler?\nProject/branch/task: main branch; PLAN.md \"Add background job retry framework\", Architecture section.\nELI10: The plan rebuilds retry scheduling inside each worker even though the job library already has retry hooks. The stated reason is \"full control over the curve\", but these libraries let you plug in your own backoff function, so you get the exact curve you want while the library keeps doing the hard parts: persisting attempt counts across worker crashes, scheduling the delayed re-run, and exposing retries to its dashboard and metrics. A hand-rolled inline scheduler has to re-solve all of that in five places.\nStakes if we pick wrong: a worker that crashes mid-sleep in a custom scheduler loses the job silently; a library-managed retry survives the crash and is visible in the queue UI.\nRecommendation: A because the library hook gives the same curve control with crash-safe attempt state and zero new scheduling code; the custom path only wins if the library has no backoff hook at all, which the plan does not claim.\nCompleteness: A=10/10, B=6/10\nPros / cons:\nA) Library retry hooks + one custom backoff function (recommended)\n \u2705 Attempt count and next-run time persist in the queue, so a worker crash mid-backoff cannot lose the job\n \u2705 Retries show up in the library's dashboard, metrics and dead-letter handling for free\n \u2705 One backoff function to test instead of five inline schedulers (human: ~1 day / CC: ~20 min)\n \u274c Curve is bounded by whatever the hook API allows (delay per attempt); exotic behavior like per-tenant curves needs a wrapper\nB) Custom inline scheduler in each worker (as proposed)\n \u2705 Unconstrained control over every retry decision inside worker code\n \u2705 No dependency on the library's hook signature or its upgrade path\n \u274c Attempt state lives in process memory unless you build persistence; a crash or deploy mid-backoff drops the retry\n \u274c Five schedulers to keep in sync, invisible to the queue UI and metrics (human: ~3 days / CC: ~45 min)\nNet: same curve either way; the trade is crash-safety and observability for free vs. rebuilding the library's job.\nHeader: Retry mech\nOptions:\nA) Library retry hooks + one custom backoff function (recommended)\n\u2705 Attempt count and next-run time persist in the queue, so a worker crash mid-backoff cannot lose the job\n\u2705 Retries show up in the library's dashboard, metrics and dead-letter handling for free\n\u2705 One backoff function to test instead of five inline schedulers (human: ~1 day / CC: ~20 min)\n\u274c Curve is bounded by whatever the hook API allows (delay per attempt); exotic behavior like per-tenant curves needs a wrapper\nB) Custom inline scheduler in each worker (as proposed)\n\u2705 Unconstrained control over every retry decision inside worker code\n\u2705 No dependency on the library's hook signature or its upgrade path\n\u274c Attempt state lives in process memory unless you build persistence; a crash or deploy mid-backoff drops the retry\n\u274c Five schedulers to keep in sync, invisible to the queue UI and metrics (human: ~3 days / CC: ~45 min)\nActual answer: A) Library retry hooks + one custom backoff function \u2014 user answer to D4\nAccepted scope: Replace the custom inline scheduler with the job library's built-in retry hooks; implement one shared custom backoff function that the hooks call; delete the plan's inline per-worker scheduler. Scope Challenge result: scope reduced per recommendation.\nHistory: original proposal (custom inline scheduler) superseded by D4.\nState: approved\n\n### R3: Delivery semantics for `processWebhookJob()` under retries\nFinding: #3, P0, confidence 9/10, PLAN.md:16-18, reviewer: Claude (plan-eng-review)\nPlan baseline: original proposal \u2014 processWebhookJob() is rewritten to retry; the prior at-most-once guarantee is neither preserved nor replaced with an explicit alternative.\nRuntime evidence: unknown \u2014 worker source not present in the reviewed repo. The plan itself states the current guarantee is at-most-once (PLAN.md:17). Web sources (Svix, Hookdeck, 2026) agree that retries imply at-least-once and require an idempotency key to be safe.\nState: approved (D5) \u2014 see Actual answer below\nComparison grid:\n\n| Choice | Current | A | B | C |\n|---|---|---|---|---|\n| R1 retry mechanism | approved: library hooks + shared backoff fn (D4) | fixed | fixed | fixed |\n| R3 webhook delivery semantics | at-most-once today; plan retries without stating new guarantee | at-least-once with idempotency key (provider event ID, or hash of raw body) claimed atomically before side effects; duplicates return success | keep at-most-once: exclude processWebhookJob from retries (max attempts = 1); other 4 workers retry | retry with no idempotency key (as proposed) |\n| R2 envelope duplication | pending | pending | pending | pending |\n| R4 retry policy bounds | pending | pending | pending | pending |\n| R5 webhook regression contract | pending | pending | pending | pending |\n| R7 graph recompute per retry | pending | pending | pending | pending |\n\nQuestion D5:\nD5 \u2014 When processWebhookJob() gains retries, what delivery guarantee does it make?\nProject/branch/task: main branch; PLAN.md \"Add background job retry framework\", Tests/Architecture sections.\nELI10: Today the webhook job runs at most once: if it fails, nothing is redelivered, so it can never double-fire. The moment you add retries, a job that timed out after the receiver already processed it will run again, and the receiver sees the same webhook twice. That is a real behavior change, not just a missing test. The standard fix is an idempotency key: record the event ID before doing the side effect, and if the key is already there, treat the retry as already done.\nStakes if we pick wrong: duplicate webhooks mean double charges, double emails, or double state changes for whoever receives them, and you find out from a customer, not a test.\nRecommendation: A because retrying is the point of this plan, and an idempotency key is a few lines plus one unique index; it turns duplicates from an incident into a no-op.\nCompleteness: A=10/10, B=7/10, C=3/10\nPros / cons:\nA) At-least-once + idempotency key (recommended)\n \u2705 Transient failures get retried and duplicates are absorbed, so receivers never see a double effect\n \u2705 Small change: atomic INSERT ... ON CONFLICT DO NOTHING (or SET NX) on the event ID before dispatch (human: ~1 day / CC: ~15 min)\n \u274c Needs a dedup store with a TTL at least as long as the retry window, and a decision on what the key is when the payload has no natural ID\nB) Keep at-most-once: no retries for the webhook job\n \u2705 Zero semantic change for webhook receivers; regression risk drops to the rewrite itself\n \u2705 Simplest to prove: max attempts = 1 for this job, existing behavior stays\n \u274c Webhook deliveries that fail on a transient blip are still lost, which is the pain this plan exists to fix\nC) Retry with no idempotency key (as proposed)\n \u2705 No extra storage or key design work up front\n \u274c Duplicate side effects on every timeout-after-success; the at-most-once guarantee silently disappears\nNet: you are choosing between fixing lost webhooks safely, not fixing them, or fixing them in a way that creates duplicates.\nHeader: Webhook sem\nOptions:\nA) At-least-once + idempotency key (recommended)\n\u2705 Transient failures get retried and duplicates are absorbed, so receivers never see a double effect\n\u2705 Small change: atomic INSERT ... ON CONFLICT DO NOTHING (or SET NX) on the event ID before dispatch (human: ~1 day / CC: ~15 min)\n\u274c Needs a dedup store with a TTL at least as long as the retry window, and a decision on what the key is when the payload has no natural ID\nB) Keep at-most-once: no retries for the webhook job\n\u2705 Zero semantic change for webhook receivers; regression risk drops to the rewrite itself\n\u2705 Simplest to prove: max attempts = 1 for this job, existing behavior stays\n\u274c Webhook deliveries that fail on a transient blip are still lost, which is the pain this plan exists to fix\nC) Retry with no idempotency key (as proposed)\n\u2705 No extra storage or key design work up front\n\u274c Duplicate side effects on every timeout-after-success; the at-most-once guarantee silently disappears\nActual answer: A) At-least-once + idempotency key \u2014 user answer to D5\nAccepted scope: processWebhookJob() becomes at-least-once. Before any side effect, atomically claim an idempotency key (provider event ID when present; otherwise SHA-256 of the raw body) via INSERT ... ON CONFLICT DO NOTHING or equivalent; if the claim loses, return success without re-dispatching. Dedup record TTL >= the full retry window. Tests and docs for this contract are carried into Test review (R5) as required proof, not a new decision.\nHistory: original proposal (retry with no idempotency key) superseded by D5.\nState: approved\n\n### R4: Retry policy completeness \u2014 bounds the plan must specify\nFinding: #5, P1, confidence 8/10, PLAN.md:6-8, reviewer: Claude (plan-eng-review)\nPlan baseline: original proposal \u2014 \"full control over the curve\" with no max attempts, delay cap, jitter, dead-letter behavior, or retryable/terminal error classification stated.\nRuntime evidence: unknown \u2014 no worker or library source in the reviewed repo. Web sources (2026) name uncapped delay and missing jitter as the two standard footguns; retry storms and ~12-day delays at attempt 20.\nState: pending\nComparison grid:\n\n| Choice | Current | A | B | C |\n|---|---|---|---|---|\n| R1 retry mechanism | approved: library hooks + shared backoff fn (D4) | fixed | fixed | fixed |\n| R3 webhook semantics | approved: at-least-once + idempotency key (D5) | fixed | fixed | fixed |\n| R4 policy: max attempts | unspecified | 5 (per-job override allowed) | 5 (per-job override allowed) | unspecified |\n| R4 policy: delay cap | unspecified | min(cap, base*2^n), cap = 10 min | unspecified | unspecified |\n| R4 policy: jitter | unspecified | full jitter: random(0, computed delay) | unspecified | unspecified |\n| R4 policy: exhaustion | unspecified | move to library dead-letter/failed set + alert | job marked failed, no alert | unspecified |\n| R4 policy: error classification | unspecified | retry only on transient errors (timeout, 5xx, connection reset); terminal errors (4xx validation, malformed payload) fail immediately | all errors retried | all errors retried |\n| R2 envelope duplication | pending | pending | pending | pending |\n| R5 webhook regression contract | pending | pending | pending | pending |\n| R7 graph recompute per retry | pending | pending | pending | pending |\n\nQuestion D6:\nD6 \u2014 How completely should the plan specify the retry policy (attempts, cap, jitter, exhaustion, error classes)?\nProject/branch/task: main branch; PLAN.md \"Add background job retry framework\", Architecture section.\nELI10: Exponential backoff is only half a policy. Without a cap, attempt 20 waits about 12 days. Without jitter, every job that failed at the same moment retries at the same moment and hammers the downstream again (a retry storm). Without a max, a permanently broken payload retries forever. And retrying a \"400 invalid payload\" is pointless: only transient failures (timeouts, 5xx, dropped connections) deserve a retry. The plan says \"full control over the curve\" but writes none of these down, so each of the five workers will pick its own answer.\nStakes if we pick wrong: a poisoned job retrying forever, or a downstream outage made worse by synchronized retries; both look like mystery load at 3am.\nRecommendation: A because every one of these is a one-line constant or a small predicate in the single shared backoff function you already approved, and the defaults below are the standard ones (adjust the numbers freely via Other).\nCompleteness: A=10/10, B=6/10, C=3/10\nPros / cons:\nA) Full policy: max 5 attempts, 10-min cap, full jitter, dead-letter + alert on exhaustion, retry only transient errors (recommended)\n \u2705 Bounded worst case: a broken job costs at most 5 attempts and ~20 minutes, then lands in a dead-letter set someone gets paged for\n \u2705 Full jitter spreads synchronized failures so a downstream blip does not become a retry storm\n \u2705 Terminal errors fail fast, so bad payloads surface immediately instead of after 5 retries (human: ~1 day / CC: ~15 min)\n \u274c Five constants and one error predicate to agree on now; per-job overrides add a little config surface\nB) Max attempts only (5), everything else unspecified\n \u2705 Prevents infinite retries with one constant\n \u2705 Smallest possible addition to the plan text\n \u274c Uncapped delay and no jitter still allow retry storms and multi-hour waits within those 5 attempts; terminal errors waste retries\nC) Curve only, as proposed\n \u2705 No policy decisions needed before coding starts\n \u274c Each worker author invents their own bounds (or none); unbounded delay and infinite retries are possible\nNet: the difference between A and C is about 15 lines in one shared function versus five different implicit policies you discover in production.\nHeader: Retry policy\nOptions:\nA) Full policy: max 5 attempts, 10-min cap, full jitter, dead-letter + alert on exhaustion, retry only transient errors (recommended)\n\u2705 Bounded worst case: a broken job costs at most 5 attempts and ~20 minutes, then lands in a dead-letter set someone gets paged for\n\u2705 Full jitter spreads synchronized failures so a downstream blip does not become a retry storm\n\u2705 Terminal errors fail fast, so bad payloads surface immediately instead of after 5 retries (human: ~1 day / CC: ~15 min)\n\u274c Five constants and one error predicate to agree on now; per-job overrides add a little config surface\nB) Max attempts only (5), everything else unspecified\n\u2705 Prevents infinite retries with one constant\n\u2705 Smallest possible addition to the plan text\n\u274c Uncapped delay and no jitter still allow retry storms and multi-hour waits within those 5 attempts; terminal errors waste retries\nC) Curve only, as proposed\n\u2705 No policy decisions needed before coding starts\n\u274c Each worker author invents their own bounds (or none); unbounded delay and infinite retries are possible\nActual answer: A) Full policy \u2014 user answer to D6\nAccepted scope: The shared backoff function implements: max 5 attempts (per-job override allowed); delay = min(10 min, base * 2^attempt) with full jitter random(0, delay); on exhaustion move the job to the library's dead-letter/failed set and emit an alert; retry only transient errors (timeout, 5xx, connection reset), fail immediately on terminal errors (4xx validation, malformed payload). Tests proving each bound are required proof carried into Test review (R6).\nHistory: original proposal (curve only, no bounds) superseded by D6.\nState: approved\n\n### R2: Retry envelope duplication across 5 worker files\nFinding: #2, P1, confidence 9/10, PLAN.md:11-13, reviewer: Claude (plan-eng-review)\nPlan baseline: original proposal \u2014 envelope (compute delay, log attempt, dispatch) copy-pasted in 5 worker files; refactor deferred to \"later\".\nRuntime evidence: unknown \u2014 worker files not present in the reviewed repo; the plan itself states the duplication exists (PLAN.md:11-12).\nState: pending\nComparison grid:\n\n| Choice | Current | A | B | C |\n|---|---|---|---|---|\n| R1 retry mechanism | approved: library hooks + shared backoff fn (D4) | fixed | fixed | fixed |\n| R3 webhook semantics | approved: at-least-once + idempotency key (D5) | fixed | fixed | fixed |\n| R4 retry policy | approved: full policy in shared fn (D6) | fixed | fixed | fixed |\n| R2 envelope: where it lives | 5 inline copies | one shared module: backoff fn + attempt-logging middleware registered once with the library | one shared module, but only workers touched by this change migrate | 5 inline copies remain |\n| R2 envelope: workers migrated in this change | 0 | all 5 | only those edited for retries (webhook + any other touched) | 0 |\n| R5 webhook regression contract | pending | pending | pending | pending |\n| R7 graph recompute per retry | pending | pending | pending | pending |\n\nQuestion D7:\nD7 \u2014 Consolidate the retry envelope into one shared module now, or leave five copies and refactor later?\nProject/branch/task: main branch; PLAN.md \"Add background job retry framework\", Code quality section.\nELI10: The same three steps (work out the delay, log the attempt, hand the job back to the queue) are pasted into five worker files. You already decided the delay logic becomes one shared backoff function the library calls, so most of that envelope disappears on its own. What remains is deciding whether the attempt logging and hook registration also live in one place, and whether all five workers move over in this change or only the ones you happen to touch. \"Refactor later\" for copy-pasted code almost always means the first backoff bug gets fixed in four files and missed in the fifth.\nStakes if we pick wrong: a policy fix (say, the jitter formula) lands in some workers and not others, and the inconsistent ones are the ones nobody tests.\nRecommendation: A because with the library-hook decision the shared module is small (one backoff function plus one logging middleware), and migrating five call sites is a mechanical change that CC does in minutes.\nCompleteness: A=10/10, B=7/10, C=3/10\nPros / cons:\nA) One shared retry module, all 5 workers migrated now (recommended)\n \u2705 Exactly one place defines delay, jitter, attempt logging and exhaustion handling; one test suite proves it\n \u2705 Deleting five inline envelopes shrinks the diff and removes the code the plan admits is copy-pasted (human: ~1.5 days / CC: ~25 min)\n \u274c Touches all five worker files in one change, so the review surface is wider than the webhook rewrite alone\nB) Shared module, migrate only workers this change already touches\n \u2705 Smaller immediate diff; untouched workers carry no risk from this change\n \u2705 The shared module exists, so later migrations are copy-and-delete\n \u274c Two retry implementations coexist in production; the inline ones keep their unbounded, unjittered behavior until someone remembers\nC) Leave five copies, refactor \"later\" (as proposed)\n \u2705 No refactor risk in this change\n \u274c Five places to apply every future policy fix; the approved policy (D6) has to be pasted five times to take effect\nNet: the shared module is the approved design's natural home; the only real question is whether all five workers adopt it now or drift.\nHeader: DRY envelope\nOptions:\nA) One shared retry module, all 5 workers migrated now (recommended)\n\u2705 Exactly one place defines delay, jitter, attempt logging and exhaustion handling; one test suite proves it\n\u2705 Deleting five inline envelopes shrinks the diff and removes the code the plan admits is copy-pasted (human: ~1.5 days / CC: ~25 min)\n\u274c Touches all five worker files in one change, so the review surface is wider than the webhook rewrite alone\nB) Shared module, migrate only workers this change already touches\n\u2705 Smaller immediate diff; untouched workers carry no risk from this change\n\u2705 The shared module exists, so later migrations are copy-and-delete\n\u274c Two retry implementations coexist in production; the inline ones keep their unbounded, unjittered behavior until someone remembers\nC) Leave five copies, refactor \"later\" (as proposed)\n\u2705 No refactor risk in this change\n\u274c Five places to apply every future policy fix; the approved policy (D6) has to be pasted five times to take effect\nActual answer: A) One shared retry module, all 5 workers migrated now \u2014 user answer to D7\nAccepted scope: Create one shared retry module (e.g. `jobs/retry/` or the project's equivalent): the backoff/policy function from D6 plus an attempt-logging middleware (job id, job type, attempt n/max, error class, next delay ms) registered once with the job library. Delete the inline envelope from all 5 worker files and register them with the shared module. Existing worker behavior other than retry scheduling is unchanged.\nHistory: original proposal (leave 5 copies, refactor later) superseded by D7.\nState: approved\n\n### R5: Regression contract for the `processWebhookJob()` rewrite\nFinding: #3, P0, confidence 9/10, PLAN.md:16-18, reviewer: Claude (plan-eng-review)\nPlan baseline: original proposal \u2014 \"No regression test for the prior at-most-once delivery guarantee is planned.\" Approved behavior to protect comes from D5 (at-least-once + idempotency key) and D6 (retry only transient errors).\nRuntime evidence: unknown \u2014 no test framework or test files exist in the reviewed repo (git ls-files: CLAUDE.md, PLAN.md only). Existing callers of processWebhookJob() cannot be traced here; treat every current behavior as at risk.\nState: pending\nComparison grid:\n\n| Choice | Current | A | B |\n|---|---|---|---|\n| R1, R3, R4, R2 | approved (D4-D7) | fixed | fixed |\n| R5 behavior preserved: success dispatches exactly once with unchanged payload/headers | untested | asserted | asserted |\n| R5 behavior preserved: terminal error (4xx/malformed) \u2192 exactly one attempt, job failed, no retry (old at-most-once holds for this class) | untested | asserted | not asserted |\n| R5 intentional change: transient error \u2192 retried per D6 policy, attempt count visible | untested | asserted | asserted |\n| R5 new contract: second run with an already-claimed key \u2192 no dispatch, returns success | untested | asserted | asserted |\n| R5 new contract: key derivation \u2014 provider event ID when present, body hash when absent | untested | asserted (both branches) | not asserted |\n| R5 new contract: crash/timeout after dispatch, before ack \u2192 retry finds claimed key \u2192 zero duplicate sends [\u2192E2E] | untested | asserted via integration test against the real queue | not asserted |\n| R5: dedup TTL >= retry window | untested | asserted | not asserted |\n| R6 policy proof depth | pending | pending | pending |\n| R7 graph recompute | pending | pending | pending |\n\nQuestion D8:\nD8 \u2014 What regression contract protects processWebhookJob() through the rewrite?\nProject/branch/task: main branch; PLAN.md \"Add background job retry framework\", Tests section.\nELI10: The webhook job is being rewritten and currently has no test guarding what it does today. A regression contract is the written list of \"what must still be true afterwards\", plus the things we are changing on purpose. For this job that is: one successful run still sends exactly one webhook with the same content; a permanently bad request still fails once and stops; a transient failure now retries (the intentional change); and a retry after a timeout-that-actually-succeeded sends nothing twice. The last one is the whole point of the idempotency key you approved, and it only shows up in a test that runs against the real queue, not a mock.\nStakes if we pick wrong: the rewrite ships, a receiver gets two payment webhooks after a slow response, and there was never a test that could have caught it.\nRecommendation: A because these seven assertions are the entire behavioral surface of the job, and the ones B drops (terminal fail-fast, key fallback, crash-after-dispatch) are exactly the failure paths that produce duplicates.\nCompleteness: A=10/10, B=6/10\nPros / cons:\nA) Full contract: 7 assertions incl. terminal fail-fast, key fallback and crash-after-dispatch integration test (recommended)\n \u2705 Every branch of the new job has a named assertion; the duplicate-send path is proven against the real queue, not a mock\n \u2705 Preserves the old at-most-once behavior where it still applies (terminal errors) and documents the intentional change for transient ones\n \u2705 Tests double as the spec for the idempotency key and TTL (human: ~2 days / CC: ~30 min)\n \u274c One integration test needs a real or in-memory queue backend in CI\nB) Happy path + duplicate-claim unit tests only\n \u2705 Fast, mock-only tests; no CI infrastructure change\n \u2705 Covers the most common two paths\n \u274c Misses terminal-error fail-fast, body-hash key fallback, TTL, and the crash-after-dispatch path that actually causes duplicates\nNet: A costs one integration test fixture; B leaves the duplicate-send path, the reason D5 exists, unproven.\nHeader: Webhook tests\nOptions:\nA) Full contract: 7 assertions incl. terminal fail-fast, key fallback and crash-after-dispatch integration test (recommended)\n\u2705 Every branch of the new job has a named assertion; the duplicate-send path is proven against the real queue, not a mock\n\u2705 Preserves the old at-most-once behavior where it still applies (terminal errors) and documents the intentional change for transient ones\n\u2705 Tests double as the spec for the idempotency key and TTL (human: ~2 days / CC: ~30 min)\n\u274c One integration test needs a real or in-memory queue backend in CI\nB) Happy path + duplicate-claim unit tests only\n\u2705 Fast, mock-only tests; no CI infrastructure change\n\u2705 Covers the most common two paths\n\u274c Misses terminal-error fail-fast, body-hash key fallback, TTL, and the crash-after-dispatch path that actually causes duplicates\nActual answer: A) Full contract \u2014 user answer to D8\nAccepted scope: CRITICAL regression contract for processWebhookJob(), 7 assertions: (1) success \u2192 exactly one dispatch, payload/headers byte-identical to pre-rewrite; (2) terminal error (4xx/malformed) \u2192 exactly one attempt, job failed, no retry scheduled; (3) transient error \u2192 retry scheduled per D6 policy, attempt count incremented; (4) already-claimed key \u2192 no dispatch, job returns success; (5) key = provider event ID when present, SHA-256(raw body) when absent; (6) crash/timeout after dispatch before ack \u2192 retry finds claimed key \u2192 zero duplicate sends, integration test against a real/in-memory queue [\u2192E2E]; (7) dedup record TTL >= max total retry window (5 attempts at 10-min cap \u2248 50 min + base; assert TTL \u2265 1h or the computed window).\nHistory: original proposal (no regression test) superseded by D8.\nState: approved\n\n### R6: Proof depth for the shared retry policy module and worker migration\nFinding: #5 and #2 follow-on, P1, confidence 8/10, PLAN.md:6-13, reviewer: Claude (plan-eng-review)\nPlan baseline: D6 approved the policy and its per-bound tests as required proof; D7 approved migrating all 5 workers. Not yet decided: whether the jitter/cap curve gets a property-based test across the attempt range and whether each migrated worker gets an integration test proving the library actually calls the shared policy for it.\nRuntime evidence: unknown \u2014 no test framework detected in the reviewed repo.\nState: pending\nComparison grid:\n\n| Choice | Current | A | B |\n|---|---|---|---|\n| R1-R5 | approved (D4-D8) | fixed | fixed |\n| R6 per-bound unit tests (max attempts, cap, jitter range, exhaustion \u2192 DLQ + alert, transient vs terminal predicate) | required by D6 | included (carried) | included (carried) |\n| R6 property test: for attempt in 0..50, 0 \u2264 delay \u2264 min(cap, base*2^n), monotone cap, no overflow | not planned | included | not included |\n| R6 per-worker integration test: each of 5 workers, forced transient failure \u2192 library reschedules via shared policy, attempt log line emitted [\u2192E2E] | not planned | included (5 tests, shared fixture) | not included |\n| R6 exhaustion E2E: 5 forced failures \u2192 job in dead-letter set, alert emitted once [\u2192E2E] | not planned | included | not included |\n| R7 graph recompute | pending | pending | pending |\n\nQuestion D9:\nD9 \u2014 How deeply do we prove the shared retry policy and the five worker migrations?\nProject/branch/task: main branch; PLAN.md \"Add background job retry framework\", Tests section.\nELI10: You already approved unit tests for each policy bound. Two extra layers are on the table. First, a property test: instead of checking attempt 1, 2, 3 by hand, run the delay function for every attempt from 0 to 50 and assert it never exceeds the cap, never goes negative, and never overflows (2^50 milliseconds is a real number that bites). Second, integration tests: for each of the five workers, force one transient failure and confirm the queue library actually reschedules it through the shared policy and logs the attempt. That last one is what proves the migration did anything; a unit test on the policy function cannot tell you a worker forgot to register it.\nStakes if we pick wrong: a worker silently keeps no-retry behavior after \"migration\", or the delay math overflows at high attempt counts, and nothing in CI notices.\nRecommendation: A because the property test is ~10 lines and the five integration tests share one fixture with the D8 crash-after-dispatch test you already approved, so the marginal CC cost is minutes.\nCompleteness: A=10/10, B=7/10\nPros / cons:\nA) Per-bound units + property test + 5 per-worker integration tests + exhaustion E2E (recommended)\n \u2705 Proves each worker is actually wired to the shared policy, not just that the policy function is correct\n \u2705 Property test catches cap/overflow/negative-delay bugs across the whole attempt range in one assertion\n \u2705 Exhaustion E2E proves the dead-letter + alert path fires exactly once (human: ~2 days / CC: ~30 min)\n \u274c Six more integration tests in CI on top of D8's fixture; slower suite by seconds\nB) Per-bound unit tests only (as carried from D6)\n \u2705 Fast, no queue backend needed for this module\n \u2705 Already covers every named bound at specific attempts\n \u274c Cannot detect an unmigrated worker, a broken hook registration, or overflow at attempts you did not hand-pick\nNet: B proves the math; A proves the math and that production uses it.\nHeader: Policy tests\nOptions:\nA) Per-bound units + property test + 5 per-worker integration tests + exhaustion E2E (recommended)\n\u2705 Proves each worker is actually wired to the shared policy, not just that the policy function is correct\n\u2705 Property test catches cap/overflow/negative-delay bugs across the whole attempt range in one assertion\n\u2705 Exhaustion E2E proves the dead-letter + alert path fires exactly once (human: ~2 days / CC: ~30 min)\n\u274c Six more integration tests in CI on top of D8's fixture; slower suite by seconds\nB) Per-bound unit tests only (as carried from D6)\n\u2705 Fast, no queue backend needed for this module\n\u2705 Already covers every named bound at specific attempts\n\u274c Cannot detect an unmigrated worker, a broken hook registration, or overflow at attempts you did not hand-pick\nActual answer: A) Per-bound units + property test + 5 per-worker integration tests + exhaustion E2E \u2014 user answer to D9\nAccepted scope: Unit tests for each D6 bound; property test over attempts 0..50 (0 \u2264 delay \u2264 min(cap, base*2^n), no negative/overflow); 5 per-worker integration tests (forced transient failure \u2192 library reschedules via shared policy, attempt log emitted); one exhaustion E2E (5 failures \u2192 dead-letter set, alert exactly once). Shares the queue fixture with D8 assertion (6).\nHistory: not planned in original proposal; added by D9.\nState: approved\n\n### Test coverage diagram (Section 3)\n\nFramework: unknown \u2014 the reviewed repo contains no source, test config or test files (`git ls-files`: CLAUDE.md, PLAN.md). Every path below is a GAP today; the approved contracts (D8, D9) define the tests to write. Test Plan Artifact: `~/.gstack/projects/gstack-plan-count-vjhjkU/vercel-sandbox-main-eng-review-test-plan-20260916-060250.md`.\n\n```\nCODE PATHS USER / OPERATOR FLOWS\n[+] jobs/retry/policy (shared, D4/D6/D7) [+] Transient webhook failure\n \u251c\u2500\u2500 backoff(attempt, err) \u251c\u2500\u2500 [GAP] [\u2192E2E] attempt 1 times out \u2192 retried \u2192 attempt 2 succeeds\n \u2502 \u251c\u2500\u2500 [GAP] attempt < max \u2192 delay = min(cap, base*2^n) \u2502 \u2192 receiver sees exactly ONE webhook\n \u2502 \u251c\u2500\u2500 [GAP] full jitter: 0 \u2264 delay \u2264 computed \u2514\u2500\u2500 [GAP] attempt log lines visible per attempt\n \u2502 \u251c\u2500\u2500 [GAP] property: attempts 0..50 in range, no overflow\n \u2502 \u251c\u2500\u2500 [GAP] attempt >= max \u2192 exhaustion [+] Poisoned job\n \u2502 \u2502 \u251c\u2500\u2500 [GAP] [\u2192E2E] job \u2192 dead-letter/failed set \u251c\u2500\u2500 [GAP] [\u2192E2E] 5 attempts \u2248 20 min \u2192 dead-letter\n \u2502 \u2502 \u2514\u2500\u2500 [GAP] [\u2192E2E] alert emitted exactly once \u2514\u2500\u2500 [GAP] operator sees it in queue UI + one alert\n \u2502 \u2514\u2500\u2500 isTransient(err)\n \u2502 \u251c\u2500\u2500 [GAP] timeout / 5xx / conn reset \u2192 true [+] Bad payload\n \u2502 \u2514\u2500\u2500 [GAP] 4xx validation / malformed \u2192 false \u2514\u2500\u2500 [GAP] fails once, no retries, error visible immediately\n \u2514\u2500\u2500 attemptLogger middleware\n \u2514\u2500\u2500 [GAP] logs job id, type, n/max, error class, delay [+] Worker crash mid-retry (D4 rationale)\n \u2514\u2500\u2500 [GAP] [\u2192E2E] queue-persisted attempt survives restart\n[+] workers/* (5 files, D7 migration)\n \u2514\u2500\u2500 [GAP] [\u2192E2E] x5: forced transient failure \u2192 library [+] Duplicate delivery (D5)\n reschedules via shared policy (one per worker) \u251c\u2500\u2500 [GAP] [\u2192E2E] timeout-after-success retry \u2192 zero duplicates\n \u2514\u2500\u2500 [GAP] same key twice \u2192 second is a no-op success\n[+] processWebhookJob() (rewrite, D5/D8)\n \u251c\u2500\u2500 deriveKey(payload)\n \u2502 \u251c\u2500\u2500 [GAP] provider event ID present \u2192 key = ID\n \u2502 \u2514\u2500\u2500 [GAP] absent \u2192 key = sha256(raw body)\n \u251c\u2500\u2500 claimKey(key) INSERT ... ON CONFLICT DO NOTHING\n \u2502 \u251c\u2500\u2500 [GAP] claim won \u2192 dispatch\n \u2502 \u2514\u2500\u2500 [GAP] claim lost \u2192 return success, no dispatch\n \u251c\u2500\u2500 dispatch()\n \u2502 \u251c\u2500\u2500 [GAP] success \u2192 exactly one send, payload/headers byte-identical\n \u2502 \u251c\u2500\u2500 [GAP] transient error \u2192 retry scheduled (D6)\n \u2502 \u2514\u2500\u2500 [GAP] terminal error \u2192 fail, one attempt (old at-most-once preserved for this class)\n \u251c\u2500\u2500 [GAP] [\u2192E2E] crash after dispatch before ack \u2192 retry finds claimed key \u2192 0 duplicates CRITICAL\n \u2514\u2500\u2500 [GAP] dedup TTL \u2265 retry window (\u2265 ~1h)\n\n[+] dependency graph per retry (R7 \u2014 pending D10)\n \u251c\u2500\u2500 [GAP] cache miss (first attempt) \u2192 compute + store\n \u251c\u2500\u2500 [GAP] cache hit, payload version unchanged \u2192 reuse\n \u2514\u2500\u2500 [GAP] payload version changed \u2192 recompute\n\nCOVERAGE: 0/32 paths tested (0%) | Code paths: 0/24 (0%) | User flows: 0/8 (0%)\nQUALITY: \u2605\u2605\u2605:0 \u2605\u2605:0 \u2605:0 | GAPS: 32 (11 E2E, 0 eval) | Approved as required tests: 29 (R7's 3 pending D10)\n```\n\nLegend: \u2605\u2605\u2605 behavior + edge + error | \u2605\u2605 happy path | \u2605 smoke | [\u2192E2E] = integration test against a real/in-memory queue. LLM/eval scope: none (no prompt or LLM code in this plan).\n\n### R7: Dependency graph recomputed on every retry\nFinding: #4, P2, confidence 7/10, PLAN.md:21-23, reviewer: Claude (plan-eng-review)\nPlan baseline: original proposal \u2014 \"On every retry we re-fetch the full job payload from the database, then iterate the payload to recompute the dependency graph. Could cache the graph on the first attempt; not planned.\"\nRuntime evidence: unknown \u2014 payload size, mutability between attempts, and graph compute cost are not measurable here; no source in the reviewed repo. With D6 approved, worst case is 5 fetch+recompute cycles per job, spread over ~20 minutes.\nState: pending\nComparison grid:\n\n| Choice | Current | A | B | C |\n|---|---|---|---|---|\n| R1-R6 | approved (D4-D9) | fixed | fixed | fixed |\n| R7 payload fetch per retry | full re-fetch every attempt | re-fetch (one row; keeps correctness if payload can change) | re-fetch | re-fetch |\n| R7 graph compute per retry | recompute every attempt | compute on attempt 1, store keyed by (job id, payload version/hash); reuse while version matches, recompute on mismatch | recompute every attempt (as proposed) | add timing instrumentation only; decide after one week of data |\n| R7 cache storage | none | alongside job (job metadata field or sidecar table), evicted when job completes/dead-letters | none | none |\n| R7 tests | none | 3 (miss, hit, version mismatch) | none | none |\n\nQuestion D10:\nD10 \u2014 Cache the dependency graph across retries, keep recomputing, or measure first?\nProject/branch/task: main branch; PLAN.md \"Add background job retry framework\", Performance section.\nELI10: Every retry currently reloads the job's payload from the database and rebuilds a dependency graph by walking that payload. With the approved policy that is at most five rebuilds per job, so this is not an outage risk on its own. It is waste, and for large payloads it is waste on the exact retry path that is already under stress. Caching the graph is safe only if a stale graph cannot be used: keying the cache on a payload version (or hash) means a changed payload forces a rebuild, and re-fetching the small payload row each time stays cheap.\nStakes if we pick wrong: caching without a version check serves a stale graph after a payload edit; not caching means retries of big jobs pay full price while the downstream is already struggling.\nRecommendation: A because a version-keyed cache is a handful of lines with three tests, keeps re-fetch correctness, and removes the O(payload) walk from attempts 2-5 where it matters most.\nCompleteness: A=10/10, B=5/10, C=6/10\nPros / cons:\nA) Cache graph on first attempt, keyed by payload version; recompute on mismatch (recommended)\n \u2705 Attempts 2-5 skip the O(payload) walk; the version key makes a stale graph impossible by construction\n \u2705 Payload re-fetch is unchanged, so any correctness that depends on fresh payload data is preserved\n \u2705 Three tests (miss, hit, mismatch) pin the behavior (human: ~1 day / CC: ~15 min)\n \u274c One more piece of per-job state to store and evict when the job finishes or dead-letters\nB) Recompute every retry (as proposed)\n \u2705 Zero new state; simplest possible code path\n \u274c Retries of large payloads do 5x the graph work exactly when the system is already failing; no upper bound on payload size stated\nC) Instrument only: add timing metrics for fetch + graph build, decide after data\n \u2705 Decision grounded in real numbers instead of guesses about payload size\n \u2705 Metrics are useful regardless of the eventual choice\n \u274c Leaves the 5x work in place for now and requires a second change later; the cache is cheap enough that measuring first buys little\nNet: A is small and self-protecting; B is fine only if payloads are known to be tiny, which the plan does not say.\nHeader: Graph cache\nOptions:\nA) Cache graph on first attempt, keyed by payload version; recompute on mismatch (recommended)\n\u2705 Attempts 2-5 skip the O(payload) walk; the version key makes a stale graph impossible by construction\n\u2705 Payload re-fetch is unchanged, so any correctness that depends on fresh payload data is preserved\n\u2705 Three tests (miss, hit, mismatch) pin the behavior (human: ~1 day / CC: ~15 min)\n\u274c One more piece of per-job state to store and evict when the job finishes or dead-letters\nB) Recompute every retry (as proposed)\n\u2705 Zero new state; simplest possible code path\n\u274c Retries of large payloads do 5x the graph work exactly when the system is already failing; no upper bound on payload size stated\nC) Instrument only: add timing metrics for fetch + graph build, decide after data\n\u2705 Decision grounded in real numbers instead of guesses about payload size\n\u2705 Metrics are useful regardless of the eventual choice\n\u274c Leaves the 5x work in place for now and requires a second change later; the cache is cheap enough that measuring first buys little\nActual answer: A) Cache graph on first attempt, keyed by payload version; recompute on mismatch \u2014 user answer to D10\nAccepted scope: On attempt 1 compute the dependency graph and store it with the job keyed by (job id, payload version or SHA-256 of payload); on later attempts re-fetch the payload row, reuse the graph when the key matches, recompute and overwrite on mismatch; evict when the job completes or dead-letters. Tests: cache miss, cache hit, version mismatch (3, carried as required proof).\nHistory: original proposal (recompute every retry) superseded by D10.\nState: approved\n\n### Outside Voice\nCODEX_MODE: disabled (`codex_reviews=disabled`). Skipped by user config; recorded `outside_status: disabled` via gstack-review-log. No native fallback dispatched (disabled is an opt-out, not a provider failure). Cross-model tension: not applicable.\n\n### T1 (TODO candidate): Retry observability beyond the single exhaustion alert\nFinding: follow-on to R4/R6 (D6, D9), P3, confidence 7/10, reviewer: Claude (plan-eng-review)\nPlan baseline: D6 approved one alert on exhaustion and per-attempt log lines. No dashboards or rate-based alerts are in scope.\nRuntime evidence: unknown \u2014 no metrics stack visible in the reviewed repo.\nState: pending\nComparison grid:\n\n| Choice | Current | A | B | C |\n|---|---|---|---|---|\n| R1-R7 | approved (D4-D10) | fixed | fixed | fixed |\n| T1 disposition | not captured | added to TODOS.md with full context | dropped | built in this change |\n\nQuestion D11:\nD11 \u2014 Capture \"retry observability dashboard + rate alerts\" as a TODO, skip it, or build it now?\nProject/branch/task: main branch; PLAN.md \"Add background job retry framework\", follow-up work.\nELI10: You approved one alert when a job exhausts its retries. That tells you about individual poisoned jobs, not about trends: a downstream that starts failing 30% of webhook attempts will show up as lots of log lines and, later, a pile of dead-letters, but nothing pages you on the retry rate itself. A small dashboard (retry rate per job type, dead-letter depth, p95 attempts to success) plus one rate alert would catch that early. It is real work with its own design questions (which metrics stack, what thresholds), so it does not belong in this change.\nStakes if we pick wrong: skipping it means the first sign of a degraded downstream is a full dead-letter queue; building it now widens a change that already touches five workers.\nRecommendation: A because the work is valuable, clearly separate, and a TODO with context keeps the reasoning from being lost.\nNote: options differ in kind, not coverage \u2014 no completeness score.\nPros / cons:\nA) Add to TODOS.md (recommended)\n \u2705 Captures the why (rate-level blindness after this change) and where to start (attempt log lines already carry the fields)\n \u2705 Keeps this change focused on the retry framework and its tests\n \u274c TODOS.md does not exist yet in this repo; creating it is a small extra file\nB) Skip, not valuable enough\n \u2705 Nothing to maintain; the exhaustion alert may be enough for current volume\n \u274c The idea and its reasoning are lost until the first slow-burn downstream incident\nC) Build it now in this change\n \u2705 Ships observability with the feature it observes\n \u274c Adds metrics-stack and threshold decisions to a change that is already five workers wide\nNet: capture now, build after the framework has a week of real attempt data to set thresholds from.\nHeader: TODO obs\nOptions:\nA) Add to TODOS.md (recommended)\n\u2705 Captures the why (rate-level blindness after this change) and where to start (attempt log lines already carry the fields)\n\u2705 Keeps this change focused on the retry framework and its tests\n\u274c TODOS.md does not exist yet in this repo; creating it is a small extra file\nB) Skip, not valuable enough\n\u2705 Nothing to maintain; the exhaustion alert may be enough for current volume\n\u274c The idea and its reasoning are lost until the first slow-burn downstream incident\nC) Build it now in this change\n\u2705 Ships observability with the feature it observes\n\u274c Adds metrics-stack and threshold decisions to a change that is already five workers wide\nActual answer: unanswered\nAccepted scope: none\nHistory: none\n", "snapshot": { "at": "2026-09-16T06:05:17.006Z", "kind": "owned-plan-or-report", "source": "/tmp/g-pyj7vm4y/gstack-paid-shard-qDJLTL/tmp/gstack-e2e-plan-eng-batching-IpzCNz/gstack-test-plan-eng-batching.md", "artifact": "objects/d3ee6f7528d1bdb29113c4651723f090069e5d70639aef47ce636e6b8657d261.md", "sha256": "d3ee6f7528d1bdb29113c4651723f090069e5d70639aef47ce636e6b8657d261", "bytes": 47356, "mtimeMs": 1789538716180.3022, "provenance": "Exact observed file bytes; never reconstructed from tool text." } } ] }