mirror of
https://github.com/garrytan/gstack.git
synced 2026-10-04 10:26:52 +02:00
* feat: bind shared-code review advice to source and branch * feat: add shared-code extraction audit and scoped review checks * test: recognize complete source reads and explicit coverage legends * chore: bump version and changelog (v1.88.0.0) Co-Authored-By: OpenAI Codex <noreply@openai.com> * test: capture native review questions and retain public evidence Capture the actual first public native question with strict ownership and display matching. Preserve terminal failures and raw evidence, and retain SDK completion checks. * test: recognize verified review evidence and complete fixtures Recognize complete source and diagram evidence, concrete design and developer-experience decisions, and the complete planted scenario contracts. Preserve negative controls and grading thresholds. * fix: preserve decision brief structure in native questions Keep the required pros-and-cons heading and final Net field in native question text. Regenerate host outputs and document the release and evaluation repairs. Co-Authored-By: OpenAI Codex <noreply@openai.com> * docs: update project documentation for v1.88.0.0 Co-Authored-By: OpenAI Codex <noreply@openai.com> * fix: correct eval retry accounting and ship workflow gates * fix: capture native eval evidence and stabilize CI fixtures * fix: keep shared-code eval skips read-only Choose explicit no-change answers instead of mixed fix/preservation options. Reuse the bounded revalidation prompt for path fixtures so required review metadata is available without repeated discovery. Preserve source checks, retry limits, and failed native terminal outcomes. Add captured-question and callback regressions, plus evaluation selection coverage for the affected fixtures. --------- Co-authored-by: OpenAI Codex <noreply@openai.com>
36 lines
7.8 KiB
JSON
36 lines
7.8 KiB
JSON
{
|
||
"publicCall": {
|
||
"sessionId": "2e0077ea-64bf-4bdf-b3a0-ae291dd57450",
|
||
"toolUseId": "toolu_01LyJhrg8sG6wbph3zeNbEtH",
|
||
"questions": [
|
||
{
|
||
"question": "D1 — What kind of developer product is this plan shipping?\nProject/branch/task: main branch, DX review of plan.md (\"developer-friendly pricing tier\" with a `mytool pricing` CLI command and setup wizard).\nELI10: A DX review only makes sense if we agree on what developers actually touch. This plan mixes a billing change (Stripe tier, Postgres entitlements, Redis cache, upgrade endpoint) with a small CLI surface (`mytool pricing` plus a setup wizard). The classification decides which personas I offer next, which competitors I benchmark against, and what \"time to hello world\" even means here.\nStakes if we pick wrong: I benchmark the wrong thing. Treating this as a CLI when the real friction is signup and billing (or the reverse) produces scores and fixes aimed at a surface developers never see.\nRecommendation: A because the only surface a developer directly operates in this plan is the `mytool pricing` command and wizard; everything else is backend plumbing behind it.\nNote: options differ in kind, not coverage — no completeness score.\nPros / cons:\nA) CLI Tool (recommended)\n ✅ Matches the one concrete developer touchpoint the plan names: `mytool pricing` and the setup wizard\n ✅ Gives the review a measurable hello-world path: install, run the command, see the tier, upgrade\n ❌ Underweights the signup, checkout, and entitlement flow where the plan's own goal (more signups) lives\nB) Platform / hosted service\n ✅ Centers the review on account creation, billing, and entitlement checks, which is where signups are won or lost\n ✅ Surfaces the missing auth check on the upgrade endpoint as a credibility and trust issue early\n ❌ The plan describes almost none of this surface, so the review would be scoring guesses rather than the written plan\nC) Both: CLI primary, Platform secondary\n ✅ Keeps the CLI hello-world path as the spine while still auditing the upgrade and entitlement flow behind it\n ✅ Lets the persona step offer both CLI-first and evaluator personas instead of forcing one\n ❌ Roughly doubles the passes and benchmarks (human: ~1 extra hour / CC: ~5 extra min) for a plan this thin\nD) No developer-facing surface, exit and use /plan-eng-review\n ✅ Honest if `mytool pricing` is a stub and the real work is a billing backend with no DX to review\n ✅ Routes the untested, unauthenticated upgrade endpoint to the review that will actually catch it\n ❌ Abandons the CLI command and wizard, which are exactly the parts a DX review exists to sharpen\nNet: We are trading a focused review of the one written developer surface against a broader review of the billing flow the plan barely describes. <gstack-qid:plan-devex-review-product-type>",
|
||
"header": "Product type",
|
||
"options": [
|
||
{
|
||
"label": "CLI Tool (recommended)",
|
||
"description": "✅ Matches the one concrete touchpoint the plan names: `mytool pricing` and the setup wizard\n✅ Gives a measurable hello-world path: install, run the command, see the tier, upgrade\n❌ Underweights the signup, checkout, and entitlement flow where the plan's goal actually lives"
|
||
},
|
||
{
|
||
"label": "Platform / hosted service",
|
||
"description": "✅ Centers the review on account creation, billing, and entitlements, where signups are won or lost\n✅ Surfaces the missing auth check on the upgrade endpoint as a trust issue early\n❌ The plan describes almost none of this surface, so the review would score guesses, not the plan"
|
||
},
|
||
{
|
||
"label": "Both: CLI + Platform",
|
||
"description": "✅ Keeps the CLI hello-world path as the spine while auditing the upgrade and entitlement flow behind it\n✅ Lets the persona step offer both CLI-first and evaluator personas\n❌ Roughly doubles passes and benchmarks (human: ~1 extra hour / CC: ~5 extra min) for a thin plan"
|
||
},
|
||
{
|
||
"label": "No DX surface, exit",
|
||
"description": "✅ Honest if `mytool pricing` is a stub and the real work is a billing backend with no DX to review\n✅ Routes the untested, unauthenticated upgrade endpoint to /plan-eng-review, which will catch it\n❌ Abandons the CLI command and wizard, exactly the parts a DX review exists to sharpen"
|
||
}
|
||
],
|
||
"multiSelect": false
|
||
}
|
||
],
|
||
"answered": false,
|
||
"failed": false,
|
||
"source": "pre_tool_use"
|
||
},
|
||
"viewport": "│ Stakes if we pick wrong: I benchmark the wrong thing. Treating this as a CLI when the real friction is signup and\n│ billing (or the reverse) produces scores and fixes aimed at a surface developers never see.\n│ Recommendation: A because the only surface a developer directly operates in this plan is the `mytool pricing` command\n│ and wizard; everything else is backend plumbing behind it.\n│ Note: options differ in kind, not coverage — no completeness score.\n│ Pros / cons:\n│ A) CLI Tool (recommended)\n│ ✅ Matches the one concrete developer touchpoint the plan names: `mytool pricing` and the setup wizard\n│ ✅ Gives the review a measurable hello-world path: install, run the command, see the tier, upgrade\n│ ❌ Underweights the signup, checkout, and entitlement flow where the plan's own goal (more signups) lives\n│ B) Platform / hosted service\n│ ✅ Centers the review on account creation, billing, and entitlement checks, which is where signups are won or lost\n│ ✅ Surfaces the missing auth check on the upgrade endpoint as a credibility and trust issue early\n│ ❌ The plan describes almost none of this surface, so the review would be scoring guesses rather than the written\n│ plan\n│ C) Both: CLI primary, Platform secondary\n│ ✅ Keeps the CLI hello-world path as the spine while still auditing the upgrade and entitlement flow behind it\n│ ✅ Lets the persona step offer both CLI-first and evaluator personas instead of forci…\n\n❯ 1. CLI Tool (recommended)\n ✅ Matches the one concrete touchpoint the plan names: `mytool pricing` and the setup wizard�✅ Gives a measurable\n hello-world path: install, run the command, see the tier, upgrade�❌ Underweights the signup, checkout, and\n entitlement flow where the plan's goal actually lives\n 2. Platform / hosted service\n ✅ Centers the review on account creation, billing, and entitlements, where signups are won or lost�✅ Surfaces the\n missing auth check on the upgrade endpoint as a trust issue early�❌ The plan describes almost none of this\n surface, so the review would score guesses, not the plan\n 3. Both: CLI + Platform\n ✅ Keeps the CLI hello-world path as the spine while auditing the upgrade and entitlement flow behind it�✅ Lets\n the persona step offer both CLI-first and evaluator personas�❌ Roughly doubles passes and benchmarks (human: ~1\n extra hour / CC: ~5 extra min) for a thin plan\n 4. No DX surface, exit\n ✅ Honest if `mytool pricing` is a stub and the real work is a billing backend with no DX to review�✅ Routes the\n untested, unauthenticated upgrade endpoint to /plan-eng-review, which will catch it�❌ Abandons the CLI command and\n wizard, exactly the parts a DX review exists to sharpen\n 5. Type something. \n────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────\n 6. Chat about this\n\nEnter to select · ↑/↓ to navigate · Esc to cancel"
|
||
}
|