mirror of
https://github.com/garrytan/gstack.git
synced 2026-09-16 18:05:31 +02:00
* feat: add a restricted and supervised Claude Code runner Preserve configured authentication and models while enforcing tool access, strict completion JSON, bounded output and process cleanup. Cover argv, failure handling, session metadata and Windows process containment. * feat: route outside reviews by harness and migrate wrapper installs Use Claude Code from Codex and Codex from other supported hosts, with shared invocation rendering, positive gate validation and per-phase provenance. Rename /claude to /claude-code, repair managed shared and copied installations safely, and generate native Kiro skills. Add installed-workflow, failure-injection and live cross-harness regression coverage. * test: recognize CEO mode labels without terminal spacing The paid workflow rendered SCOPEEXPANSION at option 4, but its driver required a literal space. Match the leading mode title without cursor-spacing artifacts and ignore adjacent preview text. Preserve missing-target failures and downstream posture assertions. * test: isolate plan-count fixtures before starting review workflows Seed the complete test plan in a private git repository before launching Claude, so a bare slash command cannot review the live workspace while a delayed fixture message remains queued. Preserve count thresholds, parsers and budgets. Add initial-context and installed-discovery tests, and retain startup/terminal diagnostics on failed evaluations. * test: stabilize review fixtures and Claude eval startup Preserve source boundaries in workflow judge inputs, isolate CEO mode plans, and wait for interactive trust input readiness. Keep startup failure evidence and retain existing models, budgets, and assertions. Co-Authored-By: OpenAI Codex <noreply@openai.com> * test: classify collapsed review modes and isolate seeded findings Keep review questions out of the setup count when terminal cursor positioning removes spaces. State existing webhook safeguards so the five-finding control measures its seeded defects without accidental extra security and concurrency gaps. Preserve question bands and the paired control. Co-Authored-By: OpenAI Codex <noreply@openai.com> * test: isolate browser daemon state across free shards Co-Authored-By: OpenAI Codex <noreply@openai.com> * test: stabilize native review counting and interactive navigation Co-Authored-By: OpenAI Codex <noreply@openai.com> * chore: prepare v1.82.0.0 release Co-Authored-By: OpenAI Codex <noreply@openai.com> * fix: eliminate browser and process-cleanup test flakes Pin every CI surface to Bun 1.4.0 to avoid extra-stdio finalizers closing reused live sockets. Add an isolated GC/listener regression that fails on Bun 1.3.13, and prevent coordinated rollback to an affected CI runtime. Check renderer cleanup against the render's own staging directory so concurrent renders cannot invalidate the assertion. Make the no-pgrep process-tree walk tolerate disappearing /proc entries, and synchronize its test fixture through child readiness and pipe EOF instead of sleeps. Validation: 9,157 passed, 31 skipped, zero failures across 556 files with retries disabled. Build, all-host generation freshness, and skill checks passed. All three races have failing-before/passing-after regressions. * fix: count completed native review questions in evals * fix: drive review navigation from confirmed native choices * fix: require complete section-loading eval reports * test: isolate telemetry HTTP transport from local assertions * fix: keep review input on the active native question * test: let tunnel revocation daemon choose an available port * test: allocate available ports for pairing and watchdog fixtures * fix: stabilize planning eval navigation and phase reporting * test: isolate installed runtime paths in planning evals * test: stabilize review evidence and concurrent refresh fixtures * fix: resolve design findings before editing the plan * fix: honor and persist disabled outside plan reviews * fix: preserve planning decisions and terminal evidence Load installed host reviews at autoplan phase entry and wait for completed reviewers and saved artifacts. Reuse approved remedies while preserving individual finding decisions. Drive interactive evals from the current terminal viewport, bind native questions across scrolling, and require complete native report evidence. Cover captured stale menus, permission lifecycles, setup classification, and disabled-review tool availability with deterministic regressions. Advance release metadata and the upgrade migration to the unclaimed 1.83.0.0 slot. * fix: drive native review questions and preserve current plans Use the native single-choice keyboard protocol and current terminal viewport, with per-question navigation inside packets and completed-call coverage. Keep permissions, multi-select menus, and Submit controls distinct. Send Autoplan reviewers the amended implementation plan, keep its review record separate, and supply retained application contracts in the chain fixture. Clarify individual DevEx decisions and complete CEO fix options; use one active plan destination for the section-loading report. * fix: preserve complete plan-review decisions * fix: recognize native plan dialogs and reviewer controls * fix: preserve review decisions and phase completion * fix: recognize completed reviews without losing findings * fix: preserve review continuity and native eval completion * test: fix native review completion and eval retry isolation * test: handle native review menus and complete eval fixtures * test: fix native review setup, completion, and isolation failures * test: limit native skill discovery to runtime assets * fix: bind Autoplan reviews to full ordered phase inputs * test: fix planning eval routing, counting, and timeout handling * chore: advance queued release to v1.84.0.0 * fix: preserve complete review inputs and planning decisions * fix: reconcile review approvals and preserve phase obligations * fix: preserve review obligations and unblock eval permissions Carry recorded Autoplan requirements into blind phase inputs, require Eng review approvals before exit, and exercise combined asynchronous flows in CEO reviews. Correct native finding and handoff classification and unblock repeated report edits using scoped request identities. * fix: retain plan requirements and complete native review dialogs * fix: complete native review prompts and retain plan references * fix: preserve review inputs and classify native eval evidence * fix: check competing completion orders in CEO reviews * fix: recognize review decisions and require phase methodology Require the current phase methodology before Autoplan snapshots. Correct substantive decision, closed handoff, and cache-finding classification, and honor the recommended implementation approach in native review dialogs. Add captured-transcript regressions without changing review thresholds, provider models, retries, or deadlines. * test: bind native review decisions and close completed handoffs * fix: complete review dialogs and verify methodology delivery * fix: preserve review evidence and unblock native eval prompts * fix: handle native review question completions * fix: recognize native review narration and controls * fix: count native review decisions and isolate eval fixtures * test: verify seeded review coverage and current artifact permissions * test: isolate model and brain-aware skill renders * fix: repair native workflow evaluation and clarify review steps * fix: stabilize workflow eval evidence and review guidance * test: repair native workflow observation and fixture isolation * fix: recognize completed workflow evidence and owned skill reads * test: repair seeded workflow delivery and completion evidence * test: recognize current review evidence across native forms * test: handle native review variants and permission redraws * fix: honor review preferences and recognize native eval evidence * test: recognize completed review decisions and queued permissions * test: match current review contracts and partial-line edits * test: recognize completed workflow evidence and bounded human waits * fix: preserve review entry gates and native eval interactions * fix: recognize native workflow evidence and preserve review gates * test: recognize current review evidence and preconfigure workflow fixtures * test: recognize completed review findings and scoped artifact permissions * fix: stabilize native workflow review and permission evidence * fix: recognize current review evidence and scoped edit confirmations Clarify Design and engineering review entry instructions and Design scoring. Recognize required legacy coverage and public Autoplan completion recaps. Bind the pending Edit confirmation to its exact file, ordered digest, and one-request approval when a preceding command display remains visible. Keep reviews within their existing size limits and preserve scope gates when extracting workflow fixtures from either supported preamble header. Keep failure outcomes, review thresholds, provider choices, and eval budgets. * fix: recover review workflow progress and eval evidence * fix: recognize valid review evidence and scope selection * test: fix review evidence parsing and repeated artifact prompts * test: recognize valid review decisions and pending native cards * fix(plan-eng-review): keep final navigation consistent with approved tasks * test: recognize valid review evidence and bind legacy diff requests * fix: stabilize review eval evidence and harness repair guidance * docs: update project documentation for v1.85.0.0 Co-Authored-By: OpenAI Codex <noreply@openai.com> * test: fix Windows CI fixtures and credential scan Rebase captured JSON values and filesystem evidence using the appropriate path convention. Compile native fake CLIs on Windows and synchronize pipe holder readiness, with cleanup retained when assertions fail. Assemble synthetic credential fixtures at runtime so the added-line scan keeps enforcing the same gate without flagging its own rejection controls. Discover generated skills directly for the empty-find regression check, avoiding a recursive scan through saved evaluation artifacts and dependencies. * fix: preserve source renders on Windows Compare canonical generator paths using native separators so an output sidecar pointing at the source cannot overwrite its skill or metadata. Keep the regression fixture isolated from the real checkout and expose freshness diagnostics before asserting subprocess status. Detach Windows drain-test pipe holders from the fake provider's automatic child cleanup while preserving the enclosing runner job and its assertions. * fix: clarify outside review fallback and CEO decisions Render one applicable own-harness fallback path and retain native review, disabled policy, and missing-coverage semantics. Align report field names and mode labels, and make the existing per-cut scope approval explicit. Regenerate skill outputs and keep the workflow judge's model, thresholds, and retry policy unchanged. * chore: move release to free version slot (v1.86.0.0) PR #2852 now claims v1.85.0.0. Align the release metadata and rename migration so upgrades from that version still receive it. Co-Authored-By: OpenAI Codex <noreply@openai.com> * fix: include engineering review prerequisites and restore branch context * fix: recognize coverage diagrams and clarify design review instructions * fix: preserve file identities and join Windows test processes --------- Co-authored-by: OpenAI Codex <noreply@openai.com>
88 lines
50 KiB
JSON
88 lines
50 KiB
JSON
{
|
||
"provenance": {
|
||
"run": "ship-source-ap-delta-paid-20260910-v1",
|
||
"captureSha256": "43995bb37c2e6fb85648a7f10422fe6249e70da02fc0be92ed1deac559474833",
|
||
"observationSha256": "ade89c2de539726204e5c827ae749d98393dc6344bd64310faf3e63ae97404b2",
|
||
"reportSha256": "be24c43a7a37aa54c056f90b58bce504e9935514bbf241f6f02362c68980a20f",
|
||
"reportMtimeMs": 1789046277623.377,
|
||
"startedAt": 1789045509540,
|
||
"projection": "Exact last two public completed AUQ calls and current pending ExitPlanMode requests; preceding calls remain in retained observation. No private reasoning."
|
||
},
|
||
"calls": [
|
||
{
|
||
"sessionId": "3e666a22-9524-42eb-9236-6488b8c8c60f",
|
||
"toolUseId": "toolu_01SH1DH99dTU1W4X5kpA9FjF",
|
||
"questions": [
|
||
{
|
||
"question": "D14 — TODO candidate 2 of 2: add an automated rewrite (one-line command or codemod) for the 1.x to 2.0 migration?\nProject/branch/task: gstack-plan-count on main, /plan-devex-review of the EvalKit beta plan, TODOS.md step.\nELI10: D8 keeps Client.evaluate() as a deprecated alias and adds a Migrating from 1.x section. D6 changes run_eval/run_batch to one positional argument plus keyword-only evaluator. Both are mechanical edits. The hall-of-fame bar (Next.js, AG Grid) is a codemod per breaking release. For two renames a full codemod is heavy, but a documented one-liner (a sed or ruff/libcst snippet in the Migrating section) gets most of the value for a fraction of the cost.\nWhat: add a tested rewrite snippet to the Migrating from 1.x section covering evaluate() -> run() and positional -> keyword evaluator; optionally grow it into python -m evalkit.migrate before 3.0 removes the alias.\nWhy: teams with many v1 scripts otherwise hand-edit each one; the deprecation warning tells them what, not how fast.\nPros: upgrades become one command; sets the precedent before 3.0, when the alias is removed and the codemod becomes necessary.\nCons: a regex rewrite can miss dynamic calls; a libcst codemod is a new dev dependency and test surface.\nContext: docs/api.md lines 15-18 describe the rename; D6 and D8 in this plan define the final shapes.\nDepends on: D6 and D8 landing; the 3.0 removal date.\nStakes if we pick wrong: low for the beta; higher at 3.0 when the alias disappears.\nRecommendation: A because the alias makes it non-urgent now, but 3.0 needs it, and recording it with the trigger avoids a scramble later.\nNote: options differ in kind, not coverage — no completeness score.\nPros / cons:\nA) Add to TODOS.md (recommended) (human: ~1 day / CC: ~30 min when built)\n ✅ Schedules the codemod against the concrete trigger: alias removal in 3.0\n ✅ Keeps the beta focused; the snippet can be added to docs any time before then\n ❌ v1 teams upgrading to 2.0 hand-edit their scripts for now, guided only by the warning\nB) Skip\n ✅ Two renames may never justify a codemod; the alias covers 2.x entirely\n ✅ No new dependency or test surface\n ❌ 3.0 arrives with no migration tooling and the removal is felt as a hard break\nC) Build it now: put the tested one-line sed/ruff snippet into the Migrating section in this release\n ✅ Cheapest possible form lands with the beta; developers upgrade in one command\n ✅ No dependency; a snippet in docs plus a test that runs it against a fixture\n ❌ Adds a docs-and-test item to a release already carrying five contract repairs\nNet: track the migration tooling for 3.0, drop it, or ship the one-liner now.",
|
||
"header": "TODO 2",
|
||
"multiSelect": false,
|
||
"options": [
|
||
{
|
||
"label": "Add to TODOS.md (recommended)",
|
||
"description": "Record as P3 with the 3.0 alias-removal trigger."
|
||
},
|
||
{
|
||
"label": "Skip",
|
||
"description": "Do not track."
|
||
},
|
||
{
|
||
"label": "Build the one-liner now",
|
||
"description": "Tested sed/ruff snippet in the Migrating section for 2.0.0b1."
|
||
}
|
||
]
|
||
}
|
||
],
|
||
"answered": true,
|
||
"failed": false,
|
||
"answers": {
|
||
"D14 — TODO candidate 2 of 2: add an automated rewrite (one-line command or codemod) for the 1.x to 2.0 migration?\nProject/branch/task: gstack-plan-count on main, /plan-devex-review of the EvalKit beta plan, TODOS.md step.\nELI10: D8 keeps Client.evaluate() as a deprecated alias and adds a Migrating from 1.x section. D6 changes run_eval/run_batch to one positional argument plus keyword-only evaluator. Both are mechanical edits. The hall-of-fame bar (Next.js, AG Grid) is a codemod per breaking release. For two renames a full codemod is heavy, but a documented one-liner (a sed or ruff/libcst snippet in the Migrating section) gets most of the value for a fraction of the cost.\nWhat: add a tested rewrite snippet to the Migrating from 1.x section covering evaluate() -> run() and positional -> keyword evaluator; optionally grow it into python -m evalkit.migrate before 3.0 removes the alias.\nWhy: teams with many v1 scripts otherwise hand-edit each one; the deprecation warning tells them what, not how fast.\nPros: upgrades become one command; sets the precedent before 3.0, when the alias is removed and the codemod becomes necessary.\nCons: a regex rewrite can miss dynamic calls; a libcst codemod is a new dev dependency and test surface.\nContext: docs/api.md lines 15-18 describe the rename; D6 and D8 in this plan define the final shapes.\nDepends on: D6 and D8 landing; the 3.0 removal date.\nStakes if we pick wrong: low for the beta; higher at 3.0 when the alias disappears.\nRecommendation: A because the alias makes it non-urgent now, but 3.0 needs it, and recording it with the trigger avoids a scramble later.\nNote: options differ in kind, not coverage — no completeness score.\nPros / cons:\nA) Add to TODOS.md (recommended) (human: ~1 day / CC: ~30 min when built)\n ✅ Schedules the codemod against the concrete trigger: alias removal in 3.0\n ✅ Keeps the beta focused; the snippet can be added to docs any time before then\n ❌ v1 teams upgrading to 2.0 hand-edit their scripts for now, guided only by the warning\nB) Skip\n ✅ Two renames may never justify a codemod; the alias covers 2.x entirely\n ✅ No new dependency or test surface\n ❌ 3.0 arrives with no migration tooling and the removal is felt as a hard break\nC) Build it now: put the tested one-line sed/ruff snippet into the Migrating section in this release\n ✅ Cheapest possible form lands with the beta; developers upgrade in one command\n ✅ No dependency; a snippet in docs plus a test that runs it against a fixture\n ❌ Adds a docs-and-test item to a release already carrying five contract repairs\nNet: track the migration tooling for 3.0, drop it, or ship the one-liner now.": "Add to TODOS.md (recommended)"
|
||
},
|
||
"unansweredQuestionIndices": [],
|
||
"answeredAt": "2026-09-10T13:16:01.210Z"
|
||
},
|
||
{
|
||
"sessionId": "3e666a22-9524-42eb-9236-6488b8c8c60f",
|
||
"toolUseId": "toolu_01TJkPfW2oXEZAtqeXgLAD7v",
|
||
"questions": [
|
||
{
|
||
"question": "D15 — DX review complete. What next?\nProject/branch/task: gstack-plan-count on main, /plan-devex-review of the EvalKit beta plan is finished; plan written to gstack-test-plan-devex.md.\nELI10: The DX review is done: overall DX 5/10 to 8/10, TTHW from 6 minutes to an estimated under 1 minute once the CI gate leaves the demo path, fourteen decisions recorded, twelve implementation tasks. The review readiness dashboard shows the DX review clean, the outside voice disabled by config, and no engineering review yet. Engineering review is the one gate that normally blocks shipping, and this plan changes runtime behavior (CI gate, signatures, error classes, alias), so it is the natural next check. After implementation, /devex-review on the live package is the boomerang that measures whether the under-2-minute target was actually hit.\nStakes if we pick wrong: low; this only routes what happens after this session. You said you will handle subsequent reviews manually.\nRecommendation: D because you stated in PLAN.md that you will handle subsequent reviews manually; the eng-review recommendation stands and is recorded in the report's verdict.\nNote: options differ in kind, not coverage — no completeness score.\nPros / cons:\nA) Run /plan-eng-review next (required gate)\n ✅ Validates the runtime changes (T3 CI gate move, T5 signatures, T6 error classes, T7 alias) architecturally before build\n ✅ Clears the only review that gates shipping under current config\n ❌ Another interactive session now, which you said you would run yourself\nB) Ready to implement; run /devex-review after shipping\n ✅ Moves straight to the twelve tasks with a concrete boomerang measurement planned\n ✅ The under-2-minute target gets verified against the real package\n ❌ Skips the eng gate for now; the dashboard stays NOT CLEARED until it runs\nC) Skip, I'll handle next steps manually (recommended)\n ✅ Matches your stated intent to run later reviews yourself\n ✅ Nothing else is launched from this session; the plan and report are complete\n ❌ Eng review remains outstanding until you start it\nNet: chain into eng review now, go build with a boomerang check, or stop here as you asked.",
|
||
"header": "Next steps",
|
||
"multiSelect": false,
|
||
"options": [
|
||
{
|
||
"label": "Skip, handle manually (recommended)",
|
||
"description": "End the DX review here; you run subsequent reviews yourself."
|
||
},
|
||
{
|
||
"label": "Run /plan-eng-review next",
|
||
"description": "Chain into the required engineering gate now."
|
||
},
|
||
{
|
||
"label": "Ready to implement",
|
||
"description": "Proceed to the tasks; run /devex-review after shipping."
|
||
}
|
||
]
|
||
}
|
||
],
|
||
"answered": true,
|
||
"failed": false,
|
||
"answers": {
|
||
"D15 — DX review complete. What next?\nProject/branch/task: gstack-plan-count on main, /plan-devex-review of the EvalKit beta plan is finished; plan written to gstack-test-plan-devex.md.\nELI10: The DX review is done: overall DX 5/10 to 8/10, TTHW from 6 minutes to an estimated under 1 minute once the CI gate leaves the demo path, fourteen decisions recorded, twelve implementation tasks. The review readiness dashboard shows the DX review clean, the outside voice disabled by config, and no engineering review yet. Engineering review is the one gate that normally blocks shipping, and this plan changes runtime behavior (CI gate, signatures, error classes, alias), so it is the natural next check. After implementation, /devex-review on the live package is the boomerang that measures whether the under-2-minute target was actually hit.\nStakes if we pick wrong: low; this only routes what happens after this session. You said you will handle subsequent reviews manually.\nRecommendation: D because you stated in PLAN.md that you will handle subsequent reviews manually; the eng-review recommendation stands and is recorded in the report's verdict.\nNote: options differ in kind, not coverage — no completeness score.\nPros / cons:\nA) Run /plan-eng-review next (required gate)\n ✅ Validates the runtime changes (T3 CI gate move, T5 signatures, T6 error classes, T7 alias) architecturally before build\n ✅ Clears the only review that gates shipping under current config\n ❌ Another interactive session now, which you said you would run yourself\nB) Ready to implement; run /devex-review after shipping\n ✅ Moves straight to the twelve tasks with a concrete boomerang measurement planned\n ✅ The under-2-minute target gets verified against the real package\n ❌ Skips the eng gate for now; the dashboard stays NOT CLEARED until it runs\nC) Skip, I'll handle next steps manually (recommended)\n ✅ Matches your stated intent to run later reviews yourself\n ✅ Nothing else is launched from this session; the plan and report are complete\n ❌ Eng review remains outstanding until you start it\nNet: chain into eng review now, go build with a boomerang check, or stop here as you asked.": "Skip, handle manually (recommended)"
|
||
},
|
||
"unansweredQuestionIndices": [],
|
||
"answeredAt": "2026-09-10T13:18:13.135Z"
|
||
}
|
||
],
|
||
"planReadyRequests": [
|
||
{
|
||
"sessionId": "3e666a22-9524-42eb-9236-6488b8c8c60f",
|
||
"toolUseId": "toolu_01DvnxQbmeG6cpv6xrqESL7U",
|
||
"timestamp": "2026-09-10T13:18:58.245Z",
|
||
"failed": false
|
||
}
|
||
],
|
||
"screen": "\n ┌──────────┬───────────────────────────┬─────────────────┬──────┬──────────┬─────────────────────────────────────┐\n │ Review │ Trigger │ Why │ Runs │ Status │ Findings │\n ├──────────┼───────────────────────────┼─────────────────┼──────┼──────────┼─────────────────────────────────────┤\n │ CEO │ /plan-ceo-review │ Scope & │ 0 │ — │ — │\n │ Review │ │ strategy │ │ │ │\n ├──────────┼───────────────────────────┼─────────────────┼──────┼──────────┼─────────────────────────────────────┤\n │ │ codex via │ │ │ │ │\n │ Outside │ /plan-devex-review (host: │ Independent 2nd │ 1 │ disabled │ codex_reviews=disabled; no native │\n │ Review │ claude, phase: │ opinion │ │ │ fallback run │\n │ │ plan-review) │ │ │ │ │\n ├──────────┼───────────────────────────┼─────────────────┼──────┼──────────┼─────────────────────────────────────┤\n │ Eng │ │ Architecture & │ │ │ │\n │ Review │ /plan-eng-review │ tests │ 0 │ — │ — │\n │ │ │ (required) │ │ │ │\n ├──────────┼───────────────────────────┼─────────────────┼──────┼──────────┼─────────────────────────────────────┤\n │ Design │ /plan-design-review │ UI/UX gaps │ 0 │ — │ — │\n │ Review │ │ │ │ │ │\n ├──────────┼───────────────────────────┼─────────────────┼──────┼──────────┼─────────────────────────────────────┤\n │ │ │ │ │ │ score: 5/10 → 8/10, TTHW: 6 min → │\n │ DX │ /plan-devex-review │ Developer │ 1 │ clean │ <2 min target (est <1 min); 5 │\n │ Review │ │ experience gaps │ │ │ contract repairs, 3 polish items, 2 │\n │ │ │ │ │ │ TODOs │\n └──────────┴───────────────────────────┴─────────────────┴──────┴──────────┴─────────────────────────────────────┘\n\n - OUTSIDE COVERAGE: provider codex, phase plan-review, outside_status disabled (user opt-out). No native fallback\n dispatched.\n - VERDICT: DX CLEARED — eng review required.\n\n NO UNRESOLVED DECISIONS\n╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌\n\n────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────\n Claude has written up a plan and is ready to execute. Would you like to proceed?\n\n ❯ 1. Yes, and use auto mode\n 2. Yes, manually approve edits\n 3. Tell Claude what to change\n shift+tab to approve with this feedback\n",
|
||
"reportContent": "# Plan: EvalKit SDK beta release polish (DX POLISH review)\n\n## Context\n\nEvalKit is a Python SDK for ML engineers evaluating LLM responses. The beta release\n(`evalkit==2.0.0b1`) is scoped as DX POLISH: no new hosted services, language\nbindings, or community programs. The persona, the under-2-minute onboarding target,\nand the terminal demo (`python -m evalkit.demo`) as the first-success vehicle are\nalready settled in README.md and docs/benchmarks.md.\n\nThe current draft ships four developer-facing contracts unchanged\n(docs/current-contracts.md, docs/api.md, docs/package-contents.txt). This review\ntraces each against the persona's actual journey and records which repairs belong\nin the release.\n\nReview conducted by `/plan-devex-review`, mode DX POLISH, session\n3158935-1789045529-82df556d. Base branch: `main` (no remote; git-native fallback).\n\n## Settled inputs (not re-litigated)\n\n- **Product type:** Library/SDK (Python). Source: README.md line 3.\n- **Persona:** ML engineer, writes Python daily, terminal user, wants a local result\n before wiring the SDK into production CI. Source: README.md lines 3-5.\n- **Benchmark:** internal study, docs/benchmarks.md. Peers A/B/C at 2/4/3 minutes.\n EvalKit at 6 minutes including the 5-minute CI wait. Target: under 2 minutes.\n- **Magical moment vehicle:** copy-paste terminal demo printing real scores. No\n playground, no new UI. Source: README.md lines 14-17.\n- **Mode:** DX POLISH. Source: PLAN.md line 8.\n\n## Housekeeping decisions\n\n- **D1 routing rules:** user accepted adding gstack skill routing rules to CLAUDE.md.\n Deferred: plan mode blocks the edit and commit. Do after plan mode exits:\n append the routing section to CLAUDE.md, then\n `git add CLAUDE.md && git commit -m \"chore: add gstack skill routing rules to CLAUDE.md\"`.\n- **D2 design doc:** user skipped /office-hours; standard review.\n\n## Developer Perspective (confirmed by user, D3)\n\n```\nTARGET DEVELOPER PERSONA\n========================\nWho: ML engineer evaluating LLM responses; writes Python daily; terminal user\nContext: wants one local evaluation result before wiring EvalKit into production CI\nTolerance: peers finish in 2-4 min (docs/benchmarks.md); expects a first result in\n under 5 minutes and no auth step before a demo\nExpects: pip install, one command, a printed score, then an API key step\n```\n\nI open the README. First heading: Getting started. It says pip install\nevalkit==2.0.0b1, then run python examples/first_eval.py. I install. I look for\nexamples/first_eval.py. It does not exist: not in site-packages, not in the examples\narchive (docs/package-contents.txt). Thirty seconds in and the quickstart is already\nlying to me. I read on and see a second first command, python -m evalkit.demo. Two\nfirst commands; I pick the one that exists. It prints \"Verifying the sample-project\nbinding with EvalKit CI\", a check URL, and \"normally completes within 300s\". Then\n\"Waiting for CI check: 30s elapsed of 300s\", every 30 seconds. I did not ask for CI.\nI have no API key yet. There is no skip flag and no offline path\n(docs/current-contracts.md). I stare at a countdown for five minutes to see three\nlines of mock scores. Peers gave me a result in 2 minutes. Afterward I export\nEVALKIT_API_KEY, paste it wrong once, and get AuthError(\"request failed\"): no code,\nno cause, no hint to rotate the key. I open the two functions and find\nrun_eval(dataset, evaluator) but run_batch(evaluator, dataset), reversed. My v1\nscript calls Client.evaluate(); v2 deleted it with no alias, warning, or migration\nguide (docs/api.md). I am now 12 minutes in and trust is gone.\n\n## Competitive DX benchmark (from docs/benchmarks.md; web research skipped, Aside not installed and the internal study is approved)\n\n```\nCOMPETITIVE DX BENCHMARK\n=========================\nTool | TTHW | Notable DX choice | Source\nPeer SDK A | 2 min | first result with no remote gate | docs/benchmarks.md\nPeer SDK B | 4 min | (not recorded) | docs/benchmarks.md\nPeer SDK C | 3 min | (not recorded) | docs/benchmarks.md\nStripe (ref) | ~30 s | test keys pre-filled in docs, 7 lines | dx-hall-of-fame.md\nEVALKIT | 6 min | keyless demo, but 5-min mandatory CI wait | README + current-contracts\n```\n\nTarget tier: Champion (< 2 min), already agreed. Current: Needs Work (5-10 min).\nThe 5-minute CI wait alone exceeds the whole target; the target is unreachable\nwhile docs/current-contracts.md lines 3-5 stand.\n\n## Magical moment (settled)\n\nVehicle: `python -m evalkit.demo` printing real per-example and overall scores from\nbundled sample data over the mock transport, no API key. The moment is \"I saw a real\nscore from my terminal\". Today the moment arrives at T+5:30 instead of T+0:30\nbecause of the CI gate.\n\n## Journey trace and friction decisions (Step 0F)\n\n### D4 INSTALL / QUICKSTART: README first command targets an unshipped file — DECIDED: single demo command\n\nEvidence: README.md line 11 runs `python examples/first_eval.py`; docs/package-contents.txt\nlines 8-9 confirm the file is absent from the wheel and the examples archive. README.md\nline 15 names a second first command, `python -m evalkit.demo`, which ships and works.\n\nPlan change:\n- README \"Getting started\" becomes exactly two steps: `python -m pip install evalkit==2.0.0b1`\n then `python -m evalkit.demo`. Remove every reference to `examples/first_eval.py`.\n- Add a release check that every command in README resolves against\n docs/package-contents.txt (a script or CI step that greps README for `python ` invocations\n and verifies the module/file exists in the built wheel). Prevents this class of drift.\n- Follow-up (not release-gating): if a copyable starter script is wanted later, ship it\n inside the package as `evalkit/examples/first_eval.py` so it is importable and testable.\n\n### D5 HELLO WORLD: mandatory 5-minute CI check before the first local result — DECIDED: remove the gate from local and keyless runs\n\nEvidence: docs/current-contracts.md lines 3-5 (blocking check, no skip, no offline path);\nREADME.md lines 17-23 (keyless demo still waits); docs/benchmarks.md lines 4-8 (EvalKit\n6 min vs peers 2-4 min; target under 2 min). The wait alone is 2.5x the target budget.\n\nPlan change:\n- Local and keyless evaluations (mock transport, bundled sample data, `python -m evalkit.demo`)\n never contact EvalKit CI. They return results immediately.\n- The sample-project binding check moves to the first **live** evaluation (one made with\n `EVALKIT_API_KEY` set and a real transport). It runs non-blocking: the evaluation\n proceeds, and the SDK emits the existing \"Verifying the sample-project binding with EvalKit\n CI; inspect <url>; normally completes within 300s\" line once, then the existing 30-second\n progress lines to stderr only while the process is still alive and the check is pending.\n- Keep the existing EVALKIT_CI_TIMEOUT message, check URL, help link, and recovery\n instruction unchanged; they now describe a warning, not a blocked result.\n- Keep the existing noninteractive CI mode behavior for pipelines.\n- Owner action before merge: confirm with the CI check's owners that nothing depends on the\n check completing before a *local* result. Record the answer in the release notes.\n- Update docs/current-contracts.md lines 3-19 and README.md lines 17-23 to describe the\n new behavior. Delete the sentence \"The keyless demo still waits for that CI check and has\n no skip or offline bypass.\"\n- Expected TTHW after change: install (~30 s) + demo (~5 s) = under 1 minute. Champion tier.\n- Verify: re-run the docs/benchmarks.md study on the same machine; the timing\n instrumentation already exists and is unchanged.\n\n### D6 REAL USAGE: run_eval / run_batch reversed positional order — DECIDED: align to (dataset, evaluator), keyword-only guard\n\nEvidence: docs/api.md lines 5-9. Same two concepts, opposite order, positional in practice.\n\nPlan change:\n- Both functions become `run_eval(dataset, /, *, evaluator)` and\n `run_batch(dataset, /, *, evaluator)`: dataset is the single positional argument,\n evaluator is keyword-only. A second positional argument raises\n `TypeError: run_batch() takes 1 positional argument. Call run_batch(dataset, evaluator=evaluator).`\n- Docstrings and the API reference show the identical call shape for both.\n- Add tests: correct call, swapped positional call raises the TypeError above, keyword call works.\n- Changelog entry under 2.0.0b1 breaking changes, alongside the Client.run rename (D8).\n- Update docs/api.md lines 3-9; delete \"The reversed positional order is intentional\".\n\n### D7 DEBUG: AuthError(\"request failed\") — DECIDED: structured auth error matching the SDK's existing pattern\n\nEvidence: docs/api.md lines 11-13. docs/current-contracts.md lines 21-24 show every other\nerror already carries cause, input, and fix; the CI timeout already has a code and help link.\n\nPlan change:\n- Three distinct auth outcomes, each with a code:\n - `EVALKIT_AUTH_MISSING_KEY`: no key found. Message: \"No API key. Set EVALKIT_API_KEY or\n pass api_key=. Create one at https://console.evalkit.example/settings/api-keys.\"\n - `EVALKIT_AUTH_INVALID_KEY`: server rejected the key. Message: \"API key ending in ...XXXX\n (from EVALKIT_API_KEY) was rejected by project <project>. Rotate or create a key at\n https://console.evalkit.example/settings/api-keys. Help: <help link>.\"\n - Network / 5xx failures stay on the existing transport error path and are never\n reported as AuthError.\n- Key hint shows only the last 4 characters; full key never appears in messages or logs\n (matches the existing \"errors redact secrets\" contract).\n- Key source (env var vs. constructor argument) is named so trailing-newline and\n wrong-shell mistakes are diagnosable.\n- Dependency: confirm the API returns distinguishable responses for missing vs. invalid\n credentials (401 vs. 403 or an error body code). If it does not, ship\n `EVALKIT_AUTH_INVALID_KEY` for all 401/403 with the same fix text and note the gap.\n- Tests: each code path, redaction, and that a 503 does not raise AuthError.\n- Update docs/api.md lines 11-13 and add the codes to the API reference error table.\n\n### D8 UPGRADE: Client.evaluate() removed with no alias or guide — DECIDED: deprecated alias through 2.x plus Migrating section\n\nEvidence: docs/api.md lines 15-18. Changelog otherwise complete.\n\nPlan change:\n- `Client.evaluate(*args, **kwargs)` remains for all 2.x releases. It calls `Client.run`\n and emits `DeprecationWarning: Client.evaluate() is deprecated since 2.0 and will be\n removed in 3.0. Use Client.run() (same arguments).` Warning fires once per process\n (`warnings.warn(..., stacklevel=2)`).\n- Removal target stated in the changelog: 3.0.\n- Add a \"Migrating from 1.x\" section to CHANGELOG and README covering both 2.0 breaks:\n `Client.evaluate` -> `Client.run`, and the `run_eval` / `run_batch` signature change (D6),\n each with a before/after snippet.\n- Docs and examples show only `Client.run()`; the alias appears only in the migration section.\n- Tests: alias delegates, warning text, warning category, fires once.\n- Update docs/api.md lines 15-18.\n\n### Journey map after decisions\n\n```\nSTAGE | DEVELOPER DOES | FRICTION POINTS | STATUS\n----------------|----------------------------------------|----------------------------------------|--------\n1. Discover | reads README \"Getting started\" | none beyond quickstart mismatch (D4) | ok\n2. Install | pip install evalkit==2.0.0b1 | first command targets unshipped file | fixed (D4)\n3. Hello World | python -m evalkit.demo | 5-min mandatory CI wait before result | fixed (D5)\n4. Real Usage | export key; run_eval / run_batch | reversed positional args | fixed (D6)\n5. Debug | bad key -> AuthError | \"request failed\", no code/cause/fix | fixed (D7)\n6. Upgrade | v1 script on 2.0 | evaluate() gone, no alias/warning/guide| fixed (D8)\n```\n\n## First-Time Developer Confusion Report (Step 0G)\n\n```\nFIRST-TIME DEVELOPER REPORT\n============================\nPersona: ML engineer, Python daily, terminal, wants a local result before CI\nAttempting: EvalKit 2.0.0b1 getting started, as documented today\n\nCONFUSION LOG:\nT+0:00 Open README. \"Getting started\": pip install evalkit==2.0.0b1. Runs clean.\nT+0:30 #1 Run python examples/first_eval.py. FileNotFoundError. The file is not in\n the wheel or the archive (package-contents.txt). First instruction is dead.\nT+0:45 #2 Re-read README; find a second first command, python -m evalkit.demo.\n Two golden paths; unclear which is canonical.\nT+1:00 #3 Demo prints \"Verifying the sample-project binding with EvalKit CI ...\n normally completes within 300s\". I have no key and asked for nothing remote.\n Why does a mock-transport demo need CI?\nT+1:30 \"Waiting for CI check: 30s elapsed of 300s\". Legible, but I am waiting.\nT+3:00 Still waiting. Peer SDK A gave me a score a minute ago (benchmarks.md).\nT+6:00 Three lines print: example 1: score=0.80 / example 2: score=1.00 /\n overall: score=0.90.\n #4 No line says what was scored, which evaluator ran, or what to do next.\n The README's key step is the next thing, but the terminal does not say so.\nT+7:00 Create a key in the console, export EVALKIT_API_KEY, run a live eval.\n #5 Pasted with a trailing newline. AuthError(\"request failed\"). Is it my key,\n my network, or the service? No code, no hint.\nT+9:00 #6 Call run_batch(dataset, evaluator) the way run_eval works. Wrong order,\n no immediate error, confusing downstream failure.\nT+11:00 #7 Port a v1 script: Client.evaluate() -> AttributeError with no mention of\n run(). Grep the changelog; the rename is not in it.\nT+12:00 Final state: succeeded, but trust spent. Would file two issues and pin 1.x.\n```\n\nAddressed by decisions so far: #1 and #2 (D4), #3 (D5), #5 (D7), #6 (D6), #7 (D8).\nOpen: #4 (demo epilogue: what was scored and what to do next).\n\n### D9 Confusion point #4: demo ends in silence — DECIDED: header + epilogue\n\nPlan change (evalkit/demo.py, print statements only):\n- Before scores: `Evaluating 2 bundled sample responses with the default evaluator (mock transport, no API key).`\n- Scores unchanged: `example 1: score=0.80`, `example 2: score=1.00`, `overall: score=0.90`.\n- After scores: `Scores are 0-1 similarity to the reference answers.` and\n `Next: export EVALKIT_API_KEY and evaluate your own dataset. See README \"Getting started\".`\n- Update the expected output block in README.md lines 31-36 to match exactly.\n- Test: snapshot the demo's stdout.\n\n## DX review passes (scores before the review -> after decisions)\n\n### Pass 1: Getting Started — 3/10 -> 9/10\nBefore: first command targets an unshipped file (D4); keyless demo blocks 5 minutes on\nremote CI (D5); TTHW 6 min against a 2-min target (docs/benchmarks.md); demo ends without\na next step (D9). Free tier and no-credit-card already fine: demo needs no key.\nAfter: `pip install evalkit==2.0.0b1` then `python -m evalkit.demo`, scores in seconds,\nepilogue names the next step. Two commands, under 1 minute. Champion tier.\nIdeal sequence and time budget:\n1. `python -m pip install evalkit==2.0.0b1` (~30 s)\n2. `python -m evalkit.demo` (~5 s) -> header, 3 score lines, epilogue\n3. export EVALKIT_API_KEY, run own dataset (~2 min, outside TTHW)\nRemaining gap to 10: the key-creation flow is 4 console sub-steps (select project, Create\nkey, copy once, export). Acceptable for a live-evaluation step; not release scope.\n\n### Pass 2: API/CLI/SDK Design — 4/10 -> 8/10\nBefore: run_eval / run_batch reversed positional order (D6); Client.evaluate removed\nwithout alias (D8). Retries, rate limits, idempotent evaluation IDs, mock transport\nalready documented (docs/current-contracts.md lines 23-24).\nAfter: one call shape for both functions, swapped calls fail loudly, rename aliased.\nRemaining gap to 10: keyword-only evaluator is a one-time edit for any beta caller who\npassed it positionally; accepted as part of the 2.0 break.\n\n### Pass 3: Error Messages & Debugging — 5/10 -> 9/10\nThree error paths traced:\n1. Invalid key. Today: `AuthError(\"request failed\")`. After D7: code, key source, redacted\n hint, console URL, help link (Rust tier).\n2. CI check timeout. Today already Rust tier: EVALKIT_CI_TIMEOUT, check URL, recovery\n steps, help link. Unchanged; after D5 it is a warning on the live path.\n3. Swapped arguments. Today: no immediate error. After D6: TypeError naming the correct call.\nSecrets already redacted. Remaining gap to 10: no verbose/debug flag is documented; not\nraised as a finding because docs/current-contracts.md does not describe one either way.\n\n### Pass 4: Documentation & Learning — 4/10 -> 8/10 (pending D11)\nBefore: quickstart command dead (D4); no migration section for the 2.0 rename (D8);\nexpected output block will drift once the demo epilogue lands (D9 updates it).\nAfter: README quickstart is copy-paste complete and matches the wheel; Migrating from 1.x\nsection exists; expected output matches the demo.\nOpen finding (D11): README.md lines 5-6 and 38-40 contain release-planning prose\n(\"The agreed review posture is DX POLISH\", \"its runtime is maintained separately from this\nrelease-planning repo\") aimed at reviewers, not developers.\n\n### Pass 5: Upgrade & Migration Path — 2/10 -> 8/10\nBefore: silent removal of Client.evaluate, no changelog entry (D8).\nAfter: deprecated alias through 2.x, warning names the replacement, removal scheduled for\n3.0, Migrating section covers both breaks. Semantic versioning respected (2.0 = breaking).\nRemaining gap to 10: no codemod. A one-line sed/ruff rewrite in the Migrating section\nwould cover the rename; candidate TODO.\n\n### Pass 6: Developer Environment & Tooling — 8/10 -> 8/10\nPython 3.10+, macOS/Linux/Windows, no Docker, type annotations, mock transport, offline\nsample data, noninteractive CI mode all exist (docs/current-contracts.md lines 26-28).\nD5 keeps noninteractive CI mode working: the binding check becomes a non-blocking warning\nthere too. No issues found.\n\n### Pass 7: Community & Ecosystem — 7/10 -> 7/10\nSupport contact, changelog, contributor guide exist. No community program, new language\nbinding, or hosted service in this release by explicit scope. Opt-in telemetry. No issues\nfound within DX POLISH scope.\n\n### Pass 8: DX Measurement & Feedback Loops — 7/10 -> (pending D12)\nTiming instrumentation and post-beta survey exist and continue (docs/benchmarks.md\nlines 9-10). Open finding (D12): the instrumentation measures one interval, install to\nfirst real result. After D5 there are two milestones worth separating: keyless demo\nresult (the under-2-minute target) and first live result (key + optional CI warning).\n\n### D11 Pass 4: README planning prose — DECIDED: move to PLAN.md\n- Move README.md lines 5-6 (\"The agreed review posture is DX POLISH ...\") and lines 38-40\n (\"These documents describe the existing SDK's behavior; its runtime is maintained\n separately ...\") into PLAN.md under a \"Review posture\" note.\n- README keeps: one-line product description, Getting started (two commands), key setup,\n expected output, links to docs/api.md and docs/benchmarks.md, Migrating from 1.x (D8).\n- Pass 4 after: 9/10.\n\n### D12 Pass 8: split timing instrumentation — DECIDED: two events + CI outcome\n- Existing opt-in instrumentation emits `demo_first_result` (install start -> demo scores\n printed) and `live_first_result` (key export -> first live result), the latter carrying\n `ci_check_outcome` in {passed, timed_out, not_run}.\n- Re-run the docs/benchmarks.md study on the same machine after D4/D5/D9 land; record both\n numbers. Target: demo_first_result under 2 minutes.\n- Post-beta survey unchanged.\n- Pass 8 after: 9/10.\n\n## Outside voice\n\nCodex review skipped (codex_reviews disabled). Re-enable: `gstack-config set codex_reviews disabled` -> `enabled`.\nNo Claude-subagent fallback is run for a disabled opt-out. Recorded `outside_status: disabled`.\n\n## Competitive DX benchmark, post-review\n\n```\nTool | TTHW | Notable DX choice\nPeer SDK A | 2 min | first result with no remote gate\nPeer SDK B | 4 min |\nPeer SDK C | 3 min |\nEVALKIT now | 6 min | 5-min CI gate before first local result\nEVALKIT plan | < 1 min est | keyless demo, no remote call, epilogue names next step\n```\nConfirm with the re-run study (D12). Tier: Champion (< 2 min) if the estimate holds.\n\n## Magical Moment Specification\n\nVehicle: `python -m evalkit.demo`. Requirements after this review:\n- No network call on the demo path (D5). Result in seconds.\n- Header line, three score lines, two-line epilogue (D9). README expected output matches.\n- README quickstart is exactly install + demo (D4).\n- Instrumented as `demo_first_result` (D12).\n\n## NOT in scope (considered, deferred)\n\n- Hosted playground or browser sandbox: excluded by the plan; terminal demo is the vehicle.\n- New language bindings, community program, hosted service: excluded by PLAN.md scope.\n- Verbose/debug flag audit: not described in the contracts either way; revisit in /devex-review.\n- Shortening the console key-creation flow (4 sub-steps): outside the SDK's control.\n- Full libcst codemod: tracked as TODO (D14), trigger is the 3.0 alias removal.\n- Starter script in the package: tracked as TODO (D13).\n\n## What already exists (reuse, do not rebuild)\n\n- CI check messaging: pre-countdown explanation, check URL, 30-second progress lines,\n EVALKIT_CI_TIMEOUT with help link (docs/current-contracts.md lines 7-19). D5 reuses all\n of it on the live path.\n- Error pattern for every non-auth error: cause, relevant input, actionable fix, redaction\n (docs/current-contracts.md lines 21-24). D7 copies this pattern for auth.\n- Mock transport, bundled sample data, demo module (docs/package-contents.txt).\n- Type annotations, noninteractive CI mode, cross-platform support, changelog,\n contributor guide, support contact, opt-in telemetry, timing instrumentation, survey.\n- Console key page with create/revoke/rotate (README.md lines 25-29); D7's fix URL points here.\n\n## TODOS.md updates (accepted; append to TODOS.md after plan mode exits)\n\n- **Starter script in package** (D13, P3). What: `evalkit/examples/first_eval.py` shipped in\n the wheel, tested, referenced from README \"Next steps\". Why: persona needs an editable file\n after the demo. Depends on: D6. Trigger: first post-beta release.\n- **Migration rewrite tooling** (D14, P3). What: tested sed/ruff snippet in \"Migrating from\n 1.x\" for `evaluate()` -> `run()` and positional -> keyword evaluator; grow into\n `python -m evalkit.migrate` before 3.0. Why: alias removal in 3.0 is otherwise a hard break.\n Depends on: D6, D8.\n\n## DX Scorecard\n\n```\n+====================================================================+\n| DX PLAN REVIEW — SCORECARD |\n+====================================================================+\n| Dimension | Before | After | Prior | Trend |\n|----------------------|--------|--------|--------|--------|\n| Getting Started | 3/10 | 9/10 | — | new |\n| API/CLI/SDK | 4/10 | 8/10 | — | new |\n| Error Messages | 5/10 | 9/10 | — | new |\n| Documentation | 4/10 | 9/10 | — | new |\n| Upgrade Path | 2/10 | 8/10 | — | new |\n| Dev Environment | 8/10 | 8/10 | — | new |\n| Community | 7/10 | 7/10 | — | new |\n| DX Measurement | 7/10 | 9/10 | — | new |\n+--------------------------------------------------------------------+\n| TTHW | 6 min | <1 min (est; verify via D12) |\n| Competitive Rank | Needs Work -> Champion (pending re-run) |\n| Magical Moment | designed via terminal demo (python -m evalkit.demo) |\n| Product Type | Library/SDK (Python) |\n| Mode | POLISH |\n| Overall DX | 5/10 | 8/10 | — | new |\n+====================================================================+\n| DX PRINCIPLE COVERAGE |\n| Zero Friction | covered (D4, D5) |\n| Learn by Doing | covered (demo; starter script deferred D13) |\n| Fight Uncertainty | covered (D7, D9, existing CI messaging) |\n| Opinionated + Escape Hatches | covered (D6 keyword-only; D8 alias) |\n| Code in Context | gap: no live-path example beyond the demo (D13)|\n| Magical Moments | covered (D5 + D9) |\n+====================================================================+\n```\n\nNo dimension below 6 after decisions. TTHW under 10 minutes before and after; not blocking.\n\n## DX Implementation Checklist\n\n```\nDX IMPLEMENTATION CHECKLIST\n============================\n[x] Time to hello world < 2 min (D5; verify D12)\n[x] Installation is one command (pip install evalkit==2.0.0b1)\n[x] First run produces meaningful output (D9 header + scores + epilogue)\n[x] Magical moment delivered via terminal demo (D4, D5)\n[x] Every error message has problem + cause + fix + docs link (D7 closes the auth gap)\n[x] API naming is guessable without docs (D6 one call shape; run() canonical)\n[x] Every parameter has a sensible default (dataset positional, evaluator keyword)\n[x] Docs have copy-paste examples that actually work (D4 command check)\n[ ] Examples show real use cases, not just hello world (deferred: D13)\n[x] Upgrade path documented with migration guide (D8)\n[~] Breaking changes have deprecation warnings + codemods (warning yes; codemod deferred D14)\n[x] Type annotations included (existing)\n[x] Works in CI/CD without special configuration (existing noninteractive mode; D5 keeps it)\n[x] Free tier, no credit card (keyless demo)\n[x] Changelog exists and is maintained (existing; D6/D8 entries added)\n[ ] Search works in documentation (not assessed; docs site out of scope)\n[x] Support channel exists (support contact exists)\n```\n\n## Implementation Tasks\nSynthesized from this review's findings. Each task derives from a specific\nfinding above. Run with Claude Code or Codex; checkbox as you ship.\n\n- [ ] **T1 (P1, human: ~1h / CC: ~5min)** — README — Make `python -m evalkit.demo` the only quickstart command; remove `examples/first_eval.py`\n - Surfaced by: D4 — README.md line 11 vs docs/package-contents.txt lines 8-9\n - Files: README.md\n - Verify: every command in README resolves against the built wheel\n- [ ] **T2 (P1, human: ~2h / CC: ~10min)** — release CI — Add a check that README commands resolve against the package inventory\n - Surfaced by: D4 — the quickstart drifted from the wheel once already\n - Files: release CI config, docs/package-contents.txt\n - Verify: check fails when a README command references a missing module or file\n- [ ] **T3 (P1, human: ~3d / CC: ~1h)** — SDK runtime — Remove the CI gate from local/keyless evaluations; run the binding check non-blocking on the first live evaluation with the existing messaging\n - Surfaced by: D5 — docs/current-contracts.md lines 3-5; README.md lines 17-23; docs/benchmarks.md\n - Files: evalkit/client.py, evalkit/demo.py, docs/current-contracts.md, README.md\n - Verify: demo prints scores with network disabled; live path logs the check and never blocks the result; noninteractive CI mode unaffected\n- [ ] **T4 (P1, human: ~1h / CC: n/a)** — release — Confirm with the CI check owners that nothing depends on the check preceding a local result; record in release notes\n - Surfaced by: D5 — accepted risk\n - Files: release notes / CHANGELOG\n - Verify: written confirmation linked from the changelog\n- [ ] **T5 (P1, human: ~1d / CC: ~30min)** — public API — Align `run_eval` and `run_batch` to `(dataset, /, *, evaluator)` with a TypeError naming the fix on swap\n - Surfaced by: D6 — docs/api.md lines 5-9\n - Files: evalkit/__init__.py, evalkit/client.py, docs/api.md, tests\n - Verify: tests for correct call, swapped call error text, keyword call\n- [ ] **T6 (P1, human: ~half day / CC: ~20min)** — errors — Structured auth errors: EVALKIT_AUTH_MISSING_KEY, EVALKIT_AUTH_INVALID_KEY with key source, redacted hint, console URL, help link\n - Surfaced by: D7 — docs/api.md lines 11-13\n - Files: evalkit/client.py, docs/api.md, tests\n - Verify: tests per code path, redaction, 503 is not AuthError; confirm server distinguishes missing vs invalid\n- [ ] **T7 (P1, human: ~half day / CC: ~20min)** — public API — Keep `Client.evaluate()` as a deprecated alias through 2.x; add \"Migrating from 1.x\" to CHANGELOG and README\n - Surfaced by: D8 — docs/api.md lines 15-18\n - Files: evalkit/client.py, CHANGELOG, README.md, docs/api.md\n - Verify: alias delegates, DeprecationWarning text names run() and 3.0, fires once\n- [ ] **T8 (P2, human: ~1h / CC: ~5min)** — demo — Header line before scores, two-line epilogue after; update README expected output\n - Surfaced by: D9 — confusion point #4\n - Files: evalkit/demo.py, README.md\n - Verify: stdout snapshot test matches README block\n- [ ] **T9 (P2, human: ~15min / CC: ~2min)** — docs — Move review-posture and repo-maintenance prose from README.md to PLAN.md\n - Surfaced by: D11 — README.md lines 5-6, 38-40\n - Files: README.md, PLAN.md\n - Verify: README contains only developer-facing sections\n- [ ] **T10 (P2, human: ~half day / CC: ~20min)** — instrumentation — Emit `demo_first_result` and `live_first_result` (+ `ci_check_outcome`); re-run the onboarding study\n - Surfaced by: D12 — docs/benchmarks.md lines 5-10\n - Files: instrumentation module, docs/benchmarks.md\n - Verify: both events appear in a local run; study records demo_first_result < 2 min\n- [ ] **T11 (P3, human: ~half day / CC: ~15min)** — package — Ship `evalkit/examples/first_eval.py` (TODO, D13)\n- [ ] **T12 (P3, human: ~1d / CC: ~30min)** — migration — Tested rewrite snippet / `evalkit.migrate` before 3.0 (TODO, D14)\n\n_No new tasks from Pass 6 (Dev Environment) or Pass 7 (Community)._\n\nHousekeeping (not DX): D1 append gstack routing rules to CLAUDE.md and commit, after plan mode exits.\n\n### Unresolved Decisions\nNone. All fourteen questions (D1-D14) were answered.\n\n## Verification (end to end, after implementation)\n\n1. Fresh virtualenv, network disabled: `pip install evalkit==2.0.0b1` from a local wheel,\n then `python -m evalkit.demo`. Expect header, three score lines, epilogue, under 10 s,\n no CI URL printed.\n2. README command check (T2) passes against the built wheel.\n3. With `EVALKIT_API_KEY` set, run a live `run_eval(dataset, evaluator=...)`. Expect the\n binding-check line once, result returned without waiting, progress lines only while the\n check is pending.\n4. Set a bad key: expect `EVALKIT_AUTH_INVALID_KEY` with key source, `...XXXX` hint, console URL.\n Unset the key: expect `EVALKIT_AUTH_MISSING_KEY`.\n5. `run_batch(evaluator, dataset)` positionally: expect the TypeError naming the correct call.\n6. v1 script calling `Client.evaluate()`: runs, emits one DeprecationWarning naming `run()` and 3.0.\n7. Re-run the docs/benchmarks.md study; `demo_first_result` under 2 minutes.\n\n## GSTACK REVIEW REPORT\n\n| Review | Trigger | Why | Runs | Status | Findings |\n|--------|---------|-----|------|--------|----------|\n| CEO Review | `/plan-ceo-review` | Scope & strategy | 0 | — | — |\n| Outside Review | codex via `/plan-devex-review` (host: claude, phase: plan-review) | Independent 2nd opinion | 1 | disabled | codex_reviews=disabled; no native fallback run |\n| Eng Review | `/plan-eng-review` | Architecture & tests (required) | 0 | — | — |\n| Design Review | `/plan-design-review` | UI/UX gaps | 0 | — | — |\n| DX Review | `/plan-devex-review` | Developer experience gaps | 1 | clean | score: 5/10 → 8/10, TTHW: 6 min → <2 min target (est <1 min); 5 contract repairs (D4-D8), 3 polish items (D9, D11, D12), 2 TODOs (D13, D14) |\n\n- **OUTSIDE COVERAGE:** provider codex, phase plan-review, outside_status disabled (user opt-out via codex_reviews). No outside findings; no native fallback dispatched for a disabled opt-out.\n- **VERDICT:** DX CLEARED — eng review required (no `/plan-eng-review` or `/review` entry within 7 days; skip_eng_review=false).\n\nNO UNRESOLVED DECISIONS\n\n\n"
|
||
}
|