mirror of
https://github.com/garrytan/gstack.git
synced 2026-09-16 18:05:31 +02:00
* feat: add a restricted and supervised Claude Code runner Preserve configured authentication and models while enforcing tool access, strict completion JSON, bounded output and process cleanup. Cover argv, failure handling, session metadata and Windows process containment. * feat: route outside reviews by harness and migrate wrapper installs Use Claude Code from Codex and Codex from other supported hosts, with shared invocation rendering, positive gate validation and per-phase provenance. Rename /claude to /claude-code, repair managed shared and copied installations safely, and generate native Kiro skills. Add installed-workflow, failure-injection and live cross-harness regression coverage. * test: recognize CEO mode labels without terminal spacing The paid workflow rendered SCOPEEXPANSION at option 4, but its driver required a literal space. Match the leading mode title without cursor-spacing artifacts and ignore adjacent preview text. Preserve missing-target failures and downstream posture assertions. * test: isolate plan-count fixtures before starting review workflows Seed the complete test plan in a private git repository before launching Claude, so a bare slash command cannot review the live workspace while a delayed fixture message remains queued. Preserve count thresholds, parsers and budgets. Add initial-context and installed-discovery tests, and retain startup/terminal diagnostics on failed evaluations. * test: stabilize review fixtures and Claude eval startup Preserve source boundaries in workflow judge inputs, isolate CEO mode plans, and wait for interactive trust input readiness. Keep startup failure evidence and retain existing models, budgets, and assertions. Co-Authored-By: OpenAI Codex <noreply@openai.com> * test: classify collapsed review modes and isolate seeded findings Keep review questions out of the setup count when terminal cursor positioning removes spaces. State existing webhook safeguards so the five-finding control measures its seeded defects without accidental extra security and concurrency gaps. Preserve question bands and the paired control. Co-Authored-By: OpenAI Codex <noreply@openai.com> * test: isolate browser daemon state across free shards Co-Authored-By: OpenAI Codex <noreply@openai.com> * test: stabilize native review counting and interactive navigation Co-Authored-By: OpenAI Codex <noreply@openai.com> * chore: prepare v1.82.0.0 release Co-Authored-By: OpenAI Codex <noreply@openai.com> * fix: eliminate browser and process-cleanup test flakes Pin every CI surface to Bun 1.4.0 to avoid extra-stdio finalizers closing reused live sockets. Add an isolated GC/listener regression that fails on Bun 1.3.13, and prevent coordinated rollback to an affected CI runtime. Check renderer cleanup against the render's own staging directory so concurrent renders cannot invalidate the assertion. Make the no-pgrep process-tree walk tolerate disappearing /proc entries, and synchronize its test fixture through child readiness and pipe EOF instead of sleeps. Validation: 9,157 passed, 31 skipped, zero failures across 556 files with retries disabled. Build, all-host generation freshness, and skill checks passed. All three races have failing-before/passing-after regressions. * fix: count completed native review questions in evals * fix: drive review navigation from confirmed native choices * fix: require complete section-loading eval reports * test: isolate telemetry HTTP transport from local assertions * fix: keep review input on the active native question * test: let tunnel revocation daemon choose an available port * test: allocate available ports for pairing and watchdog fixtures * fix: stabilize planning eval navigation and phase reporting * test: isolate installed runtime paths in planning evals * test: stabilize review evidence and concurrent refresh fixtures * fix: resolve design findings before editing the plan * fix: honor and persist disabled outside plan reviews * fix: preserve planning decisions and terminal evidence Load installed host reviews at autoplan phase entry and wait for completed reviewers and saved artifacts. Reuse approved remedies while preserving individual finding decisions. Drive interactive evals from the current terminal viewport, bind native questions across scrolling, and require complete native report evidence. Cover captured stale menus, permission lifecycles, setup classification, and disabled-review tool availability with deterministic regressions. Advance release metadata and the upgrade migration to the unclaimed 1.83.0.0 slot. * fix: drive native review questions and preserve current plans Use the native single-choice keyboard protocol and current terminal viewport, with per-question navigation inside packets and completed-call coverage. Keep permissions, multi-select menus, and Submit controls distinct. Send Autoplan reviewers the amended implementation plan, keep its review record separate, and supply retained application contracts in the chain fixture. Clarify individual DevEx decisions and complete CEO fix options; use one active plan destination for the section-loading report. * fix: preserve complete plan-review decisions * fix: recognize native plan dialogs and reviewer controls * fix: preserve review decisions and phase completion * fix: recognize completed reviews without losing findings * fix: preserve review continuity and native eval completion * test: fix native review completion and eval retry isolation * test: handle native review menus and complete eval fixtures * test: fix native review setup, completion, and isolation failures * test: limit native skill discovery to runtime assets * fix: bind Autoplan reviews to full ordered phase inputs * test: fix planning eval routing, counting, and timeout handling * chore: advance queued release to v1.84.0.0 * fix: preserve complete review inputs and planning decisions * fix: reconcile review approvals and preserve phase obligations * fix: preserve review obligations and unblock eval permissions Carry recorded Autoplan requirements into blind phase inputs, require Eng review approvals before exit, and exercise combined asynchronous flows in CEO reviews. Correct native finding and handoff classification and unblock repeated report edits using scoped request identities. * fix: retain plan requirements and complete native review dialogs * fix: complete native review prompts and retain plan references * fix: preserve review inputs and classify native eval evidence * fix: check competing completion orders in CEO reviews * fix: recognize review decisions and require phase methodology Require the current phase methodology before Autoplan snapshots. Correct substantive decision, closed handoff, and cache-finding classification, and honor the recommended implementation approach in native review dialogs. Add captured-transcript regressions without changing review thresholds, provider models, retries, or deadlines. * test: bind native review decisions and close completed handoffs * fix: complete review dialogs and verify methodology delivery * fix: preserve review evidence and unblock native eval prompts * fix: handle native review question completions * fix: recognize native review narration and controls * fix: count native review decisions and isolate eval fixtures * test: verify seeded review coverage and current artifact permissions * test: isolate model and brain-aware skill renders * fix: repair native workflow evaluation and clarify review steps * fix: stabilize workflow eval evidence and review guidance * test: repair native workflow observation and fixture isolation * fix: recognize completed workflow evidence and owned skill reads * test: repair seeded workflow delivery and completion evidence * test: recognize current review evidence across native forms * test: handle native review variants and permission redraws * fix: honor review preferences and recognize native eval evidence * test: recognize completed review decisions and queued permissions * test: match current review contracts and partial-line edits * test: recognize completed workflow evidence and bounded human waits * fix: preserve review entry gates and native eval interactions * fix: recognize native workflow evidence and preserve review gates * test: recognize current review evidence and preconfigure workflow fixtures * test: recognize completed review findings and scoped artifact permissions * fix: stabilize native workflow review and permission evidence * fix: recognize current review evidence and scoped edit confirmations Clarify Design and engineering review entry instructions and Design scoring. Recognize required legacy coverage and public Autoplan completion recaps. Bind the pending Edit confirmation to its exact file, ordered digest, and one-request approval when a preceding command display remains visible. Keep reviews within their existing size limits and preserve scope gates when extracting workflow fixtures from either supported preamble header. Keep failure outcomes, review thresholds, provider choices, and eval budgets. * fix: recover review workflow progress and eval evidence * fix: recognize valid review evidence and scope selection * test: fix review evidence parsing and repeated artifact prompts * test: recognize valid review decisions and pending native cards * fix(plan-eng-review): keep final navigation consistent with approved tasks * test: recognize valid review evidence and bind legacy diff requests * fix: stabilize review eval evidence and harness repair guidance * docs: update project documentation for v1.85.0.0 Co-Authored-By: OpenAI Codex <noreply@openai.com> * test: fix Windows CI fixtures and credential scan Rebase captured JSON values and filesystem evidence using the appropriate path convention. Compile native fake CLIs on Windows and synchronize pipe holder readiness, with cleanup retained when assertions fail. Assemble synthetic credential fixtures at runtime so the added-line scan keeps enforcing the same gate without flagging its own rejection controls. Discover generated skills directly for the empty-find regression check, avoiding a recursive scan through saved evaluation artifacts and dependencies. * fix: preserve source renders on Windows Compare canonical generator paths using native separators so an output sidecar pointing at the source cannot overwrite its skill or metadata. Keep the regression fixture isolated from the real checkout and expose freshness diagnostics before asserting subprocess status. Detach Windows drain-test pipe holders from the fake provider's automatic child cleanup while preserving the enclosing runner job and its assertions. * fix: clarify outside review fallback and CEO decisions Render one applicable own-harness fallback path and retain native review, disabled policy, and missing-coverage semantics. Align report field names and mode labels, and make the existing per-cut scope approval explicit. Regenerate skill outputs and keep the workflow judge's model, thresholds, and retry policy unchanged. * chore: move release to free version slot (v1.86.0.0) PR #2852 now claims v1.85.0.0. Align the release metadata and rename migration so upgrades from that version still receive it. Co-Authored-By: OpenAI Codex <noreply@openai.com> * fix: include engineering review prerequisites and restore branch context * fix: recognize coverage diagrams and clarify design review instructions * fix: preserve file identities and join Windows test processes --------- Co-authored-by: OpenAI Codex <noreply@openai.com>
171 lines
31 KiB
JSON
171 lines
31 KiB
JSON
{
|
|
"provenance": {
|
|
"sourceHead": "fd7698b23d0d88dfcfe8a2b0fd4c3140cdad01fa",
|
|
"sourceSnapshotSha256": "94e1f7637004b0c60dc70e38ff5bc0ef17ae07b2240354fca389d204c05f970a",
|
|
"historicalOutcome": "plan_ready; seeded-gap assertion failed",
|
|
"paidOutcomesReclassified": false
|
|
},
|
|
"calls": [
|
|
{
|
|
"sessionId": "6473b11e-784e-4c0b-9777-c7197c1c273f",
|
|
"toolUseId": "toolu_01SCW3STwTGeD6P2FAjL9Dnj",
|
|
"questions": [
|
|
{
|
|
"header": "Quickstart",
|
|
"question": "D1 \u2014 Quickstart points at examples/first_eval.py, which is not shipped\nProject/branch/task: EvalKit SDK beta (main), DX POLISH review, journey stage Hello World.\nELI10: README.md line 11 tells the developer to run `python examples/first_eval.py`. docs/package-contents.txt confirms that file is in neither the published wheel nor the examples archive. Your ML engineer hits \"No such file or directory\" at about T+0:30, on the very first command after install, and wonders whether the install is broken. The working demo, `python -m evalkit.demo`, is further down the page.\nStakes if we pick wrong: The first thing the persona runs fails. That is the minute-two abandonment moment, and it hides the demo that actually works.\nRecommendation: A because the demo module already ships and works; pointing the quickstart at it fixes the dead command with a docs edit and gives one golden path. Principle: zero friction at T0.\nCompleteness: A=10/10, B=8/10, C=3/10\nA) Make `python -m evalkit.demo` the quickstart command and drop the first_eval.py reference (recommended) (human: ~30 min / CC: ~3 min)\n \u2705 Every command in the quickstart runs against the shipped 2.0.0b1 package as published\n \u2705 One golden path: install, demo, then live eval, no \"choose your own adventure\"\n \u274c Loses the named example file some readers expect to open and edit locally\nB) Ship examples/first_eval.py in the package and archive so the current command works (human: ~2 h / CC: ~15 min)\n \u2705 Keeps the quickstart text as written and gives an editable script to start from\n \u2705 Provides a real-usage sample beyond the fixed demo output\n \u274c Two first-run paths compete on the same page, and a second file must stay in sync with demo.py\nC) Keep the reference and add a note that the file is coming later (human: ~10 min / CC: ~1 min)\n \u2705 Smallest possible edit before the beta ships\n \u2705 Signals the intended example so early users know what to expect\n \u274c The first quickstart command still fails for every developer in the beta\nNet: repair the dead command by pointing at what ships versus shipping what the text promises.",
|
|
"options": [
|
|
{
|
|
"label": "A: Point quickstart at demo (recommended)",
|
|
"description": "README quickstart runs python -m evalkit.demo; remove first_eval.py reference. Completeness 10/10."
|
|
},
|
|
{
|
|
"label": "B: Ship first_eval.py",
|
|
"description": "Add the example to the package and archive so the existing command works. Completeness 8/10."
|
|
},
|
|
{
|
|
"label": "C: Note it as coming later",
|
|
"description": "Keep the reference with a caveat; first command still fails. Completeness 3/10."
|
|
}
|
|
],
|
|
"multiSelect": false
|
|
}
|
|
],
|
|
"answered": true,
|
|
"failed": false,
|
|
"answers": {
|
|
"D1 \u2014 Quickstart points at examples/first_eval.py, which is not shipped\nProject/branch/task: EvalKit SDK beta (main), DX POLISH review, journey stage Hello World.\nELI10: README.md line 11 tells the developer to run `python examples/first_eval.py`. docs/package-contents.txt confirms that file is in neither the published wheel nor the examples archive. Your ML engineer hits \"No such file or directory\" at about T+0:30, on the very first command after install, and wonders whether the install is broken. The working demo, `python -m evalkit.demo`, is further down the page.\nStakes if we pick wrong: The first thing the persona runs fails. That is the minute-two abandonment moment, and it hides the demo that actually works.\nRecommendation: A because the demo module already ships and works; pointing the quickstart at it fixes the dead command with a docs edit and gives one golden path. Principle: zero friction at T0.\nCompleteness: A=10/10, B=8/10, C=3/10\nA) Make `python -m evalkit.demo` the quickstart command and drop the first_eval.py reference (recommended) (human: ~30 min / CC: ~3 min)\n \u2705 Every command in the quickstart runs against the shipped 2.0.0b1 package as published\n \u2705 One golden path: install, demo, then live eval, no \"choose your own adventure\"\n \u274c Loses the named example file some readers expect to open and edit locally\nB) Ship examples/first_eval.py in the package and archive so the current command works (human: ~2 h / CC: ~15 min)\n \u2705 Keeps the quickstart text as written and gives an editable script to start from\n \u2705 Provides a real-usage sample beyond the fixed demo output\n \u274c Two first-run paths compete on the same page, and a second file must stay in sync with demo.py\nC) Keep the reference and add a note that the file is coming later (human: ~10 min / CC: ~1 min)\n \u2705 Smallest possible edit before the beta ships\n \u2705 Signals the intended example so early users know what to expect\n \u274c The first quickstart command still fails for every developer in the beta\nNet: repair the dead command by pointing at what ships versus shipping what the text promises.": "A: Point quickstart at demo (recommended)"
|
|
},
|
|
"unansweredQuestionIndices": [],
|
|
"answeredAt": "2026-09-10T09:30:01.097Z"
|
|
},
|
|
{
|
|
"sessionId": "6473b11e-784e-4c0b-9777-c7197c1c273f",
|
|
"toolUseId": "toolu_012VCMJYcAdtrgKeyZvkK7Db",
|
|
"questions": [
|
|
{
|
|
"header": "CI wait",
|
|
"question": "D2 \u2014 First local evaluation blocks 5 minutes on a mandatory remote CI check\nProject/branch/task: EvalKit SDK beta (main), DX POLISH review, journey stage Hello World.\nELI10: docs/current-contracts.md lines 3-5: the first local evaluation requires a successful remote CI check and blocks for five minutes, with no skip flag and no offline path. The keyless demo inherits the same wait (README lines 17-23). Your persona wants a local score first and CI second; the SDK forces the order the other way round. docs/benchmarks.md measured 6 minutes, of which 5 is this wait. Peers land at 2 to 4 minutes and the approved target is under 2. The progress lines and timeout message already exist and are good; they do not shorten the wait.\nStakes if we pick wrong: The approved Champion target is unreachable while this gate stands. The magical moment lands at T+5:30 instead of T+0:30, and a remote CI outage turns the local demo into EVALKIT_CI_TIMEOUT.\nRecommendation: A because the check verifies a CI binding, so it belongs at the first CI run, not the first local run. The persona's stated order is local first, CI second. Principles: zero friction at T0, decide for me but let me override.\nCompleteness: A=10/10, B=7/10, C=3/10\nA) Move the binding check out of the local path: local evals and the demo return immediately; the check runs on first use in noninteractive CI mode, and locally as a non-blocking background status line (recommended) (human: ~3 days / CC: ~1 h)\n \u2705 Demo and first local result land at about T+0:30, inside the under-2-minute target\n \u2705 Check still runs where it matters, in CI, with the existing timeout and help messaging kept intact\n \u274c Changes a documented contract in the beta and needs the CI-mode path tested for the same failure states\nB) Keep the default wait, add an escape hatch: `EVALKIT_SKIP_CI_CHECK=1` or `--skip-ci-check`, honored by the demo and local runs (human: ~1 day / CC: ~20 min)\n \u2705 Existing contract preserved for anyone who relies on the gate by default\n \u2705 Developers who read the README can get the fast path in one env var\n \u274c Default TTHW stays 6 minutes; only developers who find the flag hit the target, and the docs must now teach the flag\nC) Retain the mandatory wait as documented; rely on the existing progress and timeout messaging (human: 0 / CC: 0)\n \u2705 No behavior change to ship or test in the beta\n \u2705 The wait is already well explained with URL, progress, and recovery help\n \u274c Target is missed by 3x and the demo can fail on a remote outage before any local value is shown\nNet: shipping the approved 2-minute target versus keeping a gate that guards CI but is paid at the local terminal.",
|
|
"options": [
|
|
{
|
|
"label": "A: Check runs in CI, not locally (recommended)",
|
|
"description": "Local evals and demo return at once; binding check moves to first CI-mode use plus a non-blocking local status line. Completeness 10/10."
|
|
},
|
|
{
|
|
"label": "B: Keep wait, add skip flag",
|
|
"description": "Default still waits 5 minutes; EVALKIT_SKIP_CI_CHECK / --skip-ci-check bypass for demo and local. Completeness 7/10."
|
|
},
|
|
{
|
|
"label": "C: Keep mandatory wait",
|
|
"description": "Ship the contract unchanged; TTHW stays 6 minutes. Completeness 3/10."
|
|
}
|
|
],
|
|
"multiSelect": false
|
|
}
|
|
],
|
|
"answered": true,
|
|
"failed": false,
|
|
"answers": {
|
|
"D2 \u2014 First local evaluation blocks 5 minutes on a mandatory remote CI check\nProject/branch/task: EvalKit SDK beta (main), DX POLISH review, journey stage Hello World.\nELI10: docs/current-contracts.md lines 3-5: the first local evaluation requires a successful remote CI check and blocks for five minutes, with no skip flag and no offline path. The keyless demo inherits the same wait (README lines 17-23). Your persona wants a local score first and CI second; the SDK forces the order the other way round. docs/benchmarks.md measured 6 minutes, of which 5 is this wait. Peers land at 2 to 4 minutes and the approved target is under 2. The progress lines and timeout message already exist and are good; they do not shorten the wait.\nStakes if we pick wrong: The approved Champion target is unreachable while this gate stands. The magical moment lands at T+5:30 instead of T+0:30, and a remote CI outage turns the local demo into EVALKIT_CI_TIMEOUT.\nRecommendation: A because the check verifies a CI binding, so it belongs at the first CI run, not the first local run. The persona's stated order is local first, CI second. Principles: zero friction at T0, decide for me but let me override.\nCompleteness: A=10/10, B=7/10, C=3/10\nA) Move the binding check out of the local path: local evals and the demo return immediately; the check runs on first use in noninteractive CI mode, and locally as a non-blocking background status line (recommended) (human: ~3 days / CC: ~1 h)\n \u2705 Demo and first local result land at about T+0:30, inside the under-2-minute target\n \u2705 Check still runs where it matters, in CI, with the existing timeout and help messaging kept intact\n \u274c Changes a documented contract in the beta and needs the CI-mode path tested for the same failure states\nB) Keep the default wait, add an escape hatch: `EVALKIT_SKIP_CI_CHECK=1` or `--skip-ci-check`, honored by the demo and local runs (human: ~1 day / CC: ~20 min)\n \u2705 Existing contract preserved for anyone who relies on the gate by default\n \u2705 Developers who read the README can get the fast path in one env var\n \u274c Default TTHW stays 6 minutes; only developers who find the flag hit the target, and the docs must now teach the flag\nC) Retain the mandatory wait as documented; rely on the existing progress and timeout messaging (human: 0 / CC: 0)\n \u2705 No behavior change to ship or test in the beta\n \u2705 The wait is already well explained with URL, progress, and recovery help\n \u274c Target is missed by 3x and the demo can fail on a remote outage before any local value is shown\nNet: shipping the approved 2-minute target versus keeping a gate that guards CI but is paid at the local terminal.": "A: Check runs in CI, not locally (recommended)"
|
|
},
|
|
"unansweredQuestionIndices": [],
|
|
"answeredAt": "2026-09-10T09:30:36.706Z"
|
|
},
|
|
{
|
|
"sessionId": "6473b11e-784e-4c0b-9777-c7197c1c273f",
|
|
"toolUseId": "toolu_017uV55w4iwLYU5oqwtQAexx",
|
|
"questions": [
|
|
{
|
|
"header": "Signatures",
|
|
"question": "D3 \u2014 run_eval and run_batch take the same two arguments in reversed positional order\nProject/branch/task: EvalKit SDK beta (main), DX POLISH review, journey stage Real Usage.\nELI10: docs/api.md lines 5-9: `run_eval(dataset, evaluator)` and `run_batch(evaluator, dataset)`. Same concepts, swapped order, and the draft calls the reversal intentional. Your ML engineer learns one function from the README, then calls the other with the same order. Because both arguments are plain Python objects, the call may raise a confusing error deep in the evaluator or, worse, iterate the evaluator as if it were the dataset. The good-API test is \"use it correctly after seeing one example\"; this surface fails that test by design.\nStakes if we pick wrong: Every developer who graduates from single evaluations to batches hits a swap bug once. In a beta, that is an issue filed per team and a permanent \"gotcha\" note in every internal wiki.\nRecommendation: A because consistent order plus keyword support is the pit of success, and a one-release swap detector keeps existing callers working while they move. Principle: decide for me, let me override; progressive disclosure.\nCompleteness: A=10/10, B=6/10, C=3/10\nA) Unify to `(dataset, evaluator)` in both functions, document keyword usage `run_batch(dataset=..., evaluator=...)`, and have `run_batch` detect the legacy reversed positional order for one release: emit a DeprecationWarning naming the new order, then proceed (recommended) (human: ~1 day / CC: ~20 min)\n \u2705 One order to learn; the README example transfers directly to batch calls\n \u2705 Existing 2.0.0b1 callers keep working through the beta with a warning that shows the exact fix\n \u274c Touches a public signature during the beta and needs a type-based detector plus tests for the swapped case\nB) Keep both orders as documented, add a runtime check that raises `TypeError` naming the expected order when the two arguments are swapped (human: ~half day / CC: ~10 min)\n \u2705 No signature change; the swap fails fast with a clear message instead of a deep stack trace\n \u2705 Small, contained change that only fires on the mistake\n \u274c The inconsistency stays forever, and every new developer still trips on it once\nC) Keep the reversed order, add a callout in docs/api.md and the README (human: ~15 min / CC: ~2 min)\n \u2705 Zero code change before the beta ships\n \u2705 Documents the trap for developers who read the reference\n \u274c Developers who copy from one example to the next do not read callouts; the swap bug ships\nNet: fix the API shape once now versus documenting a permanent trap.",
|
|
"options": [
|
|
{
|
|
"label": "A: Unify order + warn on legacy (recommended)",
|
|
"description": "Both take (dataset, evaluator); keywords documented; run_batch warns on detected reversed order for one release. Completeness 10/10."
|
|
},
|
|
{
|
|
"label": "B: Keep orders, TypeError on swap",
|
|
"description": "Signatures unchanged; a runtime check names the expected order when swapped. Completeness 6/10."
|
|
},
|
|
{
|
|
"label": "C: Docs callout only",
|
|
"description": "Keep the reversed order and warn in the docs. Completeness 3/10."
|
|
}
|
|
],
|
|
"multiSelect": false
|
|
}
|
|
],
|
|
"answered": true,
|
|
"failed": false,
|
|
"answers": {
|
|
"D3 \u2014 run_eval and run_batch take the same two arguments in reversed positional order\nProject/branch/task: EvalKit SDK beta (main), DX POLISH review, journey stage Real Usage.\nELI10: docs/api.md lines 5-9: `run_eval(dataset, evaluator)` and `run_batch(evaluator, dataset)`. Same concepts, swapped order, and the draft calls the reversal intentional. Your ML engineer learns one function from the README, then calls the other with the same order. Because both arguments are plain Python objects, the call may raise a confusing error deep in the evaluator or, worse, iterate the evaluator as if it were the dataset. The good-API test is \"use it correctly after seeing one example\"; this surface fails that test by design.\nStakes if we pick wrong: Every developer who graduates from single evaluations to batches hits a swap bug once. In a beta, that is an issue filed per team and a permanent \"gotcha\" note in every internal wiki.\nRecommendation: A because consistent order plus keyword support is the pit of success, and a one-release swap detector keeps existing callers working while they move. Principle: decide for me, let me override; progressive disclosure.\nCompleteness: A=10/10, B=6/10, C=3/10\nA) Unify to `(dataset, evaluator)` in both functions, document keyword usage `run_batch(dataset=..., evaluator=...)`, and have `run_batch` detect the legacy reversed positional order for one release: emit a DeprecationWarning naming the new order, then proceed (recommended) (human: ~1 day / CC: ~20 min)\n \u2705 One order to learn; the README example transfers directly to batch calls\n \u2705 Existing 2.0.0b1 callers keep working through the beta with a warning that shows the exact fix\n \u274c Touches a public signature during the beta and needs a type-based detector plus tests for the swapped case\nB) Keep both orders as documented, add a runtime check that raises `TypeError` naming the expected order when the two arguments are swapped (human: ~half day / CC: ~10 min)\n \u2705 No signature change; the swap fails fast with a clear message instead of a deep stack trace\n \u2705 Small, contained change that only fires on the mistake\n \u274c The inconsistency stays forever, and every new developer still trips on it once\nC) Keep the reversed order, add a callout in docs/api.md and the README (human: ~15 min / CC: ~2 min)\n \u2705 Zero code change before the beta ships\n \u2705 Documents the trap for developers who read the reference\n \u274c Developers who copy from one example to the next do not read callouts; the swap bug ships\nNet: fix the API shape once now versus documenting a permanent trap.": "A: Unify order + warn on legacy (recommended)"
|
|
},
|
|
"unansweredQuestionIndices": [],
|
|
"answeredAt": "2026-09-10T09:31:16.344Z"
|
|
},
|
|
{
|
|
"sessionId": "6473b11e-784e-4c0b-9777-c7197c1c273f",
|
|
"toolUseId": "toolu_01R9YdSFG3sFWNsyJgHzgDfW",
|
|
"questions": [
|
|
{
|
|
"header": "Auth error",
|
|
"question": "D4 \u2014 Invalid API key raises AuthError(\"request failed\") with no cause or fix\nProject/branch/task: EvalKit SDK beta (main), DX POLISH review, journey stage Debug.\nELI10: docs/api.md lines 11-13: an invalid key raises `AuthError(\"request failed\")` with no error code, no explanation, and no instruction for replacing the key. Every other EvalKit error already names the cause, the argument or file, and an actionable fix (current-contracts.md lines 21-24), so this is the one outlier. Your ML engineer pastes the key wrong at about T+10:00 after the demo, reads \"request failed\", and starts debugging the network instead of the key. The CI timeout message (code, URL, retry instruction, help link) is the house style to copy.\nStakes if we pick wrong: The first live-evaluation failure most developers hit is a bad key, and the SDK points them nowhere. That is ten to twenty minutes lost per developer and a support ticket per team.\nRecommendation: A because the SDK already has the pattern; one error class should not be the exception to it. Principle: fight uncertainty, every error is problem plus cause plus fix.\nCompleteness: A=10/10, B=6/10, C=2/10\nA) Structured auth errors matching the CI-timeout pattern: code `EVALKIT_AUTH_INVALID_KEY` (plus `EVALKIT_AUTH_MISSING_KEY` and `EVALKIT_AUTH_KEY_REVOKED`), message naming the redacted key suffix and project, the fix (create or rotate at the console key page, export EVALKIT_API_KEY), and a help link (recommended) (human: ~1 day / CC: ~15 min)\n \u2705 Developer sees the problem, cause, and exact fix in the traceback, no docs detour\n \u2705 Distinguishes missing, invalid, and revoked keys so rotation cases self-diagnose\n \u274c Needs the server to return a distinguishable auth reason, or a client-side fallback when it does not\nB) Single improved message: `AuthError(\"Invalid API key. Set EVALKIT_API_KEY from https://console.evalkit.example/settings/api-keys\")`, no code, no case split (human: ~1 h / CC: ~5 min)\n \u2705 Names the key as the cause and points at the console page\n \u2705 Pure message change, no server contract needed\n \u274c Missing versus revoked versus wrong-project all read the same, and no code to grep or match in CI logs\nC) Keep `AuthError(\"request failed\")` as documented (human: 0 / CC: 0)\n \u2705 No change to ship before the beta\n \u2705 Consistent with the current published behavior\n \u274c The most common live-eval failure stays undiagnosable from the error text\nNet: bring the one outlier error up to the SDK's own standard versus shipping a known dead end.",
|
|
"options": [
|
|
{
|
|
"label": "A: Structured auth errors (recommended)",
|
|
"description": "Codes for invalid, missing, revoked; redacted key suffix, console fix, help link. Matches EVALKIT_CI_TIMEOUT style. Completeness 10/10."
|
|
},
|
|
{
|
|
"label": "B: One better message",
|
|
"description": "Name the key and the console URL in a single message, no code or case split. Completeness 6/10."
|
|
},
|
|
{
|
|
"label": "C: Keep as documented",
|
|
"description": "Ship AuthError(\"request failed\") unchanged. Completeness 2/10."
|
|
}
|
|
],
|
|
"multiSelect": false
|
|
}
|
|
],
|
|
"answered": true,
|
|
"failed": false,
|
|
"answers": {
|
|
"D4 \u2014 Invalid API key raises AuthError(\"request failed\") with no cause or fix\nProject/branch/task: EvalKit SDK beta (main), DX POLISH review, journey stage Debug.\nELI10: docs/api.md lines 11-13: an invalid key raises `AuthError(\"request failed\")` with no error code, no explanation, and no instruction for replacing the key. Every other EvalKit error already names the cause, the argument or file, and an actionable fix (current-contracts.md lines 21-24), so this is the one outlier. Your ML engineer pastes the key wrong at about T+10:00 after the demo, reads \"request failed\", and starts debugging the network instead of the key. The CI timeout message (code, URL, retry instruction, help link) is the house style to copy.\nStakes if we pick wrong: The first live-evaluation failure most developers hit is a bad key, and the SDK points them nowhere. That is ten to twenty minutes lost per developer and a support ticket per team.\nRecommendation: A because the SDK already has the pattern; one error class should not be the exception to it. Principle: fight uncertainty, every error is problem plus cause plus fix.\nCompleteness: A=10/10, B=6/10, C=2/10\nA) Structured auth errors matching the CI-timeout pattern: code `EVALKIT_AUTH_INVALID_KEY` (plus `EVALKIT_AUTH_MISSING_KEY` and `EVALKIT_AUTH_KEY_REVOKED`), message naming the redacted key suffix and project, the fix (create or rotate at the console key page, export EVALKIT_API_KEY), and a help link (recommended) (human: ~1 day / CC: ~15 min)\n \u2705 Developer sees the problem, cause, and exact fix in the traceback, no docs detour\n \u2705 Distinguishes missing, invalid, and revoked keys so rotation cases self-diagnose\n \u274c Needs the server to return a distinguishable auth reason, or a client-side fallback when it does not\nB) Single improved message: `AuthError(\"Invalid API key. Set EVALKIT_API_KEY from https://console.evalkit.example/settings/api-keys\")`, no code, no case split (human: ~1 h / CC: ~5 min)\n \u2705 Names the key as the cause and points at the console page\n \u2705 Pure message change, no server contract needed\n \u274c Missing versus revoked versus wrong-project all read the same, and no code to grep or match in CI logs\nC) Keep `AuthError(\"request failed\")` as documented (human: 0 / CC: 0)\n \u2705 No change to ship before the beta\n \u2705 Consistent with the current published behavior\n \u274c The most common live-eval failure stays undiagnosable from the error text\nNet: bring the one outlier error up to the SDK's own standard versus shipping a known dead end.": "A: Structured auth errors (recommended)"
|
|
},
|
|
"unansweredQuestionIndices": [],
|
|
"answeredAt": "2026-09-10T09:31:53.981Z"
|
|
},
|
|
{
|
|
"sessionId": "6473b11e-784e-4c0b-9777-c7197c1c273f",
|
|
"toolUseId": "toolu_01AGy5jB21K3mXYjUKwrEFZW",
|
|
"questions": [
|
|
{
|
|
"header": "v1 to v2",
|
|
"question": "D5 \u2014 v2 removes Client.evaluate() immediately with no alias, warning, guide, or codemod\nProject/branch/task: EvalKit SDK beta (main), DX POLISH review, journey stage Upgrade.\nELI10: docs/api.md lines 15-18: v1 exposes `Client.evaluate()`, v2 replaces it with `Client.run()` and drops the old name at once. No compatibility alias, no DeprecationWarning, no migration guide, no codemod. Your ML engineer upgrades an existing v1 project to try the beta, every `client.evaluate()` call raises `AttributeError`, and nothing in the traceback says the method was renamed. The rest of the changelog is complete, so this one rename is the only upgrade hazard.\nStakes if we pick wrong: Upgrade fear. Teams with v1 in production will not trial the beta if the first import breaks, and the beta feedback survey only hears from greenfield users.\nRecommendation: A because a deprecated alias costs a few lines and turns a hard break into a warning that names the fix, and the guide plus one-liner make the migration boring. Principle: credible, upgrades should be boring.\nCompleteness: A=10/10, B=7/10, C=5/10\nA) Keep `Client.evaluate()` as a deprecated alias for `Client.run()` through 2.x with a DeprecationWarning that names the replacement and removal version (3.0); add a \"Upgrading from 1.x\" section to the changelog and docs; ship a documented one-line rename (`python -m evalkit.migrate` or an equivalent sed/ruff command) (recommended) (human: ~1 day / CC: ~20 min)\n \u2705 v1 projects run on 2.0.0b1 unchanged, so production teams can trial the beta safely\n \u2705 The warning tells the developer exactly what to rename and by when; the one-liner does it for them\n \u274c Carries a shim through the 2.x line and one more thing to remove in 3.0\nB) Deprecated alias plus warning only, no migration guide or codemod (human: ~2 h / CC: ~10 min)\n \u2705 Upgrade no longer breaks, and the warning names Client.run()\n \u2705 Minimal surface, easy to delete later\n \u274c Nothing in the docs explains the v1 to v2 change for someone reading before they upgrade\nC) Remove immediately as documented, but add an \"Upgrading from 1.x\" changelog section with the rename (human: ~30 min / CC: ~3 min)\n \u2705 Cleanest v2 API with no shim to maintain\n \u2705 The rename is at least documented for developers who read the changelog first\n \u274c Every v1 caller still breaks at import time in the beta, and the AttributeError does not point at the changelog\nNet: a boring upgrade with a shim to retire later versus a clean API that breaks every v1 user on day one.",
|
|
"options": [
|
|
{
|
|
"label": "A: Alias + warning + guide + codemod (recommended)",
|
|
"description": "Client.evaluate() stays as a deprecated alias through 2.x; migration guide and one-line rename shipped. Completeness 10/10."
|
|
},
|
|
{
|
|
"label": "B: Alias + warning only",
|
|
"description": "Keep the old name with a DeprecationWarning; no guide or codemod. Completeness 7/10."
|
|
},
|
|
{
|
|
"label": "C: Remove now, document in changelog",
|
|
"description": "Hard removal as drafted plus an Upgrading section. Completeness 5/10."
|
|
}
|
|
],
|
|
"multiSelect": false
|
|
}
|
|
],
|
|
"answered": true,
|
|
"failed": false,
|
|
"answers": {
|
|
"D5 \u2014 v2 removes Client.evaluate() immediately with no alias, warning, guide, or codemod\nProject/branch/task: EvalKit SDK beta (main), DX POLISH review, journey stage Upgrade.\nELI10: docs/api.md lines 15-18: v1 exposes `Client.evaluate()`, v2 replaces it with `Client.run()` and drops the old name at once. No compatibility alias, no DeprecationWarning, no migration guide, no codemod. Your ML engineer upgrades an existing v1 project to try the beta, every `client.evaluate()` call raises `AttributeError`, and nothing in the traceback says the method was renamed. The rest of the changelog is complete, so this one rename is the only upgrade hazard.\nStakes if we pick wrong: Upgrade fear. Teams with v1 in production will not trial the beta if the first import breaks, and the beta feedback survey only hears from greenfield users.\nRecommendation: A because a deprecated alias costs a few lines and turns a hard break into a warning that names the fix, and the guide plus one-liner make the migration boring. Principle: credible, upgrades should be boring.\nCompleteness: A=10/10, B=7/10, C=5/10\nA) Keep `Client.evaluate()` as a deprecated alias for `Client.run()` through 2.x with a DeprecationWarning that names the replacement and removal version (3.0); add a \"Upgrading from 1.x\" section to the changelog and docs; ship a documented one-line rename (`python -m evalkit.migrate` or an equivalent sed/ruff command) (recommended) (human: ~1 day / CC: ~20 min)\n \u2705 v1 projects run on 2.0.0b1 unchanged, so production teams can trial the beta safely\n \u2705 The warning tells the developer exactly what to rename and by when; the one-liner does it for them\n \u274c Carries a shim through the 2.x line and one more thing to remove in 3.0\nB) Deprecated alias plus warning only, no migration guide or codemod (human: ~2 h / CC: ~10 min)\n \u2705 Upgrade no longer breaks, and the warning names Client.run()\n \u2705 Minimal surface, easy to delete later\n \u274c Nothing in the docs explains the v1 to v2 change for someone reading before they upgrade\nC) Remove immediately as documented, but add an \"Upgrading from 1.x\" changelog section with the rename (human: ~30 min / CC: ~3 min)\n \u2705 Cleanest v2 API with no shim to maintain\n \u2705 The rename is at least documented for developers who read the changelog first\n \u274c Every v1 caller still breaks at import time in the beta, and the AttributeError does not point at the changelog\nNet: a boring upgrade with a shim to retire later versus a clean API that breaks every v1 user on day one.": "A: Alias + warning + guide + codemod (recommended)"
|
|
},
|
|
"unansweredQuestionIndices": [],
|
|
"answeredAt": "2026-09-10T09:32:29.623Z"
|
|
}
|
|
]
|
|
}
|