mirror of
https://github.com/garrytan/gstack.git
synced 2026-09-16 09:55:29 +02:00
* feat: add a restricted and supervised Claude Code runner Preserve configured authentication and models while enforcing tool access, strict completion JSON, bounded output and process cleanup. Cover argv, failure handling, session metadata and Windows process containment. * feat: route outside reviews by harness and migrate wrapper installs Use Claude Code from Codex and Codex from other supported hosts, with shared invocation rendering, positive gate validation and per-phase provenance. Rename /claude to /claude-code, repair managed shared and copied installations safely, and generate native Kiro skills. Add installed-workflow, failure-injection and live cross-harness regression coverage. * test: recognize CEO mode labels without terminal spacing The paid workflow rendered SCOPEEXPANSION at option 4, but its driver required a literal space. Match the leading mode title without cursor-spacing artifacts and ignore adjacent preview text. Preserve missing-target failures and downstream posture assertions. * test: isolate plan-count fixtures before starting review workflows Seed the complete test plan in a private git repository before launching Claude, so a bare slash command cannot review the live workspace while a delayed fixture message remains queued. Preserve count thresholds, parsers and budgets. Add initial-context and installed-discovery tests, and retain startup/terminal diagnostics on failed evaluations. * test: stabilize review fixtures and Claude eval startup Preserve source boundaries in workflow judge inputs, isolate CEO mode plans, and wait for interactive trust input readiness. Keep startup failure evidence and retain existing models, budgets, and assertions. Co-Authored-By: OpenAI Codex <noreply@openai.com> * test: classify collapsed review modes and isolate seeded findings Keep review questions out of the setup count when terminal cursor positioning removes spaces. State existing webhook safeguards so the five-finding control measures its seeded defects without accidental extra security and concurrency gaps. Preserve question bands and the paired control. Co-Authored-By: OpenAI Codex <noreply@openai.com> * test: isolate browser daemon state across free shards Co-Authored-By: OpenAI Codex <noreply@openai.com> * test: stabilize native review counting and interactive navigation Co-Authored-By: OpenAI Codex <noreply@openai.com> * chore: prepare v1.82.0.0 release Co-Authored-By: OpenAI Codex <noreply@openai.com> * fix: eliminate browser and process-cleanup test flakes Pin every CI surface to Bun 1.4.0 to avoid extra-stdio finalizers closing reused live sockets. Add an isolated GC/listener regression that fails on Bun 1.3.13, and prevent coordinated rollback to an affected CI runtime. Check renderer cleanup against the render's own staging directory so concurrent renders cannot invalidate the assertion. Make the no-pgrep process-tree walk tolerate disappearing /proc entries, and synchronize its test fixture through child readiness and pipe EOF instead of sleeps. Validation: 9,157 passed, 31 skipped, zero failures across 556 files with retries disabled. Build, all-host generation freshness, and skill checks passed. All three races have failing-before/passing-after regressions. * fix: count completed native review questions in evals * fix: drive review navigation from confirmed native choices * fix: require complete section-loading eval reports * test: isolate telemetry HTTP transport from local assertions * fix: keep review input on the active native question * test: let tunnel revocation daemon choose an available port * test: allocate available ports for pairing and watchdog fixtures * fix: stabilize planning eval navigation and phase reporting * test: isolate installed runtime paths in planning evals * test: stabilize review evidence and concurrent refresh fixtures * fix: resolve design findings before editing the plan * fix: honor and persist disabled outside plan reviews * fix: preserve planning decisions and terminal evidence Load installed host reviews at autoplan phase entry and wait for completed reviewers and saved artifacts. Reuse approved remedies while preserving individual finding decisions. Drive interactive evals from the current terminal viewport, bind native questions across scrolling, and require complete native report evidence. Cover captured stale menus, permission lifecycles, setup classification, and disabled-review tool availability with deterministic regressions. Advance release metadata and the upgrade migration to the unclaimed 1.83.0.0 slot. * fix: drive native review questions and preserve current plans Use the native single-choice keyboard protocol and current terminal viewport, with per-question navigation inside packets and completed-call coverage. Keep permissions, multi-select menus, and Submit controls distinct. Send Autoplan reviewers the amended implementation plan, keep its review record separate, and supply retained application contracts in the chain fixture. Clarify individual DevEx decisions and complete CEO fix options; use one active plan destination for the section-loading report. * fix: preserve complete plan-review decisions * fix: recognize native plan dialogs and reviewer controls * fix: preserve review decisions and phase completion * fix: recognize completed reviews without losing findings * fix: preserve review continuity and native eval completion * test: fix native review completion and eval retry isolation * test: handle native review menus and complete eval fixtures * test: fix native review setup, completion, and isolation failures * test: limit native skill discovery to runtime assets * fix: bind Autoplan reviews to full ordered phase inputs * test: fix planning eval routing, counting, and timeout handling * chore: advance queued release to v1.84.0.0 * fix: preserve complete review inputs and planning decisions * fix: reconcile review approvals and preserve phase obligations * fix: preserve review obligations and unblock eval permissions Carry recorded Autoplan requirements into blind phase inputs, require Eng review approvals before exit, and exercise combined asynchronous flows in CEO reviews. Correct native finding and handoff classification and unblock repeated report edits using scoped request identities. * fix: retain plan requirements and complete native review dialogs * fix: complete native review prompts and retain plan references * fix: preserve review inputs and classify native eval evidence * fix: check competing completion orders in CEO reviews * fix: recognize review decisions and require phase methodology Require the current phase methodology before Autoplan snapshots. Correct substantive decision, closed handoff, and cache-finding classification, and honor the recommended implementation approach in native review dialogs. Add captured-transcript regressions without changing review thresholds, provider models, retries, or deadlines. * test: bind native review decisions and close completed handoffs * fix: complete review dialogs and verify methodology delivery * fix: preserve review evidence and unblock native eval prompts * fix: handle native review question completions * fix: recognize native review narration and controls * fix: count native review decisions and isolate eval fixtures * test: verify seeded review coverage and current artifact permissions * test: isolate model and brain-aware skill renders * fix: repair native workflow evaluation and clarify review steps * fix: stabilize workflow eval evidence and review guidance * test: repair native workflow observation and fixture isolation * fix: recognize completed workflow evidence and owned skill reads * test: repair seeded workflow delivery and completion evidence * test: recognize current review evidence across native forms * test: handle native review variants and permission redraws * fix: honor review preferences and recognize native eval evidence * test: recognize completed review decisions and queued permissions * test: match current review contracts and partial-line edits * test: recognize completed workflow evidence and bounded human waits * fix: preserve review entry gates and native eval interactions * fix: recognize native workflow evidence and preserve review gates * test: recognize current review evidence and preconfigure workflow fixtures * test: recognize completed review findings and scoped artifact permissions * fix: stabilize native workflow review and permission evidence * fix: recognize current review evidence and scoped edit confirmations Clarify Design and engineering review entry instructions and Design scoring. Recognize required legacy coverage and public Autoplan completion recaps. Bind the pending Edit confirmation to its exact file, ordered digest, and one-request approval when a preceding command display remains visible. Keep reviews within their existing size limits and preserve scope gates when extracting workflow fixtures from either supported preamble header. Keep failure outcomes, review thresholds, provider choices, and eval budgets. * fix: recover review workflow progress and eval evidence * fix: recognize valid review evidence and scope selection * test: fix review evidence parsing and repeated artifact prompts * test: recognize valid review decisions and pending native cards * fix(plan-eng-review): keep final navigation consistent with approved tasks * test: recognize valid review evidence and bind legacy diff requests * fix: stabilize review eval evidence and harness repair guidance * docs: update project documentation for v1.85.0.0 Co-Authored-By: OpenAI Codex <noreply@openai.com> * test: fix Windows CI fixtures and credential scan Rebase captured JSON values and filesystem evidence using the appropriate path convention. Compile native fake CLIs on Windows and synchronize pipe holder readiness, with cleanup retained when assertions fail. Assemble synthetic credential fixtures at runtime so the added-line scan keeps enforcing the same gate without flagging its own rejection controls. Discover generated skills directly for the empty-find regression check, avoiding a recursive scan through saved evaluation artifacts and dependencies. * fix: preserve source renders on Windows Compare canonical generator paths using native separators so an output sidecar pointing at the source cannot overwrite its skill or metadata. Keep the regression fixture isolated from the real checkout and expose freshness diagnostics before asserting subprocess status. Detach Windows drain-test pipe holders from the fake provider's automatic child cleanup while preserving the enclosing runner job and its assertions. * fix: clarify outside review fallback and CEO decisions Render one applicable own-harness fallback path and retain native review, disabled policy, and missing-coverage semantics. Align report field names and mode labels, and make the existing per-cut scope approval explicit. Regenerate skill outputs and keep the workflow judge's model, thresholds, and retry policy unchanged. * chore: move release to free version slot (v1.86.0.0) PR #2852 now claims v1.85.0.0. Align the release metadata and rename migration so upgrades from that version still receive it. Co-Authored-By: OpenAI Codex <noreply@openai.com> * fix: include engineering review prerequisites and restore branch context * fix: recognize coverage diagrams and clarify design review instructions * fix: preserve file identities and join Windows test processes --------- Co-authored-by: OpenAI Codex <noreply@openai.com>
179 lines
12 KiB
JSON
179 lines
12 KiB
JSON
[
|
|
{
|
|
"sessionId": "4cbd2d66-f875-4fcd-b4e2-6c8a6d6716cd",
|
|
"toolUseId": "toolu_01USNDVwq7zfQWz9TRw6tnyi",
|
|
"questions": [
|
|
{
|
|
"question": "D1 \u2014 Pass 1 (Getting Started): The agreed <2 min TTHW target is mathematically impossible with the retained 5-min CI block. Which resolution belongs in the plan? <gstack-qid:devex-d1-tthw-contradiction>",
|
|
"header": "TTHW block",
|
|
"multiSelect": false,
|
|
"options": [
|
|
{
|
|
"label": "Demo-only CI bypass (Recommended)",
|
|
"description": "Add an offline mode specifically for `python -m evalkit.demo`. The demo runs bundled sample_responses.json with no real API submission, so the CI check has nothing to gate. This achieves the <2 min target for the agreed first-success experience without touching the real API contract. Scope impact: small \u2014 only evalkit/demo.py changes."
|
|
},
|
|
{
|
|
"label": "Add --offline flag to run_eval",
|
|
"description": "Give run_eval and run_batch an --offline=True flag that skips the CI block, enabling local dev testing. Achieves <2 min for both demo AND real usage in local mode. Broader fix, but expands scope beyond DX POLISH into API surface change."
|
|
},
|
|
{
|
|
"label": "Update TTHW target to reflect reality",
|
|
"description": "Change the agreed target from <2 min to <7 min in the plan. No code change, no scope expansion. Honest about the current experience, but drops EvalKit below all three competitors (A=2min, B=4min, C=3min) and accepts the 'Needs Work' tier explicitly."
|
|
},
|
|
{
|
|
"label": "Other \u2014 describe the constraint",
|
|
"description": "There's a constraint or design intent I haven't surfaced. Tell me what's actually going on with the CI block and the TTHW target."
|
|
}
|
|
]
|
|
}
|
|
],
|
|
"answered": true,
|
|
"failed": false,
|
|
"answers": {
|
|
"D1 \u2014 Pass 1 (Getting Started): The agreed <2 min TTHW target is mathematically impossible with the retained 5-min CI block. Which resolution belongs in the plan? <gstack-qid:devex-d1-tthw-contradiction>": "Demo-only CI bypass (Recommended)"
|
|
},
|
|
"unansweredQuestionIndices": [],
|
|
"answeredAt": "2026-09-09T11:06:32.086Z"
|
|
},
|
|
{
|
|
"sessionId": "4cbd2d66-f875-4fcd-b4e2-6c8a6d6716cd",
|
|
"toolUseId": "toolu_01Gxtd2wXZEDoKo7YqnZnUi7",
|
|
"questions": [
|
|
{
|
|
"question": "D2 \u2014 Pass 1+4 (Getting Started + Docs): The README tells developers to run `python examples/first_eval.py`, but docs/package-contents.txt explicitly states that file is absent from both the published package and the release examples archive. A developer's first command fails with FileNotFoundError. What's the fix? <gstack-qid:devex-d2-missing-quickstart>",
|
|
"header": "Missing file",
|
|
"multiSelect": false,
|
|
"options": [
|
|
{
|
|
"label": "Update README to point at the demo (Recommended)",
|
|
"description": "Replace the `python examples/first_eval.py` instruction in README with `python -m evalkit.demo`. The demo module IS in the package, produces real scored output, and is the agreed first-success vehicle. Zero new files needed, one line changes in README."
|
|
},
|
|
{
|
|
"label": "Add first_eval.py to the package",
|
|
"description": "Write examples/first_eval.py and include it in the package so the documented path works. More complete \u2014 developers have a real code file they can open and extend. But adds scope: file must be authored, tested, and maintained alongside demo.py."
|
|
},
|
|
{
|
|
"label": "Both: update README AND add the file",
|
|
"description": "Update README to use the demo command as the primary path, AND add examples/first_eval.py as a secondary reference for developers who want editable starter code. Most complete. Two deliverables instead of one."
|
|
},
|
|
{
|
|
"label": "Other \u2014 first_eval.py is coming from somewhere else",
|
|
"description": "The file exists or will exist in a place I haven't looked. Tell me where it is or where it's meant to come from."
|
|
}
|
|
]
|
|
}
|
|
],
|
|
"answered": true,
|
|
"failed": false,
|
|
"answers": {
|
|
"D2 \u2014 Pass 1+4 (Getting Started + Docs): The README tells developers to run `python examples/first_eval.py`, but docs/package-contents.txt explicitly states that file is absent from both the published package and the release examples archive. A developer's first command fails with FileNotFoundError. What's the fix? <gstack-qid:devex-d2-missing-quickstart>": "Update README to point at the demo (Recommended)"
|
|
},
|
|
"unansweredQuestionIndices": [],
|
|
"answeredAt": "2026-09-09T11:06:46.116Z"
|
|
},
|
|
{
|
|
"sessionId": "4cbd2d66-f875-4fcd-b4e2-6c8a6d6716cd",
|
|
"toolUseId": "toolu_012wXLgrh6oDJGHipSk1x5Ti",
|
|
"questions": [
|
|
{
|
|
"question": "D3 \u2014 Pass 3 (Error Messages): The SDK raises `AuthError(\"request failed\")` for an invalid API key. No error code, no explanation of what failed, no instruction for how to fix it. The plan explicitly retains this message. An ML engineer with a typo in EVALKIT_API_KEY sees a two-word error and has no idea what broke. Fix or retain? <gstack-qid:devex-d3-autherror-message>",
|
|
"header": "AuthError msg",
|
|
"multiSelect": false,
|
|
"options": [
|
|
{
|
|
"label": "Fix the error message (Recommended)",
|
|
"description": "Improve to: `AuthError(\"Authentication failed: EVALKIT_API_KEY is invalid or expired. Check your key at console.evalkit.com/keys and set the EVALKIT_API_KEY environment variable.\")`. Follows DX First Principle #5: every error = problem + cause + fix. One-line change in evalkit/client.py."
|
|
},
|
|
{
|
|
"label": "Add error code only",
|
|
"description": "Change to `AuthError(\"auth_invalid: request failed\")` \u2014 adds a stable machine-readable code without restructuring the message. Useful for programmatic error handling but still leaves the developer without a cause or fix."
|
|
},
|
|
{
|
|
"label": "Retain the current message",
|
|
"description": "Explicitly accept this as a known DX debt item. The plan notes it as a post-beta improvement. An ML engineer who hits this will need to consult docs or support to diagnose it."
|
|
}
|
|
]
|
|
}
|
|
],
|
|
"answered": true,
|
|
"failed": false,
|
|
"answers": {
|
|
"D3 \u2014 Pass 3 (Error Messages): The SDK raises `AuthError(\"request failed\")` for an invalid API key. No error code, no explanation of what failed, no instruction for how to fix it. The plan explicitly retains this message. An ML engineer with a typo in EVALKIT_API_KEY sees a two-word error and has no idea what broke. Fix or retain? <gstack-qid:devex-d3-autherror-message>": "Fix the error message (Recommended)"
|
|
},
|
|
"unansweredQuestionIndices": [],
|
|
"answeredAt": "2026-09-09T11:07:08.170Z"
|
|
},
|
|
{
|
|
"sessionId": "4cbd2d66-f875-4fcd-b4e2-6c8a6d6716cd",
|
|
"toolUseId": "toolu_01TSTRSbuz856wiLPpNthPtD",
|
|
"questions": [
|
|
{
|
|
"question": "D4 \u2014 Pass 2 (API Design): `run_eval(dataset, evaluator)` and `run_batch(evaluator, dataset)` have REVERSED positional order for the same concepts. The plan calls this intentional. A developer who writes `run_eval(my_dataset, my_evaluator)` then calls `run_batch` the same way silently passes wrong types \u2014 no error until runtime, potentially confusing results. Fix or keep? <gstack-qid:devex-d4-argorder-reversal>",
|
|
"header": "Arg order",
|
|
"multiSelect": false,
|
|
"options": [
|
|
{
|
|
"label": "Standardize both to (dataset, evaluator) order (Recommended)",
|
|
"description": "Change run_batch signature to `run_batch(dataset, evaluator)` in the beta. Consistent with run_eval, intuitive (dataset first, evaluator second), and v2 is a beta with breaking-change latitude. Fixes the cognitive trap before the API is widely adopted."
|
|
},
|
|
{
|
|
"label": "Enforce keyword-only arguments on both",
|
|
"description": "Add `*` to both signatures: `run_eval(*, dataset, evaluator)` and `run_batch(*, evaluator, dataset)`. Callers must use keyword args; positional mistakes become TypeErrors immediately. Preserves the current positional order while making silent misuse impossible."
|
|
},
|
|
{
|
|
"label": "Retain reversed order, add prominent docs warning",
|
|
"description": "Keep the current signatures and add a big warning box in the docs: 'Note: run_eval and run_batch take arguments in reversed order.' Accepts the cognitive trap but documents it. Lowest implementation cost."
|
|
},
|
|
{
|
|
"label": "Retain as intentional \u2014 it mirrors the underlying evaluator API convention",
|
|
"description": "The reversal follows an existing convention in the underlying evaluator library. Document the rationale in the API reference so the design looks intentional, not accidental."
|
|
}
|
|
]
|
|
}
|
|
],
|
|
"answered": true,
|
|
"failed": false,
|
|
"answers": {
|
|
"D4 \u2014 Pass 2 (API Design): `run_eval(dataset, evaluator)` and `run_batch(evaluator, dataset)` have REVERSED positional order for the same concepts. The plan calls this intentional. A developer who writes `run_eval(my_dataset, my_evaluator)` then calls `run_batch` the same way silently passes wrong types \u2014 no error until runtime, potentially confusing results. Fix or keep? <gstack-qid:devex-d4-argorder-reversal>": "Standardize both to (dataset, evaluator) order (Recommended)"
|
|
},
|
|
"unansweredQuestionIndices": [],
|
|
"answeredAt": "2026-09-09T11:07:30.245Z"
|
|
},
|
|
{
|
|
"sessionId": "4cbd2d66-f875-4fcd-b4e2-6c8a6d6716cd",
|
|
"toolUseId": "toolu_01Cnqz7g1gYHL4r6BeFYpyMV",
|
|
"questions": [
|
|
{
|
|
"question": "D5 \u2014 Pass 5 (Upgrade Path): v1 exposed `Client.evaluate()`. v2 removes it immediately and replaces it with `Client.run()`. No alias, no deprecation warning, no migration guide, no codemod. Every existing v1 user who runs `pip install evalkit==2.0.0b1` and calls `Client.evaluate()` gets an `AttributeError` with no guidance. The plan retains this. What bridge should the beta include? <gstack-qid:devex-d5-v1v2-migration>",
|
|
"header": "v1\u2192v2 upgrade",
|
|
"multiSelect": false,
|
|
"options": [
|
|
{
|
|
"label": "Deprecation alias + migration guide (Recommended)",
|
|
"description": "Add `Client.evaluate = Client.run` with a DeprecationWarning: 'evaluate() is deprecated in v2; use run() instead.' Add one section to the changelog explaining the rename. The alias makes the upgrade non-breaking for v1 users; the warning tells them to update. Removes entirely in v3."
|
|
},
|
|
{
|
|
"label": "Migration guide only, no alias",
|
|
"description": "Write a v1\u2192v2 migration section in CHANGELOG.md (and/or README) documenting `evaluate()` \u2192 `run()`. No code change. Developers who read changelogs before upgrading are fine; developers who just run pip upgrade hit an AttributeError and must consult docs."
|
|
},
|
|
{
|
|
"label": "Hard break as-is, add clear error message",
|
|
"description": "Keep the hard break but improve the AttributeError: catch `__getattr__` and raise a descriptive error: `AttributeError: Client.evaluate() was renamed to Client.run() in v2. See CHANGELOG.md.` The developer knows exactly what to do, no alias needed."
|
|
},
|
|
{
|
|
"label": "Retain as-is \u2014 it's a beta, accept the break",
|
|
"description": "Beta users accept breaking changes. The changelog mentions it. Developers who upgrade without reading the changelog chose to take the risk. No additional bridge needed."
|
|
}
|
|
]
|
|
}
|
|
],
|
|
"answered": true,
|
|
"failed": false,
|
|
"answers": {
|
|
"D5 \u2014 Pass 5 (Upgrade Path): v1 exposed `Client.evaluate()`. v2 removes it immediately and replaces it with `Client.run()`. No alias, no deprecation warning, no migration guide, no codemod. Every existing v1 user who runs `pip install evalkit==2.0.0b1` and calls `Client.evaluate()` gets an `AttributeError` with no guidance. The plan retains this. What bridge should the beta include? <gstack-qid:devex-d5-v1v2-migration>": "Deprecation alias + migration guide (Recommended)"
|
|
},
|
|
"unansweredQuestionIndices": [],
|
|
"answeredAt": "2026-09-09T11:07:44.259Z"
|
|
}
|
|
]
|