Files
gstack/test/fixtures/plan-count-crop-ak.json
Garry TanandOpenAI Codex 9f81911136 v1.86.0.0 feat: route outside reviews by harness (#2850)
* feat: add a restricted and supervised Claude Code runner

Preserve configured authentication and models while enforcing tool access, strict completion JSON, bounded output and process cleanup. Cover argv, failure handling, session metadata and Windows process containment.

* feat: route outside reviews by harness and migrate wrapper installs

Use Claude Code from Codex and Codex from other supported hosts, with shared invocation rendering, positive gate validation and per-phase provenance. Rename /claude to /claude-code, repair managed shared and copied installations safely, and generate native Kiro skills. Add installed-workflow, failure-injection and live cross-harness regression coverage.

* test: recognize CEO mode labels without terminal spacing

The paid workflow rendered SCOPEEXPANSION at option 4, but its driver required a literal space. Match the leading mode title without cursor-spacing artifacts and ignore adjacent preview text. Preserve missing-target failures and downstream posture assertions.

* test: isolate plan-count fixtures before starting review workflows

Seed the complete test plan in a private git repository before launching Claude, so a bare slash command cannot review the live workspace while a delayed fixture message remains queued. Preserve count thresholds, parsers and budgets. Add initial-context and installed-discovery tests, and retain startup/terminal diagnostics on failed evaluations.

* test: stabilize review fixtures and Claude eval startup

Preserve source boundaries in workflow judge inputs, isolate CEO mode plans, and wait for interactive trust input readiness. Keep startup failure evidence and retain existing models, budgets, and assertions.

Co-Authored-By: OpenAI Codex <noreply@openai.com>

* test: classify collapsed review modes and isolate seeded findings

Keep review questions out of the setup count when terminal cursor positioning removes spaces. State existing webhook safeguards so the five-finding control measures its seeded defects without accidental extra security and concurrency gaps. Preserve question bands and the paired control.

Co-Authored-By: OpenAI Codex <noreply@openai.com>

* test: isolate browser daemon state across free shards

Co-Authored-By: OpenAI Codex <noreply@openai.com>

* test: stabilize native review counting and interactive navigation

Co-Authored-By: OpenAI Codex <noreply@openai.com>

* chore: prepare v1.82.0.0 release

Co-Authored-By: OpenAI Codex <noreply@openai.com>

* fix: eliminate browser and process-cleanup test flakes

Pin every CI surface to Bun 1.4.0 to avoid extra-stdio finalizers closing
reused live sockets. Add an isolated GC/listener regression that fails on
Bun 1.3.13, and prevent coordinated rollback to an affected CI runtime.

Check renderer cleanup against the render's own staging directory so
concurrent renders cannot invalidate the assertion. Make the no-pgrep
process-tree walk tolerate disappearing /proc entries, and synchronize
its test fixture through child readiness and pipe EOF instead of sleeps.

Validation: 9,157 passed, 31 skipped, zero failures across 556 files with
retries disabled. Build, all-host generation freshness, and skill checks
passed. All three races have failing-before/passing-after regressions.

* fix: count completed native review questions in evals

* fix: drive review navigation from confirmed native choices

* fix: require complete section-loading eval reports

* test: isolate telemetry HTTP transport from local assertions

* fix: keep review input on the active native question

* test: let tunnel revocation daemon choose an available port

* test: allocate available ports for pairing and watchdog fixtures

* fix: stabilize planning eval navigation and phase reporting

* test: isolate installed runtime paths in planning evals

* test: stabilize review evidence and concurrent refresh fixtures

* fix: resolve design findings before editing the plan

* fix: honor and persist disabled outside plan reviews

* fix: preserve planning decisions and terminal evidence

Load installed host reviews at autoplan phase entry and wait for completed
reviewers and saved artifacts. Reuse approved remedies while preserving
individual finding decisions.

Drive interactive evals from the current terminal viewport, bind native
questions across scrolling, and require complete native report evidence.
Cover captured stale menus, permission lifecycles, setup classification,
and disabled-review tool availability with deterministic regressions.

Advance release metadata and the upgrade migration to the unclaimed
1.83.0.0 slot.

* fix: drive native review questions and preserve current plans

Use the native single-choice keyboard protocol and current terminal viewport,
with per-question navigation inside packets and completed-call coverage.
Keep permissions, multi-select menus, and Submit controls distinct.

Send Autoplan reviewers the amended implementation plan, keep its review record
separate, and supply retained application contracts in the chain fixture.
Clarify individual DevEx decisions and complete CEO fix options; use one active
plan destination for the section-loading report.

* fix: preserve complete plan-review decisions

* fix: recognize native plan dialogs and reviewer controls

* fix: preserve review decisions and phase completion

* fix: recognize completed reviews without losing findings

* fix: preserve review continuity and native eval completion

* test: fix native review completion and eval retry isolation

* test: handle native review menus and complete eval fixtures

* test: fix native review setup, completion, and isolation failures

* test: limit native skill discovery to runtime assets

* fix: bind Autoplan reviews to full ordered phase inputs

* test: fix planning eval routing, counting, and timeout handling

* chore: advance queued release to v1.84.0.0

* fix: preserve complete review inputs and planning decisions

* fix: reconcile review approvals and preserve phase obligations

* fix: preserve review obligations and unblock eval permissions

Carry recorded Autoplan requirements into blind phase inputs, require Eng
review approvals before exit, and exercise combined asynchronous flows in
CEO reviews. Correct native finding and handoff classification and unblock
repeated report edits using scoped request identities.

* fix: retain plan requirements and complete native review dialogs

* fix: complete native review prompts and retain plan references

* fix: preserve review inputs and classify native eval evidence

* fix: check competing completion orders in CEO reviews

* fix: recognize review decisions and require phase methodology

Require the current phase methodology before Autoplan snapshots. Correct
substantive decision, closed handoff, and cache-finding classification, and
honor the recommended implementation approach in native review dialogs.

Add captured-transcript regressions without changing review thresholds,
provider models, retries, or deadlines.

* test: bind native review decisions and close completed handoffs

* fix: complete review dialogs and verify methodology delivery

* fix: preserve review evidence and unblock native eval prompts

* fix: handle native review question completions

* fix: recognize native review narration and controls

* fix: count native review decisions and isolate eval fixtures

* test: verify seeded review coverage and current artifact permissions

* test: isolate model and brain-aware skill renders

* fix: repair native workflow evaluation and clarify review steps

* fix: stabilize workflow eval evidence and review guidance

* test: repair native workflow observation and fixture isolation

* fix: recognize completed workflow evidence and owned skill reads

* test: repair seeded workflow delivery and completion evidence

* test: recognize current review evidence across native forms

* test: handle native review variants and permission redraws

* fix: honor review preferences and recognize native eval evidence

* test: recognize completed review decisions and queued permissions

* test: match current review contracts and partial-line edits

* test: recognize completed workflow evidence and bounded human waits

* fix: preserve review entry gates and native eval interactions

* fix: recognize native workflow evidence and preserve review gates

* test: recognize current review evidence and preconfigure workflow fixtures

* test: recognize completed review findings and scoped artifact permissions

* fix: stabilize native workflow review and permission evidence

* fix: recognize current review evidence and scoped edit confirmations

Clarify Design and engineering review entry instructions and Design scoring.
Recognize required legacy coverage and public Autoplan completion recaps.
Bind the pending Edit confirmation to its exact file, ordered digest, and
one-request approval when a preceding command display remains visible.
Keep reviews within their existing size limits and preserve scope gates
when extracting workflow fixtures from either supported preamble header.

Keep failure outcomes, review thresholds, provider choices, and eval budgets.

* fix: recover review workflow progress and eval evidence

* fix: recognize valid review evidence and scope selection

* test: fix review evidence parsing and repeated artifact prompts

* test: recognize valid review decisions and pending native cards

* fix(plan-eng-review): keep final navigation consistent with approved tasks

* test: recognize valid review evidence and bind legacy diff requests

* fix: stabilize review eval evidence and harness repair guidance

* docs: update project documentation for v1.85.0.0

Co-Authored-By: OpenAI Codex <noreply@openai.com>

* test: fix Windows CI fixtures and credential scan

Rebase captured JSON values and filesystem evidence using the appropriate
path convention. Compile native fake CLIs on Windows and synchronize pipe
holder readiness, with cleanup retained when assertions fail.

Assemble synthetic credential fixtures at runtime so the added-line scan
keeps enforcing the same gate without flagging its own rejection controls.

Discover generated skills directly for the empty-find regression check,
avoiding a recursive scan through saved evaluation artifacts and dependencies.

* fix: preserve source renders on Windows

Compare canonical generator paths using native separators so an output
sidecar pointing at the source cannot overwrite its skill or metadata.
Keep the regression fixture isolated from the real checkout and expose
freshness diagnostics before asserting subprocess status.

Detach Windows drain-test pipe holders from the fake provider's automatic
child cleanup while preserving the enclosing runner job and its assertions.

* fix: clarify outside review fallback and CEO decisions

Render one applicable own-harness fallback path and retain native review,
disabled policy, and missing-coverage semantics. Align report field names
and mode labels, and make the existing per-cut scope approval explicit.

Regenerate skill outputs and keep the workflow judge's model, thresholds,
and retry policy unchanged.

* chore: move release to free version slot (v1.86.0.0)

PR #2852 now claims v1.85.0.0. Align the release metadata and
rename migration so upgrades from that version still receive it.

Co-Authored-By: OpenAI Codex <noreply@openai.com>

* fix: include engineering review prerequisites and restore branch context

* fix: recognize coverage diagrams and clarify design review instructions

* fix: preserve file identities and join Windows test processes

---------

Co-authored-by: OpenAI Codex <noreply@openai.com>
2026-09-14 14:32:45 -07:00

44 lines
12 KiB
JSON

{
"sourceCommit": "04c62ac678bb7bc1a22090f72f7ed51c451c22b9",
"screen": " f2 (~7.6:1) with icon and explicit text | PENDING |\n 82 \n 83 +## Information hierarchy (Pass 1, approved) \n 84 +What the user sees first, second, third on the page: \n 85 +```text \n 86 +1st Account settings (h1) + description orientation: where am I \n 87 +2nd [ Save ] Reset Cancel Export one filled button = the action \n 88 + ^ #1d4ed8 filled, white text ^ neutral ghost buttons \n 89 +3rd InlineStatus (blank / Unsaved changes / Saved at HH:mm) \n 90 +4th Profile (h2) \u2192 Display name, Email \n 91 +5th Notifications (h2) \u2192 Weekly digest, Product tips \n 92 +``` \n 93 +Constraint check: four header actions stay (Export is required by the accepted \n 94 +behavior), but only one carries fill. Ghost treatment on Reset/Cancel/Export \n 95 +keeps them findable without competing with Save. \n 96 + \n 97 ## Implementation Tasks\n 84 -_Populated only with individually approved remedies. Nothing approved yet._ \n 98 +Synthesized from this review's findings. Each task derives from a specific \n 99 +finding above. Run with Claude Code or Codex; checkbox as you ship. \n 100 \n 101 +- [ ] **T1 (P1, human: ~1h / CC: ~5min)** \u2014 Header action Buttons \u2014 Render Save with the existing filled primary v\n +ariant (#1d4ed8, white text); render Reset, Cancel and Export with the existing neutral ghost variant. \n 102 + - Surfaced by: Pass 1 Information Architecture \u2014 Issue 1 (D2 \u2192 1A), \"Save is rendered with the same size, weight\n +, and color as three other buttons\" \n 103 + - Files: settings page header action group; Button variant usage (no new tokens) \n 104 + - Verify: visual check at 1280px and 320px shows exactly one filled button; focus ring still 2px #1d4ed8 offset \n +2px on all four; disabled/pending states keep the variant \n 105 + \n 106 ## NOT in scope\n 107 _To be written after the review passes._\n\u254c\u254c\u254c\u254c\u254c\u254c\u254c\u254c\u254c\u254c\u254c\u254c\u254c\u254c\u254c\u254c\u254c\u254c\u254c\u254c\u254c\u254c\u254c\u254c\u254c\u254c\u254c\u254c\u254c\u254c\u254c\u254c\u254c\u254c\u254c\u254c\u254c\u254c\u254c\u254c\u254c\u254c\u254c\u254c\u254c\u254c\u254c\u254c\u254c\u254c\u254c\u254c\u254c\u254c\u254c\u254c\u254c\u254c\u254c\u254c\u254c\u254c\u254c\u254c\u254c\u254c\u254c\u254c\u254c\u254c\u254c\u254c\u254c\u254c\u254c\u254c\u254c\u254c\u254c\u254c\u254c\u254c\u254c\u254c\u254c\u254c\u254c\u254c\u254c\u254c\u254c\u254c\u254c\u254c\u254c\u254c\u254c\u254c\u254c\u254c\u254c\u254c\u254c\u254c\u254c\u254c\u254c\u254c\u254c\u254c\u254c\u254c\u254c\u254c\u254c\u254c\u254c\u254c\u254c\u254c\n Do you want to make this edit to gstack-test-plan-design.md?\n \u276f 1. Yes\n 2. Yes, and switch to accept edits (auto-approve file edits and common file commands) for this session; Yes, and\n always allow access to /tmp/gstack-paid-shard-m7O6cU/tmp/gstack-e2e-plan-design-B4lVyk for this session\n 3. Nohift+tab)\n\n Esc to cancel \u00b7 Tab to amend\n",
"hook": {
"cwd": "/tmp/gstack-paid-shard-m7O6cU/tmp/gstack-plan-count-0rEMmZ",
"expected": "/tmp/gstack-paid-shard-m7O6cU/tmp/gstack-e2e-plan-design-B4lVyk/gstack-test-plan-design.md",
"sessionId": "ca736f12-33d6-4ee1-831b-5924db71bf82",
"transcriptPath": "/tmp/gstack-paid-shard-m7O6cU/tmp/gstack-hermetic-1325936-1eazEZ/with-skills/.claude/projects/-tmp-gstack-paid-shard-m7O6cU-tmp-gstack-plan-count-0rEMmZ/ca736f12-33d6-4ee1-831b-5924db71bf82.jsonl",
"seenIds": [
"ca736f12-33d6-4ee1-831b-5924db71bf82:toolu_01JqWqNxJZtP7tKSd1viyqzo",
"ca736f12-33d6-4ee1-831b-5924db71bf82:toolu_014NkKMqRz1245RDiWFVMm4J",
"ca736f12-33d6-4ee1-831b-5924db71bf82:toolu_01APPQpBXrbJRqrQwHR5p5mR"
],
"pendingId": "ca736f12-33d6-4ee1-831b-5924db71bf82:toolu_01APPQpBXrbJRqrQwHR5p5mR",
"completedId": "ca736f12-33d6-4ee1-831b-5924db71bf82:toolu_014NkKMqRz1245RDiWFVMm4J",
"completedIds": [
"ca736f12-33d6-4ee1-831b-5924db71bf82:toolu_01JqWqNxJZtP7tKSd1viyqzo",
"ca736f12-33d6-4ee1-831b-5924db71bf82:toolu_014NkKMqRz1245RDiWFVMm4J"
],
"timestamp": "2026-09-10T06:42:26.280Z"
},
"cwd": "/tmp/gstack-paid-shard-m7O6cU/tmp/gstack-plan-count-0rEMmZ",
"config": "/tmp/gstack-paid-shard-m7O6cU/tmp/gstack-hermetic-1325936-1eazEZ/with-skills/.claude",
"startedAt": 1789022312505,
"transcript": {
"status": "ready",
"calls": [],
"assistantMessages": [
{
"sessionId": "ca736f12-33d6-4ee1-831b-5924db71bf82",
"text": "No prior learnings found, so I'll start with the review passes. Pass 1 (Information Architecture) scores 6/10: the DOM order is solid, but the four header buttons all render with equal visual weight, giving no clear hierarchy for finding Save first\u2014violating Design Principle 2 and Krug's satisficing rule.",
"timestamp": "2026-09-10T06:41:51.810Z"
}
]
},
"ownedBefore": "# Plan: Settings Page UI redesign (design-reviewed)\n\n## Context\nThis updates the existing account-settings form to match the checked-in\nDESIGN.md. The shell, components (Button, Field, InlineStatus, ErrorSummary,\nConfirmationDialog), interaction states, accessibility contract and responsive\nbehavior already exist and are accepted. The work is closing five documented\nvisual inconsistencies between the current form and the approved design system,\nwithout visual exploration or component replacement.\n\nSource plan: `PLAN.md`. Design system: `DESIGN.md`. Review: `/plan-design-review`\n(text-only, native voice only, per user request).\n\n## Existing product and accepted behavior (inherited, unchanged)\n- Single-column settings shell, 640px max width: persistent app navigation,\n page title and description, action group, InlineStatus, Profile and\n Notifications fieldsets. Reuse existing Button, Field, InlineStatus,\n ErrorSummary and ConfirmationDialog; no new component family.\n- Description stays verbatim: \u201cManage your display name, email address, and\n notification preferences.\u201d\n- \u201cAccount settings\u201d is h1. Profile and Notifications are h2 headings labelling\n their fieldsets via aria-labelledby; no heading level skipped.\n- DOM and visual order:\n ```text\n Persistent app navigation\n main: Account settings (h1) + description\n Save | Reset | Cancel | Export\n InlineStatus\n Profile (h2): Display name, Email\n Notifications (h2): Weekly digest, Product tips\n ```\n- Profile: Display name (text), Email (email). Notifications: Weekly digest and\n Product tips switches. New-account defaults: account name/email, Weekly digest\n on, Product tips off. Labels and order fixed. Initial load uses the existing\n form skeleton; read failure shows Retry (aria-label \u201cRetry loading\u201d).\n- InlineStatus: role=status, aria-live=polite, aria-atomic=true. After success:\n \u201cSaved at HH:mm\u201d (local 24-hour, existing formatter). Editing away from saved\n values: \u201cUnsaved changes\u201d (text, never color alone). Reverting all edits or\n confirming Reset restores the last save timestamp. Before first successful\n save: blank status; editing shows \u201cUnsaved changes\u201d; revert/Reset restores\n blank. Failed saves keep \u201cUnsaved changes\u201d beside the error.\n- Save is atomic. Validation errors appear beside fields and in the linked\n ErrorSummary (aria-describedby); focus the first invalid field. Network\n failure preserves edits and shows Retry in the status area. Retry is a\n sibling button outside the live region; visible text \u201cRetry\u201d, aria-label\n \u201cRetry save\u201d / \u201cRetry export\u201d.\n- Pending rules: repeat Save disabled while pending. Save and Export mutually\n exclusive; Reset and Cancel disabled while either is pending; all four\n re-enable on settle. Export uses the existing inline spinner beside\n \u201cExporting\u2026\u201d inside its disabled button, aria-busy=true, reduced-motion\n aware. After Save finishes, Export downloads the latest successfully saved\n preferences as JSON. Export failure uses the same error/retry area, keeps\n unsaved fields; Retry repeats Export; success clears only the Export error.\n- Reset restores saved values after confirmation. Dialog: \u201cDiscard unsaved\n changes?\u201d / \u201cYour saved preferences will be restored.\u201d Actions \u201cKeep editing\u201d\n (default) and \u201cDiscard changes\u201d. Cancel navigation uses the same dialog with\n \u201cKeep editing\u201d (default) and \u201cDiscard and leave\u201d. Dialogs trap focus; Escape\n cancels; closing while staying returns focus to the Reset/Cancel trigger;\n confirmed navigation uses the existing destination-main-heading focus.\n- Responsive: above 640px header actions in one row. At 640px and below:\n full-width Save first, then Reset/Cancel/Export in one equal-column row,\n DOM/tab order preserved, fits 320px with 44px targets, no horizontal scroll.\n- A11y: 44px targets everywhere; visible focus-visible = existing 2px solid\n #1d4ed8 outline, 2px offset on white (measured contrast above 3:1) on every\n control and dialog action; semantic fieldsets/labels/main landmark; respect\n reduced motion; Export remains a clearly labelled button.\n- Font family stays the existing system-ui, sans-serif, inherited by form\n controls (explicit scope decision in PLAN.md; not reopened by this review).\n\n## Design gaps under review (PENDING \u2014 no remedy approved yet)\nEach gap below is a documented violation of DESIGN.md. The proposed remedy is a\nrecommendation, not a decision. Each needs its own approval before it moves\ninto Implementation Tasks.\n\n| # | Gap (current form) | Proposed remedy (DESIGN.md) | Status |\n|---|--------------------|-----------------------------|--------|\n| 1 | Visual hierarchy: Save renders identically to Reset/Cancel/Export; no visible primary action | Save is the only filled primary button (#1d4ed8, white text); Reset/Cancel/Export are neutral ghost buttons | APPROVED (D2 \u2192 1A) |\n| 2 | Motion: Save takes 2-5s with no loading indicator; page looks frozen | Existing pending pattern: inline spinner beside \u201cSaving\u2026\u201d inside the disabled Save button, aria-busy=true, reduced-motion aware | PENDING |\n| 3 | Typography: form labels use 14px, 16px and 18px | Two roles only: 16px body/labels/helper text, 20px h2 section headings | PENDING |\n| 4 | Spacing: section gaps are 16px, 24px and 32px with no rhythm | 8px base scale: sections 32px, field groups 24px, label-to-input 8px | PENDING |\n| 5 | Color: error message is red on light pink at ~3:1, below WCAG AA | error.text #991b1b on error.surface #fef2f2 (~7.6:1) with icon and explicit text | PENDING |\n\n## Implementation Tasks\n_Populated only with individually approved remedies. Nothing approved yet._\n\n## NOT in scope\n_To be written after the review passes._\n\n## What already exists\n_To be written after the review passes._\n",
"provenance": {
"permissionReplaySHA": "89687a6694f5b33e2c9a2a5d4e6a5de031c658ec1850e708019d791e6c15b632",
"publicToolsProofSHA": "2c8b892a108f26ca0d05ec6a96714631e2e83fe309817c2d6c374444157b3157",
"diskFileRetained": false,
"limitation": "File bytes reconstructed only from the two acknowledged public Write calls after the natural timeout removed the fixture. No direct retained disk-file comparison."
}
}