- /ship design-lite: the probe is mandatory and any non-ready first line is
stated, matching /review (5 of 6 captured /review trials had skipped it).
- shared-libs-pr-coverage: the first PR 42 page-1 read printed only a jq error,
so the one refetch is a legitimate recovery, charged to the same budget.
- shared-libs-review-prior-coverage: the Skip option said a future review can
"reuse it once snapshot coverage holds"; a conditional tail on the recorded
decision is not product work. Captured-text regressions and negative controls.
#2994 deletes the plan-*-finding-count evals, ceo-payment-findings.ts and
design-count-review.ts. Drop the CEO throw diagnostics and Design boundary
work with them, and drop the structured completion predicate, stopReason,
review-log binding and plan/review-log evidence copy: no surviving
runPlanSkillCounting caller passes expectedPlanPath, so they would be dead
code. Keep idleFor in timeout summaries (every counting caller can time
out), asserted in the existing timeout test. W7 and W8 are unchanged.
The shared-libs shim served 2 PRs for pulls?state=all and endless full pages
for state=open. gh pr list, pulls?state=open|all|closed (per_page/page,
short last page, direction) and search/issues now page one deterministic
table: PR 7, 600 older open PRs, PR 42 and 3 closed PRs, so five 100-item
open-metadata pages still leave older open PRs unchecked. The Contents API
lists pinned directories (the captured attempt got 404 for contents/ and
contents/src while files resolved, then fell back to a raw host), unknown
endpoints return 404 instead of repo metadata, and the read-only detector
is unchanged. Free tests cover view agreement, the budget bound, gh/curl
agreement and the empty world.
Dual-voice outside-voice failures now report probeToolUseId, probeMode and
the canonical-match result with the reason the probe output was rejected.
Keep both intents: v1.91.7.0's functional QA, docsync and exploratory
paid cases and their free owners stay; this branch's deletions stay
deleted. main's new paid keys follow the derived-closure touchfile rule
(free *.test.ts paths dropped, static helper/fixture closure added), its
new helper-only tests join the ratchet baseline, and its free selection
examples that named free test files now assert the derived selection.
Periodic CI keeps seven slices without the retired Autoplan slice; the
gate census keeps seven single-worker slices with --skip-judges. Wall
and census literals are recomputed from the merged planner, durations
are re-recorded on Ubicloud, and VERSION stays 1.91.8.0 above 1.91.7.0.
touchfiles.test.ts now checks, per key, that the paid file's static
test/helpers and test/fixtures closure (plus fixture paths it names in string
literals) is covered, and names the file, path, chain and key to fix when it is
not. Free *.test.ts files are no longer touchfiles, so editing a free replay
test stops selecting paid evals: 950 entries removed, 653 real closure paths
added. The hand-copied inventories go: periodic-fixture-selection,
fake-impeccable-touchfiles and 45 per-file selection examples. Selection for
the sample edits (plan-eng-review template, claude-pty-runner,
plan-count-fixture, gstack-config) loses no case under either profile.
CONTRIBUTING documents the rule and its lower bound.
* perf: remove repeated test work and preserve AUQ execution budgets
* fix: validate native evaluation fixture evidence at its actual boundaries
* fix: clarify deployment approval and recovery state transitions
* chore: document coverage and release v1.90.2.0
* test: preserve Windows scheduling and native no-change consent
* test: recognize verified reads through fixture symlinks
* test: isolate alias-name installation from runtime assets
* fix(browse): prepare reliable cookie import wave for validation
* ci: sequence quality and behavior for validation branch
* fix(browse): isolate Windows qualification and preserve native diagnostics
* test(browse): cover cookie workflow quality and isolate Windows user paths
* test(browse): trace native member startup and initialize fresh folders
* fix(browse): keep Windows member stdin alive through EOF
* fix(browse): latch native timeouts and compare contained Edge startup
* test(browse): verify native version metadata and actual Windows argv
* test(browse): qualify Dia import on isolated macOS CI
* fix(browse): require picker origin for session mutations
* fix(browse): bound credential reads through stream completion
* test(browse): inspect owned Windows process arguments natively
* test(evals): preserve passing coverage during cookie repair reruns
* test(browse): isolate Dia qualification in a fresh macOS account
* test(browse): pass bounded integer timeouts to native Mac probes
* test(browse): distinguish Windows profile initialization from containment
* test(browse): await descendant pipe readiness before parent exit
* test(browse): initialize and restore isolated macOS Keychain state
* test(browse): initialize Windows fixture folders before qualification
* test(ci): pin the same Node runtime across Windows checks
* test(browse): distinguish native macOS browser preflight stages
* test(browse): isolate Windows descendant console lifetime
* test(browse): preserve native receipts and identify fixture lock holders
* test(browse): prepare dependency resolution before native Mac worker startup
* test(ci): include lock and close checks in native diagnostics
* test(browse): preserve native owner probe stages and subprocess deadlines
* fix(browse): classify Chromium profile-in-use exit precisely
* test(browse): retain Mac qualification evidence through cleanup failures
* test(browse): bound Mac fixture paths and retire its owned user domain
* test(browse): accept vanished fixture entries without weakening cleanup
* test(browse): identify probe-created macOS user domains safely
* test(browse): observe Mac user domains without targeting them first
* test(browse): use passive fresh-user ownership throughout Mac qualification
* test(browse): distinguish profile and registered-home Keychain lookups
* test(browse): qualify Dia under one registered account home
* test(browse): identify Dia startup and owned process-group failures
* test(browse): classify bounded Dia startup diagnostics without leaking output
* fix(test): preserve native Mac sandboxing and reap owned browser children
* fix(browse): preserve Chromium sandboxing for native profile imports
* test(browse): inspect signed Mach-O architecture without launching Xcode tools
* test(browse): sample pending Dia startup and reap on all cleanup paths
* test(browse): compare protected Dia launches in fresh Bun and Node accounts
* test(browse): inspect isolated Mac GUI readiness without browser access
* v1.90.0.0 fix: bind cookie picker actions to their document
* test: validate cookie guards and fit nested launch fixtures
* ci: configure the bundled Chromium sandbox helper
* fix(browse): classify Playwright authentication timeouts
* test: retain bounded Windows lifecycle diagnostics
* test(cso): reuse bounded NTFS precision candidates
* test(review): handle explicit preservation choices safely
* test(browse): remove owned fixture directories with explicit primitives
* test(review): distinguish descriptive reuse from edit commitments
* test: admit only the approved unscored cookie workflow refusal
* test: keep the Office Hours judge mock export-complete
* fix: keep dependency-free CI planners independent of the model SDK
* test: observe the exact holder after a native fixture unlink failure
* fix: start seeded PTY observations at owned readiness
* test: acquire identity-bound Windows deletion admission before profile resets
* test: preserve qualified Git index bits without authorizing mutations
* feat: bind shared-code review advice to source and branch
* feat: add shared-code extraction audit and scoped review checks
* test: recognize complete source reads and explicit coverage legends
* chore: bump version and changelog (v1.88.0.0)
Co-Authored-By: OpenAI Codex <noreply@openai.com>
* test: capture native review questions and retain public evidence
Capture the actual first public native question with strict ownership and display matching. Preserve terminal failures and raw evidence, and retain SDK completion checks.
* test: recognize verified review evidence and complete fixtures
Recognize complete source and diagram evidence, concrete design and developer-experience decisions, and the complete planted scenario contracts. Preserve negative controls and grading thresholds.
* fix: preserve decision brief structure in native questions
Keep the required pros-and-cons heading and final Net field in native question text. Regenerate host outputs and document the release and evaluation repairs.
Co-Authored-By: OpenAI Codex <noreply@openai.com>
* docs: update project documentation for v1.88.0.0
Co-Authored-By: OpenAI Codex <noreply@openai.com>
* fix: correct eval retry accounting and ship workflow gates
* fix: capture native eval evidence and stabilize CI fixtures
* fix: keep shared-code eval skips read-only
Choose explicit no-change answers instead of mixed fix/preservation options.
Reuse the bounded revalidation prompt for path fixtures so required review
metadata is available without repeated discovery. Preserve source checks,
retry limits, and failed native terminal outcomes.
Add captured-question and callback regressions, plus evaluation selection
coverage for the affected fixtures.
---------
Co-authored-by: OpenAI Codex <noreply@openai.com>