shared-libs-review-lifecycle ran ~88% of its 300 s session budget (12-run
census median 265 s, 4/24 sessions timed out). The fixture now executes pass 1's
Step 3 once with the real logger and Git: a real unused REVIEW_START, then the
diff, inventories, attributes/config/index flags, gstack-review-read output and
every file's bytes and sha256, saved to one observation. The model resumes at
Step 4 with an exact four-file first read, the observation named as the
authoritative pass-1 repository read, one post-fix verification, an explicit
pass-2 read list and a twelve-line summary. Pass 2 still runs its own --start,
diff, reads, fingerprint and stage actor before --finish.
The actor scope now states that a current settled final-pass actor result
supplies the replaced QA/adversarial prerequisites and that the no-credit
disclosure is a reporting label: one r1 session persisted completed:false
from that ambiguity.
New assertions: the final binding never uses the seeded token's start or tree,
and the observation was read; free controls finish the seeded token (binding
changed) and omit the observation read, and both fail.
The skill made the model retype a long safe-Git prefix on each call and a
dropped flag failed shared-libs-read-only. bin/gstack-safe-git applies the
fixed env + flag prefix, adds --no-ext-diff --no-textconv to log/show/diff,
allows diff only between two explicit object IDs and ls-files only in the
NUL-delimited overlay form, and refuses every other shape with one line
naming the allowed forms. The template now points at the installed helper
(host global runtime via {{SAFE_GIT}}) and drops the prose it enforces.
Fixtures resolve the helper to this checkout, the git shim records the safety
environment, and isGuardedGitRequest requires the complete prefix (env
included) for every repository read.
* feat: bind shared-code review advice to source and branch
* feat: add shared-code extraction audit and scoped review checks
* test: recognize complete source reads and explicit coverage legends
* chore: bump version and changelog (v1.88.0.0)
Co-Authored-By: OpenAI Codex <noreply@openai.com>
* test: capture native review questions and retain public evidence
Capture the actual first public native question with strict ownership and display matching. Preserve terminal failures and raw evidence, and retain SDK completion checks.
* test: recognize verified review evidence and complete fixtures
Recognize complete source and diagram evidence, concrete design and developer-experience decisions, and the complete planted scenario contracts. Preserve negative controls and grading thresholds.
* fix: preserve decision brief structure in native questions
Keep the required pros-and-cons heading and final Net field in native question text. Regenerate host outputs and document the release and evaluation repairs.
Co-Authored-By: OpenAI Codex <noreply@openai.com>
* docs: update project documentation for v1.88.0.0
Co-Authored-By: OpenAI Codex <noreply@openai.com>
* fix: correct eval retry accounting and ship workflow gates
* fix: capture native eval evidence and stabilize CI fixtures
* fix: keep shared-code eval skips read-only
Choose explicit no-change answers instead of mixed fix/preservation options.
Reuse the bounded revalidation prompt for path fixtures so required review
metadata is available without repeated discovery. Preserve source checks,
retry limits, and failed native terminal outcomes.
Add captured-question and callback regressions, plus evaluation selection
coverage for the affected fixtures.
---------
Co-authored-by: OpenAI Codex <noreply@openai.com>