Files
gstack/test/codex-hardening.test.ts
T
Garry TanandClaude Fable 5 394db326f2 v1.71.0.0 feat: token-load reduction — preamble runtime scripts, gated onboarding, 20 skill carves, CLAUDE.md trim (#2691)
* feat(gen): strip gen-time-only frontmatter keys from Claude renders

interactive + benefits-from are read from the .tmpl by buildContext at
generation time; no runtime, host, or test reader consumes them from the
generated SKILL.md (e2e-harness-audit reads .tmpl; benefits-from tests
assert rendered prose). gbrain: stays (bin/gstack-brain-context-load reads
it from the installed render); hooks: stays (Claude Code host wires
PreToolUse from it).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* chore(gen): regenerate SKILL.md — dead frontmatter keys removed

Mechanical regen after hosts/claude.ts stripFields change.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(test): context-budget ratchet — CI ceilings on always-on + eager token ledgers

New free test grades the two ledgers nothing else guards: the full-frontmatter
always-on catalog (aggregate) and per-skill eager tokens (SKILL.md +
forced-read refs), via checkBudget from lib/context-bill.ts. Ceilings live in
test/fixtures/context-budget.json with x1.05/x1.10 headroom; regenerate with
bun test/helpers/capture-context-budget.ts. New skills fail until consciously
budgeted; removed skills fail until the fixture is refreshed; reductions
ratchet the ceilings down so wins lock in.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs(todos): file output-template carve wave + plan-ceo doctrine revisit; mark preamble-carve P3 in flight

Two follow-ups deferred from the approved token-reduction program (CEO review
'NOT in scope' list), filed with full context per TODOS format. The existing
P3 preamble-carve entry gets a status update pointing at the program that
supersedes it.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(test): review findings — Windows path normalization, full totals rebuild, ratchet coverage

Pre-landing review (5 specialists) found one critical: the ratchet test runs
in the curated Windows lane, where path.relative yields backslash skill names
that miss the test/ filter and mismatch every POSIX fixture key. Names are now
normalized once in buildRatchetBill (toPosixName) and the fixture filter is
tightened to test/fixtures/. All eight Bill.totals fields are rebuilt from the
filtered list (no fixture-polluted perInvocation/totalMd numbers for future
consumers). New coverage: Windows-separator normalization pins, a
captureContextBudget round-trip against tree-a (headroom math exact), a
stripFields regression pin (interactive/benefits-from absent from renders,
hooks/gbrain preserved), and the ceilings test no longer double-reports
stale-fixture entries.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(test): adversarial findings — stable root key, symlink-alias dedupe, fixture-shape guard

Adversarial review (Claude subagent) verified the fixture's root-skill key was
the capture machine's checkout dirname: any non-gstack-named clone (every
Conductor worktree) failed the free suite, and the documented re-run-the-capture
recovery baked the local dirname into the committed fixture — silent corruption
through the tool's own protocol. The root skill is now pinned to ROOT_SKILL_KEY
('gstack', its frontmatter name). Symlink aliases are realpath-deduped (census
precedent): connect-chrome no longer gets its own ceiling, so Windows checkouts
that materialize the symlink as a plain file can't fail the stale-ceiling
set-equality test. New guards: fixture-shape validation (a string alwaysOnTotal
can no longer silently disable the ceiling), a mutation pin that the filter
shrinks the always-on ledger vs the raw bill, an alwaysOnTotal violation test
(the branch was load-bearing with only under-budget coverage), and an atomic
temp+rename fixture write. Fixture regenerated: 59 ceilings, alwaysOnTotal 6344.
Deferred with a TODO: anchoring transformFrontmatter's denylist strip to the
frontmatter block (latent, zero live collisions, pre-existing path).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* chore: bump version and changelog (v1.69.1.0)

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs: update project documentation for v1.69.1.0

CLAUDE.md: Token ceiling section documents the context-budget ratchet as
the third guard (test file, fixture, new-skill budgeting, capture command).
CONTRIBUTING.md: Tier 1 guard list gains a Context-budget ratchet bullet;
the Adding-a-new-skill checklist gains the budget-capture step.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs: pin exact guard semantics for the context-budget ratchet in CLAUDE.md

Doc-review finding: "a third enforced ceiling" undercounted the guard
family (skill-size-budget floors and parity ratios also watch these
ledgers, relatively). Rephrased to match the ratchet test's own header:
absolute ceilings vs relative floors/ratios.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs(changelog): heaviest-skill claim matches the fixture (land-and-deploy edges review by 0.2%)

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(bin): gstack-skill-start + gstack-skill-end — the preamble runtime, consolidated

Absorbs the ~13KB of bash every tier-2+ SKILL.md inlined twice over (bootstrap
fence + artifacts-sync fence) and the skill-end telemetry/sync fences. Same
KEY: value STATUS-line contract the prose interprets, plus SKILL_START_PROTO
handshake (OV5), SESSION_ID/TEL_START echoes, GSTACK_HOME-normalized state
paths (EOV7), --parent-pid session identity (EOV5: $PPID inside the script is
the ephemeral tool-call shell), OV4 sanitization of passthrough output, and a
receipted daily artifacts pull (_receipted_git, brain-sync class, fail-closed).
Per-line || true error style throughout (F3) — a mid-script failure never drops
later STATUS lines.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(gen): preamble resolvers emit a script invocation fence instead of inline bash

generate-preamble-bash: ~6.3KB fence -> 4-line gstack-skill-start invocation
(quoted-tilde pitfall handled: leading ~ interpolates through $HOME; env-var
hosts keep $GSTACK_BIN) + degraded-mode prose (F1/EOV8: safe defaults, consent
gates deferred-never-lost; OV5: proto rule). generate-brain-sync-block: ~6.8KB
bash -> interpretation prose + the privacy stop-gate (stays inline until
Phase 2's gated emission). generate-completion-status: telemetry fence -> one
gstack-skill-end call with SESSION_ID/TEL_START handoff.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* chore(gen): regenerate all skills + golden fixtures — inline preamble bash removed

Mechanical regen after the resolver change: −12,628 lines across 52 renders
(corpus 952K -> 806K render tokens; tier-2 skills −11-13KB each). Golden
per-host ship fixtures refreshed from the fresh claude/codex/factory renders.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* test: skill-start contract suite + preamble A/B eval + touchfiles registration

test/gstack-skill-start.test.ts (11 free tests): STATUS-key contract vs the
prose (F2), per-host fence resolution shapes (E1), proto-first, OV4 marker
sanitization, --parent-pid identity, headless suppression, skill-end duration
math + pending cleanup. test/skill-e2e-preamble-script-ab.test.ts (gate tier,
OV7): inline-bash render (pinned from 29785978) vs script render with the
fence redirected at the worktree bin (EOV2 — hermetic evals otherwise resolve
the operator install and silently exercise degraded mode). 21 touchfiles dep
lists gain the two bin scripts (EOV9) so future script edits select the
preamble evals; selection-count pin updated 23->24.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* test: repin ~70 assertions to the script contract — every literal gets a successor

Assertions that pinned inline-bash internals (update-check guard, _SESSIONS
reaping, telemetry start/end blocks, routing probe, repo-strip producer,
first-task gating, EXPLAIN_LEVEL/QUESTION_TUNING echoes, #2499 jq scope
resolution, Issue-8 CONDUCTOR gate) now pin the same invariants in their new
home: bin/gstack-skill-start / bin/gstack-skill-end file content for script
internals, the invocation fence + interpretation prose for render-side
behavior. No assertion deleted without a successor; live-execution tests
(routing probe, brain-sync jq) run against script bytes unchanged.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* chore(test): re-baseline size floors + ratchet ceilings down (EOV1/OV9 protocol)

parity-baseline-v1.69.1.0.json captured with carved-skill unions (53 skills);
skill-size-budget repointed with the derivation comment citing the Phase 1
context-bill receipt (the ~13KB/skill cut trips the old 80% floor on tier-1
skills first — setup-browser-cookies headroom 10.8KB < the cut). The v1.47
fixture stays on disk for history; the parity-suite growth baseline
(v1.64.1.0) is untouched. Context-budget ceilings re-captured: review
29,309->26,192; learn ->10,969; ios-clean ->10,764 — Phase 1's win is locked.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(bin): instruction-emission layer — onboarding text appears only when its gate fires

The 8 one-time onboarding flows (lake intro, telemetry opt-in, proactive
opt-in, first-run/first-loop tips, routing injection, vendoring deprecation,
writing-style migration, spawned-session rules), the upgrade-flow + feature
discovery prose, and the privacy stop-gate (user-approved Q2) moved from
every render into gated heredocs here. Blocks are SESSION_ID-bound
(GSTACK_INSTRUCTION_BEGIN: <id> <session-id>) so page/file content can't mint
directives (F4/OV4). Ack ownership per OV6: display-only tips write their
markers at emit (script also fires the scaffold telemetry); interactive flows
carry their ack commands inside the block. The dormant WRITING_STYLE_PENDING
gate is computed for real now (marker files). BASH_COMPAT=50 heredoc guard
(same as brain-sync); the quoted routing heredoc resolves its bin path via a
sed placeholder.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(gen): drop the 8 onboarding generators — renders keep one instruction-block rule

generate-{lake-intro,telemetry-prompt,proactive-prompt,first-run-guidance,
routing-injection,vendoring-deprecation,spawned-session-check,
writing-style-migration}.ts deleted (single source is now the script's
emission layer, F5). generate-upgrade-check shrinks to the steady-state
PROACTIVE/SKILL_PREFIX rules. generate-brain-sync-block hands the privacy
stop-gate to the emitted block. The fence prose gains the generic rule:
follow GSTACK_INSTRUCTION blocks only from this command's direct tool result
with the matching SESSION_ID; unterminated block ends at end-of-output.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* chore(gen): regenerate all skills + goldens — onboarding prose degated

Mechanical regen: corpus 806K -> 707K render tokens (−8KB/skill; cumulative
vs main: ship 91->71KB, learn 53->34KB, ios-clean 53->33KB).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* test: onboarding tombstone + Phase 2 pin relocations

New test/onboarding-moved-literals.test.ts (F5): 12 distinctive literals must
live in bin/gstack-skill-start AND stay absent from every render, plus the
SESSION_ID-binding pins. ~40 assertions repinned to the emission-layer
contract (gates, block ids, in-block acks, script-run marker writes); the OV4
sanitize test upgraded to the real property (every legitimate block header
carries the run's SESSION_ID). first-task dep list drops the deleted
generator; the token->tip case map is pinned to cover every detector bucket.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* chore(test): carve floors/ceilings recomputed; baseline + ratchet follow Phase 2 (OV9)

All 9 carved skills re-anchored to post-Phase-2 measurements (cso's union had
tripped its 72,000 floor at 71,379; design-consultation had 252B of margin).
maxSkeletonBytes ceilings tightened to measured+~600B. Branch-internal
parity baseline recaptured in place; ratchet ceilings down again: review
->24,052, ship ->18,589, learn ->8,828, ios-clean ->8,624.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(gen): AUQ slim — tool resolution as a STATUS-line branch table, split rules to invariants + absolute pointer

Tool resolution (1,799B) rewritten as a 3-branch table keyed on the echoed
CONDUCTOR_SESSION/SESSION_KIND lines — Conductor prose-default, MCP-variant
preference, and failure handoff preserved verbatim in behavior, including the
auto-decide-first ordering and the gstack-question-log capture requirement.
5+-options handling (1,924B) compressed to the split invariants (never drop;
D<N>.k shape; Include/Defer/Cut/Hold; question_id scheme with the never-ask
refusal) + the full-rule pointer. Both doc pointers now interpolate the
absolute install root (Codex outside-voice #7 convention) instead of the bare
'in the gstack repo'. Failure-fallback, Format, and self-check sections are
byte-identical — all 14 MANDATORY always-loaded pins pass with zero test
edits.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* chore(gen): regenerate all skills + goldens — AUQ slim

Mechanical regen: −1.3KB per tier-2+ skill (ship 69.9KB, learn 32.5KB).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* chore(test): baseline + ratchet follow Phase 3 (OV9); OV8 evaluated — shrink floor stays

Branch-internal baseline recaptured; ratchet ceilings down again. OV8's
floor-retirement question, evaluated as planned after Phase 3: the 80% shrink
floor stays — it uniquely catches accidental body deletion in non-carved
skills BETWEEN ratchet recaptures, and the capture command has amortized the
fixture-refresh cost that motivated retiring it.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(review): carve adversarial, plan-completion, and review-army into sections

The three resolver macros ship already carves as siblings now load on demand
for /review too: skeleton 100.2KB -> 55.0KB (-45%), union 93.4KB. Resolvers
stay the single source of truth (sections wrap the macros). Step 0/1, scope
drift, critical pass, confidence calibration, and fix-first stay always-loaded.
Fixtures and pins follow the moved content (codex-hardening wrapped-sites,
review-army E2E fixture builds skeleton+sections with an empty-fixture guard).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(codex): carve the three mutually exclusive modes into sections

Review/Challenge/Consult mode bodies (34.7KB where at most one ever runs)
load on demand: skeleton 81.0KB -> 55.2KB, union 1.04x the monolith. The mode
dispatch, filesystem boundary, and a new always-loaded 'Synthesis
recommendation (REQUIRED) — all modes' block stay skeleton-side (the AUQ
per-skill pins pass unchanged); the plan-file report + exit gate render after
the last section pointer per the gateAfterStop pattern.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(land-and-deploy): carve first-run validation, readiness gate, and merge/deploy into sections

The once-per-repo dry-run validation, the pre-merge readiness gate, and the
merge + deploy-strategy steps (37.8KB) load on demand: skeleton 91.1KB ->
55.7KB. Step 1.5 keeps its detection bash as the dispatch; the first-run
section's fingerprint-save block gained {{SLUG_EVAL}} so it is self-contained.
Zero content lost (line-coverage checked against HEAD).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(ios): demote the four ios skills to preamble-tier 2 (Phase 5)

They never consume the tier-3 sections (repo-mode ownership, search-before-
building) but do fire AskUserQuestion, which tier >=2 provides — verified by
grep before the plan review. -2.2KB per skill. Render assertions pin the
demotion (tier-3 sections absent, AUQ format present).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* chore(guards): register wave-1 carves; monolith invariants retire; baselines + ratchet follow

CARVE_GUARDS gains review/codex/land-and-deploy (12 carved skills total);
their MONOLITH_INVARIANTS entries retire (invariants now generate from the
registry, cso precedent). Touchfiles: carve-section-loading covers the three
new carves; the codex + land-and-deploy LLM-judge dep lists widen to their
sections. Regen + goldens + branch-internal baseline + ratchet ceilings
recaptured (review 24,052 -> skeleton-based ceiling; union floors hold).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* test(gen-skill-docs): review render pins read the carved union

The review carve's readSkillUnion conversions (same pattern its neighbor
carved-skill pins already use).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(autoplan): carve the four review phases + tasks aggregator into sections

Phase bodies (CEO/Design/Eng/DX consensus flows) and the Implementation Tasks
aggregator load on demand; Design and DX stay separate sections because each
is independently conditional on scope. Skeleton 83.7KB -> 58.7KB (-30%
always-loaded); the 6 decision principles, classification, sequencing, and
explicit skip-condition dispatch stay always-loaded. The chain E2E's
phase-complete markers now live only in sections, so its assertions double as
section-read proof (behavioral: external).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(spec): carve the post-confirmation gate-and-file tail into one section

Phases 1-4 are the turn-1 conversational spine — carving them would force the
Read on the first user message for zero real savings. The mechanical tail
(4.5/4.5a/4.5b redaction gates + Phase 5 filing + TTHW telemetry) fires only
after draft confirmation: a genuine lazy boundary, kept as ONE section so the
gh-issue-create bash can never load without the fail-closed redaction gate
that precedes it. Skeleton 65.4KB -> 50.7KB; all ~85 phase-structure
invariants migrated location-aware plus a new carve-shape suite (56 tests).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(setup-gbrain): carve the branch-exclusive install paths into sections

Brain-init (Paths 1/2/3/4 bodies), engine remediation, transcript gate, and
CLAUDE.md persist load on demand — at most one install route ever runs.
Skeleton 75.3KB -> 57.0KB; the Step 1 detect and Step 2 path dispatch stay
always-loaded. New buildSetupGbrainFixture helper gives the periodic E2Es
extract-don't-copy fixtures with a non-empty guard; the voyage-code-3 gate
counts scan the tmpl union (the third init site lives in engine-remediation).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* chore(guards): register wave-2 carves (15 carved skills); autoplan monolith retires; baselines follow

CARVE_GUARDS gains autoplan (behavioral: external via the chain eval), spec,
and setup-gbrain; autoplan's MONOLITH_INVARIANTS entry retires. Touchfiles:
setup-gbrain periodic dep lists gain the section tmpls + fixture helper; the
stale-brain-refs scan covers setup-gbrain/sections. Regen + goldens + branch
baseline + ratchet recaptured.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(qa): carve QA patterns + health rubric into on-demand sections (68→48KB skeleton)

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(browse): carve full command list + snapshot flags into sections/command-list.md (39→27KB skeleton)

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(retro): absorb inline git/awk metrics into bin/gstack-retro-metrics + carve report format

RETRO_METRICS_PROTO: 1 contract, local git reads only (fetch stays in the
skill prose), degraded path documented in the skeleton.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* test: register wave-3 carves (qa, browse, retro) — guards, touchfiles, pins, baselines

CARVE_GUARDS gains the three entries; qa's monolith invariant retires.
auq-format carve-safety now keys on the skeleton+sections union shipping
the AUQ block (first tier-1 carve: browse never renders it by design).
Baselines: parity v1.69.1.0 at 18 sectioned skills; ratchet recaptured.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(test): drop stale generate-lake-intro import (generator deleted in the emission-layer move)

Sol scope discipline stays pinned via the model overlay + completeness
section; the lake intro is now a single script-emitted blurb.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(office-hours): carve Phase 2A/2B into mode-exclusive sections (81→67KB skeleton)

A session runs exactly one mode, so a builder session never loads the
13KB startup diagnostic. Mode mapping and the vibe-shift upgrade rule
stay in the skeleton.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(design): carve UX doctrine + Pretext patterns into read-on-demand sections

design-html 57→49KB, design-shotgun 53→50KB. Sections wrap
{{UX_PRINCIPLES}} so scripts/resolvers/design.ts stays the source of
truth; the pretext-patterns STOP sits at the top of Step 3 so the read
provably precedes the Write.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* test: register wave-4 carves (office-hours ext, design-html, design-shotgun) — 20 carved skills

Both design entries carry requiredReads + loading-eval scenarios (D3A
condition). office-hours phase sections are mode-exclusive, so only the
always-reached design/handoff section is a deterministic requiredRead.
Baselines and ratchet recaptured.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs: trim CLAUDE.md 66.4→44.9KB — verbatim moves to docs/, pointers stay inline

Moved: browser/sidebar/server internals, CHANGELOG release-summary format
spec, project tree, hermetic-E2E detail, slop-scan reference, OpenClaw
publishing. Kept inline: every hard behavioral rule (dist/ ban, redaction
scan-at-sink, egress receipts, bisect commits, eval detach, CHANGELOG
entry rules), the machine-managed GBrain block (byte-identical), and the
'## Deploying to the active skill' header with gbrain-refresh in range
(pinned by test/gbrain-refresh-install-render.test.ts). No voice rewrites.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(test): seed onboarding markers into the hermetic child GSTACK_HOME

EOV7 made bin/gstack-skill-start honor GSTACK_HOME, so the operator-HOME
seeding in e2e-helpers.ts no longer reaches hermetic children — the
emission layer fired lake-intro/telemetry prompts that burned turns and
stalled PTY tests waiting on an answer (observed: plan-mode-no-op derailed
by the telemetry question). Onboarding-specific tests pin their own
GSTACK_HOME per-test, which merges over this seed.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* test: raise carve-section-loading wall clock to 480s SDK / 540s bun

The heavy full-workflow scenarios satisfy their required section reads
inside 60s but need 300-450s to finish the report on slower sandboxes;
the 300s default read as a loading failure when the carve invariant held
(traces: plan-eng-review read its section at 8s, office-hours all three
at 24s, design-html both at 50s — all timed out mid-report).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(security): harden the skill-start trust boundary — review-army findings

Session ID gains a urandom suffix (block binding unforgeable by reflected
content); _sanitize also neutralizes spoofed SESSION_ID: lines; branch
names are charset-clamped before JSON embedding (skill-start + skill-end);
.brain-last-push reads first line only with a charset clamp; the artifacts
URL echo routes through _sanitize; the privacy consent gate fires in
interactive sessions only (spawned auto-choose could accept consent no
human gave — emission order is not a safety property); the daily pull gets
non-interactive + slow-network git guards and stamps only when the
receipted path ran; ~/.claude.json gets a grep pre-filter before the jq
parse.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(resolvers): question-log session_id becomes a substitution placeholder + stale-comment sweep

The question-log block bound $_SESSION_ID, a shell variable the
consolidated fence never sets — hook-less hosts logged empty session_id,
breaking /plan-tune per-session grouping. It now uses the same
substitute-from-the-skill-start-echoes contract as the telemetry block.
Also: retired the pre-Phase-2 stop-gate docstring, repointed the
gbrain-local-status cross-reference at the script's inline jq, dropped an
orphaned section comment, documented retro-metrics' suffix-only census.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* chore: regenerate renders for the question-log placeholder; goldens + baselines follow

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* test: hermetic update-check, onboarding gate sequencing, seeding parity

The contract test's child did a live git ls-remote + curl to github.com on
every bun run test (update_check config now gates it off); the headless
test gets a fresh GSTACK_HOME so the suppression is actually exercised; a
new OV6 test drives the script three times to pin ack-at-emit and gate
sequencing; hermetic seeding covers the config-keyed privacy gate; the
EVALS_HERMETIC=0 debug seeding reaches marker parity.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* test(ci): demote the preamble A/B to periodic (OV7) and add it to the periodic matrix

Post-Phase-3 demotion per the plan; the eval needs fetch-depth 0 (it git
shows a pre-Phase-1 sha), which only the periodic workflow provides — and
a static matrix entry so it can't silently never run.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* chore: bump version and changelog (v1.70.0.0)

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs: update project documentation for v1.70.0.0

ARCHITECTURE.md: the preamble section now describes the v1.70 runtime —
the rendered {{PREAMBLE}} block invokes bin/gstack-skill-start and reads
STATUS lines, gstack-skill-end logs telemetry, and one-time onboarding
text arrives as gated GSTACK_INSTRUCTION blocks instead of riding in
every render.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs: doc-review fixes — repair moved-file links, drop unbacked session-count claim

docs/BROWSER_INTERNALS.md: the two ARCHITECTURE.md anchor links broke when
the section moved from repo-root CLAUDE.md into docs/ — now ../ARCHITECTURE.md.
ARCHITECTURE.md: the preamble's session-tracking item claimed an active-session
count and an "ELI16 mode" that no shipped code implements (the count
computation was deleted with the inline preamble); describe the real
touch-and-prune behavior instead.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs(changelog): correct numeric claims against measured counts

50 of 62 installed skills dropped (fixture/alias entries have no preamble);
11 new carves + a deeper office-hours carve = 9→20; test counts match the
files (13 / 11 / 3 / 7).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs: repoint the preamble-runtime version reference after the queue rebump (v1.71.0.0)

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* test(e2e-design): widen the Aesthetic synonym set — vocabulary variance, not a regression

Both attempts in run 33090283032 produced judge-praised DESIGN.md files
phrased as 'design principles'/'design language' without any of the four
original literals; inputs were identical to the prior passing run
32899975845 (design-consultation untouched by the intervening merge).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(test): stage design-consultation's sections/ into the E2E fixture

The skill has been carved since v1.57.0.0 — the DESIGN.md structure
prescription (the AESTHETIC proposal template) lives in
sections/proposal-and-preview.md behind a STOP-read. The fixture only
copied SKILL.md, so the agent improvised structure from the skeleton and
the section-synonym check has been a coin flip since the carve (CI run
33090283032 trace shows 'no sections dir'; the local eval store has the
same failure on 2026-08-25 while that day's CI run passed on lucky
vocabulary).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-27 09:50:31 -07:00

602 lines
24 KiB
TypeScript

import { describe, test, expect } from 'bun:test';
import { spawnSync } from 'child_process';
import * as path from 'path';
import * as fs from 'fs';
import * as os from 'os';
const ROOT = path.resolve(import.meta.dir, '..');
const PROBE = path.join(ROOT, 'bin/gstack-codex-probe');
// Run a bash snippet that sources the probe and evaluates one of its functions.
// Controlled env + optional tempdir for HOME isolation.
function runProbe(opts: {
snippet: string;
env?: Record<string, string | undefined>;
home?: string;
}): { stdout: string; stderr: string; status: number } {
const env: Record<string, string> = {
// Start from a clean env so test-env vars from the parent don't leak in.
PATH: process.env.PATH ?? '',
_TEL: 'off',
};
if (opts.home) env.HOME = opts.home;
// Apply overrides; undefined means "remove".
if (opts.env) {
for (const [k, v] of Object.entries(opts.env)) {
if (v === undefined) {
delete env[k];
} else {
env[k] = v;
}
}
}
const script = `set +e\nsource "${PROBE}"\n${opts.snippet}\n`;
const result = spawnSync('bash', ['-c', script], {
env,
stdio: ['pipe', 'pipe', 'pipe'],
timeout: 5000,
});
return {
stdout: (result.stdout ?? '').toString(),
stderr: (result.stderr ?? '').toString(),
status: result.status ?? -1,
};
}
function tempHome(): string {
return fs.mkdtempSync(path.join(os.tmpdir(), 'gstack-codex-probe-home-'));
}
describe('gstack-codex-probe: auth probe', () => {
test('CODEX_API_KEY set → AUTH_OK', () => {
const home = tempHome();
try {
const r = runProbe({
snippet: '_gstack_codex_auth_probe',
env: { CODEX_API_KEY: 'sk-test' },
home,
});
expect(r.stdout.trim()).toBe('AUTH_OK');
expect(r.status).toBe(0);
} finally {
fs.rmSync(home, { recursive: true, force: true });
}
});
test('OPENAI_API_KEY set → AUTH_OK', () => {
const home = tempHome();
try {
const r = runProbe({
snippet: '_gstack_codex_auth_probe',
env: { OPENAI_API_KEY: 'sk-openai' },
home,
});
expect(r.stdout.trim()).toBe('AUTH_OK');
expect(r.status).toBe(0);
} finally {
fs.rmSync(home, { recursive: true, force: true });
}
});
test('${CODEX_HOME:-~/.codex}/auth.json exists → AUTH_OK', () => {
const home = tempHome();
try {
fs.mkdirSync(path.join(home, '.codex'), { recursive: true });
fs.writeFileSync(path.join(home, '.codex', 'auth.json'), '{}');
const r = runProbe({ snippet: '_gstack_codex_auth_probe', home });
expect(r.stdout.trim()).toBe('AUTH_OK');
expect(r.status).toBe(0);
} finally {
fs.rmSync(home, { recursive: true, force: true });
}
});
test('no env + no file → AUTH_FAILED with exit 1', () => {
const home = tempHome();
try {
const r = runProbe({ snippet: '_gstack_codex_auth_probe', home });
expect(r.stdout.trim()).toBe('AUTH_FAILED');
expect(r.status).toBe(1);
} finally {
fs.rmSync(home, { recursive: true, force: true });
}
});
test('both CODEX_API_KEY and OPENAI_API_KEY set → AUTH_OK', () => {
const home = tempHome();
try {
const r = runProbe({
snippet: '_gstack_codex_auth_probe',
env: { CODEX_API_KEY: 'k1', OPENAI_API_KEY: 'k2' },
home,
});
expect(r.stdout.trim()).toBe('AUTH_OK');
expect(r.status).toBe(0);
} finally {
fs.rmSync(home, { recursive: true, force: true });
}
});
test('empty-string env vars + no file → AUTH_FAILED', () => {
const home = tempHome();
try {
const r = runProbe({
snippet: '_gstack_codex_auth_probe',
env: { CODEX_API_KEY: '', OPENAI_API_KEY: '' },
home,
});
expect(r.stdout.trim()).toBe('AUTH_FAILED');
expect(r.status).toBe(1);
} finally {
fs.rmSync(home, { recursive: true, force: true });
}
});
test('whitespace-only env vars + no file → AUTH_FAILED', () => {
const home = tempHome();
try {
const r = runProbe({
snippet: '_gstack_codex_auth_probe',
env: { CODEX_API_KEY: ' ', OPENAI_API_KEY: '\t\n' },
home,
});
expect(r.stdout.trim()).toBe('AUTH_FAILED');
expect(r.status).toBe(1);
} finally {
fs.rmSync(home, { recursive: true, force: true });
}
});
test('alternate $CODEX_HOME → checks the alternate path', () => {
const home = tempHome();
const altCodex = fs.mkdtempSync(path.join(os.tmpdir(), 'gstack-alt-codex-'));
try {
fs.writeFileSync(path.join(altCodex, 'auth.json'), '{}');
const r = runProbe({
snippet: '_gstack_codex_auth_probe',
env: { CODEX_HOME: altCodex },
home,
});
expect(r.stdout.trim()).toBe('AUTH_OK');
expect(r.status).toBe(0);
} finally {
fs.rmSync(home, { recursive: true, force: true });
fs.rmSync(altCodex, { recursive: true, force: true });
}
});
});
// --- Group 2: Version check -------------------------------------------------
// Stub `codex --version` by putting a fake `codex` executable on PATH.
function tempStubCodex(versionOutput: string, bool_command_fails = false): {
dir: string;
pathEntry: string;
} {
const dir = fs.mkdtempSync(path.join(os.tmpdir(), 'gstack-codex-stub-'));
const bin = path.join(dir, 'codex');
const script = bool_command_fails
? '#!/bin/bash\nexit 1\n'
: `#!/bin/bash\nif [ "$1" = "--version" ]; then printf '%s' ${JSON.stringify(versionOutput)}; fi\n`;
fs.writeFileSync(bin, script);
fs.chmodSync(bin, 0o755);
return { dir, pathEntry: dir };
}
function runVersionCheck(versionOutput: string): string {
const stub = tempStubCodex(versionOutput);
try {
const r = runProbe({
snippet: '_gstack_codex_version_check',
env: { PATH: `${stub.pathEntry}:${process.env.PATH}` },
});
return r.stdout + r.stderr;
} finally {
fs.rmSync(stub.dir, { recursive: true, force: true });
}
}
describe('gstack-codex-probe: version check (anchored regex per Tension I)', () => {
// Matches (should WARN)
test('codex-cli 0.120.0 → WARN', () => {
const out = runVersionCheck('codex-cli 0.120.0\n');
expect(out).toContain('WARN:');
expect(out).toContain('0.120.0');
});
test('codex-cli 0.120.1 → WARN', () => {
const out = runVersionCheck('codex-cli 0.120.1\n');
expect(out).toContain('WARN:');
});
test('codex-cli 0.120.2 → WARN', () => {
const out = runVersionCheck('codex-cli 0.120.2\n');
expect(out).toContain('WARN:');
});
// Does NOT match (should be silent)
test('codex-cli 0.116.0 → OK (no warn)', () => {
const out = runVersionCheck('codex-cli 0.116.0\n');
expect(out).not.toContain('WARN:');
});
test('codex-cli 0.121.0 → OK (no warn)', () => {
const out = runVersionCheck('codex-cli 0.121.0\n');
expect(out).not.toContain('WARN:');
});
test('codex-cli 0.120.10 → OK (anchored regex prevents substring match)', () => {
const out = runVersionCheck('codex-cli 0.120.10\n');
expect(out).not.toContain('WARN:');
});
test('codex-cli 0.120.20 → OK (anchored regex prevents substring match)', () => {
const out = runVersionCheck('codex-cli 0.120.20\n');
expect(out).not.toContain('WARN:');
});
test('codex-cli 0.120.2-beta → WARN (still a bad release family)', () => {
// 0.120.2-beta: regex (^|[^0-9.])0\.120\.(0|1|2)([^0-9.]|$) treats '-' as a
// non-digit/non-dot boundary → matches.
const out = runVersionCheck('codex-cli 0.120.2-beta\n');
expect(out).toContain('WARN:');
});
test('empty output → OK (silent, no crash)', () => {
const out = runVersionCheck('');
expect(out).not.toContain('WARN:');
});
test('v-prefixed and multiline handled', () => {
const out = runVersionCheck('codex-cli v0.116.0\nsome debug line\n');
expect(out).not.toContain('WARN:');
});
});
// --- Group 3: Timeout wrapper + namespace hygiene ---------------------------
describe('gstack-codex-probe: timeout wrapper + namespace hygiene', () => {
test('bin/gstack-codex-probe is syntactically valid bash (bash -n)', () => {
const result = spawnSync('bash', ['-n', PROBE], { timeout: 5000 });
expect(result.status).toBe(0);
});
test('timeout wrapper executes command directly when neither binary present', () => {
// Clear PATH to simulate no timeout/gtimeout. Use only /bin for `echo`.
const r = runProbe({
snippet: `_gstack_codex_timeout_wrapper 5 echo hello_world`,
env: { PATH: '/bin:/usr/bin' }, // these usually lack gtimeout; timeout may exist on linux
});
// Regardless of whether timeout is on this PATH, echo hello_world should succeed.
expect(r.stdout.trim()).toBe('hello_world');
});
test('timeout wrapper resolves gtimeout preferentially when on PATH', () => {
// Create a stub gtimeout that prints a sentinel so we can verify it was chosen.
const dir = fs.mkdtempSync(path.join(os.tmpdir(), 'gstack-gto-stub-'));
try {
const stub = path.join(dir, 'gtimeout');
fs.writeFileSync(stub, '#!/bin/bash\necho gtimeout_chosen_$1\n');
fs.chmodSync(stub, 0o755);
const r = runProbe({
snippet: `_gstack_codex_timeout_wrapper 5 echo nope`,
env: { PATH: `${dir}:/bin:/usr/bin` },
});
expect(r.stdout.trim()).toBe('gtimeout_chosen_5');
} finally {
fs.rmSync(dir, { recursive: true, force: true });
}
});
test('bash-native watchdog kills a hung command at the deadline (exit 124, no timeout binary)', () => {
// Stock macOS ships neither gtimeout nor timeout(1) — the old fallback ran
// the command unwrapped, so a hung `codex exec` blocked the calling
// workflow forever. Force the fallback everywhere (Linux /bin has timeout
// via usrmerge) with a PATH holding ONLY bash and sleep, then prove a
// 30s sleep dies at the 1s deadline with timeout(1)'s exit code. The
// runProbe 5s spawnSync cap doubles as the "actually killed fast" bound.
const dir = fs.mkdtempSync(path.join(os.tmpdir(), 'gstack-watchdog-'));
try {
const which = (tool: string) =>
spawnSync('bash', ['-c', `command -v ${tool}`]).stdout.toString().trim() || `/bin/${tool}`;
fs.symlinkSync(which('bash'), path.join(dir, 'bash'));
fs.symlinkSync(which('sleep'), path.join(dir, 'sleep'));
const r = runProbe({
snippet: `_gstack_codex_timeout_wrapper 1 sleep 30; echo "rc=$?"`,
env: { PATH: dir },
});
expect(r.stdout).toContain('rc=124');
} finally {
fs.rmSync(dir, { recursive: true, force: true });
}
});
test('sourcing probe does NOT set errexit/trap/IFS in caller shell (namespace hygiene)', () => {
// Capture `set -o` output before and after sourcing. Any drift means the
// probe polluted the caller.
const r = runProbe({
snippet: `
BEFORE=$(set -o | sort)
source "${PROBE}" # source again to catch accumulation
AFTER=$(set -o | sort)
if [ "$BEFORE" = "$AFTER" ]; then
echo "CLEAN"
else
echo "POLLUTED"
diff <(echo "$BEFORE") <(echo "$AFTER")
fi
`,
});
expect(r.stdout).toContain('CLEAN');
});
});
// --- Group 4: Telemetry event emission --------------------------------------
describe('gstack-codex-probe: telemetry event emission', () => {
test('_gstack_codex_log_event writes jsonl when _TEL != off', () => {
const home = tempHome();
try {
const r = runProbe({
snippet: `_gstack_codex_log_event "codex_test_event" "42"; cat "$HOME/.gstack/analytics/skill-usage.jsonl"`,
env: { _TEL: 'community' },
home,
});
expect(r.stdout).toContain('"event":"codex_test_event"');
expect(r.stdout).toContain('"duration_s":"42"');
} finally {
fs.rmSync(home, { recursive: true, force: true });
}
});
test('_gstack_codex_log_event skips write when _TEL = off', () => {
const home = tempHome();
try {
runProbe({
snippet: `_gstack_codex_log_event "codex_test_event" "99"`,
env: { _TEL: 'off' },
home,
});
const jsonl = path.join(home, '.gstack/analytics/skill-usage.jsonl');
expect(fs.existsSync(jsonl)).toBe(false);
} finally {
fs.rmSync(home, { recursive: true, force: true });
}
});
test('payload never contains prompt content, env values, or auth tokens (schema check)', () => {
const home = tempHome();
try {
const r = runProbe({
snippet: `_gstack_codex_log_event "codex_test_event" "1"; cat "$HOME/.gstack/analytics/skill-usage.jsonl"`,
env: {
_TEL: 'community',
CODEX_API_KEY: 'SECRET_TOKEN_SHOULD_NOT_LEAK',
OPENAI_API_KEY: 'ANOTHER_SECRET',
},
home,
});
// The emitted JSON payload should ONLY have {skill, event, duration_s, ts}.
// Specifically, it must not contain any env values or auth material.
expect(r.stdout).not.toContain('SECRET_TOKEN_SHOULD_NOT_LEAK');
expect(r.stdout).not.toContain('ANOTHER_SECRET');
// Schema: exactly these keys, in any order.
const parsed = JSON.parse(r.stdout.trim().split('\n').pop() ?? '{}');
expect(Object.keys(parsed).sort()).toEqual(['duration_s', 'event', 'skill', 'ts']);
} finally {
fs.rmSync(home, { recursive: true, force: true });
}
});
});
// ── Step 2A argv guard ─────────────────────────────────────────────────────
// Regression test for #1428: Codex CLI >=0.130.0 rejects passing a quoted
// prompt argument together with `--base <branch>`. Step 2A must never combine
// the two on the same line. Step 2A lives in the carved review-mode section
// (codex/sections/review-mode.md, generated from its .md.tmpl) — asserts
// across both the .tmpl source and the generated section so template drift
// can't silently re-introduce the bug.
describe('codex review-mode section Step 2A: PROMPT + --base mutual exclusion guard', () => {
function extractStep2A(filePath: string): string {
const content = fs.readFileSync(filePath, 'utf-8');
const startIdx = content.indexOf('## Step 2A: Review Mode');
expect(startIdx).toBeGreaterThan(-1);
// End at next `## ` heading (skill section boundary).
const tail = content.slice(startIdx);
const nextHeading = tail.slice(2).search(/\n## /);
const section = nextHeading === -1 ? tail : tail.slice(0, nextHeading + 2);
// Non-empty extraction: a carve/regen that leaves only the heading behind
// must fail here, not silently pass a vacuous scan.
expect(section.length).toBeGreaterThan(1000);
return section;
}
for (const relPath of ['codex/sections/review-mode.md.tmpl', 'codex/sections/review-mode.md']) {
test(`${relPath}: no \`codex review\` line combines a quoted prompt argument with --base`, () => {
const section = extractStep2A(path.join(ROOT, relPath));
// Find all lines invoking `codex review` (any prefix wrapper allowed).
const lines = section.split('\n');
const offendingLines: string[] = [];
for (const line of lines) {
// Skip prose lines that just discuss codex review. Only inspect lines
// that look like an actual shell invocation (codex review followed by
// a non-prose token).
const match = line.match(/\bcodex\s+review\b(.*)$/);
if (!match) continue;
const rest = match[1];
// Two regression patterns:
// codex review "..." --base <foo>
// codex review $VAR --base <foo>
// codex review -- "..." --base <foo>
// Acceptable: codex review --base <foo> (bare, no prompt arg)
const hasBase = /--base\b/.test(rest);
if (!hasBase) continue;
// Strip --base <token> and any trailing -c/--enable flags so they
// don't look like positional args. Anything that remains BEFORE
// --base and looks like a positional is the regression.
const beforeBase = rest.split(/--base\b/)[0].trim();
// Empty (or just whitespace) before --base => bare review, safe.
if (beforeBase === '') continue;
// Allow `--` separator that introduces nothing else (rare). Anything
// that looks like a quoted string OR variable expansion is the bug.
if (/^["'$]|^--\s*["']/.test(beforeBase)) {
offendingLines.push(line);
}
}
expect(offendingLines).toEqual([]);
});
test(`${relPath}: Step 2A still contains at least one fix-path invocation`, () => {
const section = extractStep2A(path.join(ROOT, relPath));
// At least one of: bare `codex review --base` OR `codex exec ...` must
// remain. Guards against accidental deletion of both fix paths.
const bareReview = /codex\s+review\s+--base\b/.test(section);
const execRoute = /codex\s+exec\b/.test(section);
expect(bareReview || execRoute).toBe(true);
});
}
});
// Regression guard for #1036. The wrapper added in #1056 was wired into
// codex/SKILL.md but not into the /review and /ship diff passes, which kept
// running under a bare 5-minute Bash gate. Measured on codex-cli 0.145.0: a
// pass was killed at 287s of a 300s budget mid-tool-call, and the same prompt
// completed in 336s. An unwrapped stall returns no exit code and no output,
// which downstream reads as "Codex reviewed and found nothing".
describe('codex timeout wrapper: /review + /ship diff passes', () => {
const WRAPPED_SITES = [
'scripts/resolvers/review.ts', // generator (source of truth)
'review/sections/adversarial.md', // review section (Step 5.7 carved out of the skeleton)
'ship/sections/adversarial.md', // ship section source
];
// Outer Bash gate for the wrapped passes. The wrapper must be strictly
// shorter so IT fires first and the failure is a diagnosable exit 124.
const BASH_GATE_MS = 600000;
for (const relPath of WRAPPED_SITES) {
const read = () => fs.readFileSync(path.join(ROOT, relPath), 'utf8');
test(`${relPath}: both diff-review Codex calls run under the wrapper`, () => {
const wrapped =
read().match(/_gstack_codex_timeout_wrapper\s+\d+\s+codex\s+(exec|review)\b/g) ?? [];
// Adversarial pass + structured review pass.
expect(wrapped.length).toBeGreaterThanOrEqual(2);
});
test(`${relPath}: does not claim \`timeout\` is unavailable on macOS`, () => {
// _gstack_codex_timeout_wrapper resolves gtimeout -> timeout -> unwrapped,
// so the coreutils-less case is already handled. The old claim is what
// steered these call sites away from the wrapper in the first place.
expect(read()).not.toMatch(/doesn't exist on macOS/);
});
test(`${relPath}: wrapper budget stays under the outer Bash gate`, () => {
const budgets = [...read().matchAll(/_gstack_codex_timeout_wrapper\s+(\d+)\s+codex\b/g)].map(
(m) => Number(m[1]) * 1000,
);
expect(budgets.length).toBeGreaterThan(0);
for (const ms of budgets) {
// Inverting this makes the wrapper unreachable: the harness kills the
// call first and the exit-124 branch below it becomes dead code.
expect(ms).toBeLessThan(BASH_GATE_MS);
}
});
}
});
// Regression guards for #2496 / #2524 / #2477 — three "guard reports success
// while doing nothing" defects in codex/SKILL.md:
// (a) the default `codex review` path set NO sandbox override, inheriting
// whatever ~/.codex/config.toml grants (write access on trusted
// projects) while the skill's Important Rules claimed read-only;
// (b) the severity-tag verdict gate could not fail on the default path — a
// non-zero exit, empty output, or untagged output all satisfied the
// "no [P1] found → PASS" branch as written;
// (c) Step 2A's Bash tool gate (300000 ms) sat BELOW the 330s wrapper
// budget, so the harness killed the call before the wrapper could emit
// its diagnosable exit-124 message — the same inversion #1036 fixed for
// /review and /ship.
// The three mode bodies are carved into codex/sections/*-mode.md (T9), so the
// sweep reads the skeleton+sections UNION on both the .tmpl side and the
// generated side — a regen or hand-edit of one but not the other can't
// silently reopen any of them. Each mode section starts with its own `## `
// heading, so the per-`## `-section split in check (c) still isolates each
// mode's gate/wrapper pair.
function readCodexUnion(kind: 'tmpl' | 'rendered'): string {
const sectionsDir = path.join(ROOT, 'codex', 'sections');
const skeleton = fs.readFileSync(
path.join(ROOT, 'codex', kind === 'tmpl' ? 'SKILL.md.tmpl' : 'SKILL.md'),
'utf-8',
);
const suffix = kind === 'tmpl' ? '.md.tmpl' : '.md';
const sections = fs.readdirSync(sectionsDir).sort()
.filter((f) => (kind === 'tmpl' ? f.endsWith('.md.tmpl') : f.endsWith('.md') && !f.endsWith('.md.tmpl')))
.map((f) => fs.readFileSync(path.join(sectionsDir, f), 'utf-8'));
expect(sections.length, `codex sections (*${suffix}) missing`).toBeGreaterThanOrEqual(3);
return [skeleton, ...sections].join('\n');
}
describe('codex skeleton+sections union: review sandbox + fail-closed gate + timeout ordering', () => {
for (const relPath of ['codex tmpl union', 'codex rendered union'] as const) {
const read = () => readCodexUnion(relPath === 'codex tmpl union' ? 'tmpl' : 'rendered');
test(`${relPath}: (a) every scoped codex review invocation pins sandbox_mode="read-only"`, () => {
const invocations = read()
.split('\n')
.filter((l) => /_gstack_codex_timeout_wrapper\s+\d+\s+codex\s+review\b/.test(l));
expect(invocations.length).toBeGreaterThanOrEqual(1);
for (const line of invocations) {
expect(line).toContain('sandbox_mode="read-only"');
// `codex review` has no -s/--sandbox flag (verified 0.147.0) — the
// config override is the only lever. `-s read-only` here would fail
// at argv parsing, which check (b) would then read as a gate FAIL.
expect(line).not.toMatch(/\s-s\s+read-only\b/);
}
});
test(`${relPath}: (b) the verdict gate fails closed — no default-PASS path`, () => {
const content = read();
// The old rule inferred PASS from the absence of a substring:
expect(content).not.toContain(
'If no `[P1]` markers are found (only `[P2]` or no findings) — the gate is **PASS**',
);
// The new rule: FAIL on non-zero exit, empty output, and untagged
// output; [P0] recognized as blocking; PASS reachable only through the
// explicit tagged-advisory-only branch.
expect(content).toContain('The gate FAILS CLOSED');
expect(content).toContain('`_CODEX_EXIT` is non-zero (including 124) → **GATE: FAIL**');
expect(content).toContain('empty or whitespace-only → **GATE: FAIL**');
expect(content).toContain('untagged output');
expect(content).toContain('`[P0]`');
expect(content).toContain('PASS is only reachable through check 5');
});
test(`${relPath}: (c) every Bash gate sits strictly above its section's wrapper budgets`, () => {
// Split on `## ` headings; within any section that declares BOTH a Bash
// tool gate (`timeout: N` in ms) and a wrapper budget
// (`_gstack_codex_timeout_wrapper S codex`), every gate must be strictly
// greater than every wrapper budget so the wrapper fires first.
const sections = read().split(/\n## /);
const inspected: string[] = [];
for (const section of sections) {
const gates = [...section.matchAll(/timeout:\s*(\d{4,})/g)].map((m) => Number(m[1]));
const wrappers = [...section.matchAll(/_gstack_codex_timeout_wrapper\s+(\d+)\s+codex\b/g)].map(
(m) => Number(m[1]) * 1000,
);
if (gates.length === 0 || wrappers.length === 0) continue;
inspected.push(section.split('\n')[0]);
for (const gate of gates) {
for (const wrapper of wrappers) {
expect(gate).toBeGreaterThan(wrapper);
}
}
}
// Review (2A), Challenge (2B), and Consult (2C) must all have been
// inspected — each declares both numbers. If a refactor drops either
// number from a section, this count catches the silent skip.
expect(inspected.length).toBeGreaterThanOrEqual(3);
});
}
});