mirror of
https://github.com/garrytan/gstack.git
synced 2026-09-09 14:38:59 +02:00
canberra
4
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
e76f65a8da |
v1.77.0.0 feat: test-infrastructure overhaul wave 1 — matrix deletion, flake telemetry, sync-spawn wedge class extinct (#2746)
* fix: pin the claude CLI to an exact version in the CI image + tripwire The image installed @anthropic-ai/claude-code UNPINNED and rebuilt weekly 'to pick up CLI updates' — while bun sat carefully pinned at 1.3.13 two RUN lines above. The PTY harness screen-scrapes this CLI's TUI, and that drift broke it three separate times (welcome-screen wedge on 2.1.233, skillify HOME discovery on 2.1.237, guard/freeze hooks on 2.1.162), each debugged as a flake first. Pin 2.1.251 (current latest), bump deliberately via a PR that runs the PTY gate, and enforce with test/ci-image-cli-pin.test.ts: any global npm install in Dockerfile.ci without an exact @X.Y.Z pin fails the free suite. The weekly ci-image cron stays as a cheap tag self-heal. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat: stamp the claude CLI version into every eval-store run record Three harness breakages were traced to claude-CLI TUI drift only after long flake hunts, because no run record said which CLI it actually exercised. EvalCollector now stamps claude_cli_version (claude --version, cached once per process, 'unknown' when the binary is absent) into both partial and finalized records — schema-additive optional field, no SCHEMA_VERSION bump. Correlating a flake wave with a CLI release becomes a grep over ~/.gstack/projects/<slug>/evals/ instead of archaeology. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix: give the spinning-shard kill test load headroom (30s -> 90s) The test spawns and group-kills three real children (one a busy-loop burning a full core) while five sibling shard processes compete for eight vCPUs. Under full-suite load it blew bun's default 30s per-test ceiling at 30,009ms — while passing in isolation in 1.4s — and red the only required lane. Every assertion in it is event-based (statuses, group-kill proof, heartbeat lines); the sole latency claim is the <30s kill-deadline sanity bound, which stays. Explicit 90s headroom, not a weakened oracle. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat: green-by-skip census — skip counts in the classifier, all-skipped labeling in the paid runner bun's 'Ran N tests' line COUNTS skipped tests, so a codex/gemini shard whose every test self-skipped (binary absent on the runner — true of every CI runner today) exits 0, dodges the hollow-shard guard, and reads as coverage in the weekly census. The classifier now parses bun's ' N skip' / ' N pass' recap lines; ShardOutcome carries skippedTests; formatSummary and the fail-closed slices report label an all-skipped pass explicitly: 'all N tests SKIPPED — verified nothing'. Status stays 'passed' (external service availability is host state, not a repo regression) but the census can no longer mistake absence for coverage. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * refactor: extract composite actions for eval-lane setup; surviving lanes gain the fail-fast registry verification 'Fix bun temp' x3, 'Restore deps' x5, 'Seed claude interactive config' x3, and 'Register gstack skills' x3 were byte-near-identical copies across the legacy matrix, the sliced lane, and the periodic lane — and only the MATRIX copy of register-skills carried the 19-line dangling-symlink + frontmatter fail-fast loop written after a silent 'Unknown command' + 35-min-timeout incident. Extract all four into .github/actions/ composites; the register composite carries the verification loop (generalized over the skill list), so the sliced and periodic lanes — the lanes that SURVIVE the matrix deletion — now inherit the check they had silently dropped. Matrix-job inline copies are left untouched: that job is deleted next. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat: delete the legacy 17-row eval matrix — the sliced lane is the only paid lane Every PR paid twice: the hand-enumerated matrix (18 test files, 22.6 min, ~$21 API measured on run 33263204465) ran serialized AHEAD of the strictly superior sliced lane via 'needs: evals' — 35.5 min wall and ~2x paid spend for the same diff. 14 of 17 rows carried no tier:, so periodic Opus benchmarks leaked into every PR (the e2e-plan row alone: 12/12 tests, 21.7 min, $7.28 — the wall-clock bound of ALL of CI). Parity receipt (static, pre-deletion): the sliced lane's gate census (49 files, derived from the runner itself) strictly contains all 18 matrix test files, plus 31 files the matrix never ran. Pure deletion — one revert restores it. The PR comment moved into slices-report (same '## E2E Evals' upsert marker, now sourced from slice artifacts + carrying the fail-closed reconciliation verdict). plan-slices loses the needs edge; the dead workflow-level EVALS_TIER env goes with it. test/evals-workflow-matrix.test.ts (and its KNOWN_MATRIX_GAPS / KNOWN_TIER_UNSET burn-down ratchets — retired: the sliced census makes 'every gate file runs' true by construction) is rewritten as test/evals-workflow-wiring.test.ts: matrix stays deleted, planner/executor/ report tier + slice-count agreement, both surviving lanes on the shared register-skills composite with its fail-fast verification loop, PR comment survival. Expected: PR eval wall 35.5 -> ~13 min, per-PR paid spend ~halved. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix: provider-runner timeouts kill the whole process GROUP; codex/gemini inherit the orphan-drain hardening All three provider runners (claude/codex/gemini) killed only the direct child on timeout: tool subprocesses the CLI spawned survived as orphans holding our pipes open and burning shared API rate (observed: a 600s timeout stretching past 1400s; a stalled run once burned a core for 15 hours). gstack-detach's watchdog had the same shape one level up — killpg SIGTERM, 5s grace, then a direct-child proc.kill() that orphaned grandchildren. Fix: spawn provider children via node:child_process with detached (own process group) and killProcessGroup(SIGKILL) in the timeout handler — runShardChild's proven pattern, EPERM/ESRCH fallbacks included. The codex and gemini copies also gain the reader.cancel() + stderr Promise.race hardening only the claude copy had (they still carried the blocked-drain hang it fixed). gstack-detach's watchdog now group-SIGKILLs after the grace. Regression net: test/session-runner-groupkill.test.ts drives the REAL runSkillTest against a fake claude shim (PATH override) that spawns a grandchild and wedges — the run must classify timeout within budget and leave neither shim nor grandchild alive — plus source pins on all three runners (detached + killProcessGroup, no bare timeout kill, no Bun.spawn reversion). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix: skill-e2e-opus-47 renders SKILL.md fixtures into a mkdtemp — never the live tree mkEvalRoot ran gen-skill-docs with cwd=ROOT, regenerating every in-repo SKILL.md mid-run while concurrent paid shards copyFileSync those same files in their beforeAll (EVALS_JOBS>=4 locally, 2 per CI slice) — a sibling could capture a half-regenerated or opus-rendered SKILL.md, and a timeout before afterAll stranded the whole tree at the wrong model for every later shard. A cross-shard race that could flake ANY concurrent paid test. Render via the --out-dir flag gen-skill-docs grew for exactly this reason (mirrors the repo layout, which is all the fixture reads), read the skill heads from the render dir, delete it, and drop the afterAll restore-regen entirely. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix: claude CLI version resolves in the runner parent, never on a test thread Eng-review finding: getClaudeCliVersion's fallback is a SYNCHRONOUS spawnSync on the same thread that polls concurrent PTY/session tests — the judgePtyState blocking class this overhaul kills elsewhere. The paid runner parent now resolves it once (cached) and stamps GSTACK_CLAUDE_CLI_VERSION into every shard's env; eval-store short-circuits on the env var, and the fallback spawn's budget tightens 10s -> 3s (bounded one-time stall, records 'unknown' on a slow CLI). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * test: wire skippedTests end-to-end through runPaidShard The census unit tests hand-built outcomes and the classifier tests parsed strings; nothing proved a real child's ' N skip' recap flows into outcome.skippedTests and the formatSummary label. A commandFor fake now prints the recap shape and the test asserts the parsed counts, the all-skipped predicate, and the 'verified nothing' label. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix: make the setup composites rerun-safe (codex diff-review hardenings) restore-deps: 'cp -r SRC node_modules' with an existing node_modules NESTS the copy and leaves stale deps active — rm first. register-gstack-skills: 'ln -snf' hard-errors under set -eu when a REAL directory occupies the gstack slot — clear a non-symlink leftover first. CI workspaces are fresh today; a reusable composite must survive dirty reruns. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix: sweep — every sync spawn in the test trees carries a timeout (436 sites, 157 files) spawnSync/execSync/Bun.spawnSync BLOCK the main thread, so bun's in-process per-test timeout can never fire while one waits — a hung child (stdin read, network probe, dead daemon) wedges the whole shard until the runner's external wall-clock SIGKILL. This exact class reached main: free-tests run 33262077256, test/gstack-memory-ingest.test.ts (normally 2.3s) held shard 2 at the 360s wall while its five siblings finished in ~65s. Mechanical sweep in two waves (12 + 4 fan-out agents, every edit verified against its call site): default timeout: 30_000 (matches the free runner's per-test budget), 120_000 for genuinely slow ops (installs, builds, playwright, provider CLIs), helper wrappers fixed ONCE where call sites route through them. Sites that only LOOK like calls (string fixtures, grep needles, comments) were skipped with reasons — the enforcement commit that follows marks them exempt. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * test: sync-spawn timeout tripwire — the wedge class stays extinct Free scanner over all test trees (test/, browse/test/, design/test/, make-pdf/test/, ios-qa, browser-skills): every spawnSync/execSync/ Bun.spawnSync call site must carry a timeout within a 30-line options window, or an explicit '// tripwire-exempt: <reason>' marker. Comment lines are skipped; exemptions are counted and ratcheted shrink-only (ceiling 6 = the 6 string-fixture/grep-needle sites where the pattern is CONTENT, not a call — marked in this commit). A scan-sanity test pins that the scanner still sees >100 real call sites so it can never rot to a vacuous green. Companion to the 436-site sweep in the previous commit. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat: paid-lane flake telemetry — record-level attempts, flaky_retries, report surfacing bun --retry leaves a retried pass INVISIBLE in its output: a fail-then-pass prints the error detail but no (fail) result line and recaps as a clean pass (probed live on 1.3.10). So attempts are recorded where they cannot lie: EvalCollector.addTest stamps a 1-based attempt on same-name re-records (a retried test runs its body again and re-records), finalized runs carry flaky_retries, printSummary warns loudly, and the fail-closed slices report lists every passed-only-on-retry test — recorded and ranked, never blocking and never silent. Cross-model confirmed (codex reached the same don't-parse -the-stream conclusion independently). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat: free-lane flake ledger — retry ON in CI, flaky-passes recorded and uploaded The runner's attribution-gated flaky-retry pass (cap 5, truncation veto) was OFF in the required lane and its FLAKY-PASS evidence was console-only — so a single timing flake red the merge gate while repeat offenders stayed unenumerable. free-tests.yml now sets GSTACK_FREE_RETRY_FLAKY=1 and points GSTACK_FLAKE_LEDGER at runner.temp; every flaky-pass appends a JSONL entry (SINGLE writer: the parent runner — no concurrent-append hazard by construction; fail-open with a loud warning so a broken ledger can never red the lane) and the artifact uploads UNCONDITIONALLY — a flaky-pass run is green, which is exactly when the evidence matters. Wiring pinned by free-tests-workflow-wiring; ledger behavior unit-tested incl. the fail-open path. Matches 2026 industry practice (retry for data, quarantine out of merge-blocking but never out of logging) with the repo's own receipts. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat: eval:flake-rank — the flake-telemetry dial Aggregates per-test series across every finalized eval-store run (shard dirs included) plus the free flake ledger: runs, fails, RETRIED PASSES (the flake signature), avg duration — ranked retries-first. This is the readable dial behind two policies: a flaky pass never blocks a merge but is always ranked here, and the WS16 required-check promotion needs weeks of clean flake-rank, not vibes. --json for machines, --dir for downloaded CI artifacts. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat: two-phase session timeout — silent APIs die at the startup grace, named The single spawn-armed timer charged API queue latency to the work budget: the recurring '0 turns / $0.00 / x3 attempts' failure with four budget-bump receipts (180->300s, 240->360s, 300->420s, 90->300s). Split: startup phase (no NDJSON byte yet) kills EARLY at min(grace, timeout) with the distinct exitReason 'timeout_startup' — an availability verdict, not transcript archaeology — and the work phase arms on the first byte for the REMAINING budget, so total wall never exceeds the timeout (tier envelopes are margin-free: tests pass timeout: CAPTURE_MS and bun-budget the same tier). Local grace 90s (observed queue latency 60-90s), CI floor 300s (TODOS-filed; shared runners queue harder), both pinned by the new grace tests with fake -claude shims covering the late-first-byte and silent-API paths. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix: census integrity — 17 phantom selection keys deleted, reverse invariant added, gitignored dep patterns replaced, local map forks derived The merge-blocking gate census counted tests that could not run. Deleted (critic-verified against both quoted-occurrence and dep-registration liveness): 7 *-prosons-format keys with no declaring test, ship-plan- completion/-verification, review-plan-completion, design-shotgun-path/ session/full, autoplan-core (dead ~10 months), e2e-harness-audit (its namesake is a FREE-suite file), plus 2 dead LLM-judge keys and 2 free-file keys (budget-regression-pty, global-discover) misplaced in the PAID maps. Census: 191 -> 174 keys, gate 86 -> 78 honest. The new reverse invariant in touchfiles.test.ts makes the class structurally impossible: every key must be quoted in a living paid test file OR registered to an existing paid test file via its dep list (the constructed- name binding the 2026-08 self-registration sweep established) — zero exceptions needed today, with a live-file check on any future exception. Also: '.agents/skills/**' dep patterns replaced with the generator (scripts/gen-skill-docs.ts) — .agents/ is gitignored, so those patterns could NEVER match a git diff and review-template edits silently stopped selecting codex/gemini tests; the codex/gemini local touchfile maps are now DERIVED from the canonical map (loud throw if a key vanishes) instead of hand-forked copies that had already drifted. ios-qa-e2e demoted gate -> periodic: its gate declaration was never executable in CI (hardware exclusion only applies at tier=periodic), so every Linux PR planned a hollow shard. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix: routing journeys lose their answer key and end at the routing decision The journey tests exist to catch skill-DESCRIPTION regressions (touchfiles: */SKILL.md.tmpl), but the fixture CLAUDE.md shipped an explicit prompt->skill lookup table — with the answer key in context, a badly regressed frontmatter description still routed correctly, so the tests could not fail on the exact class they select for. The fixture now carries only the generic invoke-skills nudge; the frontmatter carries the routing load. Also capped all 10 journeys at maxTurns 2 / tools [Skill, Read]: only the FIRST Skill call is asserted, so 5 turns of Read/Bash/Glob/Grep was pure spend — roughly halves each journey's cost. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix: retire decided A/B experiments; vendor the pre-cut fixture; ban raw-SHA fixtures Three one-shot decision experiments kept re-running weekly as N=1 stochastic comparisons — flaky by construction with near-zero remaining information: skill-e2e-auq-repetition-cut-ab (its own header: gate "passed pre-landing, approved 2026-08-25"), skill-e2e-preamble-script-ab ("demoted post-Phase-3"), and opus-47's fanout arm-vs-arm (parA >= parB across two SINGLE stochastic runs — a coin flip). Deleted, with their selection keys; the SDK overlay-harness stays as the maintained instrument for the next experiment, and opus-47 keeps its routing-precision cases. verboseSkill() now reads the VENDORED test/fixtures/auq-pre-cut-...-SKILL.md instead of `git show ab66193e^:...` — a branch-local ref that dies on branch prune and already failed on shallow clones. New free tripwire (test/git-ref-fixture-tripwire.test.ts) bans the raw-SHA fixture class outright: quoted SHA:path rev-specs and gitRef-style hex defaults in the test trees fail the suite with the vendor-instead instruction. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix: demote plan-ceo-review-expansion-energy to periodic Opus generator + a subjective 2-axis >=4/5 LLM-judge threshold sat in the MERGE-BLOCKING gate — the exact class its sibling posture tests were demoted for, with a receipt (a +21-line preamble change once flipped the score). CLAUDE.md's own tiering rule: Opus model test -> periodic. The weekly lane keeps the regression signal; merges stop paying a judge- temperament tax. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix: paid shards get per-shard TMPDIR + CHROMIUM_PROFILE isolation and a kill-path cleanup backstop The free runner treats this isolation as MANDATORY (two concurrent shards on one Chromium profile kill each other's browser; shared tmp cross-contaminates) — the paid lane had none of it. Doubly load-bearing here: a shard that hits its 30-min wall is group-SIGKILLed, so per-test afterAll cleanup never runs; the rmSync backstop is the only thing keeping wedged runs from accumulating full git-repo workspaces in the shared tmpdir forever. This is the DAG prerequisite for raising EVALS_JOBS (next commit) — more concurrency on shared state amplifies exactly the shared-tree race class opus-47 exhibited. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat: paid-runner defaults 4x4 -> 8x2 — halve the local gate worst case 39 of 75 skill-e2e files hold exactly ONE test, so within-shard concurrency was dead weight for most shards: 4 jobs x 4 concurrency yielded only ~4-6 real in-flight sessions and a 13-wave local gate worst case (~6.5h). 8 jobs x 2 gives ~10-13 in-flight — under the documented-safe ~15 — and ~7 waves (~3.3h worst case). CI lanes keep their explicit EVALS_JOBS env (2 per slice; 4 for gate-census); this changes local defaults. Rollback trigger: sustained 429 storms in the WS1 telemetry across 2 PR cycles. test/eval-detach-timeout-floor.test.ts recomputed green (the raise LOWERS the worst-case floor). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix: SHA-pin every action in the secrets-bearing eval lanes evals.yml and evals-periodic.yml execute PR-authored code with three provider API keys in env, yet rode mutable action tags (@v7/@v8/@v2/@v4) — while quality-gate.yml, osv-scanner.yml, and dependency-review.yml already model the SHA-pin pattern. All 30 uses sites across both lanes now pin the exact commit (tag noted in a trailing comment); dependabot's github-actions ecosystem keeps them fresh via PRs instead of silent tag moves. Pulled forward from the plan's endgame on the CEO-review + outside- voice agreement: supply-chain pins on secret lanes go first, not last. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix: sweep wave 3 — the execFileSync family gets timeouts (90 sites, 17 files) The tripwire's regex covered spawnSync/execSync/Bun.spawnSync but not execFileSync — an entire blocking sync-spawn API family that could reintroduce the shard-wedge class undetected (ship review army). Same mechanical recipe as waves 1-2: timeout: 30_000 default, 120_000 for slow ops, shared wrappers fixed once, string-needle sites skipped with reasons. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * test: review-army + adversarial test hardening - Tripwire scans execFileSync too (ceiling 8: two more grep-needle string exemptions); merge-introduced timeout-less spawnSync in question-preference-hook fixed — the tripwire caught a site that landed on main AFTER the sweep, on its first day. - gstack-detach gains TWO watchdog kill regression tests: TERM-immune grandchild (the killpg-after-grace escalation) and the leader-dies variant (the pgid-at-spawn fix — the case the first test cannot see). - eval-flake-rank gets its unit suite (final-attempt accounting, artifact exclusion, shard recursion, recency bound). - Groupkill/startup-grace shim markers are per-run unique (pid-suffixed sleep durations): sibling Conductor worktrees run free suites with no machine lock, and fixed markers let one run pgrep/pkill the other's shims — a cross-run flake inside the anti-flake tests. - flake-ledger test pins the project-scoped local default; stale empty section headers in touchfiles-data deleted (they invited entries under deliberately retired categories). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix: adversarial-review runtime fixes across the telemetry + kill paths - session-runner: exit-labeling keys off 'exit', not 'close' — an orphan holding the pipes could relabel a REAL exit (auth failure) as 'timeout_startup' availability noise; the kill path still always group-kills and cancels the reader (labeling and unblocking are separate concerns). Work phase arms on a flag, not firstResponseMs===0 (a same-ms first byte left the startup timer live all run). The CI startup grace is now a real FLOOR (Math.max), matching its name and pinning test. - gstack-detach: pgid captured AT SPAWN (== child pid under start_new_session) — resolving it after the grace raised ESRCH once the leader died on SIGTERM, orphaning TERM-immune grandchildren forever. - test-free-shards: ledger entries carry branch + git_sha (rev-parse split: '--abbrev-ref HEAD HEAD' printed the branch twice and recorded it as the sha); local ledger default is per-PROJECT, not the machine-global tmpdir. - eval-flake-rank: per-LINE ledger parse (one torn JSONL line vanished the whole series), 60-day recency bound (transcript-bearing files are MBs), shared isFinalizedEvalResultFile predicate (the artifact-taxonomy rule lived in three places); eval-store exports the predicate and finalize stops computing flakyRetries twice; paid-shards cleanup uses async rm (a SIGKILLed shard's git-workspace teardown blocked every sibling's stream classification on the parent event loop). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix: CI trust-boundary + fail-closed repairs (adversarial findings) - Token/exec separation restored: slices-report (runs PR-authored code: bun install + the reconcile runner) drops to contents:read; the PR comment moves to a NEW slices-comment job holding the write token with ZERO repo code — no checkout, no bun, only downloaded artifacts + jq/gh. $GITHUB_ENV/BASH_ENV persistence is job-scoped, so the split is the boundary. The matrix-era report job had this property; the consolidation had regressed it. Pinned by the wiring test. - Reconcile exit captured via PIPESTATUS[0] in BOTH lanes: GitHub's default run-step shell has no pipefail, so `$?` after `| tee` was tee's exit — the fail-closed gate was silently fail-open. Wiring test pins it. - PR comment: final-attempt accounting restored the dropped COST accumulation (the dial read $0 forever), flaky passes render as the warning they are (never as failures), and a malformed tests[] artifact skips that file instead of aborting the whole comment under bash -e. - Remaining mutable action tags pinned (free-tests upload-artifact, ci-image checkout/docker trio — the image publisher holds packages:write and feeds the secret-bearing lanes). restore-deps fallback installs --frozen-lockfile; register-gstack-skills validates skill names before its rm -rf. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * chore: bump version and changelog (v1.77.0.0) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs: update project documentation for v1.77.0.0 Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs: cross-model doc-review fixes — flake-ledger env knobs, CI retry-on note, stale version comment Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs: correct CHANGELOG receipt numbers to measured values Gate census keys: 78 -> 77 (bun-imported E2E_TIERS count). Sweep receipt: 586 sites/176 files -> 499 sites/146 files, measured by running this branch's spawnsync-timeout-tripwire against origin/main (exit 1, 499 violations across 146 unique files; green on this branch). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix: slices-comment creates the PR comment via REST — the write-token job has no git context The token/exec split gives slices-comment NO checkout by design, and gh's pr-comment subcommand resolves the repo FROM git — it died with 'not a git repository' on PR #2746's first run (the update-existing PATCH path was already explicit-repo REST and worked). Create now posts through gh api repos/.../issues/N/comments, and the wiring test pins that no git-context-requiring comment call can creep back into the job. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix: startup-grace probes clear CI for local semantics; new probe pins the floor clamp The two shim probes pass explicit 2s/4s graces, but in CI the runner clamps any explicit grace up to the 300s floor (deliberate adversarial-review fix), so 'silent API killed at the grace' died at the 30s work cap instead of 2s — a deterministic red on every CI run, green locally. The probes now pin LOCAL semantics with CI cleared (same save/restore pattern as their PATH shim), and a fourth probe pins the clamp itself: CI=1 + 2s grace + 6s timeout must kill at the 6s cap, still in the startup phase — proof an explicit low grace cannot bypass the floor. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
b1485d8897 |
v1.74.0.0 test/CI overhaul: green means green, suites restructured for speed (#2721)
* fix(ci): free-tests lane actually runs the make-pdf e2e gates The 9 make-pdf/test/e2e gate tests probe make-pdf/dist/pdf, browse/dist/browse, and the diagram-render bundle, then self-skip when absent. The required free-tests lane never built any of them, so the gates silently skipped on Linux for their entire life (verified: 9 of 14 skip, exit 0). make-pdf-gate.yml's justification for deleting its Linux leg claimed the free lane covered this — it didn't. - new build:gates script: exactly the three artifacts the gates probe (full bun run build compiles five binaries; ~60-90s tax on the only required check is not warranted) - free-tests.yml: build:gates step + poppler-utils + fonts-noto-color-emoji (fonts must precede the first browse daemon launch — Chromium snapshots fontconfig at startup; verified live: a warm daemon renders tofu, a fresh one embeds NotoColorEmoji) - make-pdf/test/e2e/ci-prereqs.test.ts: GSTACK_EXPECT_BINARIES=1 (set by the workflow) inverts the skip polarity in CI — dropping the build step or poppler fails the lane instead of re-opening the silent-skip hole Pre-flight: all 9 gates green on Linux locally. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(ci): kill the three zero-test eval jobs (hollow green) - delete the vestigial e2e-codex / e2e-gemini matrix rows: both files are whole-file periodic-tier, so with no row tier: they ran ZERO tests and reported green on every PR (~2 min of runner each, pure false confidence; the periodic lane owns those suites) - e2e-pty-plan-smoke gains tier: gate — its two files are whole-file describeE2ETier('gate'), so the job burned ~7 min of container setup then skipped every describe - KNOWN_TIER_UNSET burned down to empty; the ratchet stays armed so a future row/file tier mismatch fails the suite instead of shipping hollow green Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(ci): least-privilege permissions + fork-safe concurrency keys - evals.yml / evals-periodic.yml evals jobs: explicit contents:read + packages:read (container-image pull) and persist-credentials:false — the jobs that execute PR-authored code with three provider API keys ran on the repo-default token grant with the token written into .git/config - permissions blocks for the 4 workflows that had none (skill-docs, make-pdf-gate, windows-free-tests, windows-setup-e2e) - fork-safe concurrency keys: actionlint, skill-docs, make-pdf-gate, windows-setup-e2e switch from head_ref to PR-number keying — a bare branch name carries no fork prefix, so same-name branches from two forks shared one group and cancelled each other's runs Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(ci): one bun version everywhere + drift tripwire Lanes disagreed four ways: 1.3.13 (free-tests, windows, Dockerfile.ci), latest (quality-gate, make-pdf-gate), unpinned (skill-docs, version-gate — setup-bun installs latest), 1.3.10 (.gitlab-ci.yml). Different Bun versions change the runner output shapes the strict classifiers regex-match, spawn semantics, and shell parsing — a lane on a different Bun tests a different product; Dockerfile.ci's own comment records this class biting once already (silent 1.3.13/1.3.14 drift). All surfaces pinned to 1.3.13; test/bun-version-drift.test.ts scans every workflow setup-bun stanza + Dockerfile.ci + .gitlab-ci.yml and fails on any mismatch or unpinned stanza. skill-docs also gains --frozen-lockfile (was bare bun install). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * test(ci): bind the three-way image-tag hashFiles() expressions evals.yml, evals-periodic.yml, and ci-image.yml each compute the CI image tag from hashFiles('.github/docker/Dockerfile.ci', 'bun.lock', 'patches/**') — synced by comment only (TODOS.md 'CI three-way image-tag drift'). If one input list drifts, that workflow computes a different tag for the same content: eval lanes silently rebuild the image every run, or ci-image prebuilds a tag nobody looks up. The test extracts each tag-computation site and fails on any mismatch. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(ci): ci-image stops rebuilding the identical image every ship - package.json out of the trigger paths: the tag hash deliberately excludes it (version bumps every ship), so every merge rebuilt and re-pushed the IDENTICAL tag (~2m26s for zero content change); patches/** added (it IS a tag input) - manifest existence check (mirrors evals.yml): tag already exists → skip the build - concurrency group: two rapid main pushes raced pushing the same :latest/:buildcache tags - cron staggered 06:00→04:00 Monday: it shared the exact minute with evals-periodic, which could race a half-pushed tag or duplicate the build - timeout-minutes: 30 (was unbounded → 360-min default for a hung docker build) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(ci): quality-gate drops the 74s full-history checkout fetch-depth:0 cost 74 of the job's 92 seconds; the three gates it feeds take ~12s combined. Shallow checkout + exact-SHA fetches for the diff's base/head (an exact-SHA fetch, not a guessed depth — long-lived branches and merge queues still resolve), with a --deepen fallback for push events whose 'before' is unusable. timeout right-sized 20→10 min. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(ci): small-lane batch — timeouts, right-sizing, windows cache warm-start - timeout-minutes on the 6 remaining unbounded jobs (actionlint 5, skill-docs 10, version-gate 10, make-pdf-gate 15, pr-title-sync 5, evals build-image 15) — a hung step sat on GitHub's 360-min default - right-size measured-over-long timeouts: dependency-review 10→5, windows-setup-e2e 15→10 - dependency-review: 2-core runner (28s API call on an 8-core box) and drop .github/workflows/** from its trigger paths (workflow edits have no dependencies to review) - windows caches gain restore-keys: a lockfile bump paid the 26s/43s restore for a guaranteed cold miss Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(test): scope GSTACK_HOME to each file's execution window Five files assigned process.env.GSTACK_HOME at module scope. Shard processes evaluate sibling modules before running their tests, so the assignment leaked into every other file in the shard — the damage was already visible in defensive workarounds (relink.test.ts:28 'fresh install test saw a neighbor's skill_prefix'; cdp-e2e's own comment documents a sibling's temp dir baked into artifacts). Pattern: save original, assign in beforeAll, restore in afterAll (cdp-e2e already restored but still assigned at load — its window now matches the others). GSTACK_TELEMETRY_OFF and GSTACK_PROJECT_SLUG get the same treatment where they rode along. Victim files' defenses stay in place (cheap insurance). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * test: tripwire against module-scope GSTACK_HOME assignments Column-0 assignment of GSTACK_HOME / GSTACK_STATE_ROOT in any tracked *.test.ts fails with the file:line and the fix (beforeAll + afterAll restore). Kills the cross-file env-leak class the previous commit swept. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(test): e2e-harness-audit derives its skill census from disk The hand-maintained 39-name SKILL_GLOBS list had drifted to 39 of 54 SKILL.md.tmpl on disk. No live gap today (none of the 15 unlisted skills is interactive), but the next interactive skill would have landed unguarded with zero signal. The audit now walks top-level dirs for SKILL.md.tmpl (statSync so symlinked dirs like connect-chrome count), so new skills are in scope the commit they appear. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(evals): judges honor the eval-model resolution chain + real 429 backoff callJudge inlined GSTACK_EVAL_MODEL_JUDGE || sonnet, silently ignoring the global GSTACK_EVAL_MODEL override every other eval call site honors via lib/eval-model.ts. New 'judge' kind in DEFAULTS (sonnet — the D1a pin-on-regressors calibration stands; model CHOICE unchanged) and callJudge resolves through it: explicit arg > GSTACK_EVAL_MODEL_JUDGE > GSTACK_EVAL_MODEL > default. 429 handling upgraded from one fixed 1s retry (reliably lost races at CI concurrency) to three jittered exponential retries (~1s/4s/16s), honoring the server's retry-after when present. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(test): the two expect(true) paid stubs become test.todo skill-e2e-spec-execute (600s budget) and skill-llm-eval-spec (300s) reported PASS on every periodic run while asserting nothing. Deleting them would remove the periodic-tier selector surface they exist to register (diff-based selection for spec/ changes), so they become test.todo — reported as todo/skip, never pass — with the v1.1 implementation specs kept in-file. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(test): reactivate 5 quarantined browse tests (2 security) extension-sender-auth's two privileged-message denial tests (content script + missing sender.url — the extension's security boundary) and snapshot's three skips were quarantined 'pre-existing' failures. Root cause: machine-local state on the quarantining dev machines — the test and gate code are byte-identical between the quarantining commit ( |
||
|
|
c118e2402e |
v1.64.1.0 v1.64.1.0: the code-smell fix wave — every pipeline guard now provably fires (net −24,943 lines) (#2572)
* fix(ci): skill-docs freshness gate covers all 10 hosts and can actually fail The Codex/Factory gates ran 'git diff --exit-code -- .agents/' / '-- .factory/', but both paths are gitignored (.gitignore:16-17) — git diff on ignored untracked paths is always empty, so those two gates were structurally incapable of failing and 7 of 10 hosts had no gate at all. New shape: one 'gen:skill-docs --host all' pass (the generator hard-fails on any per-host error, gating all 10 hosts on generates-cleanly), byte-freshness via git diff for tracked output, plus a porcelain check that fails on untracked generated strays (git diff can't see brand-new files). The gitignored-hosts byte-freshness limitation is documented in the workflow comment. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(test): exorcise the sidebar-agent ghost from the test suite browse/src/sidebar-agent.ts was deleted in the v1.14 sidebar refactor, but the test suite kept testing it for 48 versions. Nothing noticed because the free suite runs in no CI job and Bun-era module-load errors were suppressed in the Windows shard runner via an exclusion pattern whose own comment documented the breakage ('broken on every platform since v1.14 ... exit 0'). - Delete sidebar-security.test.ts + security-source-contracts.test.ts: crashed at module load (unguarded readFileSync of the deleted file); per-assertion triage confirmed every SERVER_SRC pin targeted the deleted chat prompt builder (zero hits in today's server.ts) — nothing to port. - Delete sidebar-integration.test.ts: 11 of 13 tests exercised deleted endpoints (/sidebar-command queue, /sidebar-agent/event, chat buffer); the 2 passing tests pinned only the blanket auth gate, covered by server-auth.test.ts + dual-listener.test.ts. - Delete test/skill-e2e-sidebar.test.ts: E2E for the deleted queue flow. - sidebar-ux.test.ts 1,669 -> 830 lines: 20 dead-chat describes + 15 dead tests removed (incl. 10 vacuous passes asserting on empty indexOf slices); 2 stale pins on LIVE features fixed (content.js typed-catch CSSOM fallback, arrow-hint window widened). 95 pass / 0 fail. - sidebar-tabs.test.ts: both failures were stale pins, not regressions — forceRestart's deliberate ws.close(4001) and the terminal-agent spawn that moved into spawnTerminalAgent() (identity-based kill refactor). 28 pass. - touchfiles.ts: drop the three sidebar E2E entries from BOTH maps (E2E_TOUCHFILES + E2E_TIERS) — they pointed diff-selection at the deleted file, so those tests were unreachable by any diff. - test-free-shards.ts: remove the now-dead sidebar-agent exclusion pattern. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat(ci): run the free test suite in CI (it ran nowhere) The full free suite (bun test: browse/test/ + test/ + make-pdf/test/) had no CI job on any Linux/macOS runner — only Windows curated shards, paid evals, and doc-freshness gates existed. That's how two module-load-crashing test files survived 48 versions. Same cached Dockerfile.ci image and container wiring as evals.yml (deps restore, build, Chromium verify). Includes a module-load-error guard: older Bun reported test-file import crashes with exit 0 on macOS/Linux, so the job also fails on any nonzero 'N errors' count in the summary — future crash-class regressions can't hide from the exact job built to catch them. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat(test): validate touchfile dependency paths exist on disk New guard in touchfiles.test.ts: every non-glob dep path must exist, and every glob's anchor directory must exist. This is the axis the 181-key two-map sync discipline never covered — an entry can point at a long-deleted file and diff-based selection then silently never triggers those tests (the sidebar trio sat rotted for 48 versions). First run immediately caught a fourth rotted entry: 'spec authored quality' referenced test/fixtures/spec/** (directory does not exist) and selected for a judge test that exists nowhere in the repo. Removed. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(security): remove deleted /sidebar-chat endpoint from tunnel allowlist TUNNEL_PATHS is the audited tunnel attack surface — its own comment says every addition widens it. '/sidebar-chat' stayed in the set after the endpoint was deleted with the chat-queue path, meaning any future route matching that path would have been silently tunnel-exposed. The set is now exactly the pair ceremony (/connect) and the scoped command endpoint (/command), and the dual-listener closed-set pin enforces that. Also repairs a pre-existing red pin in dual-listener.test.ts: v1.63.0.0 made the tunnel allowlist args-aware (canDispatchOverTunnel gained a second param) without updating the test — red on main since then, invisible because the free suite had no CI job. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(security): delete chain's shadow dispatcher that skipped every security gate meta-commands.ts carried a 'CLI mode' fallback that re-implemented command routing without the server pipeline's gates: no scope check, no domain check, no tab ownership, no rate limit, no hidden-element stripping, no scoped-token enveloping — and it called handleReadCommand without a BrowserManager, which also skipped the JS-origin cookie-exfiltration assertion. It was unreachable in production (server.ts always passes executeCommand) and one boolean away from being live. chain now hard-errors without a server context. handleReadCommand's bm param is required and assertJsOriginAllowed runs unconditionally. The chain tests that exercised the deleted fallback now route through a server-shaped executeCommand adapter (real handlers + trust wrapping + {status,result} envelope), so their behavioral coverage — sequencing, trust markers, pipe format, aliases, error reporting — survives on the production-shaped path. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(extension): delete the dead chat-queue client surface The sidebar-command handler in background.js POSTed to a server endpoint that no longer exists (deleted with the chat queue) — ~35 lines of fully-wired dead code including error handling for the permanent 404, plus its allowlist entry. No sender in the extension ever emitted the message type. chatEnabled leaves the /health contract (server hardcoded false, background.js re-derived it, nothing consumed it — the chat input element it guarded is gone from sidepanel.html). BROWSE_SIDEBAR_CHAT env flag had zero readers. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(security): delete dead exports the ripped chat path left behind Three-way split by importer class: (a) Zero importers, deleted: the whole attack-attempt logging cluster in security.ts (logAttempt, AttemptRecord, salted hashPayload + device-salt, attempts.jsonl rotation, telemetry spawn plumbing incl. buildTelemetrySpawnCommand/resolveBashBinary — the LIVE attempts.jsonl writer is tunnel-denial-log.ts with its own rotation); the decision-file handshake (writeDecision/readDecision/clearDecision/excerptForReview — written for sidebar-agent's poll loop, which no longer exists); sidebar-utils.ts (whole module — its sanitizeExtensionUrl 'sanitized before embedding in a prompt' for the deleted prompt builder); 8 dead server.ts imports (sanitizeExtensionUrl, generateCanary, injectCanary, writeDecision, rotateRoot, serializeRegistry, restoreRegistry, clearAgentRecord); buildPtyClearCookie + buildSseClearCookie; WEBDRIVER_MASK_SCRIPT (orphaned by the D7 stealth narrowing — applyStealth never used it). (b) Dead-pin tests edited with their exports: the 'still exported' pin in stealth-layer-c, the string-content describe in stealth-webdriver (its live applyStealth behavioral coverage untouched), the clear-cookie assertions, security-review-flow.test.ts deleted whole (all 4 describes exercised the dead decision mechanism, incl. a 'simulated sidebar-agent poll loop'). (c) KEPT deliberately: leaseCount (live behavioral coverage), extractPtyCookie + validatePtySessionToken (extractPtyCookie is adopted by the terminal-agent cookie-parse unification later in this wave), resetSessionMarker + clearContentFilters (test-support API for the live content-security layer). Also fixes two pre-existing red pins found while here, invisible until the free suite got a CI job: the v1.44 spawnClaude->maybeSpawnPty rename in terminal-agent.test.ts, and a cross-file test-isolation bug where content-security.test.ts's clearContentFilters() wiped the auto-registered url-blocklist filter for every later file in the same bun process (security-integration.test.ts failed on co-run; afterAll now restores it). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(security): delete the dead ML layers — transcript classifier and DeBERTa ensemble The L4b Haiku transcript classifier and the opt-in DeBERTa ensemble (GSTACK_SECURITY_ENSEMBLE=deberta, a documented 721MB download) had ZERO production callers since the chat-path agent that invoked them was ripped. The only live ML path is scanPageContent (testsavant) inside the security sidecar subprocess. Deleted by import graph: - security-classifier.ts 614 -> 265 lines: HAIKU_MODEL, checkTranscript, shouldRunTranscriptCheck, loadDeberta, scanPageContentDeberta, ToolCallInput, all DEBERTA_* consts + load state. Header now states the live truth (imported only by security-sidecar-entry.ts). downloadFile kept, name intact — it is an enumerated egress sink (HF model download). - security-bunnative.ts + test: a research skeleton self-described as 'NOT a production replacement', shipped into src/ with zero importers. - security-bench-ensemble{,-live}.test.ts + the Haiku response fixture: a paid live-model benchmark for a layer that could not fire. The security-classifier-tdz test's only case exercised checkTranscript — gone. - security.ts: layer-model header rewritten to the live architecture; StatusDetail.layers -> {testsavant, canary}; getStatus() no longer requires the impossible transcript==='ok' for 'protected' (old on-disk session state with a transcript key is tolerated on read, never re-emitted). - security-sidecar-entry.ts needed zero changes: it serializes getClassifierStatus() verbatim and no consumer read .transcript (verified in sidecar-client + server.ts). - BROWSER.md security section matches reality (ensemble knob gone, 112MB not 22MB, sidecar hosting documented). combineVerdict/THRESHOLDS retained as the pure, tested combiner of record — comments now flag transcript/deberta votes as producer-less. Net: 26 pass in security.test.ts incl. a NEW regression test for stale- transcript disk tolerance; egress-receipt tripwire green. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs: scrub the sidebar-agent ghost from comments and CLAUDE.md 20+ comments across 10 files still described the deleted sidebar-agent.ts as a live process — including load-bearing architecture claims ('IMPORTED ONLY BY sidebar-agent.ts', 'sidebar-agent fills this in on first prompt-injection load', 'kill sidebar-agent' in shutdown docs) and ~60 lines of tombstone blocks in server.ts enumerating deleted identifiers by name (a false grep surface: searching processAgentEvent hit server.ts and looked live). CLAUDE.md's security-stack section now documents the LIVE architecture: L1-L3 content filters + testsavant via the security sidecar subprocess; the L4b/ensemble rows, the GSTACK_SECURITY_ENSEMBLE knob, and the 721MB DeBERTa download are gone (deleted as dead code this wave) with an explicit do-not-re-document note; attempts.jsonl is correctly attributed to tunnel-denial-log.ts; the no-live-writer status of classifierStatus is stated. Comments that survive now describe what IS, not what WAS: the promotion gate in domain-skills.ts explains why classifier_score>0 is load-bearing given no L4 load-time scan exists; file-permissions.ts names real sensitive files. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(gen): delete the codex-helpers shadow module gen-skill-docs.ts imported externalSkillName (unaliased) from resolvers/codex-helpers.ts at line 21 and then re-declared the same function locally — the import was silently shadowed, and the imported copy was the STALE one (it lacked the frontmatterName param the local copy grew). Three more functions were byte-identical duplicates, imported only under _-prefixed aliases to keep the module 'referenced', and transformFrontmatter was a superseded hardcoded-Codex variant. Nothing else imported the module. Also drops three dead top-of-file imports (COMMAND_DESCRIPTIONS, SNAPSHOT_FLAGS — which pulled the whole browse/src module graph into every generator run for nothing — and an unused review-resolver trio). Proof: bun run gen:skill-docs exits 0 with a byte-identical tree (zero-diff regen); gen-skill-docs.test.ts 405/405 green. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(server): delete ServerConfig.idleTimeoutMs + chromiumProfile — documented, never read Both fields carried JSDoc asserting embedder behavior that did not exist: the idle check reads the module-level IDLE_TIMEOUT_MS env constant, and both resolveChromiumProfile() call sites pass no argument. Worse than absent — an embedder passing idleTimeoutMs: 5000 silently got 30 minutes. Wiring them honestly is impossible today: the idle timer, activity state, and shutdown target are module-global, so a per-factory value would lie for any process running more than one handler. Deleted instead, with a ServerConfig note pointing at the deferred singleton/route-table refactor where real support belongs. BROWSE_IDLE_TIMEOUT and CHROMIUM_PROFILE env remain the honest knobs. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(security): wire appendSecureFile at the four real log-append sites file-permissions.ts carries a 24-line rationale for why POSIX mode bits are insufficient on Windows and implements appendSecureFile (0600 at create, Windows ACL on first write only) — but its single caller was the dead logAttempt, while the four REAL page-content log writers (console/network/ dialog logs in server.ts, the command audit log) used raw fs.appendFileSync with no mode. Page-content-derived logs now get owner-only permissions from birth on every platform. Verified before wiring: mode applies atomically at create via appendFileSync {mode}, and the ACL pass runs only on first write — no per-append subprocess cost on the hot console-log path. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(stealth): handoff() uses the shared profile resolution + lock cleanup The headless-to-headed handoff path hardcoded ~/.gstack/chromium-profile, silently ignoring $CHROMIUM_PROFILE and $GSTACK_HOME (gbrowser's gbd sets per-workspace profiles), and skipped cleanSingletonLocks() — so a handoff into a profile with a stale SingletonLock could hang where launchHeaded() would have recovered. This was the third live drift between the three Chromium launch paths; the first two are documented in comments as shipped stealth regressions. Minimal targeted fix — the full buildLaunchConfig() extraction stays in the deferred queue. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(gen): resolver registry describes the template language again Seven registered {{PLACEHOLDER}}s had zero uses in any .tmpl (checked in both bare and :arg forms): REDACT_TAXONOMY_TABLE, TEST_COVERAGE_AUDIT_REVIEW, MODEL_OVERLAY, QUESTION_PREFERENCE_CHECK, QUESTION_LOG, INLINE_TUNE_FEEDBACK, MAKE_PDF_SETUP. The last two of those families are invoked programmatically by preamble.ts (functions kept, registry entries dropped); the question-tuning trio and the review coverage-audit wrapper were documented by their own module as existing 'for unit testing' that no test performed — deleted, along with generateRedactTaxonomyTable + its EXAMPLE/TIER_BLURB constants (its '/cso renders the full table' comment was itself stale) and its test describe. Also deletes the gated-resolver mechanism (ResolverEntry/appliesTo/ unwrapResolver + test/resolver-entry.test.ts): fully built, fully tested, used by zero of the 65 registry entries — the generator loop simplifies to a direct function call. CLAUDE.md's redact-doc line stops advertising the dead token. Proof: zero-diff regen (0 SKILL.md changed); gen-skill-docs + skill-validation 737 tests green. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(gen): wire boundaryInstruction from host config; drop three no-op binDir ternaries hosts/codex.ts declared boundaryInstruction and nothing read it — review.ts kept its own byte-identical CODEX_BOUNDARY literal (verified equal + trailing escaped newlines). The resolver now reads the config, so the boundary has one owner. (autoplan's template carries deliberately generic variants, enforced by gen-skill-docs.test.ts:1358 — untouched by design.) The 'ctx.host === codex ? $GSTACK_BIN : ctx.paths.binDir' ternary appeared in three resolvers and could never change the result: resolvers/types.ts already sets binDir to $GSTACK_BIN for every usesEnvVars host including codex. Proof: zero-diff regen for claude AND codex hosts; gen-skill-docs + host-config suites green. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(test-infra): judge uses resolveClaudeBinary; eval:watch reads the real partials dir judgePtyState spawned the bare string 'claude' three definitions below the resolveClaudeBinary() helper this same file exports — broken under hermetic PATHs where every other launch in the file resolves correctly. eval:watch read _partial-e2e.json from the legacy global ~/.gstack-dev/evals/ while EvalCollector writes it into the per-project eval dir (or GSTACK_EVAL_DIR) — so the dashboard's completed-tests panel was empty whenever slug detection succeeded, i.e. the normal case. The heartbeat and per-run progress logs stay global by design (session-runner.ts: 'heartbeat stays global'). The three eval-CLI docstrings stop claiming the legacy dir is the primary location. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(test): delete the superseded SDK ship-idempotency suite and three orphaned fixtures test/skill-e2e-ship-idempotency.test.ts's own header documented that the monolith's SDK-harness version tests a synthetic prompt while it exercises the real /ship skill — the author knew the old suite was superseded and left both running, two paid LLM runs for one behavior. The weaker copy is gone; its 'ship-idempotency' diff-selection key goes with it (the dedicated file is periodic-tier, which always runs under EVALS_ALL — the key had no remaining consumer). Fixture rot: test/fixtures/golden-ship-claude.md was a 128KB zero-reader orphan that had drifted 46KB from its live successor (test/fixtures/golden/claude-ship-SKILL.md) while looking authoritative; parity-baseline-v1.46.0.0.json and v1.53.0.0.json had zero readers (three tests pin three OTHER baseline versions — consolidation is queued, deletion of the unreferenced two is free). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(bin): delete zero-caller scripts; make host-config-export's docstring honest - bin/gstack-open-url (14 lines): announced in a CHANGELOG entry, wired into nothing, ever. bin/gstack-platform-detect (27 lines): zero callers, and its hand-rolled host list was already stale (SLATE_HOST.md cites it as a problem). Note: the deprecated gstack-brain-consumer/reader pair the audit flagged was already deleted upstream in v1.63 with a stay-deleted tripwire. - scripts/task-emission-schema.ts (61 lines): a typed schema module nothing imported; the tasks-section comment now documents the JSONL fields inline. - scripts/host-config-export.ts claimed to be the 'shell bridge for the bash setup script' — setup never calls it (its hand-rolled host lists drifting is a known follow-up). Docstring now states what it IS: a standalone, test-pinned query CLI not yet wired into setup. Its validateValue + CLI_REGEX/PATH_REGEX internals were dead (defined for a guarantee the header claimed but nothing enforced). - KEPT deliberately: scripts/preflight-agent-sdk.ts — a documented manual diagnostic (CONTRIBUTING.md + USING_GBRAIN_WITH_GSTACK.md reference it). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(server): one lone-surrogate sanitizer, one sanitizeReplacer, one startTunnel Three copies of the surrogate sanitizer existed with two algorithms (sanitize.ts regex vs a hand-rolled charCodeAt walk in server.ts — verified byte-identical across 11 edge cases before converging) plus two identical sanitizeReplacer definitions each wrapping a different copy. sanitize.ts is now the single source of truth; the runs-INSIDE-JSON.stringify egress invariant is unchanged at every call site and its pin tests were adapted to the new import shape without losing intent. The ngrok tunnel-start sequence existed three times in server.ts — the /tunnel/start route and the BROWSE_TUNNEL=1 autostart were line-for-line equivalent (a comment admitted 'Same cleanup as /tunnel/start's error path'). One startTunnel() now owns the ephemeral loopback bind, the pre-send egress receipt, the state-file RMW via tmpStatePath(), and the ordered error-path cleanup; callers keep their distinct response surfaces. The BROWSE_TUNNEL_LOCAL_ONLY test path shares nothing (no ngrok, different state field) and deliberately stays separate. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(security): one session-cookie registry implementation, two instances pty-session-cookie.ts and sse-session-cookie.ts were byte-identical modulo the cookie name — mint/validate/parse/prune/TTL, the exact code a security fix would have to land in twice (and a third hand-rolled cookie parse in terminal-agent.ts had already diverged; unified next commit). createSessionCookieStore() owns the implementation; both modules become thin instantiations keeping every exported name, their distinct threat-model docstrings, and separate token spaces (an SSE-read cookie must never grant PTY access). pty-session-lease.ts deliberately stays out — different contract (sessionId/secret split, refresh, env TTL). The factory imports nothing from token-registry (cookie-picker-auth-isolation invariant, still pinned by sse-session-cookie.test.ts). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(security): terminal-agent uses the shared PTY cookie parser The /ws upgrade's cookie fallback hand-parsed the Cookie header inline — the fourth copy of the session-cookie parse, and the one that had already diverged from the others. Parsing now goes through extractPtyCookie; validation deliberately stays against the agent's own in-process validTokens map (the server's registry lives in a different process). The ws-handler pin test now pins the shared-parser call instead of the raw cookie-name literal. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * refactor(hosts): defineHost() factory — 10 copy-paste host files become declarations hosts/*.ts were ten copies of one file: runtimeRoot byte-identical in 9/10, pathRewrites mechanically derivable from the host name for 7/10, the 11-entry toolRewrites map byte-identical between openclaw and gbrain, and every asset change a 10-file edit (cursor and slate had already fallen out of three other hand-maintained lists). defineHost() owns the defaults; each host file now declares only what makes it different (slate/cursor: 8 lines each). Shared constants: CROSS_MODEL_RESOLVERS, GBRAIN_RESOLVERS, EXEC_STYLE_TOOL_REWRITES. Genuinely-different things stayed explicit: codex/factory $GSTACK_ROOT rewrites, hermes's tool vocabulary, claude's denylist+prefixable install, opencode's wider runtimeRoot. Proof: JSON.stringify(ALL_HOST_CONFIGS) dump-diff before/after EMPTY (and a runtime walk confirmed no function-valued or undefined-keyed fields, so the JSON diff is complete); gen:skill-docs --host all zero-diff; host-config + gen-skill-docs + idempotency suites 485/485. Host files 595 -> 285 lines. docs/ADDING_A_HOST.md teaches the factory pattern. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat(lib): fs-atomic — one atomic-write implementation, with the race actually fixed Atomic tmp-write-then-rename was reimplemented ~20 times across lib/, bin/, and browse/src with three tmp-suffix conventions. One of them was a latent bug this commit closes: lib/worktree.ts used a bare '.tmp' suffix — the deterministic-tmp collision race browse/src/server.ts documents having hit in production (its fix, pid+random, was trapped in a comment at one site). lib/fs-atomic.ts: atomicWriteSync (always throws, best-effort tmp cleanup, pid+random suffix, optional mode applied at tmp creation so the file never exists with looser permissions) + atomicWriteQuiet (shutdown paths only). Unit tests pin the throw/quiet contracts, 0600 mode, tmp-name uniqueness (captured via the read-only-dir failure path — Bun's fs exports are readonly, no monkeypatching), and no-stray-tmp cleanup. Migrated: lib/worktree.ts (the bare-.tmp bug), lib/gstack-decision.ts (snapshot + compact log), lib/gbrain-local-status.ts (probe cache). browse sites follow separately. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(lib): jsonl-store's docstring stops lying; mode option added; lib bypasses adopted The header claimed 'single source of truth... the ONLY copy' with write-time injection REJECTION — while appendJsonl never screened anything, only 1 of ~10 JSONL stores imported it, and a bypass appender lived in the same directory. Now: the contract is explicit (screening is the CALLER's job via hasInjection/firstInjectionMatch; the enforcing callers are named), a option applies 0600 at create for sensitive stores, and the lib bypasses are adopted (gstack-memory-helpers ×2, redact-audit-log — which keeps its chmod backstop for files created looser by pre-mode versions). browse/src keeps its own appenders by design (compiled-binary surface, own secure-append helper) and the header now says so. gstack-decision's batched archive append stays deliberate (single-write crash-window semantics appendJsonl's one-record contract can't express). New pins: 0600-at-create, and a test that documents appendJsonl does NOT self-screen — so nobody can re-document it as self-screening without making it true. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(browse): migrate hand-rolled atomic writes to lib/fs-atomic Seven sites, each audited for its existing throw-vs-swallow contract before migrating: writeSessionState + the four fire-and-forget tab/state writers use atomicWriteQuiet (they swallowed before); writeAgentRecord + the boot-time port-file write use atomicWriteSync (they threw before — and writeAgentRecord previously leaked its tmp file on rename failure, which the helper cleans). All carry {mode: 0o600} plus restrictFilePermissions after successful writes, preserving the Windows ACL hardening that writeSecureFile provided (mode bits are POSIX-only). server.ts untouched: its three state writes route through tmpStatePath(), pinned by server-tmp-state-path.test.ts. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(hosts): delete five dead HostConfig fields metadataFormat (generator hardcodes openai.yaml), sidecar (behavior lives in setup's create_agents_sidecar — knowledge preserved as a comment in codex.ts), install.prefixable (skill_prefix is implemented entirely in bin/gstack-config), staticFiles (docstring cited a SOUL.md that never existed anywhere), and adapter (its only would-be consumer, openclaw-adapter.ts, was fully dead — with a test asserting the field was undefined). Kept: learningsMode (wired next), linkingStrategy (validation reads it), coAuthorTrailer (consumed by resolvers/utility.ts). Proof: JSON dump diff shows ONLY the deleted keys vanishing; zero-diff regen across all 10 hosts; host-config + gen-skill-docs suites green. Note: this commit also carries chunk-23 edits to the shared hosts/claude.ts + define-host.ts + host-config.test.ts files (skipSkills collapse, stale line-number comment drops) — pathspec commits, concurrent prep. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(gen): preamble tiers are explicit; silent ?? 4 default becomes an error; spec stops rendering its preamble twice Eight skills (scrape, diagram, spec, skillify, pair-agent, landing-report, open-gstack-browser + its connect-chrome symlink) silently received the HEAVIEST tier-4 preamble because a missing frontmatter field defaulted to 4. Tiers are now declared in every {{PREAMBLE}} template's frontmatter and a missing declaration throws at generation time with the template path (the 5 templates without {{PREAMBLE}} never invoke the resolver). The stale hand-written tier-map comment (wrong in 3 of 4 rows) is gone. Bonus bug fixed: spec/SKILL.md.tmpl mentioned {{PREAMBLE}} in prose, so the generator inlined the ENTIRE preamble a second time — spec/SKILL.md shrinks 127,462 -> 80,924 bytes (-46,538) from de-duplication alone. skill-size-budget gains a reasoned INTENTIONAL_SHRINKS entry (its frozen baseline had measured the doubled-preamble bug). New tests: missing-tier throw carries the path; every {{PREAMBLE}} template declares a tier. (Carries chunk-23 edits in the shared test/gen-skill-docs.test.ts.) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(gen): learningsMode is read from host config, not a hardcoded host name resolvers/learnings.ts branched on ctx.host === 'codex' while every host declared learningsMode — the field was decorative, and the 7 hosts configured 'basic' (cursor, slate, kiro, opencode, openclaw, hermes, gbrain) silently received the 'full' cross-project flow their runtimes can't execute (it depends on AskUserQuestion + gstack-config plumbing). Output now matches declaration: basic hosts get the project-scoped search block. Blast radius proof: all committed Claude SKILL.md files and the three golden fixtures are byte-identical; the behavior diff lands only in the gitignored external-host trees (hand-verified: .cursor review's learnings section swaps the cross-project AskUserQuestion block for the project-scoped search). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(gen): small config scrubs — openclaw blobs to real files, setup host drift, dead artifacts - The three openclaw markdown blobs hardcoded inside gen-skill-docs.ts (which silently reverted any hand edit to their tracked outputs on regen) move to openclaw/templates/*.md source files; output shasums byte-identical. - setup's --host allowlists gain cursor + slate — both fully registered hosts with generated output, but './setup --host cursor' exited 1 because two hand-rolled lists in setup had drifted from hosts/index.ts. - scripts/proactive-suggestions.json deleted: 31KB regenerated on every run, read by nobody (the catalog-trim design's reader was never built); its emitter and three determinism tests (which guaranteed a file nothing reads didn't churn) retired with stays-retired pins. - claude/SKILL.md.tmpl deleted: a complete 8.9KB skill that never generated output (directory name collides with the host id 'claude'), in no registry. Recoverable from git if ever wanted under a non-colliding name. - openclaw's frozen extraFields.version '0.15.2.0' stamp dropped; includeSkills: [] no-ops omitted (the generator treats [] as absent); llms.txt 55 -> 54 skills. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat(gen): correct preamble tiers for the 8 silently-heaviest skills With tiers now explicit, set them RIGHT by analogy to the tiered population: scrape/diagram/open-gstack-browser (+ the connect-chrome symlink) -> tier 1 (launchers and artifact generators, like browse and make-pdf); landing-report/pair-agent/skillify -> tier 2 (dashboards and session tools, like health and canary); spec -> tier 3 (interactive planning, like the plan-*-review family). Each tier-1 skill sheds 271 lines of onboarding prose it never needed; tier-2 shed 20 each. Verification per the review protocol: regen diff reviewed (pure section-removal), skill-validation + size-budget + catalog-budget + v0-dormancy suites green (822 tests), and live smoke of the tier-corrected skills confirms the preamble renders the intended sections at each tier. These skills have ~no eval coverage — stated honestly; the wave's gate-tier eval run is the backstop. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat(test): e2e-gate — one tier-gate implementation, side-effect-free, with the trap pinned The EVALS/EVALS_TIER gate was copy-pasted into ~40 test files and had drifted into six different predicates — the drift that made 'eval:bg:all runs everything' silently false. test/helpers/e2e-gate.ts owns the semantics now: describeE2ETier(tier) + e2eTierEnabled(tier), env read at call time, zero side effects (the existing e2e-helpers module runs a ~30s claude ping at import under EVALS=1, so the gate lives in its own module; purity is pinned by tests that scan imports and comment-stripped source). The unit matrix pins all four env combos — including EVALS=1 with EVALS_TIER unset -> SKIP, the exact trap that made eval:bg:all a non-run. The tier-alignment tripwire gains a second regex for the helper shape (old shape still detected — stragglers can't hide), and the sharded paid runner's PRE-SPAWN tier classifier learns the helper shape too: without that, every gate-sharded run would have spawned all 28 periodic shards just to skip them, each paying the e2e-helpers import ping (~15 min of dead wall clock in the CI-blocking lane). Verified: gate runs exclude the 29 periodic files, periodic excludes the 8 gate files — identical to pre-migration. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * refactor(test): migrate the 36 tier-gated eval files to describeE2ETier Mechanical two-liner swap in 34 files (each keeping its declared tier — all 36 predicates verified against E2E_TIERS before migrating); the two files with compound gates (overlay-harness's EvalCollector feed, codex-e2e's CODEX_AVAILABLE) keep their extra conditions via e2eTierEnabled. Tier rationale comments preserved. codex-e2e/gemini-e2e/benchmark-providers keep their distinct stderr-message gate shapes by design. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * refactor(test): skill-e2e + skill-llm-eval adopt the shared selection machinery Both files re-implemented the diff-selection machinery e2e-helpers already exported. The helper gained computeDiffSelection() (extracted, identical behavior) and a trailing optional selection param on the *IfSelected helpers (defaults preserve all 30+ existing importers). skill-e2e.test.ts drops ~120 duplicated lines; skill-llm-eval keeps its LLM_JUDGE_TOUCHFILES selection and test.concurrent semantics via testConcurrentIfSelected. Deliberate deltas, stated: skill-e2e.test.ts now honors the EVALS_TIER intersection its local copy lacked (affects only direct bun test invocations of that file — it matches no eval-script glob); its recordE2E gains the helper's three diagnostic fields; skill-llm-eval sharded solo now runs e2e-helpers' module-scope preflight it already ran in combined processes. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(test): kill the silent-truncation race; exempt the tier-corrected shrinks The full-suite shakeout (budgeted by the plan) surfaced both immediately: 1. server-embedder-terminal-port.test.ts stubbed process.exit and restored the REAL exit in its finally — but shutdown() schedules async work that can call process.exit AFTER restoration, killing the entire bun process mid-suite with exit 0 and NO summary. This is the silent-truncation class the new free-suite CI job guards against, reproduced locally on the first full run. Exit now stays a logging no-op between tests (late async exits become visible stderr lines, not process death); the true exit returns in afterAll. 2. The 80%-of-baseline shrink guard correctly flagged the six tier-corrected skills — their baseline was measured at the silent tier-4 default. Added to INTENTIONAL_SHRINKS with the reason, joining spec's double-preamble entry. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * release: v1.64.0.0 — the code-smell fix wave 35 commits, one PR: guard repairs (free suite in CI per-file, all-host freshness gates, tunnel allowlist, diff-selection validation), the sidebar-agent ghost exorcism (dead ML layers, dead endpoints, dead exports, ghost comments), config honesty (defineHost factory, dead fields deleted, preamble tiers explicit, spec double-render fixed), and dedup with safety nets (session-cookie factory, fs-atomic, jsonl-store contract, one eval tier-gate). Net -24,943 lines across 183 files. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(ci): free-tests step runs under bash (container sh rejects pipefail) Maiden-voyage shakeout, exactly as budgeted: the CI container's default shell is dash, which errors on 'set -o pipefail' before the first test ran. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(ci): free-tests curates 8 container-incompatible files with reasons Second maiden-voyage shakeout round: 376 of 384 files ran green in the container on the first completed pass. The 8 that can't run there yet are excluded the same way the Windows shards curate POSIX-bound files — each with its reason inline (headed-Chrome handoff, real-PTY round-trip, X server management, extension-origin identity, the job's own TMPDIR override, and three pre-existing env failures that fail on dev machines too). Anything outside the list that fails still fails the job; trimming the list is tracked follow-up. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(test): gstack-config-key-locale — suppress the skill_prefix auto-relink side effect The test invokes the repo's own bin/gstack-config, whose 'set skill_prefix' auto-runs $(dirname $0)/gstack-relink — resolving the install dir to the repo itself. In any environment where the loop shares a working tree (the free-tests CI container, a fresh-HOME run), gstack-patch-names rewrote all 52 tracked SKILL.md names to gstack- prefixed, poisoning five unrelated suites downstream (hermetic-skills-seeding, host-config golden, skill-census, skill-validation, spec-template-sync). GSTACK_SETUP_RUNNING=1 is the documented suppression; relink behavior stays covered by relink.test.ts's mock install. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(bin): gstack-codex-session-import — empty sessions dir exits 0 on Linux GNU xargs runs 'ls -t' once even on empty input, listing the cwd and producing a bogus LATEST from the repo root; BSD xargs (macOS) skips the run, which is why the NO_SESSIONS path only broke on Linux. xargs -r pins the BSD behavior on both platforms. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * test(parity): rebaseline v1.57.7.0 → v1.64.1.0 + skeleton-cap headroom The two parallel v1.64 waves (code-smell fix wave + main's #2571) each added shared-preamble prose, pushing document-release / design-consultation / cso past their size ratios on the v1.57.7.0 anchor and four carved skeletons (plan-ceo-review, plan-eng-review, office-hours, design-consultation) 22-280 B over their absolute caps. New baseline is union-normalized (skeleton + sections/*.md, matching what the harness measures); caps get +~1 KB headroom each with per-cap rationale. The v1.57.7.0 fixture stays in test/fixtures/ for the audit trail, and capture-parity-baseline.ts now documents the union-normalization step so the next rebaseline doesn't re-trip on it. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(ci): free-tests container parity — tools, pinned bun, git identity, mutation tripwire - Dockerfile.ci: add python3 (gstack-jsonl-merge/brain-sync/detach shell out to it), file (skill-validation's binary check), poppler-utils (make-pdf e2e gates hard-require pdftotext/pdffonts/pdfinfo), fonts-noto-color-emoji (emoji render gate, mirrors make-pdf-gate.yml). Fix the bun pin: the bun.sh installer ignores a BUN_VERSION env var, so the old form silently installed latest on every rebuild (observed 1.3.13/1.3.14 drift vs the 1.3.10 devs run locally); pass the version as the positional arg. - free-tests.yml: git identity + safe.directory for the git-exercising tests (container checkout is owned by a different uid than runner); post-loop tree-mutation tripwire that names a tracked-file-mutating test instead of letting downstream collateral confuse the report; skip the documented variants-retry-after timing flake. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(bin): gstack-session-update — detached updater owns its stdio (SIGPIPE) The backgrounded update subshell inherited the session hook's stdout/stderr pipes. Once the hook exits and the caller closes them, any child that writes — git pull's autostash notice, setup output — dies of SIGPIPE, logged as PULL_FAILED exit=141 with an empty stderr capture (observed in the free-tests container, and reachable by any production hook runner that closes stdio promptly). Redirect the fork to /dev/null; all observability already flows through the session-update log file. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(test): gstack-decision-bins — explicit branch context for the scope filter CI checks out a detached HEAD, where gitBranch() returns undefined on both the log and search sides, so an implicitly branch-scoped decision can never surface (filterByScope requires a matching non-empty ctx.branch). Pass the branch explicitly on both sides — the filter logic is what's under test, not git branch detection. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(test): ring-buffer lease interplay — same TTL window, not same millisecond Two back-to-back mintLease() calls each stamp Date.now() + TTL; when they straddle a millisecond boundary the exact-equality assertion flakes (observed in CI: expiries of ...525 vs ...526). Assert the expiries are within a 50 ms window instead — the invariant under test is that leases share a TTL policy, not that they mint in the same clock tick. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
cab774cced |
v1.56.0.0 Token-reduction Phase B + AUQ paranoid safety net (#1849)
* refactor(plan-ceo-review): carve review body into on-demand section
Carve the largest skill (138,838 B) into a skeleton + one on-demand
section, the documented next Phase B target after /ship (v2_PLAN.md:216).
- sections/review-sections.md(.tmpl): the 11-section deep review, codex/
outside-voice rules, how-to-ask, Required Outputs, registries, Completion
Summary, Review Log, REVIEW_DASHBOARD, PLAN_FILE_REVIEW_REPORT, Next Steps,
docs/designs promotion, Formatting Rules, and the Mode Quick Reference.
- sections/manifest.json: passive registry (CM2), one entry.
- SKILL.md.tmpl: {{SECTION_INDEX}} after the system audit, a single
{{SECTION:review-sections}} STOP-Read after Step 0 mode selection, and a
Section self-check. All of Step 0 (the scope/mode conversation) stays in
the always-loaded skeleton; only EXIT_PLAN_MODE_GATE follows the section.
Measured: always-loaded skeleton 138,838 -> 80,731 B (-42%, ~14.4K tokens
off every invocation). Union (skeleton + section) 139,110 B, behavior held.
Boundary honors Codex P1: nothing review-governing (formatting rules, mode
reference, how-to-ask, required outputs) sits in the skeleton below the
STOP. Housekeeping resolvers ride in the section, matching the ship
precedent (adversarial.md carries LEARNINGS_LOG + GBRAIN_SAVE_RESULTS).
Tests (atomic with the carve — skill-docs.yml gates gen:skill-docs
freshness on every push, so source + regen + tests must land together):
- parity-harness: plan-ceo flipped to sectioned, maxSkeletonBytes 90_000
(measured 80,731 + headroom); content/minBytes run against the union.
- skill-size-budget: plan-ceo-review added to SECTIONS_EXTRACTED.
- section-manifest-consistency: generalized to discover every carved skill,
vars computed per-skill-case (Codex P2).
- skill-ceo-section-ordering (new, gate): per-PR static guard — STOP after
Step 0, review body absent from skeleton, report writer in the section,
nothing review-governing below the STOP.
- skill-e2e-plan-ceo-review-section-loading (new, periodic): refreshes the
installed skill first (Codex P1), drives full Step 0, asserts the section
is Read before the report.
- gen-skill-docs + skill-validation: read the skeleton+sections union for
carved skills so relocated prose still counts.
- touchfiles: plan-ceo-section-loading registered (periodic).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* chore: bump VERSION + CHANGELOG for plan-ceo-review carve (v1.56.0.0)
MINOR: carves the largest skill into skeleton + on-demand section,
dropping plan-ceo-review's always-loaded cost 42% (138,838 -> 80,731 B,
~14.4K tokens off every invocation). User-facing release notes lead with
the measured token win.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* docs(todos): file P3 follow-up — carve the shared {{PREAMBLE}} reference blocks
Surfaced by /plan-eng-review on the plan-ceo-review carve: per-skill section
carves stay modest because the ~40-50KB shared preamble dominates the
always-loaded surface. A single preamble-reference carve would help every
tier->=2 skill at once. Records the why, the cold-vs-hot split to measure,
and the guards it needs. Not implemented this PR.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* test(auq): Layer 0 — guarantee AUQ format spec is always-loaded
Deterministic, free, per-PR keystone for the token-reduction era. For every
interactive (tier>=2) skill, asserts the full AskUserQuestion decision-brief
format (ELI10/Recommendation/Pros-cons/checks/Net/(recommended)/Stakes/
self-check) lives in the always-loaded SKILL.md skeleton, NOT only in an
on-demand section. Plus a roster guard (a carve can't silently drop the block)
and per-skill rule survival in the skeleton+sections union. 51 cases + a
negative control. Fails the instant a future carve strands AUQ-governing text
where it won't be loaded when a question fires.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* test(auq): SDK capture engine + verbose-vs-carved no-degradation A/B
Adds the reusable SDK $OUT_FILE capture engine (auq-sdk-capture.ts): drives a
skill to its AUQ and captures the verbatim text the model GENERATES, cleanly
(real-PTY mangles plan-mode AUQs via cursor escapes). Pins the skill to an
absolute path with Read/Write-only tools so the agent can't wander to the
global install. gradeAuqRecommendation normalizes a non-"because" connective
before grading so substantive reasons aren't false-flagged (without touching
the pinned shared judge).
The A/B drives the same prompt through the carved 80KB skeleton and the
pre-carve 137KB monolith and fails if carved scores worse. Result: both 7/7
format, substance 5 — proven no degradation, transcript-verified each side read
its own planted SKILL.md. Periodic tier.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* test(auq): consistency — same trigger N runs, stable format + substance
Drives the carved /plan-ceo-review AUQ N=3 times and fails if any format
element appears in one run but not another, or substance craters. Targets the
"fine one run, broken the next" failure class a single snapshot can't see.
Result: 3/3 stable, 7/7 + substance 5 every run. Periodic tier.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* test(auq): behavioral matrix across AUQ-heavy skills
Data-driven test that drives each AUQ-heavy skill (plan-eng/design/devex,
office-hours, cso, spec, design-consultation) to its first AskUserQuestion and
grades it to the plan-ceo bar: 7/7 decision-brief format + recommendation
substance >=4. One case per skill (isolated failures), env-subsettable via
AUQ_MATRIX_ONLY. Browser/design-binary skills are intentionally excluded
(comparison boards, not format-AUQs; Layer 0 covers their spec). All targeted
skills pass 7/7 with substance 4-5. Periodic tier.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* test(codex): live recommendation-substance grade for /codex
Closes the gap where /codex's synthesis recommendation was only checked
statically (template grep) and via fixtures. Drives the real /codex skill over
a flawed diff and grades the emitted "Recommendation: ... because ..." line
with judgeRecommendation (present/commits/has_because/substance>=4). The named
weak spot holds up: substance 5. Periodic tier.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* test(auq): deterministic trigger for format-compliance gate
A bare /plan-ceo-review against a repo whose work is already implemented makes
the model improvise an off-script "what should I review?" scope question that
skips the decision-brief format, which the gate test then times out waiting for.
Hand it a concrete plan to review (FORCING_FLOOR_CEO) so it reaches the real
Step 0 mode-selection AUQ that is the intended format check.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* refactor(office-hours): carve Phase 5+6 into on-demand section
Third Phase B carve (v2_PLAN.md:216, after ship and plan-ceo-review). Moves
Phase 5 (Design Doc templates) + Phase 6 (tiered relationship handoff) — the
session's output + closing tail, only reached after the conversation and
alternatives are done — into sections/design-and-handoff.md, behind a single
STOP-Read after Phase 4.5. The live conversation (Phases 1-4.5) and the
always-run Important Rules stay in the always-loaded skeleton.
Measured: always-loaded skeleton 118,280 -> 88,975 B (-24.8%). Union preserved.
The carved AUQ is identical to pre-carve (matrix: 7/7 format, substance 5),
and Layer 0 confirms the AUQ format spec stays in the skeleton — the AUQ
paranoid suite de-risked this carve end to end.
Atomic with tests + regen (skill-docs.yml gates gen:skill-docs freshness on
every push, so source + regen + tests land together; --host all regenerates
the inlined non-Claude variants):
- sections/manifest.json: passive registry, one entry.
- parity-harness: office-hours flipped to sectioned, maxSkeletonBytes 96_000
(measured 88,975 + headroom); content/minBytes run against the union.
- skill-size-budget: office-hours added to SECTIONS_EXTRACTED.
- gen-skill-docs + skill-validation: read the skeleton+sections union for
office-hours so relocated Phase 5/6 prose still counts.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* chore: bump VERSION + CHANGELOG for office-hours carve + AUQ suite (v1.57.0.0)
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* refactor(preamble): carve CJK-escaping manual to on-demand doc
The AskUserQuestion format block is inlined into every interactive skill (~33).
It carried the full multi-paragraph non-ASCII/CJK escaping manual inline, but
that rationale only matters when a question contains CJK text and the operative
rule already lives in the always-loaded self-check. Moved the justification to
docs/askuserquestion-cjk.md (read on demand); kept the rule + a pointer.
Corpus: Claude-host SKILL.md total 3,087,499 -> 3,057,975 B (-29,524 B, ~900 B
x ~33 skills). Layer 0 still passes — the core decision-brief format stays
always-loaded; only the rare CJK rationale moved. Atomic with the all-host
regen (skill-docs.yml freshness gate). VERSION + package.json -> 1.58.0.0.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* refactor(plan-eng-review): carve review body into on-demand section
Fourth Phase B carve (v2_PLAN.md:220). Moves the 4-section review (Architecture,
Code Quality, Tests, Performance), outside voice, required outputs, and review
report — everything after Step 0 scope — into sections/review-sections.md behind
a single STOP-Read. Step 0 (scope challenge) and EXIT_PLAN_MODE_GATE stay in the
always-loaded skeleton.
Measured: skeleton 106,984 -> 54,892 B (-48.7%). Union preserved. Atomic with
tests + all-host regen (freshness gate): parity flipped to sectioned
(maxSkeletonBytes 62K), plan-eng-review added to SECTIONS_EXTRACTED, gen-skill-docs
reads the union for relocated review/TEST_COVERAGE/dashboard prose. Layer 0 green.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* refactor(plan-design-review): carve review body into on-demand section
Fifth Phase B carve (v2_PLAN.md:220, bundled with plan-eng). Moves the 7 design
passes, required outputs, and review report — everything after Step 0 scope and
the mockup/rating phase — into sections/review-sections.md behind a STOP-Read.
Step 0, Step 0.5 mockups, the rating method, and EXIT_PLAN_MODE_GATE stay in the
always-loaded skeleton.
Measured: skeleton 112,057 -> 76,024 B (-32.2%). Union preserved. Atomic with
tests + all-host regen: parity sectioned (maxSkeletonBytes 82K), added to
SECTIONS_EXTRACTED, gen-skill-docs reads the union. Layer 0 green.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* refactor(plan-devex-review): carve review body into on-demand section
Sixth Phase B carve. Moves the 8 DX passes, required outputs, and review report
— everything after the Step 0 DX investigation — into sections/review-sections.md
behind a STOP-Read. All of Step 0 (persona, empathy, benchmark, journey trace,
roleplay) + the rating method + EXIT_PLAN_MODE_GATE stay always-loaded.
Measured: skeleton 110,621 -> 69,658 B (-37%). Union preserved. Atomic with
tests + all-host regen: added to SECTIONS_EXTRACTED, gen-skill-docs reads the
union. Layer 0 green. (No parity invariant entry for plan-devex-review.)
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* chore: bump VERSION + CHANGELOG for plan-* family carves (v1.59.0.0)
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* test: refresh ship golden baselines + gbrain-detection union after carves
Two follow-ups the carve commits should have carried (caught by the full suite,
missed by targeted subsets):
- ship golden baselines (claude/codex/factory) regenerated: the preamble CJK
trim (v1.58) changed ship's always-loaded AskUserQuestion block.
- gbrain-detection-override probes the office-hours skeleton+section union:
GBRAIN_SAVE_RESULTS moved into sections/design-and-handoff.md when office-hours
was carved, so the detection assertions now check both files.
Full `bun test` green.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* test(auq): grade format-compliance gate from SDK capture, not the TUI
The real-PTY version grepped the stripAnsi'd interactive AUQ picker. Verified
directly that this cannot work: plan-mode AUQs render as a cursor picker whose
cursor-positioning escapes stripAnsi can't flatten — the picker renders fine for
a human (cursorSeen=45) but the flattened text drops ELI10:/(recommended) and
parseNumberedOptions returns 0. The test was grading a lossy projection and
failed by construction.
Rewritten to drive /plan-ceo-review via the SDK $OUT_FILE capture (the agent
writes the verbatim question it would have shown — clean text, no rendering
loss) and grade 7/7 format + kind-note + recommendation substance >=4. Same
property, reliable, environment-independent; shares the engine with the periodic
A/B and matrix evals. Result: 7/7 format, substance 5. Touchfiles key renamed
ask-user-question-format-pty -> auq-format-gate (no longer a PTY test).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* test: fix carve-broken CI evals (union reads + section fixtures)
Two CI eval jobs failed on the carved plan-* skills because they read content
that moved into sections/:
- llm-judge (skill-llm-eval): runWorkflowJudge sliced SKILL.md between markers
like "## Review Sections" / "## CRITICAL RULE" that now live in
sections/review-sections.md. The markers vanished from the skeleton, so the
judge scored empty/wrong content. Fix: read the skeleton+sections union.
Verified: plan-ceo modes / plan-eng sections / plan-design passes all PASS
(25/25).
- e2e-plan (skill-e2e-plan): setupPlanDir copied only <skill>/SKILL.md into the
fixture, not sections/. The carved skill's STOP pointed at a section file that
was absent, so the model improvised a compressed report table instead of the
canonical "| Review | Trigger | Why | Runs | Status | Findings |". Fix: copy
sections/ alongside SKILL.md in all 6 setup sites. Verified: report test PASS,
canonical table emitted.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* test: copy carved sections into all e2e fixtures (prevent more carve-blind CI fails)
Proactive sweep beyond the two CI logs: every e2e test that copies a carved
skill's SKILL.md into a temp fixture must also copy its sections/, or the
model hits a STOP pointing at a missing section file and improvises/degrades.
- skill-e2e.test.ts: plan-ceo/plan-eng/plan-design/office-hours copies across
planDir/reviewDir/ohDir/benefitsDir dests now copy sections/.
- skill-e2e-plan.test.ts: the office-hours copy + the 4-skill codex-offering
loop now copy sections/.
- skill-e2e-design.test.ts: plan-design-review copy now copies sections/.
- skill-e2e-office-hours.test.ts: both office-hours copies now copy sections/.
- skill-e2e-office-hours-brain-writeback.test.ts: GBRAIN_SAVE_RESULTS moved into
the section, so check the regenerated skeleton+section UNION for the gbrain put
block, ship both into the workdir, and restore both (the section regen was also
leaking into the working tree — finally now restores it).
ship copies (single-file Step-0 slices) and review/retro (not carved) untouched.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* test: migrate section-loading E2E to lossless SDK tool-stream detection
The /ship and /plan-ceo-review section-loading tests drove a real PTY and
scraped the ANSI screen buffer for sections/<file>.md paths. That silently
saw nothing in a Conductor PTY (cursor-positioned tool renders and an
unanswered Step 0 question loop both defeat the regex), so both reported
read: [] even when the agent did the work.
They now run the skill through claude -p (the same SDK path the AUQ matrix
uses) and detect section reads from the tool-use stream — Read calls whose
file_path contains sections/<file>.md — with no rendering layer to mangle.
The run is also hermetic: the freshly-generated worktree skeleton + sections
are copied into a throwaway fixture with the absolute path pinned, so the
test validates this branch's carve without mutating the user's ~/.claude
install.
Validated EVALS_TIER=periodic: both pass (plan-ceo Reads review-sections.md;
ship Reads review-army.md + changelog.md), ~6.5 min for both vs ~23 min
combined on the old PTY path where both were failing.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* chore: consolidate branch to v1.56.0.0 (single MINOR above main)
The branch bumped VERSION several times during development (1.56 → 1.57 →
1.58 → 1.59), but none of those landed on main (main is at 1.55.1.0). Per
the "never orphan branch-internal versions" discipline, collapse all four
into a single 1.56.0.0 entry — one MINOR release covering the whole branch:
five skills carved (plan-ceo, office-hours, plan-eng, plan-design,
plan-devex), the shared AskUserQuestion preamble CJK trim, and the paranoid
AUQ no-degradation test suite + lossless section-loading tests.
VERSION and package.json set to 1.56.0.0; main's 1.55.1.0 entry preserved
below the consolidated entry. No SKILL.md drift (VERSION is not embedded in
generated bodies).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
|