Files
gstack/scripts/test-free-shards.ts
Garry TanandClaude Fable 5 e76f65a8da v1.77.0.0 feat: test-infrastructure overhaul wave 1 — matrix deletion, flake telemetry, sync-spawn wedge class extinct (#2746)
* fix: pin the claude CLI to an exact version in the CI image + tripwire

The image installed @anthropic-ai/claude-code UNPINNED and rebuilt weekly
'to pick up CLI updates' — while bun sat carefully pinned at 1.3.13 two RUN
lines above. The PTY harness screen-scrapes this CLI's TUI, and that drift
broke it three separate times (welcome-screen wedge on 2.1.233, skillify
HOME discovery on 2.1.237, guard/freeze hooks on 2.1.162), each debugged as
a flake first. Pin 2.1.251 (current latest), bump deliberately via a PR
that runs the PTY gate, and enforce with test/ci-image-cli-pin.test.ts:
any global npm install in Dockerfile.ci without an exact @X.Y.Z pin fails
the free suite. The weekly ci-image cron stays as a cheap tag self-heal.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat: stamp the claude CLI version into every eval-store run record

Three harness breakages were traced to claude-CLI TUI drift only after long
flake hunts, because no run record said which CLI it actually exercised.
EvalCollector now stamps claude_cli_version (claude --version, cached once
per process, 'unknown' when the binary is absent) into both partial and
finalized records — schema-additive optional field, no SCHEMA_VERSION bump.
Correlating a flake wave with a CLI release becomes a grep over
~/.gstack/projects/<slug>/evals/ instead of archaeology.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix: give the spinning-shard kill test load headroom (30s -> 90s)

The test spawns and group-kills three real children (one a busy-loop
burning a full core) while five sibling shard processes compete for eight
vCPUs. Under full-suite load it blew bun's default 30s per-test ceiling at
30,009ms — while passing in isolation in 1.4s — and red the only required
lane. Every assertion in it is event-based (statuses, group-kill proof,
heartbeat lines); the sole latency claim is the <30s kill-deadline sanity
bound, which stays. Explicit 90s headroom, not a weakened oracle.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat: green-by-skip census — skip counts in the classifier, all-skipped labeling in the paid runner

bun's 'Ran N tests' line COUNTS skipped tests, so a codex/gemini shard
whose every test self-skipped (binary absent on the runner — true of every
CI runner today) exits 0, dodges the hollow-shard guard, and reads as
coverage in the weekly census. The classifier now parses bun's ' N skip' /
' N pass' recap lines; ShardOutcome carries skippedTests; formatSummary and
the fail-closed slices report label an all-skipped pass explicitly:
'all N tests SKIPPED — verified nothing'. Status stays 'passed' (external
service availability is host state, not a repo regression) but the census
can no longer mistake absence for coverage.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* refactor: extract composite actions for eval-lane setup; surviving lanes gain the fail-fast registry verification

'Fix bun temp' x3, 'Restore deps' x5, 'Seed claude interactive config' x3,
and 'Register gstack skills' x3 were byte-near-identical copies across the
legacy matrix, the sliced lane, and the periodic lane — and only the MATRIX
copy of register-skills carried the 19-line dangling-symlink + frontmatter
fail-fast loop written after a silent 'Unknown command' + 35-min-timeout
incident. Extract all four into .github/actions/ composites; the register
composite carries the verification loop (generalized over the skill list),
so the sliced and periodic lanes — the lanes that SURVIVE the matrix
deletion — now inherit the check they had silently dropped. Matrix-job
inline copies are left untouched: that job is deleted next.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat: delete the legacy 17-row eval matrix — the sliced lane is the only paid lane

Every PR paid twice: the hand-enumerated matrix (18 test files, 22.6 min,
~$21 API measured on run 33263204465) ran serialized AHEAD of the strictly
superior sliced lane via 'needs: evals' — 35.5 min wall and ~2x paid spend
for the same diff. 14 of 17 rows carried no tier:, so periodic Opus
benchmarks leaked into every PR (the e2e-plan row alone: 12/12 tests,
21.7 min, $7.28 — the wall-clock bound of ALL of CI).

Parity receipt (static, pre-deletion): the sliced lane's gate census (49
files, derived from the runner itself) strictly contains all 18 matrix test
files, plus 31 files the matrix never ran. Pure deletion — one revert
restores it. The PR comment moved into slices-report (same '## E2E Evals'
upsert marker, now sourced from slice artifacts + carrying the fail-closed
reconciliation verdict). plan-slices loses the needs edge; the dead
workflow-level EVALS_TIER env goes with it.

test/evals-workflow-matrix.test.ts (and its KNOWN_MATRIX_GAPS /
KNOWN_TIER_UNSET burn-down ratchets — retired: the sliced census makes
'every gate file runs' true by construction) is rewritten as
test/evals-workflow-wiring.test.ts: matrix stays deleted, planner/executor/
report tier + slice-count agreement, both surviving lanes on the shared
register-skills composite with its fail-fast verification loop, PR comment
survival. Expected: PR eval wall 35.5 -> ~13 min, per-PR paid spend ~halved.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix: provider-runner timeouts kill the whole process GROUP; codex/gemini inherit the orphan-drain hardening

All three provider runners (claude/codex/gemini) killed only the direct
child on timeout: tool subprocesses the CLI spawned survived as orphans
holding our pipes open and burning shared API rate (observed: a 600s
timeout stretching past 1400s; a stalled run once burned a core for 15
hours). gstack-detach's watchdog had the same shape one level up — killpg
SIGTERM, 5s grace, then a direct-child proc.kill() that orphaned
grandchildren.

Fix: spawn provider children via node:child_process with detached (own
process group) and killProcessGroup(SIGKILL) in the timeout handler —
runShardChild's proven pattern, EPERM/ESRCH fallbacks included. The codex
and gemini copies also gain the reader.cancel() + stderr Promise.race
hardening only the claude copy had (they still carried the blocked-drain
hang it fixed). gstack-detach's watchdog now group-SIGKILLs after the
grace.

Regression net: test/session-runner-groupkill.test.ts drives the REAL
runSkillTest against a fake claude shim (PATH override) that spawns a
grandchild and wedges — the run must classify timeout within budget and
leave neither shim nor grandchild alive — plus source pins on all three
runners (detached + killProcessGroup, no bare timeout kill, no Bun.spawn
reversion).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix: skill-e2e-opus-47 renders SKILL.md fixtures into a mkdtemp — never the live tree

mkEvalRoot ran gen-skill-docs with cwd=ROOT, regenerating every in-repo
SKILL.md mid-run while concurrent paid shards copyFileSync those same files
in their beforeAll (EVALS_JOBS>=4 locally, 2 per CI slice) — a sibling
could capture a half-regenerated or opus-rendered SKILL.md, and a timeout
before afterAll stranded the whole tree at the wrong model for every later
shard. A cross-shard race that could flake ANY concurrent paid test.

Render via the --out-dir flag gen-skill-docs grew for exactly this reason
(mirrors the repo layout, which is all the fixture reads), read the skill
heads from the render dir, delete it, and drop the afterAll restore-regen
entirely.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix: claude CLI version resolves in the runner parent, never on a test thread

Eng-review finding: getClaudeCliVersion's fallback is a SYNCHRONOUS
spawnSync on the same thread that polls concurrent PTY/session tests — the
judgePtyState blocking class this overhaul kills elsewhere. The paid runner
parent now resolves it once (cached) and stamps GSTACK_CLAUDE_CLI_VERSION
into every shard's env; eval-store short-circuits on the env var, and the
fallback spawn's budget tightens 10s -> 3s (bounded one-time stall, records
'unknown' on a slow CLI).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* test: wire skippedTests end-to-end through runPaidShard

The census unit tests hand-built outcomes and the classifier tests parsed
strings; nothing proved a real child's ' N skip' recap flows into
outcome.skippedTests and the formatSummary label. A commandFor fake now
prints the recap shape and the test asserts the parsed counts, the
all-skipped predicate, and the 'verified nothing' label.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix: make the setup composites rerun-safe (codex diff-review hardenings)

restore-deps: 'cp -r SRC node_modules' with an existing node_modules NESTS
the copy and leaves stale deps active — rm first. register-gstack-skills:
'ln -snf' hard-errors under set -eu when a REAL directory occupies the
gstack slot — clear a non-symlink leftover first. CI workspaces are fresh
today; a reusable composite must survive dirty reruns.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix: sweep — every sync spawn in the test trees carries a timeout (436 sites, 157 files)

spawnSync/execSync/Bun.spawnSync BLOCK the main thread, so bun's in-process
per-test timeout can never fire while one waits — a hung child (stdin read,
network probe, dead daemon) wedges the whole shard until the runner's
external wall-clock SIGKILL. This exact class reached main: free-tests run
33262077256, test/gstack-memory-ingest.test.ts (normally 2.3s) held shard 2
at the 360s wall while its five siblings finished in ~65s.

Mechanical sweep in two waves (12 + 4 fan-out agents, every edit verified
against its call site): default timeout: 30_000 (matches the free runner's
per-test budget), 120_000 for genuinely slow ops (installs, builds,
playwright, provider CLIs), helper wrappers fixed ONCE where call sites
route through them. Sites that only LOOK like calls (string fixtures, grep
needles, comments) were skipped with reasons — the enforcement commit that
follows marks them exempt.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* test: sync-spawn timeout tripwire — the wedge class stays extinct

Free scanner over all test trees (test/, browse/test/, design/test/,
make-pdf/test/, ios-qa, browser-skills): every spawnSync/execSync/
Bun.spawnSync call site must carry a timeout within a 30-line options
window, or an explicit '// tripwire-exempt: <reason>' marker. Comment
lines are skipped; exemptions are counted and ratcheted shrink-only
(ceiling 6 = the 6 string-fixture/grep-needle sites where the pattern is
CONTENT, not a call — marked in this commit). A scan-sanity test pins that
the scanner still sees >100 real call sites so it can never rot to a
vacuous green. Companion to the 436-site sweep in the previous commit.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat: paid-lane flake telemetry — record-level attempts, flaky_retries, report surfacing

bun --retry leaves a retried pass INVISIBLE in its output: a fail-then-pass
prints the error detail but no (fail) result line and recaps as a clean
pass (probed live on 1.3.10). So attempts are recorded where they cannot
lie: EvalCollector.addTest stamps a 1-based attempt on same-name re-records
(a retried test runs its body again and re-records), finalized runs carry
flaky_retries, printSummary warns loudly, and the fail-closed slices report
lists every passed-only-on-retry test — recorded and ranked, never blocking
and never silent. Cross-model confirmed (codex reached the same don't-parse
-the-stream conclusion independently).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat: free-lane flake ledger — retry ON in CI, flaky-passes recorded and uploaded

The runner's attribution-gated flaky-retry pass (cap 5, truncation veto)
was OFF in the required lane and its FLAKY-PASS evidence was console-only —
so a single timing flake red the merge gate while repeat offenders stayed
unenumerable. free-tests.yml now sets GSTACK_FREE_RETRY_FLAKY=1 and points
GSTACK_FLAKE_LEDGER at runner.temp; every flaky-pass appends a JSONL entry
(SINGLE writer: the parent runner — no concurrent-append hazard by
construction; fail-open with a loud warning so a broken ledger can never
red the lane) and the artifact uploads UNCONDITIONALLY — a flaky-pass run
is green, which is exactly when the evidence matters. Wiring pinned by
free-tests-workflow-wiring; ledger behavior unit-tested incl. the fail-open
path. Matches 2026 industry practice (retry for data, quarantine out of
merge-blocking but never out of logging) with the repo's own receipts.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat: eval:flake-rank — the flake-telemetry dial

Aggregates per-test series across every finalized eval-store run (shard
dirs included) plus the free flake ledger: runs, fails, RETRIED PASSES
(the flake signature), avg duration — ranked retries-first. This is the
readable dial behind two policies: a flaky pass never blocks a merge but
is always ranked here, and the WS16 required-check promotion needs weeks
of clean flake-rank, not vibes. --json for machines, --dir for downloaded
CI artifacts.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat: two-phase session timeout — silent APIs die at the startup grace, named

The single spawn-armed timer charged API queue latency to the work budget:
the recurring '0 turns / $0.00 / x3 attempts' failure with four budget-bump
receipts (180->300s, 240->360s, 300->420s, 90->300s). Split: startup phase
(no NDJSON byte yet) kills EARLY at min(grace, timeout) with the distinct
exitReason 'timeout_startup' — an availability verdict, not transcript
archaeology — and the work phase arms on the first byte for the REMAINING
budget, so total wall never exceeds the timeout (tier envelopes are
margin-free: tests pass timeout: CAPTURE_MS and bun-budget the same tier).
Local grace 90s (observed queue latency 60-90s), CI floor 300s (TODOS-filed;
shared runners queue harder), both pinned by the new grace tests with fake
-claude shims covering the late-first-byte and silent-API paths.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix: census integrity — 17 phantom selection keys deleted, reverse invariant added, gitignored dep patterns replaced, local map forks derived

The merge-blocking gate census counted tests that could not run. Deleted
(critic-verified against both quoted-occurrence and dep-registration
liveness): 7 *-prosons-format keys with no declaring test, ship-plan-
completion/-verification, review-plan-completion, design-shotgun-path/
session/full, autoplan-core (dead ~10 months), e2e-harness-audit (its
namesake is a FREE-suite file), plus 2 dead LLM-judge keys and 2 free-file
keys (budget-regression-pty, global-discover) misplaced in the PAID maps.
Census: 191 -> 174 keys, gate 86 -> 78 honest.

The new reverse invariant in touchfiles.test.ts makes the class structurally
impossible: every key must be quoted in a living paid test file OR
registered to an existing paid test file via its dep list (the constructed-
name binding the 2026-08 self-registration sweep established) — zero
exceptions needed today, with a live-file check on any future exception.

Also: '.agents/skills/**' dep patterns replaced with the generator
(scripts/gen-skill-docs.ts) — .agents/ is gitignored, so those patterns
could NEVER match a git diff and review-template edits silently stopped
selecting codex/gemini tests; the codex/gemini local touchfile maps are now
DERIVED from the canonical map (loud throw if a key vanishes) instead of
hand-forked copies that had already drifted. ios-qa-e2e demoted gate ->
periodic: its gate declaration was never executable in CI (hardware
exclusion only applies at tier=periodic), so every Linux PR planned a
hollow shard.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix: routing journeys lose their answer key and end at the routing decision

The journey tests exist to catch skill-DESCRIPTION regressions (touchfiles:
*/SKILL.md.tmpl), but the fixture CLAUDE.md shipped an explicit
prompt->skill lookup table — with the answer key in context, a badly
regressed frontmatter description still routed correctly, so the tests
could not fail on the exact class they select for. The fixture now carries
only the generic invoke-skills nudge; the frontmatter carries the routing
load. Also capped all 10 journeys at maxTurns 2 / tools [Skill, Read]:
only the FIRST Skill call is asserted, so 5 turns of Read/Bash/Glob/Grep
was pure spend — roughly halves each journey's cost.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix: retire decided A/B experiments; vendor the pre-cut fixture; ban raw-SHA fixtures

Three one-shot decision experiments kept re-running weekly as N=1
stochastic comparisons — flaky by construction with near-zero remaining
information: skill-e2e-auq-repetition-cut-ab (its own header: gate "passed
pre-landing, approved 2026-08-25"), skill-e2e-preamble-script-ab ("demoted
post-Phase-3"), and opus-47's fanout arm-vs-arm (parA >= parB across two
SINGLE stochastic runs — a coin flip). Deleted, with their selection keys;
the SDK overlay-harness stays as the maintained instrument for the next
experiment, and opus-47 keeps its routing-precision cases.

verboseSkill() now reads the VENDORED test/fixtures/auq-pre-cut-...-SKILL.md
instead of `git show ab66193e^:...` — a branch-local ref that dies on
branch prune and already failed on shallow clones. New free tripwire
(test/git-ref-fixture-tripwire.test.ts) bans the raw-SHA fixture class
outright: quoted SHA:path rev-specs and gitRef-style hex defaults in the
test trees fail the suite with the vendor-instead instruction.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix: demote plan-ceo-review-expansion-energy to periodic

Opus generator + a subjective 2-axis >=4/5 LLM-judge threshold sat in the
MERGE-BLOCKING gate — the exact class its sibling posture tests were
demoted for, with a receipt (a +21-line preamble change once flipped the
score). CLAUDE.md's own tiering rule: Opus model test -> periodic. The
weekly lane keeps the regression signal; merges stop paying a judge-
temperament tax.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix: paid shards get per-shard TMPDIR + CHROMIUM_PROFILE isolation and a kill-path cleanup backstop

The free runner treats this isolation as MANDATORY (two concurrent shards
on one Chromium profile kill each other's browser; shared tmp
cross-contaminates) — the paid lane had none of it. Doubly load-bearing
here: a shard that hits its 30-min wall is group-SIGKILLed, so per-test
afterAll cleanup never runs; the rmSync backstop is the only thing keeping
wedged runs from accumulating full git-repo workspaces in the shared
tmpdir forever. This is the DAG prerequisite for raising EVALS_JOBS (next
commit) — more concurrency on shared state amplifies exactly the
shared-tree race class opus-47 exhibited.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat: paid-runner defaults 4x4 -> 8x2 — halve the local gate worst case

39 of 75 skill-e2e files hold exactly ONE test, so within-shard
concurrency was dead weight for most shards: 4 jobs x 4 concurrency
yielded only ~4-6 real in-flight sessions and a 13-wave local gate worst
case (~6.5h). 8 jobs x 2 gives ~10-13 in-flight — under the
documented-safe ~15 — and ~7 waves (~3.3h worst case). CI lanes keep
their explicit EVALS_JOBS env (2 per slice; 4 for gate-census); this
changes local defaults. Rollback trigger: sustained 429 storms in the WS1
telemetry across 2 PR cycles. test/eval-detach-timeout-floor.test.ts
recomputed green (the raise LOWERS the worst-case floor).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix: SHA-pin every action in the secrets-bearing eval lanes

evals.yml and evals-periodic.yml execute PR-authored code with three
provider API keys in env, yet rode mutable action tags (@v7/@v8/@v2/@v4)
— while quality-gate.yml, osv-scanner.yml, and dependency-review.yml
already model the SHA-pin pattern. All 30 uses sites across both lanes now
pin the exact commit (tag noted in a trailing comment); dependabot's
github-actions ecosystem keeps them fresh via PRs instead of silent tag
moves. Pulled forward from the plan's endgame on the CEO-review + outside-
voice agreement: supply-chain pins on secret lanes go first, not last.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix: sweep wave 3 — the execFileSync family gets timeouts (90 sites, 17 files)

The tripwire's regex covered spawnSync/execSync/Bun.spawnSync but not
execFileSync — an entire blocking sync-spawn API family that could
reintroduce the shard-wedge class undetected (ship review army). Same
mechanical recipe as waves 1-2: timeout: 30_000 default, 120_000 for slow
ops, shared wrappers fixed once, string-needle sites skipped with reasons.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* test: review-army + adversarial test hardening

- Tripwire scans execFileSync too (ceiling 8: two more grep-needle string
  exemptions); merge-introduced timeout-less spawnSync in
  question-preference-hook fixed — the tripwire caught a site that landed
  on main AFTER the sweep, on its first day.
- gstack-detach gains TWO watchdog kill regression tests: TERM-immune
  grandchild (the killpg-after-grace escalation) and the leader-dies
  variant (the pgid-at-spawn fix — the case the first test cannot see).
- eval-flake-rank gets its unit suite (final-attempt accounting, artifact
  exclusion, shard recursion, recency bound).
- Groupkill/startup-grace shim markers are per-run unique (pid-suffixed
  sleep durations): sibling Conductor worktrees run free suites with no
  machine lock, and fixed markers let one run pgrep/pkill the other's
  shims — a cross-run flake inside the anti-flake tests.
- flake-ledger test pins the project-scoped local default; stale empty
  section headers in touchfiles-data deleted (they invited entries under
  deliberately retired categories).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix: adversarial-review runtime fixes across the telemetry + kill paths

- session-runner: exit-labeling keys off 'exit', not 'close' — an orphan
  holding the pipes could relabel a REAL exit (auth failure) as
  'timeout_startup' availability noise; the kill path still always
  group-kills and cancels the reader (labeling and unblocking are separate
  concerns). Work phase arms on a flag, not firstResponseMs===0 (a same-ms
  first byte left the startup timer live all run). The CI startup grace is
  now a real FLOOR (Math.max), matching its name and pinning test.
- gstack-detach: pgid captured AT SPAWN (== child pid under
  start_new_session) — resolving it after the grace raised ESRCH once the
  leader died on SIGTERM, orphaning TERM-immune grandchildren forever.
- test-free-shards: ledger entries carry branch + git_sha (rev-parse split:
  '--abbrev-ref HEAD HEAD' printed the branch twice and recorded it as the
  sha); local ledger default is per-PROJECT, not the machine-global tmpdir.
- eval-flake-rank: per-LINE ledger parse (one torn JSONL line vanished the
  whole series), 60-day recency bound (transcript-bearing files are MBs),
  shared isFinalizedEvalResultFile predicate (the artifact-taxonomy rule
  lived in three places); eval-store exports the predicate and finalize
  stops computing flakyRetries twice; paid-shards cleanup uses async rm
  (a SIGKILLed shard's git-workspace teardown blocked every sibling's
  stream classification on the parent event loop).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix: CI trust-boundary + fail-closed repairs (adversarial findings)

- Token/exec separation restored: slices-report (runs PR-authored code:
  bun install + the reconcile runner) drops to contents:read; the PR
  comment moves to a NEW slices-comment job holding the write token with
  ZERO repo code — no checkout, no bun, only downloaded artifacts + jq/gh.
  $GITHUB_ENV/BASH_ENV persistence is job-scoped, so the split is the
  boundary. The matrix-era report job had this property; the consolidation
  had regressed it. Pinned by the wiring test.
- Reconcile exit captured via PIPESTATUS[0] in BOTH lanes: GitHub's default
  run-step shell has no pipefail, so `$?` after `| tee` was tee's exit —
  the fail-closed gate was silently fail-open. Wiring test pins it.
- PR comment: final-attempt accounting restored the dropped COST
  accumulation (the dial read $0 forever), flaky passes render as the
  warning they are (never as failures), and a malformed tests[] artifact
  skips that file instead of aborting the whole comment under bash -e.
- Remaining mutable action tags pinned (free-tests upload-artifact,
  ci-image checkout/docker trio — the image publisher holds packages:write
  and feeds the secret-bearing lanes). restore-deps fallback installs
  --frozen-lockfile; register-gstack-skills validates skill names before
  its rm -rf.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* chore: bump version and changelog (v1.77.0.0)

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs: update project documentation for v1.77.0.0

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs: cross-model doc-review fixes — flake-ledger env knobs, CI retry-on note, stale version comment

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs: correct CHANGELOG receipt numbers to measured values

Gate census keys: 78 -> 77 (bun-imported E2E_TIERS count). Sweep receipt:
586 sites/176 files -> 499 sites/146 files, measured by running this
branch's spawnsync-timeout-tripwire against origin/main (exit 1, 499
violations across 146 unique files; green on this branch).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix: slices-comment creates the PR comment via REST — the write-token job has no git context

The token/exec split gives slices-comment NO checkout by design, and gh's
pr-comment subcommand resolves the repo FROM git — it died with 'not a git
repository' on PR #2746's first run (the update-existing PATCH path was
already explicit-repo REST and worked). Create now posts through
gh api repos/.../issues/N/comments, and the wiring test pins that no
git-context-requiring comment call can creep back into the job.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix: startup-grace probes clear CI for local semantics; new probe pins the floor clamp

The two shim probes pass explicit 2s/4s graces, but in CI the runner clamps
any explicit grace up to the 300s floor (deliberate adversarial-review fix),
so 'silent API killed at the grace' died at the 30s work cap instead of 2s —
a deterministic red on every CI run, green locally. The probes now pin LOCAL
semantics with CI cleared (same save/restore pattern as their PATH shim),
and a fourth probe pins the clamp itself: CI=1 + 2s grace + 6s timeout must
kill at the 6s cap, still in the startup phase — proof an explicit low grace
cannot bypass the floor.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-31 08:55:30 -07:00

1617 lines
75 KiB
TypeScript
Executable File
Raw Permalink Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
#!/usr/bin/env bun
/**
* test-free-shards — enumerate, shard, curate, and run the free test suite.
*
* Four jobs:
* 1. Enumeration. Walk `browse/test/`, `test/`, `make-pdf/test/` and return
* every `*.test.{ts,tsx,js,jsx,mjs,cjs}` that isn't a paid-eval test.
* 2. Sharding. Stable-hash assign each test to one of N shards. Used by CI
* to parallelize the free suite when needed.
* 3. Curation (Windows-safe filter). Scan each test's content for POSIX-only
* patterns (`/bin/bash`, `sh -c`, raw `/tmp/`, `chmod`, `xargs`). Files
* that match are excluded from the Windows-safe subset — they would fail
* on `windows-latest` no matter how the runner shards them.
* 4. Execution. Spawn `bun test` children and refuse to trust their exit
* code alone: every byte of output is classified through
* scripts/test-strict-output.ts, so a child that exits 0 without bun's
* terminal summary (a mid-suite process.exit truncation), with `(fail)`
* result lines, or with fewer files run than planned is a FAILURE. An
* external wall-clock timeout SIGKILLs the child's process group and
* reports the shard as timed-out — distinct from failed.
*
* Execution strategy (decision ledger V3/D6 — evaluate the Bun built-in
* first; probed 2026-08 on Bun 1.3.13):
* - Full-suite runs (`bun test` via package.json, `bun run test:free`) use
* N CONCURRENT SHARD PROCESSES, serial within each (the paid runner's
* model). A single `--parallel` invocation was probed and initially
* adopted, then abandoned: three distinct Bun 1.3.13 worker pathologies
* (segfault + crash-retry wedge, skipped-file hooks stalling a worker,
* spawn-heavy files hanging under load) each stalled the whole
* invocation, while process shards isolate any wedge to its own shard.
* Original --parallel probe results, kept for the record: it
* showed --parallel (a) prints the standard `Ran N tests across M files`
* terminal summary, (b) exits non-zero when any file fails, (c) runs each
* file in its own worker process (distinct pids, no shared globals), and
* (d) converts a mid-suite process.exit(0) — which silently truncates a
* serial run at exit 0 — into a per-file `(crashed: exited)` failure with
* a complete summary and exit 1. Strictly SAFER than the serial path and
* ~2x faster on a 6-file probe (0.22s -> 0.11s wall, 280% CPU); the win
* grows with suite size since the serial suite measured 454s.
* - CI-matrix runs (`--shards M --shard i`) keep the hash-partitioned
* one-child-per-shard path. Cross-runner partitioning must be
* deterministic and per-file stable, so bun's own `--shard=M/N`
* (round-robin over sorted paths — every assignment shifts when a file
* lands) is not used, and there are no static per-file weight lists.
* Shard indices are STABLE: assignFilesToShards never renumbers on
* occupancy, and an empty shard is a fast no-op success.
*
* Adapted from the McGluut/gstack fork's test-free-shards.ts (190 LOC). The
* Windows-safe filter is upstream-original — codex flagged that sharding alone
* doesn't fix POSIX-bound tests, so we curate the subset that actually runs
* on the windows-latest CI job.
*
* Output contract (v1.66): the full child stream ALWAYS lands in a per-run
* log file under os.tmpdir() (path printed once at start and again in the
* epilogue). The console is quiet by default — only the runner's own
* [test:free] lines, `(fail)` result lines, bun error/crash markers
* (`error:`, `panic:`, `crashed`, `Unhandled error`), and the terminal
* `Ran N tests across M files` summary reach it; `--verbose` restores full
* forwarding. After every run a stable epilogue names the failing tests
* (attributed to files via bun's `path/to/file.test.ts:` chunk headers),
* crashed+retried workers, and — on a wall-timeout kill — the wedge-suspect
* files. The strict classifier consumes the FULL stream regardless of what
* the console shows.
*
* Exit codes: 0 pass, 1 fail, 124 wall-clock timeout.
*
* Usage:
* bun run scripts/test-free-shards.ts # full suite, N concurrent shard processes
* bun run scripts/test-free-shards.ts --list # show all
* bun run scripts/test-free-shards.ts --windows-only --list # show curated
* bun run scripts/test-free-shards.ts --windows-only # run curated
* bun run scripts/test-free-shards.ts --shards 4 --shard 1 # one shard (CI matrix)
* bun run scripts/test-free-shards.ts --wall-timeout 600 # override the kill deadline
* bun run scripts/test-free-shards.ts --verbose # forward the full child stream
*/
import * as fs from 'fs';
import * as os from 'os';
import * as path from 'path';
import { spawn, spawnSync } from 'child_process';
import { StringDecoder } from 'node:string_decoder';
import { isPaidTestFile } from '../test/helpers/paid-test-set';
import {
BunTestOutputClassifier,
exactTestFileSelectors,
installChildSignalForwarding,
isTerminationRequested,
killProcessGroup,
strictTestExitCode,
stripAnsiLine,
} from './test-strict-output';
const ROOT = path.resolve(import.meta.dir, '..');
// design/test was silently absent from BOTH the package.json test script and
// this list — design tests (including a teardown bomb) never ran in any CI
// or local free run. Keep the two lists in sync. This list is the single
// source of truth for free-suite roots: package.json's `test` script routes
// through this runner rather than passing its own directory globs.
export const TEST_ROOTS = [
'browse/test',
'test',
'make-pdf/test',
'design/test',
// v1.65 orphan wire-in (decision D3a): these ran under NO script or CI —
// written coverage that caught nothing. All were green on arrival.
'ios-qa/daemon/test',
'ios-qa/scripts',
'browser-skills',
] as const;
const TEST_FILE_REGEX = /\.test\.(?:[cm]?[jt]s|tsx|jsx)$/;
// POSIX-only patterns that indicate a test will fail on windows-latest no
// matter how the runner shards. Codex's v1.18.0.0 review flagged the first
// three as concrete examples in the existing free suite (test/ship-version-sync.test.ts:72,
// test/helpers/providers/claude.ts:22, package.json:12). We scan the test's
// own content here so the filter stays automatic as new tests land. The
// "Windows-incompatible APIs" patterns at the bottom were added after the
// first windows-free-tests CI run surfaced concrete failure modes.
const WINDOWS_FRAGILE_PATTERNS: Array<{ pattern: RegExp; reason: string }> = [
// Hardcoded POSIX shells / commands.
{ pattern: /['"`]\/bin\/(?:ba)?sh/, reason: 'hardcoded /bin/sh or /bin/bash' },
{ pattern: /spawnSync\(['"]sh['"],|spawn\(['"]sh['"],|exec\(['"]sh /, reason: 'spawn("sh", ...)' },
{ pattern: /['"]bash -c['"]|['"]sh -c['"]/, reason: 'bash -c / sh -c' },
{ pattern: /['"`]\/tmp\//, reason: 'raw /tmp/ path (use os.tmpdir())' },
{ pattern: /['"]chmod\b/, reason: 'chmod shell command' },
{ pattern: /['"]xargs\b/, reason: 'xargs pipeline' },
{ pattern: /\bwhich claude\b/, reason: 'which claude (use Bun.which)' },
// Windows-incompatible APIs.
{ pattern: /\.mode\s*&\s*0o[0-7]+/, reason: 'POSIX file mode bitmask (mode & 0o600 etc — Windows fakes mode bits)' },
{ pattern: /\.endsWith\(['"]\//, reason: 'hardcoded forward-slash path assertion (Windows uses \\\\)' },
{ pattern: /['"]\.\/[a-zA-Z][^"']*['"]\)\s*\.\s*toBe\(true\)/, reason: 'forward-slash path comparison' },
// Tests that spawn a bash shebang script in bin/ via spawnSync. Git Bash on
// Windows can run `bash /path/to/script` but spawnSync(scriptPath, ...)
// tries to execute the file directly via CreateProcess, which fails on the
// shebang. The pattern matches `, 'bin'` as a path-join argument (closing
// OR followed by another segment), which catches:
// - path.join(ROOT, 'bin', 'script-name') — typical
// - join(import.meta.dir, '..', 'bin', 'name') — destructured (diff-scope)
// - path.join(ROOT, 'bin') — bare BIN constant (brain-sync)
{ pattern: /,\s*['"]bin['"]\s*[,)]|['"]\.?\/?bin\/[a-z][\w-]+['"]/, reason: 'spawns bin/ shebang script (Windows CreateProcess does not parse shebangs)' },
// Tests that launch a real Playwright browser. The windows-free-tests CI job
// runs a curated subset that intentionally does NOT install Chromium —
// browser bring-up on Windows is a separate concern (see PR #1238). Tests
// matching `await foo.launch(` need Chromium and fail with "Executable
// doesn't exist" on the runner.
{ pattern: /await\s+\w+\.launch\(/, reason: 'launches Playwright browser (Chromium not installed in windows-free CI)' },
// Tests that spawn the browse server as a subprocess via `bun run server.ts`.
// The Bun → server.ts → Playwright path is the same one that doesn't work
// on Windows (PR #1238 windows-pty-bun-pty-fix). Tests typically set
// BROWSE_HEADLESS_SKIP=1 to skip the browser launch but still need a working
// server, which they don't get on Windows.
{ pattern: /BROWSE_HEADLESS_SKIP|spawn\(\[['"]bun['"],\s*['"]run['"]/, reason: 'spawns the browse server subprocess (Bun-driven path is Windows-broken)' },
];
// Explicit known-Windows-incompatible test files that don't fit a regex
// pattern. Listed here with the precise reason. Prefer adding a pattern above
// when possible; this list is for environment-/runtime-specific tests where
// the failure mode is structural rather than detectable via source-file scan.
export const KNOWN_WINDOWS_INCOMPATIBLE: Array<{ file: string; reason: string }> = [
{
file: 'test/host-config.test.ts',
reason: 'asserts "claude" binary on PATH (only true when running inside Claude Code, not on bare CI runner)',
},
{
file: 'browse/test/findport.test.ts',
reason: 'asserts Bun.serve.stop() is fire-and-forget — Bun behavior differs on Windows for this polyfill',
},
// First full run of the expanded lane (v1.66, 13 → ~258 files) surfaced
// seven POSIX-bound files the content patterns cannot see (their
// POSIX-ness is what they TEST, or arrives via a variable). Receipts:
// PR #2593 windows-free-tests run 31918591602.
{
file: 'test/codex-under-codex-detection.test.ts',
reason: 'drives the rendered preflight bash under a hardcoded POSIX PATH (/usr/bin:/bin) — bash is unreachable through that PATH on Windows, so every case sees empty output (v1.67 windows lane run 95234224148)',
},
{
file: 'test/regression-pr1169-build-app-sed.test.ts',
reason: 'tests sed escape sequences in build-app.sh — sed/bash are the subject under test',
},
{
file: 'test/setup-conductor-worktree.test.ts',
reason: 'tests ln -snf symlink semantics in the setup script — POSIX ln is the subject under test',
},
{
file: 'test/artifacts-init-migration.test.ts',
reason: 'runs a bash migration script + jq against a scaffolded git state — POSIX toolchain paths break under cmd spawn',
},
{
file: 'test/gstack-decision-semantic.test.ts',
reason: 'installs a fake gbrain SHEBANG SHIM on PATH; Windows spawn cannot exec shebang scripts',
},
{
file: 'test/question-log-hook.test.ts',
reason: 'spawns the PostToolUse hook script (bash shebang) directly; Windows spawn cannot exec it',
},
{
file: 'browse/test/browser-skills-e2e.test.ts',
reason: 'asserts forward-slash tier paths (<repo>/browser-skills/) that resolve with backslashes on Windows',
},
{
file: 'design/test/variants-retry-after.test.ts',
reason: 'wall-clock retry-timing assertions — flaky on the slow windows-latest runner even with widened bounds',
},
// Round-2 census (PR #2593 run 31919227507) after the first seven:
{
file: 'test/skill-census.test.ts',
reason: 'census walk throws at module load on Windows (skill-census.ts:63) — the skills-tree symlink layout needs Developer Mode that CI runners lack',
},
{
file: 'browse/test/browser-manager-unit.test.ts',
reason: 'wedges the shard to its wall deadline on windows-latest (in-flight at kill); needs a Windows repro to diagnose — macOS + Linux lanes cover the file',
},
// Round-3 census (PR #2593 run 31919871680): the round-2 wedge had been
// TRUNCATING its shard, so these seven only surfaced once shard 2 completed.
// All the same POSIX-environment classes: PID/cmdline identity probing,
// bash scripts as the subject under test, env-scrubbed child spawns.
{
file: 'browse/test/server-embedder-terminal-port.test.ts',
reason: 'identity-based terminal-agent kill probes PID/cmdline with POSIX semantics; teardown asserts fail on windows-latest',
},
{
file: 'design/test/daemon-discovery.test.ts',
reason: 'verifyIdentity matches a spawned daemon via /proc-style cmdline probing — POSIX identity semantics',
},
{
file: 'test/context-save-hardening.test.ts',
reason: 'bash context-save/migration scripts (HOME-unset semantics, random-suffix path) are the subject under test',
},
{
file: 'test/eval-list-cli.test.ts',
reason: 'spawns the eval:list CLI via bun with a constructed env — bun resolution fails under Windows spawn',
},
{
file: 'test/memory-cache-injection.test.ts',
reason: 'exercises hook/deny-enforcement shell scripts — POSIX toolchain is the subject under test',
},
{
file: 'test/migrations-v1.65.0.0.test.ts',
reason: 'bash migration script (bunx re-fetch, .done markers) is the subject under test',
},
{
file: 'test/question-preference-hook.test.ts',
reason: 'spawns the PreToolUse preference hook (shebang script) directly; Windows spawn cannot exec it',
},
// Round-4 census (PR #2593 run 31920052810): unhandled errors with no
// (fail) lines — attributed statically (the lane had no log artifact yet).
{
file: 'browse/test/browser-skill-commands.test.ts',
reason: 'spawnSkill spawns bun with a constructed env — bun resolution fails under Windows spawn (unhandled, no (fail) line)',
},
{
file: 'browse/test/security-audit-r2.test.ts',
reason: 'symlink-attack fixtures (evil-link) need Developer Mode CI runners lack; expect(toThrow) fires unhandled on Windows',
},
];
// Force-include overrides: files a WINDOWS_FRAGILE_PATTERNS regex excludes for
// a reason that does not actually apply to them. Each entry documents WHY the
// pattern hit is a false positive — the point of these files is Windows
// coverage, so auto-excluding them defeats the regression tests they carry.
const KNOWN_WINDOWS_SAFE: Array<{ file: string; reason: string }> = [
{
file: 'test/setup-windows-rerun-refresh.test.ts',
// Trips the "spawns bin/ shebang script" pattern via path.join(..., 'bin',
// 'tool.sh') fixture paths, but every spawn goes through spawnSync('bash',
// ['-c', ...]) — Git Bash executes it fine on windows-latest. This file IS
// the #2444 Windows regression coverage (IS_WINDOWS=1 copy-refresh path);
// excluding it here would keep the bug class unexercised on the one
// platform it bites.
reason: 'bin/ hits are fixture path segments; spawns bash explicitly — the IS_WINDOWS=1 refresh path must run on windows-latest',
},
{
file: 'test/uninstall-windows-copies.test.ts',
// Trips the "spawns bin/ shebang script" pattern via the
// path.join(ROOT, 'bin', 'gstack-uninstall') constant, but the script is
// always spawned through spawnSync('bash', [UNINSTALL, ...]). This file
// carries the #2563 Windows real-dir-copy uninstall coverage — the bug
// ONLY reproduces on the copy install shape windows-latest exercises.
// The symlink-shape describe block self-skips on win32.
reason: 'bin/ hit is a bash-spawned script path; #2563 real-dir uninstall coverage must run on windows-latest',
},
{
file: 'browse/test/file-permissions.test.ts',
// Trips the POSIX-mode-bitmask pattern, but every `mode & 0o777` assertion
// is platform-guarded: win32-only tests return early, POSIX-only tests
// guard the bitmask behind `process.platform !== 'win32'`, and the
// symlink-skip regression test both wraps symlinkSync in try/catch
// (runners without Developer Mode can't create symlinks) and guards its
// bitmask — on win32 it asserts behavior (warns, skips, doesn't throw,
// target stays usable), never fake Windows mode bits (dirs stat 0o777
// there, so a 0o755 expectation fails on runner semantics, not our code).
// This file carries the win32-only icacls-by-SID regression tests, which
// can ONLY execute on windows-latest — excluding it here means the
// machine-account ACL lockout regression is never exercised on the one
// platform it bricks.
reason: 'every mode-bitmask assertion is guarded off win32 (behavior asserted instead); win32-only ACL regression tests must run on windows-latest',
},
{
file: 'browse/test/terminal-agent-owner-watchdog.test.ts',
// Trips the spawn(['bun','run',...]) pattern, whose reason is the
// Playwright-bound browse server. This test spawns terminal-agent.ts,
// which imports only fs/path/crypto + local helpers (no Playwright, no
// PTY at module scope) and boots under Bun on Windows — the owner-PID
// orphan leak it pins was reported on Windows (#2019).
reason: 'spawns terminal-agent (no Playwright), not the browse server; owner-orphan leak is a Windows defect',
},
];
export const DEFAULT_SHARD_COUNT = 20;
// Per-test timeout passed to `bun test --timeout`. 30s matches what
// package.json's `test` script used before it was repointed at this runner —
// the runner is now the single owner of that semantic.
export const FREE_TEST_TIMEOUT_MS = 30_000;
// External wall-clock deadline per spawned child (whole shard or the single
// full-suite --parallel invocation). A wedged child — a spinning main thread
// no in-process --timeout timer can interrupt — is SIGKILLed at the group
// level and reported 'timed-out', distinct from 'failed'.
// ~3.5x the observed full-suite wall (~100-160s). A wedged run should be
// killed-and-diagnosed (the epilogue prints the in-flight suspects) in
// minutes, not sat out — 15min of silence was pure diagnosis latency.
// Override per run with --wall-timeout <secs>.
export const DEFAULT_WALL_TIMEOUT_MS = 6 * 60_000;
/**
* Full-suite shards scale their wall deadline with shard size:
* max(DEFAULT_WALL_TIMEOUT_MS, files × PER_FILE_WALL_MS). The 6-min floor
* keeps wedge diagnosis fast on a typical ~70-file local shard, while a
* low-core machine (jobs=1 → the whole suite in one shard) or the Windows
* lane (~130 files/shard) gets proportional headroom instead of a false
* timed-out kill of a healthy run. Explicit --wall-timeout disables scaling.
*/
export const PER_FILE_WALL_MS = 5_000;
export function wallTimeoutForShard(fileCount: number, baseMs = DEFAULT_WALL_TIMEOUT_MS): number {
return Math.max(baseMs, fileCount * PER_FILE_WALL_MS);
}
/**
* Wall for a duration-packed shard. The count heuristic above assumes count
* approximates cost; LPT packing breaks that BY DESIGN (a shard may hold six
* slow Playwright files), so packed shards get max(base, predicted x 3) —
* generous against seed drift, still bounded.
*/
export function wallTimeoutForPackedShard(predictedMs: number, baseMs = DEFAULT_WALL_TIMEOUT_MS, fileCount = 0): number {
// Predictions transfer badly across machines: the committed duration seed
// is recorded on fast CI, and a syscall-supervised sandbox replays those
// files 2-4x slower (observed: a 253-file shard predicted ~242s wall-killed
// at its 725s predicted-x3 wall while genuinely still progressing). The
// packed wall may therefore be LOOSER than the count heuristic, never
// tighter — it keeps the per-file floor the runner has always guaranteed.
return Math.max(baseMs, Math.ceil(predictedMs * 3), fileCount * PER_FILE_WALL_MS);
}
/**
* Full-suite parallelism: leave RESERVED_CPUS cores for the parent runner +
* OS, cap at MAX_FULL_SUITE_JOBS — beyond ~6 concurrent bun processes the
* playwright-heavy shards contend on browser launches instead of finishing
* sooner (measured on an M-series dev box).
*
* GSTACK_FREE_JOBS overrides the computed count (the free runner's analogue
* of the paid runner's EVALS_JOBS). Exists for syscall-supervised sandboxes:
* on Vercel sandboxes, PID 1 (sandbox-init) installs a seccomp filter whose
* user-space supervisor saturates under ~6 concurrent bun+playwright shards
* and starts returning EACCES from plain file syscalls (measured: 200/200
* `git init` probes in fresh mktemp dirs fail with
* "Cannot access work tree: Permission denied" while the suite runs, 0/200
* when idle — access(dir, X_OK) = EACCES under strace). Fewer shards keep
* the supervisor inside its budget. Not clamped by MAX_FULL_SUITE_JOBS so a
* beefy box can also raise it deliberately.
*/
export const MAX_FULL_SUITE_JOBS = 6;
export const RESERVED_CPUS = 2;
export function fullSuiteJobs(): number {
const raw = process.env.GSTACK_FREE_JOBS;
if (raw !== undefined && raw !== '') {
// Strict digits-only: parseInt would silently truncate "2abc" -> 2 and
// "3.7" -> 3, defeating the loud-failure contract the error text claims.
if (!/^\d+$/.test(raw.trim()) || Number.parseInt(raw, 10) <= 0) {
throw new Error(`GSTACK_FREE_JOBS must be a positive integer, got: ${raw}`);
}
return Number.parseInt(raw, 10);
}
return Math.max(1, Math.min(MAX_FULL_SUITE_JOBS, os.cpus().length - RESERVED_CPUS));
}
/**
* Files that crash or wedge Bun's --parallel WORKERS but run fine in a plain
* serial process. Full-suite mode now uses shard PROCESSES (no workers), so
* this list is inert placement-wise — retained as the paper trail of why the
* one-invocation --parallel strategy was abandoned, and as the exclusion list
* should anyone re-attempt it on a newer Bun.
*/
export const WORKER_HOSTILE: Record<string, string> = {
'browse/test/security-live-playwright.test.ts':
'Bun 1.3.13 segfaults running this file in a --parallel worker ("panic: '
+ 'Segmentation fault ... a bug in Bun"), and the crashed-worker retry then '
+ 'wedges the whole invocation past the wall clock. Passes serially.',
};
/**
* TREE-SERIAL files: run in ONE serial shard AFTER the parallel shards.
* EMPTY since the 2026-08 dissolution — kept as a mechanism, not a museum:
* a test that must regenerate shared repo artifacts IN PLACE (and cannot
* render into an out-dir instead) earns an entry here with a reason, and
* the runner will serialize it again.
*
* How it emptied: gen-skill-docs gained a main() guard (imports stopped
* regenerating 71 files at load) and --out-dir grew to every host, so all
* eight mutators now render into mkdtemps — the live tree is never written
* by the suite (pinned by gen-skill-docs-import-purity + each migrated
* file's own porcelain/mtime assertions). With zero mutators, the four
* ratchet READERS (parity caps, size budgets, carve parity/ordering) get a
* quiet tree by construction in any shard, so they rejoined the parallel
* phase — the ~35-40s serial tail on every full-suite run is gone.
* Keys are pinned against the live file census by test-free-shards.test.ts —
* a renamed file fails the suite instead of silently dropping serialization.
*/
export const TREE_MUTATING: Record<string, string> = {};
export function normalizeRelativePath(filePath: string): string {
return filePath.replace(/\\/g, '/');
}
export function isFreeTestFile(relativePath: string): boolean {
const normalized = normalizeRelativePath(relativePath);
if (!TEST_FILE_REGEX.test(normalized)) return false;
return !isPaidTestFile(normalized);
}
/**
* Returns the first POSIX-only pattern hit in the file, or null if Windows-safe.
*/
export function detectWindowsFragility(absolutePath: string): { reason: string } | null {
let content: string;
try {
content = fs.readFileSync(absolutePath, 'utf-8');
} catch {
return null;
}
for (const { pattern, reason } of WINDOWS_FRAGILE_PATTERNS) {
if (pattern.test(content)) return { reason };
}
return null;
}
function walkTestFiles(dirPath: string): string[] {
const entries = fs.readdirSync(dirPath, { withFileTypes: true });
const files: string[] = [];
for (const entry of entries) {
const fullPath = path.join(dirPath, entry.name);
if (entry.isDirectory()) {
files.push(...walkTestFiles(fullPath));
continue;
}
if (TEST_FILE_REGEX.test(entry.name)) {
files.push(fullPath);
}
}
return files;
}
export function collectFreeTestFiles(rootDir = ROOT): string[] {
const discovered = new Set<string>();
for (const testRoot of TEST_ROOTS) {
const absoluteRoot = path.join(rootDir, testRoot);
if (!fs.existsSync(absoluteRoot)) continue;
for (const fullPath of walkTestFiles(absoluteRoot)) {
const relativePath = normalizeRelativePath(path.relative(rootDir, fullPath));
if (isFreeTestFile(relativePath)) {
discovered.add(relativePath);
}
}
}
return [...discovered].sort();
}
export interface CurationResult {
safe: string[];
excluded: Array<{ file: string; reason: string }>;
}
export function curateWindowsSafe(files: string[], rootDir = ROOT): CurationResult {
const safe: string[] = [];
const excluded: Array<{ file: string; reason: string }> = [];
const knownBad = new Map(KNOWN_WINDOWS_INCOMPATIBLE.map((e) => [e.file, e.reason]));
const knownSafe = new Set(KNOWN_WINDOWS_SAFE.map((e) => e.file));
for (const relativePath of files) {
const knownReason = knownBad.get(relativePath);
if (knownReason) {
excluded.push({ file: relativePath, reason: knownReason });
continue;
}
if (knownSafe.has(relativePath)) {
safe.push(relativePath);
continue;
}
const absolute = path.join(rootDir, relativePath);
const fragility = detectWindowsFragility(absolute);
if (fragility) {
excluded.push({ file: relativePath, reason: fragility.reason });
} else {
safe.push(relativePath);
}
}
return { safe, excluded };
}
export function stableHash(input: string): number {
let hash = 0x811c9dc5;
for (let index = 0; index < input.length; index += 1) {
hash ^= input.charCodeAt(index);
hash = Math.imul(hash, 0x01000193);
}
return hash >>> 0;
}
/**
* Hash-partition files across EXACTLY shardCount shards. Empty shards are
* preserved: a file's shard index is a pure function of its own path and the
* shard count, never of which other files happen to exist. A CI matrix keys
* runners off the index, so filtering empty shards (the old behavior) would
* renumber every later shard whenever occupancy shifted — runner 3 silently
* running shard 4's files. An empty shard is instead a fast no-op success at
* run time.
*/
export function assignFilesToShards(files: string[], shardCount: number): string[][] {
if (!Number.isInteger(shardCount) || shardCount <= 0) {
throw new Error(`Shard count must be a positive integer. Received: ${shardCount}`);
}
const shards = Array.from({ length: shardCount }, () => [] as string[]);
for (const file of files) {
const shardIndex = stableHash(file) % shardCount;
shards[shardIndex].push(file);
}
return shards.map(filesInShard => filesInShard.sort());
}
// ─── Duration-aware packing (full-suite path ONLY) ─────────────────────────
// Hash sharding balances file COUNTS (~1.15x spread) but not cost: the 15
// Playwright-launching files land 4/3/4/1/2/1 across 6 shards, giving a
// measured 28s97s shard spread and ~40s of idle tail on every run. LPT
// packing over recorded per-file durations reclaims most of it. The `--shard`
// CI-matrix path is deliberately untouched — its contract is stable indices
// via assignFilesToShards/stableHash (empty shards no-op; see above).
//
// One store, no overlay: durations come from the committed seed
// (scripts/free-test-durations.json), refreshed occasionally via
// `--record-durations` (each file timed in its own child — exact, and immune
// to bun's stream buffering, where silent passers print no header to
// timestamp). GSTACK_FREE_TEST_DURATIONS overrides the path for experiments.
// The seed is a HINT, not a contract: missing file → hash-shard fallback;
// unknown file → 75th-percentile pessimism (placed early by LPT, bounding
// tail risk). Successor note: bun ≥1.3.14 ships native --timings/--shard LPT
// scheduling — when the repo unpins 1.3.13, this packer is the code to
// replace (keep it swappable).
export const FREE_TEST_DURATIONS_FILE = 'scripts/free-test-durations.json';
export function loadFreeTestDurations(rootDir = ROOT): Record<string, number> | null {
const file = process.env.GSTACK_FREE_TEST_DURATIONS
?? path.join(rootDir, FREE_TEST_DURATIONS_FILE);
let raw: string;
try {
raw = fs.readFileSync(file, 'utf-8');
} catch {
return null; // no seed — hash sharding, silently (fresh checkouts are normal)
}
try {
const parsed = JSON.parse(raw) as { durations?: Record<string, unknown> };
const entries = Object.entries(parsed.durations ?? {})
.filter((entry): entry is [string, number] =>
typeof entry[1] === 'number' && Number.isFinite(entry[1]) && entry[1] >= 0);
if (entries.length === 0) return null;
return Object.fromEntries(entries);
} catch (error) {
// A corrupt seed (bad merge) must cost a warning, never the suite.
console.error(`[test:free] WARNING: corrupt durations seed ${file} (${(error as Error).message}) — falling back to hash sharding`);
return null;
}
}
export interface PackedShards {
shards: string[][];
/** Predicted total per shard, aligned with `shards` — feeds walls + logs. */
predictedMs: number[];
}
/**
* Longest-processing-time-first bin packing: files sorted by predicted
* duration (desc, path-stable tiebreak) each go to the currently-lightest
* shard. Deterministic for a given (files, shardCount, durations).
*/
export function packShardsByDuration(
files: string[],
shardCount: number,
durations: Record<string, number>,
): PackedShards {
if (!Number.isInteger(shardCount) || shardCount <= 0) {
throw new Error(`Shard count must be a positive integer. Received: ${shardCount}`);
}
const known = files
.map((f) => durations[normalizeRelativePath(f)])
.filter((v): v is number => typeof v === 'number')
.sort((a, b) => a - b);
// Unknown files get the 75th percentile of known durations: pessimistic, so
// LPT places them early and a surprise long-runner can't recreate the tail.
const fallback = known.length > 0 ? known[Math.min(known.length - 1, Math.floor(known.length * 0.75))] : 1;
const predicted = (f: string): number => durations[normalizeRelativePath(f)] ?? fallback;
const ordered = [...files].sort((a, b) => predicted(b) - predicted(a) || (a < b ? -1 : 1));
const shards = Array.from({ length: shardCount }, () => [] as string[]);
const loads = new Array<number>(shardCount).fill(0);
for (const file of ordered) {
let lightest = 0;
for (let i = 1; i < shardCount; i += 1) {
if (loads[i] < loads[lightest]) lightest = i;
}
shards[lightest].push(file);
loads[lightest] += predicted(file);
}
return { shards: shards.map((s) => s.sort()), predictedMs: loads };
}
export interface BuildShardArgsOptions {
/**
* Pass bun's --parallel (worker-per-file, implies --isolate). No production
* caller today — full-suite mode uses N shard PROCESSES after the worker
* pathologies documented in main(); retained for a future re-attempt on a
* newer Bun (see WORKER_HOSTILE).
*/
parallel?: boolean;
rootDir?: string;
}
export function buildShardArgs(files: string[], options: BuildShardArgsOptions = {}): string[] {
// Exact absolute selectors: bun treats positional test paths as substring
// filters, so a relative `test/x.test.ts` would ALSO select
// `browse/test/x.test.ts` — shard bleed that double-runs files.
const selectors = exactTestFileSelectors(files, options.rootDir ?? ROOT);
const args = ['test', ...selectors, `--timeout=${FREE_TEST_TIMEOUT_MS}`];
if (options.parallel) args.push('--parallel');
else args.push('--max-concurrency=1');
return args;
}
type CliOptions = {
dryRun: boolean;
listOnly: boolean;
recordDurations: boolean;
windowsOnly: boolean;
verbose: boolean;
shardCount: number;
shardIndex: number | null;
wallTimeoutMs: number;
/** True when --wall-timeout was passed explicitly; full-suite mode only auto-scales the default. */
wallTimeoutExplicit: boolean;
};
function parseCliOptions(argv: string[]): CliOptions {
let dryRun = false;
let listOnly = false;
let recordDurations = false;
let windowsOnly = false;
let verbose = false;
let shardCount = DEFAULT_SHARD_COUNT;
let shardIndex: number | null = null;
let wallTimeoutMs = DEFAULT_WALL_TIMEOUT_MS;
let wallTimeoutExplicit = false;
for (let index = 0; index < argv.length; index += 1) {
const arg = argv[index];
if (arg === '--dry-run') { dryRun = true; continue; }
if (arg === '--list') { listOnly = true; continue; }
if (arg === '--record-durations') { recordDurations = true; continue; }
if (arg === '--windows-only') { windowsOnly = true; continue; }
if (arg === '--verbose') { verbose = true; continue; }
if (arg === '--shards') {
const value = argv[index + 1];
if (!value) throw new Error('Missing value for --shards');
shardCount = Number.parseInt(value, 10);
index += 1;
continue;
}
if (arg === '--shard') {
const value = argv[index + 1];
if (!value) throw new Error('Missing value for --shard');
shardIndex = Number.parseInt(value, 10);
index += 1;
continue;
}
if (arg === '--wall-timeout') {
const value = Number.parseInt(argv[index + 1] ?? '', 10);
if (!Number.isInteger(value) || value <= 0) throw new Error('--wall-timeout needs a positive integer (seconds)');
wallTimeoutMs = value * 1000;
wallTimeoutExplicit = true;
index += 1;
continue;
}
throw new Error(`Unknown argument: ${arg}`);
}
return { dryRun, listOnly, recordDurations, windowsOnly, verbose, shardCount, shardIndex, wallTimeoutMs, wallTimeoutExplicit };
}
function formatShardSummary(shards: string[][]): string[] {
return shards.map((files, index) => {
const preview = files.slice(0, 3).join(', ');
const suffix = files.length > 3 ? ', ...' : '';
return `Shard ${index + 1}/${shards.length}: ${files.length} files${preview ? ` -> ${preview}${suffix}` : ''}`;
});
}
/**
* True when a shard's output shows the run ended WITHOUT bun's final summary
* ("Ran N tests across ..."). A process.exit() fired mid-suite skips the
* summary AND hands back whatever code the caller passed — historically 0,
* which made a truncated shard indistinguishable from a green one. Exit code
* alone is therefore not evidence of completion; the summary line is.
*
* The runner itself now enforces this (and more) through
* scripts/test-strict-output.ts inside runFreeShard; this predicate remains
* the minimal documented primitive that test/exit-propagation.test.ts drives
* with genuine truncated and genuine complete bun runs.
*/
export function shardRunLooksTruncated(status: number | null, output: string): boolean {
if (status !== 0) return false; // already failing — not the silent case
return !/Ran \d+ tests? across \d+ files?/.test(output);
}
// ---------------------------------------------------------------------------
// Output contract: console filtering + per-file failure attribution.
//
// Bun groups each file's output under a `path/to/file.test.ts:` header line
// (cwd-relative, sometimes ../-prefixed through a symlinked cwd). The
// reporter tracks the current header while consuming the stream, attributes
// `(fail)` lines and crash markers to files, and decides which lines reach
// the console in the default quiet mode. All matching happens on
// ANSI-stripped lines — colored `(fail)` lines defeated a prior grep.
// ---------------------------------------------------------------------------
const TEST_PATH_SOURCE = String.raw`\.test\.(?:[cm]?[jt]s|tsx|jsx)`;
/** A file chunk header: the path bun printed, terminated by a bare colon. */
const FILE_HEADER_RE = new RegExp(`^(\\S.*${TEST_PATH_SOURCE}):$`);
/** Same shape strict-output classifies as failed-test, with the name captured. */
const FAIL_RESULT_CAPTURE_RE = /^\(fail\) (.+) \[\d+(?:\.\d+)?(?:ns|us|µs|ms|s)\]$/;
/** bun --parallel retries a crashed worker once: `<icon> crashed running <path>, retrying`. */
const CRASH_RETRY_RE = new RegExp(`crashed running (\\S*${TEST_PATH_SOURCE}), retrying`);
/** The give-up marker after the retry also crashes: `✗ <path> (crashed: exited)`. */
const CRASH_FINAL_RE = new RegExp(`(\\S*${TEST_PATH_SOURCE}) \\(crashed: [^)]+\\)`);
const TERMINAL_SUMMARY_CAPTURE_RE = /^Ran (\d+) tests? across (\d+) files?\. \[/;
/** Substrings that must reach the console even in the default quiet mode. */
const CONSOLE_ALWAYS_MARKERS = ['error:', 'panic:', 'Unhandled error', 'crashed'] as const;
export type StreamOrigin = 'stdout' | 'stderr';
export interface FreeRunFailure {
/** Planned relative path when attributable, else the raw header path, else null. */
file: string | null;
testName: string;
}
export interface FreeRunReport {
testsRan: number | null;
filesRan: number | null;
sawTerminalSummary: boolean;
/** Deduped `(fail)` lines in arrival order, attributed to the current file header. */
failures: FreeRunFailure[];
/** Files that crashed a worker (bun retries once; a second crash is final). Deduped. */
crashedFiles: string[];
/**
* "# Unhandled error between tests" markers, attributed to the chunk they
* appeared in. These fail the shard via the strict classifier but produce
* NO (fail) lines — without surfacing them here, the epilogue reads
* "FAIL — 0 failing test(s)" and the culprit is undiscoverable from CI
* output (first Windows lane run: a module-load throw in skill-census).
*/
unhandledErrors: Array<{ file: string | null }>;
/**
* Wedge-suspect heuristic for a wall-timeout kill: files whose header was
* seen but whose chunk never ENDED (chunk end = the next file's header, or
* a final crash marker) before the terminal summary — i.e. "started but
* never produced a result chunk end". Result lines deliberately do NOT end
* a chunk: a file that printed a fail and then wedged stays listed. Known
* limits of the approximation:
* - Serial (--shard CI path): bun streams live but prints a file's header
* lazily, on its first output line — a wedged file that printed ANY
* line is listed; a fully silent wedge is not.
* - Parallel (full-suite path): bun buffers a file's whole chunk until it
* COMPLETES, so a wedged file usually never prints a header (see
* filesWithNoOutput), and the LAST flushed chunk before the kill has no
* closing header, so one completed noisy file can be over-listed.
*/
inFlight: string[];
/** Planned files never observed in the stream (silent passers + never-flushed wedges). */
filesWithNoOutput: number;
}
interface FileProgress {
headerSeen: boolean;
/** The file's chunk ended: a later file's header arrived, or it crashed out. */
ended: boolean;
}
/**
* Incrementally consumes the child's stdout/stderr (chunk boundaries need not
* align to lines), attributing results to files and forwarding only
* always-visible lines to `forward` (omit `forward` for verbose/quiet modes —
* attribution still runs so the epilogue works in every mode).
*/
export class FreeRunReporter {
private readonly decoders: Record<StreamOrigin, StringDecoder> = {
stdout: new StringDecoder('utf8'),
stderr: new StringDecoder('utf8'),
};
private readonly pending: Record<StreamOrigin, string> = { stdout: '', stderr: '' };
private readonly plannedSet: Set<string>;
private readonly canonicalCache = new Map<string, string>();
private readonly progress = new Map<string, FileProgress>();
private readonly failureKeys = new Set<string>();
private readonly failures: FreeRunFailure[] = [];
private readonly crashed = new Set<string>();
private currentFile: string | null = null;
private inRecap = false;
private readonly unhandled: Array<{ file: string | null }> = [];
private testsRan: number | null = null;
private filesRan: number | null = null;
private sawSummary = false;
constructor(
private readonly plannedFiles: string[],
private readonly forward?: (text: string, origin: StreamOrigin) => void,
) {
this.plannedSet = new Set(plannedFiles.map(normalizeRelativePath));
}
write(chunk: Uint8Array | string, origin: StreamOrigin): void {
this.pending[origin] += typeof chunk === 'string'
? chunk
: this.decoders[origin].write(Buffer.from(chunk));
let newline = this.pending[origin].indexOf('\n');
while (newline !== -1) {
this.handleLine(this.pending[origin].slice(0, newline), origin);
this.pending[origin] = this.pending[origin].slice(newline + 1);
newline = this.pending[origin].indexOf('\n');
}
}
/** Flush partial trailing lines (a stream killed mid-line still classifies). */
end(): void {
for (const origin of ['stdout', 'stderr'] as const) {
this.pending[origin] += this.decoders[origin].end();
if (this.pending[origin].length > 0) this.handleLine(this.pending[origin], origin);
this.pending[origin] = '';
}
}
report(): FreeRunReport {
const inFlight = this.sawSummary
? []
: [...this.progress.entries()]
.filter(([, p]) => p.headerSeen && !p.ended)
.map(([file]) => file)
.sort();
return {
testsRan: this.testsRan,
filesRan: this.filesRan,
sawTerminalSummary: this.sawSummary,
failures: [...this.failures],
crashedFiles: [...this.crashed].sort(),
unhandledErrors: [...this.unhandled],
inFlight,
filesWithNoOutput: this.plannedFiles.filter((f) => !this.progress.has(normalizeRelativePath(f))).length,
};
}
private handleLine(rawLine: string, origin: StreamOrigin): void {
// GitHub Actions: bun wraps each file's section in ::group::<header>.
// Without stripping, the real header fails FILE_HEADER_RE, failures get
// attributed to the PREVIOUS file, and the terminal recap's re-printed
// (fail) lines land under a second phantom file (observed on the first
// Linux run: 5 real failures reported as 10 across 2 files).
const line = stripAnsiLine(rawLine).replace(/^::group::/, '');
let visible = false;
// Bun's terminal recap ("N tests failed:") re-prints every (fail) line
// WITHOUT re-printing file headers. Attributing those to the stale
// currentFile invented a phantom failing file on the first Linux run
// (5 real failures reported as 10 across 2 files, one innocent).
if (/^\d+ tests? failed:$/.test(line)) {
this.inRecap = true;
if (this.currentFile) this.progressFor(this.currentFile).ended = true;
this.currentFile = null;
}
if (line === '# Unhandled error between tests') {
this.unhandled.push({ file: this.currentFile });
}
const header = FILE_HEADER_RE.exec(line);
if (header) {
const file = this.canonicalize(header[1]);
// A new header ends the previous file's chunk — that file is no longer
// a wedge suspect. (Bun 1.3.x prints NO (pass) lines, so chunk
// delimiters, not result lines, are the completion signal.)
if (this.currentFile && this.currentFile !== file) this.progressFor(this.currentFile).ended = true;
this.currentFile = file;
this.progressFor(file).headerSeen = true;
} else {
const fail = FAIL_RESULT_CAPTURE_RE.exec(line);
const retry = fail ? null : CRASH_RETRY_RE.exec(line);
const final = fail || retry ? null : CRASH_FINAL_RE.exec(line);
if (fail) {
visible = true;
// In the recap, a (fail) line only records a failure the main run
// somehow never attributed (belt and braces); known names dedupe.
const recapDuplicate = this.inRecap
&& this.failures.some((f) => f.testName === fail[1]);
const key = `${this.currentFile ?? ''}\u0000${fail[1]}`;
if (!recapDuplicate && !this.failureKeys.has(key)) {
this.failureKeys.add(key);
this.failures.push({ file: this.currentFile, testName: fail[1] });
}
} else if (retry) {
// The file will run again — a crash+retry does not end its chunk.
visible = true;
this.crashed.add(this.canonicalize(retry[1]));
} else if (final) {
visible = true;
const file = this.canonicalize(final[1]);
this.crashed.add(file);
this.progressFor(file).ended = true;
} else {
const summary = TERMINAL_SUMMARY_CAPTURE_RE.exec(line);
if (summary) {
visible = true;
this.sawSummary = true;
this.testsRan = Number.parseInt(summary[1], 10);
this.filesRan = Number.parseInt(summary[2], 10);
}
}
}
if (!visible) visible = CONSOLE_ALWAYS_MARKERS.some((marker) => line.includes(marker));
if (visible && this.forward) this.forward(`${rawLine.replace(/\r$/, '')}\n`, origin);
}
private progressFor(file: string): FileProgress {
let entry = this.progress.get(file);
if (!entry) {
entry = { headerSeen: false, ended: false };
this.progress.set(file, entry);
}
return entry;
}
/**
* Map a printed path back to its planned relative path. Bun prints paths
* relative to the child's (real)cwd, so a symlinked cwd (macOS /tmp) yields
* `../..`-prefixed forms — strip the prefix and suffix-match.
*/
private canonicalize(printedPath: string): string {
const cached = this.canonicalCache.get(printedPath);
if (cached) return cached;
const stripped = normalizeRelativePath(printedPath).replace(/^(?:\.{1,2}\/)+/, '');
let resolved = stripped;
if (!this.plannedSet.has(stripped)) {
const match = this.plannedFiles.find(
(planned) => stripped.endsWith(`/${planned}`) || planned.endsWith(`/${stripped}`),
);
if (match) resolved = match;
}
this.canonicalCache.set(printedPath, resolved);
return resolved;
}
}
/**
* The stable post-run epilogue. Success is one line; failure names every
* failing test (deduped, attributed) and crashed worker; a wall-timeout kill
* additionally prints the wedge-suspect list (see FreeRunReport.inFlight for
* the heuristic and its limits).
*/
export function buildRunEpilogue(
status: FreeShardStatus,
report: FreeRunReport,
elapsedMs: number,
logPath: string,
): string[] {
const seconds = Math.round(elapsedMs / 1000);
if (status === 'passed') {
return [
`[test:free] PASS — ${report.testsRan ?? '?'} tests, ${report.filesRan ?? '?'} files, ${seconds}s. Full log: ${logPath}`,
];
}
const failingFiles = new Set(report.failures.map((f) => f.file ?? '(unattributed)'));
const lines = [
`[test:free] FAIL — ${report.failures.length} failing test(s) in ${failingFiles.size} file(s), `
+ `${report.crashedFiles.length} crashed worker(s)${report.unhandledErrors.length > 0 ? `, ${report.unhandledErrors.length} unhandled error(s) between tests` : ''}. Full log: ${logPath}`,
];
for (const failure of report.failures) {
lines.push(` ✗ ${failure.file ?? '(unattributed)'}${failure.testName}`);
}
for (const file of report.crashedFiles) {
lines.push(` ⚠ crashed+retried: ${file}`);
}
for (const u of report.unhandledErrors) {
lines.push(` ⚠ unhandled error between tests (around ${u.file ?? 'unknown file'})`);
}
if (status === 'timed-out') {
if (report.inFlight.length > 0) {
lines.push(` ⏱ in flight at kill: ${report.inFlight.join(', ')}`);
} else {
lines.push(
' ⏱ in flight at kill: unknown — no open file chunk was observed '
+ '(bun --parallel buffers a file\'s output until it completes, so a silent wedge never prints); '
+ `${report.filesWithNoOutput} planned file(s) produced no output before the kill.`,
);
}
}
return lines;
}
export type FreeShardStatus = 'passed' | 'failed' | 'timed-out';
// ─── Flake ledger (WS1 telemetry) ───────────────────────────────────────────
// Single-writer JSONL: ONLY this parent runner appends (never shards, never
// tests — no concurrent-append hazard by construction). CI points
// GSTACK_FLAKE_LEDGER at $RUNNER_TEMP and uploads it as an artifact every
// run, so repeat offenders become an enumerable series instead of console
// scrollback. Fail-open with a loud stderr warning: a broken ledger must
// never red the only required lane.
export interface FlakeLedgerEntry {
ts: string;
runner: 'free';
kind: 'flaky-pass';
file: string;
/** Shard the original failure surfaced in, when attributable. */
shard?: number;
/** Code-state attribution (review finding): without branch/sha the series
* can't tie an entry to the state that produced it, and the WS16
* promotion evidence needs exactly that. */
branch?: string;
git_sha?: string;
}
export function flakeLedgerPath(env: NodeJS.ProcessEnv = process.env): string {
if (env.GSTACK_FLAKE_LEDGER) return env.GSTACK_FLAKE_LEDGER;
// Local default: per-PROJECT, not the machine-global tmpdir — sibling
// Conductor worktrees of DIFFERENT repos must not interleave into one
// series (review finding). CI always sets GSTACK_FLAKE_LEDGER explicitly.
try {
const slug = spawnSync('bash', ['-c', '~/.claude/skills/gstack/bin/gstack-slug 2>/dev/null'], { stdio: 'pipe', timeout: 3000 })
.stdout?.toString().match(/^SLUG=(.+)$/m)?.[1];
if (slug) {
const dir = path.join(os.homedir(), '.gstack', 'projects', slug);
fs.mkdirSync(dir, { recursive: true });
return path.join(dir, 'flake-ledger.jsonl');
}
} catch { /* fall through */ }
return path.join(os.tmpdir(), 'gstack-flake-ledger.jsonl');
}
export function appendFlakeLedger(
entries: FlakeLedgerEntry[],
ledgerPath: string,
warn: (line: string) => void = (line) => console.error(line),
): boolean {
if (entries.length === 0) return true;
try {
fs.mkdirSync(path.dirname(ledgerPath), { recursive: true });
fs.appendFileSync(ledgerPath, entries.map((e) => JSON.stringify(e)).join('\n') + '\n');
return true;
} catch (error) {
warn(`[test:free] WARNING: could not append flake ledger at ${ledgerPath} `
+ `(${error instanceof Error ? error.message : String(error)}) — flaky-pass telemetry lost for this run, verdict unaffected`);
return false;
}
}
export interface FreeShardOutcome {
shard: number;
files: string[];
status: FreeShardStatus;
exitCode: number | null;
elapsedMs: number;
groupPid: number | null;
/**
* Repo-relative files with attributed test failures or crashes, deduped.
* Feeds the opt-in flaky retry pass (GSTACK_FREE_RETRY_FLAKY) — empty on
* pass, and empty when every failure was unattributed (retry would be
* meaningless without knowing what to re-run).
*/
failingFiles: string[];
/**
* Count of failure evidence the retry pass CANNOT re-run by file: fail
* lines seen before any file-chunk header, unhandled errors between tests,
* and a truncated run (no terminal summary). Nonzero vetoes the flaky
* retry for the whole run — retrying only failingFiles would re-run a
* subset and mask the rest as a FLAKY-PASS, re-opening the silent-truncation
* hole the strict classifier exists to close.
*/
unattributedFailures: number;
}
export interface ShardCommand {
command: string;
args: string[];
}
export interface RunFreeShardOptions {
/** External wall-clock deadline; on expiry the child's process GROUP is SIGKILLed. */
wallTimeoutMs?: number;
rootDir?: string;
env?: NodeJS.ProcessEnv;
/** Pass bun's --parallel. No production caller today (see BuildShardArgsOptions.parallel). */
parallel?: boolean;
/** Override the spawned command. Tests inject fake pass/fail/slow commands. */
commandFor?: (files: string[]) => ShardCommand;
/** Suppress ALL child output from the console (tests). The classifier and the log file still see every byte. */
quiet?: boolean;
/** Forward the full child stream to the console (legacy firehose). Default: the quiet filtered console. */
verbose?: boolean;
/**
* Console sink for child-stream output (tests inject to assert quiet vs
* verbose behavior). Default: process.stdout / process.stderr by origin.
* Runner-owned [test:free] lines go through `log`, not this sink.
*/
consoleWrite?: (text: string) => void;
/** Per-run full-stream log path (tests inject). Default: a timestamped file under os.tmpdir(). */
logFilePath?: string;
log?: (line: string) => void;
}
const EPILOGUE_WORD: Record<FreeShardStatus, string> = {
passed: 'pass',
failed: 'fail',
'timed-out': 'timed-out',
};
/** One line per shard, printed after the run: `[test:free] shard i/N: M files, XXs, pass|fail|timed-out`. */
function shardEpilogue(outcome: FreeShardOutcome, totalShards: number): string {
return `[test:free] shard ${outcome.shard}/${totalShards}: ${outcome.files.length} files, `
+ `${Math.round(outcome.elapsedMs / 1000)}s, ${EPILOGUE_WORD[outcome.status]}`;
}
/**
* Run one shard (or the whole suite, in --parallel full-suite mode) in its own
* bun process and classify the result strictly.
*
* Verdict integrity: the child's exit code is never trusted alone. Output is
* fed through BunTestOutputClassifier, and strictTestExitCode requires bun's
* terminal summary to report EXACTLY the planned file count — a shard that
* exits 0 without the summary (mid-suite process.exit truncation), with
* `(fail)` result lines, or having run fewer files than planned is a FAILURE.
* This is enforced for injected fake commands too (unlike the paid runner),
* so tests can pin the summary-missing => failure backstop; fake passing
* commands must print a synthetic `Ran N tests across M files. [Xms]` line.
*
* Per-shard temp isolation: each spawned child gets its own throwaway TMPDIR
* (TEMP/TMP on Windows) so shards can't trip over each other's temp files.
* Deliberately NOT GSTACK_HOME: injecting one shared scratch home for a whole
* invocation made 6,900 tests share a MUTABLE state dir — config tests wrote
* keys into it and relink/update-check tests then read them (measured: 12
* cross-contamination failures on the first full run). Tests that need
* GSTACK_HOME isolation mkdtemp their own per test — the repo convention —
* and the hermetic-env machinery covers E2E children.
*/
export async function runFreeShard(
files: string[],
shardNumber: number,
totalShards: number,
options: RunFreeShardOptions = {},
): Promise<FreeShardOutcome> {
const log = options.log ?? ((line: string) => console.log(line));
const label = `[test:free] shard ${shardNumber}/${totalShards}`;
// Empty shard = fast no-op SUCCESS. Indices are stable for the CI matrix,
// so an unoccupied index must not fail or shift work to a different runner.
if (files.length === 0) {
const outcome: FreeShardOutcome = {
shard: shardNumber, files: [], status: 'passed', exitCode: 0, elapsedMs: 0, groupPid: null, failingFiles: [], unattributedFailures: 0,
};
log(shardEpilogue(outcome, totalShards));
return outcome;
}
const rootDir = options.rootDir ?? ROOT;
const wallTimeoutMs = options.wallTimeoutMs ?? DEFAULT_WALL_TIMEOUT_MS;
log(`${label} (${files.length} files${options.parallel ? ', bun --parallel' : ''})`);
// Full-stream capture: EVERY child byte lands here, whatever the console
// shows. Printed once at start so a wedged or noisy run is inspectable
// without a re-run.
const logPath = options.logFilePath ?? nextDefaultLogPath();
const logStream = fs.createWriteStream(logPath);
let logWriteFailed = false;
logStream.on('error', (err) => {
if (logWriteFailed) return;
logWriteFailed = true;
console.error(`${label} could not write the full log at ${logPath}: ${err.message}`);
});
log(`[test:free] full log: ${logPath}`);
const { command, args } = options.commandFor
? options.commandFor(files)
: { command: process.execPath, args: buildShardArgs(files, { parallel: options.parallel, rootDir }) };
const env = { ...(options.env ?? process.env) };
const stateDir = fs.mkdtempSync(path.join(os.tmpdir(), 'gstack-free-shard-'));
const childTmp = path.join(stateDir, 'tmp');
fs.mkdirSync(childTmp);
env.TMPDIR = childTmp;
env.TEMP = childTmp;
env.TMP = childTmp;
// Per-shard Chromium profile (same isolation idea as TMPDIR): nine test
// files launch in-process persistent contexts or daemons that default to
// the SHARED ~/.gstack/chromium-profile, and two concurrent shards on one
// profile dir kill each other's browser — observed live on CI once
// duration packing recomposed shards (handoff's launchPersistentContext
// died "Target page, context or browser has been closed" while a sibling
// shard's daemon logged "Chromium process crashed"). Hash sharding had
// masked the collision by chance placement. Within a shard, files run
// serially, so sharing the per-shard profile is safe; config tests that
// assert resolution order save/restore this env around their assertions.
env.CHROMIUM_PROFILE = path.join(stateDir, 'chromium-profile');
const startedAt = Date.now();
const child = spawn(command, args, {
cwd: rootDir,
env,
stdio: ['ignore', 'pipe', 'pipe'],
detached: process.platform !== 'win32',
windowsHide: true,
});
const groupPid = child.pid ?? null;
// Group-kill on parent SIGINT/SIGTERM too, not just on timeout.
const forwarding = installChildSignalForwarding({
kill: (signal?: NodeJS.Signals | number) => {
killProcessGroup(child, (signal as NodeJS.Signals) ?? 'SIGTERM');
return true;
},
});
const classifier = new BunTestOutputClassifier();
// Console policy: quiet => nothing; verbose => the raw firehose; default =>
// only always-visible lines (fail results, crash markers, error/panic
// markers, the terminal summary), selected by the reporter. The reporter
// consumes the stream in EVERY mode so the epilogue can attribute failures.
const emitToConsole = (text: string, origin: StreamOrigin): void => {
if (options.quiet) return;
if (options.consoleWrite) {
options.consoleWrite(text);
return;
}
(origin === 'stdout' ? process.stdout : process.stderr).write(text);
};
const reporter = new FreeRunReporter(files, options.verbose ? undefined : emitToConsole);
const consumeStream = (stream: NodeJS.ReadableStream, origin: StreamOrigin): Promise<void> =>
new Promise((resolve, reject) => {
stream.on('data', (chunk: Buffer | string) => {
classifier.write(chunk, origin); // strict verdict ALWAYS sees the full stream
if (!logWriteFailed) logStream.write(chunk);
reporter.write(chunk, origin);
if (options.verbose) emitToConsole(typeof chunk === 'string' ? chunk : chunk.toString('utf8'), origin);
});
stream.on('end', resolve);
stream.on('error', reject);
});
let timedOut = false;
const killTimer = setTimeout(() => {
timedOut = true;
killProcessGroup(child, 'SIGKILL');
}, wallTimeoutMs);
let exitCode: number | null = null;
try {
const streams: Array<Promise<void>> = [];
if (child.stdout) streams.push(consumeStream(child.stdout, 'stdout'));
if (child.stderr) streams.push(consumeStream(child.stderr, 'stderr'));
exitCode = await new Promise<number | null>((resolve, reject) => {
child.once('error', reject);
child.once('close', (code) => resolve(code));
});
await Promise.all(streams);
} finally {
clearTimeout(killTimer);
forwarding.dispose();
// Reap survivors of this shard even on the clean path.
killProcessGroup(child, 'SIGKILL');
reporter.end();
await new Promise<void>((resolve) => logStream.end(() => resolve()));
try {
fs.rmSync(stateDir, { recursive: true, force: true });
} catch {
// Best-effort cleanup of a throwaway temp dir — a locked file on
// Windows must not turn a real verdict into an exception.
}
}
const summary = classifier.end();
const status: FreeShardStatus = timedOut
? 'timed-out'
: strictTestExitCode(exitCode ?? 1, summary, files.length) === 0 ? 'passed' : 'failed';
if (status === 'timed-out') {
console.error(
`${label} exceeded the ${Math.round(wallTimeoutMs / 1000)}s wall-clock deadline — `
+ 'killed the process group. Reporting as TIMED-OUT (distinct from failed).',
);
} else if (status === 'failed' && (exitCode ?? 1) === 0) {
const reason = summary.failedTests > 0 || summary.unhandledBetweenTests > 0
? `printed ${summary.failedTests} failing result(s) and ${summary.unhandledBetweenTests} unhandled error(s) between tests`
: summary.terminalFileCounts.length === 0
? "never printed bun's terminal summary — the run was truncated (a process.exit fired mid-suite)"
: `bun's summary reported ${summary.terminalFileCounts.join(', ')} file(s), expected ${files.length}`;
console.error(`${label} exited 0 but ${reason}. Treating as FAILED.`);
} else if (status === 'failed') {
console.error(`${label} failed with exit code ${exitCode ?? 'signal'}`);
}
const report = reporter.report();
const failingFiles = status === 'passed' ? [] : [...new Set([
...report.failures.map((f) => f.file).filter((f): f is string => !!f),
...report.crashedFiles,
])];
const unattributedFailures = status === 'passed' ? 0
: report.failures.filter((f) => !f.file).length
+ report.unhandledErrors.length
+ (report.sawTerminalSummary ? 0 : 1);
const outcome: FreeShardOutcome = {
shard: shardNumber, files, status, exitCode, elapsedMs: Date.now() - startedAt, groupPid, failingFiles, unattributedFailures,
};
log(shardEpilogue(outcome, totalShards));
for (const line of buildRunEpilogue(status, report, outcome.elapsedMs, logPath)) log(line);
return outcome;
}
let logPathSequence = 0;
/** Timestamped per-run log file under os.tmpdir(); pid+sequence defeat same-ms collisions. */
function nextDefaultLogPath(): string {
const stamp = new Date().toISOString().replace(/[:.]/g, '-');
logPathSequence += 1;
return path.join(os.tmpdir(), `gstack-free-test-${stamp}-${process.pid}-${logPathSequence}.log`);
}
function exitCodeFor(status: FreeShardStatus): number {
if (status === 'passed') return 0;
return status === 'timed-out' ? 124 : 1;
}
/**
* `--record-durations`: time every file in its own child (exact per-file wall,
* immune to bun's stream buffering) and write the committed seed atomically.
* Occasional + manual by design — CI never records (a hint refreshed by a
* human beats per-run churn), and the runtime (~serial suite / jobs) is fine
* for an operation run a few times a quarter.
*/
async function recordFreeTestDurations(files: string[], jobs: number): Promise<number> {
const durations: Record<string, number> = {};
const failed: string[] = [];
let cursor = 0;
console.log(`[test:free] recording per-file durations: ${files.length} files across ${jobs} workers`);
const worker = async (): Promise<void> => {
for (;;) {
const index = cursor;
cursor += 1;
if (index >= files.length) return;
const file = files[index];
const started = Date.now();
const child = spawn('bun', ['test', file, `--timeout=${FREE_TEST_TIMEOUT_MS}`], {
cwd: ROOT,
stdio: ['ignore', 'ignore', 'ignore'],
env: { ...process.env, GSTACK_HEADLESS: '1' },
});
const code = await new Promise<number>((resolve) => {
const timer = setTimeout(() => { child.kill('SIGKILL'); }, wallTimeoutForShard(1));
child.on('close', (c) => { clearTimeout(timer); resolve(c ?? 1); });
child.on('error', () => { clearTimeout(timer); resolve(1); });
});
durations[normalizeRelativePath(file)] = Date.now() - started;
if (code !== 0) failed.push(file);
}
};
await Promise.all(Array.from({ length: Math.max(1, jobs) }, () => worker()));
const target = process.env.GSTACK_FREE_TEST_DURATIONS ?? path.join(ROOT, FREE_TEST_DURATIONS_FILE);
const payload = {
version: 1,
recordedAt: new Date().toISOString(),
durations: Object.fromEntries(Object.entries(durations).sort(([a], [b]) => (a < b ? -1 : 1))),
};
// Atomic temp+rename (capture-context-budget's pattern): a killed recorder
// must never leave a truncated seed for loadFreeTestDurations to warn on.
const tmp = `${target}.tmp-${process.pid}`;
fs.writeFileSync(tmp, `${JSON.stringify(payload, null, 2)}\n`);
fs.renameSync(tmp, target);
console.log(`[test:free] wrote ${Object.keys(durations).length} durations to ${path.relative(ROOT, target)}`);
if (failed.length > 0) {
// Failures still recorded (a red file's duration is still a real cost),
// but surfaced loudly — recording from a broken tree deserves a look.
console.error(`[test:free] WARNING: ${failed.length} file(s) failed while recording:`);
for (const f of failed) console.error(` ✗ ${f}`);
return 1;
}
return 0;
}
async function main(): Promise<number> {
const options = parseCliOptions(process.argv.slice(2));
const allFiles = collectFreeTestFiles();
if (allFiles.length === 0) {
throw new Error('No free test files were discovered.');
}
let files = allFiles;
let curationReport: CurationResult | null = null;
if (options.windowsOnly) {
curationReport = curateWindowsSafe(allFiles);
files = curationReport.safe;
console.log(`[test:free] curated ${files.length} Windows-safe tests (${curationReport.excluded.length} excluded)`);
if (options.listOnly && curationReport.excluded.length > 0) {
console.log('\nExcluded (POSIX-fragile):');
for (const { file, reason } of curationReport.excluded) {
console.log(` - ${file} [${reason}]`);
}
}
}
if (options.listOnly) {
console.log(`\nDiscovered ${files.length} test files.`);
for (const file of files) console.log(` ${file}`);
return 0;
}
if (options.recordDurations) {
const jobs = Math.max(1, Math.min(MAX_FULL_SUITE_JOBS, os.cpus().length - RESERVED_CPUS));
return recordFreeTestDurations(files, jobs);
}
if (options.dryRun) {
const shards = assignFilesToShards(files, options.shardCount);
const occupied = shards.filter((s) => s.length > 0).length;
console.log(
`\nWould run ${files.length} files across ${shards.length} shards (${occupied} occupied). `
+ 'Without --shard, the full suite runs as N concurrent shard processes '
+ '(plus a serial tree-mutating shard) instead.',
);
for (const line of formatShardSummary(shards)) console.log(line);
return 0;
}
if (options.shardIndex !== null) {
// Bounds-check against the REQUESTED shard count, not post-assignment
// occupancy — indices must be stable for a CI matrix, and an empty shard
// is a valid fast no-op.
if (!Number.isInteger(options.shardIndex) || options.shardIndex < 1 || options.shardIndex > options.shardCount) {
throw new Error(`--shard must be between 1 and ${options.shardCount}. Received: ${options.shardIndex}`);
}
const shards = assignFilesToShards(files, options.shardCount);
const outcome = await runFreeShard(shards[options.shardIndex - 1], options.shardIndex, options.shardCount, {
wallTimeoutMs: options.wallTimeoutMs,
verbose: options.verbose,
});
return exitCodeFor(outcome.status);
}
// Full-suite mode: N concurrent shard PROCESSES, serial within each — the
// paid runner's proven model. One `bun test --parallel` invocation was
// tried first (decision V3) and abandoned after three distinct
// worker-runtime pathologies in a single day on Bun 1.3.13: a segfault
// whose crashed-worker retry wedged the run (security-live-playwright), a
// gated file's still-running file-level hooks stalling a worker
// (compare-board), and spawn-heavy files hanging workers under load
// (session-runner-timeout). Plain child processes have none of these:
// proven spawn semantics, per-shard group-kill, per-shard logs, and a
// wedge only ever costs its own shard. WORKER_HOSTILE files are moot in
// process shards (no workers) and fold back into normal assignment.
const jobs = fullSuiteJobs();
// Phase split: tree-mutating tests run AFTER the parallel shards, in one
// serial shard, so no concurrent shard ever reads a half-regenerated tree.
const mutators = files.filter((f) => f in TREE_MUTATING);
const readers = files.filter((f) => !(f in TREE_MUTATING));
const durations = loadFreeTestDurations();
const packed = durations ? packShardsByDuration(readers, jobs, durations) : null;
const shards = packed ? packed.shards : assignFilesToShards(readers, jobs);
const totalShards = jobs + (mutators.length > 0 ? 1 : 0);
console.log(`[test:free] full suite: ${readers.length} files across ${jobs} shard processes`
+ (packed ? ' (duration-packed)' : '')
+ (mutators.length > 0 ? `, then ${mutators.length} tree-mutating file(s) serially` : ''));
if (packed) {
// One line per shard so a packing regression is diagnosable from any log.
packed.predictedMs.forEach((ms, i) => {
console.log(`[test:free] shard ${i + 1}: ${shards[i].length} files, predicted ~${Math.round(ms / 1000)}s`);
});
}
const shardTimeout = (fileCount: number): number =>
options.wallTimeoutExplicit ? options.wallTimeoutMs : wallTimeoutForShard(fileCount, options.wallTimeoutMs);
const outcomes = await Promise.all(
shards.map((shardFiles, index) => runFreeShard(shardFiles, index + 1, totalShards, {
// Packed shards get duration-aware walls: LPT decouples file count from
// cost BY DESIGN, so the 5s/file heuristic would undersize a shard
// holding few expensive files.
wallTimeoutMs: packed && !options.wallTimeoutExplicit
? wallTimeoutForPackedShard(packed.predictedMs[index], options.wallTimeoutMs, shardFiles.length)
: shardTimeout(shardFiles.length),
verbose: options.verbose,
})),
);
let worst = Math.max(...outcomes.map((o) => exitCodeFor(o.status)));
// Cancellation stops the run: don't launch the serial tree-mutating shard
// after a SIGINT/SIGTERM already killed the parallel phase.
if (mutators.length > 0 && !isTerminationRequested()) {
const mutatorOutcome = await runFreeShard(mutators, totalShards, totalShards, {
wallTimeoutMs: shardTimeout(mutators.length),
verbose: options.verbose,
});
worst = Math.max(worst, exitCodeFor(mutatorOutcome.status));
if (mutatorOutcome.status !== 'passed') {
// Mutator safety rests on each test restoring default state itself; a
// SIGKILL at the wall deadline (or a mid-regeneration crash) defeats
// that by construction. Say so, loudly, before someone commits
// regenerated SKILL.md / .agents artifacts by accident.
const dirty = spawnSyncGitStatusGenerated();
if (dirty.length > 0) {
console.error('[test:free] ⚠ tree-mutating shard did not finish cleanly — generated artifacts may be mid-regeneration:');
for (const line of dirty.slice(0, 20)) console.error(`[test:free] ${line}`);
console.error('[test:free] restore with: bun run gen:skill-docs (or git checkout -- <paths>)');
}
}
outcomes.push(mutatorOutcome);
}
// Opt-in flaky retry (GSTACK_FREE_RETRY_FLAKY=1): when every failure is an
// attributed test failure (no timeouts, no unattributed carnage), re-run
// just the failing files ONCE in a fresh serial shard. A clean retry
// downgrades the run to a loud flaky-pass; a repeat failure stays a
// failure. Default OFF: dev boxes should see flakes, not absorb them.
// Exists for syscall-supervised sandboxes (see fullSuiteJobs) where a run
// lands 0-1 spurious browser-timing failures under an otherwise-green
// suite. Capped so a genuinely broken tree never masquerades as flaky.
const RETRY_CAP = 5;
if (
worst !== 0
&& process.env.GSTACK_FREE_RETRY_FLAKY === '1'
&& !isTerminationRequested()
&& outcomes.every((o) => o.status !== 'timed-out')
) {
const flakyFiles = [...new Set(outcomes.flatMap((o) => o.failingFiles))]
.filter((f): f is string => typeof f === 'string' && f.length > 0);
// "Fully attributed" is per-failure, not per-shard: a shard with one
// attributed failure PLUS a headerless failure / unhandled error /
// truncated run must veto the retry — re-running only failingFiles would
// mask the unattributable evidence as a FLAKY-PASS.
const allAttributed = outcomes.every((o) => o.status === 'passed'
|| (o.failingFiles.length > 0 && o.unattributedFailures === 0));
if (allAttributed && flakyFiles.length > 0 && flakyFiles.length <= RETRY_CAP) {
console.log(`[test:free] flaky-retry: re-running ${flakyFiles.length} failing file(s) once, serially: ${flakyFiles.join(', ')}`);
const retryOutcome = await runFreeShard(flakyFiles, totalShards + 1, totalShards + 1, {
wallTimeoutMs: shardTimeout(flakyFiles.length),
verbose: options.verbose,
});
if (retryOutcome.status === 'passed') {
console.log(`[test:free] FLAKY-PASS — ${flakyFiles.length} file(s) failed once and passed on serial retry: ${flakyFiles.join(', ')}`);
console.log('[test:free] treat repeat offenders as real flakes worth fixing, not noise.');
// Durable record (WS1): console lines vanish with the scrollback; the
// ledger makes repeat offenders rankable across runs (eval:flake-rank).
const ts = new Date().toISOString();
// Two separate calls: `rev-parse --abbrev-ref HEAD HEAD` abbreviates
// BOTH revs, printing the branch twice — git_sha recorded the branch
// name (codex adversarial finding).
const ledgerBranch = (spawnSync('git', ['rev-parse', '--abbrev-ref', 'HEAD'], { cwd: ROOT, encoding: 'utf8', timeout: 5000 }).stdout ?? '').trim();
const ledgerSha = (spawnSync('git', ['rev-parse', 'HEAD'], { cwd: ROOT, encoding: 'utf8', timeout: 5000 }).stdout ?? '').trim();
appendFlakeLedger(
flakyFiles.map((file) => ({
ts,
runner: 'free' as const,
kind: 'flaky-pass' as const,
file,
shard: outcomes.find((o) => o.failingFiles.includes(file))?.shard,
...(ledgerBranch ? { branch: ledgerBranch } : {}),
...(ledgerSha ? { git_sha: ledgerSha.slice(0, 12) } : {}),
})),
flakeLedgerPath(),
);
worst = 0;
} else {
console.error('[test:free] flaky-retry FAILED — the failures reproduce serially; not flaky.');
}
} else {
console.log(`[test:free] flaky-retry skipped: ${allAttributed ? `${flakyFiles.length} failing file(s) exceeds cap ${RETRY_CAP}` : 'failures not fully attributed'}.`);
}
}
return worst;
}
/** Dirty generated artifacts (SKILL.md / host outputs) after a failed mutator shard. */
function spawnSyncGitStatusGenerated(): string[] {
const result = spawnSync('git', ['status', '--porcelain'], { cwd: ROOT, encoding: 'utf8' });
if (result.status !== 0 || !result.stdout) return [];
return result.stdout.split('\n').filter((line) =>
/SKILL\.md$/.test(line) || line.includes('.agents/') || line.includes('.factory/'));
}
if (import.meta.main) {
try {
process.exitCode = await main();
} catch (error) {
console.error(`[test:free] ${error instanceof Error ? error.message : String(error)}`);
process.exitCode = 1;
}
}