mirror of
https://github.com/garrytan/gstack.git
synced 2026-09-26 06:41:13 +02:00
* fix: acknowledge seeded plans before invoking review skills * fix: distinguish current plan input from conversation history * fix: keep hermetic plan reviews on manual permissions * fix: distinguish tool discovery from file permission ownership * fix: preserve initial plan mode in observation tests * fix: wait for scope decisions before writing review findings * fix: carry autoplan decisions consistently into review artifacts * test: retain native failure context in periodic assertions * fix: advance active file permissions before queued questions * fix: finish red-team attempts before retry and cleanup * fix: finalize plan format captures and judges before retry * fix: cancel setup-gbrain SDK attempts before fixture cleanup * test: select periodic consumers of the bounded attempt helper * fix native Bash permission cards and queued questions * fix: preserve independent decisions and review scope Keep CEO approach, engineering scope and outside-review choices from approving independent remedies together. Carry declared contracts through DX polish and resolve new gaps before editing the plan. Regenerate every host and retain existing stop boundaries. Validation: 654 focused tests passed across nine files; all-host generation passed. Full free and periodic validation pending. Co-Authored-By: OpenAI Codex <noreply@openai.com> * fix: require approval before design plan amendments Align the Design review philosophy and rating recipe with its section protocol: resolve one proposed fix, then apply only that approved decision and retain honest scores for declined fixes. Validation: 469 focused tests passed across four files; all-host generation passed. Co-Authored-By: OpenAI Codex <noreply@openai.com> * fix: observe native question completion before transcript persistence Match owned completion hooks to submitted choices, reject conflicting or late answers, and retain bounded failure evidence. * test: recognize review posture in acknowledged native questions Require the selected mode acknowledgement, a completed follow-up question, and its current decoded display while preserving existing posture assertions. * fix: preserve settled CEO choices and isolate pending remedies Resolve established approach gates with cited authority and keep independent fixes out of unrelated option commitments and plan amendments. * fix: carry approved DX choices through later review steps Choose documentation approaches within the accepted scope and map resolved confusion points without reopening them through a bulk menu. * test: handle native settings-file edit prompts Keep one-time owned-file approvals and retain the actual sampled Autoplan permission frame with its matching barrier state. * test: accept standard CEO reply directives with tuning footers Recognize the exact trailing preference footer and letter-list directive while preserving current-display and exact acknowledgement checks. * test: scope split reviewers to their generated plan artifacts * test: observe native Bash permissions and invocation results * test: handle owned Bash prompts during mode preference checks * test: preserve synchronous subprocess rejection in Codex fixture * Fix periodic review handoff navigation Recognize review-first and explicit manual-next-step labels while preserving exact action families, manual preference, and ambiguous-menu rejection. Co-authored-by: OpenAI Codex <noreply@openai.com> * Bind pending file permissions to distinct current targets Allow one captured file request to own the complete current dialog while unrelated file work is pending. Preserve same-path ambiguity, exact input ownership, and one-time grant checks. Co-authored-by: OpenAI Codex <noreply@openai.com> * Make paired CEO verification choices genuinely unresolved Start the positive control with proposed manual checks so its unchanged oracle measures two new coverage decisions. Preserve runtime contracts, targets, count bounds, and all assertions. Co-authored-by: OpenAI Codex <noreply@openai.com> * Keep CEO review options and verification within approved scope Audit every offered option for independent add-ons and keep new verification depth pending until accepted. Preserve already requested coverage and trace plan changes to the actual decision. Co-authored-by: OpenAI Codex <noreply@openai.com> * Assemble DX review artifacts before appending the final report Keep early DX evidence above decisions, update artifact sections in place, and append the report using the actual current file suffix. Re-read after deleting an existing report before choosing the append anchor. Co-authored-by: OpenAI Codex <noreply@openai.com> * Keep outside plan reviews exclusive and invocation-owned Follow one preflight-selected backend, terminate failed Codex work before fallback, and allocate extra prompt/output files uniquely. Consume only the current invocation’s completed output. Co-authored-by: OpenAI Codex <noreply@openai.com> * Select periodic completion evaluations for report writer changes Register the shared review resolver for eight missing consumers and regress selection for all nine completion cases without changing their IDs or tiers. Co-authored-by: OpenAI Codex <noreply@openai.com> * Keep permission ambiguity fixtures on the same normalized target Use distinct raw spellings of one target in the four negative fixtures so they exercise the normalized duplicate-owner guard after exact current-file disambiguation. Preserve the existing exception, no-input, diagnostic and cleanup assertions. Co-authored-by: OpenAI Codex <noreply@openai.com> * Clarify preserved contracts in engineering review fixture Co-authored-by: OpenAI Codex <noreply@openai.com> * Recognize the offered DX follow-up handoff Co-authored-by: OpenAI Codex <noreply@openai.com> * Check independent commitments before presenting review options Co-authored-by: OpenAI Codex <noreply@openai.com> * Keep Codex review output and status in one shell invocation Co-authored-by: OpenAI Codex <noreply@openai.com> * Distinguish seeded plans from reports written by a test attempt Co-authored-by: OpenAI Codex <noreply@openai.com> * Recover clipped Autoplan file approvals with bounded viewport resizing Co-authored-by: OpenAI Codex <noreply@openai.com> * Recover clipped Bash approvals before binding the complete command Co-authored-by: OpenAI Codex <noreply@openai.com> * Isolate setup message tests from the shared checkout Run the real installer in a temporary payload with private config, require successful completion, and guard source and binary contents and mtimes. Co-authored-by: OpenAI Codex <noreply@openai.com> * Fix periodic native permission and report completion handling Match the pinned CLI's soft wraps and clipped headings without granting from incomplete frames. Retire completed file requests, retain mode annotations, and ask section captures for a short final acknowledgement after their full report is saved. Co-authored-by: OpenAI Codex <noreply@openai.com> * Preserve review approvals and validate DX comparison artifacts Keep independent remedies and approved amendments explicit. Give the synthetic DX review its existing documentation and validate peer comparison as required analysis alongside four native decisions. Add positive and negative semantic calibrations while preserving review counts, model budgets and prompt size limits. Co-authored-by: OpenAI Codex <noreply@openai.com> * Make the five-finding CEO fixture's application boundary explicit Materialize the request adapter and service composition used by the synthetic payment application. Explicitly declare the revised unregistered-event and mail-telemetry assumptions while preserving uncaught handler errors, the original invoice path and all five unresolved findings. Co-authored-by: OpenAI Codex <noreply@openai.com> * Keep CEO state-path checks scoped to directory preparation Co-authored-by: OpenAI Codex <noreply@openai.com> * Use checked ports and bounded cleanup in pair-agent tests Discover the daemon port from its owned state file, retain startup diagnostics, and await failed-start cleanup. Add occupied-port, early-exit, deadline, and foreign-state regressions while preserving the existing HTTP assertions and hook budgets. Co-authored-by: Codex <noreply@openai.com> * Preserve queued edit identity and recover clipped Bash permissions Distinguish separately queued unfinished edits from mutation of one native tool ID. Keep grants bound to an exact owned request and reject reused IDs, ambiguous inputs, and competing owners. Support the pinned renderer's literal em dash and request a repaint when only the Bash card's top rule is clipped. Grants still require the complete fresh card and an exact native acknowledgment. Validation: 413 integrated parser/event tests passed; private repaint controls and joint source review passed. Full canonical suite and native periodic rerun remain pending. Co-authored-by: Codex <noreply@openai.com> * Keep periodic reviews within their approved contracts and deliverables Carry exact approvals through engineering review, preserve declared contracts when amending CEO plans, and keep prioritization at the requested decision level. Materialize the revised synthetic SDK reference contract while retaining the five original documentation gaps. Accept the observed semicolon in the finite DX handoff menu and register the direct source dependencies used by the engineering cases. Regenerate canonical review documents without changing model budgets, retries, count bands, or native completion assertions. Validation: all-host generation and 275 review, fixture, selection and parity tests passed. Full free-suite and native periodic validation remain pending. Co-authored-by: Codex <noreply@openai.com> * Keep Eng approval cadence and independence guards explicit * Accept ordinary punctuation in manual review handoffs * Recover file permissions alongside queued Bash calls * Carry approved DX work through later review findings * Clarify the synthetic auth internal failure decision * Bound the periodic DX fixture to onboarding changes * Recognize native Design review handoff labels * Hold scope in the integration-choice review fixture * Carry approved Design decisions through review evidence * Capture listener state when feedback reload fails * Exclude workspace caches before checking deprecated flags * Verify Design UI scope against a seeded review plan * Clarify plan review decisions and outside-voice approval flow * Reject setup menus in the Design UI gate * docs: require focused repair validation before final acceptance * fix: separate review commitments within existing prompt budgets * docs: align generation and contributor validation guidance * fix: advance native review prompts and count acknowledged findings * chore: bump version and changelog (v1.87.1.0) Co-Authored-By: OpenAI Codex <noreply@openai.com> * chore: enforce cheap checks and side-effect-free validation previews * fix: handle owned Fetch permissions and oversized native cards * test: ground review fixtures in independent executable contracts * fix: preserve review decisions and verify reports before completion * test: construct the synthetic credential URL without a scanner false positive * test: materialize DX examples and verify their actual local behavior * fix: clarify CEO review decisions and execution order * fix: clarify review workflow ordering and select Design quality checks * Fix review decision gates and incomplete evaluation fixtures Persist CEO and engineering commitment ledgers before menus, preserve exact approvals, and distinguish implementation structure from feature scope. Route Autoplan through the canonical CEO Step 0 ordering. Classify DX findings before requesting approval and ground runtime claims in actual evidence. Complete neutral non-target fixture contracts and accept the captured Design handoff purpose without relaxing its ownership or acknowledgment checks. Record runtime-capability verification in AGENTS.md validation discipline. Validation: 1,335 focused tests passed across 21 files; build, all-host freshness, skill validation (647 artifacts / 107 tracked), and credential checks passed. Prior paid failures are preserved; behavioral acceptance remains pending. * Fix review decision boundaries and owned Read prompts Preserve exact approvals across review options, compare consistent DX milestones, and keep proposed implementation separate from review evidence. Bind modern Read prompts to one immutable native request and wait for its result. Retain captured regression verdicts, correct fixture error names, improve import probe diagnostics, and record focused-first validation discipline in AGENTS.md. * Clarify CEO and engineering review decisions Use explicit decision steps, one engineering ledger, and clear scope/write transitions. Preserve exact approvals and distinguish pending test requirements. Keep unrelated generated content unchanged. * Fix review decision ordering and native evaluation interactions * Clarify engineering decisions and test artifact order * Clarify pending choices and approvals in CEO reviews * Make CEO review phases sequential and clarify completion * Fix Design board submission intent matching * Seed an existing browser test baseline for Autoplan * Document decision-log payloads before state initialization * Preserve exact review scope and decide one change before drafting options * Require input identity before repeating passing model judges * Honor permitted storage throughout CEO review completion * Match complete native permission text within the pinned renderer contract * Align review approvals, independent choices, and bounded validation * fix: preserve reopened approvals and declare fixture interfaces * fix: isolate review artifacts and audit complete questions * fix: match detector artifact permissions to configured storage * fix: complete native permissions and review fixture workflows * fix: order CEO review work and separate engineering guarantees * fix: preserve native validation and separate review choices * fix: clarify review decisions and judge complete report context * fix: constrain review judgments and retain parse failures * fix: compare each affected value before review decisions * fix: make engineering review decisions and completion order explicit * fix: give the complete Autoplan evaluation a bounded chain budget * fix(cso): diagnose forbidden Docker endpoints before tool lookup * fix(reviews): reconcile workflow contracts and generated artifacts after main integration * fix(evals): migrate retained regressions to the native review harness * fix(tests): close native harness and workflow integration regressions * fix(evals): preserve complete permission context and native menu contracts * fix(tests): capture synchronous command output without pipe drain stalls * fix(reviews): clarify decision and completion ordering * fix(reviews): separate decision readiness from final completion checks * refactor(reviews): consolidate decision rules and completion branches * fix(plan-eng-review): order preparation and clarify decision routing * fix(plan-eng-review): restore size and question-format guard parity * fix(plan-eng-review): clarify scope phases and blocked completion * fix(plan-eng-review): unify review flow and report destination * fix(plan-eng-review): define bootstrap and question stage ownership * fix(plan-eng-review): clarify review structure and design lookup * fix(plan-eng-review): render report examples and show saved decisions * fix: consolidate Eng review decisions and select their evaluations * test: cover overlapping terminal attachments and clean merged runner type * fix: preserve Office Hours relationship closings during review updates * fix: retain pasted review targets across slash invocations * docs: preserve validation traces and correct release scope * test: cover pasted targets in both review skills * fix: validate report artifacts before recording success * fix: redact source roots at CSO report boundaries * fix: bind native Design questions before answering * test: select report privacy and native recovery regressions * test: bind rejection predicate in extracted observers * fix: bind complete boxed native questions * test: keep the Design UI fixture on native review * fix: preserve review decisions and evaluation completion outcomes * fix: clarify CEO approval and report completion order * fix: align native review evaluation ownership and completion * fix: bind review evaluators to native decisions and owned artifacts * fix: validate review decisions against native outcomes * fix: preserve review evidence and Autoplan phase handoffs * test: bind review evidence to owned decisions and completion * fix: retain owned native history across compaction * fix(evals): validate current review decisions and setup choices * fix: bind Autoplan reviews and phase completion to current amended input * fix: reconcile native review evidence and close Autoplan phases * test: recognize owned whole-candidate complexity decisions * test: preserve report freshness for approved investigation handoffs * fix: recognize scoped review findings and isolate dual voice fixtures * fix: make review handoffs and question dispatch self-contained * test: recognize complete CEO decisions and procedural pauses * fix: bind current CEO comparison options and risk intervals * test: bind engineering decisions and completion to owned evidence * fix: publish Autoplan phase reports before continuing tools * test: verify actual Autoplan dual-review dispatch evidence * test: select dual review when shared evidence fixtures change * fix: clarify plan review decisions and completion gates * fix: make CEO review decisions and return paths explicit * test: keep Autoplan prompt files inside attempt state * test: preserve source whitespace across permission dialog wraps * fix: publish Autoplan phase reports before continuing * test: recognize current CEO comparisons and reject inactive records * fix: reconcile engineering decision states before completion * test: recognize complete Design decisions and reports * test: verify current engineering decisions before navigation * Recognize source-owned component reduction choices * fix: recognize current CEO ledger and commitment grids * test: supply RequestPolicy context to Eng count fixture * fix: save complete engineering decisions before asking * fix: bind Autoplan publication to the complete phase readback * chore: prepare 1.87.5.0 reliability release * fix: clarify engineering review completion and preserve log failures * fix: bind CEO saved choices and current section ancestry * fix(evals): bind review execution and completion evidence * fix(plan-ceo-review): verify complete decisions before asking * fix(evals): preserve complete engineering choice records * fix(evals): preserve complete review outcomes and bounded fixtures * fix(autoplan): publish phase reports before advancing * fix(plan-ceo-review): validate option fields before asking * fix(plan-eng-review): verify current decisions after answers * fix(evals): bind review decisions and bound fixture scope * fix(plan-ceo-review): verify decision rows and edit saved checkpoints * fix(evals): bind review evidence and scope document lookup * fix(plan-eng-review): update resolution state with its answer * fix(reviews): preserve complete questions through dispatch * fix(evals): recognize completed mode declarations * fix(evals): define cache consistency at wrapper completion * fix(evals): validate owned initial scope and completed review handoffs * fix: assemble complete CEO decision fields before saving * fix: authenticate automatic mode decisions without guessing selectors * fix: bind engineering coverage to approved regression contracts * fix(evals): supply review helpers to native Eng capture * fix(plan-eng-review): preserve the full selected option scope * fix(evals): recognize owned engineering seed and regression evidence * fix(evals): bind engineering retry reports to native approvals * docs: clarify release guarantees (v1.87.5.0) Co-Authored-By: OpenAI Codex <noreply@openai.com> * fix(evals): recognize owned engineering decisions and handoffs * fix(evals): bind engineering decisions and completion evidence * fix(tests): align review contracts and selection fixtures * fix(skills): restore review prompt size limits * fix(plan-eng-review): clarify review execution and completion * fix(evals): preserve configured retries through all supervision layers * Clarify Engineering decisions and report completion * Keep native decision assertions within their source boundary * fix: recognize owned engineering decisions and completed navigation * fix: bind completed auto decisions to their current review * fix: recognize explicit CEO source attribution * fix: dispatch verified CEO decisions without recomposing fields * test: expose existing execution deadlines to review actors * fix: distinguish CEO decision records from incidental headings * test: bind split-scope choices to the registered native actor * test: connect reviewed regressions to required evaluation coverage * Clarify CEO decision routing and completion stages * test: expose existing section review deadlines to fixture actors * test: recognize complete native CEO pacing inventories * test: exclude answered history from current CEO payloads * test: detect phase entry through owned skill HOME aliases * test: validate native review completion and owned report permissions * fix: make Autoplan close packets carry the parent handoff steps * test: assess source-bound HOLD decisions within the existing deadline * fix: keep CEO native decision fields under one formatting authority * test: register integrated review and permission dependencies * test: align native review adapters and finding coverage Preserve explicit AUTO decisions, apply native single-select defaults, and bind complete cropped questions and report permissions to their owned requests. Require seeded review findings instead of crediting setup menus. Keep captured failure controls and additive selection dependencies. The integrated candidate passed 3,099 focused tests across 65 files; affected paid validation remains required before publication. * fix(autoplan): require phase reports before advancing * fix(evals): bind setup and evidence to complete attempts * fix(evals): bind native answers and pending writes to fixture scope Preserve complete option rows when native descriptions wrap, retain current owned Write arguments before journal publication, and keep engineering and DX answers within their declared fixture interfaces. Add captured free regressions without increasing model budgets or relaxing completion checks. * fix(autoplan): verify phase reports across native tool paths Guard owned methodology reads and reviewer dispatches, detect complete driver loads through Bash, and distinguish report-only edits from implementation changes. Follow authenticated native UUID ancestry when journal writes arrive out of order and verify earlier native content for cached phase reads. Keep current close acknowledgment and parent publication in order, require CEO entry before later phases, and register captured failure regressions. * fix(evals): honor native input and collection lifecycles Match complete native Edit panes and truncated question borders, reject stderr close before EOF, and stop the CEO split fixture once its acknowledged scope decisions are collected. Keep semantic validation, process failures, report requirements, and absolute deadlines authoritative. Add captured-event and real-process regressions with selection dependencies. Focused checks pass; final integrated paid and full-suite acceptance remain pending. * fix(autoplan): retain native session ownership across directory changes Recover missed native UUID ancestry through the existing strict graph while preserving ordinary event order and legacy scoping. Bind publication hooks to Claude's original project directory while retaining current cwd for requested file paths. Captured public-event regressions, existing caller checks, and a pinned native CLI loopback verify both fixes. Preserve failed attempts and require fresh paid and final full-suite acceptance. * docs: align evaluation limits and completion version * fix(autoplan): allow authenticated phase reads during journal streaming * fix(evals): bind clipped native questions and owned edit dialogs * fix: preserve overlay retries and bounded cleanup * fix: recognize owned planning preludes in native questions * docs: explain overlay scheduling and cleanup guarantees * fix: require fresh publication after Autoplan phase reruns * Release gstack 1.87.6 * fix: preserve CI paths, process identity, and test deadlines * fix: keep informational setup commands independent of install probes * fix: clarify plan review decisions and bound source audit reports * Fix remaining Windows identity and native path CI failures * Clarify CEO review decision and reviewer-result routing * test: accept no-install planner in retry supervision * fix(ceo-review): make review decisions and report completion explicit * perf(test): add fast PR gates, input-keyed judge reuse and isolated free shards * fix(test): start isolated CEO smoke from its existing project plan * fix(test): repair CI fixture races and preserve retry evidence * fix(ceo-review): clarify approvals, depth and saved completion --------- Co-authored-by: OpenAI Codex <noreply@openai.com>
1903 lines
90 KiB
TypeScript
Executable File
1903 lines
90 KiB
TypeScript
Executable File
#!/usr/bin/env bun
|
||
/**
|
||
* test-free-shards — enumerate, shard, curate, and run the free test suite.
|
||
*
|
||
* Four jobs:
|
||
* 1. Enumeration. Walk `browse/test/`, `test/`, `make-pdf/test/` and return
|
||
* every `*.test.{ts,tsx,js,jsx,mjs,cjs}` that isn't a paid-eval test.
|
||
* 2. Sharding. Duration-pack local and isolated CI runs. Legacy --shard
|
||
* selection retains stable hash assignment.
|
||
* 3. Curation (Windows-safe filter). Scan each test's content for POSIX-only
|
||
* patterns (`/bin/bash`, `sh -c`, raw `/tmp/`, `chmod`, `xargs`). Files
|
||
* that match are excluded from the Windows-safe subset — they would fail
|
||
* on `windows-latest` no matter how the runner shards them.
|
||
* 4. Execution. Spawn `bun test` children and refuse to trust their exit
|
||
* code alone: every byte of output is classified through
|
||
* scripts/test-strict-output.ts, so a child that exits 0 without bun's
|
||
* terminal summary (a mid-suite process.exit truncation), with `(fail)`
|
||
* result lines, or with fewer files run than planned is a FAILURE. An
|
||
* external wall-clock timeout SIGKILLs the child's process group and
|
||
* reports the shard as timed-out — distinct from failed.
|
||
*
|
||
* Execution strategy (decision ledger V3/D6 — evaluate the Bun built-in
|
||
* first; probed 2026-08 on Bun 1.3.13):
|
||
* - Full-suite runs (`bun test` via package.json, `bun run test:free`) use
|
||
* N CONCURRENT SHARD PROCESSES, serial within each (the paid runner's
|
||
* model). A single `--parallel` invocation was probed and initially
|
||
* adopted, then abandoned: three distinct Bun 1.3.13 worker pathologies
|
||
* (segfault + crash-retry wedge, skipped-file hooks stalling a worker,
|
||
* spawn-heavy files hanging under load) each stalled the whole
|
||
* invocation, while process shards isolate any wedge to its own shard.
|
||
* Original --parallel probe results, kept for the record: it
|
||
* showed --parallel (a) prints the standard `Ran N tests across M files`
|
||
* terminal summary, (b) exits non-zero when any file fails, (c) runs each
|
||
* file in its own worker process (distinct pids, no shared globals), and
|
||
* (d) converts a mid-suite process.exit(0) — which silently truncates a
|
||
* serial run at exit 0 — into a per-file `(crashed: exited)` failure with
|
||
* a complete summary and exit 1. Strictly SAFER than the serial path and
|
||
* ~2x faster on a 6-file probe (0.22s -> 0.11s wall, 280% CPU); the win
|
||
* grows with suite size since the serial suite measured 454s.
|
||
* - Legacy runs (`--shards M --shard i`) keep the hash-partitioned
|
||
* one-child-per-shard path. Cross-runner partitioning must be
|
||
* deterministic and per-file stable, so bun's own `--shard=M/N`
|
||
* (round-robin over sorted paths — every assignment shifts when a file
|
||
* lands) is not used, and there are no static per-file weight lists.
|
||
* Shard indices are STABLE: assignFilesToShards never renumbers on
|
||
* occupancy, and an empty shard is a fast no-op success.
|
||
*
|
||
* Adapted from the McGluut/gstack fork's test-free-shards.ts (190 LOC). The
|
||
* Windows-safe filter is upstream-original — codex flagged that sharding alone
|
||
* doesn't fix POSIX-bound tests, so we curate the subset that actually runs
|
||
* on the windows-latest CI job.
|
||
*
|
||
* Output contract (v1.66): the full child stream ALWAYS lands in a per-run
|
||
* log file under os.tmpdir() (path printed once at start and again in the
|
||
* epilogue). The console is quiet by default — only the runner's own
|
||
* [test:free] lines, `(fail)` result lines, bun error/crash markers
|
||
* (`error:`, `panic:`, `crashed`, `Unhandled error`), and the terminal
|
||
* `Ran N tests across M files` summary reach it; `--verbose` restores full
|
||
* forwarding. After every run a stable epilogue names the failing tests
|
||
* (attributed to files via bun's `path/to/file.test.ts:` chunk headers),
|
||
* crashed+retried workers, and — on a wall-timeout kill — the wedge-suspect
|
||
* files. The strict classifier consumes the FULL stream regardless of what
|
||
* the console shows.
|
||
*
|
||
* Exit codes: 0 pass, 1 fail, 124 wall-clock timeout.
|
||
*
|
||
* Usage:
|
||
* bun run scripts/test-free-shards.ts # full suite, N concurrent shard processes
|
||
* bun run scripts/test-free-shards.ts --list # show all
|
||
* bun run scripts/test-free-shards.ts --windows-only --list # show curated
|
||
* bun run scripts/test-free-shards.ts --windows-only # run curated
|
||
* bun run scripts/test-free-shards.ts --shards 4 --shard 1 # one shard (CI matrix)
|
||
* bun run scripts/test-free-shards.ts --wall-timeout 600 # override the kill deadline
|
||
* bun run scripts/test-free-shards.ts --verbose # forward the full child stream
|
||
* bun run scripts/test-free-shards.ts --quick # explicit fast subset, not acceptance
|
||
* bun run scripts/test-free-shards.ts --ci-plan plan.json --shards 20
|
||
* bun run scripts/test-free-shards.ts --ci-run plan.json --shard 1 --result result.json
|
||
* bun run scripts/test-free-shards.ts --ci-verify plan.json --results results/
|
||
*/
|
||
|
||
import * as fs from 'fs';
|
||
import * as os from 'os';
|
||
import * as path from 'path';
|
||
import { spawn, spawnSync } from 'child_process';
|
||
import { StringDecoder } from 'node:string_decoder';
|
||
import { createHash } from 'node:crypto';
|
||
import { isPaidTestFile } from '../test/helpers/paid-test-set';
|
||
import {
|
||
BunTestOutputClassifier,
|
||
exactTestFileSelectors,
|
||
installChildSignalForwarding,
|
||
isTerminationRequested,
|
||
killProcessGroup,
|
||
strictTestExitCode,
|
||
stripAnsiLine,
|
||
} from './test-strict-output';
|
||
|
||
const ROOT = path.resolve(import.meta.dir, '..');
|
||
// design/test was silently absent from BOTH the package.json test script and
|
||
// this list — design tests (including a teardown bomb) never ran in any CI
|
||
// or local free run. Keep the two lists in sync. This list is the single
|
||
// source of truth for free-suite roots: package.json's `test` script routes
|
||
// through this runner rather than passing its own directory globs.
|
||
export const TEST_ROOTS = [
|
||
'browse/test',
|
||
'test',
|
||
'make-pdf/test',
|
||
'design/test',
|
||
// v1.65 orphan wire-in (decision D3a): these ran under NO script or CI —
|
||
// written coverage that caught nothing. All were green on arrival.
|
||
'ios-qa/daemon/test',
|
||
'ios-qa/scripts',
|
||
'browser-skills',
|
||
] as const;
|
||
const TEST_FILE_REGEX = /\.test\.(?:[cm]?[jt]s|tsx|jsx)$/;
|
||
|
||
// POSIX-only patterns that indicate a test will fail on windows-latest no
|
||
// matter how the runner shards. Codex's v1.18.0.0 review flagged the first
|
||
// three as concrete examples in the existing free suite (test/ship-version-sync.test.ts:72,
|
||
// test/helpers/providers/claude.ts:22, package.json:12). We scan the test's
|
||
// own content here so the filter stays automatic as new tests land. The
|
||
// "Windows-incompatible APIs" patterns at the bottom were added after the
|
||
// first windows-free-tests CI run surfaced concrete failure modes.
|
||
const WINDOWS_FRAGILE_PATTERNS: Array<{ pattern: RegExp; reason: string }> = [
|
||
// Hardcoded POSIX shells / commands.
|
||
{ pattern: /['"`]\/bin\/(?:ba)?sh/, reason: 'hardcoded /bin/sh or /bin/bash' },
|
||
{ pattern: /spawnSync\(['"]sh['"],|spawn\(['"]sh['"],|exec\(['"]sh /, reason: 'spawn("sh", ...)' },
|
||
{ pattern: /['"]bash -c['"]|['"]sh -c['"]/, reason: 'bash -c / sh -c' },
|
||
{ pattern: /['"`]\/tmp\//, reason: 'raw /tmp/ path (use os.tmpdir())' },
|
||
{ pattern: /['"]chmod\b/, reason: 'chmod shell command' },
|
||
{ pattern: /['"]xargs\b/, reason: 'xargs pipeline' },
|
||
{ pattern: /\bwhich claude\b/, reason: 'which claude (use Bun.which)' },
|
||
// Windows-incompatible APIs.
|
||
{ pattern: /\.mode\s*&\s*0o[0-7]+/, reason: 'POSIX file mode bitmask (mode & 0o600 etc — Windows fakes mode bits)' },
|
||
{ pattern: /\.endsWith\(['"]\//, reason: 'hardcoded forward-slash path assertion (Windows uses \\\\)' },
|
||
{ pattern: /['"]\.\/[a-zA-Z][^"']*['"]\)\s*\.\s*toBe\(true\)/, reason: 'forward-slash path comparison' },
|
||
// Tests that spawn a bash shebang script in bin/ via spawnSync. Git Bash on
|
||
// Windows can run `bash /path/to/script` but spawnSync(scriptPath, ...)
|
||
// tries to execute the file directly via CreateProcess, which fails on the
|
||
// shebang. The pattern matches `, 'bin'` as a path-join argument (closing
|
||
// OR followed by another segment), which catches:
|
||
// - path.join(ROOT, 'bin', 'script-name') — typical
|
||
// - join(import.meta.dir, '..', 'bin', 'name') — destructured (diff-scope)
|
||
// - path.join(ROOT, 'bin') — bare BIN constant (brain-sync)
|
||
{ pattern: /,\s*['"]bin['"]\s*[,)]|['"]\.?\/?bin\/[a-z][\w-]+['"]/, reason: 'spawns bin/ shebang script (Windows CreateProcess does not parse shebangs)' },
|
||
// Tests that launch a real Playwright browser. The windows-free-tests CI job
|
||
// runs a curated subset that intentionally does NOT install Chromium —
|
||
// browser bring-up on Windows is a separate concern (see PR #1238). Tests
|
||
// matching `await foo.launch(` need Chromium and fail with "Executable
|
||
// doesn't exist" on the runner.
|
||
{ pattern: /await\s+\w+\.launch\(/, reason: 'launches Playwright browser (Chromium not installed in windows-free CI)' },
|
||
// Tests that spawn the browse server as a subprocess via `bun run server.ts`.
|
||
// The Bun → server.ts → Playwright path is the same one that doesn't work
|
||
// on Windows (PR #1238 windows-pty-bun-pty-fix). Tests typically set
|
||
// BROWSE_HEADLESS_SKIP=1 to skip the browser launch but still need a working
|
||
// server, which they don't get on Windows.
|
||
{ pattern: /BROWSE_HEADLESS_SKIP|spawn\(\[['"]bun['"],\s*['"]run['"]/, reason: 'spawns the browse server subprocess (Bun-driven path is Windows-broken)' },
|
||
];
|
||
|
||
// Explicit known-Windows-incompatible test files that don't fit a regex
|
||
// pattern. Listed here with the precise reason. Prefer adding a pattern above
|
||
// when possible; this list is for environment-/runtime-specific tests where
|
||
// the failure mode is structural rather than detectable via source-file scan.
|
||
export const KNOWN_WINDOWS_INCOMPATIBLE: Array<{ file: string; reason: string }> = [
|
||
{
|
||
file: 'test/setup-gbrain-fixture.test.ts',
|
||
reason: 'the fixture invokes real POSIX detector/verifier helpers through executable shebang wrappers',
|
||
},
|
||
{
|
||
file: 'test/hermetic-skills-seeding.test.ts',
|
||
reason: 'seeds the POSIX PTY skill runtime, whose embedded shell paths require a POSIX temporary root',
|
||
},
|
||
{
|
||
file: 'test/hermetic-wiring.test.ts',
|
||
reason: 'its runtime contract check seeds the POSIX PTY skill runtime; the curated Windows lane does not run that harness',
|
||
},
|
||
{
|
||
file: 'test/pty-workspace-trust.test.ts',
|
||
reason: 'launches the POSIX PTY harness with a fake executable and bound skill runtime',
|
||
},
|
||
{
|
||
file: 'test/host-config.test.ts',
|
||
reason: 'asserts "claude" binary on PATH (only true when running inside Claude Code, not on bare CI runner)',
|
||
},
|
||
{
|
||
file: 'browse/test/findport.test.ts',
|
||
reason: 'asserts Bun.serve.stop() is fire-and-forget — Bun behavior differs on Windows for this polyfill',
|
||
},
|
||
// First full run of the expanded lane (v1.66, 13 → ~258 files) surfaced
|
||
// seven POSIX-bound files the content patterns cannot see (their
|
||
// POSIX-ness is what they TEST, or arrives via a variable). Receipts:
|
||
// PR #2593 windows-free-tests run 31918591602.
|
||
{
|
||
file: 'test/codex-under-codex-detection.test.ts',
|
||
reason: 'drives the rendered preflight bash under a hardcoded POSIX PATH (/usr/bin:/bin) — bash is unreachable through that PATH on Windows, so every case sees empty output (v1.67 windows lane run 95234224148)',
|
||
},
|
||
{
|
||
file: 'test/regression-pr1169-build-app-sed.test.ts',
|
||
reason: 'tests sed escape sequences in build-app.sh — sed/bash are the subject under test',
|
||
},
|
||
{
|
||
file: 'test/setup-conductor-worktree.test.ts',
|
||
reason: 'tests ln -snf symlink semantics in the setup script — POSIX ln is the subject under test',
|
||
},
|
||
{
|
||
file: 'test/artifacts-init-migration.test.ts',
|
||
reason: 'runs a bash migration script + jq against a scaffolded git state — POSIX toolchain paths break under cmd spawn',
|
||
},
|
||
{
|
||
file: 'test/gstack-decision-semantic.test.ts',
|
||
reason: 'installs a fake gbrain SHEBANG SHIM on PATH; Windows spawn cannot exec shebang scripts',
|
||
},
|
||
{
|
||
file: 'test/question-log-hook.test.ts',
|
||
reason: 'spawns the PostToolUse hook script (bash shebang) directly; Windows spawn cannot exec it',
|
||
},
|
||
{
|
||
file: 'browse/test/browser-skills-e2e.test.ts',
|
||
reason: 'asserts forward-slash tier paths (<repo>/browser-skills/) that resolve with backslashes on Windows',
|
||
},
|
||
{
|
||
file: 'design/test/variants-retry-after.test.ts',
|
||
reason: 'wall-clock retry-timing assertions — flaky on the slow windows-latest runner even with widened bounds',
|
||
},
|
||
// Round-2 census (PR #2593 run 31919227507) after the first seven:
|
||
{
|
||
file: 'test/skill-census.test.ts',
|
||
reason: 'census walk throws at module load on Windows (skill-census.ts:63) — the skills-tree symlink layout needs Developer Mode that CI runners lack',
|
||
},
|
||
{
|
||
file: 'browse/test/browser-manager-unit.test.ts',
|
||
reason: 'wedges the shard to its wall deadline on windows-latest (in-flight at kill); needs a Windows repro to diagnose — macOS + Linux lanes cover the file',
|
||
},
|
||
// Round-3 census (PR #2593 run 31919871680): the round-2 wedge had been
|
||
// TRUNCATING its shard, so these seven only surfaced once shard 2 completed.
|
||
// All the same POSIX-environment classes: PID/cmdline identity probing,
|
||
// bash scripts as the subject under test, env-scrubbed child spawns.
|
||
{
|
||
file: 'browse/test/server-embedder-terminal-port.test.ts',
|
||
reason: 'identity-based terminal-agent kill probes PID/cmdline with POSIX semantics; teardown asserts fail on windows-latest',
|
||
},
|
||
{
|
||
file: 'design/test/daemon-discovery.test.ts',
|
||
reason: 'verifyIdentity matches a spawned daemon via /proc-style cmdline probing — POSIX identity semantics',
|
||
},
|
||
{
|
||
file: 'test/context-save-hardening.test.ts',
|
||
reason: 'bash context-save/migration scripts (HOME-unset semantics, random-suffix path) are the subject under test',
|
||
},
|
||
{
|
||
file: 'test/eval-list-cli.test.ts',
|
||
reason: 'spawns the eval:list CLI via bun with a constructed env — bun resolution fails under Windows spawn',
|
||
},
|
||
{
|
||
file: 'test/memory-cache-injection.test.ts',
|
||
reason: 'exercises hook/deny-enforcement shell scripts — POSIX toolchain is the subject under test',
|
||
},
|
||
{
|
||
file: 'test/migrations-v1.65.0.0.test.ts',
|
||
reason: 'bash migration script (bunx re-fetch, .done markers) is the subject under test',
|
||
},
|
||
{
|
||
file: 'test/question-preference-hook.test.ts',
|
||
reason: 'spawns the PreToolUse preference hook (shebang script) directly; Windows spawn cannot exec it',
|
||
},
|
||
// Round-4 census (PR #2593 run 31920052810): unhandled errors with no
|
||
// (fail) lines — attributed statically (the lane had no log artifact yet).
|
||
{
|
||
file: 'browse/test/browser-skill-commands.test.ts',
|
||
reason: 'spawnSkill spawns bun with a constructed env — bun resolution fails under Windows spawn (unhandled, no (fail) line)',
|
||
},
|
||
{
|
||
file: 'browse/test/security-audit-r2.test.ts',
|
||
reason: 'symlink-attack fixtures (evil-link) need Developer Mode CI runners lack; expect(toThrow) fires unhandled on Windows',
|
||
},
|
||
// CSO comprehensive execution is qualified only for Linux containers behind
|
||
// the POSIX watchdog and Unix-domain registry broker. Keep the portable
|
||
// static/parser contracts in the Windows lane while leaving these exact
|
||
// containment suites to the Linux and macOS gates.
|
||
{
|
||
file: 'test/cso-preparation-adversarial.test.ts',
|
||
reason: 'exercises POSIX prepared-tree and archive-cache containment for qualified Linux Docker execution, which Windows does not admit',
|
||
},
|
||
{
|
||
file: 'test/cso-preparation-container.test.ts',
|
||
reason: 'asserts POSIX permission and symlink semantics for inert exports consumed by qualified Linux Docker execution',
|
||
},
|
||
{
|
||
file: 'test/cso-preparation-executor.test.ts',
|
||
reason: 'executes the Linux Docker acquisition path and its Unix-domain registry broker; comprehensive execution is unavailable on Windows',
|
||
},
|
||
{
|
||
file: 'test/cso-verification-cleanup.test.ts',
|
||
reason: 'spawns the POSIX detached watchdog used by contained repair verification, which Windows intentionally leaves unavailable',
|
||
},
|
||
{
|
||
file: 'test/cso-witness.test.ts',
|
||
reason: 'tests the contained repair witness with POSIX private-directory and compiled-helper assumptions; comprehensive execution is unavailable on Windows',
|
||
},
|
||
{
|
||
file: 'test/cso-scanner-cli.test.ts',
|
||
reason: 'drives the prebuilt POSIX CSO launcher with /usr/bin/git and a POSIX-only PATH; native Windows launcher behavior is covered by the dedicated cso-windows-launcher gate',
|
||
},
|
||
];
|
||
|
||
// Force-include overrides: files a WINDOWS_FRAGILE_PATTERNS regex excludes for
|
||
// a reason that does not actually apply to them. Each entry documents WHY the
|
||
// pattern hit is a false positive — the point of these files is Windows
|
||
// coverage, so auto-excluding them defeats the regression tests they carry.
|
||
const KNOWN_WINDOWS_SAFE: Array<{ file: string; reason: string }> = [
|
||
{
|
||
file: 'test/claude-code-windows-job.test.ts',
|
||
reason: 'invokes Bun directly; verifies Windows job containment at the standalone CLI boundary',
|
||
},
|
||
{
|
||
file: 'test/claude-code-runner.test.ts',
|
||
// The bin/ path is launched through process.execPath (Bun), never as a
|
||
// shebang executable. Keep taskkill tree supervision in the Windows lane.
|
||
reason: 'invokes the runner via Bun argv; fake CLI and timeout descendant assertions cover native Windows taskkill',
|
||
},
|
||
{
|
||
file: 'test/setup-gbrain-remote-caller.test.ts',
|
||
// bin is an expected PATH component; the adapter injects the SDK boundary
|
||
// and never launches a shebang. Keep the native delimiter cases in CI.
|
||
reason: 'replays the registered SDK callback with fixture-only bin paths; covers native Windows PATH composition',
|
||
},
|
||
{
|
||
file: 'test/cso-windows-build-contract.test.ts',
|
||
// bin is a temporary staging directory. The adapter injects spawnSync;
|
||
// real PowerShell/native execution remains in cso-windows-launcher.
|
||
reason: 'replays the native build callbacks with an injected subprocess; bin paths are staging fixtures, not shebang launches',
|
||
},
|
||
{
|
||
file: 'test/setup-windows-rerun-refresh.test.ts',
|
||
// Trips the "spawns bin/ shebang script" pattern via path.join(..., 'bin',
|
||
// 'tool.sh') fixture paths, but every spawn goes through test/helpers/bash-script.ts
|
||
// (bash <tempfile>) — Git Bash executes it fine on windows-latest, with no argv-length ceiling. This file IS
|
||
// the #2444 Windows regression coverage (IS_WINDOWS=1 copy-refresh path);
|
||
// excluding it here would keep the bug class unexercised on the one
|
||
// platform it bites.
|
||
reason: 'bin/ hits are fixture path segments; spawns bash explicitly — the IS_WINDOWS=1 refresh path must run on windows-latest',
|
||
},
|
||
{
|
||
file: 'test/uninstall-windows-copies.test.ts',
|
||
// Trips the "spawns bin/ shebang script" pattern via the
|
||
// path.join(ROOT, 'bin', 'gstack-uninstall') constant, but the script is
|
||
// always spawned through spawnSync('bash', [UNINSTALL, ...]). This file
|
||
// carries the #2563 Windows real-dir-copy uninstall coverage — the bug
|
||
// ONLY reproduces on the copy install shape windows-latest exercises.
|
||
// The symlink-shape describe block self-skips on win32.
|
||
reason: 'bin/ hit is a bash-spawned script path; #2563 real-dir uninstall coverage must run on windows-latest',
|
||
},
|
||
{
|
||
file: 'browse/test/file-permissions.test.ts',
|
||
// Trips the POSIX-mode-bitmask pattern, but every `mode & 0o777` assertion
|
||
// is platform-guarded: win32-only tests return early, POSIX-only tests
|
||
// guard the bitmask behind `process.platform !== 'win32'`, and the
|
||
// symlink-skip regression test both wraps symlinkSync in try/catch
|
||
// (runners without Developer Mode can't create symlinks) and guards its
|
||
// bitmask — on win32 it asserts behavior (warns, skips, doesn't throw,
|
||
// target stays usable), never fake Windows mode bits (dirs stat 0o777
|
||
// there, so a 0o755 expectation fails on runner semantics, not our code).
|
||
// This file carries the win32-only icacls-by-SID regression tests, which
|
||
// can ONLY execute on windows-latest — excluding it here means the
|
||
// machine-account ACL lockout regression is never exercised on the one
|
||
// platform it bricks.
|
||
reason: 'every mode-bitmask assertion is guarded off win32 (behavior asserted instead); win32-only ACL regression tests must run on windows-latest',
|
||
},
|
||
{
|
||
file: 'browse/test/terminal-agent-owner-watchdog.test.ts',
|
||
// Trips the spawn(['bun','run',...]) pattern, whose reason is the
|
||
// Playwright-bound browse server. This test spawns terminal-agent.ts,
|
||
// which imports only fs/path/crypto + local helpers (no Playwright, no
|
||
// PTY at module scope) and boots under Bun on Windows — the owner-PID
|
||
// orphan leak it pins was reported on Windows (#2019).
|
||
reason: 'spawns terminal-agent (no Playwright), not the browse server; owner-orphan leak is a Windows defect',
|
||
},
|
||
];
|
||
|
||
export const DEFAULT_SHARD_COUNT = 20;
|
||
// Per-test timeout passed to `bun test --timeout`. 30s matches what
|
||
// package.json's `test` script used before it was repointed at this runner —
|
||
// the runner is now the single owner of that semantic.
|
||
export const FREE_TEST_TIMEOUT_MS = 30_000;
|
||
// External wall-clock deadline per spawned child (whole shard or the single
|
||
// full-suite --parallel invocation). A wedged child — a spinning main thread
|
||
// no in-process --timeout timer can interrupt — is SIGKILLed at the group
|
||
// level and reported 'timed-out', distinct from 'failed'.
|
||
// ~3.5x the observed full-suite wall (~100-160s). A wedged run should be
|
||
// killed-and-diagnosed (the epilogue prints the in-flight suspects) in
|
||
// minutes, not sat out — 15min of silence was pure diagnosis latency.
|
||
// Override per run with --wall-timeout <secs>.
|
||
export const DEFAULT_WALL_TIMEOUT_MS = 6 * 60_000;
|
||
/**
|
||
* Full-suite shards scale their wall deadline with shard size:
|
||
* max(DEFAULT_WALL_TIMEOUT_MS, files × PER_FILE_WALL_MS). The 6-min floor
|
||
* keeps wedge diagnosis fast on a typical ~70-file local shard, while a
|
||
* low-core machine (jobs=1 → the whole suite in one shard) or the Windows
|
||
* lane (~130 files/shard) gets proportional headroom instead of a false
|
||
* timed-out kill of a healthy run. Explicit --wall-timeout disables scaling.
|
||
*/
|
||
export const PER_FILE_WALL_MS = 5_000;
|
||
export function wallTimeoutForShard(fileCount: number, baseMs = DEFAULT_WALL_TIMEOUT_MS): number {
|
||
return Math.max(baseMs, fileCount * PER_FILE_WALL_MS);
|
||
}
|
||
|
||
/**
|
||
* Wall for a duration-packed shard. The count heuristic above assumes count
|
||
* approximates cost; LPT packing breaks that BY DESIGN (a shard may hold six
|
||
* slow Playwright files), so packed shards get max(base, predicted x 3) —
|
||
* generous against seed drift, still bounded.
|
||
*/
|
||
export function wallTimeoutForPackedShard(predictedMs: number, baseMs = DEFAULT_WALL_TIMEOUT_MS, fileCount = 0): number {
|
||
// Predictions transfer badly across machines: the committed duration seed
|
||
// is recorded on fast CI, and a syscall-supervised sandbox replays those
|
||
// files 2-4x slower (observed: a 253-file shard predicted ~242s wall-killed
|
||
// at its 725s predicted-x3 wall while genuinely still progressing). The
|
||
// packed wall may therefore be LOOSER than the count heuristic, never
|
||
// tighter — it keeps the per-file floor the runner has always guaranteed.
|
||
return Math.max(baseMs, Math.ceil(predictedMs * 3), fileCount * PER_FILE_WALL_MS);
|
||
}
|
||
/**
|
||
* Full-suite parallelism: leave RESERVED_CPUS cores for the parent runner +
|
||
* OS, cap at MAX_FULL_SUITE_JOBS — beyond ~6 concurrent bun processes the
|
||
* playwright-heavy shards contend on browser launches instead of finishing
|
||
* sooner (measured on an M-series dev box).
|
||
*
|
||
* GSTACK_FREE_JOBS overrides the computed count (the free runner's analogue
|
||
* of the paid runner's EVALS_JOBS). Exists for syscall-supervised sandboxes:
|
||
* on Vercel sandboxes, PID 1 (sandbox-init) installs a seccomp filter whose
|
||
* user-space supervisor saturates under ~6 concurrent bun+playwright shards
|
||
* and starts returning EACCES from plain file syscalls (measured: 200/200
|
||
* `git init` probes in fresh mktemp dirs fail with
|
||
* "Cannot access work tree: Permission denied" while the suite runs, 0/200
|
||
* when idle — access(dir, X_OK) = EACCES under strace). Fewer shards keep
|
||
* the supervisor inside its budget. Not clamped by MAX_FULL_SUITE_JOBS so a
|
||
* beefy box can also raise it deliberately.
|
||
*/
|
||
export const MAX_FULL_SUITE_JOBS = 6;
|
||
export const RESERVED_CPUS = 2;
|
||
|
||
export function fullSuiteJobs(): number {
|
||
const raw = process.env.GSTACK_FREE_JOBS;
|
||
if (raw !== undefined && raw !== '') {
|
||
// Strict digits-only: parseInt would silently truncate "2abc" -> 2 and
|
||
// "3.7" -> 3, defeating the loud-failure contract the error text claims.
|
||
if (!/^\d+$/.test(raw.trim()) || Number.parseInt(raw, 10) <= 0) {
|
||
throw new Error(`GSTACK_FREE_JOBS must be a positive integer, got: ${raw}`);
|
||
}
|
||
return Number.parseInt(raw, 10);
|
||
}
|
||
return Math.max(1, Math.min(MAX_FULL_SUITE_JOBS, os.cpus().length - RESERVED_CPUS));
|
||
}
|
||
|
||
/**
|
||
* Files that crash or wedge Bun's --parallel WORKERS but run fine in a plain
|
||
* serial process. Full-suite mode now uses shard PROCESSES (no workers), so
|
||
* this list is inert placement-wise — retained as the paper trail of why the
|
||
* one-invocation --parallel strategy was abandoned, and as the exclusion list
|
||
* should anyone re-attempt it on a newer Bun.
|
||
*/
|
||
export const WORKER_HOSTILE: Record<string, string> = {
|
||
'browse/test/security-live-playwright.test.ts':
|
||
'Bun 1.3.13 segfaults running this file in a --parallel worker ("panic: '
|
||
+ 'Segmentation fault ... a bug in Bun"), and the crashed-worker retry then '
|
||
+ 'wedges the whole invocation past the wall clock. Passes serially.',
|
||
};
|
||
|
||
/**
|
||
* TREE-SERIAL files: run in ONE serial shard AFTER the parallel shards.
|
||
* EMPTY since the 2026-08 dissolution — kept as a mechanism, not a museum:
|
||
* a test that must regenerate shared repo artifacts IN PLACE (and cannot
|
||
* render into an out-dir instead) earns an entry here with a reason, and
|
||
* the runner will serialize it again.
|
||
*
|
||
* How it emptied: gen-skill-docs gained a main() guard (imports stopped
|
||
* regenerating 71 files at load) and --out-dir grew to every host, so all
|
||
* eight mutators now render into mkdtemps — the live tree is never written
|
||
* by the suite (pinned by gen-skill-docs-import-purity + each migrated
|
||
* file's own porcelain/mtime assertions). With zero mutators, the four
|
||
* ratchet READERS (parity caps, size budgets, carve parity/ordering) get a
|
||
* quiet tree by construction in any shard, so they rejoined the parallel
|
||
* phase — the ~35-40s serial tail on every full-suite run is gone.
|
||
* Keys are pinned against the live file census by test-free-shards.test.ts —
|
||
* a renamed file fails the suite instead of silently dropping serialization.
|
||
*/
|
||
export const TREE_MUTATING: Record<string, string> = {};
|
||
|
||
export function normalizeRelativePath(filePath: string): string {
|
||
return filePath.replace(/\\/g, '/');
|
||
}
|
||
|
||
export function isFreeTestFile(relativePath: string): boolean {
|
||
const normalized = normalizeRelativePath(relativePath);
|
||
if (!TEST_FILE_REGEX.test(normalized)) return false;
|
||
return !isPaidTestFile(normalized);
|
||
}
|
||
|
||
/**
|
||
* Returns the first POSIX-only pattern hit in the file, or null if Windows-safe.
|
||
*/
|
||
export function detectWindowsFragility(absolutePath: string): { reason: string } | null {
|
||
let content: string;
|
||
try {
|
||
content = fs.readFileSync(absolutePath, 'utf-8');
|
||
} catch {
|
||
return null;
|
||
}
|
||
for (const { pattern, reason } of WINDOWS_FRAGILE_PATTERNS) {
|
||
if (pattern.test(content)) return { reason };
|
||
}
|
||
return null;
|
||
}
|
||
|
||
function walkTestFiles(dirPath: string): string[] {
|
||
const entries = fs.readdirSync(dirPath, { withFileTypes: true });
|
||
const files: string[] = [];
|
||
for (const entry of entries) {
|
||
const fullPath = path.join(dirPath, entry.name);
|
||
if (entry.isDirectory()) {
|
||
files.push(...walkTestFiles(fullPath));
|
||
continue;
|
||
}
|
||
if (TEST_FILE_REGEX.test(entry.name)) {
|
||
files.push(fullPath);
|
||
}
|
||
}
|
||
return files;
|
||
}
|
||
|
||
export function collectFreeTestFiles(rootDir = ROOT): string[] {
|
||
const discovered = new Set<string>();
|
||
for (const testRoot of TEST_ROOTS) {
|
||
const absoluteRoot = path.join(rootDir, testRoot);
|
||
if (!fs.existsSync(absoluteRoot)) continue;
|
||
for (const fullPath of walkTestFiles(absoluteRoot)) {
|
||
const relativePath = normalizeRelativePath(path.relative(rootDir, fullPath));
|
||
if (isFreeTestFile(relativePath)) {
|
||
discovered.add(relativePath);
|
||
}
|
||
}
|
||
}
|
||
return [...discovered].sort();
|
||
}
|
||
|
||
export interface CurationResult {
|
||
safe: string[];
|
||
excluded: Array<{ file: string; reason: string }>;
|
||
}
|
||
|
||
export function curateWindowsSafe(files: string[], rootDir = ROOT): CurationResult {
|
||
const safe: string[] = [];
|
||
const excluded: Array<{ file: string; reason: string }> = [];
|
||
const knownBad = new Map(KNOWN_WINDOWS_INCOMPATIBLE.map((e) => [e.file, e.reason]));
|
||
const knownSafe = new Set(KNOWN_WINDOWS_SAFE.map((e) => e.file));
|
||
for (const relativePath of files) {
|
||
const knownReason = knownBad.get(relativePath);
|
||
if (knownReason) {
|
||
excluded.push({ file: relativePath, reason: knownReason });
|
||
continue;
|
||
}
|
||
if (knownSafe.has(relativePath)) {
|
||
safe.push(relativePath);
|
||
continue;
|
||
}
|
||
const absolute = path.join(rootDir, relativePath);
|
||
const fragility = detectWindowsFragility(absolute);
|
||
if (fragility) {
|
||
excluded.push({ file: relativePath, reason: fragility.reason });
|
||
} else {
|
||
safe.push(relativePath);
|
||
}
|
||
}
|
||
return { safe, excluded };
|
||
}
|
||
|
||
export function stableHash(input: string): number {
|
||
let hash = 0x811c9dc5;
|
||
for (let index = 0; index < input.length; index += 1) {
|
||
hash ^= input.charCodeAt(index);
|
||
hash = Math.imul(hash, 0x01000193);
|
||
}
|
||
return hash >>> 0;
|
||
}
|
||
|
||
/**
|
||
* Hash-partition files across EXACTLY shardCount shards. Empty shards are
|
||
* preserved: a file's shard index is a pure function of its own path and the
|
||
* shard count, never of which other files happen to exist. A CI matrix keys
|
||
* runners off the index, so filtering empty shards (the old behavior) would
|
||
* renumber every later shard whenever occupancy shifted — runner 3 silently
|
||
* running shard 4's files. An empty shard is instead a fast no-op success at
|
||
* run time.
|
||
*/
|
||
export function assignFilesToShards(files: string[], shardCount: number): string[][] {
|
||
if (!Number.isInteger(shardCount) || shardCount <= 0) {
|
||
throw new Error(`Shard count must be a positive integer. Received: ${shardCount}`);
|
||
}
|
||
|
||
const shards = Array.from({ length: shardCount }, () => [] as string[]);
|
||
for (const file of files) {
|
||
const shardIndex = stableHash(file) % shardCount;
|
||
shards[shardIndex].push(file);
|
||
}
|
||
|
||
return shards.map(filesInShard => filesInShard.sort());
|
||
}
|
||
|
||
// ─── Duration-aware packing (local full suite and explicit CI plans) ───────
|
||
// Hash sharding balances file COUNTS (~1.15x spread) but not cost: the 15
|
||
// Playwright-launching files land 4/3/4/1/2/1 across 6 shards, giving a
|
||
// measured 28s–97s shard spread and ~40s of idle tail on every run. LPT
|
||
// packing over recorded per-file durations reclaims most of it. The `--shard`
|
||
// legacy path is deliberately untouched — its contract is stable indices
|
||
// via assignFilesToShards/stableHash (empty shards no-op; see above).
|
||
//
|
||
// One store, no overlay: durations come from the committed seed
|
||
// (scripts/free-test-durations.json), refreshed occasionally via
|
||
// `--record-durations` (each file timed in its own child — exact, and immune
|
||
// to bun's stream buffering, where silent passers print no header to
|
||
// timestamp). GSTACK_FREE_TEST_DURATIONS overrides the path for experiments.
|
||
// The seed is a HINT, not a contract: missing file → hash-shard fallback;
|
||
// unknown file → 75th-percentile pessimism (placed early by LPT, bounding
|
||
// tail risk). CI shares one plan rather than independently recomputing it.
|
||
|
||
export const FREE_TEST_DURATIONS_FILE = 'scripts/free-test-durations.json';
|
||
|
||
export function loadFreeTestDurations(rootDir = ROOT): Record<string, number> | null {
|
||
const file = process.env.GSTACK_FREE_TEST_DURATIONS
|
||
?? path.join(rootDir, FREE_TEST_DURATIONS_FILE);
|
||
let raw: string;
|
||
try {
|
||
raw = fs.readFileSync(file, 'utf-8');
|
||
} catch {
|
||
return null; // no seed — hash sharding, silently (fresh checkouts are normal)
|
||
}
|
||
try {
|
||
const parsed = JSON.parse(raw) as { durations?: Record<string, unknown> };
|
||
const entries = Object.entries(parsed.durations ?? {})
|
||
.filter((entry): entry is [string, number] =>
|
||
typeof entry[1] === 'number' && Number.isFinite(entry[1]) && entry[1] >= 0);
|
||
if (entries.length === 0) return null;
|
||
return Object.fromEntries(entries);
|
||
} catch (error) {
|
||
// A corrupt seed (bad merge) must cost a warning, never the suite.
|
||
console.error(`[test:free] WARNING: corrupt durations seed ${file} (${(error as Error).message}) — falling back to hash sharding`);
|
||
return null;
|
||
}
|
||
}
|
||
|
||
export interface PackedShards {
|
||
shards: string[][];
|
||
/** Predicted total per shard, aligned with `shards` — feeds walls + logs. */
|
||
predictedMs: number[];
|
||
}
|
||
|
||
/**
|
||
* Longest-processing-time-first bin packing: files sorted by predicted
|
||
* duration (desc, path-stable tiebreak) each go to the currently-lightest
|
||
* shard. Deterministic for a given (files, shardCount, durations).
|
||
*/
|
||
export function packShardsByDuration(
|
||
files: string[],
|
||
shardCount: number,
|
||
durations: Record<string, number>,
|
||
): PackedShards {
|
||
if (!Number.isInteger(shardCount) || shardCount <= 0) {
|
||
throw new Error(`Shard count must be a positive integer. Received: ${shardCount}`);
|
||
}
|
||
const known = files
|
||
.map((f) => durations[normalizeRelativePath(f)])
|
||
.filter((v): v is number => typeof v === 'number')
|
||
.sort((a, b) => a - b);
|
||
// Unknown files get the 75th percentile of known durations: pessimistic, so
|
||
// LPT places them early and a surprise long-runner can't recreate the tail.
|
||
const fallback = known.length > 0 ? known[Math.min(known.length - 1, Math.floor(known.length * 0.75))] : 1;
|
||
const predicted = (f: string): number => durations[normalizeRelativePath(f)] ?? fallback;
|
||
|
||
const ordered = [...files].sort((a, b) => predicted(b) - predicted(a) || (a < b ? -1 : 1));
|
||
const shards = Array.from({ length: shardCount }, () => [] as string[]);
|
||
const loads = new Array<number>(shardCount).fill(0);
|
||
for (const file of ordered) {
|
||
let lightest = 0;
|
||
for (let i = 1; i < shardCount; i += 1) {
|
||
if (loads[i] < loads[lightest]) lightest = i;
|
||
}
|
||
shards[lightest].push(file);
|
||
loads[lightest] += predicted(file);
|
||
}
|
||
return { shards: shards.map((s) => s.sort()), predictedMs: loads };
|
||
}
|
||
|
||
export interface FreeCiPlan {
|
||
version: 1;
|
||
revision: string;
|
||
id: string;
|
||
shards: Array<{ shard: number; files: string[]; predictedMs: number }>;
|
||
}
|
||
|
||
const planDigest = (plan: Omit<FreeCiPlan, 'id'>): string =>
|
||
createHash('sha256').update(JSON.stringify(plan)).digest('hex');
|
||
|
||
/** One immutable plan is shared by isolated CI machines; never repack per job. */
|
||
export function createFreeCiPlan(files: string[], count: number, durations: Record<string, number>, revision: string): FreeCiPlan {
|
||
const readers = files.filter(file => !(file in TREE_MUTATING));
|
||
const mutators = files.filter(file => file in TREE_MUTATING).sort();
|
||
const packed = packShardsByDuration(readers, count, durations);
|
||
const shards = packed.shards.map((files, index) => ({ shard: index + 1, files, predictedMs: packed.predictedMs[index] }));
|
||
if (mutators.length) shards.push({ shard: shards.length + 1, files: mutators, predictedMs: mutators.reduce((ms, file) => ms + (durations[file] ?? 0), 0) });
|
||
const body = { version: 1 as const, revision, shards };
|
||
return { ...body, id: planDigest(body) };
|
||
}
|
||
|
||
export function validateFreeCiPlan(plan: FreeCiPlan, files: string[], revision: string): void {
|
||
const { id, version, shards } = plan;
|
||
if (version !== 1 || plan.revision !== revision || !Array.isArray(shards) || !shards.length
|
||
|| id !== planDigest({ version, revision: plan.revision, shards })) throw new Error('CI plan identity or revision mismatch');
|
||
if (shards.some((shard, index) => shard.shard !== index + 1 || !Array.isArray(shard.files)
|
||
|| !Number.isFinite(shard.predictedMs) || shard.predictedMs < 0)) throw new Error('Invalid CI shard plan');
|
||
const planned = shards.flatMap(shard => shard.files).sort();
|
||
if (new Set(planned).size !== planned.length || JSON.stringify(planned) !== JSON.stringify([...files].sort())) throw new Error('CI plan must cover every free file exactly once');
|
||
}
|
||
|
||
export interface FreeCiResult {
|
||
planId: string;
|
||
revision: string;
|
||
outcome: FreeShardOutcome;
|
||
retry: FreeShardOutcome | null;
|
||
}
|
||
|
||
/** Preserve the full-suite retry cap across independently running CI jobs. */
|
||
export function eligibleFreeRetryFiles(outcomes: FreeShardOutcome[]): string[] | null {
|
||
if (!outcomes.every(outcome => hasScopedFailureAttribution(outcome) && (outcome.status === 'passed'
|
||
|| (outcome.status === 'failed' && outcome.failingFiles.length > 0 && outcome.unattributedFailures === 0)))) return null;
|
||
const files = [...new Set(outcomes.flatMap(outcome => outcome.failingFiles))];
|
||
return files.length > 0 && files.length <= 5 ? files : null;
|
||
}
|
||
|
||
function hasScopedFailureAttribution(outcome: FreeShardOutcome): boolean {
|
||
return new Set(outcome.failingFiles).size === outcome.failingFiles.length
|
||
&& outcome.failingFiles.every(file => outcome.files.includes(file));
|
||
}
|
||
|
||
export function verifyFreeCiResults(plan: FreeCiPlan, results: FreeCiResult[]): void {
|
||
if (results.length !== plan.shards.length) throw new Error('Missing or duplicate CI shard results');
|
||
const seen = new Set<number>();
|
||
for (const result of results) {
|
||
const outcome = result.outcome;
|
||
const shard = plan.shards[outcome.shard - 1];
|
||
if (result.planId !== plan.id || result.revision !== plan.revision || !shard || seen.has(outcome.shard)
|
||
|| JSON.stringify(outcome.files) !== JSON.stringify(shard.files)) throw new Error('CI result identity, shard or file coverage mismatch');
|
||
seen.add(outcome.shard);
|
||
if (!hasCompleteCiSummary(outcome)) throw new Error('Missing or incomplete CI execution summary');
|
||
if (outcome.status === 'passed') {
|
||
if (outcome.exitCode !== 0 || outcome.failingFiles.length || outcome.unattributedFailures || result.retry) throw new Error('Inconsistent passing CI result');
|
||
} else {
|
||
const retryFiles = eligibleFreeRetryFiles([outcome]);
|
||
const retry = result.retry;
|
||
if (!retryFiles || !retry || retry.status !== 'passed' || retry.exitCode !== 0
|
||
|| retry.failingFiles.length || retry.unattributedFailures
|
||
|| !hasCompleteCiSummary(retry)
|
||
|| JSON.stringify([...retry.files].sort()) !== JSON.stringify(retryFiles.sort())) throw new Error('Failed or incomplete CI shard');
|
||
}
|
||
}
|
||
if (results.some(result => result.retry) && !eligibleFreeRetryFiles(results.map(result => result.outcome))) {
|
||
throw new Error('CI retries exceed the full-suite attribution or five-file limit');
|
||
}
|
||
}
|
||
|
||
function hasCompleteCiSummary(outcome: FreeShardOutcome): boolean {
|
||
const summary = outcome.summary;
|
||
if (!summary || !Number.isInteger(summary.testsRan) || summary.testsRan! < 0
|
||
|| summary.filesRan !== outcome.files.length) return false;
|
||
// Empty assigned shards deliberately do not launch Bun or invent a summary.
|
||
return outcome.files.length === 0
|
||
? summary.testsRan === 0 && summary.sawTerminalSummary === false
|
||
: summary.sawTerminalSummary === true;
|
||
}
|
||
|
||
export const QUICK_CORE = [
|
||
'test/strict-output.test.ts', 'test/gen-skill-docs.test.ts',
|
||
'test/skill-check-driver.test.ts', 'test/ceo-native-ledger-replay.test.ts',
|
||
'test/skill-ceo-section-ordering.test.ts',
|
||
];
|
||
|
||
export function selectQuickFreeFiles(files: string[], durations: Record<string, number>): string[] {
|
||
return files.filter(file => isFreeTestFile(file)
|
||
&& (QUICK_CORE.includes(file) || (durations[file] !== undefined && durations[file] <= 2_000)));
|
||
}
|
||
|
||
export interface BuildShardArgsOptions {
|
||
/**
|
||
* Pass bun's --parallel (worker-per-file, implies --isolate). No production
|
||
* caller today — full-suite mode uses N shard PROCESSES after the worker
|
||
* pathologies documented in main(); retained for a future re-attempt on a
|
||
* newer Bun (see WORKER_HOSTILE).
|
||
*/
|
||
parallel?: boolean;
|
||
rootDir?: string;
|
||
}
|
||
|
||
export function buildShardArgs(files: string[], options: BuildShardArgsOptions = {}): string[] {
|
||
// Exact absolute selectors: bun treats positional test paths as substring
|
||
// filters, so a relative `test/x.test.ts` would ALSO select
|
||
// `browse/test/x.test.ts` — shard bleed that double-runs files.
|
||
const selectors = exactTestFileSelectors(files, options.rootDir ?? ROOT);
|
||
const args = ['test', ...selectors, `--timeout=${FREE_TEST_TIMEOUT_MS}`];
|
||
if (options.parallel) args.push('--parallel');
|
||
else args.push('--max-concurrency=1');
|
||
return args;
|
||
}
|
||
|
||
type CliOptions = {
|
||
dryRun: boolean;
|
||
listOnly: boolean;
|
||
recordDurations: boolean;
|
||
windowsOnly: boolean;
|
||
verbose: boolean;
|
||
shardCount: number;
|
||
shardIndex: number | null;
|
||
wallTimeoutMs: number;
|
||
/** True when --wall-timeout was passed explicitly; full-suite mode only auto-scales the default. */
|
||
wallTimeoutExplicit: boolean;
|
||
quick: boolean;
|
||
ciPlan: string | null;
|
||
ciRun: string | null;
|
||
ciVerify: string | null;
|
||
result: string | null;
|
||
results: string | null;
|
||
};
|
||
|
||
function parseCliOptions(argv: string[]): CliOptions {
|
||
let dryRun = false;
|
||
let listOnly = false;
|
||
let recordDurations = false;
|
||
let windowsOnly = false;
|
||
let verbose = false;
|
||
let shardCount = DEFAULT_SHARD_COUNT;
|
||
let shardIndex: number | null = null;
|
||
let wallTimeoutMs = DEFAULT_WALL_TIMEOUT_MS;
|
||
let wallTimeoutExplicit = false;
|
||
let quick = false;
|
||
const paths: Record<'ciPlan' | 'ciRun' | 'ciVerify' | 'result' | 'results', string | null> = {
|
||
ciPlan: null, ciRun: null, ciVerify: null, result: null, results: null,
|
||
};
|
||
|
||
for (let index = 0; index < argv.length; index += 1) {
|
||
const arg = argv[index];
|
||
if (arg === '--dry-run') { dryRun = true; continue; }
|
||
if (arg === '--list') { listOnly = true; continue; }
|
||
if (arg === '--record-durations') { recordDurations = true; continue; }
|
||
if (arg === '--windows-only') { windowsOnly = true; continue; }
|
||
if (arg === '--verbose') { verbose = true; continue; }
|
||
if (arg === '--quick') { quick = true; continue; }
|
||
const pathKey = ({ '--ci-plan': 'ciPlan', '--ci-run': 'ciRun', '--ci-verify': 'ciVerify', '--result': 'result', '--results': 'results' } as const)[arg];
|
||
if (pathKey) {
|
||
const value = argv[++index];
|
||
if (!value || value.startsWith('--')) throw new Error(`Missing path for ${arg}`);
|
||
paths[pathKey] = value;
|
||
continue;
|
||
}
|
||
if (arg === '--shards') {
|
||
const value = argv[index + 1];
|
||
if (!value) throw new Error('Missing value for --shards');
|
||
shardCount = Number.parseInt(value, 10);
|
||
index += 1;
|
||
continue;
|
||
}
|
||
if (arg === '--shard') {
|
||
const value = argv[index + 1];
|
||
if (!value) throw new Error('Missing value for --shard');
|
||
shardIndex = Number.parseInt(value, 10);
|
||
index += 1;
|
||
continue;
|
||
}
|
||
if (arg === '--wall-timeout') {
|
||
const value = Number.parseInt(argv[index + 1] ?? '', 10);
|
||
if (!Number.isInteger(value) || value <= 0) throw new Error('--wall-timeout needs a positive integer (seconds)');
|
||
wallTimeoutMs = value * 1000;
|
||
wallTimeoutExplicit = true;
|
||
index += 1;
|
||
continue;
|
||
}
|
||
throw new Error(`Unknown argument: ${arg}`);
|
||
}
|
||
|
||
const ciModes = [paths.ciPlan, paths.ciRun, paths.ciVerify].filter(Boolean).length;
|
||
if (ciModes > 1 || (ciModes && (quick || listOnly || dryRun || recordDurations || windowsOnly))) throw new Error('CI modes cannot be combined with other selection modes');
|
||
if (paths.ciRun && (shardIndex === null || !paths.result)) throw new Error('--ci-run requires --shard and --result');
|
||
if (paths.ciVerify && !paths.results) throw new Error('--ci-verify requires --results');
|
||
if (quick && (recordDurations || windowsOnly || shardIndex !== null)) throw new Error('--quick cannot change recording, Windows or shard selection');
|
||
return { dryRun, listOnly, recordDurations, windowsOnly, verbose, shardCount, shardIndex, wallTimeoutMs, wallTimeoutExplicit, quick, ...paths };
|
||
}
|
||
|
||
function formatShardSummary(shards: string[][]): string[] {
|
||
return shards.map((files, index) => {
|
||
const preview = files.slice(0, 3).join(', ');
|
||
const suffix = files.length > 3 ? ', ...' : '';
|
||
return `Shard ${index + 1}/${shards.length}: ${files.length} files${preview ? ` -> ${preview}${suffix}` : ''}`;
|
||
});
|
||
}
|
||
|
||
/**
|
||
* True when a shard's output shows the run ended WITHOUT bun's final summary
|
||
* ("Ran N tests across ..."). A process.exit() fired mid-suite skips the
|
||
* summary AND hands back whatever code the caller passed — historically 0,
|
||
* which made a truncated shard indistinguishable from a green one. Exit code
|
||
* alone is therefore not evidence of completion; the summary line is.
|
||
*
|
||
* The runner itself now enforces this (and more) through
|
||
* scripts/test-strict-output.ts inside runFreeShard; this predicate remains
|
||
* the minimal documented primitive that test/exit-propagation.test.ts drives
|
||
* with genuine truncated and genuine complete bun runs.
|
||
*/
|
||
export function shardRunLooksTruncated(status: number | null, output: string): boolean {
|
||
if (status !== 0) return false; // already failing — not the silent case
|
||
return !/Ran \d+ tests? across \d+ files?/.test(output);
|
||
}
|
||
|
||
// ---------------------------------------------------------------------------
|
||
// Output contract: console filtering + per-file failure attribution.
|
||
//
|
||
// Bun groups each file's output under a `path/to/file.test.ts:` header line
|
||
// (cwd-relative, sometimes ../-prefixed through a symlinked cwd). The
|
||
// reporter tracks the current header while consuming the stream, attributes
|
||
// `(fail)` lines and crash markers to files, and decides which lines reach
|
||
// the console in the default quiet mode. All matching happens on
|
||
// ANSI-stripped lines — colored `(fail)` lines defeated a prior grep.
|
||
// ---------------------------------------------------------------------------
|
||
|
||
const TEST_PATH_SOURCE = String.raw`\.test\.(?:[cm]?[jt]s|tsx|jsx)`;
|
||
/** A file chunk header: the path bun printed, terminated by a bare colon. */
|
||
const FILE_HEADER_RE = new RegExp(`^(\\S.*${TEST_PATH_SOURCE}):$`);
|
||
/** Same shape strict-output classifies as failed-test, with the name captured. */
|
||
const FAIL_RESULT_CAPTURE_RE = /^\(fail\) (.+) \[\d+(?:\.\d+)?(?:ns|us|µs|ms|s)\]$/;
|
||
/** bun --parallel retries a crashed worker once: `<icon> crashed running <path>, retrying`. */
|
||
const CRASH_RETRY_RE = new RegExp(`crashed running (\\S*${TEST_PATH_SOURCE}), retrying`);
|
||
/** The give-up marker after the retry also crashes: `✗ <path> (crashed: exited)`. */
|
||
const CRASH_FINAL_RE = new RegExp(`(\\S*${TEST_PATH_SOURCE}) \\(crashed: [^)]+\\)`);
|
||
const TERMINAL_SUMMARY_CAPTURE_RE = /^Ran (\d+) tests? across (\d+) files?\. \[/;
|
||
/** Substrings that must reach the console even in the default quiet mode. */
|
||
const CONSOLE_ALWAYS_MARKERS = ['error:', 'panic:', 'Unhandled error', 'crashed'] as const;
|
||
|
||
export type StreamOrigin = 'stdout' | 'stderr';
|
||
|
||
export interface FreeRunFailure {
|
||
/** Planned relative path when attributable, else the raw header path, else null. */
|
||
file: string | null;
|
||
testName: string;
|
||
}
|
||
|
||
export interface FreeRunReport {
|
||
testsRan: number | null;
|
||
filesRan: number | null;
|
||
sawTerminalSummary: boolean;
|
||
/** Deduped `(fail)` lines in arrival order, attributed to the current file header. */
|
||
failures: FreeRunFailure[];
|
||
/** Files that crashed a worker (bun retries once; a second crash is final). Deduped. */
|
||
crashedFiles: string[];
|
||
/**
|
||
* "# Unhandled error between tests" markers, attributed to the chunk they
|
||
* appeared in. These fail the shard via the strict classifier but produce
|
||
* NO (fail) lines — without surfacing them here, the epilogue reads
|
||
* "FAIL — 0 failing test(s)" and the culprit is undiscoverable from CI
|
||
* output (first Windows lane run: a module-load throw in skill-census).
|
||
*/
|
||
unhandledErrors: Array<{ file: string | null }>;
|
||
/**
|
||
* Wedge-suspect heuristic for a wall-timeout kill: files whose header was
|
||
* seen but whose chunk never ENDED (chunk end = the next file's header, or
|
||
* a final crash marker) before the terminal summary — i.e. "started but
|
||
* never produced a result chunk end". Result lines deliberately do NOT end
|
||
* a chunk: a file that printed a fail and then wedged stays listed. Known
|
||
* limits of the approximation:
|
||
* - Serial (--shard CI path): bun streams live but prints a file's header
|
||
* lazily, on its first output line — a wedged file that printed ANY
|
||
* line is listed; a fully silent wedge is not.
|
||
* - Parallel (full-suite path): bun buffers a file's whole chunk until it
|
||
* COMPLETES, so a wedged file usually never prints a header (see
|
||
* filesWithNoOutput), and the LAST flushed chunk before the kill has no
|
||
* closing header, so one completed noisy file can be over-listed.
|
||
*/
|
||
inFlight: string[];
|
||
/** Planned files never observed in the stream (silent passers + never-flushed wedges). */
|
||
filesWithNoOutput: number;
|
||
}
|
||
|
||
interface FileProgress {
|
||
headerSeen: boolean;
|
||
/** The file's chunk ended: a later file's header arrived, or it crashed out. */
|
||
ended: boolean;
|
||
}
|
||
|
||
/**
|
||
* Incrementally consumes the child's stdout/stderr (chunk boundaries need not
|
||
* align to lines), attributing results to files and forwarding only
|
||
* always-visible lines to `forward` (omit `forward` for verbose/quiet modes —
|
||
* attribution still runs so the epilogue works in every mode).
|
||
*/
|
||
export class FreeRunReporter {
|
||
private readonly decoders: Record<StreamOrigin, StringDecoder> = {
|
||
stdout: new StringDecoder('utf8'),
|
||
stderr: new StringDecoder('utf8'),
|
||
};
|
||
private readonly pending: Record<StreamOrigin, string> = { stdout: '', stderr: '' };
|
||
private readonly plannedSet: Set<string>;
|
||
private readonly canonicalCache = new Map<string, string>();
|
||
private readonly progress = new Map<string, FileProgress>();
|
||
private readonly failureKeys = new Set<string>();
|
||
private readonly failures: FreeRunFailure[] = [];
|
||
private readonly crashed = new Set<string>();
|
||
private currentFile: string | null = null;
|
||
private inRecap = false;
|
||
private readonly unhandled: Array<{ file: string | null }> = [];
|
||
private testsRan: number | null = null;
|
||
private filesRan: number | null = null;
|
||
private sawSummary = false;
|
||
|
||
constructor(
|
||
private readonly plannedFiles: string[],
|
||
private readonly forward?: (text: string, origin: StreamOrigin) => void,
|
||
) {
|
||
this.plannedSet = new Set(plannedFiles.map(normalizeRelativePath));
|
||
}
|
||
|
||
write(chunk: Uint8Array | string, origin: StreamOrigin): void {
|
||
this.pending[origin] += typeof chunk === 'string'
|
||
? chunk
|
||
: this.decoders[origin].write(Buffer.from(chunk));
|
||
let newline = this.pending[origin].indexOf('\n');
|
||
while (newline !== -1) {
|
||
this.handleLine(this.pending[origin].slice(0, newline), origin);
|
||
this.pending[origin] = this.pending[origin].slice(newline + 1);
|
||
newline = this.pending[origin].indexOf('\n');
|
||
}
|
||
}
|
||
|
||
/** Flush partial trailing lines (a stream killed mid-line still classifies). */
|
||
end(): void {
|
||
for (const origin of ['stdout', 'stderr'] as const) {
|
||
this.pending[origin] += this.decoders[origin].end();
|
||
if (this.pending[origin].length > 0) this.handleLine(this.pending[origin], origin);
|
||
this.pending[origin] = '';
|
||
}
|
||
}
|
||
|
||
report(): FreeRunReport {
|
||
const inFlight = this.sawSummary
|
||
? []
|
||
: [...this.progress.entries()]
|
||
.filter(([, p]) => p.headerSeen && !p.ended)
|
||
.map(([file]) => file)
|
||
.sort();
|
||
return {
|
||
testsRan: this.testsRan,
|
||
filesRan: this.filesRan,
|
||
sawTerminalSummary: this.sawSummary,
|
||
failures: [...this.failures],
|
||
crashedFiles: [...this.crashed].sort(),
|
||
unhandledErrors: [...this.unhandled],
|
||
inFlight,
|
||
filesWithNoOutput: this.plannedFiles.filter((f) => !this.progress.has(normalizeRelativePath(f))).length,
|
||
};
|
||
}
|
||
|
||
private handleLine(rawLine: string, origin: StreamOrigin): void {
|
||
// GitHub Actions: bun wraps each file's section in ::group::<header>.
|
||
// Without stripping, the real header fails FILE_HEADER_RE, failures get
|
||
// attributed to the PREVIOUS file, and the terminal recap's re-printed
|
||
// (fail) lines land under a second phantom file (observed on the first
|
||
// Linux run: 5 real failures reported as 10 across 2 files).
|
||
const line = stripAnsiLine(rawLine).replace(/^::group::/, '');
|
||
let visible = false;
|
||
|
||
// Bun's terminal recap ("N tests failed:") re-prints every (fail) line
|
||
// WITHOUT re-printing file headers. Attributing those to the stale
|
||
// currentFile invented a phantom failing file on the first Linux run
|
||
// (5 real failures reported as 10 across 2 files, one innocent).
|
||
if (/^\d+ tests? failed:$/.test(line)) {
|
||
this.inRecap = true;
|
||
if (this.currentFile) this.progressFor(this.currentFile).ended = true;
|
||
this.currentFile = null;
|
||
}
|
||
|
||
if (line === '# Unhandled error between tests') {
|
||
this.unhandled.push({ file: this.currentFile });
|
||
}
|
||
|
||
const header = FILE_HEADER_RE.exec(line);
|
||
if (header) {
|
||
const file = this.canonicalize(header[1]);
|
||
// A new header ends the previous file's chunk — that file is no longer
|
||
// a wedge suspect. (Bun 1.3.x prints NO (pass) lines, so chunk
|
||
// delimiters, not result lines, are the completion signal.)
|
||
if (this.currentFile && this.currentFile !== file) this.progressFor(this.currentFile).ended = true;
|
||
this.currentFile = file;
|
||
this.progressFor(file).headerSeen = true;
|
||
} else {
|
||
const fail = FAIL_RESULT_CAPTURE_RE.exec(line);
|
||
const retry = fail ? null : CRASH_RETRY_RE.exec(line);
|
||
const final = fail || retry ? null : CRASH_FINAL_RE.exec(line);
|
||
if (fail) {
|
||
visible = true;
|
||
// In the recap, a (fail) line only records a failure the main run
|
||
// somehow never attributed (belt and braces); known names dedupe.
|
||
const recapDuplicate = this.inRecap
|
||
&& this.failures.some((f) => f.testName === fail[1]);
|
||
const key = `${this.currentFile ?? ''}\u0000${fail[1]}`;
|
||
if (!recapDuplicate && !this.failureKeys.has(key)) {
|
||
this.failureKeys.add(key);
|
||
this.failures.push({ file: this.currentFile, testName: fail[1] });
|
||
}
|
||
} else if (retry) {
|
||
// The file will run again — a crash+retry does not end its chunk.
|
||
visible = true;
|
||
this.crashed.add(this.canonicalize(retry[1]));
|
||
} else if (final) {
|
||
visible = true;
|
||
const file = this.canonicalize(final[1]);
|
||
this.crashed.add(file);
|
||
this.progressFor(file).ended = true;
|
||
} else {
|
||
const summary = TERMINAL_SUMMARY_CAPTURE_RE.exec(line);
|
||
if (summary) {
|
||
visible = true;
|
||
this.sawSummary = true;
|
||
this.testsRan = Number.parseInt(summary[1], 10);
|
||
this.filesRan = Number.parseInt(summary[2], 10);
|
||
}
|
||
}
|
||
}
|
||
|
||
if (!visible) visible = CONSOLE_ALWAYS_MARKERS.some((marker) => line.includes(marker));
|
||
if (visible && this.forward) this.forward(`${rawLine.replace(/\r$/, '')}\n`, origin);
|
||
}
|
||
|
||
private progressFor(file: string): FileProgress {
|
||
let entry = this.progress.get(file);
|
||
if (!entry) {
|
||
entry = { headerSeen: false, ended: false };
|
||
this.progress.set(file, entry);
|
||
}
|
||
return entry;
|
||
}
|
||
|
||
/**
|
||
* Map a printed path back to its planned relative path. Bun prints paths
|
||
* relative to the child's (real)cwd, so a symlinked cwd (macOS /tmp) yields
|
||
* `../..`-prefixed forms — strip the prefix and suffix-match.
|
||
*/
|
||
private canonicalize(printedPath: string): string {
|
||
const cached = this.canonicalCache.get(printedPath);
|
||
if (cached) return cached;
|
||
const stripped = normalizeRelativePath(printedPath).replace(/^(?:\.{1,2}\/)+/, '');
|
||
let resolved = stripped;
|
||
if (!this.plannedSet.has(stripped)) {
|
||
const match = this.plannedFiles.find(
|
||
(planned) => stripped.endsWith(`/${planned}`) || planned.endsWith(`/${stripped}`),
|
||
);
|
||
if (match) resolved = match;
|
||
}
|
||
this.canonicalCache.set(printedPath, resolved);
|
||
return resolved;
|
||
}
|
||
}
|
||
|
||
/**
|
||
* The stable post-run epilogue. Success is one line; failure names every
|
||
* failing test (deduped, attributed) and crashed worker; a wall-timeout kill
|
||
* additionally prints the wedge-suspect list (see FreeRunReport.inFlight for
|
||
* the heuristic and its limits).
|
||
*/
|
||
export function buildRunEpilogue(
|
||
status: FreeShardStatus,
|
||
report: FreeRunReport,
|
||
elapsedMs: number,
|
||
logPath: string,
|
||
): string[] {
|
||
const seconds = Math.round(elapsedMs / 1000);
|
||
if (status === 'passed') {
|
||
return [
|
||
`[test:free] PASS — ${report.testsRan ?? '?'} tests, ${report.filesRan ?? '?'} files, ${seconds}s. Full log: ${logPath}`,
|
||
];
|
||
}
|
||
const failingFiles = new Set(report.failures.map((f) => f.file ?? '(unattributed)'));
|
||
const lines = [
|
||
`[test:free] FAIL — ${report.failures.length} failing test(s) in ${failingFiles.size} file(s), `
|
||
+ `${report.crashedFiles.length} crashed worker(s)${report.unhandledErrors.length > 0 ? `, ${report.unhandledErrors.length} unhandled error(s) between tests` : ''}. Full log: ${logPath}`,
|
||
];
|
||
for (const failure of report.failures) {
|
||
lines.push(` ✗ ${failure.file ?? '(unattributed)'} — ${failure.testName}`);
|
||
}
|
||
for (const file of report.crashedFiles) {
|
||
lines.push(` ⚠ crashed+retried: ${file}`);
|
||
}
|
||
for (const u of report.unhandledErrors) {
|
||
lines.push(` ⚠ unhandled error between tests (around ${u.file ?? 'unknown file'})`);
|
||
}
|
||
if (status === 'timed-out') {
|
||
if (report.inFlight.length > 0) {
|
||
lines.push(` ⏱ in flight at kill: ${report.inFlight.join(', ')}`);
|
||
} else {
|
||
lines.push(
|
||
' ⏱ in flight at kill: unknown — no open file chunk was observed '
|
||
+ '(bun --parallel buffers a file\'s output until it completes, so a silent wedge never prints); '
|
||
+ `${report.filesWithNoOutput} planned file(s) produced no output before the kill.`,
|
||
);
|
||
}
|
||
}
|
||
return lines;
|
||
}
|
||
|
||
export type FreeShardStatus = 'passed' | 'failed' | 'timed-out';
|
||
|
||
// ─── Flake ledger (WS1 telemetry) ───────────────────────────────────────────
|
||
// Single-writer JSONL: ONLY this parent runner appends (never shards, never
|
||
// tests — no concurrent-append hazard by construction). CI points
|
||
// GSTACK_FLAKE_LEDGER at $RUNNER_TEMP and uploads it as an artifact every
|
||
// run, so repeat offenders become an enumerable series instead of console
|
||
// scrollback. Fail-open with a loud stderr warning: a broken ledger must
|
||
// never red the only required lane.
|
||
|
||
export interface FlakeLedgerEntry {
|
||
ts: string;
|
||
runner: 'free';
|
||
kind: 'flaky-pass';
|
||
file: string;
|
||
/** Shard the original failure surfaced in, when attributable. */
|
||
shard?: number;
|
||
/** Code-state attribution (review finding): without branch/sha the series
|
||
* can't tie an entry to the state that produced it, and the WS16
|
||
* promotion evidence needs exactly that. */
|
||
branch?: string;
|
||
git_sha?: string;
|
||
}
|
||
|
||
export function flakeLedgerPath(env: NodeJS.ProcessEnv = process.env): string {
|
||
if (env.GSTACK_FLAKE_LEDGER) return env.GSTACK_FLAKE_LEDGER;
|
||
// Local default: per-PROJECT, not the machine-global tmpdir — sibling
|
||
// Conductor worktrees of DIFFERENT repos must not interleave into one
|
||
// series (review finding). CI always sets GSTACK_FLAKE_LEDGER explicitly.
|
||
try {
|
||
const slug = spawnSync('bash', ['-c', '~/.claude/skills/gstack/bin/gstack-slug 2>/dev/null'], { stdio: 'pipe', timeout: 3000 })
|
||
.stdout?.toString().match(/^SLUG=(.+)$/m)?.[1];
|
||
if (slug) {
|
||
const dir = path.join(os.homedir(), '.gstack', 'projects', slug);
|
||
fs.mkdirSync(dir, { recursive: true });
|
||
return path.join(dir, 'flake-ledger.jsonl');
|
||
}
|
||
} catch { /* fall through */ }
|
||
return path.join(os.tmpdir(), 'gstack-flake-ledger.jsonl');
|
||
}
|
||
|
||
export function appendFlakeLedger(
|
||
entries: FlakeLedgerEntry[],
|
||
ledgerPath: string,
|
||
warn: (line: string) => void = (line) => console.error(line),
|
||
): boolean {
|
||
if (entries.length === 0) return true;
|
||
try {
|
||
fs.mkdirSync(path.dirname(ledgerPath), { recursive: true });
|
||
fs.appendFileSync(ledgerPath, entries.map((e) => JSON.stringify(e)).join('\n') + '\n');
|
||
return true;
|
||
} catch (error) {
|
||
warn(`[test:free] WARNING: could not append flake ledger at ${ledgerPath} `
|
||
+ `(${error instanceof Error ? error.message : String(error)}) — flaky-pass telemetry lost for this run, verdict unaffected`);
|
||
return false;
|
||
}
|
||
}
|
||
|
||
export interface FreeShardOutcome {
|
||
shard: number;
|
||
files: string[];
|
||
status: FreeShardStatus;
|
||
exitCode: number | null;
|
||
elapsedMs: number;
|
||
groupPid: number | null;
|
||
/** Required by CI receipts; optional for existing local caller fixtures. */
|
||
summary?: Pick<FreeRunReport, 'testsRan' | 'filesRan' | 'sawTerminalSummary'>;
|
||
/**
|
||
* Repo-relative files with attributed test failures or crashes, deduped.
|
||
* Feeds the opt-in flaky retry pass (GSTACK_FREE_RETRY_FLAKY) — empty on
|
||
* pass, and empty when every failure was unattributed (retry would be
|
||
* meaningless without knowing what to re-run).
|
||
*/
|
||
failingFiles: string[];
|
||
/**
|
||
* Count of failure evidence the retry pass CANNOT re-run by file: fail
|
||
* lines seen before any file-chunk header, unhandled errors between tests,
|
||
* and a truncated run (no terminal summary). Nonzero vetoes the flaky
|
||
* retry for the whole run — retrying only failingFiles would re-run a
|
||
* subset and mask the rest as a FLAKY-PASS, re-opening the silent-truncation
|
||
* hole the strict classifier exists to close.
|
||
*/
|
||
unattributedFailures: number;
|
||
}
|
||
|
||
export interface ShardCommand {
|
||
command: string;
|
||
args: string[];
|
||
}
|
||
|
||
export interface RunFreeShardOptions {
|
||
/** External wall-clock deadline; on expiry the child's process GROUP is SIGKILLed. */
|
||
wallTimeoutMs?: number;
|
||
rootDir?: string;
|
||
env?: NodeJS.ProcessEnv;
|
||
/** Pass bun's --parallel. No production caller today (see BuildShardArgsOptions.parallel). */
|
||
parallel?: boolean;
|
||
/** Override the spawned command. Tests inject fake pass/fail/slow commands. */
|
||
commandFor?: (files: string[]) => ShardCommand;
|
||
/** Suppress ALL child output from the console (tests). The classifier and the log file still see every byte. */
|
||
quiet?: boolean;
|
||
/** Forward the full child stream to the console (legacy firehose). Default: the quiet filtered console. */
|
||
verbose?: boolean;
|
||
/**
|
||
* Console sink for child-stream output (tests inject to assert quiet vs
|
||
* verbose behavior). Default: process.stdout / process.stderr by origin.
|
||
* Runner-owned [test:free] lines go through `log`, not this sink.
|
||
*/
|
||
consoleWrite?: (text: string) => void;
|
||
/** Per-run full-stream log path (tests inject). Default: a timestamped file under os.tmpdir(). */
|
||
logFilePath?: string;
|
||
log?: (line: string) => void;
|
||
}
|
||
|
||
const EPILOGUE_WORD: Record<FreeShardStatus, string> = {
|
||
passed: 'pass',
|
||
failed: 'fail',
|
||
'timed-out': 'timed-out',
|
||
};
|
||
|
||
/** One line per shard, printed after the run: `[test:free] shard i/N: M files, XXs, pass|fail|timed-out`. */
|
||
function shardEpilogue(outcome: FreeShardOutcome, totalShards: number): string {
|
||
return `[test:free] shard ${outcome.shard}/${totalShards}: ${outcome.files.length} files, `
|
||
+ `${Math.round(outcome.elapsedMs / 1000)}s, ${EPILOGUE_WORD[outcome.status]}`;
|
||
}
|
||
|
||
/**
|
||
* Run one shard (or the whole suite, in --parallel full-suite mode) in its own
|
||
* bun process and classify the result strictly.
|
||
*
|
||
* Verdict integrity: the child's exit code is never trusted alone. Output is
|
||
* fed through BunTestOutputClassifier, and strictTestExitCode requires bun's
|
||
* terminal summary to report EXACTLY the planned file count — a shard that
|
||
* exits 0 without the summary (mid-suite process.exit truncation), with
|
||
* `(fail)` result lines, or having run fewer files than planned is a FAILURE.
|
||
* This is enforced for injected fake commands too (unlike the paid runner),
|
||
* so tests can pin the summary-missing => failure backstop; fake passing
|
||
* commands must print a synthetic `Ran N tests across M files. [Xms]` line.
|
||
*
|
||
* Per-shard temp isolation: each spawned child gets its own throwaway TMPDIR
|
||
* (TEMP/TMP on Windows) so shards can't trip over each other's temp files.
|
||
* Deliberately NOT GSTACK_HOME: injecting one shared scratch home for a whole
|
||
* invocation made 6,900 tests share a MUTABLE state dir — config tests wrote
|
||
* keys into it and relink/update-check tests then read them (measured: 12
|
||
* cross-contamination failures on the first full run). Tests that need
|
||
* GSTACK_HOME isolation mkdtemp their own per test — the repo convention —
|
||
* and the hermetic-env machinery covers E2E children.
|
||
*/
|
||
export async function runFreeShard(
|
||
files: string[],
|
||
shardNumber: number,
|
||
totalShards: number,
|
||
options: RunFreeShardOptions = {},
|
||
): Promise<FreeShardOutcome> {
|
||
const log = options.log ?? ((line: string) => console.log(line));
|
||
const label = `[test:free] shard ${shardNumber}/${totalShards}`;
|
||
|
||
// Empty shard = fast no-op SUCCESS. Indices are stable for the CI matrix,
|
||
// so an unoccupied index must not fail or shift work to a different runner.
|
||
if (files.length === 0) {
|
||
const outcome: FreeShardOutcome = {
|
||
shard: shardNumber, files: [], status: 'passed', exitCode: 0, elapsedMs: 0, groupPid: null, failingFiles: [], unattributedFailures: 0,
|
||
summary: { testsRan: 0, filesRan: 0, sawTerminalSummary: false },
|
||
};
|
||
log(shardEpilogue(outcome, totalShards));
|
||
return outcome;
|
||
}
|
||
|
||
const rootDir = options.rootDir ?? ROOT;
|
||
const wallTimeoutMs = options.wallTimeoutMs ?? DEFAULT_WALL_TIMEOUT_MS;
|
||
log(`${label} (${files.length} files${options.parallel ? ', bun --parallel' : ''})`);
|
||
|
||
// Full-stream capture: EVERY child byte lands here, whatever the console
|
||
// shows. Printed once at start so a wedged or noisy run is inspectable
|
||
// without a re-run.
|
||
const logPath = options.logFilePath ?? nextDefaultLogPath();
|
||
const logStream = fs.createWriteStream(logPath);
|
||
let logWriteFailed = false;
|
||
logStream.on('error', (err) => {
|
||
if (logWriteFailed) return;
|
||
logWriteFailed = true;
|
||
console.error(`${label} could not write the full log at ${logPath}: ${err.message}`);
|
||
});
|
||
log(`[test:free] full log: ${logPath}`);
|
||
|
||
const { command, args } = options.commandFor
|
||
? options.commandFor(files)
|
||
: { command: process.execPath, args: buildShardArgs(files, { parallel: options.parallel, rootDir }) };
|
||
|
||
const env = { ...(options.env ?? process.env) };
|
||
const stateDir = fs.mkdtempSync(path.join(os.tmpdir(), 'gstack-free-shard-'));
|
||
const childTmp = path.join(stateDir, 'tmp');
|
||
fs.mkdirSync(childTmp);
|
||
env.TMPDIR = childTmp;
|
||
env.TEMP = childTmp;
|
||
env.TMP = childTmp;
|
||
// CLI renders otherwise attach to the repo's shared .gstack/browse.json,
|
||
// even with distinct Chromium profiles. Concurrent shards and surviving
|
||
// daemons from prior runs can then replace or remove each other's state.
|
||
// Override inherited state too; the shard owns this directory's cleanup.
|
||
env.BROWSE_STATE_FILE = path.join(stateDir, '.gstack', 'browse.json');
|
||
// Per-shard Chromium profile (same isolation idea as TMPDIR): nine test
|
||
// files launch in-process persistent contexts or daemons that default to
|
||
// the SHARED ~/.gstack/chromium-profile, and two concurrent shards on one
|
||
// profile dir kill each other's browser — observed live on CI once
|
||
// duration packing recomposed shards (handoff's launchPersistentContext
|
||
// died "Target page, context or browser has been closed" while a sibling
|
||
// shard's daemon logged "Chromium process crashed"). Hash sharding had
|
||
// masked the collision by chance placement. Within a shard, files run
|
||
// serially, so sharing the per-shard profile is safe; config tests that
|
||
// assert resolution order save/restore this env around their assertions.
|
||
env.CHROMIUM_PROFILE = path.join(stateDir, 'chromium-profile');
|
||
|
||
const startedAt = Date.now();
|
||
const child = spawn(command, args, {
|
||
cwd: rootDir,
|
||
env,
|
||
stdio: ['ignore', 'pipe', 'pipe'],
|
||
detached: process.platform !== 'win32',
|
||
windowsHide: true,
|
||
});
|
||
const groupPid = child.pid ?? null;
|
||
// Group-kill on parent SIGINT/SIGTERM too, not just on timeout.
|
||
const forwarding = installChildSignalForwarding({
|
||
kill: (signal?: NodeJS.Signals | number) => {
|
||
killProcessGroup(child, (signal as NodeJS.Signals) ?? 'SIGTERM');
|
||
return true;
|
||
},
|
||
});
|
||
|
||
const classifier = new BunTestOutputClassifier();
|
||
|
||
// Console policy: quiet => nothing; verbose => the raw firehose; default =>
|
||
// only always-visible lines (fail results, crash markers, error/panic
|
||
// markers, the terminal summary), selected by the reporter. The reporter
|
||
// consumes the stream in EVERY mode so the epilogue can attribute failures.
|
||
const emitToConsole = (text: string, origin: StreamOrigin): void => {
|
||
if (options.quiet) return;
|
||
if (options.consoleWrite) {
|
||
options.consoleWrite(text);
|
||
return;
|
||
}
|
||
(origin === 'stdout' ? process.stdout : process.stderr).write(text);
|
||
};
|
||
const reporter = new FreeRunReporter(files, options.verbose ? undefined : emitToConsole);
|
||
|
||
const captureFailures = new Map<StreamOrigin, Error>();
|
||
const consumeStream = (stream: typeof child.stdout, origin: StreamOrigin): Promise<void> =>
|
||
new Promise((resolve) => {
|
||
if (!stream) {
|
||
captureFailures.set(origin, new Error('configured pipe is missing'));
|
||
resolve();
|
||
return;
|
||
}
|
||
let ended = stream.readableEnded;
|
||
const incomplete = (error?: Error | null): void => {
|
||
// A delayed error replaces the initial destroyed-stream diagnostic
|
||
// with its original cause. Failures are data, never early rejections
|
||
// while the caller is still waiting for the child's real exit.
|
||
if (error) captureFailures.set(origin, error);
|
||
else if (!captureFailures.has(origin)) captureFailures.set(origin, new Error('stream closed before end'));
|
||
resolve();
|
||
};
|
||
// Even an already-destroyed pipe can emit error on the next tick.
|
||
stream.on('error', incomplete);
|
||
stream.once('end', () => { ended = true; resolve(); });
|
||
stream.once('close', () => {
|
||
if (!ended) incomplete(stream.errored);
|
||
else resolve();
|
||
});
|
||
stream.on('data', (chunk: Buffer | string) => {
|
||
classifier.write(chunk, origin); // strict verdict ALWAYS sees the full stream
|
||
if (!logWriteFailed) logStream.write(chunk);
|
||
reporter.write(chunk, origin);
|
||
if (options.verbose) emitToConsole(typeof chunk === 'string' ? chunk : chunk.toString('utf8'), origin);
|
||
});
|
||
if (ended) resolve();
|
||
else if (stream.destroyed) incomplete(stream.errored);
|
||
});
|
||
|
||
let timedOut = false;
|
||
const killTimer = setTimeout(() => {
|
||
timedOut = true;
|
||
killProcessGroup(child, 'SIGKILL');
|
||
}, wallTimeoutMs);
|
||
|
||
let exitCode: number | null = null;
|
||
try {
|
||
const streams = [consumeStream(child.stdout, 'stdout'), consumeStream(child.stderr, 'stderr')];
|
||
exitCode = await new Promise<number | null>((resolve, reject) => {
|
||
child.once('error', reject);
|
||
child.once('close', (code) => resolve(code));
|
||
});
|
||
await Promise.all(streams);
|
||
} finally {
|
||
clearTimeout(killTimer);
|
||
forwarding.dispose();
|
||
// Reap survivors of this shard even on the clean path.
|
||
killProcessGroup(child, 'SIGKILL');
|
||
reporter.end();
|
||
for (const [origin, error] of captureFailures) {
|
||
const diagnostic = `${label} ${origin} capture incomplete: ${error.message} `
|
||
+ `(child exit ${exitCode ?? 'signal'}). Full log: ${logPath}`;
|
||
console.error(diagnostic);
|
||
if (!logWriteFailed) logStream.write(diagnostic + '\n');
|
||
}
|
||
await new Promise<void>((resolve) => logStream.end(() => resolve()));
|
||
try {
|
||
fs.rmSync(stateDir, { recursive: true, force: true });
|
||
} catch {
|
||
// Best-effort cleanup of a throwaway temp dir — a locked file on
|
||
// Windows must not turn a real verdict into an exception.
|
||
}
|
||
}
|
||
|
||
const summary = classifier.end();
|
||
const status: FreeShardStatus = timedOut
|
||
? 'timed-out'
|
||
: captureFailures.size === 0 && strictTestExitCode(exitCode ?? 1, summary, files.length) === 0 ? 'passed' : 'failed';
|
||
|
||
if (status === 'timed-out') {
|
||
console.error(
|
||
`${label} exceeded the ${Math.round(wallTimeoutMs / 1000)}s wall-clock deadline — `
|
||
+ 'killed the process group. Reporting as TIMED-OUT (distinct from failed).',
|
||
);
|
||
} else if (status === 'failed' && captureFailures.size === 0 && (exitCode ?? 1) === 0) {
|
||
const reason = summary.failedTests > 0 || summary.unhandledBetweenTests > 0
|
||
? `printed ${summary.failedTests} failing result(s) and ${summary.unhandledBetweenTests} unhandled error(s) between tests`
|
||
: summary.terminalFileCounts.length === 0
|
||
? "never printed bun's terminal summary — the run was truncated (a process.exit fired mid-suite)"
|
||
: `bun's summary reported ${summary.terminalFileCounts.join(', ')} file(s), expected ${files.length}`;
|
||
console.error(`${label} exited 0 but ${reason}. Treating as FAILED.`);
|
||
} else if (status === 'failed' && (exitCode ?? 1) !== 0) {
|
||
console.error(`${label} failed with exit code ${exitCode ?? 'signal'}`);
|
||
}
|
||
|
||
const report = reporter.report();
|
||
const failingFiles = status === 'passed' ? [] : [...new Set([
|
||
...report.failures.map((f) => f.file).filter((f): f is string => !!f),
|
||
...report.crashedFiles,
|
||
])];
|
||
const unattributedFailures = status === 'passed' ? 0
|
||
: report.failures.filter((f) => !f.file).length
|
||
+ report.unhandledErrors.length
|
||
+ captureFailures.size
|
||
+ (report.sawTerminalSummary ? 0 : 1);
|
||
const outcome: FreeShardOutcome = {
|
||
shard: shardNumber, files, status, exitCode, elapsedMs: Date.now() - startedAt, groupPid, failingFiles, unattributedFailures,
|
||
summary: { testsRan: report.testsRan, filesRan: report.filesRan, sawTerminalSummary: report.sawTerminalSummary },
|
||
};
|
||
log(shardEpilogue(outcome, totalShards));
|
||
for (const line of buildRunEpilogue(status, report, outcome.elapsedMs, logPath)) log(line);
|
||
return outcome;
|
||
}
|
||
|
||
let logPathSequence = 0;
|
||
|
||
/** Timestamped per-run log file under os.tmpdir(); pid+sequence defeat same-ms collisions. */
|
||
function nextDefaultLogPath(): string {
|
||
const stamp = new Date().toISOString().replace(/[:.]/g, '-');
|
||
logPathSequence += 1;
|
||
return path.join(os.tmpdir(), `gstack-free-test-${stamp}-${process.pid}-${logPathSequence}.log`);
|
||
}
|
||
|
||
function exitCodeFor(status: FreeShardStatus): number {
|
||
if (status === 'passed') return 0;
|
||
return status === 'timed-out' ? 124 : 1;
|
||
}
|
||
|
||
/**
|
||
* `--record-durations`: time every file in its own child (exact per-file wall,
|
||
* immune to bun's stream buffering) and write the committed seed atomically.
|
||
* Uses the same isolated, strictly classified children as a normal run.
|
||
* Run against an immutable checkout; never time while editing its inputs.
|
||
*/
|
||
async function recordFreeTestDurations(files: string[], jobs: number): Promise<number> {
|
||
const durations: Record<string, number> = {};
|
||
const failed: string[] = [];
|
||
const expectedCount = files.length;
|
||
let cursor = 0;
|
||
console.log(`[test:free] recording per-file durations: ${files.length} files across ${jobs} workers`);
|
||
const worker = async (): Promise<void> => {
|
||
for (;;) {
|
||
if (isTerminationRequested()) return;
|
||
const index = cursor;
|
||
cursor += 1;
|
||
if (index >= files.length) return;
|
||
const file = files[index];
|
||
const outcome = await runFreeShard([file], index + 1, files.length, {
|
||
wallTimeoutMs: wallTimeoutForShard(1),
|
||
quiet: true,
|
||
});
|
||
durations[normalizeRelativePath(file)] = outcome.elapsedMs;
|
||
if (outcome.status !== 'passed') failed.push(file);
|
||
}
|
||
};
|
||
// As in full-suite mode, finish readers before any checkout-mutating tests.
|
||
const mutators = files.filter(file => file in TREE_MUTATING);
|
||
files = files.filter(file => !(file in TREE_MUTATING));
|
||
await Promise.all(Array.from({ length: Math.max(1, jobs) }, () => worker()));
|
||
files = mutators;
|
||
cursor = 0;
|
||
await worker();
|
||
if (Object.keys(durations).length !== expectedCount) {
|
||
throw new Error('Duration recording was interrupted; the seed was not replaced.');
|
||
}
|
||
|
||
const target = process.env.GSTACK_FREE_TEST_DURATIONS ?? path.join(ROOT, FREE_TEST_DURATIONS_FILE);
|
||
const payload = {
|
||
version: 1,
|
||
recordedAt: new Date().toISOString(),
|
||
durations: Object.fromEntries(Object.entries(durations).sort(([a], [b]) => (a < b ? -1 : 1))),
|
||
};
|
||
// Atomic temp+rename (capture-context-budget's pattern): a killed recorder
|
||
// must never leave a truncated seed for loadFreeTestDurations to warn on.
|
||
const tmp = `${target}.tmp-${process.pid}`;
|
||
fs.writeFileSync(tmp, `${JSON.stringify(payload, null, 2)}\n`);
|
||
fs.renameSync(tmp, target);
|
||
console.log(`[test:free] wrote ${Object.keys(durations).length} durations to ${path.relative(ROOT, target)}`);
|
||
if (failed.length > 0) {
|
||
// Failures still recorded (a red file's duration is still a real cost),
|
||
// but surfaced loudly — recording from a broken tree deserves a look.
|
||
console.error(`[test:free] WARNING: ${failed.length} file(s) failed while recording:`);
|
||
for (const f of failed) console.error(` ✗ ${f}`);
|
||
return 1;
|
||
}
|
||
return 0;
|
||
}
|
||
|
||
async function retryFailedFreeFiles(
|
||
outcomes: FreeShardOutcome[], totalShards: number,
|
||
options: Pick<CliOptions, 'wallTimeoutExplicit' | 'wallTimeoutMs' | 'verbose'>,
|
||
): Promise<{ exitCode: number; retry: FreeShardOutcome | null }> {
|
||
let worst = Math.max(...outcomes.map(outcome => exitCodeFor(outcome.status)));
|
||
let retry: FreeShardOutcome | null = null;
|
||
const shardTimeout = (count: number) => options.wallTimeoutExplicit
|
||
? options.wallTimeoutMs : wallTimeoutForShard(count, options.wallTimeoutMs);
|
||
// Opt-in flaky retry (GSTACK_FREE_RETRY_FLAKY=1): when every failure is an
|
||
// attributed test failure (no timeouts, no unattributed carnage), re-run
|
||
// just the failing files ONCE in a fresh serial shard. A clean retry
|
||
// downgrades the run to a loud flaky-pass; a repeat failure stays a
|
||
// failure. Default OFF: dev boxes should see flakes, not absorb them.
|
||
// Exists for syscall-supervised sandboxes (see fullSuiteJobs) where a run
|
||
// lands 0-1 spurious browser-timing failures under an otherwise-green
|
||
// suite. Capped so a genuinely broken tree never masquerades as flaky.
|
||
const RETRY_CAP = 5;
|
||
if (
|
||
worst !== 0
|
||
&& process.env.GSTACK_FREE_RETRY_FLAKY === '1'
|
||
&& !isTerminationRequested()
|
||
&& outcomes.every((o) => o.status !== 'timed-out')
|
||
) {
|
||
const flakyFiles = [...new Set(outcomes.flatMap((o) => o.failingFiles))]
|
||
.filter((f): f is string => typeof f === 'string' && f.length > 0);
|
||
// "Fully attributed" is per-failure, not per-shard: a shard with one
|
||
// attributed failure PLUS a headerless failure / unhandled error /
|
||
// truncated run must veto the retry — re-running only failingFiles would
|
||
// mask the unattributable evidence as a FLAKY-PASS.
|
||
const allAttributed = outcomes.every((o) => hasScopedFailureAttribution(o) && (o.status === 'passed'
|
||
|| (o.failingFiles.length > 0 && o.unattributedFailures === 0)));
|
||
if (allAttributed && flakyFiles.length > 0 && flakyFiles.length <= RETRY_CAP) {
|
||
console.log(`[test:free] flaky-retry: re-running ${flakyFiles.length} failing file(s) once, serially: ${flakyFiles.join(', ')}`);
|
||
const retryOutcome = await runFreeShard(flakyFiles, totalShards + 1, totalShards + 1, {
|
||
wallTimeoutMs: shardTimeout(flakyFiles.length),
|
||
verbose: options.verbose,
|
||
});
|
||
retry = retryOutcome;
|
||
if (retryOutcome.status === 'passed') {
|
||
console.log(`[test:free] FLAKY-PASS — ${flakyFiles.length} file(s) failed once and passed on serial retry: ${flakyFiles.join(', ')}`);
|
||
console.log('[test:free] treat repeat offenders as real flakes worth fixing, not noise.');
|
||
// Durable record (WS1): console lines vanish with the scrollback; the
|
||
// ledger makes repeat offenders rankable across runs (eval:flake-rank).
|
||
const ts = new Date().toISOString();
|
||
// Two separate calls: `rev-parse --abbrev-ref HEAD HEAD` abbreviates
|
||
// BOTH revs, printing the branch twice — git_sha recorded the branch
|
||
// name (codex adversarial finding).
|
||
const ledgerBranch = (spawnSync('git', ['rev-parse', '--abbrev-ref', 'HEAD'], { cwd: ROOT, encoding: 'utf8', timeout: 5000 }).stdout ?? '').trim();
|
||
const ledgerSha = (spawnSync('git', ['rev-parse', 'HEAD'], { cwd: ROOT, encoding: 'utf8', timeout: 5000 }).stdout ?? '').trim();
|
||
appendFlakeLedger(
|
||
flakyFiles.map((file) => ({
|
||
ts,
|
||
runner: 'free' as const,
|
||
kind: 'flaky-pass' as const,
|
||
file,
|
||
shard: outcomes.find((o) => o.failingFiles.includes(file))?.shard,
|
||
...(ledgerBranch ? { branch: ledgerBranch } : {}),
|
||
...(ledgerSha ? { git_sha: ledgerSha.slice(0, 12) } : {}),
|
||
})),
|
||
flakeLedgerPath(),
|
||
);
|
||
worst = 0;
|
||
} else {
|
||
console.error('[test:free] flaky-retry FAILED — the failures reproduce serially; not flaky.');
|
||
}
|
||
} else {
|
||
console.log(`[test:free] flaky-retry skipped: ${allAttributed ? `${flakyFiles.length} failing file(s) exceeds cap ${RETRY_CAP}` : 'failures not fully attributed'}.`);
|
||
}
|
||
}
|
||
return { exitCode: worst, retry };
|
||
}
|
||
|
||
async function main(): Promise<number> {
|
||
const options = parseCliOptions(process.argv.slice(2));
|
||
const allFiles = collectFreeTestFiles();
|
||
if (allFiles.length === 0) {
|
||
throw new Error('No free test files were discovered.');
|
||
}
|
||
|
||
if (options.ciPlan || options.ciRun || options.ciVerify) {
|
||
const git = spawnSync('git', ['rev-parse', 'HEAD'], { cwd: ROOT, encoding: 'utf8', timeout: 5_000 });
|
||
if (git.status !== 0 || !git.stdout.trim()) throw new Error('Cannot bind CI plan to the checkout revision');
|
||
const revision = git.stdout.trim();
|
||
const writeJson = (file: string, value: unknown) => {
|
||
fs.mkdirSync(path.dirname(path.resolve(file)), { recursive: true });
|
||
const temporary = `${file}.tmp-${process.pid}`;
|
||
fs.writeFileSync(temporary, JSON.stringify(value, null, 2) + '\n');
|
||
fs.renameSync(temporary, file);
|
||
};
|
||
if (options.ciPlan) {
|
||
const plan = createFreeCiPlan(allFiles, options.shardCount, loadFreeTestDurations() ?? {}, revision);
|
||
validateFreeCiPlan(plan, allFiles, revision);
|
||
writeJson(options.ciPlan, plan);
|
||
console.log(JSON.stringify({ shard: plan.shards.map(shard => shard.shard) }));
|
||
return 0;
|
||
}
|
||
const plan = JSON.parse(fs.readFileSync((options.ciRun ?? options.ciVerify)!, 'utf8')) as FreeCiPlan;
|
||
validateFreeCiPlan(plan, allFiles, revision);
|
||
if (options.ciVerify) {
|
||
const results = fs.readdirSync(options.results!).filter(file => file.endsWith('.json'))
|
||
.map(file => JSON.parse(fs.readFileSync(path.join(options.results!, file), 'utf8')) as FreeCiResult);
|
||
verifyFreeCiResults(plan, results);
|
||
console.log(`[test:free] CI PASS: ${allFiles.length} files across ${results.length} isolated shards; slowest ${Math.round(Math.max(...results.map(result => result.outcome.elapsedMs + (result.retry?.elapsedMs ?? 0))) / 1000)}s including retries`);
|
||
return 0;
|
||
}
|
||
const shard = plan.shards[options.shardIndex! - 1];
|
||
if (!shard || shard.shard !== options.shardIndex) throw new Error('CI shard index is outside the plan');
|
||
const outcome = await runFreeShard(shard.files, shard.shard, plan.shards.length, {
|
||
wallTimeoutMs: options.wallTimeoutExplicit ? options.wallTimeoutMs
|
||
: wallTimeoutForPackedShard(shard.predictedMs, options.wallTimeoutMs, shard.files.length),
|
||
verbose: options.verbose,
|
||
});
|
||
const retried = await retryFailedFreeFiles([outcome], plan.shards.length, options);
|
||
writeJson(options.result!, { planId: plan.id, revision, outcome, retry: retried.retry } satisfies FreeCiResult);
|
||
return retried.exitCode;
|
||
}
|
||
|
||
let files = allFiles;
|
||
if (options.quick) {
|
||
const missing = QUICK_CORE.filter(file => !allFiles.includes(file));
|
||
if (missing.length) throw new Error(`Quick core files missing: ${missing.join(', ')}`);
|
||
const durations = loadFreeTestDurations() ?? {};
|
||
files = selectQuickFreeFiles(allFiles, durations);
|
||
const unknown = allFiles.filter(file => durations[file] === undefined && !QUICK_CORE.includes(file)).length;
|
||
console.log(`[test:free] QUICK SUBSET: ${files.length}/${allFiles.length} files; ${unknown} unclassified and ${allFiles.length - files.length - unknown} slow files excluded. Full CI remains required; this is not release acceptance.`);
|
||
}
|
||
let curationReport: CurationResult | null = null;
|
||
if (options.windowsOnly) {
|
||
curationReport = curateWindowsSafe(allFiles);
|
||
files = curationReport.safe;
|
||
console.log(`[test:free] curated ${files.length} Windows-safe tests (${curationReport.excluded.length} excluded)`);
|
||
if (options.listOnly && curationReport.excluded.length > 0) {
|
||
console.log('\nExcluded (POSIX-fragile):');
|
||
for (const { file, reason } of curationReport.excluded) {
|
||
console.log(` - ${file} [${reason}]`);
|
||
}
|
||
}
|
||
}
|
||
|
||
if (options.listOnly) {
|
||
console.log(`\nDiscovered ${files.length} test files.`);
|
||
for (const file of files) console.log(` ${file}`);
|
||
return 0;
|
||
}
|
||
|
||
if (options.recordDurations) {
|
||
return recordFreeTestDurations(files, fullSuiteJobs());
|
||
}
|
||
|
||
if (options.dryRun) {
|
||
const shards = assignFilesToShards(files, options.shardCount);
|
||
const occupied = shards.filter((s) => s.length > 0).length;
|
||
console.log(
|
||
`\nWould run ${files.length} files across ${shards.length} shards (${occupied} occupied). `
|
||
+ 'Without --shard, the full suite runs as N concurrent shard processes '
|
||
+ '(plus a serial tree-mutating shard) instead.',
|
||
);
|
||
for (const line of formatShardSummary(shards)) console.log(line);
|
||
return 0;
|
||
}
|
||
|
||
if (options.shardIndex !== null) {
|
||
// Bounds-check against the REQUESTED shard count, not post-assignment
|
||
// occupancy — indices must be stable for a CI matrix, and an empty shard
|
||
// is a valid fast no-op.
|
||
if (!Number.isInteger(options.shardIndex) || options.shardIndex < 1 || options.shardIndex > options.shardCount) {
|
||
throw new Error(`--shard must be between 1 and ${options.shardCount}. Received: ${options.shardIndex}`);
|
||
}
|
||
const shards = assignFilesToShards(files, options.shardCount);
|
||
const outcome = await runFreeShard(shards[options.shardIndex - 1], options.shardIndex, options.shardCount, {
|
||
wallTimeoutMs: options.wallTimeoutMs,
|
||
verbose: options.verbose,
|
||
});
|
||
return exitCodeFor(outcome.status);
|
||
}
|
||
|
||
// Full-suite mode: N concurrent shard PROCESSES, serial within each — the
|
||
// paid runner's proven model. One `bun test --parallel` invocation was
|
||
// tried first (decision V3) and abandoned after three distinct
|
||
// worker-runtime pathologies in a single day on Bun 1.3.13: a segfault
|
||
// whose crashed-worker retry wedged the run (security-live-playwright), a
|
||
// gated file's still-running file-level hooks stalling a worker
|
||
// (compare-board), and spawn-heavy files hanging workers under load
|
||
// (session-runner-timeout). Plain child processes have none of these:
|
||
// proven spawn semantics, per-shard group-kill, per-shard logs, and a
|
||
// wedge only ever costs its own shard. WORKER_HOSTILE files are moot in
|
||
// process shards (no workers) and fold back into normal assignment.
|
||
const jobs = fullSuiteJobs();
|
||
// Phase split: tree-mutating tests run AFTER the parallel shards, in one
|
||
// serial shard, so no concurrent shard ever reads a half-regenerated tree.
|
||
const mutators = files.filter((f) => f in TREE_MUTATING);
|
||
const readers = files.filter((f) => !(f in TREE_MUTATING));
|
||
const durations = loadFreeTestDurations();
|
||
const packed = durations ? packShardsByDuration(readers, jobs, durations) : null;
|
||
const shards = packed ? packed.shards : assignFilesToShards(readers, jobs);
|
||
const totalShards = jobs + (mutators.length > 0 ? 1 : 0);
|
||
console.log(`[test:free] full suite: ${readers.length} files across ${jobs} shard processes`
|
||
+ (packed ? ' (duration-packed)' : '')
|
||
+ (mutators.length > 0 ? `, then ${mutators.length} tree-mutating file(s) serially` : ''));
|
||
if (packed) {
|
||
// One line per shard so a packing regression is diagnosable from any log.
|
||
packed.predictedMs.forEach((ms, i) => {
|
||
console.log(`[test:free] shard ${i + 1}: ${shards[i].length} files, predicted ~${Math.round(ms / 1000)}s`);
|
||
});
|
||
}
|
||
const shardTimeout = (fileCount: number): number =>
|
||
options.wallTimeoutExplicit ? options.wallTimeoutMs : wallTimeoutForShard(fileCount, options.wallTimeoutMs);
|
||
const outcomes = await Promise.all(
|
||
shards.map((shardFiles, index) => runFreeShard(shardFiles, index + 1, totalShards, {
|
||
// Packed shards get duration-aware walls: LPT decouples file count from
|
||
// cost BY DESIGN, so the 5s/file heuristic would undersize a shard
|
||
// holding few expensive files.
|
||
wallTimeoutMs: packed && !options.wallTimeoutExplicit
|
||
? wallTimeoutForPackedShard(packed.predictedMs[index], options.wallTimeoutMs, shardFiles.length)
|
||
: shardTimeout(shardFiles.length),
|
||
verbose: options.verbose,
|
||
})),
|
||
);
|
||
let worst = Math.max(...outcomes.map((o) => exitCodeFor(o.status)));
|
||
// Cancellation stops the run: don't launch the serial tree-mutating shard
|
||
// after a SIGINT/SIGTERM already killed the parallel phase.
|
||
if (mutators.length > 0 && !isTerminationRequested()) {
|
||
const mutatorOutcome = await runFreeShard(mutators, totalShards, totalShards, {
|
||
wallTimeoutMs: shardTimeout(mutators.length),
|
||
verbose: options.verbose,
|
||
});
|
||
worst = Math.max(worst, exitCodeFor(mutatorOutcome.status));
|
||
if (mutatorOutcome.status !== 'passed') {
|
||
// Mutator safety rests on each test restoring default state itself; a
|
||
// SIGKILL at the wall deadline (or a mid-regeneration crash) defeats
|
||
// that by construction. Say so, loudly, before someone commits
|
||
// regenerated SKILL.md / .agents artifacts by accident.
|
||
const dirty = spawnSyncGitStatusGenerated();
|
||
if (dirty.length > 0) {
|
||
console.error('[test:free] ⚠ tree-mutating shard did not finish cleanly — generated artifacts may be mid-regeneration:');
|
||
for (const line of dirty.slice(0, 20)) console.error(`[test:free] ${line}`);
|
||
console.error('[test:free] restore with: bun run gen:skill-docs (or git checkout -- <paths>)');
|
||
}
|
||
}
|
||
outcomes.push(mutatorOutcome);
|
||
}
|
||
|
||
return (await retryFailedFreeFiles(outcomes, totalShards, options)).exitCode;
|
||
}
|
||
|
||
/** Dirty generated artifacts (SKILL.md / host outputs) after a failed mutator shard. */
|
||
function spawnSyncGitStatusGenerated(): string[] {
|
||
const result = spawnSync('git', ['status', '--porcelain'], { cwd: ROOT, encoding: 'utf8' });
|
||
if (result.status !== 0 || !result.stdout) return [];
|
||
return result.stdout.split('\n').filter((line) =>
|
||
/SKILL\.md$/.test(line) || line.includes('.agents/') || line.includes('.factory/'));
|
||
}
|
||
|
||
if (import.meta.main) {
|
||
try {
|
||
process.exitCode = await main();
|
||
} catch (error) {
|
||
console.error(`[test:free] ${error instanceof Error ? error.message : String(error)}`);
|
||
process.exitCode = 1;
|
||
}
|
||
}
|