mirror of
https://github.com/garrytan/gstack.git
synced 2026-09-28 07:32:14 +02:00
* fix(memory-ingest): --scan-secrets scans the rendered page and fails closed --scan-secrets ran gitleaks on the raw transcript .jsonl, then imported a page rendered from it. gitleaks' assignment rules don't match across a JSON-escaped quote (KEY=\"v\" on disk), so a secret the rendered page shows as KEY="v" was imported unflagged. And the gate skipped a file only on scanner "gitleaks" with findings, so a scan that errored (non-zero exit, 16MB maxBuffer overflow on a file with many findings, unparseable report) or could not run (gitleaks missing, slow-probe cooldown) imported the file unscanned. Scan the rendered page body, the exact bytes writeStaged() writes, via a new secretScanText() helper, and skip the file whenever the scan did not complete. Skipped files stay out of the state file, so the next run retries them. Reword the helper warnings and setup-gbrain/memory.md, which described the fail-open as intended. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(test): reconcile Bun failure markers and footer counts * fix(sync-gbrain): verify source-scoped reads without mutation * fix(test): recognize grounded TTHW target choices structurally * fix(aside): make the readiness probe work under zsh and report why it failed The probe built its deadline into `_T` and expanded it unquoted, so `$_T aside repl …` only worked in a shell that word-splits. zsh does not: it looked for a command literally named "gtimeout 30", the probe answered ASIDE_NOT_RUNNING with Aside installed and ready, and every browsing skill fell back to the bundled Chromium in silence. zsh is the macOS default and Aside is macOS-only, so on a stock Mac the probe could never report READY. The deadline becomes a function, `_gs_d`. It receives the command as "$@", already split, so sh, bash and zsh all behave the same, and the gtimeout → timeout → perl alarm chain is unchanged. A 4th arm runs the call unbounded when none of the three is present, which is what the empty `_T` did before. Not `eval`: it re-parses the string, so the parens and `;` of the perl arm become syntax and that arm dies in bash *and* zsh — on a stock Mac, the arm that actually runs. On failure the probe now prints the CLI's reason after ASIDE_NOT_RUNNING:, the shape gstack-render already uses: the first line that starts with a capital letter, i.e. the CLI's own sentence or Node's `Error:` line below its loader frame. "Not running" covers states with different fixes — no window open for the profile, a NODE_OPTIONS preload that kills the CLI — and a bare verdict sent all of them to "open the Aside app". The BROWSER SETUP prose quotes that reason before asking the user to open the app. The text pin asserted the broken invocation verbatim, so it now pins the function and asserts neither `$_T aside repl` nor an eval form comes back. A second test executes the rendered probe in sh, bash and zsh on each of the four deadline arms with stubbed binaries on a narrowed PATH, plus two failing CLIs: one that prints its own sentence, one that crashes like Node with the useful line below the frame. The deadline function costs zero bytes against the lines it replaces; the reason costs 53 per copy of the probe (44 where the reworded BROWSER SETUP line gives 9 back). That moves four guards by the measured amount: plan-devex-review's skeleton cap to 68,550 (measured 68,544), plan-ceo-review's skeleton cap to 80,150 (measured 80,111) and union ratio to 1.081 (measured 1.0803), and plan-eng-review's union ratio to 1.151 (measured 1.1504). Fixes #2842, #2941. * Clarify engineering review startup and decision flow * Fix Windows readiness fixture PATH and command shim * fix(test): recognize grounded TTHW target choices structurally * Clarify engineering review startup and decision flow * fix(test): restrict QA-only fixture tools to its no-Edit contract * v1.90.0.0 fix(sync-gbrain): guard readiness verdicts and refresh metadata * fix(browse): validate canonical upload targets * fix(gbrain): classify structured PGLite busy response * fix(browse): preserve native extension runtime APIs * Fix displayless browser handoff ownership * Accept unique installed autoplan methodology aliases * fix(skills): preserve positional literals during installation * fix(browse): checksum installer contents through stdin * fix(test): normalize Windows checksum fixture paths * test: emulate unavailable shasum in Windows checksum fixture * fix(investigate): preserve owned freeze lifecycle * fix(review): preserve N+1 retry and Red Team completion * fix: bound Aside readiness and preserve safe fallback * test: exercise setup and Chromium on native ARM * fix: preserve install ownership and ARM browser selection * Fix gbrain ingest scan boundaries and seed observation * Refresh managed ship hooks and supervise expanded paid census * Reject resumed gbrain pages excluded by current policy * Recover zombie agent locks safely and enable CI Python venv * Repair paid actor declarations and Aside pitch assertions * Bump consolidated wave to next free minor release * Clarify CEO review admin choices and option tradeoffs * Preserve CEO mode handoff anchors in clarified workflow * Make Windows portability fixtures use shell-native paths * Restore ARM Bun alias and clarify ship review gates * Refresh ship workflow golden snapshots * Fix Windows DX documentation controls without piped stdin * Decode Codex child pipes without Bun's encoded-stream stall * Bound DX pre-review audit before product questions * Clarify trusted review-start read in paid revalidation * Bump consolidated wave to next free minor release * Clarify CEO review admin choices and option tradeoffs * Preserve CEO mode handoff anchors in clarified workflow * Make Windows portability fixtures use shell-native paths * Restore ARM Bun alias and clarify ship review gates * Refresh ship workflow golden snapshots * Fix Windows DX documentation controls without piped stdin * Decode Codex child pipes without Bun's encoded-stream stall * Bound DX pre-review audit before product questions * Clarify trusted review-start read in paid revalidation * Reconcile new main planning flow and paid judge census * fix: reconcile rebased planning and source-bound validation * test: pin cookie workflow judge to scored Sonnet model * fix: keep terminal agent boot out of module imports * fix: preserve pending-question uncertainty in engineering review * fix: stabilize Windows reliability-wave fixtures * fix: clarify design consultation research workflow * fix: preserve independent design consultation inputs * fix: resolve design taste scope and browser research guidance * fix: make consultation opt-in preflight unambiguous * test: await native Edge owner readiness or terminal result --------- Co-authored-by: Bruce Krysiak <brucek@alum.mit.edu> Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> Co-authored-by: Antonio Vitalic <antoninte99@gmail.com>
474 lines
18 KiB
TypeScript
474 lines
18 KiB
TypeScript
/**
|
|
* Strict Bun-test output classification + child lifecycle helpers.
|
|
*
|
|
* Works around a Bun test runner bug where failures can be printed even though
|
|
* the child exits successfully: output is forwarded byte-for-byte as it
|
|
* arrives, and only complete Bun result lines and terminal summaries are
|
|
* classified. `strictTestExitCode` then refuses to trust a zero exit when the
|
|
* output shows failures (or when fewer files ran than expected).
|
|
*
|
|
* Shared by the sharded paid-tier runner (scripts/test-paid-shards.ts) and any
|
|
* future strict wrapper around `bun test`.
|
|
*/
|
|
|
|
import { spawn, type ChildProcess } from 'node:child_process';
|
|
import { StringDecoder } from 'node:string_decoder';
|
|
import * as path from 'node:path';
|
|
|
|
const ROOT = path.resolve(import.meta.dir, '..');
|
|
const ANSI_ESCAPE = /\u001B\[[0-?]*[ -/]*[@-~]/g;
|
|
const BUN_FAIL_RESULT = /^(?:\(fail\)|✗) (.+) \[(?:\d+(?:\.\d+)?)(?:ns|us|µs|ms|s)\]$/;
|
|
const BUN_BETWEEN_TESTS_ERROR = '# Unhandled error between tests';
|
|
const BUN_TERMINAL_SUMMARY = /^Ran (\d+) tests? across (\d+) files?\. \[(?:\d+(?:\.\d+)?)(?:ns|us|µs|ms|s)\]$/;
|
|
// The counts block bun prints just before the terminal summary (" 1 pass",
|
|
// " 2 skip", " 0 fail"). "Ran N tests" COUNTS skipped tests, so N alone
|
|
// cannot distinguish a shard that verified work from one whose every test
|
|
// self-skipped (external-service binary missing, tier mismatch) — the
|
|
// green-by-skip class. Anchored to whole-line matches; nested bun-test
|
|
// children can still contribute counts (same known limit as the terminal
|
|
// summary — see the last-summary-anchoring TODO in the audit).
|
|
const BUN_SKIP_COUNT = /^\s*(\d+) skip$/;
|
|
const BUN_PASS_COUNT = /^\s*(\d+) pass$/;
|
|
const BUN_FAIL_COUNT = /^\s*(\d+) fail$/;
|
|
|
|
export type BunTestOutputFinding = 'failed-test' | 'unhandled-between-tests';
|
|
|
|
export interface BunTestOutputSummary {
|
|
failedTests: number;
|
|
unhandledBetweenTests: number;
|
|
terminalFileCounts: number[];
|
|
/** Test counts from the same terminal lines — feeds the hollow-shard guard. */
|
|
terminalTestCounts: number[];
|
|
/** Sum of bun's " N skip" count lines. "Ran N tests" includes skips, so
|
|
* this is what separates verified work from green-by-skip. */
|
|
skippedTests: number;
|
|
/** Sum of bun's " N pass" count lines. */
|
|
passedTests: number;
|
|
}
|
|
|
|
export type ForwardedTerminationSignal = 'SIGINT' | 'SIGTERM';
|
|
|
|
export interface TerminationSignalSource {
|
|
on(event: string, listener: () => void): unknown;
|
|
off(event: string, listener: () => void): unknown;
|
|
}
|
|
|
|
export interface TerminationTimerApi {
|
|
schedule(callback: () => void, delayMs: number): unknown;
|
|
cancel(handle: unknown): void;
|
|
}
|
|
|
|
export interface ChildSignalForwarding {
|
|
readonly receivedSignal: ForwardedTerminationSignal | null;
|
|
dispose(): void;
|
|
}
|
|
|
|
const DEFAULT_TERMINATION_TIMER: TerminationTimerApi = {
|
|
schedule: (callback, delayMs) => setTimeout(callback, delayMs),
|
|
cancel: (handle) => clearTimeout(handle as ReturnType<typeof setTimeout>),
|
|
};
|
|
|
|
/**
|
|
* Per-source termination bookkeeping, shared across every forwarder bound to
|
|
* the same source. Installing ANY signal listener suppresses Node's default
|
|
* terminate-on-SIGINT/SIGTERM, so without this the parent runner survived
|
|
* cancellation: it killed the current child, then kept LAUNCHING new shards
|
|
* (observed: paid runs continuing to burn API spend after Ctrl-C). The first
|
|
* signal now also schedules the parent's own exit after the children's
|
|
* SIGKILL grace, and runners consult isTerminationRequested() before
|
|
* launching more work.
|
|
*/
|
|
interface SourceTerminationState {
|
|
requested: boolean;
|
|
exitScheduled: boolean;
|
|
}
|
|
const SOURCE_TERMINATION_STATE = new WeakMap<TerminationSignalSource, SourceTerminationState>();
|
|
function terminationStateFor(source: TerminationSignalSource): SourceTerminationState {
|
|
let state = SOURCE_TERMINATION_STATE.get(source);
|
|
if (!state) {
|
|
state = { requested: false, exitScheduled: false };
|
|
SOURCE_TERMINATION_STATE.set(source, state);
|
|
}
|
|
return state;
|
|
}
|
|
export function isTerminationRequested(source: TerminationSignalSource = process): boolean {
|
|
return SOURCE_TERMINATION_STATE.get(source)?.requested ?? false;
|
|
}
|
|
const signalExitCode = (signal: ForwardedTerminationSignal): number =>
|
|
128 + (signal === 'SIGINT' ? 2 : 15);
|
|
|
|
/**
|
|
* Bind one active child to the parent's termination lifecycle. SIGINT and
|
|
* SIGTERM get a grace period so Bun can clean up; a repeated signal, timeout,
|
|
* or synchronous parent exit uses SIGKILL so the child cannot be orphaned.
|
|
* The parent itself exits shortly after the grace window (or immediately on
|
|
* a repeated signal) — cancellation must terminate the RUN, not just the
|
|
* currently-running children.
|
|
*/
|
|
export function installChildSignalForwarding(
|
|
child: Pick<ChildProcess, 'kill'>,
|
|
source: TerminationSignalSource = process,
|
|
timer: TerminationTimerApi = DEFAULT_TERMINATION_TIMER,
|
|
graceMs = 5_000,
|
|
exitImpl: (code: number) => void = (code) => process.exit(code),
|
|
): ChildSignalForwarding {
|
|
let receivedSignal: ForwardedTerminationSignal | null = null;
|
|
let forceTimer: unknown = null;
|
|
let disposed = false;
|
|
|
|
const scheduleParentExit = (signal: ForwardedTerminationSignal, delayMs: number): void => {
|
|
const state = terminationStateFor(source);
|
|
state.requested = true;
|
|
if (state.exitScheduled) return;
|
|
state.exitScheduled = true;
|
|
// Never cancelled by dispose(): once cancellation is requested, the run
|
|
// is going down even if this particular shard finishes cleanly first.
|
|
timer.schedule(() => exitImpl(signalExitCode(signal)), delayMs);
|
|
};
|
|
|
|
const forward = (signal: ForwardedTerminationSignal): void => {
|
|
if (disposed) return;
|
|
if (receivedSignal !== null) {
|
|
child.kill('SIGKILL');
|
|
scheduleParentExit(signal, 0);
|
|
return;
|
|
}
|
|
receivedSignal = signal;
|
|
child.kill(signal);
|
|
forceTimer = timer.schedule(() => {
|
|
forceTimer = null;
|
|
child.kill('SIGKILL');
|
|
}, graceMs);
|
|
// Exit AFTER the children's SIGKILL grace so the group kills land first.
|
|
scheduleParentExit(signal, graceMs + 1_000);
|
|
};
|
|
const onSigint = () => forward('SIGINT');
|
|
const onSigterm = () => forward('SIGTERM');
|
|
const onExit = () => { child.kill('SIGKILL'); };
|
|
|
|
source.on('SIGINT', onSigint);
|
|
source.on('SIGTERM', onSigterm);
|
|
source.on('exit', onExit);
|
|
|
|
return {
|
|
get receivedSignal() {
|
|
return receivedSignal;
|
|
},
|
|
dispose() {
|
|
if (disposed) return;
|
|
disposed = true;
|
|
source.off('SIGINT', onSigint);
|
|
source.off('SIGTERM', onSigterm);
|
|
source.off('exit', onExit);
|
|
if (forceTimer !== null) timer.cancel(forceTimer);
|
|
forceTimer = null;
|
|
},
|
|
};
|
|
}
|
|
|
|
/**
|
|
* SIGKILL the shard's whole process group. Orphaned grandchildren (browsers,
|
|
* claude sessions) are how a stalled run once burned a core for 15.7 hours.
|
|
*/
|
|
export function killProcessGroup(child: ChildProcess, signal: NodeJS.Signals): void {
|
|
if (process.platform === 'win32' || typeof child.pid !== 'number') {
|
|
child.kill(signal);
|
|
return;
|
|
}
|
|
try {
|
|
process.kill(-child.pid, signal);
|
|
} catch (err) {
|
|
const code = (err as NodeJS.ErrnoException).code;
|
|
if (code === 'ESRCH') return; // group already gone
|
|
if (code !== 'EPERM') throw err;
|
|
// Observed on macOS after a SIGKILLed group is reaped: signalling the
|
|
// now-empty group id returns EPERM, not ESRCH. Throwing here loses the
|
|
// shard's real outcome (a timeout gets recorded as a failure) and, from
|
|
// the timeout timer, leaves the shard promise unsettled — a hang, which
|
|
// is the exact failure class this runner exists to kill. Fall back to the
|
|
// direct pid so a genuinely-live child is still signalled.
|
|
try {
|
|
child.kill(signal);
|
|
} catch {
|
|
// Best-effort reap: nothing actionable is left if this fails too.
|
|
}
|
|
}
|
|
}
|
|
|
|
/**
|
|
* Strip ANSI escapes and a trailing CR from one output line. Every line
|
|
* matcher (here and in the free runner's console filter / failure
|
|
* attribution) MUST match against this form — a prior grep for `(fail)`
|
|
* lines missed real failures because color codes sat inside the line.
|
|
*/
|
|
export function stripAnsiLine(rawLine: string): string {
|
|
return rawLine.replace(ANSI_ESCAPE, '').replace(/\r$/, '');
|
|
}
|
|
|
|
export function classifyBunTestOutputLine(rawLine: string): BunTestOutputFinding | null {
|
|
const line = stripAnsiLine(rawLine);
|
|
if (parseBunFailureResult(line) !== null) return 'failed-test';
|
|
if (line === BUN_BETWEEN_TESTS_ERROR) return 'unhandled-between-tests';
|
|
return null;
|
|
}
|
|
|
|
export function parseBunFailureResult(rawLine: string): string | null {
|
|
return BUN_FAIL_RESULT.exec(stripAnsiLine(rawLine))?.[1] ?? null;
|
|
}
|
|
|
|
export function parseBunTerminalSummaryLine(rawLine: string): number | null {
|
|
return parseBunTerminalSummary(rawLine)?.files ?? null;
|
|
}
|
|
|
|
export function parseBunTerminalSummary(rawLine: string): { tests: number; files: number } | null {
|
|
const line = stripAnsiLine(rawLine);
|
|
const match = BUN_TERMINAL_SUMMARY.exec(line);
|
|
return match
|
|
? { tests: Number.parseInt(match[1], 10), files: Number.parseInt(match[2], 10) }
|
|
: null;
|
|
}
|
|
|
|
/**
|
|
* Incrementally classifies output without assuming process chunks align to
|
|
* lines. Buffers are PER ORIGIN: stdout and stderr are independent pipes, so
|
|
* a chunk from one can arrive between two halves of a line from the other.
|
|
* A single shared buffer would glue those fragments into garbled lines — a
|
|
* sheared `(fail)` line goes uncounted and a sheared terminal summary reads
|
|
* as truncation. Counters are shared; only line assembly is per-stream.
|
|
*/
|
|
export type ClassifierOrigin = 'stdout' | 'stderr';
|
|
|
|
export class BunFailureSummaryParser {
|
|
private readonly pending: Partial<Record<ClassifierOrigin, { failures: number | null }>> = {};
|
|
|
|
consume(rawLine: string, origin: ClassifierOrigin): number | null {
|
|
const line = stripAnsiLine(rawLine);
|
|
if (BUN_PASS_COUNT.test(line)) {
|
|
this.pending[origin] = { failures: null };
|
|
return null;
|
|
}
|
|
const pending = this.pending[origin];
|
|
if (!pending) return null;
|
|
const fail = BUN_FAIL_COUNT.exec(line);
|
|
if (fail) {
|
|
pending.failures = Math.max(pending.failures ?? 0, Number.parseInt(fail[1], 10));
|
|
return null;
|
|
}
|
|
if (parseBunTerminalSummary(line) !== null) {
|
|
delete this.pending[origin];
|
|
return pending.failures;
|
|
}
|
|
return null;
|
|
}
|
|
}
|
|
|
|
export class BunTestOutputClassifier {
|
|
private readonly decoders: Record<ClassifierOrigin, StringDecoder> = {
|
|
stdout: new StringDecoder('utf8'),
|
|
stderr: new StringDecoder('utf8'),
|
|
};
|
|
private pending: Record<ClassifierOrigin, string> = { stdout: '', stderr: '' };
|
|
private failedTests = 0;
|
|
private reportedFailedTests = 0;
|
|
private readonly failureSummary = new BunFailureSummaryParser();
|
|
private unhandledBetweenTests = 0;
|
|
private terminalFileCounts: number[] = [];
|
|
private terminalTestCounts: number[] = [];
|
|
private skippedTests = 0;
|
|
private passedTests = 0;
|
|
|
|
write(chunk: Uint8Array | string, origin: ClassifierOrigin = 'stdout'): void {
|
|
this.pending[origin] += typeof chunk === 'string'
|
|
? chunk
|
|
: this.decoders[origin].write(Buffer.from(chunk));
|
|
this.consumeCompleteLines(origin);
|
|
}
|
|
|
|
end(): BunTestOutputSummary {
|
|
for (const origin of ['stdout', 'stderr'] as const) {
|
|
this.pending[origin] += this.decoders[origin].end();
|
|
if (this.pending[origin].length > 0) this.classify(this.pending[origin], origin);
|
|
this.pending[origin] = '';
|
|
}
|
|
return this.summary();
|
|
}
|
|
|
|
summary(): BunTestOutputSummary {
|
|
return {
|
|
failedTests: Math.max(this.failedTests, this.reportedFailedTests),
|
|
unhandledBetweenTests: this.unhandledBetweenTests,
|
|
terminalFileCounts: [...this.terminalFileCounts],
|
|
terminalTestCounts: [...this.terminalTestCounts],
|
|
skippedTests: this.skippedTests,
|
|
passedTests: this.passedTests,
|
|
};
|
|
}
|
|
|
|
private consumeCompleteLines(origin: ClassifierOrigin): void {
|
|
let newline = this.pending[origin].indexOf('\n');
|
|
while (newline !== -1) {
|
|
this.classify(this.pending[origin].slice(0, newline), origin);
|
|
this.pending[origin] = this.pending[origin].slice(newline + 1);
|
|
newline = this.pending[origin].indexOf('\n');
|
|
}
|
|
}
|
|
|
|
private classify(line: string, origin: ClassifierOrigin): void {
|
|
const finding = classifyBunTestOutputLine(line);
|
|
if (finding === 'failed-test') this.failedTests += 1;
|
|
if (finding === 'unhandled-between-tests') this.unhandledBetweenTests += 1;
|
|
const stripped = stripAnsiLine(line);
|
|
const skip = BUN_SKIP_COUNT.exec(stripped);
|
|
if (skip !== null) this.skippedTests += Number.parseInt(skip[1], 10);
|
|
const pass = BUN_PASS_COUNT.exec(stripped);
|
|
if (pass !== null) this.passedTests += Number.parseInt(pass[1], 10);
|
|
const fail = this.failureSummary.consume(stripped, origin);
|
|
if (fail !== null) this.reportedFailedTests = Math.max(this.reportedFailedTests, fail);
|
|
const terminal = parseBunTerminalSummary(line);
|
|
if (terminal !== null) {
|
|
this.terminalFileCounts.push(terminal.files);
|
|
this.terminalTestCounts.push(terminal.tests);
|
|
}
|
|
}
|
|
}
|
|
|
|
export function strictTestExitCode(
|
|
childExitCode: number,
|
|
summary: BunTestOutputSummary,
|
|
expectedFiles?: number,
|
|
): number {
|
|
if (childExitCode !== 0) return childExitCode;
|
|
if (summary.failedTests > 0 || summary.unhandledBetweenTests > 0) return 1;
|
|
if (expectedFiles !== undefined && !summary.terminalFileCounts.includes(expectedFiles)) return 1;
|
|
return 0;
|
|
}
|
|
|
|
export function normalizeRelativePath(filePath: string): string {
|
|
return filePath.replace(/\\/g, '/');
|
|
}
|
|
|
|
/**
|
|
* Bun treats positional test paths as substring filters. Resolve every
|
|
* canonical relative path before spawning so `test/foo.test.ts` cannot also
|
|
* select `browse/test/foo.test.ts`.
|
|
*/
|
|
export function exactTestFileSelectors(files: string[], rootDir = ROOT): string[] {
|
|
return files.map((file) => path.isAbsolute(file) ? path.normalize(file) : path.resolve(rootDir, file));
|
|
}
|
|
|
|
export function forwardAndClassify(
|
|
stream: NodeJS.ReadableStream,
|
|
destination: NodeJS.WriteStream,
|
|
classifier: BunTestOutputClassifier,
|
|
origin: ClassifierOrigin = 'stdout',
|
|
): Promise<void> {
|
|
return new Promise((resolve, reject) => {
|
|
let ended = false;
|
|
const incomplete = () => reject(new Error(`incomplete ${origin} capture: stream closed before end`));
|
|
stream.on('data', (chunk: Buffer | string) => {
|
|
classifier.write(chunk, origin);
|
|
destination.write(chunk);
|
|
});
|
|
stream.once('end', () => { ended = true; resolve(); });
|
|
stream.on('error', reject);
|
|
stream.once('close', () => { if (!ended) incomplete(); });
|
|
// Bun can return an already-destroyed pipe whose close event is past.
|
|
if ('destroyed' in stream && stream.destroyed && !ended) {
|
|
if ('errored' in stream && stream.errored) reject(stream.errored);
|
|
else incomplete();
|
|
}
|
|
});
|
|
}
|
|
|
|
// --- Shared shard-child lifecycle ---
|
|
|
|
export interface RunShardChildOptions {
|
|
command: string;
|
|
args: string[];
|
|
cwd: string;
|
|
env: NodeJS.ProcessEnv;
|
|
/** External wall-clock deadline; on expiry the child's process GROUP is SIGKILLed. */
|
|
timeoutMs: number;
|
|
/**
|
|
* Hook the freshly-spawned child's stdout/stderr. Stream POLICY (classifier
|
|
* tees, log spooling, console forwarding, reporters) is entirely the
|
|
* caller's. Runs synchronously right after spawn; the returned promises are
|
|
* awaited AFTER the child closes, so trailing output is fully drained
|
|
* before the caller reads its classifier/reporter state.
|
|
*/
|
|
hookStreams: (child: ChildProcess) => Array<Promise<void>>;
|
|
}
|
|
|
|
export interface ShardChildResult {
|
|
exitCode: number | null;
|
|
/** True when the wall timer fired and SIGKILLed the group. */
|
|
timedOut: boolean;
|
|
/** The child's pid — the process-GROUP id on POSIX (detached spawn). */
|
|
groupPid: number | null;
|
|
}
|
|
|
|
/**
|
|
* The child lifecycle both sharded runners need, extracted from
|
|
* scripts/test-paid-shards.ts runPaidShard (scripts/test-free-shards.ts
|
|
* runFreeShard duplicates the same ~35 lines verbatim today and is designed
|
|
* to migrate here in a later change):
|
|
*
|
|
* - spawn detached on POSIX so the child owns its process group,
|
|
* - forward parent SIGINT/SIGTERM to the whole group (not just the child),
|
|
* - arm an EXTERNAL wall-clock timer that SIGKILLs the group — a spinning
|
|
* child main thread never fires its own in-process timer,
|
|
* - in EVERY exit path: disarm the timer, detach the signal forwarder, and
|
|
* reap group survivors with SIGKILL.
|
|
*
|
|
* Caller-side cleanup that must run even on a spawn failure (log streams,
|
|
* reporters, temp dirs) belongs in the caller's own try/finally around this
|
|
* call: a spawn 'error' event THROWS from here after the finally block runs,
|
|
* preserving the runners' existing could-not-run handling.
|
|
*/
|
|
export async function runShardChild(options: RunShardChildOptions): Promise<ShardChildResult> {
|
|
const child = spawn(options.command, options.args, {
|
|
cwd: options.cwd,
|
|
env: options.env,
|
|
stdio: ['ignore', 'pipe', 'pipe'],
|
|
detached: process.platform !== 'win32',
|
|
windowsHide: true,
|
|
});
|
|
const groupPid = child.pid ?? null;
|
|
// Group-kill on parent SIGINT/SIGTERM too, not just on timeout.
|
|
const forwarding = installChildSignalForwarding({
|
|
kill: (signal?: NodeJS.Signals | number) => {
|
|
killProcessGroup(child, (signal as NodeJS.Signals) ?? 'SIGTERM');
|
|
return true;
|
|
},
|
|
});
|
|
|
|
let timedOut = false;
|
|
const killTimer = setTimeout(() => {
|
|
timedOut = true;
|
|
killProcessGroup(child, 'SIGKILL');
|
|
}, options.timeoutMs);
|
|
|
|
let exitCode: number | null = null;
|
|
try {
|
|
const streams = options.hookStreams(child);
|
|
// Observe failures now; a pipe can reject before the child closes. Keep
|
|
// that first error until close so final process-group cleanup still runs.
|
|
const drainage = Promise.all(streams).then(
|
|
() => ({ ok: true as const }),
|
|
(error: unknown) => ({ ok: false as const, error }),
|
|
);
|
|
exitCode = await new Promise<number | null>((resolve, reject) => {
|
|
child.once('error', reject);
|
|
child.once('close', (code) => resolve(code));
|
|
});
|
|
const captured = await drainage;
|
|
if (!captured.ok) throw captured.error;
|
|
} finally {
|
|
clearTimeout(killTimer);
|
|
forwarding.dispose();
|
|
// Reap survivors of this shard even on the clean path.
|
|
killProcessGroup(child, 'SIGKILL');
|
|
}
|
|
return { exitCode, timedOut, groupPid };
|
|
}
|