refactor(test): mechanical sweep — 298 paid-test timeouts onto eval-budget tiers

69 files, both shapes (trailing bun-test budgets and runner
timeout/timeoutMs options), ROUND-UP ONLY so nothing that passed can
start failing: 75 → JUDGE_MS, 137 → CAPTURE_MS, 74 → CAPTURE_LONG_MS,
9 → PTY_MS, 3 → PTY_LONG_MS. Raw >=60s literal count in the paid scope:
395 → 97, of which 51 are non-timeout noise (fixture dates, run IDs)
and 46 are enumerated justified holds (comment-carrying calibrated
budgets, poll-loop constants, utility spawn waits, and the seven
physical-ceiling 1_500_000 sites). The eval-budgets policy ratchet
keeps the residue from regrowing.

Known collapse: where an inner runner budget and its enclosing test
budget now share a tier, the old stagger is gone — an overrun surfaces
as a bun test timeout instead of a graceful runner timeout
(diagnosability trade, not a correctness one).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
This commit is contained in:
Garry Tan
2026-08-29 05:25:06 +00:00
co-authored by Claude Fable 5
parent 74ae357e0f
commit 6841183c35
69 changed files with 367 additions and 298 deletions
+9 -8
View File
@@ -18,6 +18,7 @@
* accordingly.
*/
import { describe, test, expect, beforeAll, afterAll } from 'bun:test';
import { CAPTURE_MS } from './helpers/eval-budgets';
import { runSkillTest } from './helpers/session-runner';
import {
ROOT, runId,
@@ -134,7 +135,7 @@ ${captureInstruction(outFile)}
After writing the file, stop. Do not continue the review.`,
workingDirectory: planDir,
maxTurns: 10,
timeout: 240_000,
timeout: CAPTURE_MS,
testName: 'plan-ceo-review-format-mode',
runId,
model: 'claude-opus-4-7',
@@ -160,7 +161,7 @@ After writing the file, stop. Do not continue the review.`,
result,
passed: ['success', 'error_max_turns'].includes(result.exitReason),
});
}, 300_000);
}, CAPTURE_MS);
});
// --- Case 2: plan-ceo-review approach menu (coverage-differentiated) ---
@@ -191,7 +192,7 @@ ${captureInstruction(outFile)}
After writing the file, stop. Do not continue the review.`,
workingDirectory: planDir,
maxTurns: 10,
timeout: 240_000,
timeout: CAPTURE_MS,
testName: 'plan-ceo-review-format-approach',
runId,
model: 'claude-opus-4-7',
@@ -216,7 +217,7 @@ After writing the file, stop. Do not continue the review.`,
result,
passed: ['success', 'error_max_turns'].includes(result.exitReason),
});
}, 300_000);
}, CAPTURE_MS);
});
// --- Case 3: plan-eng-review coverage-differentiated per-issue AskUserQuestion ---
@@ -250,7 +251,7 @@ ${captureInstruction(outFile)}
After writing the file with that ONE question, stop. Do not continue the review.`,
workingDirectory: planDir,
maxTurns: 10,
timeout: 240_000,
timeout: CAPTURE_MS,
testName: 'plan-eng-review-format-coverage',
runId,
model: 'claude-opus-4-7',
@@ -275,7 +276,7 @@ After writing the file with that ONE question, stop. Do not continue the review.
result,
passed: ['success', 'error_max_turns'].includes(result.exitReason),
});
}, 300_000);
}, CAPTURE_MS);
});
// --- Case 4: plan-eng-review kind-differentiated per-issue AskUserQuestion ---
@@ -306,7 +307,7 @@ ${captureInstruction(outFile)}
After writing the file with that ONE question, stop. Do not continue the review.`,
workingDirectory: planDir,
maxTurns: 10,
timeout: 240_000,
timeout: CAPTURE_MS,
testName: 'plan-eng-review-format-kind',
runId,
model: 'claude-opus-4-7',
@@ -332,7 +333,7 @@ After writing the file with that ONE question, stop. Do not continue the review.
result,
passed: ['success', 'error_max_turns'].includes(result.exitReason),
});
}, 300_000);
}, CAPTURE_MS);
});
afterAll(async () => {