Files
gstack/test/eng-scheduled-regression.test.ts
T
Garry TanandOpenAI Codex 9f81911136 v1.86.0.0 feat: route outside reviews by harness (#2850)
* feat: add a restricted and supervised Claude Code runner

Preserve configured authentication and models while enforcing tool access, strict completion JSON, bounded output and process cleanup. Cover argv, failure handling, session metadata and Windows process containment.

* feat: route outside reviews by harness and migrate wrapper installs

Use Claude Code from Codex and Codex from other supported hosts, with shared invocation rendering, positive gate validation and per-phase provenance. Rename /claude to /claude-code, repair managed shared and copied installations safely, and generate native Kiro skills. Add installed-workflow, failure-injection and live cross-harness regression coverage.

* test: recognize CEO mode labels without terminal spacing

The paid workflow rendered SCOPEEXPANSION at option 4, but its driver required a literal space. Match the leading mode title without cursor-spacing artifacts and ignore adjacent preview text. Preserve missing-target failures and downstream posture assertions.

* test: isolate plan-count fixtures before starting review workflows

Seed the complete test plan in a private git repository before launching Claude, so a bare slash command cannot review the live workspace while a delayed fixture message remains queued. Preserve count thresholds, parsers and budgets. Add initial-context and installed-discovery tests, and retain startup/terminal diagnostics on failed evaluations.

* test: stabilize review fixtures and Claude eval startup

Preserve source boundaries in workflow judge inputs, isolate CEO mode plans, and wait for interactive trust input readiness. Keep startup failure evidence and retain existing models, budgets, and assertions.

Co-Authored-By: OpenAI Codex <noreply@openai.com>

* test: classify collapsed review modes and isolate seeded findings

Keep review questions out of the setup count when terminal cursor positioning removes spaces. State existing webhook safeguards so the five-finding control measures its seeded defects without accidental extra security and concurrency gaps. Preserve question bands and the paired control.

Co-Authored-By: OpenAI Codex <noreply@openai.com>

* test: isolate browser daemon state across free shards

Co-Authored-By: OpenAI Codex <noreply@openai.com>

* test: stabilize native review counting and interactive navigation

Co-Authored-By: OpenAI Codex <noreply@openai.com>

* chore: prepare v1.82.0.0 release

Co-Authored-By: OpenAI Codex <noreply@openai.com>

* fix: eliminate browser and process-cleanup test flakes

Pin every CI surface to Bun 1.4.0 to avoid extra-stdio finalizers closing
reused live sockets. Add an isolated GC/listener regression that fails on
Bun 1.3.13, and prevent coordinated rollback to an affected CI runtime.

Check renderer cleanup against the render's own staging directory so
concurrent renders cannot invalidate the assertion. Make the no-pgrep
process-tree walk tolerate disappearing /proc entries, and synchronize
its test fixture through child readiness and pipe EOF instead of sleeps.

Validation: 9,157 passed, 31 skipped, zero failures across 556 files with
retries disabled. Build, all-host generation freshness, and skill checks
passed. All three races have failing-before/passing-after regressions.

* fix: count completed native review questions in evals

* fix: drive review navigation from confirmed native choices

* fix: require complete section-loading eval reports

* test: isolate telemetry HTTP transport from local assertions

* fix: keep review input on the active native question

* test: let tunnel revocation daemon choose an available port

* test: allocate available ports for pairing and watchdog fixtures

* fix: stabilize planning eval navigation and phase reporting

* test: isolate installed runtime paths in planning evals

* test: stabilize review evidence and concurrent refresh fixtures

* fix: resolve design findings before editing the plan

* fix: honor and persist disabled outside plan reviews

* fix: preserve planning decisions and terminal evidence

Load installed host reviews at autoplan phase entry and wait for completed
reviewers and saved artifacts. Reuse approved remedies while preserving
individual finding decisions.

Drive interactive evals from the current terminal viewport, bind native
questions across scrolling, and require complete native report evidence.
Cover captured stale menus, permission lifecycles, setup classification,
and disabled-review tool availability with deterministic regressions.

Advance release metadata and the upgrade migration to the unclaimed
1.83.0.0 slot.

* fix: drive native review questions and preserve current plans

Use the native single-choice keyboard protocol and current terminal viewport,
with per-question navigation inside packets and completed-call coverage.
Keep permissions, multi-select menus, and Submit controls distinct.

Send Autoplan reviewers the amended implementation plan, keep its review record
separate, and supply retained application contracts in the chain fixture.
Clarify individual DevEx decisions and complete CEO fix options; use one active
plan destination for the section-loading report.

* fix: preserve complete plan-review decisions

* fix: recognize native plan dialogs and reviewer controls

* fix: preserve review decisions and phase completion

* fix: recognize completed reviews without losing findings

* fix: preserve review continuity and native eval completion

* test: fix native review completion and eval retry isolation

* test: handle native review menus and complete eval fixtures

* test: fix native review setup, completion, and isolation failures

* test: limit native skill discovery to runtime assets

* fix: bind Autoplan reviews to full ordered phase inputs

* test: fix planning eval routing, counting, and timeout handling

* chore: advance queued release to v1.84.0.0

* fix: preserve complete review inputs and planning decisions

* fix: reconcile review approvals and preserve phase obligations

* fix: preserve review obligations and unblock eval permissions

Carry recorded Autoplan requirements into blind phase inputs, require Eng
review approvals before exit, and exercise combined asynchronous flows in
CEO reviews. Correct native finding and handoff classification and unblock
repeated report edits using scoped request identities.

* fix: retain plan requirements and complete native review dialogs

* fix: complete native review prompts and retain plan references

* fix: preserve review inputs and classify native eval evidence

* fix: check competing completion orders in CEO reviews

* fix: recognize review decisions and require phase methodology

Require the current phase methodology before Autoplan snapshots. Correct
substantive decision, closed handoff, and cache-finding classification, and
honor the recommended implementation approach in native review dialogs.

Add captured-transcript regressions without changing review thresholds,
provider models, retries, or deadlines.

* test: bind native review decisions and close completed handoffs

* fix: complete review dialogs and verify methodology delivery

* fix: preserve review evidence and unblock native eval prompts

* fix: handle native review question completions

* fix: recognize native review narration and controls

* fix: count native review decisions and isolate eval fixtures

* test: verify seeded review coverage and current artifact permissions

* test: isolate model and brain-aware skill renders

* fix: repair native workflow evaluation and clarify review steps

* fix: stabilize workflow eval evidence and review guidance

* test: repair native workflow observation and fixture isolation

* fix: recognize completed workflow evidence and owned skill reads

* test: repair seeded workflow delivery and completion evidence

* test: recognize current review evidence across native forms

* test: handle native review variants and permission redraws

* fix: honor review preferences and recognize native eval evidence

* test: recognize completed review decisions and queued permissions

* test: match current review contracts and partial-line edits

* test: recognize completed workflow evidence and bounded human waits

* fix: preserve review entry gates and native eval interactions

* fix: recognize native workflow evidence and preserve review gates

* test: recognize current review evidence and preconfigure workflow fixtures

* test: recognize completed review findings and scoped artifact permissions

* fix: stabilize native workflow review and permission evidence

* fix: recognize current review evidence and scoped edit confirmations

Clarify Design and engineering review entry instructions and Design scoring.
Recognize required legacy coverage and public Autoplan completion recaps.
Bind the pending Edit confirmation to its exact file, ordered digest, and
one-request approval when a preceding command display remains visible.
Keep reviews within their existing size limits and preserve scope gates
when extracting workflow fixtures from either supported preamble header.

Keep failure outcomes, review thresholds, provider choices, and eval budgets.

* fix: recover review workflow progress and eval evidence

* fix: recognize valid review evidence and scope selection

* test: fix review evidence parsing and repeated artifact prompts

* test: recognize valid review decisions and pending native cards

* fix(plan-eng-review): keep final navigation consistent with approved tasks

* test: recognize valid review evidence and bind legacy diff requests

* fix: stabilize review eval evidence and harness repair guidance

* docs: update project documentation for v1.85.0.0

Co-Authored-By: OpenAI Codex <noreply@openai.com>

* test: fix Windows CI fixtures and credential scan

Rebase captured JSON values and filesystem evidence using the appropriate
path convention. Compile native fake CLIs on Windows and synchronize pipe
holder readiness, with cleanup retained when assertions fail.

Assemble synthetic credential fixtures at runtime so the added-line scan
keeps enforcing the same gate without flagging its own rejection controls.

Discover generated skills directly for the empty-find regression check,
avoiding a recursive scan through saved evaluation artifacts and dependencies.

* fix: preserve source renders on Windows

Compare canonical generator paths using native separators so an output
sidecar pointing at the source cannot overwrite its skill or metadata.
Keep the regression fixture isolated from the real checkout and expose
freshness diagnostics before asserting subprocess status.

Detach Windows drain-test pipe holders from the fake provider's automatic
child cleanup while preserving the enclosing runner job and its assertions.

* fix: clarify outside review fallback and CEO decisions

Render one applicable own-harness fallback path and retain native review,
disabled policy, and missing-coverage semantics. Align report field names
and mode labels, and make the existing per-cut scope approval explicit.

Regenerate skill outputs and keep the workflow judge's model, thresholds,
and retry policy unchanged.

* chore: move release to free version slot (v1.86.0.0)

PR #2852 now claims v1.85.0.0. Align the release metadata and
rename migration so upgrades from that version still receive it.

Co-Authored-By: OpenAI Codex <noreply@openai.com>

* fix: include engineering review prerequisites and restore branch context

* fix: recognize coverage diagrams and clarify design review instructions

* fix: preserve file identities and join Windows test processes

---------

Co-authored-by: OpenAI Codex <noreply@openai.com>
2026-09-14 14:32:45 -07:00

155 lines
15 KiB
TypeScript

import { expect, test } from 'bun:test';
import { E2E_TOUCHFILES } from './helpers/touchfiles';
import { evaluateEngSeedCoverage } from './helpers/eng-seeded-coverage';
// Minimal verbatim public requirements/tasks/verification from the two failed
// September 11 Eng runs. Replays diagnose the oracle; they do not credit those runs.
const reports: string[] = [
"# Current reviewed plan\n\n### REGRESSION RULE (mandatory, no decision needed)\n\n`legacyAuthFlow()` is existing behavior being rewritten with no existing test\non the changed path. **CRITICAL:** add\n`test/auth/legacyAuthFlow.regression.test` before any rewrite. It records\nthe observable outcomes (status, reason code, cache side effects) of\n`legacyAuthFlow()` for a fixture matrix (valid, expired, wrong issuer, wrong\naudience, revoked, suspended tenant, malformed) and asserts `AuthBroker`\nproduces identical outcomes on the same fixtures.\n\n## Implementation Tasks\n- [ ] **T4 (P1, human: ~4h / CC: ~10min)** — tests — CRITICAL regression test pinning legacyAuthFlow() behavior\n - Surfaced by: Test review REGRESSION RULE — PLAN.md:27-28\n - Files: test/auth/legacyAuthFlow.regression.test\n - Verify: passes against legacy before any rewrite; passes against AuthBroker after\n\n## Verification\n1. Write T4 first and run it against the untouched `legacyAuthFlow()`; it must pass before any other change.\n",
"# Current reviewed plan\n\n### CRITICAL: regression suite for `legacyAuthFlow()` (regression rule, mandatory)\n\n`test/auth/legacyAuthFlow.regression.test.ts`. Captures current behavior\nbefore any rewrite: every success path, every error path, cache interactions,\nand the invalidation hooks it triggers. Runs against the flag-off path after\nthe refactor. This is the highest-priority test in the plan.\n\n## Implementation Tasks\n- [ ] **T4 (P1, human: ~4 hr / CC: ~15 min)** — legacyAuthFlow — CRITICAL regression suite for prior behavior\n - Surfaced by: Test review REGRESSION RULE — PLAN.md:14-16, 27-28\n - Files: `test/auth/legacyAuthFlow.regression.test.ts`\n - Verify: suite green before and after the refactor on the flag-off path\n\n## Verification\n1. Run the regression suite (T4) against the current `legacyAuthFlow()` before touching it; it must be green on the unmodified code.\n"
];
const regression = (plan: string) => evaluateEngSeedCoverage({ status: 'ready', calls: [], assistantMessages: [] }, plan, 0, 1).regression;
const inlineRequired = `# Current reviewed plan
## Required tests
- **CRITICAL regression (T4)** \`auth/legacy-parity.test.ts\`:
Record legacyAuthFlow() outputs before any change. Run the same fixtures
against the new path; assert identical session shape and identical rejection class.
- **Other test** unrelated.test.ts: tests another feature.
## Implementation Tasks
- [ ] **T4 (P1)** — auth/tests — Regression: pin legacyAuthFlow() behavior
- Files: auth/legacy-parity.test.ts
- Verify: suite green on legacy before any refactor commit; green on both paths before rollout
## Verification
1. Run T4 against the untouched legacy path and commit the fixtures first.
2. Land the replacement and run T4 on both paths.
`;
test('inline required regression binds its own task, baseline and same-fixture parity', () => {
expect(regression(inlineRequired)).toBe('plan');
expect(regression(inlineRequired.replaceAll('T4', 'T17').replaceAll('auth/legacy-parity.test.ts', 'spec/old-path.test.ts'))).toBe('plan');
expect(regression(inlineRequired.replace('CRITICAL regression', 'MANDATORY characterization').replace('Record', 'Capture')
.replace('Run the same', 'Replay the same').replace('identical session shape', 'matching outputs')
.replace('suite green on legacy', 'tests pass on the legacy path').replace('untouched', 'unmodified'))).toBe('plan');
for (const change of [
(s: string) => s.replace('CRITICAL regression', 'Optional regression'),
(s: string) => s.replace('CRITICAL regression', 'CRITICAL regression withdrawn'),
(s: string) => s.replace('## Required tests', '## Historical required tests'),
(s: string) => s.replace(' Record', ' If approved, record'),
(s: string) => s.replace('outputs before', 'behavior after'),
(s: string) => s.replace('same fixtures', 'different fixtures'),
(s: string) => s.replace('new path', 'unrelated path'),
(s: string) => s.replace('identical rejection class', 'unspecified behavior'),
(s: string) => s.replace(' - Files: auth/legacy-parity.test.ts', ' - Files: auth/other.test.ts'),
(s: string) => s.replace('green on legacy before', 'red on legacy before'),
(s: string) => s.replace('green on both paths', 'green on new path'),
(s: string) => s.replace('untouched legacy', 'rewritten legacy'),
(s: string) => s.replace('1. Run T4', '1. Run T9'),
(s: string) => s + '\n## Current status\nT4 is "withdrawn".\n',
(s: string) => s + '\n## Current status\nChange T4 assertions to match the new behavior.\n',
(s: string) => s.split('\n').map(line => '> ' + line).join('\n'),
]) expect(regression(change(inlineRequired))).toBeUndefined();
});
test('mandatory named regression suites bind the task to an untouched baseline', () => {
for (const report of reports) {
expect(regression(report)).toBe('plan');
for (const change of [
(s: string) => s.replaceAll('T4', 'T17'),
(s: string) => s.replaceAll('test/auth/legacyAuthFlow.regression.test', 'specs/old-auth.test'),
(s: string) => s.replaceAll('AuthBroker', 'ReplacementBroker'),
(s: string) => s.replace(/[`*]/g, ''),
(s: string) => s + '\n## Future cleanup\nAfter 100% rollout for two weeks, delete legacyAuthFlow() and replace the parity test with a behavioral test.\n',
(s: string) => s + '\n## History\nT4 is withdrawn.\n',
(s: string) => s + '\n## Current assessment\n"T4 is withdrawn."\n',
(s: string) => s + '\n## Payment regression suite\nThe regression suite is withdrawn.\n',
(s: string) => s + '\n## Current assessment\nAfter committing the green baseline, run T4 after rewriting legacyAuthFlow().\n',
(s: string) => s.replace('before any rewrite:', 'before any rewrite: A token receives success if accepted by legacyAuthFlow(). Rejected inputs receive the recorded error.'),
]) expect(regression(change(report))).toBe('plan');
}
});
const controls: Array<[string, (s: string) => string]> = [
['missing declaration', s => s.replace(/### [\s\S]*?(?=## Implementation Tasks)/, '')],
['optional declaration', s => s.replaceAll('mandatory', 'optional')],
['never mandatory', s => s.replaceAll('mandatory', 'never mandatory')],
['missing legacy subject', s => s.replaceAll('legacyAuthFlow', 'otherAuthFlow')],
['no baseline capture', s => s.replace(/records|Captures/g, 'describes')],
['late baseline', s => s.replaceAll('before any rewrite', 'after any rewrite')],
['missing task', s => s.replace(/- \[ \] \*\*T4 [\s\S]*?(?=## Verification)/, '')],
['different task file', s => s.replace(/ - Files: .*/, ' - Files: test/other.test.ts')],
['missing verification', s => s.replace(/ - Verify: .*/, '')],
['missing baseline step', s => s.replace(/^1\. .*$/m, '')],
['rewritten baseline', s => s.replace(/untouched|unmodified/g, 'rewritten')],
['wrong task in baseline', s => s.replace(/^(1\. .*)T4/m, '$1T9')],
['quoted report', s => s.split('\n').map(l => '> ' + l).join('\n')],
['quoted baseline', s => s.replace(/^(1\. )(.*)$/m, '$1"$2"')],
['quoted declaration', s => s.replace(/(### [^\n]+\n)([\s\S]*?)(?=\n## Implementation Tasks)/, '$1"$2"')],
['fenced report', s => '```\n' + s + '\n```'],
['historical report', s => s.replace('Current reviewed plan', 'Historical reviewed plan')],
['source declaration', s => s.replace(/(### [^\n]+\n)/, '$1Source:\n')],
['conditional declaration', s => s.replace(/(### [^\n]+\n)/, '$1If approved,\n')],
['conditional task', s => s.replace('## Implementation Tasks\n', '## Implementation Tasks\nOnce approved,\n')],
['source task', s => s.replace('## Implementation Tasks\n', '## Implementation Tasks\nSource:\n')],
['explicit cancelled task', s => s + '\n## Current assessment\nDo not run T4.\n'],
['withdrawn task', s => s + '\n## Current assessment\nT4 is withdrawn.\n'],
['withdrawn verification', s => s + '\n## Current assessment\nT4 verification is optional.\n'],
['withdrawn legacy suite', s => s + '\n## Current assessment\nThe legacy regression suite is not required.\n'],
['quoted status', s => s + '\n## Current assessment\nT4 is "withdrawn".\n'],
['changed before baseline', s => s + '\n## Current assessment\nlegacyAuthFlow() is rewritten before T4.\n'],
['conditional mandatory heading', s => s.replace('mandatory', 'mandatory if approved')],
['conditional numbered baseline', s => s.replace(/^1\. /m, '1. Once approved, ')],
['conditional verification line', s => s.replace(' - Verify: ', ' - Verify: If approved, ')],
['hypothetical baseline', s => s + '\n## Current assessment\nT4 baseline verification is hypothetical.\n'],
['current post-rewrite instruction', s => s + '\n## Current assessment\nRun T4 only after rewriting legacyAuthFlow().\n'],
['post-rewrite baseline row', s => s.replace(/^1\. .*$/m, '1. Rewrite legacyAuthFlow() first, then run T4 against the unmodified legacyAuthFlow() snapshot; it must be green before rollout.')],
];
test.each(controls)('%s cannot provide mandatory regression coverage', (_, change) => {
for (const report of reports) {
const changed = change(report);
expect(changed).not.toBe(report);
expect(regression(changed)).toBeUndefined();
}
});
test('the regression evidence test selects its existing Eng workflow', () => {
expect(Object.entries(E2E_TOUCHFILES).filter(([, paths]) => paths.includes('test/eng-scheduled-regression.test.ts')).map(([name]) => name)).toEqual(['plan-eng-finding-count']);
});
// Verbatim owned rule, T1 and ordered verification from the failed AZ report.
const orderedRuleReport = "# Current reviewed plan\n\n### REGRESSION RULE — CRITICAL, no approval needed (skill iron rule)\n\nPLAN.md:27-28 rewrites `legacyAuthFlow()` with no regression test;\nPLAN.md:14-16 excluded it from coverage. That is modified existing behavior\nwith no covering test. **Before** the rewrite, add\n`legacyAuthFlow.characterization.test.ts` capturing current outputs for:\nvalid token, expired token, wrong tenant, wrong audience, revoked token,\nIDP unavailable. The rewrite must pass the same suite unchanged.\n\n## Implementation Tasks\n- [ ] **T1 (P1, human: ~half day / CC: ~15min)** — legacyAuthFlow — Write characterization suite for 6 prior behaviors BEFORE rewrite\n - Surfaced by: Test review — REGRESSION RULE, PLAN.md:27-28 and 14-16\n - Files: `legacyAuthFlow.characterization.test.ts`\n - Verify: suite green on current code; green again after rewrite\n\n## Verification (end to end)\n1. Run T1's characterization suite on unmodified code: green.\n2. Implement T2-T6; run unit suites: green, no shared-state ordering flakes (run with shuffled order).\n3. Run T7 E2E: A/B isolation, suspend-mid-mint denial, double-submit consistency all green.\n4. Re-run T1 after the rewrite: green, unchanged.\n";
test('an iron-rule declaration and ordered task verification establish the mandatory baseline',()=>{
expect(regression(orderedRuleReport)).toBe('plan');
for(const change of [(s:string)=>s.replaceAll('T1','T17'),(s:string)=>s.replaceAll('legacyAuthFlow.characterization.test.ts','spec/legacy-golden.test.ts'),(s:string)=>s.replace(/[`*]/g,''),
(s:string)=>s+'\n## History\nT1 is withdrawn.\n',(s:string)=>s+'\n## Current assessment\n"T1 is withdrawn."\n',(s:string)=>s+'\n## Payment regression suite\nThe regression suite is withdrawn.\n'])expect(regression(change(orderedRuleReport))).toBe('plan');
});
test('the ordered baseline stays owned, required, and unchanged across the rewrite',()=>{
for(const [before,after] of [
['no approval needed','optional if approved'],['skill iron rule','hypothetical example'],['legacyAuthFlow','otherAuthFlow'],
['**Before** the rewrite','After the rewrite'],['capturing current outputs','capturing expected outputs'],
['The rewrite must pass the same suite unchanged.','The rewrite may update the expectations.'],
['suite green on current code; green again after rewrite','suite green on changed code; green again after rewrite'],
['suite green on current code','suite is not green on current code'],['green again after rewrite','not green again after rewrite'],
['on unmodified code: green.','on unmodified code: not green.'],['after the rewrite: green, unchanged.','after the rewrite: failing, unchanged.'],
['1. Run T1','1. Run T9'],['on unmodified code: green','on changed code: green'],['1. Run ','1. If approved, Run '],
['4. Re-run T1','4. Re-run T9'],['green, unchanged.','green, with updated expectations.'],
['## Implementation Tasks\n','## Implementation Tasks\nSource:\n'],['## Implementation Tasks\n','## Implementation Tasks\nOnce approved,\n'],
['PLAN.md:27-28','Source:\nPLAN.md:27-28'],['Current reviewed plan','Historical reviewed plan'],
]){const changed=orderedRuleReport.replaceAll(before!,after!);expect(changed).not.toBe(orderedRuleReport);expect(regression(changed)).toBeUndefined();}
for(const change of [(s:string)=>s.replace(/^ - Files: .*$/m,' - Files: different.test.ts'),(s:string)=>s.replace(/^1\. .*$/m,''),
(s:string)=>s.replace(/^1\. .*$/m,'1. Rewrite legacyAuthFlow() before recording T1.'),(s:string)=>s.replace(/^(1\. )(.*)$/m,'$1"$2"'),
(s:string)=>s.replace(/^4\. .*$/m,''),(s:string)=>s.split('\n').map(l=>'> '+l).join('\n'),(s:string)=>'```\n'+s+'\n```',
(s:string)=>s+'\n## Current assessment\nT1 is withdrawn.\n',(s:string)=>s+'\n## Current assessment\nT1 verification is "optional".\n',
(s:string)=>s+'\n## Current assessment\nDo not run T1.\n',(s:string)=>s+'\n## Current assessment\nlegacyAuthFlow() is rewritten before T1.\n',
(s:string)=>s+'\n## Current assessment\nUpdate T1 assertions.\n']){const changed=change(orderedRuleReport);expect(changed).not.toBe(orderedRuleReport);expect(regression(changed)).toBeUndefined();}
});
test('required regression relations survive heading, task and verification paraphrases',()=>{
const changed=orderedRuleReport.replace('REGRESSION RULE — CRITICAL, no approval needed (skill iron rule)','Required characterization baseline')
.replace('The rewrite must pass the same suite unchanged.','The same suite must remain green unchanged after the rewrite.')
.replace('Write characterization suite for 6 prior behaviors BEFORE rewrite','Add characterization tests for existing outputs')
.replace('suite green on current code; green again after rewrite','current implementation passes; after the rewrite the suite passes again')
.replace("1. Run T1's characterization suite on unmodified code: green.",'1) Execute characterization task T1 against untouched code; it must pass.')
.replace('4. Re-run T1 after the rewrite: green, unchanged.','4) Execute the same T1 tests unchanged after the refactor; they must pass.');
expect(regression(changed)).toBe('plan');
});