Commit Graph
512 Commits
Author SHA1 Message Date
garrytan 18ea34b949 ci(evals): name the PR-comment loop's unused fields so shellcheck passes (SC2034) 2026-09-29 22:02:03 +00:00
garrytan 58b5e3f977 Merge origin/main (v1.91.9.0, #2998) into capy/audit-fix-wave; release as v1.91.10.0
Main shipped v1.91.9.0, so this wave becomes v1.91.10.0. Conflicts kept this
branch's planner-derived assertions. Main's new test-value eval gets rule
kinds for its three gate cases and its measured 332 s duration from #2998's
CI; census counts and the PR fallback floor follow the new file. The packing
test now weighs each tier's recorded durations the way the planner does.
2026-09-29 22:00:14 +00:00
garrytan bfad7fd37d docs: final census numbers in the v1.91.9.0 entry; file the paid-eval follow-ups 2026-09-29 21:50:57 +00:00
garrytan 8dd3c44fc6 fix(qa): checkpoint receipts print the report link for their exploration file
qa-functional-webhook-report failed in two of three censuses because the
report linked .qa-evidence/NNN capture folders as "checkpoints" and never
linked exploration-NNN.json. The checkpoint receipt now prints
link: [checkpoint NNN](exploration-NNN.json), and the functional report
template says capture folders are not checkpoints.
2026-09-29 21:49:46 +00:00
Garry Tan 96764e80a6 v1.91.9.0 feat: test value bar in plan-eng-review, review, qa and ship, plus /test-audit (#2998) 2026-09-29 14:35:00 -07:00
garrytan 958d1ceba6 test: supply holdDeferKeepIndex to the CEO routing mocks and follow split-overflow into the marathon lane
The registered-callback fixtures mock ceo-mode-option and lacked the new
export; the split fixtures asserted the periodic tier; the registered-budget
check looked for split-overflow only in the periodic manifest.
2026-09-29 21:33:13 +00:00
garrytan bf4667146c test: attribute quoted prior-record field lists, state the judge reason bound in its schema, move split-overflow to marathon
Census 36629958451 reds:
- outside-plan-disabled-no-fallback: the model quoted the pre-existing record
  as a parenthesized field list with its exact timestamp; attribution now
  requires that exact timestamp and the record's own field values.
- plan-devex-peer-comparison-classification: the judge correctly returned
  missing but wrote a 1069-character reason, voiding the judgment; structured
  outputs cannot enforce maxLength, so the bound is stated on the field.
- plan-ceo-split-overflow ran 504-1188 s as one PTY flow and set the
  periodic lane's wall clock; it now runs weekly in the marathon lane.
2026-09-29 21:26:37 +00:00
garrytan f63e1fb7cb ci(image): keep Claude Code 2.1.251; test(ceo-mode-routing): keep HOLD's own deferrals in scope before assessing its rigor decision
2.1.284 enables per-turn effort for claude-fable-5-1: in gate census
36626737820, 66 of 84 sessions ran longer than on 2.1.251 (+20% session time,
+32% thinking tokens) and 11 cases timed out on unchanged budgets.

HOLD SCOPE's 0G step asks its own defer/keep menu; the actor answered it
Defer and the assessment then judged that scope question as the rigor
decision. The actor now answers that menu Keep and assesses the next one.
2026-09-29 20:57:04 +00:00
garrytan d05d161713 test(autoplan-dual-voice): unwrap Claude Code 2.1.284 subagent hand-back frames; accept read-only probe diagnostics; record before asserting
Census run 36626737820: the native CEO report arrived framed and indented, so
its INPUT line never matched, and the model's exact probe plus two variable
echoes was not canonical. A column-zero line inside a frame, command
substitution, backticks, redirects, assignments, CODEX_MODE echoes and output
line-count mismatches stay rejected. The failure now records before asserting.
2026-09-29 20:48:37 +00:00
garrytan 3fc05932b7 test: settle the post-response composer before seeding; give the TPA recorder adapter its infra helper
submitPlanSeed accepted a stale empty composer when the transcript recorded
end_turn before the CLI repainted (late-repaint-typed-current fails 5/5 on the
old helper, passes 5/5 now). The TPA recording fixture extracted recordE2E
without isPreTurnInfraFailure, so every failed case threw before recording.
2026-09-29 20:43:00 +00:00
garrytan 80c92c86ed chore(release): v1.91.9.0 2026-09-29 20:29:34 +00:00
garrytan 581e603313 fix(evals): --case list mode and name precheck; case-shard qa-callers; refresh batching and design-with-ui seeds 2026-09-29 20:28:31 +00:00
garrytan 2dae4bf944 Merge remote-tracking branch 'origin/capy/rel-c' into capy/audit-fix-wave 2026-09-29 20:23:51 +00:00
garrytan 475667ff5d test(eng-batching): read the report target as a field, not a spelling
The next targeted rerun (Claude Code 2.1.284) again asked eleven separate
native questions and again counted zero: its briefs named no plan and its
report declared '- **Review target (fixed):** `/abs/PLAN.md`' under
'# Eng Review — PLAN.md: <plan>'. An unsourced brief now inherits the one
current target field that names a PLAN.md file, whatever its list or
emphasis markup; its ledger record still supplies the cited finding and
must reproduce the brief exactly. A brief that names its plan must still
match the report title. Replays of all three captures count 9, 9 and 3;
controls reject a foreign, duplicate or missing target and an archived
title.
2026-09-29 20:12:02 +00:00
garrytan dc640acd8a test: pin every-record outcome counts and the twelve doc-sync callbacks 2026-09-29 19:55:37 +00:00
garrytan e758fbe96d test(evals): record a pre-turn API or CLI failure as infra
recordE2E sets failure_class 'infra' on a failed session whose runner
reports error_api, timeout_startup, error_output_stream or a non-zero CLI
exit with zero turns and no assistant event. A model refusal, a timeout
after model work, max turns, or an explicit caller pass/class keeps its
ordinary classification.
2026-09-29 19:51:35 +00:00
garrytan f4ab5ee75f test(judges): sample the recommendation rubric as a panel; never re-ask armJudge
llm-judge-recommendation is a judge case: each fixture now draws a
3-sample judgePanel, gates reason_substance on the panel mean and the
present/commits/has_because checks on a 2-of-3 majority, thresholds
unchanged. armJudge no longer re-asks on a malformed verdict; it is a
failed sample, as the judge policy requires.
2026-09-29 19:51:35 +00:00
garrytan 2fb55e4c7a test(evals): record the read-only and detector-row invariants as contracts
shared-libs-opportunity-judgment and review-design-lite are behavior
cases: their recommendation and checklist judgments may vary, but the
read-only invariant (commands, provider requests, fixture bytes, hooks,
state) and the deterministic fake-engine detector rows are contracts.
Both now go through expectContract, so any failure vetoes the panel.
2026-09-29 19:51:35 +00:00
garrytan 1e640bbfb7 fix(evals): tsx-safe generics in eval-flake-rank, legacy artifact names, no-retry wall docs 2026-09-29 19:47:44 +00:00
garrytan b2ca207cf0 Merge remote-tracking branch 'origin/capy/rel-a' into capy/rel-c 2026-09-29 19:47:17 +00:00
garrytan 74c3184dcc test(pty): grant an owned Create pane whose title row is cropped
The targeted batching rerun on Claude Code 2.1.284 left its first report
Write unanswered for 1,372 s and timed out: the viewport began at the
pane's relative file row and rule, with the 'Create file' title cropped
above, so the preview parser rejected the file row as foreign. That row
must now resolve to the owned path and is skipped before the unchanged
line-by-line preview match. Replay controls reject another file, another
directory and an edited preview row.
2026-09-29 19:47:15 +00:00
garrytan 32772f535d feat(evals): --case/--trials local diagnosis and panels in local sharded runs
bun run scripts/test-paid-shards.ts --case <id> [--trials N] runs N
independent trials of one case through the CI panel runner (trial shards,
TRIAL_ENV identity, name-pattern isolation) and prints its panelVerdict();
N defaults to the case's policy panel and CI never reads it. The local
sharded path (test:gate:sharded, test:periodic:sharded) now plans the same
trial shards and exclusions as CI and exits on execution completeness plus
panel verdicts.
2026-09-29 19:44:55 +00:00
garrytan 3bae8e33da feat(evals): planner-side whole-panel reuse and negative receipts
The planner job restores this PR's receipt store once and ships a single
filtered set with the plan: a pass or panel receipt with a same-or-newer
FAIL for its input identity is dropped, and a panel receipt ships only as
a whole PASS panel (re-verified with panelVerdict) from one run. Executors
read only that set (no per-slice cache restore or save), so every trial of
a panel sees the same receipts; a trial reuses its own record from the
panel receipt, keeping a split PASS's failed trial.

Trial identities drop the trial index (run-scoped) and bind the panel
policy. Executed shards carry their input identity; the report turns a
whole fresh PASS panel into a panel receipt and a FAIL panel or failed rule
shard into a negative receipt, and marks a panel that mixes reused and
fresh trials INCOMPLETE. The report job merges plan, slice and report
receipts (newest per file) and saves one store per run.

Also fixes two TS2352 casts in browse/test/dia-macos-qualification.test.ts
whose diagnostic text drifted with program order (baseline locked, fix only).
2026-09-29 19:41:39 +00:00
garrytan 8622535b90 ci(evals): attempt-scoped artifacts, verdict-v2 PR comment, weekly pass-rate gate and one INFRA re-dispatch
- Slice, census and marathon artifacts carry -a<run_attempt>; reports
  download them per artifact (no merge), so records never overwrite and a
  re-run never replaces the first attempt's verdict.
- Planners pass --max-parallel for the capacity preflight (24/16 unchanged:
  the refreshed periodic plan needs 24 slices, the gate census 12).
- PR comment: jq-only job reads collector-outcomes v2 (headline, sanitized
  failure block); the group_by(.name)|last recomputation is gone.
- Reports stamp series identities, upload trial-outcomes-* for history, and
  shard logs upload always (a failed trial no longer reds its runner).
- Weekly report: headline + failure block of both lanes in the issue body,
  the eval:pass-rates --gate step (fails closed without history), close the
  issue on a green run, and UC-E1: when every red is machine-classified
  INFRA/INCOMPLETE, one re-dispatch as a new run in its own concurrency
  group (redispatch_of), both runs reported.
2026-09-29 19:34:33 +00:00
garrytan 62fb9a255d feat(evals): stamp trial series identities and fit panels to the live registry
- scripts/eval-trial-series.ts stamps series_identity (eval-flake-rank's
  caseSeriesIdentities) on a report's trial-outcomes JSONL as its own step,
  keeping the history tool out of the paid runner's closure;
  TrialOutcomeRecord gains the optional series_identity field.
- Slice-count plans let a registered trial spill into an ordinary lane when
  its siblings hold every long lane, so panels never share a runner.
- Re-audited test-selection.ts (Stream B added the E2E_KINDS/BEHAVIOR_WHY
  map-diff; no new module loading) and repinned its hash.
- Detach and release floors now count trial shards (66 periodic trials in
  22 panels): periodic floor 33,821s, still under eval:bg:periodic's 67,380s.
- Coordination fixtures supply the executor's trial records.
2026-09-29 19:30:43 +00:00
garrytan d5876efa5d chore(evals): refresh paid duration seeds from proof runs 36597762183 and 36606688266
Both tiers, merged in run order (the later run wins). Notable: split-overflow
1332s -> 504s, section-loading 604s -> 342s, mode-routing 575s -> 444s;
multi-finding-batching 734s -> 1318s (its red path in run 36606688266).
2026-09-29 19:30:43 +00:00
garrytan a385e5de18 Merge remote-tracking branch 'origin/capy/rel-b' into capy/rel-a 2026-09-29 19:24:03 +00:00
garrytan 8eb55b6ebf feat(evals): trial planner, slice exit split and panel-verdict report
Planner: behavior and quarantined cases become panels of isolated trial
shards (<file>#<id>~t<N>) bound by EVALS_SELECTION_JSON=[id] and the exact
test name; the file shard excludes them by name. Trials of one case never
share a slice, result slugs are unique, panels are validated whole, unknown
registrations throw, and the planner prints a capacity preflight.

Executor: each trial shard gets its TRIAL_ENV identity and a trial record
(outcome, failure class, cause, cost); every shard writes a JUnit report.
The slice exit now means execution completeness: a failed rule shard or a
trial without a record reds the runner, a failed trial does not.

Report: panelVerdict() decides every panel of the first run attempt (later
attempts are reported, never replacing it); rule shards keep the unchanged
fail-closed checks; collector records all count (no last-attempt wins);
census runs enforce the quarantine cap and expiry. It writes
collector-outcomes v2, trial-outcomes.jsonl (trials plus JUnit rule/judge
cases), report-summary.md, and one headline + failure block with rerun
commands, and flags INFRA/INCOMPLETE-only reds for the one re-dispatch.

The fail-open suite gains the panel cases: behavior 1/3 red, 2/3 green
with its failed trial shown, missing trial INCOMPLETE, contract at 2/3 red,
quarantined 1/3 green, 0/3 and contract red, missing slice red, and a later
attempt never replacing the first.
2026-09-29 19:23:55 +00:00
garrytan ccb5f3c07c docs(evals): document the pre-registered verdict policy, quarantine, pass-rate history and arithmetic
AGENTS.md replaces the retry rule with the approved policy text (no retries;
kind fixes trials; no added trials, samples or dispatches after a result;
quarantine by CASE_QUARANTINE only; one INFRA/INCOMPLETE re-dispatch) and
notes that a pre-registered fixed panel is not rejudging. CONTRIBUTING gains
the kind rules, the judge panel, eval:pass-rates and an 'Add a paid eval'
checklist. TESTING_INTERNALS describes verdicts, quarantine, history and the
arithmetic, including the rule term: 1 trial vs 2-of-3 red rates at
p = 0.99/0.95/0.90/0.70/0.30 and lane all-green probabilities for the
current 191 rule / 22 behavior / 25 judge registry.
2026-09-29 19:17:12 +00:00
garrytan 49761c97d6 test(ceo-mode-routing): submit a mode review that scrolled past the viewport
Run 36606688266 bundled routing, learnings and the mode choice into one
native call. Its review panel was taller than the terminal, so the tab
bar scrolled off, ceoModeSubmissionInput returned null for 240 s and HOLD
SCOPE was never submitted ('no posture match'). With no bar on screen the
viewport must still end at the focused Submit prompt, and the accumulated
screen text supplies the one complete panel, authenticated exactly as
before. Replay controls reject another mode, an unoffered answer, an
altered question, a quoted panel, trailing output, a moved cursor and an
answered or changed call.
2026-09-29 19:16:01 +00:00
garrytan 69cb4c8e9d fix(deslop-shared-libs): read related sources together within the turn limit
Run 36606688266's opportunity audit read sixteen sources one per turn and
stopped at error_max_turns; the passing run 36597762183 read the same
files in three batched commands. The skill now says turns are bounded and
asks for parallel reads or one read-only command per step.
2026-09-29 19:16:00 +00:00
garrytan 65f28fa9e4 fix(plan-design-review): treat a designer with no API key as unavailable
Both proof runs (36597762183, 36606688266) printed DESIGN_READY, hit
'No OpenAI API key found' on the first $D variants call, then hand-built
HTML/CSS wireframes, screenshots and a comparison board for ~195-245 s
before the first review question; the second run timed out at 600 s.
A failed first generation now takes the existing text-only path, and the
skill forbids substituting hand-built mockups.
2026-09-29 19:16:00 +00:00
garrytan 8bd53faa6c test(eng-batching): bind unsourced native briefs through the report's target
Run 36606688266 asked ten separate native review questions (D1-D9 bound
to ledger records R1-R9) and failed reviewCount=0 < FLOOR=3: its briefs
named the plan by title instead of citing PLAN.md, its report declared
'Review target (fixed): PLAN.md' under '# Engineering review: <plan>', and
it kept an unfenced copy of the plan's own H1. The named-source route now
accepts those spellings and non-inline ledger briefs. The same replay
rejects a foreign, mixed, duplicate or missing target, another plan's
title or copied H1, a brief naming another plan or file, a mismatched
saved brief, and re-asks. The run-36597762183 capture still counts 3.
2026-09-29 19:15:06 +00:00
garrytan 5269452983 test(eng-batching): grade the floor once the review report is complete
A completed GSTACK REVIEW REPORT ends the review, so the review-question
count is final there. Run 36606688266 wrote its report at 1,248 s and
closed the session at 1,318 s; the case now stops collection and applies
the unchanged floor at the report instead of waiting out the session.
No budget changes.
2026-09-29 19:15:06 +00:00
garrytan 392d63a537 ci(image): pin Claude Code 2.1.284 so the eval model is recognized
2.1.251 logs [claude-code:unrecognized_model] for claude-fable-5-1, the
eval capture/judge default. 2.1.284 does not. The gate PTY smoke subset
(plan-ceo/plan-devex plan-mode, plan-mode-no-op) parses on the new TUI;
plan-design-review-plan-mode passed at 293 s on 2.1.284 and timed out at
300 s on 2.1.251 on the same tree.
2026-09-29 19:14:54 +00:00
garrytan 8ee9f887ed feat(eval-pass-rates): attribute legacy records by the exact slug of their display name 2026-09-29 19:13:54 +00:00
garrytan d6b3559b78 feat(evals): per-case pass rates with Wilson intervals, identity series and quarantine policy
scripts/eval-flake-rank.ts becomes eval:pass-rates (eval:flake-rank stays an
alias, and the legacy aggregate stays exported). It reads eval-store's
trial-outcomes JSONL from the last N completed evals-periodic runs on this
branch and main (gh, downloading only the trial-outcomes artifact, cached and
size-capped, parsed as data), plus local eval dirs, and prints per-case
per-trial pass rates with 95% Wilson intervals.

A series is a case's own touchfiles minus GLOBAL_TOUCHFILES
(caseSeriesIdentities, for the report job to stamp), per model, CLI version
and policy version. Labels: INCONCLUSIVE, BROKEN, FLAKY, FAILING, PASSING.
--backfill imports legacy slice artifacts as pre-policy trials (first
attempt only, attributed by registry id, never guessed) for display only.

--gate fails with ACTION REQUIRED on post-policy evidence only: drift below
the quarantine entry rule, a rule case behaving like behavior, a one-sided
Fisher drop against the previous identity (Holm-controlled), and quarantine
entries that met their exit rule, expired after 8 weekly runs, broke the
10% tier cap or are invalid. CASE_QUARANTINE entries now carry a
failureClass (detector, harness or model-latency); a product defect has no
class and is never quarantined. The policy test pins EVAL_POLICY's approved
constants.
2026-09-29 19:11:27 +00:00
garrytan 0292ee3f6f test(evals): classify every live case and re-select a case when its kind changes
E2E_KINDS: rule by default (191 E2E ids), 22 behavior cases whose verdict is
a live model choice with an acceptable sub-100% per-trial rate, each with a
BEHAVIOR_WHY tolerance, and 25 judge entries (the 24 workflow judges plus the
fixed-fixture llm-judge-recommendation rubric check). Contract-shaped cases
(ask-before-decide, plan-mode no-writes, mandated steps, secrets, the batching
floor) stay rule. Behavior requires a known literal registration and an exact
Bun test name so the case runs as its own trial shard.

Map-diff selection now diffs E2E_KINDS and BEHAVIOR_WHY per key, and a base
revision without them selects every key, so a kind flip runs the panel it
introduces. test/eval-kinds.test.ts enforces coverage, tolerances,
isolatability and the reviewed counts, printing the literal to add.
2026-09-29 19:11:04 +00:00
garrytan 3f68572cab test(llm-judge): sample every judge as a pre-registered 3-sample panel
Each of the 24 skill-llm-eval judges now draws EVAL_POLICY.judge.samples
independent samples of the same prompt concurrently inside the unchanged
JUDGE_MS budget. Numeric dimensions gate on the per-dimension panel mean
against the unchanged threshold; booleans (would_browse, consistent) on a
strict majority. An erroring sample fails the whole panel and is never
resampled; a refusal is an unscored panel only when every sample refused.
callJudge's 429 backoff stays: it is transport before any model output.

The workflow-judge cache stores and validates only complete panels, and its
identity now records the panel and zero file retries. Harness tests that
pinned one provider call per case now pin the panel size.
2026-09-29 19:11:04 +00:00
garrytan b1f5bc0032 test(evals): retire every paid automatic retry
Paid evals never retry (approved 2026-09-29): delete SHORT_CASE_RETRY_FILES
and retriesWithinCaseCap, drop the retry fields from the registered wall rows
(walls now cover one run plus reserve), make retriesForFiles return 0, pass
--retry 0 explicitly, and drop --retry 1 from the package.json paid scripts.
Add the eval:pass-rates alias. Tests that pinned the old retry allowance are
updated as a policy change; review-finalization-budget now proves late-result
recording under the production zero-retry arguments.
2026-09-29 19:09:51 +00:00
garrytan a427f05c58 test(evals): pin the fail-closed rule-shard gate through the real --report path
Synthetic slice artifacts for rule fail, timeout, missing slice, unreported
entry, hollow, never-started, collector failure and wrong-slice reports all
exit red before the panel-verdict gate change lands.
2026-09-29 19:02:34 +00:00
garrytan adced7e046 test(evals): add trial records, panelVerdict, expectContract and trial-outcomes JSONL
EvalTestEntry gains case_id, kind, trial, panel, failure_class and
policy_version, stamped from the runner's TRIAL_ENV on isolated trial
shards. panelVerdict() is the single verdict function (INCOMPLETE on
missing or duplicate trials, contract veto at any count, quarantine
hard-break rule, INFRA/INCOMPLETE machine classification). expectContract()
records failure_class 'contract' on the collector entry and a sidecar
before throwing. trial-outcomes JSONL has a fail-closed writer and a
data-only reader.
2026-09-29 18:56:27 +00:00
garrytan 3c441fe311 test(evals): add E2E_KINDS, BEHAVIOR_WHY, EVAL_POLICY and CASE_QUARANTINE skeletons
Every E2E_TIERS and LLM_JUDGE_TOUCHFILES key starts as 'rule'; BEHAVIOR_WHY
and CASE_QUARANTINE start empty. EVAL_POLICY pre-registers the approved
panel (3, majority 2), quarantine entry 0.95/10 and exit 0.97/10, 10% cap,
8-weekly-run expiry, Fisher drift alarm and one INFRA re-dispatch.
2026-09-29 18:56:27 +00:00
garrytan 4c5fc0e2ce test(qa-bugs): keep claude-opus-4-7 after qa-b6-static stalled on the default model
qa-b6-static timed out on claude-fable-5-1 in census 36597762183 and in one
of two targeted reruns. Both times the stream stopped mid-message with no
pending tool, right after the model found the disabled submit button, and
stayed silent until the 300 s deadline. Per the B8 fallback, re-pin with a
TODOS entry; budgets and retries are unchanged. A rerun on opus-4-7 passed
(125 s, 5/5 detected).
2026-09-29 17:28:32 +00:00
garrytan 1c32c7d16a test(ceo-mode-routing): accept the skill-mandated Note form and Recommendation reason as HOLD posture
HOLD Defer/Keep briefs must use 'Note: options differ in kind' (preamble),
but the answered-HOLD path demanded a Completeness score, rejected a
one-line Net with a semicolon, and read posture only from ELI10. The rerun's
brief applied HOLD SCOPE in its Recommendation reason. Revert the
ineffective 'always'/'handoff chat' wording: two runs still skipped the
mode handoff.
2026-09-29 17:10:33 +00:00
garrytan 41dca7609d test(design-floor): answer the seed-declared all-seven 0D focus menu while it is pending
The actor declares 'Design: review all seven dimensions', but its picker
reused designReviewSetupAUQ, which only matches already-answered calls
(and a narrower header/label set), so the pending D1 focus menu was never
answered and the case waited out its 609 s deadline. The skill's Step 0D
requires asking; the fixture now answers it.
2026-09-29 17:07:38 +00:00
garrytan 3a888896db test(ceo-section-loading): recognize an arrow-ordered stale-fill execution by event roles
The census review traced the seeded race as 'R1 miss -> R1 store read (v1)
-> W commit v2 -> W cache.delete -> W fulfills -> R1 cache.set(v1) -> R2
(begun after W) hits v1', but the in-flight gate only accepted race
vocabulary or fixed sentence shapes. Order, actor, version and dismissal
mutations still fail.
2026-09-29 17:05:15 +00:00
garrytan 66d49280a5 test(outside-plan-disabled): compare named record timestamps as instants; negated authorship is not a current claim
The repair rerun named the seeded record by its ISO second
(2026-09-29T16:58:52Z vs .727Z) and said 'I did not write'; both were
misread as a foreign timestamp and a current write.
2026-09-29 17:02:31 +00:00
garrytan 97f0eee33d Merge remote-tracking branch 'origin/capy/audit-fix-wave' into capy/fixwave-baseline-repairs 2026-09-29 16:57:56 +00:00
garrytan 2bf97ab077 test(outside-plan-disabled): bind quoted prior-record values by their sentence, not phrase order
The parent obeyed the off switch and twice named the seeded completed
record as pre-existing, once with the quotation after its owner and once
with slash separators; the order-specific stripper counted both as current
completion. Timestamp, location, current-claim and value-match controls
still reject.
2026-09-29 16:57:54 +00:00