mirror of
https://github.com/garrytan/gstack.git
synced 2026-10-04 18:36:54 +02:00
* test: delete test-infrastructure dead code (G) - exit-propagation drives the runner's real strict verdict (BunTestOutputClassifier + strictTestExitCode); delete the unused shardRunLooksTruncated predicate. - delete skill-coverage-matrix registry + its gate (nothing reads it; the floor already iterates skillCensus()). - delete touchfiles-facade export-parity tests (Bun fails missing imports at link time) and the duplicated E2E_TIERS tier-value test. - delete brain-cache-spec TRANSPORT_DEFAULT_POLICY, SKILL_RUN_RETENTION_DAYS and the now-unused BrainTrustPolicy type with their literal tests. AUTOPLAN_PREFLIGHT_BUDGET_BYTES stays: skill-preflight-budget enforces it against real resolver output. - delete audit-compliance's JSDoc-comment grep. * test: replace product tests that fake the product with real-boundary tests (F) - design: serve.test.ts drove an inline mirror server; now two tests run the real serve() on an ephemeral port (reload confinement, submit exit 0). - setup-gbrain: rollback + voyage tests execute the template-extracted init blocks (3 sites) instead of drifted local bash copies. - terminal-agent: internalHandler source greps replaced by a behavioral /internal/grant + /internal/revoke auth matrix (no/wrong/valid token). - /health: server-security-surface and the server-auth / security-audit-r2 / sidebar-tabs source greps fold into one liveness-only check on the real body; the L4 sidecar wiring gets a behavioral /pty-inject-scan test. - delete tautologies (browser-manager onDisconnect, memory-command #12), ios swiftui tap fixture self-check, memory-ingest put_page grep, detach source greps, sidebar-agent absence pins, dead-CSS pins + the dead CSS, security-audit-r2 Task 1 + the test-only meta-commands re-export, duplicate generated-SKILL.md checks. - make-pdf coverage-gaps cases move into their owner test files. * test: delete tests of dead eval code (A) - A1: the retired Eng lexical oracle (evaluateEngSeedCoverage, isEngSeedDecisionAUQ), the completion-handoff detector and the retained corpus had no paid caller since v1.87.6; delete their 26 replay files, ~2.6k helper LOC and fixtures, and the dead blocks in 8 mixed files (live hasNativePlanTerminal / batching assertions stay). - A2: dead viewport approvers in autoplan-artifact-permission and their 11 replay files + fixtures; recorder/launcher cases stay. - A3: never-wired oracles and seeders (autoplan-phase-order, eng-finding-fixture, ceo-paired-fixture, design-ui-scope, plan-skill-completion, pty-current-screen, required-reads, transcript-section-logger); plan-seed-submission now decodes through the production createPtyScreen; section manifests name their actual guard. - A4: zero-reference helper exports, plus execGit and invokeAndObserve found by the reachability pass. - 52 fixtures orphaned by the deletions; touchfile and selection-table entries for every deleted path. * test: clean up the paid eval lane (B1-B4, B6, B7) - B1: delete paid files that assert nothing or cannot pass meaningfully: skill-llm-eval-spec and skill-e2e-spec-execute (test.todo), gemini-e2e (+ gemini-session-runner; no gemini CLI in CI), ship-idempotency (red since v1.63), the two opus-4-7 *-sonnet overlay wrappers, conductor-prose (+ its source-evaluation replay), codex-e2e-plan-format; drop their keys, scripts and census rows. - B2: skill-llm-eval grades browse/sections/command-list.md with one union judge that also carries the baseline score pin; regression-vs-baseline deleted (paid run: pass, c4/c4/a4). - B3: memory-pipeline, ios-qa, ios-qa-swift-build and plan-tune-cathedral make no model calls; renamed out of the paid glob so they run on every PR. Swift builds need GSTACK_TEST_SWIFT=1; device stub deleted. - B4: codex-e2e*, outside-voice, aside and ios-device cannot run in the CI image; excluded from the weekly lane with a tracked re-entry condition. - B6: fold opus-47's negative routing controls into skill-routing-e2e journey-negatives (paid run: 3/3 unrouted) and delete the file. - B7: delete the never-green brain-privacy-gate eval; a free gstack-skill-start test now proves consent precedes artifacts egress. * test: retire the finding-count cluster and trim its helpers (C) - C0/C1: the five never-green evals (skill-e2e-autoplan-chain and skill-e2e-plan-{ceo,eng,design,devex}-finding-count) failed on harness and budget, never on skill behavior; delete them, their touchfile/tier ids, AUTOPLAN_CHAIN_BUDGET and the dedicated eighth periodic slice (--slices 7). - C2: delete the helper groups whose only paid consumers were those files (11 modules), trim claude-pty-runner and eng-seeded-coverage to the paid closure, and delete the free replay tests whose assertions exercised only that dead code (89 files, 135 orphaned fixtures). Blocks that used dead code only as input for a live subject keep their assertions: the multiSelect default moved to plan-review-decisions, runner PTY tests use inline caller policies, and the timer-safe budget checks moved to eng-finding-retry-budget. - The eight production-touching files stay except ceo-current-decision-record (its template read only feeds the retired counter). - CARVE_GUARDS.autoplan is behavioral 'none'; TODOS records the lost chain and per-finding cadence coverage with their re-entry tests. * test: fold per-incident replay series into their detector owners (D) Twelve detector families move into one owner test each: 73 incident files become describe blocks in ceo-section-loading-fixture (stale-fill race), model-overlays, coverage-audit-evidence, autoplan-phase-observer, native-auto-decide, outside-voice-evidence, eng-first-review, plan-count-completion, plan-count-file-permission, ceo-mode-option, plan-scope-selection and plan-count-prerequisite. Each block keeps its original code and fixture, so every case still runs; only tests asserting the incident file's own touchfile registration are dropped (41). Touchfile lists that named an incident now name its owner. * test: start the plan-count history PTY on its readiness marker (H) The fake CLI prints a startup marker and the runner waits for it instead of the fixed 8 s startup sleep (8.6 s -> 0.9 s locally). eng-semantic-terminal's sleeping registration cases went with C; plan-count-timeout keeps the fixed wait because it asserts deadline behavior. * test: derive paid touchfiles from each eval's static closure (E) touchfiles.test.ts now checks, per key, that the paid file's static test/helpers and test/fixtures closure (plus fixture paths it names in string literals) is covered, and names the file, path, chain and key to fix when it is not. Free *.test.ts files are no longer touchfiles, so editing a free replay test stops selecting paid evals: 950 entries removed, 653 real closure paths added. The hand-copied inventories go: periodic-fixture-selection, fake-impeccable-touchfiles and 45 per-file selection examples. Selection for the sample edits (plan-eng-review template, claude-pty-runner, plan-count-fixture, gstack-config) loses no case under either profile. CONTRIBUTING documents the rule and its lower bound. * test: skip hollow tier shards and census judges in the paid planner (B5) A paid file is now skipped for a tier lane only when every E2E id it registers is known statically and none has that tier; ids come from the touchfile registrations and literal testName/*IfSelected arguments, so a comment or skill path that quotes another id cannot unschedule it, and computed names keep today's scheduling. --list and the manifest show each skip as "skipped: no E2E_TIERS id has tier <tier>". The weekly gate census drops the LLM judges (--skip-judges); they still run in the periodic census and PR gate lanes. Gate lane 52 -> 42 files, census 41; periodic 77 -> 69. * test: run seven paid evals on the current default capture model (B8) skill-e2e-{auq-matrix,plan-format,qa-bugs,retro,workflow} pinned claude-opus-4-7 and skill-e2e-office-hours plus -brain-writeback pinned claude-sonnet-4-6; none tests a historical model, so they now capture with resolveEvalModel('capture'), and the free harness tests that execute these registrations receive the same resolver. The paid re-pin run passed all of them. skill-e2e-{design,office-hours-phase4,plan-prosons,plan} keep claude-opus-4-7: six of their cases failed on the default model (three timeouts, a missing report file, a format miss and a posture score of 3), so per the plan's fallback they keep their pins with a TODOS entry. The pre-spend estimate and drop threshold are in docs/test-audit-2026-09.md. * test: guard the reduced suite against new test-of-test files - test/test-of-test-ratchet.test.ts records the 228 free tests that import only test/ code and fails on a new one, naming the owner test to extend instead; a stale baseline entry fails with the remove instruction. - test/helpers/resolve-repo-path.ts is the one specifier/literal resolver for the ratchet and the touchfile closure invariant, with its own unit tests. - CONTRIBUTING "Test tiers" describes the paid-failure workflow (fix, then one row in the detector's owner test) and the ratchet; TEST_PORTFOLIO gains the detector -> owner-test table and no longer claims an Autoplan chain eval. - TODOS: automatic exclusion policy for chronically red periodic files (P3), the deferred native-completion table collapse, the unused CEO payment seeder; the PTY readiness item is narrowed to the paid runner. - docs/test-audit-2026-09.md collects the triage, security mapping, inventories, selection proof, behavior-commit decisions and retained false positives. * v1.91.8.0 test: smaller suite, derived paid selection, retired never-green evals Release metadata for the test-reduction branch: VERSION 1.91.8.0 (1.91.7.0 is claimed by #2983), CHANGELOG with the measured before/after table and a contributor section, durations re-recorded on Ubicloud standard-16 (857 files, 0 failures), the agents digest, CONTRIBUTING's after-measurement row, the B8 fallback TODOS entry, and the after metrics, kept-vs-plan notes, B8 run and census estimate in docs/test-audit-2026-09.md. * fix(ubicloud): skip retrieval globs that match nothing instead of reporting a failed pull * test: pin DISABLE_AUTOUPDATER in hermetic env and capture corrupt-seed warning Both EVALS_HERMETIC branches of buildHermeticEnv now carry DISABLE_AUTOUPDATER=1 (the allowlist scrubbed the workflow's copy, so every PTY screen showed the updater's npm-prefix failure). Per-test overrides still win. The corrupt durations-seed test now captures its expected warning and restores the console spy. * style(cso): format lib/cso TypeScript with pinned Prettier Mechanical reformat only. Minified transpile output is byte-identical for 21 of 22 files; witness.ts differs only in three regex flag orders (/mi -> /im), which JavaScript canonicalizes. Source-text assertions over lib/cso now compare whitespace-insensitively with the same tokens. * fix(cso): import join for compiled-launcher assertion witnesses Compiled installs always take the non-Bun branch, which called an unimported join and threw before any runtime-tested assertion could be witnessed. The child command selection is now a pure, platform-aware function; a missing sibling launcher fails with its expected path. * fix(browse): make connect --supervise actually respawn a crashed server The supervisor respawned with a block-scoped env that no longer existed, so every attempt threw and the loop gave up after five tries. The headed env is now one pure helper used by connect and respawn, the loop is an injectable runHeadedSupervisor with behavioral tests, failures name the daemon log and relaunch command, and connect's usage advertises --supervise. * test: one finite PR world for the shared-libs fixture; name dual-voice probe evidence The shared-libs shim served 2 PRs for pulls?state=all and endless full pages for state=open. gh pr list, pulls?state=open|all|closed (per_page/page, short last page, direction) and search/issues now page one deterministic table: PR 7, 600 older open PRs, PR 42 and 3 closed PRs, so five 100-item open-metadata pages still leave older open PRs unchecked. The Contents API lists pinned directories (the captured attempt got 404 for contents/ and contents/src while files resolved, then fell back to a raw host), unknown endpoints return 404 instead of repo metadata, and the read-only detector is unchanged. Free tests cover view agreement, the budget bound, gh/curl agreement and the empty world. Dual-voice outside-voice failures now report probeToolUseId, probeMode and the canonical-match result with the reason the probe output was rejected. * feat: require a zero-error product typecheck and a test type-debt ratchet Adds tsconfig.json (strict) over product code, fixes its remaining 90 diagnostics (type-only, interface corrections, and explicit narrowing), and adds a typecheck job to the required free-tests aggregate running bun run typecheck, the test-code ratchet (identity -> count baseline, fails on new, repeated, or unlocked fixed diagnostics), and the lib/cso format check. Reuses fixes from #2447 where they still applied. * test: follow the headed env helper and the typecheck gate in source-shape checks * fix(test): pin the package.json change kind in shared-input selection tests computePaidCaseSelection read the version-only exemption from git even when changed files were injected, so the shared-input test failed on main and on version-only branches. The exemption is now an optional input; the test pins a real package.json change and covers the version-only case. * test: judge plan-count completion on structured evidence, not wording Replaying run 36385945043's two Design attempts showed the existing routes rejected correct endings: attempt 1 at the typed-completion path field ('- Reviewed plan written to …' is not a 'Plan written to' line), attempt 2 at the leading-fence veto (its final message opens with the dashboard). nativePlanTerminalPreconditions is the structural prefix of hasNativePlanTerminal (behavior unchanged). structuredPlanCompletion adds, inside the existing nativeSummary branch: a complete report (Design binding for Design), a completed review-log row for the expected skill appended during this attempt under the child's GSTACK_HOME/project slug (resolved with bin/gstack-slug) and stamped with the fixture commit, timed between the report/last answer (second resolution) and the final native message, a final message with stop_reason end_turn (now carried on public transcript messages), and no visible question or permission prompt. Timeout summaries add idleFor and lastTerminalCandidate. Terminal and throw captures copy the plan file and review-log rows into the artifact directory; copies are best-effort and recorded in evidence-copy.json. Free regressions: both captured Design endings (trimmed fixture with provenance; report, row and end_turn reconstructed and labelled), the negative controls, and real-PTY completion/timeout runs through the real review logger. * test: structural Design count boundary; TODO proposals are not findings Replaying run 36385945043 through the Design count predicates: routing, focus and learnings setup was not recognized as setup, Issue 1 was counted pre-review in both attempts (the boundary fired on it), and attempt 2 counted the Font TODO proposal as a finding (review=4 and review=5 for five issues). The paid caller now starts review at the first answered native decision that is not setup (recognized packet, or setup header/question ID), a completion handoff, artifact rendering or a TODO proposal (the review's Add to TODOS.md / Skip / Build it now menu). TODO proposals are recorded as administrative extra decisions. The replay asserts each counted call: both attempts review=5 (Issues 1-5). isDesignCountFirstReview and its controls are unchanged. * test: CEO classifier throws name the question and matched predicates Replaying run 36385945043's FAN-1 and ERR-1 throws (ledger rows reconstructed from rendered diffs) through ceoPaymentFinding: the email obligation's row, subject, option and proposal predicates pass and the ELI10 explanation-defect predicate fails first ('lets that exception fly out', 'the error bubbles up'). Binding the defect to the named ledger row instead (the planned fix) was tried and reverted: scoped to the email seed it flips 30+ existing cf74 still-rejects replays, which require a vocabulary-free, ledger-bound email question to earn credit only through a complete saved comparison. With FAN-1's rendered currentDecision payload reconstructed, the recorded- decision path counts it, so the real saved plan (not uploaded) must have differed; failure artifacts now retain it. The classifier stays fail-closed and unchanged. Its throw now prints the header, the first 200 question characters and each obligation's predicate results. Free regressions with provenance and negative controls: an unrelated question, an email question whose row says it is already rescued, and a ledger ID whose row belongs to another seed. * chore: regenerate the test type-debt baseline on top of #2994 * fix(typecheck): strip the checkout root from ratchet diagnostic identities * fix(test): recognize ledger row-ID split candidates so collection stops at the last ACK Run 36385945043's split-overflow case asked all five candidate decisions by 8m55s, but the live candidate check required the question to open with "E1:" and every option to be a known disposition. The skill cited ledger row IDs ("D2.1 — R-E1: …") and offered "Hold, discuss first", so no candidate was recognized and the attempt ran the whole review (1302s). Identity now comes from the native header; the question must open with that candidate's ledger reference, name only that candidate, and offer exactly one include, defer and cut disposition. The selected answer must still be one of those three. The semantic evaluator and every existing negative control are unchanged; a trimmed capture from the run adds the positive case and four row-ID negative controls. * fix(test): stop the eng batching eval once its floor is proven The case's only verdict is reviewCount >= FLOOR (3). Run 36385945043 had three distinct acknowledged review decisions at 6m41s but kept answering until the ceiling (7) at 12m13s. The registration now passes the runner's existing isCollectionComplete stop once FLOOR non-setup, non-administrative review decisions are acknowledged; the floor check, ceiling, budget and counter are unchanged. A child-process registration test proves the stop predicate and that below-floor and timeout outcomes still fail. * test: add the non-blocking 'marathon' E2E tier Full start-to-finish flows move out of the blocking lanes. E2E_TIERS and E2ETier gain 'marathon'; describeE2ETier('marathon') is enabled only when EVALS_TIER=marathon, so the gate/PR and periodic lanes (and the gate census) never run those cases. The PR profile accepts marathon ids as scheduled elsewhere and defers them with their own reason, even on full fallback. * test: move the full office-hours workflow to marathon; add a periodic design-draft checkpoint The full startup workflow runs 1–3 real spec-review rounds (~280s each) and hit its 1200s capture in run 36385945043 at finalize. Review depth is the product's loop, so the case cannot fit a blocking lane without cutting rounds. It is now marathon tier with every assertion unchanged. skill-e2e-office-hours-design-draft.test.ts (periodic) runs the same fixed interview only through the Write that creates the design (269s in that run) and applies the full validator's design-draft checks, the required section reads and the launch/foreign-skill-read guards. validateOfficeHoursDesignDraft is extracted from validateOfficeHoursCompletion, which still applies it. Selection: office-hours-design-draft is registered periodic; the marathon-only file is already excluded from the gate and periodic plans by the B5 planner rule. Tier-alignment regexes and the valid-tier check accept 'marathon'. A type-only cast in plan-scope-selection.test.ts removes a diagnostic whose union print order made the ratchet identity unstable; baseline tightened. * test: supply the split-overflow fixture's HOLD SCOPE mode as a prerequisite The split actor always answered 0E's mode question with HOLD SCOPE. The skill skips that question on an explicit choice, so the fixture now states it and the attempt starts at the five candidate decisions (about 1.5 min earlier in run 36385945043). Candidates, actor policy, floor and semantic evaluation are unchanged; the fixture test pins the supplied choice. * test: start the eng batching eval with its setup prerequisites supplied Routing setup and cross-project learnings (D1/D2 in run 36385945043) are never counted and are not what the case measures. The registration now uses the runner's existing preconfiguredReviewActor so the attempt starts at the review; engSetupAUQ still vetoes any late setup question. The registration test pins the option. * test: count the design-draft paid file and defer marathon ids in PR selection pins The discovered paid-file census grows by one (skill-e2e-office-hours-design-draft). Full-fallback PR selection defers every non-gate id; the shared-input pins now expect periodic and marathon ids there. * fix(review): resolve the judged revalidation, setup-authority, plan-gate and findings-record ambiguities The census review workflow judge scored clarity/actionability 3 on both attempts: smoke-clock limits appeared to forbid post-repair revalidation, the caller deadline was undefined, 'ask for setup' conflicted with the report-only browser rule, fallback-sourced HIGH discrepancies had no gate decision, and the Step 5.8 record omitted adversarial findings. * fix(office-hours): load the builder section for every builder-mode reply Both census builder-wildness attempts answered a direct request for adjacent unlocks without reading phase-2b-builder-brainstorm.md, whose trigger read as applying only to the generative questions. * fix(sync-gbrain): define Step 4 helper args and one atomic write path Both census read-ready attempts spent turns reading the helper source to resolve <user-args>, inspecting fixture internals kept inside the repo, and reconciling 'Read + Edit' with the tmp+mv atomic write, then hit max turns before the verdict. * refactor(evals): share the import-closure walker and add the E2E shard reuse identity sourceDependencyClosure moves from the workflow-judge adapter into scripts/eval-input-cache.ts unchanged, so judge keys stay byte-identical. scripts/e2e-shard-reuse.ts builds the consumed-input identity of one PR-lane E2E shard (test import closure, every registered case's touchfiles, globals, runner/workflow/setup actions, child env pins, CI image, Claude CLI) and fails closed on anything unknown. Marathon joins the always-fresh purposes. * feat(evals): ~12-minute blocking paid lanes and a non-blocking marathon lane - Planner budget mode (--slice-budget S --jobs J): recorded per-tier wall times pack into as many ~9-minute executors as the work needs; the plan records per-slice estimates and the CI job timeout (supervised worst case + 20 min). evals.yml and evals-periodic.yml derive matrix size and timeout-minutes from it; max-parallel covers every slice at once. - Case shards: plan/design/review-army/shared-libs(-paths) run one registered case per process (<file>#<case id>, exact name pattern, exactly one case). - Retry rule: a timed-out attempt is a verdict. Only files whose every case budget is CAPTURE tier or shorter keep one retry; walls shrink to match. - Marathon tier: positive selection, excluded from gate/periodic planners, run by the new evals-marathon.yml (weekly + dispatch, fresh, own report). - PR-lane E2E reuse of verified first-attempt passes on identical inputs; the report rejects reuse outside the fast PR profile. - Duration seed from census run 36385945043, per tier and per case shard. * docs: blocking lane budget, marathon lane, retry policy and E2E reuse * chore(typecheck): lock in two fixed test diagnostics * fix(ci): drop a duplicated env/jobs block in evals-marathon.yml * test(ship-docsync): shard the doc-sync lifecycle by case and drop the duplicate dispatch-only case ship-docsync ran the same fixture and prompt as ship-docsync-completion and asserted a subset of it. The file now runs one case per process, so its lane wall is its longest case instead of half the sum of thirteen. * fix(evals): plan CI-unrunnable cases as excluded entries, not empty case shards design-review-fix drives the Aside browser and registers test.skip on Linux runners, so its case shard executed zero cases and failed the exact-one-case check in proof census 36597762183 (eval-slices 6). CASE_CI_EXCLUDE (reason + tracking, beside PERIODIC_CI_EXCLUDE) now turns such cases into excluded manifest entries that --list and the manifest surface; every planned case shard still must execute exactly its case. * docs(todos): list the case-level Aside exclusion with the CI-unrunnable evals * fix(plan-ceo-review): restore experience-first expansion framing, require the mode handoff, skip pacing menus Census 36597762183: both mode-routing runs logged provenance and moved on without the mandated handoff chat; the EXPANSION run asked an unauthorized batch/narrow pacing menu instead of the first per-addition question; the expansion-energy proposals led with the spec because v1.87.6.0 dropped 'lead with the felt experience'. The HOLD review detector also rejected a decision whose grounding line named no plan file although the owned source Read binds it. * test(outside-plan-disabled): bind quoted prior-record values by their sentence, not phrase order The parent obeyed the off switch and twice named the seeded completed record as pre-existing, once with the quotation after its owner and once with slash separators; the order-specific stripper counted both as current completion. Timestamp, location, current-claim and value-match controls still reject. * test(outside-plan-disabled): compare named record timestamps as instants; negated authorship is not a current claim The repair rerun named the seeded record by its ISO second (2026-09-29T16:58:52Z vs .727Z) and said 'I did not write'; both were misread as a foreign timestamp and a current write. * test(ceo-section-loading): recognize an arrow-ordered stale-fill execution by event roles The census review traced the seeded race as 'R1 miss -> R1 store read (v1) -> W commit v2 -> W cache.delete -> W fulfills -> R1 cache.set(v1) -> R2 (begun after W) hits v1', but the in-flight gate only accepted race vocabulary or fixed sentence shapes. Order, actor, version and dismissal mutations still fail. * test(design-floor): answer the seed-declared all-seven 0D focus menu while it is pending The actor declares 'Design: review all seven dimensions', but its picker reused designReviewSetupAUQ, which only matches already-answered calls (and a narrower header/label set), so the pending D1 focus menu was never answered and the case waited out its 609 s deadline. The skill's Step 0D requires asking; the fixture now answers it. * test(ceo-mode-routing): accept the skill-mandated Note form and Recommendation reason as HOLD posture HOLD Defer/Keep briefs must use 'Note: options differ in kind' (preamble), but the answered-HOLD path demanded a Completeness score, rejected a one-line Net with a semicolon, and read posture only from ELI10. The rerun's brief applied HOLD SCOPE in its Recommendation reason. Revert the ineffective 'always'/'handoff chat' wording: two runs still skipped the mode handoff. * test(qa-bugs): keep claude-opus-4-7 after qa-b6-static stalled on the default model qa-b6-static timed out on claude-fable-5-1 in census 36597762183 and in one of two targeted reruns. Both times the stream stopped mid-message with no pending tool, right after the model found the disabled submit button, and stayed silent until the 300 s deadline. Per the B8 fallback, re-pin with a TODOS entry; budgets and retries are unchanged. A rerun on opus-4-7 passed (125 s, 5/5 detected). * test(evals): add E2E_KINDS, BEHAVIOR_WHY, EVAL_POLICY and CASE_QUARANTINE skeletons Every E2E_TIERS and LLM_JUDGE_TOUCHFILES key starts as 'rule'; BEHAVIOR_WHY and CASE_QUARANTINE start empty. EVAL_POLICY pre-registers the approved panel (3, majority 2), quarantine entry 0.95/10 and exit 0.97/10, 10% cap, 8-weekly-run expiry, Fisher drift alarm and one INFRA re-dispatch. * test(evals): add trial records, panelVerdict, expectContract and trial-outcomes JSONL EvalTestEntry gains case_id, kind, trial, panel, failure_class and policy_version, stamped from the runner's TRIAL_ENV on isolated trial shards. panelVerdict() is the single verdict function (INCOMPLETE on missing or duplicate trials, contract veto at any count, quarantine hard-break rule, INFRA/INCOMPLETE machine classification). expectContract() records failure_class 'contract' on the collector entry and a sidecar before throwing. trial-outcomes JSONL has a fail-closed writer and a data-only reader. * test(evals): pin the fail-closed rule-shard gate through the real --report path Synthetic slice artifacts for rule fail, timeout, missing slice, unreported entry, hollow, never-started, collector failure and wrong-slice reports all exit red before the panel-verdict gate change lands. * test(evals): retire every paid automatic retry Paid evals never retry (approved 2026-09-29): delete SHORT_CASE_RETRY_FILES and retriesWithinCaseCap, drop the retry fields from the registered wall rows (walls now cover one run plus reserve), make retriesForFiles return 0, pass --retry 0 explicitly, and drop --retry 1 from the package.json paid scripts. Add the eval:pass-rates alias. Tests that pinned the old retry allowance are updated as a policy change; review-finalization-budget now proves late-result recording under the production zero-retry arguments. * test(llm-judge): sample every judge as a pre-registered 3-sample panel Each of the 24 skill-llm-eval judges now draws EVAL_POLICY.judge.samples independent samples of the same prompt concurrently inside the unchanged JUDGE_MS budget. Numeric dimensions gate on the per-dimension panel mean against the unchanged threshold; booleans (would_browse, consistent) on a strict majority. An erroring sample fails the whole panel and is never resampled; a refusal is an unscored panel only when every sample refused. callJudge's 429 backoff stays: it is transport before any model output. The workflow-judge cache stores and validates only complete panels, and its identity now records the panel and zero file retries. Harness tests that pinned one provider call per case now pin the panel size. * test(evals): classify every live case and re-select a case when its kind changes E2E_KINDS: rule by default (191 E2E ids), 22 behavior cases whose verdict is a live model choice with an acceptable sub-100% per-trial rate, each with a BEHAVIOR_WHY tolerance, and 25 judge entries (the 24 workflow judges plus the fixed-fixture llm-judge-recommendation rubric check). Contract-shaped cases (ask-before-decide, plan-mode no-writes, mandated steps, secrets, the batching floor) stay rule. Behavior requires a known literal registration and an exact Bun test name so the case runs as its own trial shard. Map-diff selection now diffs E2E_KINDS and BEHAVIOR_WHY per key, and a base revision without them selects every key, so a kind flip runs the panel it introduces. test/eval-kinds.test.ts enforces coverage, tolerances, isolatability and the reviewed counts, printing the literal to add. * feat(evals): per-case pass rates with Wilson intervals, identity series and quarantine policy scripts/eval-flake-rank.ts becomes eval:pass-rates (eval:flake-rank stays an alias, and the legacy aggregate stays exported). It reads eval-store's trial-outcomes JSONL from the last N completed evals-periodic runs on this branch and main (gh, downloading only the trial-outcomes artifact, cached and size-capped, parsed as data), plus local eval dirs, and prints per-case per-trial pass rates with 95% Wilson intervals. A series is a case's own touchfiles minus GLOBAL_TOUCHFILES (caseSeriesIdentities, for the report job to stamp), per model, CLI version and policy version. Labels: INCONCLUSIVE, BROKEN, FLAKY, FAILING, PASSING. --backfill imports legacy slice artifacts as pre-policy trials (first attempt only, attributed by registry id, never guessed) for display only. --gate fails with ACTION REQUIRED on post-policy evidence only: drift below the quarantine entry rule, a rule case behaving like behavior, a one-sided Fisher drop against the previous identity (Holm-controlled), and quarantine entries that met their exit rule, expired after 8 weekly runs, broke the 10% tier cap or are invalid. CASE_QUARANTINE entries now carry a failureClass (detector, harness or model-latency); a product defect has no class and is never quarantined. The policy test pins EVAL_POLICY's approved constants. * feat(eval-pass-rates): attribute legacy records by the exact slug of their display name * ci(image): pin Claude Code 2.1.284 so the eval model is recognized 2.1.251 logs [claude-code:unrecognized_model] for claude-fable-5-1, the eval capture/judge default. 2.1.284 does not. The gate PTY smoke subset (plan-ceo/plan-devex plan-mode, plan-mode-no-op) parses on the new TUI; plan-design-review-plan-mode passed at 293 s on 2.1.284 and timed out at 300 s on 2.1.251 on the same tree. * test(eng-batching): grade the floor once the review report is complete A completed GSTACK REVIEW REPORT ends the review, so the review-question count is final there. Run 36606688266 wrote its report at 1,248 s and closed the session at 1,318 s; the case now stops collection and applies the unchanged floor at the report instead of waiting out the session. No budget changes. * test(eng-batching): bind unsourced native briefs through the report's target Run 36606688266 asked ten separate native review questions (D1-D9 bound to ledger records R1-R9) and failed reviewCount=0 < FLOOR=3: its briefs named the plan by title instead of citing PLAN.md, its report declared 'Review target (fixed): PLAN.md' under '# Engineering review: <plan>', and it kept an unfenced copy of the plan's own H1. The named-source route now accepts those spellings and non-inline ledger briefs. The same replay rejects a foreign, mixed, duplicate or missing target, another plan's title or copied H1, a brief naming another plan or file, a mismatched saved brief, and re-asks. The run-36597762183 capture still counts 3. * fix(plan-design-review): treat a designer with no API key as unavailable Both proof runs (36597762183, 36606688266) printed DESIGN_READY, hit 'No OpenAI API key found' on the first $D variants call, then hand-built HTML/CSS wireframes, screenshots and a comparison board for ~195-245 s before the first review question; the second run timed out at 600 s. A failed first generation now takes the existing text-only path, and the skill forbids substituting hand-built mockups. * fix(deslop-shared-libs): read related sources together within the turn limit Run 36606688266's opportunity audit read sixteen sources one per turn and stopped at error_max_turns; the passing run 36597762183 read the same files in three batched commands. The skill now says turns are bounded and asks for parallel reads or one read-only command per step. * test(ceo-mode-routing): submit a mode review that scrolled past the viewport Run 36606688266 bundled routing, learnings and the mode choice into one native call. Its review panel was taller than the terminal, so the tab bar scrolled off, ceoModeSubmissionInput returned null for 240 s and HOLD SCOPE was never submitted ('no posture match'). With no bar on screen the viewport must still end at the focused Submit prompt, and the accumulated screen text supplies the one complete panel, authenticated exactly as before. Replay controls reject another mode, an unoffered answer, an altered question, a quoted panel, trailing output, a moved cursor and an answered or changed call. * docs(evals): document the pre-registered verdict policy, quarantine, pass-rate history and arithmetic AGENTS.md replaces the retry rule with the approved policy text (no retries; kind fixes trials; no added trials, samples or dispatches after a result; quarantine by CASE_QUARANTINE only; one INFRA/INCOMPLETE re-dispatch) and notes that a pre-registered fixed panel is not rejudging. CONTRIBUTING gains the kind rules, the judge panel, eval:pass-rates and an 'Add a paid eval' checklist. TESTING_INTERNALS describes verdicts, quarantine, history and the arithmetic, including the rule term: 1 trial vs 2-of-3 red rates at p = 0.99/0.95/0.90/0.70/0.30 and lane all-green probabilities for the current 191 rule / 22 behavior / 25 judge registry. * feat(evals): trial planner, slice exit split and panel-verdict report Planner: behavior and quarantined cases become panels of isolated trial shards (<file>#<id>~t<N>) bound by EVALS_SELECTION_JSON=[id] and the exact test name; the file shard excludes them by name. Trials of one case never share a slice, result slugs are unique, panels are validated whole, unknown registrations throw, and the planner prints a capacity preflight. Executor: each trial shard gets its TRIAL_ENV identity and a trial record (outcome, failure class, cause, cost); every shard writes a JUnit report. The slice exit now means execution completeness: a failed rule shard or a trial without a record reds the runner, a failed trial does not. Report: panelVerdict() decides every panel of the first run attempt (later attempts are reported, never replacing it); rule shards keep the unchanged fail-closed checks; collector records all count (no last-attempt wins); census runs enforce the quarantine cap and expiry. It writes collector-outcomes v2, trial-outcomes.jsonl (trials plus JUnit rule/judge cases), report-summary.md, and one headline + failure block with rerun commands, and flags INFRA/INCOMPLETE-only reds for the one re-dispatch. The fail-open suite gains the panel cases: behavior 1/3 red, 2/3 green with its failed trial shown, missing trial INCOMPLETE, contract at 2/3 red, quarantined 1/3 green, 0/3 and contract red, missing slice red, and a later attempt never replacing the first. * chore(evals): refresh paid duration seeds from proof runs 36597762183 and 36606688266 Both tiers, merged in run order (the later run wins). Notable: split-overflow 1332s -> 504s, section-loading 604s -> 342s, mode-routing 575s -> 444s; multi-finding-batching 734s -> 1318s (its red path in run 36606688266). * feat(evals): stamp trial series identities and fit panels to the live registry - scripts/eval-trial-series.ts stamps series_identity (eval-flake-rank's caseSeriesIdentities) on a report's trial-outcomes JSONL as its own step, keeping the history tool out of the paid runner's closure; TrialOutcomeRecord gains the optional series_identity field. - Slice-count plans let a registered trial spill into an ordinary lane when its siblings hold every long lane, so panels never share a runner. - Re-audited test-selection.ts (Stream B added the E2E_KINDS/BEHAVIOR_WHY map-diff; no new module loading) and repinned its hash. - Detach and release floors now count trial shards (66 periodic trials in 22 panels): periodic floor 33,821s, still under eval:bg:periodic's 67,380s. - Coordination fixtures supply the executor's trial records. * ci(evals): attempt-scoped artifacts, verdict-v2 PR comment, weekly pass-rate gate and one INFRA re-dispatch - Slice, census and marathon artifacts carry -a<run_attempt>; reports download them per artifact (no merge), so records never overwrite and a re-run never replaces the first attempt's verdict. - Planners pass --max-parallel for the capacity preflight (24/16 unchanged: the refreshed periodic plan needs 24 slices, the gate census 12). - PR comment: jq-only job reads collector-outcomes v2 (headline, sanitized failure block); the group_by(.name)|last recomputation is gone. - Reports stamp series identities, upload trial-outcomes-* for history, and shard logs upload always (a failed trial no longer reds its runner). - Weekly report: headline + failure block of both lanes in the issue body, the eval:pass-rates --gate step (fails closed without history), close the issue on a green run, and UC-E1: when every red is machine-classified INFRA/INCOMPLETE, one re-dispatch as a new run in its own concurrency group (redispatch_of), both runs reported. * feat(evals): planner-side whole-panel reuse and negative receipts The planner job restores this PR's receipt store once and ships a single filtered set with the plan: a pass or panel receipt with a same-or-newer FAIL for its input identity is dropped, and a panel receipt ships only as a whole PASS panel (re-verified with panelVerdict) from one run. Executors read only that set (no per-slice cache restore or save), so every trial of a panel sees the same receipts; a trial reuses its own record from the panel receipt, keeping a split PASS's failed trial. Trial identities drop the trial index (run-scoped) and bind the panel policy. Executed shards carry their input identity; the report turns a whole fresh PASS panel into a panel receipt and a FAIL panel or failed rule shard into a negative receipt, and marks a panel that mixes reused and fresh trials INCOMPLETE. The report job merges plan, slice and report receipts (newest per file) and saves one store per run. Also fixes two TS2352 casts in browse/test/dia-macos-qualification.test.ts whose diagnostic text drifted with program order (baseline locked, fix only). * feat(evals): --case/--trials local diagnosis and panels in local sharded runs bun run scripts/test-paid-shards.ts --case <id> [--trials N] runs N independent trials of one case through the CI panel runner (trial shards, TRIAL_ENV identity, name-pattern isolation) and prints its panelVerdict(); N defaults to the case's policy panel and CI never reads it. The local sharded path (test:gate:sharded, test:periodic:sharded) now plans the same trial shards and exclusions as CI and exits on execution completeness plus panel verdicts. * test(pty): grant an owned Create pane whose title row is cropped The targeted batching rerun on Claude Code 2.1.284 left its first report Write unanswered for 1,372 s and timed out: the viewport began at the pane's relative file row and rule, with the 'Create file' title cropped above, so the preview parser rejected the file row as foreign. That row must now resolve to the owned path and is skipped before the unchanged line-by-line preview match. Replay controls reject another file, another directory and an edited preview row. * fix(evals): tsx-safe generics in eval-flake-rank, legacy artifact names, no-retry wall docs * test(evals): record the read-only and detector-row invariants as contracts shared-libs-opportunity-judgment and review-design-lite are behavior cases: their recommendation and checklist judgments may vary, but the read-only invariant (commands, provider requests, fixture bytes, hooks, state) and the deterministic fake-engine detector rows are contracts. Both now go through expectContract, so any failure vetoes the panel. * test(judges): sample the recommendation rubric as a panel; never re-ask armJudge llm-judge-recommendation is a judge case: each fixture now draws a 3-sample judgePanel, gates reason_substance on the panel mean and the present/commits/has_because checks on a 2-of-3 majority, thresholds unchanged. armJudge no longer re-asks on a malformed verdict; it is a failed sample, as the judge policy requires. * test(evals): record a pre-turn API or CLI failure as infra recordE2E sets failure_class 'infra' on a failed session whose runner reports error_api, timeout_startup, error_output_stream or a non-zero CLI exit with zero turns and no assistant event. A model refusal, a timeout after model work, max turns, or an explicit caller pass/class keeps its ordinary classification. * test: pin every-record outcome counts and the twelve doc-sync callbacks * test(eng-batching): read the report target as a field, not a spelling The next targeted rerun (Claude Code 2.1.284) again asked eleven separate native questions and again counted zero: its briefs named no plan and its report declared '- **Review target (fixed):** `/abs/PLAN.md`' under '# Eng Review — PLAN.md: <plan>'. An unsourced brief now inherits the one current target field that names a PLAN.md file, whatever its list or emphasis markup; its ledger record still supplies the cited finding and must reproduce the brief exactly. A brief that names its plan must still match the report title. Replays of all three captures count 9, 9 and 3; controls reject a foreign, duplicate or missing target and an archived title. * fix(evals): --case list mode and name precheck; case-shard qa-callers; refresh batching and design-with-ui seeds * chore(release): v1.91.9.0 * test: settle the post-response composer before seeding; give the TPA recorder adapter its infra helper submitPlanSeed accepted a stale empty composer when the transcript recorded end_turn before the CLI repainted (late-repaint-typed-current fails 5/5 on the old helper, passes 5/5 now). The TPA recording fixture extracted recordE2E without isPreTurnInfraFailure, so every failed case threw before recording. * test(autoplan-dual-voice): unwrap Claude Code 2.1.284 subagent hand-back frames; accept read-only probe diagnostics; record before asserting Census run 36626737820: the native CEO report arrived framed and indented, so its INPUT line never matched, and the model's exact probe plus two variable echoes was not canonical. A column-zero line inside a frame, command substitution, backticks, redirects, assignments, CODEX_MODE echoes and output line-count mismatches stay rejected. The failure now records before asserting. * ci(image): keep Claude Code 2.1.251; test(ceo-mode-routing): keep HOLD's own deferrals in scope before assessing its rigor decision 2.1.284 enables per-turn effort for claude-fable-5-1: in gate census 36626737820, 66 of 84 sessions ran longer than on 2.1.251 (+20% session time, +32% thinking tokens) and 11 cases timed out on unchanged budgets. HOLD SCOPE's 0G step asks its own defer/keep menu; the actor answered it Defer and the assessment then judged that scope question as the rigor decision. The actor now answers that menu Keep and assesses the next one. * test: attribute quoted prior-record field lists, state the judge reason bound in its schema, move split-overflow to marathon Census 36629958451 reds: - outside-plan-disabled-no-fallback: the model quoted the pre-existing record as a parenthesized field list with its exact timestamp; attribution now requires that exact timestamp and the record's own field values. - plan-devex-peer-comparison-classification: the judge correctly returned missing but wrote a 1069-character reason, voiding the judgment; structured outputs cannot enforce maxLength, so the bound is stated on the field. - plan-ceo-split-overflow ran 504-1188 s as one PTY flow and set the periodic lane's wall clock; it now runs weekly in the marathon lane. * test: supply holdDeferKeepIndex to the CEO routing mocks and follow split-overflow into the marathon lane The registered-callback fixtures mock ceo-mode-option and lacked the new export; the split fixtures asserted the periodic tier; the registered-budget check looked for split-overflow only in the periodic manifest. * fix(qa): checkpoint receipts print the report link for their exploration file qa-functional-webhook-report failed in two of three censuses because the report linked .qa-evidence/NNN capture folders as "checkpoints" and never linked exploration-NNN.json. The checkpoint receipt now prints link: [checkpoint NNN](exploration-NNN.json), and the functional report template says capture folders are not checkpoints. * docs: final census numbers in the v1.91.9.0 entry; file the paid-eval follow-ups * ci(evals): name the PR-comment loop's unused fields so shellcheck passes (SC2034) * fix(plan-ceo-review): tighten expansion pacing wording to fit the skeleton cap after the main merge The merged skeleton measured 80,166 bytes against its unchanged 80,150 cap. Same instructions: ask separately for each addition, in turn, with no pacing menu; lead each proposal with the felt experience, then shape, effort and impact. * fix(eval-pass-rates): match trial-outcome files by basename so Windows backslash paths are read * fix(evals): repair proof-run reds in design-consultation, document-release, design and QA fixtures - design-consultation Phase 1 asks one brief that confirms context and decides research; the confirm-only first question scored substance 2. - document-release defines ship-owned inputs, exact steps and the JSON result, and drops stale spawned-from-/ship text (judge actionability 3.67 -> 4/4/4). - plan-design-with-ui accepts the Step 0D focus menu the same way the shared picker does ("focus on specific ones?"). - plan-design-review plan-mode saves in three Edits instead of one final Write. - QA functional annotations ask for the full 40-character revision. - Outside-disabled attribution judges quoted prior-record data by its exact timestamp or a dated, pre-existing-record sentence; four captured phrasings replay clean and current claims still fail. - --case can select autoplan-dual-voice by its literal test name. * test(design): revert the three-Edit plan-mode flow A focused paid run still timed out at 300 s: the first three passes alone took 150 s of thinking. The case stays a named timeout red rather than cutting review depth. * test: accept 'review mode = X' auto-decide declarations and parenthetical scope exclusions in the shared-libs actor auto-decide-preserved: the product auto-decided HOLD SCOPE and said "Decision: review mode = HOLD SCOPE"; the grammar knew only "is" and ":". shared-libs-plan-callers: the recommended option said "(no hardening)" and the actor read "hardening" as an expansion. Both replay the captured text, keep negative controls, and passed focused paid runs. * fix(review): pass Review Army checklists by path, run research alongside dispatch, always probe the design detector; state review-log invocation and statuses in the caller fixture - review-army-perf-n-plus-one: the parent copied full checklists into agent prompts and ran web research before dispatch (290 s on a 12-line diff); 212 s now. - review-design-lite: 5 of 6 captured trials reported the detector absent without probing; the probe is mandatory and its first line is reported, and the contract credits only fake-engine rule ids the checklist never names. - review-exploratory-small-cli: the fixture never gave review-log's direct invocation or status vocabulary; the model ran it through bun and wrote status "blocked". The prompt states both and the validator rejects out-of-vocabulary review statuses. Each case passed a focused paid run after repair. * docs(changelog): proof-run product fixes * fix(ship): always run the design-lite detector probe; test(shared-libs): credit a failed first file view and deferred-reuse Skip wording - /ship design-lite: the probe is mandatory and any non-ready first line is stated, matching /review (5 of 6 captured /review trials had skipped it). - shared-libs-pr-coverage: the first PR 42 page-1 read printed only a jq error, so the one refetch is a legitimate recovery, charged to the same budget. - shared-libs-review-prior-coverage: the Skip option said a future review can "reuse it once snapshot coverage holds"; a conditional tail on the recorded decision is not product work. Captured-text regressions and negative controls. * fix(ship,qa,document-release): repair proof-run regressions and fixture gaps - ship-docsync-completion: yesterday's audit-scope result dropped the section's status, so /ship spliced one in; the section now opens with **Status:**. - ship-docsync-missing-asset: a missing section or old Ship-owned mode blocks before launch. - ship-docsync-late-result: the invocation record says prepare already saves the candidate selection (no extra Read; budget unchanged). - qa exploratory: await the method Reads before the first probe. - qa-callers fixture: quote the real review-log record template; allow the git log command plan-completion prescribes. - qa functional observer: a receipt caught mid-link(2) is checked at stop instead of failing with ENOENT (reproduced from CI). Each repaired case passed a focused paid run. * ci(image): pin Claude Code 2.1.284, the version users run Request-body capture shows both 2.1.251 and 2.1.284 send effort "high" to claude-fable-5-1; 2.1.284 adds the model's own profile. The slower 2.1.284 census was mostly API latency: its SDK-only judges were 25% slower too. Nine previously slow cases pass on 2.1.284 within unchanged budgets. * test: one owner per case id, a structural devex 0B setup rule, and correct design/gbrain actors - plan-design-review-plan-mode was registered by two files; the PTY smoke is now plan-design-review-plan-mode-smoke, and a registry test requires one owner per case in case-sharded files. - plan-devex-finding-floor: the template's 0B narrative-confirmation question is classified as setup structurally instead of timing out a Haiku assessor. - setup-gbrain-remote: the actor accepted 'skip' on the MCP-registration question the test asserts; it now accepts that question and declines others. - design-review-plugin-handoff: the fake engine cited a file absent from the fixture repo and index.html linked a missing styles.css. Captured-question regressions with negative controls; each case passed a focused paid run. * test: PTY harness handles clipped reviews and bundled setup tabs; AUQ judge uses structured output; design-consultation carve declines optional outside voices - ceo mode routing: a Submit review taller than the viewport, a setup tab bundled after the mode tab, and a clip through the mode question each hung or misread the run; the native answer is still verified after Submit. - judgeRecommendation requests a 1-5 enum schema; a malformed Haiku reply had scored substance 0 for a 4/5 brief. Judge failures now propagate. - carve section-loading for design-consultation declines the optional outside voices (a supported path) and treats DESIGN.md as the report; timeout unchanged. The Step 0E handoff defect is not fixed (0/15 samples across four wordings, none shipped) and is filed in TODOS. * test: fold the design-consultation completion replay into carve-section-sharding (test-of-test ratchet) * docs(todos): record the pre-push hook shard-order hang * test(qa-callers): disable git auto maintenance in the fixture repo (same guard as shared-libs; from #3002) * test(office-hours-attempt): the fake judge SDK response carries stop_reason like the real API (structured judge requires end_turn) * fix(qa): the caller STOP line says to await the method Reads before any probe ship-exploratory-plan-checks: the model read exploratory.md and sent a capture in the same response, before seeing the section's own await rule. * fix(qa): number the qa value-bar questions from 1 and say reproduced bugs already answer the first two * fix(qa): define evidence.json where it is built, point the preparation gate at the next section, name measured command durations in the report template Recurring qa/qa-only workflow-judge complaints in CI (clarity/actionability 3.33). * fix(plan-eng-review,review): a disallowed question tool is not headless; report kept tests only when some were skipped * fix(plan-eng-review): keep the headless-rule contract phrases adjacent * fix(evals): cut path variance at its measured sources - gstack-qa-evidence capture prints startedAt/completedAt/durationMs and, for --deadline captures, remainingMs; the functional report takes durations from them. The section clock notice asks for one clock read up front instead of one after every checkpoint (QA runs spent 7-14% of tool calls on date -u). - ship plan-completion: skip the audit dispatch when discovery already found no plan (the dispatch-vs-skip conflict produced an optional 60-100 s subagent). - materialize/checkpoint validation errors state the expected schema, so a rejected annotations file is fixable in one call instead of blocking the phase. - session-runner counts turns from the transcript when a run times out, so timeouts stop reporting 'turn 0'. * fix(evals): count timeout turns only from object transcript events * test(qa-callers): deterministic child transport, completion-time handoff reads, compact phase report The exploratory caller cases exist to prove the caller starts and bounds exploratory QA. Their native adversarial reviewer (review) and plan audit (ship plan-checks) now come from recorded child outputs instead of a live subagent, handoff freshness reads are required before completion records rather than every bookkeeping log, and the phase report is compact. Measured: 194-257 s per case against 208-284 s before, no subagent calls. * test(ship-docsync): seed fault cases at their gate instead of replaying attempt 1 The post-dispatch fault cases (missing-marker, launch-failure, timeout-unsettled, late-result, stale-before, stale-after, recovery) now start from a fixture-owned attempt 1: the real actor prepares and dispatches it, its verbatim output is saved once, and the invocation journal carries its pre-dispatch entry with the child asset hashes. The model resumes at Parent processing with a trimmed read list, inspect named as the authoritative repository observation, and recovery's intermediate checkpoint folded into the next attempt's pre-dispatch entry. Assertions count only parent-issued transport events and require a read of the saved attempt-1 output; missing-asset and the legacy failure case keep the full model-driven first attempt, and their prompts are byte-identical. * test(ship-docsync): name the seeded read list and cap journal/report length The first seeded stale-before run spent calls locating documentation.md (two ls sweeps), reading through cat and re-Reading the record before Edit, and ~40 s composing 1.5-2.2 KB entries and report. Name every seeded read path, ask for native Read, and bound entry/report length. * test(ship-docsync): trim the seeded parent's measured model time Measured on the seeded runs: one read the 78 KB ship/SKILL.md, the post-child freshness comparison spent 18-32 s of thinking over full inspect contents, and the final response restated the report (~1.1 KB). Say the phase excerpt stands in for ship/SKILL.md, compare hashes first and read content only for changed paths, and end with one status line. * feat(qa-evidence): enforce the checkpoint sequence and fill report bookkeeping in code - capture refuses to run another probe until a checkpoint anchored on the latest complete capture names this capture as its next command, and every complete capture prints that requirement. - materialize fills revision, runtime, cwd and learning (checkpoints whose next native command differs) when omitted and prints the reportLinks the report must include; the QA section shrinks accordingly. * test(qa-callers): hand the caller phase its invocation-start observations and review token; fix(next-version): fetch without auto maintenance - Every caller case receives the diff, status, log, untracked list, HEAD and an already-captured review start token, so the phase spends its budget on the contract under test instead of re-running setup reads. - gstack-next-version's fetches pass --no-auto-maintenance. On git 2.55 a completed fetch forks detached maintenance in the caller's repository; the free suite's live smoke test ran it inside the CI checkout, and every shard-12 pre-push hook hang so far followed a completed smoke fetch. * feat(deslop-shared-libs): route every Git read through bin/gstack-safe-git The skill made the model retype a long safe-Git prefix on each call and a dropped flag failed shared-libs-read-only. bin/gstack-safe-git applies the fixed env + flag prefix, adds --no-ext-diff --no-textconv to log/show/diff, allows diff only between two explicit object IDs and ls-files only in the NUL-delimited overlay form, and refuses every other shape with one line naming the allowed forms. The template now points at the installed helper (host global runtime via {{SAFE_GIT}}) and drops the prose it enforces. Fixtures resolve the helper to this checkout, the git shim records the safety environment, and isGuardedGitRequest requires the complete prefix (env included) for every repository read. * test(shared-libs): tee to a discard device is not a file write Paid shared-libs-opportunity-judgment t1 on 1213b01 failed read-only on '... | tee /dev/null | sha256sum'. The detector flagged any tee operand while the same devices are allowed for redirection. tee now fails only when an operand is a real file; tee to a file, -a file and -- -a stay violations. * fix(qa-evidence,observer): reject placeholder metadata and replay-only learning; declare the docs atomic-write target - materialize measures revision, runtime and cwd itself and rejects supplied values that differ (CI run wrote revision "HEAD" and runtime "bun"), and refuses learning checkpoints that replay the same probe, naming the fix. - The docs write observer treats Claude Code's atomic temp for the authorized doc target as transient, so a temp renamed before its per-file watch no longer marks the observation incomplete (ship-docsync-completion flake). Per-file monitoring outside declared targets stays fail-closed. * test(qa-functional): fix mode requires only the happy scenario from the model (carried byte-identical from #3002 183b01f4..3e6074b4) verifyQANativeRegression already reruns all eight webhook scenarios on the repaired source, so the model-side eight-scenario requirement in fix mode duplicated harness coverage and pushed qa-functional-webhook-fix past its budget. qa-only still requires every scenario. * fix(deslop-shared-libs): probe the audited repository with -C <repo> A CI run probed safe-git from the session directory above the target repo, so the capability probe never touched the repository and the run fell back to the API without a local attempt. The probe (and any call from elsewhere) now names the audited repository. * test(qa-deadline): never attach a reader to the full-pipe fixture's stdout The full-pipe receipt test attached a 'data' listener (flowing mode) and then paused; on CI the reader could drain the 2 MB write before the pause, so the receipt write never blocked and the helper exited 0 in ~126 ms. The stdout pipe now stays unread until the assertion, which is what the test means to model. * feat(qa): helpers answer --help, and the QA eval interfaces declare it Approved by Garry: asking gstack-qa-evidence or gstack-qa-deadline for usage is read-only, so both helpers print usage and exit 0 on --help (the evidence usage now names the annotation shape), and the functional and caller command allowlists accept exactly 'bun <path>/bin/gstack-qa-{evidence,deadline} --help'. Two CI runs failed only on that call. * fix(qa): after an input change, a probe is affected unless shown otherwise CI late-input run finished in time but revalidated only the happy probe after the locale input changed and reported the stale adverse probe green. The revalidation step now treats any probe not shown to be unaffected as affected. * test(shared-libs): seed the lifecycle replay's first Step 3 pass instead of replaying it shared-libs-review-lifecycle ran ~88% of its 300 s session budget (12-run census median 265 s, 4/24 sessions timed out). The fixture now executes pass 1's Step 3 once with the real logger and Git: a real unused REVIEW_START, then the diff, inventories, attributes/config/index flags, gstack-review-read output and every file's bytes and sha256, saved to one observation. The model resumes at Step 4 with an exact four-file first read, the observation named as the authoritative pass-1 repository read, one post-fix verification, an explicit pass-2 read list and a twelve-line summary. Pass 2 still runs its own --start, diff, reads, fingerprint and stage actor before --finish. The actor scope now states that a current settled final-pass actor result supplies the replaced QA/adversarial prerequisites and that the no-credit disclosure is a reporting label: one r1 session persisted completed:false from that ambiguity. New assertions: the final binding never uses the seeded token's start or tree, and the observation was read; free controls finish the seeded token (binding changed) and omit the observation read, and both fail. * test(shared-libs): trim the resumed review replays' setup and report Every sibling review session (revalidation, path-eligibility, index-flags, prior-coverage) loaded qa/sections/exploratory.md and often scope.md although its QA and native adversarial results are supplied synthetic inputs, then spent a second request on shared-code-reuse.md and base metadata. The resumed scope now states that the supplied results replace Step 4's QA method loading; the revalidation contract names one first response (workflow, checklist, finding, prerequisites, shared-code-reuse.md, base metadata) and caps the summary at twelve lines. Receipt order, direct source reads, the checker, the question and final persistence are unchanged. * fix(review): define what a Step 5c Skip option says Step 5c named "B) Skip" without saying what its description may claim. Two CI captures (path-eligibility on131d43be, index-flags on4643cb85) offered a Skip whose description added effects beyond declining: "The extraction can be applied in a later editing review pass" and "replacing the invalidated prior Skip". Those read as change commitments, so the no-change actor refused both. Step 5c now says to describe Skip only as no code/index change with the Skip recorded; adjacent lines are compacted so the review parity caps hold unchanged. Both exact packets are kept as a free regression: still refused, and accepted once Skip follows the rule. The actor's classifier is unchanged. * fix(qa-evidence): every complete capture needs an evidence row; test(tpa): accept the hyphenated app-specific-password spelling - materialize refuses when a complete capture has no evidence row and is not named in limits (CI cli-report omitted capture 004), naming the missing IDs. - tpa-apple-ban's detector required 'app-specific password' with a space; the CI answer said 'app-specific-password path' and was otherwise correct. * test(qa-observer): fix mode treats atomic temps of authorized src/test writes as transient CI webhook-fix failed with 'Could not watch test/worker.regression-1.test.ts.tmp...': Claude Code's Write renamed its temp before the per-file watch was added. The functional eval now tells the observer its mode, and a temp whose target that mode may write is observed through its directory watch. Report-only mode and undeclared paths keep failing closed. * feat(qa-evidence): refuse evidence observed on an older input snapshot than the latest capture When native probe output declares a top-level input snapshot, materialize compares each evidence row with the latest capture's snapshot and refuses stale rows unless they are classified superseded, naming the captures to rerun. ship-exploratory-late-input kept reporting a pre-change adverse probe green after the input changed. * test(qa-functional): point the fixture at the helper's --help instead of its source A CI webhook-fix run spent three turns reading lib/qa-evidence.ts to learn the interface and timed out just before materialize (agreed with #3002's owner). * feat(qa-evidence): captures list the caller's declared-but-unrun required probes GSTACK_QA_REQUIRED_PROBES (a JSON array of native child commands) makes every capture print requiredRemaining; it never judges pass or fail. The functional eval passes the webhook list from QA_WEBHOOK_REQUIRED_SCENARIOS, which the verdict now reads too, so the nudge and the verdict share one source (agreed with #3002's owner). CI webhook-report kept stopping with scenarios unrun. * test(review-army): record N+1's pre-dispatch stages and scope the session to Step 4.5 review-army-perf-n-plus-one timed out in 7 of 13 CI runs on this branch (passing 245-280 s of 300). Each session spent ~95 s on setup (the full extracted SKILL, checklist, section greps, exploratory.md, diff-scope/stats/learnings, tooling checks), ran Step 4's core pass, a search-before-recommending WebSearch, and wrote a 10-16 KB report (~100 s after the Red Team returned). The fixture now stages only review/sections/review-army.md plus the performance and red-team checklists, and hands the session the recorded detect-scope, specialist-stats and learnings outputs and the diff. The caller passes --performance (every CI parent already treated the prompt as that force flag against the <50-line skip), declares the core pass, QA, adversarial review, web research, Fix-First and persistence out of scope, and caps the report at the selection line, the SPECIALIST REVIEW block and the Red Team result (30 lines). The Performance specialist and the conditional Red Team are still real foreground subagents, and the report still has to surface the N+1. New assertion: a foreground Performance specialist dispatch precedes the Red Team dispatch. Free controls omit the Performance dispatch or background it, and both fail; the budget lifecycle adapter supplies the current result shape. Touchfiles now include the .rb fixture the case reads. * test(review-army): share the recorded Step 4.5 staging with consensus and supply its Red Team review-army-consensus (periodic) timed out in 2 of 13 census sessions; passing runs took 213-297 s of 300. Like N+1 it spent ~30-50 s reading the whole extracted SKILL, checklist and every specialist file, sometimes dispatched an unrequested Maintainability specialist, then ran a Red Team (60-70 s) and a second merge before writing a 9-15 KB report. The N+1 staging and scope text move into stageReviewArmySession / reviewArmyScope / reviewArmyChecklists (the N+1 prompt renders byte-identical). Consensus now records its detect-scope, stats, learnings and diff, stages the Review Army section with the security and testing checklists, forces --security --testing, and caps the report like N+1. Its Red Team is outside the multi-specialist contract, so the fixture supplies a labeled synthetic NO FINDINGS result instead of a dispatch. The existing SQL-finding and browser-error assertions are unchanged; the lifecycle adapter's spawnSync now returns the git output the staging reads. * docs(changelog): v1.91.10.0 records the flake census and its repairs * test(strict-output): give the spool-prefix child time to finish before the pending stream times out windows-free-tests failed on9a7a7e54: the 150 ms shared deadline raced Bun startup on Windows, so the child was killed mid-write and the spool held a partial payload. Only the never-released extra stream should time out; the child now has 3 s. * fix(qa-evidence): accept a single limits string; test(qa-callers): read the handoff first when a probe snapshot changes CI late-input spent a turn rewriting limits as an array after materialize refused a string, and a ten-read sweep hunting for the changed input before it read reports/HANDOFF.md, then timed out at 300 s. * test(autoplan-dual-voice): unwrap the framed native report before Claude Code 2.1.284's agentId/usage trailer * test(section-loading): credit a Bash print that contains every line of the carved section * test(auto-decide): ask for the selected mode in the skill's mode handoff line, not a separate public decision * test(plan-ceo floor): scope preservation approves no premise, approach or remedy * test(autoplan-dual-voice): the fixture declares that delivered bash blocks run alone, diagnostics separately * test(coverage-audit): a fenced plain-word caption in a successful && read chain is display only Census 36776104571 plan-eng capture read both owned files with cat -n in one successful && chain; the caption 'echo "=== git diff main --stat ==="' fell outside the two-token caption grammar, so both reads lost credit. Accept a fenced caption of plain words; unfenced command strings, expansions, redirection, -e escapes and ; / || tails stay rejected. * test(office-hours): a fork whose outer options are the seeded shapes is the Phase 4 question Census trials 1-2 captured complete Phase 4 forks (A) Server-side B) Client-side C) Hybrid, recommendation with because) whose prose used none of the vocabulary words. Accept two seeded shapes as outer options as Phase 4 specificity; the earlier-phase, nested, fenced and single-shape controls still fail. * fix(review): design-lite rows keep the detector's [rule-id]; the e2e detector rows point at the diff The output template had no rule-id slot, so rows merged with checklist items dropped the detector id (census t2, local t1). Rows now carry [rule-id]. The fake engine's sample rows named a foreign fixture path at line 0; the e2e remaps them to landing.html/styles.css so trials stop spending turns reconciling it. * test(shared-libs): the plan actor reads scheduler parity and unchanged-scope lists Census 36776104571's question preserved the contract ('behaving exactly like the scheduler', 'scheduler parity holds by construction') and excluded work with 'Existing copies and helper hardening stay unchanged'. Accept exactly/parity as preservation (negated forms refuse) and a bare noun list that stays unchanged as an exclusion for the expansion scan only; verb-led clauses still refuse. * fix(qa-only,qa): name the exploratory read point and finalization order; judge qa with its browser assets qa-only judges cited 'next section' pointing at the wrong heading, an exploratory trigger that contradicted its read point, clock ownership in mixed runs and the unstated order of exploratory section 4 vs reporting. The qa judge penalized the absent qa-report-template and issue-taxonomy that qa-patterns loads; with them in, it found issue-taxonomy's dangling 'rule 13' (the consent rule is browser rule 3). * test(ship-docsync): seeded attempt 1 counts toward the limit; transport counts ignore calls that never reached the state file - CI launch-failure retried after the seeded attempt 1 as if that attempt were the fixture's; the seeded prompt now says attempt 1 is this invocation's and a further attempt needs what Blocked recovery requires. - A late-result run typo'd the state path once (ENOENT, the actor never ran), then repeated the call correctly; the per-action count compared both calls with one actor event. Only calls naming the real state file are counted. * fix(plan-eng-review): show the accepted dedicated read form for coverage-diagram sources CI plan-eng-coverage-audit mixed package/config and git diff into the source read; the review variant, whose prompt shows the && display form, does not. The plan trace step now shows it too, within the unchanged size cap. * test(sync-gbrain-readiness): a negation earlier in the claim clause is not a search/write readiness claim The census unknown actor wrote 'nothing about read, search, or write capability is confirmed either way' after a YELLOW/WARN verdict. The claim window started at 'write', so the leading 'nothing' was outside it. Check the clause subject for nothing/neither/none/no; keep the original in-claim negations. Replay of the captured output passes; positive controls still flag an unnegated claim. * fix(office-hours): a forcing question's recommendation takes the position the founder's words support auq-matrix office-hours asked D1 Demand as options about the founder's own evidence and, with no rule for that shape, recommended 'answer whichever is TRUE — A is marked recommended only because it is the strongest position' (substance 2). Say what such a recommendation is: the option the founder's own words support, why it matters for the next step, and what would change it. * fix(plan-ceo-review): name the mode preference command and the exact handoff line auto-decide-preserved at6fcb0981: the model never ran the preference check, read 'check ... through the preamble' as already done, auto-selected 'per your preference setting', and wrote 'Selected mode: HOLD SCOPE, auto-decided from your tuned preference' instead of the AUTO_DECIDE handoff line. At9a7a7e54it ran the check but wrote 'Decision: HOLD SCOPE is the review mode for ...'. Neither matched the handoff template the observer recognizes. Name gstack-question-preference --check at the point of use and say the handoff begins with the exact matching line. Collapse the audit block's comment padding to stay within the unchanged 80150-byte skeleton cap. * test(section-loading): record the CEO capture's report and transcript The6fcb0981census failed hasStaleFillRaceFinding (line 98), but the case records nothing beyond junit, so the report the detector judged is gone. Return the SkillTestResult from captureSectionReads and record it, with the full saved report, through the eval collector on pass and fail. * test(design): plan-mode names its read list and caps its additions and summary At6fcb0981plan-design-review-plan-mode timed out at 300 s (9 turns): 22 cat/sed chunk reads (~50 s), then a 28 KB plan Write (~150 s), before the read-back finished. The9a7a7e54pass took 240 s with a 24.6 KB Write. Read SKILL.md, review-sections.md and plan.md natively in one response, keep additions under 14,000 characters and the summary within ten lines. Budgets unchanged. * test(plan-mode-no-op): require prose evidence before a waiting verdict ends eng/design runs (carried byte-identical from #3002) With the prose fallback forced, the gate renders as a lettered menu; a judge 'waiting' verdict on a spinner-only frame ended the run as 'asked' before the menu rendered, so the scope-gate check failed on unchanged behavior. * feat(qa-evidence): materialize computes the phase verdict; callers must report it Approved by Garry: the helper, not the model, decides whether evidence can pass. materialize writes verdict {status, open} into evidence.json and prints it: fail or blocked from row classifications, inconclusive while any row is superseded, a complete capture is withheld, a declared required probe is unrun or there is no evidence, else pass. The caller fixture requires receipt.status to equal that verdict. CI late-input kept reporting pass with a superseded happy probe. * test(qa-callers): compare the receipt with the helper verdict only when evidence.json was materialized The producer free tests run captures without materialize; evidence.json is optional for callers, so its absence is not a verdict mismatch. * test(llm-judge): run the ship workflow judge at medium effort so its panel fits JUDGE_MS claude-fable-5-1 accepts only adaptive thinking (thinking.type.enabled with budget_tokens returns 400), so effort is the available thinking control. Measured on the exact ship judge request (105,301 input tokens): - default effort, 18 samples: thinking 5,086-10,881 tokens, 75.9-144.7 s; 3 of 18 passed the 120 s deadline (about 42% of 3-sample panels). - medium effort, 18 samples: thinking 2,749-5,762, output at most 6,144 tokens, 43.1-77.9 s; scores 4/4/4 in 16 of 18 (clarity 3 in two), versus 14 of 18 at default. callJudge gains an effort option sent as output_config.effort; only the ship judge sets it. Rubric, floors, panel size, deadline, model and max_tokens are unchanged. The cache identity records effort. * test(llm-judge): ask frontier workflow judges for 120-word reasoning under the unchanged 150-word check Told "under 150 words", the ship judge's reasoning landed at 130-156 words (3 of 18 probe samples at 152-156), so the structured-response check failed about one panel in three independent of effort. The prompt's frontier block and the response schema description now say under 120 words; the validator still rejects 150 words or more. The changed prompt bytes reach only the two frontier judges: ship/SKILL.md workflow (prompt and schema) and review/SKILL.md workflow (prompt). * test(llm-judge): type the stream transport mock call * test(plan-ceo floor): the request answers only the questions it names PR lane 36794871032 (head20d6e98f): the CEO floor ran 608 s without a question. Its Step 0 recorded the premise gap and approach choice as unresolved ledger rows, then said "this session supplies all answers up front, so no decision brief was dispatched" and wrote Sections 1-11.2734e203stopped scope preservation from approving the premise; this time the actor block (declined setup, recall, outside voices, HOLD SCOPE) and the fixture's "complete user request is available from the start" were read as pre-answering every review question. The CEO actor now states that the request answers only the routing, recall, outside-reviewer and review-mode questions it names. * test(plan-devex floor): a 'Partly wrong' 0B answer is the narrative confirmation PR lane 36794871032: the DX floor asked its D1 narrative confirmation (Accurate, proceed / Partly wrong, correct it / Way off, actual is...). The deterministic setup rule accepted only 'Some ... wrong', so the question went to the LLM assessor, which hit its 30 s spawnSync cap (ETIMEDOUT) and ended the case as assessment_error at 141 s, the same failure as census 36641820398. The rule now accepts 'partly' beside 'some'; the captured question is a free regression and the remedy-option controls still go to the assessor. * test(design-review plugin handoff): quoted report text is not an install command PR lane 36794871032: every behavioral check passed except noInstallOrOverride, which matched "no `npx impeccable`" inside the quoted heredoc that wrote detector-output.md. Nothing was installed or downloaded. The check now drops quoted-delimiter heredoc bodies (literal data) before matching; unquoted bodies, which can expand $(...), and unterminated bodies stay checked. Free controls cover the captured write, bare npx, an IMPECCABLE_BIN override, an unquoted $(npx ...), npx after the delimiter and an unterminated body. * test(review-army delivery audit): stage only the plan-completion section and record its git reads PR lane 36794871032: the case timed out at its 120 s budget after 7 turns (previous lane passed in 45 s). The session read the 46 KB extracted SKILL in three passes (cat to persisted output, grep, sed), ran its own git reads, wrote a 74-line report, then inspected and ran gstack-learnings-log and rewrote the report's Learnings section. As in the Step 4.5 cases (17ee2e54/2bd4651c), the fixture now stages only review/sections/plan-completion.md, hands the session the recorded git log and diff, declares the HIGH-impact question, its Scope Check, learnings logging and later steps outside the capture, and caps the report at the audit block and its DISCREPANCY entries (30 lines). The NOT DONE and email assertions are unchanged. * feat(qa-evidence): one capture call records the causal note for the previous capture capture R NNN [--public] (--deadline D|--timeout-ms MS) --after PREV --hypothesis 'TEXT' -- CMD publishes exploration-NNN.json {observationCapture, observationArgv, observed, hypothesis, nextCapture, nextArgv} before running CMD, refusing unless PREV is the latest complete capture. The receipt carries checkpoint/checkpointSha256; validators bind the note to the transcript's capture calls by capture ID and receipt hash instead of exact command strings. The separate checkpoint command and the capture guard keep working; materialize learning accepts both note shapes and still rejects same-probe replays. Prose and eval fixture prompts teach the merged form. * fix(qa-evidence): a superseded row stops holding the verdict open once its probe is rerun on current inputs materialize requires an old-snapshot row to be classified superseded, and its verdict kept every superseded row open, so rerunning the probe (what its own error tells the model to do) could never reach pass; late-input reran 3 and 9 on the new snapshot and still got inconclusive. A superseded row now closes only when a non-superseded row with the same captured argv observed the current snapshot. Re-materializing an already-published evidence.json names the cause instead of failing generically. * test(plan-eng batching): count saved decisions whose label drops the (recommended) marker or whose report is titled 'Eng Review Report — <plan>' * fix(qa): browser-only runs skip annotations/materialize; only Q captures can anchor evidence rows * test(design): plan-mode length is a drafting target, not a check to measure and trim * test(llm-judge): structured output for doc, outcome and posture judges so reasoning quotes cannot break JSON * test(ship-docsync): steer skill file reads to Read; large cat output becomes an unpageable preview * docs(changelog): browser-only QA evidence and structured judge output * test(qa-only cleanup): refusal scenarios get a 1 s budget and an absolute worker deadline; 300 ms starved under parallel load * fix(office-hours, design-consultation): ask the goal question and read the mode section first; ask the memorable-thing question on its own * test(outside-disabled): a record named by the retained record's own clock and then disowned owns its completed status * test(context-skills): install gstack-paths in the fixture bin; without it the model guessed the checkpoint root * test(ceo mode routing): SCOPE EXPANSION posture credits plural 'expansions' * test(ship-docsync): name the unmet atomic-replacement check on a forbidden temp-file write * fix(qa): browser-only runs materialize an empty evidence list with checkpoints in limits, matching /qa-only * test(qa callers): an accepted review-log record may cite checkpoints as finding evidence * fix(plan-eng-review): state that a disallowed question tool never qualifies as headless before the headless action * merge follow-up: re-record paid CLI parity for #2999's flags; trim merged review, qa-only and plan-eng wording toward the size caps * test(golden): refresh codex/factory ship goldens for the trimmed caller QA wording * test(coverage-audit fixture): disable git auto maintenance so cleanup is not racing a detached git writer * test(parity): raise review, qa and plan-eng caps to the measured merged size of #2999 and #3002 (each fit alone), documented per cap * fix(qa-evidence): materialize rejects an unrecognized classification before publishing, so the one-shot verdict cannot be locked inconclusive by a descriptive label
1911 lines
121 KiB
TypeScript
1911 lines
121 KiB
TypeScript
import lifetimeFixture from './fixtures/ceo-fill-lifetime.json';
|
|
import { describe, expect, test } from 'bun:test';
|
|
import { readFileSync } from 'node:fs';
|
|
import { join } from 'node:path';
|
|
import {
|
|
CACHE_READ_WRITE_SKETCH,
|
|
CEO_SECTION_CACHE_PLAN,
|
|
hasStaleFillRaceFinding,
|
|
} from './helpers/ceo-section-loading-fixture';
|
|
import captured_sdk_columnar_af from './fixtures/sdk-columnar-af.json';
|
|
import captured_sdk_compact_sequence_aj from './fixtures/sdk-compact-sequence-aj.json';
|
|
import captured_sdk_order_b_ag from './fixtures/sdk-order-b-ag.json';
|
|
import fs_sdk_ordered_schedule_ar from 'node:fs';
|
|
import fixture_sdk_ordering_ae from './fixtures/sdk-ordering-ae.json';
|
|
import captured_sdk_original_order_ai from './fixtures/sdk-original-order-ai.json';
|
|
import fixture_sdk_schedule_continuation_ah from './fixtures/sdk-schedule-continuation-ah.json';
|
|
import fixture_sdk_stale_table_ad_v3 from './fixtures/sdk-stale-table-ad-v3.json';
|
|
|
|
describe('future-reader vocabulary in the actual AA finding', () => {
|
|
const report = require('node:fs').readFileSync(require('node:path').join(import.meta.dir, 'fixtures/ceo-section-aa-report.md'), 'utf8');
|
|
const paragraph = report.slice(report.indexOf('**S4-1 (CRITICAL'), report.indexOf('\n\nNo UI scope.', report.indexOf('**S4-1 (CRITICAL')));
|
|
|
|
test('recognizes the exact complete report and its same-paragraph post-write consequence', () => {
|
|
expect(paragraph).toContain('A read in-flight when a write');
|
|
expect(paragraph).toContain('stale snapshot after `cache.delete` fires');
|
|
expect(paragraph).toContain('future callers with stale data');
|
|
expect(hasStaleFillRaceFinding(paragraph)).toBe(true);
|
|
expect(hasStaleFillRaceFinding(report)).toBe(true);
|
|
expect(hasStaleFillRaceFinding(paragraph.replace('future callers', 'subsequent callers'))).toBe(true);
|
|
});
|
|
|
|
test.each(['future reads', 'future requests', 'future callers'])('recognizes a later consumer: %s', reader => {
|
|
expect(hasStaleFillRaceFinding(`An in-flight read inserts a stale snapshot after write invalidation, leaving ${reader} with stale data.`)).toBe(true);
|
|
});
|
|
|
|
test.each([
|
|
'An in-flight read inserts a stale snapshot after write invalidation. Future work documents the cache.',
|
|
'An in-flight read inserts a fresh snapshot after write invalidation, leaving future callers with fresh data.',
|
|
'An in-flight read inserts a stale snapshot before write invalidation, leaving future callers with stale data.',
|
|
'A completed read inserts a stale snapshot after write invalidation, leaving future callers with stale data.',
|
|
'An in-flight read returns a stale snapshot after write invalidation to its original pending caller.',
|
|
'An in-flight read inserts a stale snapshot after write invalidation. Future callers seeing stale data is allowed behavior.',
|
|
'An in-flight read inserts a stale snapshot after write invalidation, leaving future callers with stale data. This is the accepted consistency model.',
|
|
'An in-flight read cannot refill stale data after write invalidation. Future callers observe committed data.',
|
|
'An in-flight read inserts a stale snapshot after write invalidation, leaving future callers with stale data. This is not a bug; no guard is required.',
|
|
'> An in-flight read inserts a stale snapshot after write invalidation, leaving future callers with stale data.',
|
|
'```text\nAn in-flight read inserts a stale snapshot after write invalidation, leaving future callers with stale data.\n```',
|
|
'An in-flight read inserts a stale snapshot after write invalidation.\n\n## A different section\nFuture callers need documentation.',
|
|
'* An in-flight read inserts a stale snapshot after write invalidation.\n* Future callers need documentation.',
|
|
])('retains ordering, stale-value, source and dismissal boundaries: %s', text => {
|
|
expect(hasStaleFillRaceFinding(text)).toBe(false);
|
|
});
|
|
|
|
test('the original-caller exception cannot permit the same stale value for future callers', () => {
|
|
expect(hasStaleFillRaceFinding('An in-flight read refills stale data after write invalidation, so new reads see old data. The original pending caller may receive an old snapshot and future callers observe it; this is permitted. Guard cache fills with a generation token.')).toBe(false);
|
|
expect(hasStaleFillRaceFinding('The original pending caller may receive an old snapshot; that return is permitted. However, an in-flight read refills stale data after write invalidation, so future callers violate the contract. Guard cache fills with a generation token.')).toBe(true);
|
|
});
|
|
});
|
|
|
|
describe('restore vocabulary in the actual Y finding', () => {
|
|
const report = require('node:fs').readFileSync(require('node:path').join(import.meta.dir, 'fixtures/ceo-section-y-report.md'), 'utf8');
|
|
const amendment = report.slice(report.indexOf('**AMENDMENT (Finding 1'), report.indexOf('```javascript')).trim();
|
|
|
|
test('recognizes the delivered report and its explicit original-invariant failure', () => {
|
|
expect(amendment).toContain('original pseudocode did not satisfy');
|
|
expect(amendment).toContain('in-flight read from restoring a stale cache entry');
|
|
expect(hasStaleFillRaceFinding(amendment)).toBe(true);
|
|
expect(hasStaleFillRaceFinding(report)).toBe(true);
|
|
});
|
|
|
|
test.each(['can restore', 'restores', 'restored', 'is restoring'])('recognizes the cache-fill verb %s', verb => {
|
|
expect(hasStaleFillRaceFinding(`An in-flight read ${verb} stale data after write invalidation. A subsequent read sees the old value, violating the contract.`)).toBe(true);
|
|
});
|
|
|
|
test.each(['cannot restore', "can't restore", 'never restores', 'does not restore', "doesn't restore", 'will not restore', "won't restore", 'did not restore', "didn't restore", 'is not restoring', "isn't restoring", 'was not restoring', 'has not restored', "hasn't restored", 'had not restored'])('rejects a current prevention assertion: %s', denied => {
|
|
expect(hasStaleFillRaceFinding(`An in-flight read ${denied} stale data after write invalidation. A subsequent read observes the committed value.`)).toBe(false);
|
|
});
|
|
|
|
test.each([
|
|
'An in-flight read restores stale data after write invalidation. This is not a defect; no guard is required.',
|
|
'An in-flight read restores stale data after write invalidation. Subsequent stale reads are permitted by the contract.',
|
|
'An in-flight read restores stale data after write invalidation. This is the accepted consistency model.',
|
|
'An in-flight read restores the committed new value after write invalidation. A subsequent read observes it.',
|
|
'The original pending caller receives an old snapshot after the write; that return is permitted.',
|
|
'If deletion throws after a write, a restore operation leaves stale cache data. Log and bypass the adapter.',
|
|
'> An in-flight read restores stale data after write invalidation; a subsequent read violates the contract.',
|
|
'```text\nAn in-flight read restores stale data after write invalidation; a subsequent read violates the contract.\n```',
|
|
'* An in-flight read restores stale data after write invalidation.\n* Telemetry has a bug.',
|
|
])('preserves dismissal, source and separate-finding boundaries: %s', text => {
|
|
expect(hasStaleFillRaceFinding(text)).toBe(false);
|
|
});
|
|
});
|
|
|
|
describe('pre-write snapshot vocabulary in the actual U finding', () => {
|
|
const report = require('node:fs').readFileSync(require('node:path').join(import.meta.dir, 'fixtures/ceo-section-u-report.md'), 'utf8');
|
|
const paragraph = report.slice(report.indexOf('After T4 the cache correctly reflects'), report.indexOf('**Recommended fix (auto-decided):**')).trim();
|
|
|
|
test('recognizes the exact delivered report and its complete asserted paragraph independently of the bad remedy', () => {
|
|
expect(paragraph).toContain('re-populates the cache with the pre-write snapshot');
|
|
expect(paragraph).toContain('This violates the invariant:');
|
|
expect(hasStaleFillRaceFinding(paragraph)).toBe(true);
|
|
expect(hasStaleFillRaceFinding(report)).toBe(true);
|
|
});
|
|
|
|
test.each(['pre-write snapshot', 'pre write snapshot', 'pre-write value', 'pre-write data', 'pre-write version'])('recognizes an old snapshot synonym: %s', value => {
|
|
expect(hasStaleFillRaceFinding(`An in-flight read re-populates the cache with the ${value} after write invalidation. A new reader sees it, violating the contract.`)).toBe(true);
|
|
});
|
|
|
|
test.each([
|
|
'A pending read returns the pre-write snapshot to its original caller; that return is permitted.',
|
|
'An in-flight read re-populates the cache with the post-write snapshot after invalidation.',
|
|
'The pre-write snapshot expires after 30 seconds. The LRU byte cap is adequate.',
|
|
'If invalidation throws after a write, the cache retains the pre-write snapshot. Log the failure and bypass the cache.',
|
|
'An in-flight read re-populates the cache with the pre-write snapshot after invalidation. This is allowed behavior for subsequent reads.',
|
|
'An in-flight read re-populates the cache with the pre-write snapshot after invalidation. This is the accepted consistency model.',
|
|
'An in-flight read re-populates the cache with the pre-write snapshot after invalidation, so a new reader receives that version. This is the accepted consistency model.',
|
|
'An in-flight read re-populates the cache with the pre-write snapshot after invalidation. It is not a bug; no guard is required.',
|
|
'An in-flight read cannot re-populate the cache with the pre-write snapshot after invalidation. No race remains.',
|
|
'> An in-flight read re-populates the pre-write snapshot after write invalidation; a new read sees it, violating the contract.',
|
|
'```text\nAn in-flight read re-populates the pre-write snapshot after write invalidation; a new read sees it, violating the contract.\n```',
|
|
'* An in-flight read re-populates the pre-write snapshot after invalidation.\n* Telemetry retry handling has a bug.',
|
|
])('retains original-caller, freshness, dismissal and source boundaries: %s', value => {
|
|
expect(hasStaleFillRaceFinding(value)).toBe(false);
|
|
});
|
|
|
|
test('permission for the original caller still cannot excuse a later-reader violation', () => {
|
|
expect(hasStaleFillRaceFinding('The original pending caller may receive the pre-write snapshot; that return is permitted. However, an in-flight read refills the cache with the pre-write snapshot after write invalidation, so a new reader violates the contract. Guard cache fills with a generation token.')).toBe(true);
|
|
});
|
|
});
|
|
|
|
describe('CEO section-loading cache fixture', () => {
|
|
test.each([
|
|
{ retires: false, rejects: false },
|
|
{ retires: true, rejects: false },
|
|
{ retires: true, rejects: true },
|
|
])('cohort admission is distinct from the actual wrapper fill: %j', async ({ retires, rejects }) => {
|
|
const pending: Array<{ value: string; finish: () => void }> = [];
|
|
const flights = new Map<string, Promise<string>>();
|
|
let stored = 'old';
|
|
const failure = new Error('rolled back');
|
|
const repository = {
|
|
read(key: string) {
|
|
if (flights.has(key)) return flights.get(key)!;
|
|
const value = stored;
|
|
let finish!: () => void;
|
|
const flight = new Promise<string>(resolve => { finish = () => resolve(value); })
|
|
.finally(() => { if (flights.get(key) === flight) flights.delete(key); });
|
|
pending.push({ value, finish }); flights.set(key, flight);
|
|
return flight;
|
|
},
|
|
async write(key: string, value: string) {
|
|
if (rejects) throw failure;
|
|
stored = value;
|
|
if (retires) flights.delete(key);
|
|
return value;
|
|
},
|
|
};
|
|
const cache = new Map<string, string>();
|
|
const { readProfile, writeProfile } = new Function('cache', 'repository',
|
|
CACHE_READ_WRITE_SKETCH + '\nreturn { readProfile, writeProfile };')(cache, repository);
|
|
const earlier = readProfile('tenant:profile');
|
|
if (rejects) await expect(writeProfile('tenant:profile', 'new')).rejects.toBe(failure);
|
|
else await writeProfile('tenant:profile', 'new');
|
|
const later = readProfile('tenant:profile');
|
|
expect(pending).toHaveLength(retires && !rejects ? 2 : 1);
|
|
if (retires && !rejects) {
|
|
pending[1]!.finish();
|
|
expect(await later).toBe('new');
|
|
}
|
|
pending[0]!.finish();
|
|
expect(await earlier).toBe('old');
|
|
if (!retires || rejects) expect(await later).toBe('old');
|
|
// Even correct repository admission cannot stop this exact new sketch
|
|
// from caching its older result after the committed write. Keep that gap.
|
|
expect(await readProfile('tenant:profile')).toBe('old');
|
|
expect(stored).toBe(rejects ? 'old' : 'new');
|
|
expect(CEO_SECTION_CACHE_PLAN).toContain('single-flight wrapper sits inside\n repository.read');
|
|
expect(CEO_SECTION_CACHE_PLAN).toContain('A committed repository.write retires');
|
|
expect(CEO_SECTION_CACHE_PLAN).toContain('This admission rule does not inspect cache fills');
|
|
expect(CEO_SECTION_CACHE_PLAN).toContain(CACHE_READ_WRITE_SKETCH);
|
|
expect(hasStaleFillRaceFinding(CEO_SECTION_CACHE_PLAN)).toBe(false);
|
|
});
|
|
|
|
test('author bounds implementation depth without preapproving the wrapper or weakening required proof', () => {
|
|
expect(CEO_SECTION_CACHE_PLAN).toContain('wrapper itself remains unapproved');
|
|
expect(CEO_SECTION_CACHE_PLAN).toContain('actual\ncontradiction or missing proof must be reported and resolved');
|
|
expect(CEO_SECTION_CACHE_PLAN).toContain('exact data\nstructures, full function bodies and executable test code belong to subsequent\nengineering planning');
|
|
expect(CEO_SECTION_CACHE_PLAN).toContain('not tests already\nimplemented or passing');
|
|
expect(CEO_SECTION_CACHE_PLAN).toContain('Preserve all 11 review outcomes');
|
|
expect(CEO_SECTION_CACHE_PLAN).toContain('full GSTACK REVIEW REPORT');
|
|
});
|
|
|
|
test('declared absence decoding and atomic write failure do not add independent wrapper defects', async () => {
|
|
const missing = Object.freeze({ found: false });
|
|
const absent = Symbol('adapter-private absence');
|
|
const stored = new Map<string, unknown>();
|
|
const cache = {
|
|
get: (key: string) => stored.get(key) === absent ? missing : stored.get(key),
|
|
set: (key: string, value: unknown) => stored.set(key, value === missing ? absent : value),
|
|
delete: (key: string) => stored.delete(key),
|
|
};
|
|
const rejected = new Error('atomic write rejected before commit');
|
|
let reads = 0;
|
|
const repository = {
|
|
read: async () => { reads++; return missing; },
|
|
write: async () => { throw rejected; },
|
|
};
|
|
const { readProfile, writeProfile } = new Function('cache', 'repository',
|
|
CACHE_READ_WRITE_SKETCH + '\nreturn { readProfile, writeProfile };')(cache, repository);
|
|
expect(await readProfile('tenant:missing')).toBe(missing);
|
|
expect(stored.get('tenant:missing')).toBe(absent);
|
|
expect(await readProfile('tenant:missing')).toBe(missing);
|
|
expect(reads).toBe(1);
|
|
await expect(writeProfile('tenant:missing', { found: true })).rejects.toBe(rejected);
|
|
expect(await readProfile('tenant:missing')).toBe(missing);
|
|
expect(reads).toBe(1);
|
|
expect(CEO_SECTION_CACHE_PLAN).toContain('every\n rejected promise guarantees no commit');
|
|
expect(CEO_SECTION_CACHE_PLAN).toContain('cache.get decodes it back to the same absent-result DTO');
|
|
expect(CEO_SECTION_CACHE_PLAN).toContain('cannot fill the new one');
|
|
expect(CEO_SECTION_CACHE_PLAN).toContain('does not coordinate\nan ordinary DB write');
|
|
});
|
|
|
|
test.each([false, true])('rollout publication fences admitted old writes: %s', async (fenceWrites) => {
|
|
let stored = 'old';
|
|
let releaseWrite!: () => void;
|
|
const gate = new Promise<void>(resolve => { releaseWrite = resolve; });
|
|
const repository = {
|
|
read: async () => stored,
|
|
write: async (_key: string, value: string) => { await gate; stored = value; return value; },
|
|
};
|
|
const instance = () => new Function('cache', 'repository',
|
|
CACHE_READ_WRITE_SKETCH + '\nreturn { readProfile, writeProfile };')(new Map(), repository);
|
|
const old = instance();
|
|
const writing = old.writeProfile('tenant:profile', 'new');
|
|
let published = false;
|
|
const publish = async () => {
|
|
if (fenceWrites) await writing;
|
|
published = true;
|
|
return instance();
|
|
};
|
|
const publishing = publish();
|
|
await Promise.resolve();
|
|
expect(published).toBe(!fenceWrites);
|
|
if (!fenceWrites) {
|
|
const next = await publishing;
|
|
expect(await next.readProfile('tenant:profile')).toBe('old');
|
|
releaseWrite(); await writing;
|
|
// The initial isolation-only contract still permits a stale new cache.
|
|
expect(await next.readProfile('tenant:profile')).toBe('old');
|
|
} else {
|
|
releaseWrite(); await writing;
|
|
const next = await publishing;
|
|
expect(await next.readProfile('tenant:profile')).toBe('new');
|
|
}
|
|
expect(CEO_SECTION_CACHE_PLAN).toContain('awaits every admitted old-instance write');
|
|
expect(CEO_SECTION_CACHE_PLAN).toContain('fresh single-flight cohort before admitting new work');
|
|
});
|
|
|
|
test('an internal store commit still overlaps the unfinished public wrapper write', async () => {
|
|
let stored = 'old';
|
|
let commitWrite!: () => void;
|
|
let wrapperReturned = false;
|
|
const cache = new Map([['tenant:profile', 'old']]);
|
|
const repository = {
|
|
read: async () => stored,
|
|
write: () => new Promise<string>(resolve => {
|
|
commitWrite = () => { stored = 'new'; resolve('new'); };
|
|
}),
|
|
};
|
|
const { readProfile, writeProfile } = new Function('cache', 'repository',
|
|
CACHE_READ_WRITE_SKETCH + '\nreturn { readProfile, writeProfile };')(cache, repository);
|
|
const writing = writeProfile('tenant:profile', 'new').then((value: string) => {
|
|
wrapperReturned = true;
|
|
return value;
|
|
});
|
|
commitWrite();
|
|
expect(stored).toBe('new');
|
|
expect(wrapperReturned).toBe(false);
|
|
// Invocation precedes the wrapper's invalidation/return continuation.
|
|
// Its old cache hit is permitted; a later caller is still protected.
|
|
const overlappingRead = readProfile('tenant:profile');
|
|
await writing;
|
|
expect(wrapperReturned).toBe(true);
|
|
expect(cache.has('tenant:profile')).toBe(false);
|
|
expect(await overlappingRead).toBe('old');
|
|
expect(await readProfile('tenant:profile')).toBe('new');
|
|
const contract = CEO_SECTION_CACHE_PLAN.replace(/\s+/g, ' ');
|
|
expect(contract).toContain("when writeProfile's promise fulfills after cache.delete, not when repository.write commits or resolves");
|
|
expect(contract).toContain('Reads that overlap an unfinished writeProfile may return an earlier snapshot');
|
|
});
|
|
|
|
test('an old miss filled before the completed write is correctly invalidated', async () => {
|
|
let stored = 'old';
|
|
let releaseRead!: () => void;
|
|
let first = true;
|
|
const cache = new Map<string, string>();
|
|
const repository = {
|
|
read: () => {
|
|
if (!first) return Promise.resolve(stored);
|
|
first = false;
|
|
const snapshot = stored;
|
|
return new Promise<string>(resolve => { releaseRead = () => resolve(snapshot); });
|
|
},
|
|
write: async (_key: string, value: string) => { stored = value; return value; },
|
|
};
|
|
const { readProfile, writeProfile } = new Function('cache', 'repository',
|
|
CACHE_READ_WRITE_SKETCH + '\nreturn { readProfile, writeProfile };')(cache, repository);
|
|
const earlierRead = readProfile('tenant:profile');
|
|
releaseRead();
|
|
expect(await earlierRead).toBe('old');
|
|
expect(cache.get('tenant:profile')).toBe('old');
|
|
await writeProfile('tenant:profile', 'new');
|
|
expect(stored).toBe('new');
|
|
expect(cache.has('tenant:profile')).toBe(false);
|
|
expect(await readProfile('tenant:profile')).toBe('new');
|
|
});
|
|
|
|
test('the exact proposed wrapper retains a reproducible stale-fill race', async () => {
|
|
let releaseRead!: (value: string) => void;
|
|
let stored = 'old';
|
|
const cache = new Map<string, string>();
|
|
const repository = {
|
|
read: () => new Promise<string>((resolve) => { releaseRead = resolve; }),
|
|
write: async (_key: string, value: string) => { stored = value; return value; },
|
|
};
|
|
// Execute the same sketch the live reviewer receives, not a second model
|
|
// of its ordering. Holding the old read exposes the intended interleaving.
|
|
const { readProfile, writeProfile } = new Function('cache', 'repository',
|
|
CACHE_READ_WRITE_SKETCH + '\nreturn { readProfile, writeProfile };')(cache, repository);
|
|
const pending = readProfile('tenant:profile');
|
|
await writeProfile('tenant:profile', 'new');
|
|
releaseRead('old');
|
|
const earlierResult = await pending;
|
|
expect(stored).toBe('new');
|
|
// Returning the earlier snapshot to the already-pending caller is
|
|
// explicitly permitted. Reusing it for this new reader is the defect.
|
|
expect(earlierResult).toBe('old');
|
|
expect(CEO_SECTION_CACHE_PLAN).toContain('Every read begun after that write completes must');
|
|
expect(await readProfile('tenant:profile')).toBe('old');
|
|
expect(CEO_SECTION_CACHE_PLAN).toContain(CACHE_READ_WRITE_SKETCH);
|
|
expect(hasStaleFillRaceFinding(CEO_SECTION_CACHE_PLAN)).toBe(false);
|
|
});
|
|
|
|
test.each([
|
|
// Actual finding in the unchanged fixture's successful 38 KB live report.
|
|
'**Missing: What happens to in-flight requests during invalidation?** If a write invalidates a key and 10 requests are simultaneously loading it (cache miss, in-flight DB fetch), all 10 will cache the same value after the invalidation. The invalidated key may get re-populated with a stale value if any of those fetches started before the write. No mention of this race.',
|
|
'P1: An in-flight read can repopulate stale data after a committed write invalidates the key. Guard fills with a generation token.',
|
|
'| Cache fill race | An older value fetched before the write is inserted after eviction, so the next read is stale. | Add a per-key epoch. |',
|
|
'**Invalidation race:** the pending fetch stores an outdated snapshot after cache.delete. Serialize the fill with mutation.',
|
|
])('recognizes the actual ordering defect: %s', (report) => {
|
|
expect(hasStaleFillRaceFinding(report)).toBe(true);
|
|
});
|
|
|
|
test.each([
|
|
'The full review is complete. No issues found.',
|
|
'A stale value expires after 30 seconds. The LRU byte cap is adequate.',
|
|
'If invalidation throws after a write, the cache retains stale data. Log the failure and bypass the cache.',
|
|
'Read and write concurrency is covered. No stale data can be returned.',
|
|
'| Reads | Coalesced concurrent misses |\n| Writes | Invalidation failure leaves stale data |',
|
|
'```javascript\n// An in-flight read can cache stale data after invalidation.\n```',
|
|
'> An in-flight read can cache stale data after invalidation.',
|
|
])('rejects completion, unrelated text, and quoted source: %s', (report) => {
|
|
expect(hasStaleFillRaceFinding(report)).toBe(false);
|
|
});
|
|
});
|
|
|
|
|
|
const CAPTURED_ACCEPTED_RACE_REPORT = `**Shadow paths:**
|
|
1. Nil key: Programming error — caught by auth/key-validation before wrapper.
|
|
2. Empty key: Same — upstream validation gate.
|
|
3. Upstream error: Single-flight releases all waiters with the error. Cache
|
|
not populated. Next request retries DB. Correct.
|
|
4. Concurrent write during read in-flight: The plan documents this explicitly.
|
|
The stale read is an accepted invariant, bounded by 30s TTL.
|
|
|
|
**Async ordering — critical race:**
|
|
\`\`\`
|
|
1. Request A: cache.get(key) → miss → enters single-flight
|
|
2. Request B: cache.get(key) → miss → joins single-flight (awaiting)
|
|
3. fn: repository.read(key) → suspend (await)
|
|
4. Write commits → cache.delete(key) [nothing to delete — key not set yet]
|
|
5. repository.read(key) returns OLD snapshot (pre-write)
|
|
6. cache.set(key, OLD_VALUE) ← stale value in cache for up to 30s
|
|
7. Requests A and B both return OLD_VALUE ← accepted by plan
|
|
\`\`\`
|
|
|
|
This is the one documented asymmetry. It is not a gap — it is a named invariant.
|
|
The TTL bounds the stale window to 30 seconds.
|
|
|
|
`;
|
|
|
|
|
|
describe('CEO concurrency finding requires a violation, not an accepted trace', () => {
|
|
test('rejects the captured accepted-invariant report that passed the old keyword oracle', () => {
|
|
expect(hasStaleFillRaceFinding(CAPTURED_ACCEPTED_RACE_REPORT)).toBe(false);
|
|
});
|
|
|
|
test.each([
|
|
'An in-flight fetch can refill the cache with old data after a write invalidates it. The next read sees that stale snapshot, violating the post-write contract.',
|
|
'The pending read stores an older value after invalidation.\n\nGuard cache fills with a version check so a later request cannot observe pre-write state.',
|
|
'Returning the old snapshot to the pending caller is permitted. But a late cache.set after concurrent write invalidation exposes stale data to a new reader. Serialize mutation and cache fills.',
|
|
'Returning an old snapshot to the original pending caller is an accepted invariant. But an in-flight read can repopulate stale cache data after write invalidation, so a new reader violates the post-write contract. Guard cache fills with a generation token.',
|
|
'The original caller may receive the old snapshot; that return is permitted. However, a pending fetch refills stale data after write invalidation, breaking consistency for a later reader. Skip the cache fill when its version changed.',
|
|
'An in-flight read can repopulate stale data after write invalidation, so a new reader gets the old value. This is not permitted by the contract. Guard cache fills with a version check.',
|
|
'| Late cache fill | A concurrent read repopulates an outdated result after eviction. | Reject the fill when its generation token changed. |',
|
|
])('accepts the later-reader consequence or a concrete ordering remedy: %s', report => {
|
|
expect(hasStaleFillRaceFinding(report)).toBe(true);
|
|
});
|
|
|
|
test.each([
|
|
'An in-flight read repopulates stale data after write invalidation. This is an accepted invariant bounded by the TTL.',
|
|
'An in-flight read repopulates stale data after write invalidation. This is permitted by the contract; a later read may be stale for 30 seconds.',
|
|
'An in-flight read refills stale data after write invalidation. This is allowed behavior for the next read because the TTL bounds it.',
|
|
'The original caller and the new reader may both observe the old snapshot as an accepted invariant. A pending read refills stale data after write invalidation; no guard is required.',
|
|
'An in-flight read repopulates stale data after write invalidation.\n\nIt is not a gap. No change is needed.',
|
|
'No race: a pending read cannot repopulate stale cache data after write invalidation; the existing version check rejects it.',
|
|
'An in-flight read repopulates stale data after write invalidation, but does not violate the contract. No guard is required.',
|
|
'1. Cache population after a miss is safe.\n2. Concurrent writes can return an older snapshot to their original pending reader.\n3. Guard unrelated network retries.',
|
|
'An in-flight read stores stale data after write invalidation.\n\n**Finding S9:** Guard telemetry delivery with a version token.',
|
|
'An in-flight read stores stale data after write invalidation.\n\nGuard unrelated telemetry delivery with a version token.',
|
|
'1. An in-flight read repopulates stale data after write invalidation.\n2. Telemetry retry handling has a bug.',
|
|
'An in-flight read repopulates stale data after write invalidation.\n\n#2 — Unrelated telemetry delivery bug',
|
|
'* An in-flight read repopulates stale data after write invalidation.\n* Telemetry retry handling has a bug.',
|
|
'> P1: An in-flight read refills stale data after invalidation; a new read gets the old value.',
|
|
'```text\nP1: An in-flight read refills stale data after invalidation; a new read gets the old value.\n```',
|
|
])('rejects dismissals, negations, unrelated findings and quoted examples: %s', report => {
|
|
expect(hasStaleFillRaceFinding(report)).toBe(false);
|
|
});
|
|
});
|
|
|
|
|
|
describe('proposed cache-fill prevention remains an unresolved finding', () => {
|
|
test('an imperative remedy describes the behavior it must prevent', () => {
|
|
expect(hasStaleFillRaceFinding('An in-flight read repopulates stale data after write invalidation. Guard cache fills so pending reads cannot repopulate stale values after invalidation.')).toBe(true);
|
|
});
|
|
|
|
test('an existing guard remains a dismissal, not a proposed fix', () => {
|
|
expect(hasStaleFillRaceFinding('An in-flight read cannot repopulate stale data after write invalidation because the existing guard rejects that fill. No race remains.')).toBe(false);
|
|
});
|
|
|
|
test('an imperative does not erase a separate explicit dismissal', () => {
|
|
expect(hasStaleFillRaceFinding('An in-flight read repopulates stale data after write invalidation. Guard cache fills so pending reads cannot repopulate stale values after invalidation. This is not a bug; no fix is needed.')).toBe(false);
|
|
});
|
|
});
|
|
|
|
|
|
// The native report separates an asserted finding, its ordered trace and its
|
|
// explicit contract violation. Detection does not certify the offered fix.
|
|
describe('structured native stale-fill finding', () => {
|
|
const report = require('node:fs').readFileSync(require('node:path').join(import.meta.dir, 'fixtures/ceo-section-loading-l-report.md'), 'utf8');
|
|
const finding = report.slice(report.indexOf('**CRITICAL FINDING — Write-then-read stale-set race**'), report.indexOf('**Required fix:**'));
|
|
test('retains the exact positive later-reader finding even though the proposed mitigation is wrong', () => {
|
|
expect(hasStaleFillRaceFinding(report)).toBe(true);
|
|
expect(hasStaleFillRaceFinding(finding)).toBe(true);
|
|
});
|
|
test.each([
|
|
['standalone trace', finding.slice(finding.indexOf('```'), finding.lastIndexOf('```') + 3)],
|
|
['quoted finding', finding.split('\n').map((line: string) => '> ' + line).join('\n')],
|
|
['source example', 'Example of report format:\n' + finding],
|
|
['outer fenced source', '````text\n' + finding + '\n````'],
|
|
['explicit accepted trace', finding.replace('This violates the stated invariant:', 'This is not a gap. The following behavior is accepted:')],
|
|
['negated violation', finding.replace('This violates the stated invariant:', 'This does not violate the stated invariant:')],
|
|
['separate dismissal', finding + '\nThis is not a defect; no fix is required.\n'],
|
|
['no new reader', finding.replace(/T3: readProfile[\s\S]*?\n```/, '```')],
|
|
['reverse ordering', finding.replace('cache.delete(key)', 'cache.get(key)')],
|
|
['unrelated heading', finding.replace('CRITICAL FINDING', 'EXAMPLE')],
|
|
...['~~~', '````'].map(fence => ['nontriple fenced violation', finding.replace(/This violates[\s\S]*$/, text => fence + 'text\n' + text + '\n' + fence)]),
|
|
['unclosed fenced violation', finding.replace('This violates', '```text\nThis violates')],
|
|
['later named finding', finding.replace('This violates', '**CRITICAL FINDING — unrelated documentation defect**\nThis violates')],
|
|
['later heading', finding.replace('This violates', '## Unrelated finding\nThis violates')],
|
|
|
|
])('rejects %s', (_name, text) => expect(hasStaleFillRaceFinding(text)).toBe(false));
|
|
});
|
|
|
|
|
|
// Actual Q report amended the original contract to accept later stale reads.
|
|
// The oracle must not count that permission paragraph as an unresolved defect.
|
|
describe('accepted consistency model is not a stale-fill finding', () => {
|
|
const report = require('node:fs').readFileSync(require('node:path').join(import.meta.dir, 'fixtures/ceo-section-loading-q-report.md'), 'utf8');
|
|
const accepted = report.slice(report.indexOf('- **AMENDED (stale-fill race):**'), report.indexOf('\n\n', report.indexOf('- **AMENDED (stale-fill race):**')));
|
|
test('rejects the exact amended consistency paragraph', () => {
|
|
expect(accepted).toMatch(/accepted consistency\s+model/);
|
|
expect(hasStaleFillRaceFinding(accepted)).toBe(false);
|
|
});
|
|
test('rejects the complete Q report that accepts the late-fill race', () => {
|
|
expect(hasStaleFillRaceFinding(report)).toBe(false);
|
|
});
|
|
test.each([
|
|
'An in-flight read refills stale data after write invalidation. A subsequent read sees the old value. This is the accepted consistency model; TTL expiry is the consistency deadline.',
|
|
'An in-flight read refills stale data after write invalidation. A subsequent read sees the old value. This remains the documented consistency contract.',
|
|
'An in-flight read refills stale data after write invalidation. A subsequent read sees the old value. This is an intentional consistency policy.',
|
|
])('rejects a declared consistency allowance: %s', text => {
|
|
expect(hasStaleFillRaceFinding(text)).toBe(false);
|
|
});
|
|
test.each([
|
|
'An in-flight read refills stale data after write invalidation. A subsequent read sees the old value. Every later read must observe the committed version; this is the accepted consistency model. The stale refill violates that contract.',
|
|
'An in-flight read refills stale data after write invalidation. A subsequent read sees the old value. Every later read must observe the committed version. This is the accepted consistency model. The stale refill violates that contract.',
|
|
'An in-flight read refills stale data after write invalidation. A subsequent read sees the old value, violating the accepted consistency model.',
|
|
'An in-flight read refills stale data after write invalidation. The accepted consistency model requires a subsequent read to observe the committed version; this violates that contract.',
|
|
'An in-flight read refills stale data after write invalidation. A subsequent read sees the old value. This is not the accepted consistency model.',
|
|
'An in-flight read refills stale data after write invalidation. A subsequent read sees the old value. This is not an accepted consistency model.',
|
|
])('retains an unresolved later-reader violation without demanding a correct remedy: %s', text => {
|
|
expect(hasStaleFillRaceFinding(text)).toBe(true);
|
|
});
|
|
});
|
|
|
|
|
|
describe('CEO R report requirement weakening remains rejected', () => {
|
|
test('a race trace followed by acceptance and a weaker guarantee is not an unresolved defect', async () => {
|
|
// Exact delivered R retry report; the existing oracle already rejects it.
|
|
// This pins the policy failure without changing the paid fixture or oracle.
|
|
const report = await Bun.file(new URL('./fixtures/ceo-section-r-rejected-report.md', import.meta.url)).text();
|
|
expect(report).toContain('Every read begun after that write completes must');
|
|
expect(report).toContain('this is an accepted design choice, not a quality gap');
|
|
expect(report).toContain('Verify: new reads see stale value until TTL');
|
|
expect(hasStaleFillRaceFinding(report)).toBe(false);
|
|
});
|
|
});
|
|
|
|
// Exact S output names the original-contract defect but proposes an ineffective
|
|
// guard. Detection and remedy correctness remain separate assertions.
|
|
describe('S native same-finding ordered trace', () => {
|
|
const report = require('node:fs').readFileSync(require('node:path').join(import.meta.dir, 'fixtures/ceo-section-s-trace-report.md'), 'utf8');
|
|
const finding = report.slice(report.indexOf('### Critical Finding: Stale Re-insertion After Write Invalidation'), report.indexOf('### State Machine: Cache Entry'));
|
|
test('recognizes the exact delivered S finding without certifying its remedy', () => {
|
|
expect(hasStaleFillRaceFinding(report)).toBe(true);
|
|
expect(hasStaleFillRaceFinding(finding)).toBe(true);
|
|
expect(report).toContain('if (cache.get(key) === undefined)');
|
|
});
|
|
test('the same ordered evidence tolerates whitespace and a consistent key name', () => {
|
|
expect(hasStaleFillRaceFinding(finding.replaceAll('(key', '(accountKey').replaceAll(' t', ' t'))).toBe(true);
|
|
expect(hasStaleFillRaceFinding(finding.replaceAll('(key', '($key'))).toBe(true);
|
|
});
|
|
test.each([
|
|
['optional single-flight label', finding.replace('single-flight → ', '')],
|
|
['old/stale value vocabulary', finding.replace('OLD snapshot', 'stale value').replace('OLD_VALUE', 'STALE_VALUE').replace('stale value re-inserted', 'old snapshot refilled').replace('next readProfile', 'subsequent readProfile')],
|
|
['prose payload vocabulary', finding.replace('OLD_VALUE', 'old value').replace('returns stale value', 'returns old snapshot')],
|
|
['ASCII arrows and compact spacing', finding.replaceAll(' → ', '->').replaceAll(' ← ', '<-')],
|
|
['trace keyword case', finding.replace(/t[1-6]:[^\n]*/g, (event: string) => event.toLowerCase())],
|
|
['call whitespace and optional suspension annotation', finding.replaceAll('(key)', '( key )').replace('(key, OLD_VALUE)', '( key , OLD_VALUE )').replaceAll(' (suspends)', '')],
|
|
])('accepts equivalent %s', (_name, text) => expect(hasStaleFillRaceFinding(text)).toBe(true));
|
|
test.each([
|
|
['only the trace', finding.slice(finding.indexOf('```'), finding.indexOf('```', finding.indexOf('```') + 3) + 3)],
|
|
['quoted whole finding', finding.split('\n').map((line: string) => '> ' + line).join('\n')],
|
|
['indented whole finding', finding.split('\n').map((line: string) => ' ' + line).join('\n')],
|
|
['source format preface', 'Example of report format:\n' + finding],
|
|
['outer fenced source', '````text\n' + finding + '\n````'],
|
|
['missing independent violation', finding.replace('The proposed wrapper violates this invariant.', '')],
|
|
['negated independent violation', finding.replace('The proposed wrapper violates this invariant.', 'The proposed wrapper does not violate this invariant.')],
|
|
['unrelated heading', finding.replace('### Critical Finding:', '### Example:')],
|
|
['separate named finding', finding.replace('Race sequence', '### Another finding\nRace sequence')],
|
|
['separate bold finding', finding.replace('Race sequence', '**HIGH FINDING — unrelated issue**\nRace sequence')],
|
|
['missing write completion', finding.replace('DB write completes', 'DB write remains pending')],
|
|
['no invalidation', finding.replace('cache.delete(key) → writeProfile returns', 'cache.get(key) → writeProfile returns')],
|
|
['missing late old fill', finding.replace('cache.set(key, OLD_VALUE)', 'cache.set(key, NEW_VALUE)')],
|
|
['different filled key', finding.replace('cache.set(key, OLD_VALUE)', 'cache.set(otherKey, OLD_VALUE)')],
|
|
['case-distinct filled key', finding.replace('cache.set(key, OLD_VALUE)', 'cache.set(KEY, OLD_VALUE)')],
|
|
['different later key', finding.replace('next readProfile(key)', 'next readProfile(otherKey)')],
|
|
['only original pending reader', finding.replace('next readProfile(key)', 'original pending readProfile(key)')],
|
|
['later reader misses', finding.replace('cache HIT → returns stale value', 'cache MISS → returns committed value')],
|
|
['nonviolating trace', finding.replace('← INVARIANT VIOLATED', '← INVARIANT PRESERVED')],
|
|
['reverse order labels', finding.replace('t3:', 't4:').replace('t4: DB read', 't3: DB read')],
|
|
['missing event', finding.replace(/^.*t4:.*\n/m, '')],
|
|
['unclosed trace', finding.replace('```\n\nThe plan says', '\nThe plan says')],
|
|
['source-code trace fence', finding.replace('```\n t1:', '```javascript\n t1:')],
|
|
['split traces', finding.replace(' t4:', '```\n\n```\n t4:')],
|
|
['accepted stale trace', finding + '\nThis stale-read behavior is accepted; no guard is required.\n'],
|
|
['allowed new-reader consequence', finding + '\nA subsequent stale read is permitted by the amended contract.\n'],
|
|
['explicit defect dismissal', finding + '\nThis is not a defect; no fix is required.\n'],
|
|
])('rejects %s', (_name, text) => expect(hasStaleFillRaceFinding(text)).toBe(false));
|
|
});
|
|
|
|
// Actual v2 SDK review identifies the missing fill/write coordination directly.
|
|
// The full delivered report is retained in run evidence; this is its exact finding.
|
|
describe('explicit uncoordinated cache-fill freshness violation', () => {
|
|
const finding = [
|
|
"**[Amended: D2, D3, D4, D5]** The original sketch had no coordination between a",
|
|
"cache fill and a write and omitted the single-flight wrapper and the absence",
|
|
"sentinel; finding F1 showed that violates the read-after-write rule. The",
|
|
"ordering rules below replace it. `flight` is the existing per-key single-flight",
|
|
"wrapper extended with `invalidate(key)` and `invalidateAll()`; a fill ticket is",
|
|
"`live()` until its key is invalidated. `ProfileNotFound` stands for the",
|
|
"repository's existing typed not-found error class.",
|
|
].join('\n');
|
|
test('accepts the actual finding without requiring its separate execution diagram', () => {
|
|
expect(hasStaleFillRaceFinding(finding)).toBe(true);
|
|
expect(hasStaleFillRaceFinding('## Proposed wrapper integration\nKeep the current read-through repository interface and shared adapters.\n' + finding)).toBe(true);
|
|
});
|
|
test.each([
|
|
['current wrapper', finding.replace('original sketch had', 'current wrapper has')],
|
|
['proposed implementation', finding.replace('original sketch had', 'proposed implementation has')],
|
|
['freshness contract', finding.replace('read-after-write rule', 'read-after-write contract')],
|
|
])('recognizes equivalent %s evidence', (_name, text) => expect(hasStaleFillRaceFinding(text)).toBe(true));
|
|
test.each([
|
|
['no violation asserted', finding.replace('finding F1 showed that violates the read-after-write rule.', '')],
|
|
['negated violation', finding.replace('that violates', 'that does not violate')],
|
|
['uncertain violation', finding.replace('that violates', 'that might violate')],
|
|
['conditional premise', 'If ' + finding],
|
|
['coordination exists', finding.replace('had no coordination', 'had coordination')],
|
|
['different operations', finding.replace('cache fill and a write', 'cache hit and a read')],
|
|
['wrong contract', finding.replace('read-after-write', 'read-before-write')],
|
|
['dismissed defect', finding + '\n\nThis is not a defect; no fix is required.'],
|
|
['accepted stale consequence', finding + '\n\nA subsequent stale read is permitted by the amended contract.'],
|
|
['quoted finding', finding.split('\n').map(line => '> ' + line).join('\n')],
|
|
['indented finding', finding.split('\n').map(line => ' ' + line).join('\n')],
|
|
['quoted paragraph', '"' + finding + '"'],
|
|
['source preface', 'Example of report format:\n' + finding],
|
|
['separate source preface', 'Example of report format:\n\n' + finding],
|
|
['backtick source fence', '````text\n' + finding + '\n````'],
|
|
['tilde source fence', '~~~text\n' + finding + '\n~~~'],
|
|
['unclosed source fence', '```text\n' + finding],
|
|
['split unrelated paragraphs', finding.replace('sentinel; finding', 'sentinel.\n\n### Separate issue\nFinding')],
|
|
])('rejects %s', (_name, text) => expect(hasStaleFillRaceFinding(text)).toBe(false));
|
|
});
|
|
|
|
// Peer counterexamples: embedded, uncertain and hypothetical assertions stay closed.
|
|
describe('coordination findings require directly asserted premises and conclusions', () => {
|
|
const premise = 'The original sketch had no coordination between a cache fill and a write. ';
|
|
const claim = 'This violates the read-after-write rule.';
|
|
test('accepts a direct assertion', () => expect(hasStaleFillRaceFinding(premise + claim)).toBe(true));
|
|
test.each([
|
|
['negated embedded conclusion', premise + 'It is false that this violates the read-after-write rule.'],
|
|
['unproven conclusion', premise + 'We have not shown that it violates the read-after-write rule.'],
|
|
['uncertain conclusion', premise + 'It is unclear whether this violates the read-after-write rule.'],
|
|
['question rather than assertion', premise + claim.replace('.', '?')],
|
|
['separate issue without blank line', premise + '\n## A different issue\nReplica lag is high. ' + claim],
|
|
['hypothetical premise', 'Suppose ' + premise + claim],
|
|
['suggested report', 'A suggested report sentence: ' + premise + claim],
|
|
['nested unmatched fence', '````text\n' + premise + claim + '\n```\n' + premise + claim + '\n````'],
|
|
['mismatched fence', '~~~text\n' + premise + claim + '\n```\n' + premise + claim],
|
|
])('rejects %s', (_name, text) => expect(hasStaleFillRaceFinding(text)).toBe(false));
|
|
});
|
|
|
|
// AM first SDK report, exact public Write53 acknowledged by native result54.
|
|
// The original report remains in run evidence; this is its complete asserted paragraph.
|
|
describe('reported original coordination violation with an owned finding citation', () => {
|
|
const finding = "`[Amended: F1, F2, F3, F6]` The original sketch stated that no coordination\nbetween a cache fill and a write was proposed. Review showed that sketch\nviolates the retained read-after-write invariant (see F1). The ordering rules\nbelow replace it. They are the complete new read/write ordering rules.";
|
|
test('accepts the exact owned paragraph without requiring its separate amended diagram', () => {
|
|
expect(hasStaleFillRaceFinding(finding)).toBe(true);
|
|
expect(hasStaleFillRaceFinding('## Proposed wrapper integration\n\n' + finding)).toBe(true);
|
|
});
|
|
test('binds the named violation to its original subject and finding identity', () => {
|
|
expect(hasStaleFillRaceFinding(finding.replaceAll('F1', 'F7'))).toBe(true);
|
|
expect(hasStaleFillRaceFinding(finding.replaceAll('sketch', 'wrapper'))).toBe(true);
|
|
expect(hasStaleFillRaceFinding('## Historical example\n\nA copied example.\n\n## Current review\n\n' + finding)).toBe(true);
|
|
expect(hasStaleFillRaceFinding(finding + '\n\n## Assessment of F2\nF2 is rejected.')).toBe(true);
|
|
expect(hasStaleFillRaceFinding(finding + '\n\n## Assessment of F1\n> F1 is rejected.')).toBe(true);
|
|
});
|
|
test.each([
|
|
['negated missing coordination', finding.replace('no coordination', 'coordination')],
|
|
['unrelated operations', finding.replace('cache fill and a write', 'cache hit and a read')],
|
|
['uncertain absence', finding.replace('stated that no', 'might have stated that no')],
|
|
['unproven violation', finding.replace('Review showed', 'Review may show')],
|
|
['negated violation', finding.replace('sketch\nviolates', 'sketch\ndoes not violate')],
|
|
['hypothetical violation', finding.replace('sketch\nviolates', 'sketch\nmight violate')],
|
|
['wrong contract', finding.replace('read-after-write', 'read-before-write')],
|
|
['different subject', finding.replace('Review showed that sketch', 'Review showed that wrapper')],
|
|
['missing finding identity', finding.replace(' (see F1)', '')],
|
|
['question rather than conclusion', finding.replace('(see F1).', '(see F1)?')],
|
|
['conditional finding', 'If approved: ' + finding],
|
|
['historical owner', '## Historical example\n\n' + finding],
|
|
['source owner', '## Quoted source\n\n' + finding],
|
|
['hypothetical owner', '## Hypothetical example\n\n' + finding],
|
|
['explicit source preface', 'The following is a quoted source excerpt.\n\n' + finding],
|
|
['quoted paragraph', '"' + finding + '"'],
|
|
['block quote', finding.split('\n').map(line => '> ' + line).join('\n')],
|
|
['fenced source', '````text\n' + finding + '\n````'],
|
|
['literal assertion', finding.replace('The original sketch', '`The original sketch').replace('(see F1).', '(see F1).`')],
|
|
['split unrelated section', finding.replace('Review showed', '\n\n## Another finding\nReview showed')],
|
|
['direct same-finding withdrawal', finding + '\n\nF1 is rejected.'],
|
|
['later named same-finding withdrawal', finding + '\n\n## Assessment of F1\nF1 is withdrawn.'],
|
|
['dismissed defect', finding + '\n\nThis is not a defect; no fix is required.'],
|
|
['accepted stale consequence', finding + '\n\nA subsequent stale read is permitted by the amended contract.'],
|
|
])('rejects %s', (_name, text) => expect(hasStaleFillRaceFinding(text)).toBe(false));
|
|
});
|
|
|
|
// Exact public Write at d30620e8, session 59f999d1-de6b-4c67-ae6c-210efa05f0cc,
|
|
// toolu_01CWeW6wi4V6YEMNtd5dVdz2 acknowledged at native line 3343.
|
|
// PLAN.md SHA-256 577b669134977a17c779765c4a979a5fc1e9bc672cd77948d04d6bf066bfd7ff:
|
|
// retained requirement lines 38-40, current amendment 45-50, earlier-caller allowance 229.
|
|
// The recorded paid failure remains a failure; these are free detector regressions.
|
|
describe('attributed original coordination premise with a current freshness finding', () => {
|
|
const evidence = "- A read already in progress when a write commits may return its earlier DB\n snapshot to that caller. Every read begun after that write completes must\n observe the committed version. TTL expiry is not a substitute for this rule.\n\n## Proposed wrapper integration\nKeep the current read-through repository interface and shared adapters.\n**[Amended: D1]** The original sketch stated \"no additional version checks or\ncoordination between a cache fill and a write\". That is withdrawn: the review\nshowed it violates the freshness invariant above (schedule in Section 4). The\naccepted ordering rules are:";
|
|
const allowance = "Waiters coalesced on R1 receive V1 \u2014 permitted by PLAN.md:27-28 (they began before W completed)";
|
|
const withAllowance = (claim = allowance) => evidence + '\n\n## Assessment of D1\n' + claim + '.';
|
|
test('accepts the captured current assertion against its retained requirement', () => {
|
|
expect(hasStaleFillRaceFinding(evidence)).toBe(true);
|
|
expect(hasStaleFillRaceFinding(withAllowance())).toBe(true);
|
|
});
|
|
test('preserves premise ownership through equivalent labels, quotes and rule wording', () => {
|
|
for (const text of [
|
|
evidence.replaceAll('D1', 'F7'),
|
|
evidence.replace('original sketch', 'original wrapper'),
|
|
evidence.replace('original sketch', 'original implementation'),
|
|
evidence.replace('"no additional', '“no additional').replace('a write"', 'a write”'),
|
|
evidence.replace('no additional version checks or\ncoordination', 'no coordination'),
|
|
evidence.replace('That is withdrawn: the review\nshowed it violates', 'This is withdrawn: the review shows it breaks'),
|
|
evidence.replace('freshness invariant above', 'retained read-after-write contract above'),
|
|
evidence.replace('Every read begun', 'Every read started').replace('that write completes', 'the write returns').replace('committed version', 'committed value'),
|
|
evidence.replace(' (schedule in Section 4)', ''),
|
|
]) expect(hasStaleFillRaceFinding(text)).toBe(true);
|
|
});
|
|
test('requires the original missing coordination and the reviewer current assertion together', () => {
|
|
for (const text of [
|
|
evidence.replace('no additional version checks or\ncoordination', 'coordination'),
|
|
evidence.replace('cache fill and a write', 'cache hit and a read'),
|
|
evidence.replace('original sketch stated', 'original sketch may have stated'),
|
|
evidence.replace('That is withdrawn:', 'That is retained:'),
|
|
evidence.replace('showed it violates', 'did not show it violates'),
|
|
evidence.replace('showed it violates', 'showed another wrapper violates'),
|
|
evidence.replace('showed it violates', 'showed it might violate'),
|
|
evidence.replace('freshness invariant', 'formatting invariant'),
|
|
evidence.replace('Section 4).', 'Section 4)?'),
|
|
evidence.replace('That is withdrawn:', '\n\n## Other finding\nThat is withdrawn:'),
|
|
evidence.replace('That is withdrawn:', '| That is withdrawn:'),
|
|
evidence.replace('**[Amended: D1]**', 'If approved: **[Amended: D1]**'),
|
|
]) expect(hasStaleFillRaceFinding(text)).toBe(false);
|
|
});
|
|
test('requires a retained current requirement from this plan', () => {
|
|
const finding = evidence.slice(evidence.indexOf('## Proposed wrapper integration'));
|
|
for (const text of [
|
|
finding,
|
|
evidence.replace('Every read begun after that write completes must\n observe the committed version.', 'Later reads may observe an earlier value.'),
|
|
evidence.replace('must\n observe', 'might\n observe'),
|
|
evidence.replace('Every read begun', 'Not every read begun'),
|
|
evidence.replace('## Proposed wrapper integration', 'This rule is withdrawn.\n\n## Proposed wrapper integration'),
|
|
'## Finding F9: unrelated cache\n' + evidence,
|
|
'## Historical source\n' + evidence.slice(0, evidence.indexOf('## Proposed wrapper integration')) + '\n## Current review\n' + finding,
|
|
'> Every read begun after that write completes must observe the committed version.\n\n' + finding,
|
|
'"Every read begun after that write completes must observe the committed version."\n\n' + finding,
|
|
'For another cache. Every read begun after that write completes must observe the committed version.\n\n' + finding,
|
|
]) expect(hasStaleFillRaceFinding(text)).toBe(false);
|
|
});
|
|
test('does not promote copied, fenced or quoted review assertions', () => {
|
|
for (const text of [
|
|
'Source:\n\n' + evidence,
|
|
'Earlier review:\n\n' + evidence,
|
|
'## Hypothetical example\n' + evidence,
|
|
'The following is a quoted source excerpt.\n\n' + evidence,
|
|
evidence.split('\n').map(line => '> ' + line).join('\n'),
|
|
evidence.split('\n').map(line => ' ' + line).join('\n'),
|
|
'````text\n' + evidence + '\n````',
|
|
'~~~text\n' + evidence + '\n~~~',
|
|
evidence.replace('**[Amended: D1]**', '"Copied sentence. **[Amended: D1]**') + '"',
|
|
evidence.replace('**[Amended: D1]**', 'Example of report format:\n\n**[Amended: D1]**'),
|
|
evidence.replace('That is withdrawn:', '"That is withdrawn:').replace('Section 4).', 'Section 4)."'),
|
|
evidence.replace('The original sketch', '`The original sketch').replace('Section 4).', 'Section 4).`'),
|
|
]) expect(hasStaleFillRaceFinding(text)).toBe(false);
|
|
});
|
|
test('same-finding rejection and stale-result permission remain failures', () => {
|
|
for (const tail of [
|
|
'D1 is rejected.', 'D1 is "withdrawn".', 'This finding is dismissed.',
|
|
'| D1 | Withdrawn |', '| D1 | "rejected" |',
|
|
'A subsequent stale read is permitted by the amended contract.',
|
|
'This stale-fill behavior is accepted.', 'No coordination is required.',
|
|
]) expect(hasStaleFillRaceFinding(withAllowance() + '\n\n## Assessment of D1\n' + tail)).toBe(false);
|
|
});
|
|
test('foreign or copied rejections do not override the current finding', () => {
|
|
for (const tail of [
|
|
'## Assessment of D2\nD2 is rejected.',
|
|
'## Assessment of D2\n| D2 | Withdrawn |',
|
|
'## Historical assessment\nD1 is withdrawn.',
|
|
'## Assessment of D1\n> D1 is rejected.',
|
|
]) expect(hasStaleFillRaceFinding(evidence + '\n\n' + tail)).toBe(true);
|
|
});
|
|
test('coalesced earlier callers may use different symbolic writer and snapshot names', () => {
|
|
for (const text of [
|
|
allowance.replaceAll('R1', 'R17').replaceAll('V1', 'snapshot-A').replaceAll('W ', 'W9 '),
|
|
allowance.replace('Waiters', 'Readers').replace('began', 'started').replace('W completed', 'the write returned'),
|
|
allowance.replace('Waiters coalesced on R1', 'Callers').replace('receive', 'observe').replace('permitted by PLAN.md:27-28', 'allowed'),
|
|
]) expect(hasStaleFillRaceFinding(withAllowance(text))).toBe(true);
|
|
});
|
|
test('earlier-call allowance cannot credit a later reader, uncertain ordering or a fill', () => {
|
|
for (const text of [
|
|
allowance.replace('before W', 'after W'),
|
|
allowance.replace('they began', 'they never began'),
|
|
allowance.replace('they began', 'they may have begun'),
|
|
allowance.replace('they began', 'another reader began'),
|
|
allowance.replace('before W completed', 'before R2 completed'),
|
|
allowance.replace('Waiters', 'Later readers'),
|
|
allowance.replace('receive V1', 'fill the cache with V1'),
|
|
allowance.replace('receive V1', 'return V1 to later callers'),
|
|
allowance.replace('W completed)', 'W completed only if the write failed)'),
|
|
allowance + ' and later readers may reuse V1',
|
|
allowance + '. A subsequent stale read is permitted',
|
|
allowance + '. The stale cache fill is acceptable',
|
|
]) expect(hasStaleFillRaceFinding(withAllowance(text))).toBe(false);
|
|
});
|
|
});
|
|
|
|
|
|
describe('attributed coordination phrase classes and ownership', () => {
|
|
// Independently written forms: no captured sentence, schedule schema or fixed parenthetical.
|
|
const reports = [
|
|
`## Retained contract
|
|
Any request started after the write returns shall receive the committed value.
|
|
|
|
## Wrapper review
|
|
[Amended: F8] Our original wrapper assumed “cache population proceeds without synchronization with writes”.
|
|
We reject that assumption. Our review established that this approach contradicts the existing freshness guarantee.`,
|
|
`## Existing contract
|
|
All reads that begin after write completion must return the newly committed version.
|
|
|
|
## Implementation review
|
|
[Amended: D4] The proposed implementation specifies "no coordination for writes and cache fills".
|
|
This proposal was rejected. Review found it fails to preserve the read-after-write requirement.`,
|
|
`## Contract retained
|
|
Once a write has completed, new reads must see the value it committed.
|
|
|
|
## Current review
|
|
[Amended: F3] The current sketch states "cache fills and writes run without coordination".
|
|
That sketch breaks the current freshness rule.`,
|
|
`## Required behavior
|
|
The retained requirement: all requests that start after that write finishes are required to receive the committed snapshot.
|
|
|
|
## Current review
|
|
[Amended: D17] Our original implementation assumed "cache repopulation and writes lacked ordering guards".
|
|
That assumption has been retracted. The implementation does not preserve the existing freshness contract.`,
|
|
];
|
|
|
|
for (const [index, report] of reports.entries()) {
|
|
test(`phrase classes recognize independent current review form ${index + 1}`, () => expect(hasStaleFillRaceFinding(report)).toBe(true));
|
|
}
|
|
|
|
test('the attributed premise and direct violation need no fixed rejection sentence or above reference', () => {
|
|
expect(hasStaleFillRaceFinding(reports[0]!.replace('We reject that assumption. ', ''))).toBe(true);
|
|
expect(hasStaleFillRaceFinding(reports[1]!.replace('This proposal was rejected. ', ''))).toBe(true);
|
|
expect(hasStaleFillRaceFinding(reports[2]!.replace('That sketch breaks', 'We found that this sketch violates'))).toBe(true);
|
|
});
|
|
|
|
test('different finding and artifact subjects cannot borrow the attributed premise', () => {
|
|
for (const text of [
|
|
reports[0]!.replace('We reject that assumption.', 'Another unrelated finding concerns replica lag.'),
|
|
reports[0]!.replace('this approach contradicts', 'another approach contradicts'),
|
|
reports[0]!.replace('this approach contradicts', 'the implementation contradicts'),
|
|
reports[1]!.replace('Review found it fails', 'Review found another issue fails'),
|
|
reports[2]!.replace('That sketch breaks', 'It is unclear whether that sketch breaks'),
|
|
reports[2]!.replace('That sketch breaks', 'It is false that that sketch breaks'),
|
|
reports[2]!.replace('That sketch breaks', 'That sketch does not break'),
|
|
]) expect(hasStaleFillRaceFinding(text)).toBe(false);
|
|
});
|
|
|
|
test('normative freshness rules cannot be replaced by conditional or permissive statements', () => {
|
|
for (const report of reports) {
|
|
for (const text of [
|
|
report.replace(/shall receive|must return|must see|are required to receive/, 'may receive'),
|
|
report.startsWith('## Contract retained') ? report.replace('Once a write has completed', 'Before a write has completed') : report.replace('after', 'before'),
|
|
report.replace('committed value', 'earlier value').replace('newly committed version', 'old version').replace('the value it committed', 'an older snapshot').replace('committed snapshot', 'stale snapshot'),
|
|
]) expect(hasStaleFillRaceFinding(text)).toBe(false);
|
|
}
|
|
});
|
|
|
|
test('quoted current assertions, source owners and split findings remain closed across phrase forms', () => {
|
|
for (const report of reports) {
|
|
const assertion = report.slice(report.lastIndexOf('\n') + 1);
|
|
for (const text of [
|
|
'> ' + report.replaceAll('\n', '\n> '),
|
|
'````text\n' + report + '\n````',
|
|
'## Historical source\n' + report.replaceAll('## ', '### '),
|
|
report.replace(assertion, '"' + assertion + '"'),
|
|
report.replace(assertion, '### Unrelated finding F99\n' + assertion),
|
|
]) expect(hasStaleFillRaceFinding(text)).toBe(false);
|
|
}
|
|
});
|
|
|
|
const finding = reports[0]!;
|
|
const withAllowance = (claim: string) => finding + '\n\n## Assessment of F8\n' + claim;
|
|
const earlierAllowances = [
|
|
'The coalesced readers receive their earlier snapshot; that return is permitted because they started before the write completed.',
|
|
'Readers coalesced on R8 return SNAPSHOT_X (allowed, because each call began before W9 returned).',
|
|
'Waiters that began before write completion are permitted to receive V7.',
|
|
];
|
|
|
|
test('earlier-group permission depends on ownership and chronology rather than exact punctuation', () => {
|
|
for (const claim of earlierAllowances) expect(hasStaleFillRaceFinding(withAllowance(claim))).toBe(true);
|
|
});
|
|
|
|
test('earlier-group phrases cannot permit a fill, another caller or uncertain start', () => {
|
|
for (const claim of [
|
|
earlierAllowances[0]!.replace('before the write completed', 'after the write completed'),
|
|
earlierAllowances[0]!.replace('they started', 'another reader started'),
|
|
earlierAllowances[0]!.replace('they started', 'they might have started'),
|
|
earlierAllowances[0]!.replace('receive their earlier snapshot', 'store their earlier snapshot in the cache'),
|
|
earlierAllowances[1]!.replace('each call began', 'some other call began'),
|
|
earlierAllowances[1]!.replace('W9 returned', 'R3 returned'),
|
|
earlierAllowances[2]!.replace('before write completion', 'before another write completed'),
|
|
earlierAllowances[2]!.replace('receive V7', 'return V7 to future callers'),
|
|
earlierAllowances[0]!.replace(/\.$/, ' and future consumers may reuse that snapshot.'),
|
|
]) expect(hasStaleFillRaceFinding(withAllowance(claim))).toBe(false);
|
|
});
|
|
|
|
test('a legitimate earlier group never overrides a same-finding rejection or accepted stale fill', () => {
|
|
for (const allowance of earlierAllowances) {
|
|
expect(hasStaleFillRaceFinding(withAllowance(allowance) + '\nF8 is rejected.')).toBe(false);
|
|
expect(hasStaleFillRaceFinding(withAllowance(allowance) + '\nA later stale read is permitted.')).toBe(false);
|
|
expect(hasStaleFillRaceFinding(withAllowance(allowance) + '\nThis stale-fill behavior is accepted.')).toBe(false);
|
|
expect(hasStaleFillRaceFinding(withAllowance(allowance) + '\n\n## Assessment of F9\nF9 is rejected.')).toBe(true);
|
|
}
|
|
});
|
|
|
|
|
|
test('current assertions cannot attribute the retained-rule violation to a different finding', () => {
|
|
const report = reports[2]!;
|
|
expect(hasStaleFillRaceFinding(report.replace('freshness rule.', 'freshness rule (see F3).'))).toBe(true);
|
|
expect(hasStaleFillRaceFinding(report.replace('freshness rule.', 'freshness rule (see F9).'))).toBe(false);
|
|
expect(hasStaleFillRaceFinding(report.replace('freshness rule.', 'freshness rule (see D3).'))).toBe(false);
|
|
});
|
|
});
|
|
|
|
describe('current fill-lifetime overlap', () => {
|
|
const lifetime = 'a fill that started before a write and stored after it caches the pre-write snapshot';
|
|
const current = (text: string, suffix = '') => `### Current findings\n\nWithout coordination, ${text}, violating the retained read-after-write rule. ${suffix}`;
|
|
|
|
test('captured current lifetime claim is independent of the ambiguous inline schedule', () => {
|
|
expect(hasStaleFillRaceFinding(lifetimeFixture.claim)).toBe(true);
|
|
expect(hasStaleFillRaceFinding(lifetimeFixture.ambiguousFinding)).toBe(false);
|
|
});
|
|
|
|
test.each([
|
|
lifetime,
|
|
'the cache fill which began before the write and completed after that write, storing the old value',
|
|
'a fill that begins before a write and finishes after the same write stores the stale data',
|
|
'the original fill starts before this write and completes after that same write caches the pre-write snapshot',
|
|
'the fill began before the write completes and stored after it settles caches the pre-write value',
|
|
])('recognizes an explicit same-fill lifetime: %s', text => {
|
|
expect(hasStaleFillRaceFinding(current(text))).toBe(true);
|
|
});
|
|
|
|
test.each(['F1', 'R7', 'BUG-cache', '17'])('current ownership does not depend on row-ID spelling: %s', id => {
|
|
expect(hasStaleFillRaceFinding(`### Current findings\n\n| ${id} | CRITICAL GAP | ${current(lifetime).split('\n\n')[1]} |`)).toBe(true);
|
|
});
|
|
|
|
test.each([
|
|
['starts after the write', lifetime.replace('started before', 'started after')],
|
|
['stores before the write', lifetime.replace('stored after', 'stored before')],
|
|
['different writer', lifetime.replace('after it', 'after another write')],
|
|
['different filling actor', lifetime.replace('and stored', 'and another fill stored')],
|
|
['foreign key', lifetime.replace('a fill', 'a fill for another key')],
|
|
['return to original caller only', lifetime.replace('stored after it caches', 'returned after it with')],
|
|
['fresh value', lifetime.replace('pre-write snapshot', 'committed snapshot')],
|
|
['missing start', lifetime.replace('that started before a write and ', '')],
|
|
['missing late storage', lifetime.replace('and stored after it ', '')],
|
|
['missing cache storage', lifetime.replace('caches the pre-write snapshot', 'returns the pre-write snapshot')],
|
|
['explicit conditional', 'if ' + lifetime],
|
|
['explicit hypothesis', 'hypothetical execution: ' + lifetime],
|
|
['possible execution only', 'it might be that ' + lifetime],
|
|
['negated execution', 'it is not true that ' + lifetime],
|
|
['prevented execution', 'the guard prevents ' + lifetime],
|
|
['impossible execution', 'it is impossible that ' + lifetime],
|
|
['quoted execution', '"' + lifetime + '"'],
|
|
['code literal execution', '`' + lifetime + '`'],
|
|
['quote cannot join phase fragments', lifetime.replace('and stored after it', 'and "stored after it"')],
|
|
['second subject cannot inherit write', lifetime.replace('after it', 'after a separate write')],
|
|
])('rejects incomplete or unasserted overlap: %s', (_name, text) => {
|
|
expect(hasStaleFillRaceFinding(current(text))).toBe(false);
|
|
});
|
|
|
|
test.each([
|
|
['quoted block', current(lifetime).split('\n').map(line => '> ' + line).join('\n')],
|
|
['fenced block', '```text\n' + current(lifetime) + '\n```'],
|
|
['unclosed fence', '~~~text\n' + current(lifetime)],
|
|
['source heading', current(lifetime).replace('Current findings', 'Quoted source')],
|
|
['historical owner', current(lifetime).replace('Current findings', 'Historical review')],
|
|
['source introduction', 'Source:\n' + current(lifetime).replace('Current findings', 'Findings')],
|
|
['accepted staleness', current(lifetime, 'This staleness is the accepted consistency model.')],
|
|
['later stale read permitted', current(lifetime, 'Later stale reads are permitted by the contract.')],
|
|
['current prevention', current(lifetime, 'The current wrapper cannot refill old data after invalidation.')],
|
|
['finding withdrawn', current(lifetime, 'This finding is withdrawn.')],
|
|
['finding rejected', current(lifetime, 'This finding is rejected.')],
|
|
['no repair required', current(lifetime, 'No guard is required.')],
|
|
])('preserves current ownership and dismissal: %s', (_name, text) => {
|
|
expect(hasStaleFillRaceFinding(text)).toBe(false);
|
|
});
|
|
|
|
const seeded = () => {
|
|
const events: string[] = [];
|
|
let cached: string | undefined;
|
|
let releaseRead!: () => void;
|
|
const pendingRead = new Promise<string>(resolve => { releaseRead = () => { events.push('DB read v1 completes'); resolve('v1'); }; });
|
|
const cache = {
|
|
get: () => cached,
|
|
set: (_key: string, value: string) => { events.push('cache set ' + value); cached = value; },
|
|
delete: () => { events.push('cache delete'); cached = undefined; },
|
|
};
|
|
const repository = {
|
|
read: () => pendingRead,
|
|
write: async () => { events.push('DB write commits v2'); return 'v2'; },
|
|
};
|
|
const functions = new Function('cache', 'repository', CACHE_READ_WRITE_SKETCH + '\nreturn { readProfile, writeProfile };')(cache, repository);
|
|
return { events, releaseRead, ...functions } as { events: string[]; releaseRead: () => void; readProfile: (key: string) => Promise<string>; writeProfile: (key: string, update: unknown) => Promise<string> };
|
|
};
|
|
|
|
test('the actual seeded wrapper can fill a pre-write snapshot when its pending read completes after the writer', async () => {
|
|
const fixture = seeded();
|
|
const original = fixture.readProfile('profile');
|
|
await fixture.writeProfile('profile', {});
|
|
fixture.releaseRead();
|
|
expect(await original).toBe('v1');
|
|
expect(await fixture.readProfile('profile')).toBe('v1');
|
|
expect(fixture.events).toEqual(['DB write commits v2', 'cache delete', 'DB read v1 completes', 'cache set v1']);
|
|
});
|
|
|
|
test('the seeded await continuation stores synchronously before a later writer invalidates it', async () => {
|
|
const fixture = seeded();
|
|
const original = fixture.readProfile('profile');
|
|
fixture.releaseRead();
|
|
expect(await original).toBe('v1');
|
|
await fixture.writeProfile('profile', {});
|
|
expect(fixture.events).toEqual(['DB read v1 completes', 'cache set v1', 'DB write commits v2', 'cache delete']);
|
|
});
|
|
});
|
|
|
|
|
|
describe('section fixture rollout metrics retain final acceptance without an impossible early-stage gate',()=>{
|
|
test('early-stage hit rate counts admitted requests while aggregate metrics use baseline limits',()=>{
|
|
const requests=9000, admitted=requests*0.1, hits=admitted*0.6;
|
|
const cohortHitRate=hits/admitted, aggregateHitRate=hits/requests;
|
|
expect(cohortHitRate).toBe(0.6); expect(aggregateHitRate).toBe(0.06);
|
|
expect(70*(1-aggregateHitRate)).toBeCloseTo(65.8); // uniform traffic, linear read CPU: above final 50%, below baseline 70%
|
|
expect(CEO_SECTION_CACHE_PLAN).toContain('among requests admitted to the cache path');
|
|
expect(CEO_SECTION_CACHE_PLAN).toContain('tracked separately, not as misses');
|
|
expect(CEO_SECTION_CACHE_PLAN).toContain('DB CPU and read p95 are service-wide metrics, including bypassed requests');
|
|
expect(CEO_SECTION_CACHE_PLAN).toContain('At the 10% and 50% stages');
|
|
expect(CEO_SECTION_CACHE_PLAN).toContain('no worse than their 70%/120 ms pre-rollout baselines');
|
|
expect(CEO_SECTION_CACHE_PLAN).not.toContain('A healthy hour means the stated hit-rate, CPU, latency and error targets hold');
|
|
});
|
|
test('full rollout keeps all original absolute targets and the seeded race still needs repair',()=>{
|
|
expect(CEO_SECTION_CACHE_PLAN).toContain('At 100%, the original');
|
|
expect(CEO_SECTION_CACHE_PLAN).toContain('absolute acceptance targets (DB CPU below 50%, read p95 below 60 ms, hits at');
|
|
expect(CEO_SECTION_CACHE_PLAN).toContain('least 60%) must all hold with unchanged correctness/error SLOs and no alerts');
|
|
expect(CEO_SECTION_CACHE_PLAN).toContain('Every read begun after that write completes must');
|
|
expect(CEO_SECTION_CACHE_PLAN).toContain('TTL expiry is not a substitute for this rule');
|
|
expect(CEO_SECTION_CACHE_PLAN).toContain('no additional version checks or');
|
|
expect(CEO_SECTION_CACHE_PLAN).toContain(CACHE_READ_WRITE_SKETCH);
|
|
expect(CEO_SECTION_CACHE_PLAN).toContain('repository.read returns\n an immutable absent-result DTO for a missing record, never undefined');
|
|
});
|
|
});
|
|
|
|
describe('sdk-columnar-af', () => {
|
|
const captured = captured_sdk_columnar_af;
|
|
const evidence = () => captured.retryFinding + '\n\n' + captured.retrySchedule;
|
|
function replace(text: string, before: string, after: string) {
|
|
expect(text).toContain(before);
|
|
return text.replace(before, after);
|
|
}
|
|
|
|
test('AF exact retry columnar schedule establishes a post-write stale reader', () => {
|
|
expect(hasStaleFillRaceFinding(captured.retryReport)).toBe(true);
|
|
expect(hasStaleFillRaceFinding(captured.firstGuardedEvidence)).toBe(false);
|
|
});
|
|
|
|
test('AF columnar evidence binds named actors, keys and distinct versions independently of their spelling', () => {
|
|
expect(hasStaleFillRaceFinding(evidence())).toBe(true);
|
|
const varied = evidence().replace(/\bR1\b/g, 'R4').replace(/\bR2\b/g, 'R8').replace(/\bW\b/g, 'W3')
|
|
.replace(/\bK\b/g, 'profileKey').replace(/\bv1\b/g, 'oldVersion').replace(/\bv2\b/g, 'newVersion')
|
|
.replace(/->/g, '→');
|
|
expect(hasStaleFillRaceFinding(varied)).toBe(true);
|
|
expect(hasStaleFillRaceFinding(evidence().replace(/\bv2\b/g, 'v1'))).toBe(false);
|
|
});
|
|
|
|
test('AF every ordered operation and version witness is required', () => {
|
|
for (const [before, after] of [
|
|
['get(K) -> undefined', 'get(K) -> v1'],
|
|
['await repository.read -> v1', 'await repository.read -> v2'],
|
|
['await write commits v2', 'await write fails'],
|
|
['delete(K) (no entry)', 'keep(K)'],
|
|
['| returns |', '| still pending |'],
|
|
['resume: set(K, v1); return v1', 'resume: set(K, v2); return v2'],
|
|
['get(K) -> v1; return v1', 'get(K) -> v2; return v2'],
|
|
['v1 STALE | v2', 'v1 STALE | v1'],
|
|
['6 | resume:', '8 | resume:'],
|
|
]) expect(hasStaleFillRaceFinding(replace(evidence(), before!, after!))).toBe(false);
|
|
for (let event = 1; event <= 7; event++) {
|
|
expect(hasStaleFillRaceFinding(evidence().split('\n').filter(line => !line.trim().startsWith(`${event} |`)).join('\n'))).toBe(false);
|
|
}
|
|
});
|
|
|
|
test('AF a different key, reader, write or column cannot lend ownership', () => {
|
|
for (const [before, after] of [
|
|
['cache[K] | DB[K]', 'cache[K] | DB[J]'],
|
|
['delete(K) (no entry)', 'delete(J) (no entry)'],
|
|
['set(K, v1)', 'set(J, v1)'],
|
|
['get(K) -> v1; return v1', 'get(J) -> v1; return v1'],
|
|
['R2 read (begins after W)', 'R1 read (begins after W)'],
|
|
['R2 read (begins after W)', 'R2 read (begins after W2)'],
|
|
['R2 began after W completed (t5)', 'R1 began after W completed (t5)'],
|
|
['R2 began after W completed (t5)', 'R2 began before W completed (t5)'],
|
|
['R2 began after W completed (t5)', 'R2 began after W completed (t6)'],
|
|
['observes v1 for up to 30 s.', 'observes v2 for up to 30 s.'],
|
|
]) expect(hasStaleFillRaceFinding(replace(evidence(), before!, after!))).toBe(false);
|
|
});
|
|
|
|
test('AF current declarative execution cannot borrow a conditional, negated or quoted schedule', () => {
|
|
for (const [before, after] of [
|
|
['await write commits v2', 'write might commit v2'],
|
|
['resume: set(K, v1); return v1', 'resume: no set(K, v1); return v1'],
|
|
['VIOLATION t7:', 'If VIOLATION t7:'],
|
|
['VIOLATION t7:', 'Quoted VIOLATION t7:'],
|
|
['observes v1 for up to 30 s.', 'observes v1 for up to 30 s.?'],
|
|
]) expect(hasStaleFillRaceFinding(replace(evidence(), before!, after!))).toBe(false);
|
|
});
|
|
|
|
test('AF the same named finding and a real top-level fence own the schedule', () => {
|
|
const text = evidence();
|
|
for (const value of [
|
|
captured.retrySchedule,
|
|
text.replace('(F1 evidence)', '(F2 evidence)'),
|
|
text.replace('Schedule S1 below', 'Schedule S2 below'),
|
|
captured.retryFinding + '\n' + text,
|
|
text.split('\n').map(line => '> ' + line).join('\n'),
|
|
'````text\n' + text + '\n````',
|
|
text.replace('```\n t |', '```javascript\n t |'),
|
|
text.slice(0, text.lastIndexOf('```')),
|
|
'Example:\n\n' + text,
|
|
captured.retryFinding + '\n\nTemplate:\n' + captured.retrySchedule,
|
|
]) expect(hasStaleFillRaceFinding(value)).toBe(false);
|
|
});
|
|
|
|
test('AF an original-caller allowance cannot excuse a stale cache or later caller', () => {
|
|
const allowed = 'Allowed by contract: R1 itself returns v1 (read in progress when write committed).';
|
|
for (const value of [
|
|
replace(evidence(), allowed, 'Allowed by contract: R2 itself returns v1 (read in progress when write committed).'),
|
|
replace(evidence(), allowed, 'The stale-fill behavior is accepted.'),
|
|
replace(evidence(), allowed, 'There is no stale-fill race.'),
|
|
replace(evidence(), allowed, 'The trace is impossible.'),
|
|
evidence() + '\n\nThis is not a violation. No guard is required.',
|
|
]) expect(hasStaleFillRaceFinding(value)).toBe(false);
|
|
});
|
|
test('AF every same-row assessment and an unproven source frame remain authoritative', () => {
|
|
for (const [cell, value] of [
|
|
[6, 'Rejected: there is no stale-fill race.'],
|
|
[5, 'The stale-fill behavior is accepted. No guard is required.'],
|
|
[6, 'Rejected: “There is no stale-fill race.”'],
|
|
[6, 'Rejected: "The stale-fill behavior is accepted. No guard is required."'],
|
|
] as const) {
|
|
const cells = captured.retryFinding.split('|'); cells[cell] = value;
|
|
expect(hasStaleFillRaceFinding(cells.join('|') + '\n\n' + captured.retrySchedule)).toBe(false);
|
|
}
|
|
expect(hasStaleFillRaceFinding('An unproven hypothesis:\n\n' + evidence())).toBe(false);
|
|
});
|
|
});
|
|
|
|
describe('sdk-compact-sequence-aj', () => {
|
|
const captured = captured_sdk_compact_sequence_aj;
|
|
const sequence = 'fill starts, write commits, write deletes (no-op), fill sets pre-commit v1, later read hits v1.';
|
|
const report = captured.finding;
|
|
|
|
test('recognizes the captured current original-plan sequence without borrowing the amended diagram', () => {
|
|
expect(hasStaleFillRaceFinding(report)).toBe(true);
|
|
expect(hasStaleFillRaceFinding(report.replaceAll('v1', 'snapshot_A'))).toBe(true);
|
|
expect(hasStaleFillRaceFinding(report.replace('Schedule Diagram 2b: ', ''))).toBe(true);
|
|
});
|
|
|
|
test('requires the ordered original fill, commit, invalidation, old cache value and same later value', () => {
|
|
for (const changed of [
|
|
sequence.replace('fill starts, ', ''),
|
|
sequence.replace('write commits, ', ''),
|
|
sequence.replace('write deletes (no-op), ', ''),
|
|
sequence.replace('fill sets pre-commit v1, ', ''),
|
|
sequence.replace(', later read hits v1', ''),
|
|
sequence.replace('later read hits v1', 'later read hits v2'),
|
|
sequence.replace('pre-commit v1', 'post-commit v1'),
|
|
sequence.replace('fill starts, write commits', 'write commits, fill starts'),
|
|
sequence.replace('write deletes (no-op), fill sets pre-commit v1', 'fill sets pre-commit v1, write deletes (no-op)'),
|
|
sequence.replace('later read hits', 'another key later read hits'),
|
|
]) expect(hasStaleFillRaceFinding(report.replace(sequence, changed))).toBe(false);
|
|
expect(hasStaleFillRaceFinding(report.replace('Original plan', 'Amended plan'))).toBe(false);
|
|
expect(hasStaleFillRaceFinding(report.replace(sequence, '"' + sequence + '"'))).toBe(false);
|
|
});
|
|
|
|
test('preserves accepted-staleness and explicit dismissal boundaries', () => {
|
|
for (const suffix of [
|
|
'This is not a gap; no guard is needed.',
|
|
'This staleness is the accepted consistency model.',
|
|
'This finding is withdrawn.',
|
|
'F1 is rejected.',
|
|
]) expect(hasStaleFillRaceFinding(report.trimEnd() + '\n\n' + suffix)).toBe(false);
|
|
expect(hasStaleFillRaceFinding(report.replace('fill starts', 'fill never starts'))).toBe(false);
|
|
});
|
|
|
|
test('source, quotes and hypothetical framing cannot supply current coverage', () => {
|
|
for (const text of [
|
|
'```text\n' + report + '```',
|
|
report.split('\n').map(line => '> ' + line).join('\n'),
|
|
report.split('\n').map(line => ' ' + line).join('\n'),
|
|
'## Historical example\n\n' + report,
|
|
'## Quoted source\n\n' + report,
|
|
'An unproven hypothesis.\n\n' + report,
|
|
'The following is a hypothetical example.\n\n' + report,
|
|
]) expect(hasStaleFillRaceFinding(text)).toBe(false);
|
|
expect(hasStaleFillRaceFinding('## Historical example\nOld material.\n\n## Current review\n' + report)).toBe(true);
|
|
expect(hasStaleFillRaceFinding(report + '\n## Unrelated issue\nF2 is rejected.')).toBe(true);
|
|
});
|
|
|
|
test('source framing remains attached to descendant registry headings', () => {
|
|
for (const prefix of [
|
|
'## Copied material\nThe following subsections reproduce source examples, not current findings.\n\n',
|
|
'## Input material\nThe following sections quote historical examples.\n\n',
|
|
'The following subsections reproduce source examples, not current findings.\n\n',
|
|
]) expect(hasStaleFillRaceFinding(prefix + report)).toBe(false);
|
|
expect(hasStaleFillRaceFinding('## Source notes\nThe following material quotes historical examples.\n\n## Current findings\n' + report)).toBe(true);
|
|
});
|
|
|
|
test('same finding assessments retain identity across sections and unrelated findings', () => {
|
|
for (const suffix of [
|
|
'## F1 assessment\nThis finding is withdrawn.',
|
|
'## Final assessment\nF1 is rejected.',
|
|
'## F2\nUnrelated issue accepted.\n\n## Final assessment\nF1 is dismissed.',
|
|
]) expect(hasStaleFillRaceFinding(report + '\n\n' + suffix)).toBe(false);
|
|
for (const suffix of [
|
|
'## F2 assessment\nThis finding is withdrawn.',
|
|
'## Final assessment\nF2 is rejected.',
|
|
'## Quoted source\nF1 is rejected.',
|
|
'## Source notes\nThe following subsections quote historical examples.\n\n### F1 assessment\nThis finding is withdrawn.',
|
|
]) expect(hasStaleFillRaceFinding(report + '\n\n' + suffix)).toBe(true);
|
|
});
|
|
});
|
|
|
|
describe('sdk-order-b-ag', () => {
|
|
const captured = captured_sdk_order_b_ag;
|
|
const compactFirst = () => `${captured.first.finding}\n\n${captured.first.heading}\n\`\`\`\n${captured.first.trace}\n\`\`\``;
|
|
const compactRetry = () => `${captured.retry.finding}\n\n${captured.retry.heading}\n\`\`\`\n${captured.retry.trace}\n\`\`\``;
|
|
|
|
test('actual first completed report proves a later stale cache hit', () => {
|
|
expect(hasStaleFillRaceFinding(captured.first.report)).toBe(true);
|
|
});
|
|
|
|
test('isolated Order B proves a later stale cache hit', () => {
|
|
expect(hasStaleFillRaceFinding(compactFirst())).toBe(true);
|
|
});
|
|
|
|
test('actual retry and its explicit original-sketch override establish the unsafe execution', () => {
|
|
expect(hasStaleFillRaceFinding(captured.retry.report)).toBe(true);
|
|
expect(hasStaleFillRaceFinding(compactRetry())).toBe(true);
|
|
});
|
|
|
|
function replaceOnce(text: string, before: string, after: string): string {
|
|
expect(text.includes(before)).toBe(true);
|
|
return text.replace(before, after);
|
|
}
|
|
|
|
test('original-caller return or flight joining alone cannot supply the later cache reader', () => {
|
|
for (const [before, after] of [
|
|
[' Order B: R2 begins after t5 -> cache hit v1 VIOLATION (until TTL or next write)\n', ''],
|
|
['Order B: R2 begins after t5', 'Order B: R1 begins after t5'],
|
|
['Order B: R2 begins after t5', 'Order B: R2 begins before t3'],
|
|
['Order B: R2 begins after t5', 'Order B: R2 begins after t2'],
|
|
['cache hit v1 VIOLATION', 'fresh DB read v2'],
|
|
['cache hit v1 VIOLATION', 'cache hit v2 SAFE'],
|
|
]) expect(hasStaleFillRaceFinding(replaceOnce(compactFirst(), before!, after!))).toBe(false);
|
|
});
|
|
|
|
test('all read, commit, invalidation and late-fill operations retain shared key and version ownership', () => {
|
|
for (const [before, after] of [
|
|
['R1 readProfile(k)', 'R1 readProfile(other)'],
|
|
['W writeProfile(k, v2)', 'W writeProfile(other, v2)'],
|
|
['R2 readProfile(k)', 'R2 readProfile(other)'],
|
|
['cache[k]', 'cache[other]'],
|
|
['inflight[k]', 'inflight[other]'],
|
|
['set(k, v1)', 'set(other, v1)'],
|
|
['set(k, v1)', 'set(k, v2)'],
|
|
['read resolves v1; set(k, v1)', 'read resolves v2; set(k, v1)'],
|
|
['miss; flight f1; await read', 'cache hit v1; return'],
|
|
['await write ... commit v2', 'await write ... abort'],
|
|
['delete(k) no-op; return', 'delete(other) no-op; return'],
|
|
['delete(k) no-op; return', 'write still pending'],
|
|
['read resolves v1; set(k, v1)', 'read resolves v1; return to R1 only'],
|
|
['v1 BAD | -', 'v2 SAFE | -'],
|
|
]) expect(hasStaleFillRaceFinding(replaceOnce(compactFirst(), before!, after!))).toBe(false);
|
|
});
|
|
|
|
test('quoted, conditional and impossible schedules are not actual asserted execution', () => {
|
|
const report = compactFirst();
|
|
for (const changed of [
|
|
report.split('\n').map(line => `> ${line}`).join('\n'),
|
|
`\`\`\`markdown\n${report}\n\`\`\``,
|
|
`An unproven hypothesis:\n${report}`,
|
|
replaceOnce(report, 'Schedule below shows', 'An unproven hypothesis: Schedule below shows'),
|
|
replaceOnce(report, 'Order B: R2 begins', 'Order B: If R2 begins'),
|
|
replaceOnce(report, 'Order B: R2 begins', 'Order B: R2 never begins'),
|
|
replaceOnce(report, 'Order B: R2 begins after t5 -> cache hit v1 VIOLATION', 'Order B: R2 begins after t5 -> cache hit v1 VIOLATION?'),
|
|
report + '\nThis trace is impossible.',
|
|
report + '\n\nThe trace is impossible.',
|
|
]) expect(hasStaleFillRaceFinding(changed)).toBe(false);
|
|
});
|
|
|
|
test('one finding owns the original trace and every same-row assessment', () => {
|
|
const report = compactFirst();
|
|
for (const changed of [
|
|
replaceOnce(report, 'Async schedule (F1)', 'Async schedule (F9)'),
|
|
replaceOnce(report, '| F1 |', '| F9 |'),
|
|
replaceOnce(report, 'Fills overlapping a write are not cached (bounded hit-rate cost, visible in metric)', 'There is no stale-fill race.'),
|
|
replaceOnce(report, 'Fills overlapping a write are not cached (bounded hit-rate cost, visible in metric)', 'Rejected: "There is no stale-fill race."'),
|
|
replaceOnce(report, 'D3: single-flight `invalidate(key)` before and after the write; invalidated fills never `set`; `fill_discarded` metric', 'The stale-fill behavior is accepted. No guard is required.'),
|
|
]) expect(hasStaleFillRaceFinding(changed)).toBe(false);
|
|
});
|
|
|
|
test('retry amendment alone and unasserted original-sketch annotations cannot prove a stale fill', () => {
|
|
const original = 'Original sketch: step 6 fills v1 after step 4 → R2 hits v1 → VIOLATION (S1).';
|
|
for (const replacement of [
|
|
'',
|
|
`"${original}"`,
|
|
`> ${original}`,
|
|
`If ${original}`,
|
|
`Example: ${original}`,
|
|
original.replace('fills v1', 'does not fill v1'),
|
|
original.replace('VIOLATION (S1).', 'VIOLATION (S1)?'),
|
|
original.replace('fills v1', 'fills v2'),
|
|
original.replace('after step 4', 'before step 4'),
|
|
original.replace('after step 4', 'after step 3'),
|
|
original.replace('step 6 fills', 'step 7 fills'),
|
|
original.replace('R2 hits v1', 'R1 receives v1'),
|
|
original.replace('R2 hits v1', 'R2 hits v2'),
|
|
original.replace('(S1)', '(S9)'),
|
|
]) expect(hasStaleFillRaceFinding(replaceOnce(compactRetry(), original, replacement))).toBe(false);
|
|
expect(hasStaleFillRaceFinding(compactRetry() + '\n\nThe trace is impossible.')).toBe(false);
|
|
});
|
|
|
|
test('retry original override is bound to the same actors, cancelled token and completed write', () => {
|
|
for (const [before, after] of [
|
|
['R1 read (began before commit)', 'R1 read (began after commit)'],
|
|
['R2 read (began after W resolves)', 'R2 read (began before W resolves)'],
|
|
['R2 read (began after W resolves)', 'R2 read (other key, began after W resolves)'],
|
|
['invalidate: cancel t1, detach, delete', 'invalidate: cancel other, detach, delete'],
|
|
['writeProfile resolves (write "complete")', 'writeProfile still pending'],
|
|
['await repo.write → v2 committed', 'await repo.write → aborted'],
|
|
['read resolves v1; t1✗ → no fill', 'read resolves v2; t1✗ → no fill'],
|
|
['read resolves v1; t1✗ → no fill', 'read resolves v1; other✗ → no fill'],
|
|
['S1: R1 misses, W commits and deletes, R1 fills stale v1, R2 hits v1.', 'S1: R1 misses, W commits and deletes, R1 fills stale v1, R1 receives v1.'],
|
|
['### 4. Async schedule (F1)', '### 4. Async schedule (F9)'],
|
|
]) expect(hasStaleFillRaceFinding(replaceOnce(compactRetry(), before!, after!))).toBe(false);
|
|
});
|
|
|
|
test('consistent actor, key, version and pending-identity renaming preserves each causal proof', () => {
|
|
const names: Record<string, string> = {
|
|
R1: 'R7', R2: 'R8', W: 'W9', k: 'profile_key', v1: 'oldValue', v2: 'newValue',
|
|
f1: 'flight_old', f2: 'flight_new', t1: 'token_old', t2: 'token_new',
|
|
};
|
|
for (const report of [compactFirst(), compactRetry()]) {
|
|
const renamed = report.replace(/\b(?:R1|R2|W|k|v1|v2|f1|f2|t1|t2)\b/g, token => names[token]!);
|
|
expect(hasStaleFillRaceFinding(renamed)).toBe(true);
|
|
}
|
|
});
|
|
|
|
test('current findings cannot borrow assertion authority from a hypothetical preceding frame', () => {
|
|
for (const report of [compactFirst(), compactRetry()]) {
|
|
for (const prefix of ['An unproven hypothesis.', 'Historical example only.', 'The following is a hypothetical example.']) {
|
|
expect(hasStaleFillRaceFinding(`${prefix}\n\n${report}`)).toBe(false);
|
|
}
|
|
expect(hasStaleFillRaceFinding(`## Prior example\nA completed historical illustration.\n\n## Current findings\n${report}`)).toBe(true);
|
|
}
|
|
});
|
|
});
|
|
|
|
describe('sdk-ordered-schedule-ar', () => {
|
|
const fs = fs_sdk_ordered_schedule_ar;
|
|
const report = fs.readFileSync(new URL('./fixtures/sdk-ordered-schedule-ar.md', import.meta.url), 'utf8');
|
|
const row = report.split('\n').find(line => line.startsWith('| F1 |'))!;
|
|
const schedule = 'Schedule: read misses, write commits and deletes (no-op), read resolves and stores the pre-write snapshot. A later read hits the stale value';
|
|
|
|
test('an actual review supplies the stale-fill ordering without a concurrency keyword', () => {
|
|
expect(row).toContain(schedule);
|
|
expect(row).not.toMatch(/\b(?:race|concurrent|in-flight|pending)\b/i);
|
|
expect(hasStaleFillRaceFinding(row)).toBe(true);
|
|
expect(hasStaleFillRaceFinding(report)).toBe(true);
|
|
expect(hasStaleFillRaceFinding(row.replace('Original sketch', 'Original wrapper'))).toBe(true);
|
|
expect(hasStaleFillRaceFinding(row.replace('pre-write snapshot', 'old value'))).toBe(true);
|
|
expect(hasStaleFillRaceFinding(row.replace('A later read', 'The subsequent read'))).toBe(true);
|
|
expect(hasStaleFillRaceFinding(row.replaceAll('"', ''))).toBe(true);
|
|
});
|
|
|
|
test('every operation and the stale value observed by a later read are required', () => {
|
|
for (const [from, to] of [
|
|
['read misses, ', ''],
|
|
['write commits and deletes (no-op), ', ''],
|
|
['write commits and deletes', 'write rolls back and deletes'],
|
|
['write commits and deletes', 'write commits without deleting'],
|
|
['read resolves and stores the pre-write snapshot', 'read resolves and skips the fill'],
|
|
['read resolves and stores the pre-write snapshot', 'read resolves and stores the fresh snapshot'],
|
|
['A later read hits the stale value', 'The original read returns its own pre-write snapshot'],
|
|
['A later read hits the stale value', 'A later read hits the fresh value'],
|
|
['write commits and deletes (no-op), read resolves and stores the pre-write snapshot', 'read resolves and stores the pre-write snapshot, write commits and deletes (no-op)'],
|
|
['write commits and deletes (no-op), read resolves', 'write commits and deletes (no-op) | read resolves'],
|
|
['read resolves and stores', 'another reader resolves and stores'],
|
|
['write commits and deletes (no-op)', 'write commits and deletes another key'],
|
|
]) {
|
|
expect(row).toContain(from);
|
|
expect(hasStaleFillRaceFinding(row.replace(from, to))).toBe(false);
|
|
}
|
|
});
|
|
|
|
test('copied, conditional, quoted and hypothetical schedules cannot supply current evidence', () => {
|
|
for (const text of [
|
|
'> ' + row,
|
|
'```text\n' + row + '\n```',
|
|
'## Historical example\n' + row,
|
|
'Source:\n' + row,
|
|
'Earlier review:\n' + row,
|
|
row.replace('Original sketch fills', 'Original sketch source excerpt only: fills'),
|
|
row.replace('Original sketch fills', 'Original sketch from an earlier review fills'),
|
|
row.replace('Schedule:', '\nFinding F2. Schedule:'),
|
|
'## Source notes\nThe following material is copied from a template.\n' + row,
|
|
row.replace('Original sketch', 'Quoted original sketch'),
|
|
row.replace('Schedule: read misses', 'Schedule: if a read misses'),
|
|
row.replace('Schedule: read misses', 'Hypothetical schedule: read misses'),
|
|
row.replace('read resolves and stores', 'read never resolves and stores'),
|
|
row.replace(schedule, '"' + schedule + '"'),
|
|
row.replace(schedule, '`' + schedule + '`'),
|
|
row.replace('read resolves and stores the pre-write snapshot', '`read resolves and stores the pre-write snapshot`'),
|
|
row.replace('Flag flip mid-read has the same shape.', 'This sequence is impossible.'),
|
|
]) expect(hasStaleFillRaceFinding(text)).toBe(false);
|
|
});
|
|
|
|
test('a current dismissal stays a dismissal even when the original schedule is complete', () => {
|
|
for (const suffix of [
|
|
'F1 is withdrawn.',
|
|
'F1 is "withdrawn".',
|
|
'F1 is “withdrawn”.',
|
|
'F1 is rejected.',
|
|
'This finding is dismissed.',
|
|
'This is not a bug; no fix is needed.',
|
|
'The stale-fill behavior is permitted.',
|
|
]) expect(hasStaleFillRaceFinding(row + '\n\n' + suffix)).toBe(false);
|
|
expect(hasStaleFillRaceFinding(row + '\n\nF2 is rejected.')).toBe(true);
|
|
expect(hasStaleFillRaceFinding('## Historical example\nOld material.\n\n## Current findings\n' + row)).toBe(true);
|
|
});
|
|
});
|
|
|
|
describe('sdk-ordering-ae', () => {
|
|
const fixture = fixture_sdk_ordering_ae;
|
|
const found = hasStaleFillRaceFinding;
|
|
const trace = fixture.f1.split('|')[4]!.trim();
|
|
function withTrace(value: string): string {
|
|
const cells = fixture.f1.split('|');
|
|
cells[4] = ` ${value} `;
|
|
return cells.join('|');
|
|
}
|
|
|
|
test('actual completed F1 report row supplies ordered stale-fill evidence without a race keyword', () => {
|
|
expect(found(fixture.f1)).toBe(true);
|
|
expect(fixture.provenance.historicalOutcome).toContain('timeout480032ms');
|
|
expect(trace).not.toMatch(/\b(?:race|in-flight|concurrent|pending)\b/i);
|
|
expect(found(`F1 — P1: ${trace}`)).toBe(true);
|
|
});
|
|
|
|
test('ordering evidence requires miss, committed invalidation, stale refill and later stale readers', () => {
|
|
for (const value of [
|
|
'Reader fills the pre-commit snapshot; write commits and deletes; read misses; every later reader sees stale data.',
|
|
'Read misses; reader then fills the pre-commit snapshot; write commits and deletes; every later reader sees stale data.',
|
|
'Write commits and deletes; read misses; reader then fills the pre-commit snapshot; every later reader sees stale data.',
|
|
'Read misses; reader then fills the pre-commit snapshot; every later reader sees stale data.',
|
|
'Read misses, write commits; reader then fills the pre-commit snapshot; every later reader sees stale data.',
|
|
'Read misses, write commits and deletes; every later reader sees stale data.',
|
|
'Read misses, write commits and deletes; reader then fills the post-commit snapshot; every later reader sees fresh data.',
|
|
'Read misses, write commits and deletes; the original reader returns its pre-commit snapshot to its own caller; every later reader sees fresh data.',
|
|
]) expect(found(withTrace(value))).toBe(false);
|
|
});
|
|
|
|
test('explicit other cache, key or reader references cannot borrow the anonymous same-read trace', () => {
|
|
for (const value of [
|
|
'Read misses cache A, write commits and deletes cache B, reader then fills cache A with the pre-commit snapshot; every later reader sees stale data in cache A.',
|
|
'Read misses key u1, write commits and deletes key u2, reader then fills key u1 with the pre-commit snapshot; every later reader sees stale data for key u1.',
|
|
'Read R1 misses, write commits and deletes, reader R2 then fills the pre-commit snapshot; every later reader sees stale data.',
|
|
]) expect(found(withTrace(value))).toBe(false);
|
|
});
|
|
|
|
test('hypothetical, negated and unestablished traces do not assert a current defect', () => {
|
|
for (const value of [
|
|
`If ${trace[0]!.toLowerCase()}${trace.slice(1)}`,
|
|
`A hypothetical example: ${trace}`,
|
|
`An unproven hypothesis: ${trace}`,
|
|
`The following trace is impossible: ${trace}`,
|
|
`An unrelated illustration: ${trace}`,
|
|
`It is unclear whether this happens: ${trace}`,
|
|
`This trace did not occur: ${trace}`,
|
|
trace.replace('Read misses', 'Read may miss'),
|
|
trace.replace('write commits and deletes', 'write does not commit or delete'),
|
|
trace.replace('reader then fills', 'reader never fills'),
|
|
trace.replace('every later reader sees stale data', 'every later reader never sees stale data'),
|
|
'Read misses, write commits and deletes, reader then fills the pre-commit snapshot; every later reader sees stale data?',
|
|
'Read misses, write commits and deletes, reader then fills the pre-commit snapshot; every later reader sees stale data. This scenario is impossible.',
|
|
]) expect(found(withTrace(value))).toBe(false);
|
|
});
|
|
|
|
test('copied source and independent rows or cells cannot supply missing ordered operations', () => {
|
|
for (const value of [`> ${fixture.f1}`, ` ${fixture.f1}`, `\t${fixture.f1}`,
|
|
`\`\`\`text\n${fixture.f1}\n\`\`\``, `~~~text\n${fixture.f1}\n~~~`]) expect(found(value)).toBe(false);
|
|
const first = withTrace('Read misses; write commits and deletes.');
|
|
const last = withTrace('Reader then fills the pre-commit snapshot; every later reader sees stale data.').replace('| F1 |', '| F2 |');
|
|
expect(found(first + '\n' + last)).toBe(false);
|
|
const cells = fixture.f1.split('|');
|
|
cells[4] = ' Read misses; write commits and deletes. ';
|
|
cells[6] = ' Reader then fills the pre-commit snapshot; every later reader sees stale data. ';
|
|
expect(found(cells.join('|'))).toBe(false);
|
|
expect(found(first + '\n\n> ' + trace)).toBe(false);
|
|
expect(found(withTrace(`"${trace}" is a copied source example, not an observed defect.`))).toBe(false);
|
|
});
|
|
|
|
test('a real trace still rejects dismissal or acceptance of the later stale consequence', () => {
|
|
for (const suffix of [' No fix is required.', ' This stale-read behavior is accepted.', ' There is no stale-fill race.',
|
|
' Later readers may return stale data and that is permitted.']) expect(found(withTrace(trace + suffix))).toBe(false);
|
|
expect(found(withTrace(trace + ' Original reader returns v1 to its own caller (allowed: it began before commit).'))).toBe(true);
|
|
expect(found(withTrace(trace + ' Later reader returns v1 to its own caller (allowed: it began after commit).'))).toBe(false);
|
|
});
|
|
});
|
|
|
|
describe('sdk-original-order-ai', () => {
|
|
const captured = captured_sdk_original_order_ai;
|
|
const compact = () => `### Findings registry\n\n${captured.finding}\n\n${captured.heading}\n\`\`\`\n${captured.trace}\n\`\`\``;
|
|
const rejects = (changes: Array<[string, string]>) => {
|
|
for (const [before, after] of changes) {
|
|
expect(compact()).toContain(before);
|
|
expect(hasStaleFillRaceFinding(compact().replace(before, after))).toBe(false);
|
|
}
|
|
};
|
|
|
|
describe('asserted original order beside an amended cache schedule', () => {
|
|
test('exact completed report and its owned finding/schedule show the original late-fill violation', () => {
|
|
expect(hasStaleFillRaceFinding(captured.report)).toBe(true);
|
|
expect(hasStaleFillRaceFinding(compact())).toBe(true);
|
|
});
|
|
|
|
test('amended behavior or the original caller allowance cannot replace the original stale-fill evidence', () => {
|
|
const original = 'Original sketch, order A: fill V1 at 6 after delete at 4 -> R2 reads V1 for <=30 s VIOLATION';
|
|
rejects([
|
|
[original, ''], [original, 'Not ' + original], [original, '> ' + original],
|
|
[original, '"' + original + '"'], [original, 'If ' + original],
|
|
[original, original.replace('VIOLATION', 'PERMITTED')],
|
|
[original, original.replace('R2 reads', 'R1 reads')],
|
|
[original, original.replace('fill V1', 'skip fill V1')],
|
|
[original, original.replace('after delete at 4', 'before delete at 4')],
|
|
]);
|
|
});
|
|
|
|
test('reader, writer, cache key, versions and completion order must all refer to the same execution', () => {
|
|
rejects([
|
|
['inflight[k]', 'inflight[foreign]'], ['R1 (began before W)', 'R1 (began after W)'],
|
|
['DB write commits V2', 'DB write commits V1'], ['DB returns V1', 'DB returns V2'],
|
|
['resume: invalidate(E1), delete', 'resume: invalidate(E9), delete'],
|
|
['settles -> W complete', 'settles -> W pending'],
|
|
['resume: E1.stale -> skip fill', 'resume: E9.stale -> skip fill'],
|
|
['7 | R2 begins:', '4.5 | R2 begins:'], ['R2 reads V1 for', 'R2 reads V2 for'],
|
|
['fill V1 at 6 after delete at 4', 'fill V1 at 3 after delete at 4'],
|
|
['cache[k]', 'cache[foreign]'],
|
|
]);
|
|
});
|
|
|
|
test('the current finding owns the trace and must independently assert the invariant violation', () => {
|
|
rejects([
|
|
['schedule (F1,', 'schedule (F2,'], ['| F1 | CRITICAL |', '| F2 | CRITICAL |'],
|
|
['| F1 | CRITICAL |', '| F1 | LOW |'],
|
|
['Schedule in Section 4 shows', 'A hypothetical Schedule in Section 4 shows'],
|
|
['filled after `cache.delete`', 'filled before `cache.delete`'],
|
|
['every read begun after that write completes must observe the committed version', 'earlier values are accepted for later readers'],
|
|
]);
|
|
expect(hasStaleFillRaceFinding(compact().replace(captured.finding, captured.finding + '\n' + captured.finding))).toBe(false);
|
|
});
|
|
|
|
test('source and hypothetical framing cannot supply the assertion', () => {
|
|
for (const prefix of ['An unproven hypothesis.', 'Historical example only.', 'The following is a hypothetical example.']) {
|
|
expect(hasStaleFillRaceFinding(prefix + '\n' + compact())).toBe(false);
|
|
expect(hasStaleFillRaceFinding(compact().replace(captured.heading, prefix + '\n' + captured.heading))).toBe(false);
|
|
}
|
|
expect(hasStaleFillRaceFinding(compact().split('\n').map(line => '> ' + line).join('\n'))).toBe(false);
|
|
expect(hasStaleFillRaceFinding('````text\n' + compact() + '\n````')).toBe(false);
|
|
expect(hasStaleFillRaceFinding(compact().replace('### Findings registry', '### Quoted source'))).toBe(false);
|
|
});
|
|
|
|
test('same-finding direct and quoted withdrawals remain authoritative inside or after the trace', () => {
|
|
for (const withdrawal of ['F1 is withdrawn.', 'F1 is rejected.', 'The original schedule is impossible.', 'There is no stale-fill race.', 'Rejected: "There is no stale-fill race."']) {
|
|
expect(hasStaleFillRaceFinding(compact() + '\n\n' + withdrawal)).toBe(false);
|
|
expect(hasStaleFillRaceFinding(compact().replace(captured.trace, captured.trace + '\n' + withdrawal))).toBe(false);
|
|
expect(hasStaleFillRaceFinding(compact().replace('Ordering tests, both orders + late joiner + sentinel variant', withdrawal))).toBe(false);
|
|
}
|
|
expect(hasStaleFillRaceFinding(compact().replace('Readers that began before the write may still see the old snapshot (permitted by contract)', 'Later readers may see old snapshots; this stale-fill behavior is accepted.'))).toBe(false);
|
|
});
|
|
|
|
test('unrelated sections and consistently renamed identities do not change valid evidence', () => {
|
|
expect(hasStaleFillRaceFinding('### Prior example\nHistorical example only.\n\n### Current review\n' + compact())).toBe(true);
|
|
expect(hasStaleFillRaceFinding(compact() + '\n\n### Other finding\nF2 is rejected.')).toBe(true);
|
|
const renamed = compact().replaceAll('R1', 'R7').replaceAll('R2', 'R8').replaceAll('R3', 'R9')
|
|
.replaceAll('V1', 'oldSnapshot').replaceAll('V2', 'newSnapshot').replaceAll('E1', 'pendingA').replaceAll('E2', 'pendingB')
|
|
.replaceAll('[k]', '[profileKey]').replace(/\bW\b/g, 'W2');
|
|
expect(hasStaleFillRaceFinding(renamed)).toBe(true);
|
|
});
|
|
|
|
test('owning source headings and same-finding assessments survive intervening structure', () => {
|
|
for (const heading of ['## Hypothetical example', '## Quoted source', '## Historical example only']) {
|
|
expect(hasStaleFillRaceFinding(heading + '\n' + compact())).toBe(false);
|
|
}
|
|
expect(hasStaleFillRaceFinding(compact().replace(captured.heading,
|
|
'F1 is rejected.\n\nUnrelated diagram:\n```\nA -> B\n```\n\n' + captured.heading))).toBe(false);
|
|
expect(hasStaleFillRaceFinding(compact() + '\n\n### Assessment of F1\nF1 is rejected.')).toBe(false);
|
|
});
|
|
});
|
|
|
|
const retry = () => `## Findings Registry\n\n${captured.retry.finding}\n\n${captured.retry.heading}\n\`\`\`\n${captured.retry.trace}\n\`\`\``;
|
|
describe('version-labelled original prose with its owned schedule', () => {
|
|
test('the exact retry and compact evidence require the original sequence, not amended prevention', () => {
|
|
expect(hasStaleFillRaceFinding(captured.retry.report)).toBe(true);
|
|
expect(hasStaleFillRaceFinding(retry())).toBe(true);
|
|
});
|
|
|
|
test('each version and shared key must agree, with write completion before the later reader', () => {
|
|
for (const [before, after] of [
|
|
['DB returns v1', 'DB returns v2'], ['write commits v2 and', 'write commits v1 and'],
|
|
['read then fills v1;', 'read then fills v2;'], ['every later read gets v1', 'every later read gets v2'],
|
|
['write commits v2 and', 'write commits v3 and'], ['writeGen[key]', 'writeGen[foreign]'],
|
|
['cache[key]', 'cache[foreign]'], ['R2 (read, began after W)', 'R2 (read, began before W)'],
|
|
['delete (no-op), return', 'delete (no-op), pending'], ['DB SELECT -> v1', 'DB SELECT -> v2'],
|
|
['DB UPDATE commits v2', 'DB UPDATE commits v3'], ['promise resolves, set(v1)', 'promise resolves, set(v2)'],
|
|
['get -> v1 VIOLATION', 'get -> v2 VIOLATION'], ['6 sketch', '3 sketch'],
|
|
['3 DB UPDATE commits v2', '3 DB UPDATE commits v2'],
|
|
]) {
|
|
expect(retry()).toContain(before);
|
|
expect(hasStaleFillRaceFinding(retry().replace(before, after))).toBe(false);
|
|
}
|
|
expect(hasStaleFillRaceFinding(retry().replace(captured.retry.trace, captured.retry.trace.split('\n').filter(line => !/\b[456] sketch\b/.test(line)).join('\n')))).toBe(false);
|
|
});
|
|
|
|
test('conditional, quoted, obsolete or withdrawn evidence cannot become a current finding', () => {
|
|
for (const prefix of ['An unproven hypothesis.', 'Historical example only.', 'The following is a hypothetical example.']) {
|
|
expect(hasStaleFillRaceFinding(prefix + '\n' + retry())).toBe(false);
|
|
expect(hasStaleFillRaceFinding(retry().replace('Late fill after write.', prefix + ' Late fill after write.'))).toBe(false);
|
|
}
|
|
for (const heading of ['## Hypothetical example', '## Quoted source', '## Historical example only']) {
|
|
expect(hasStaleFillRaceFinding(heading + '\n' + retry().replace('## Findings Registry', '### Findings Registry'))).toBe(false);
|
|
}
|
|
for (const withdrawal of ['F1 is rejected.', 'S1 is withdrawn.', 'The original schedule is impossible.', 'There is no stale-fill race.', 'Rejected: "There is no stale-fill race."']) {
|
|
expect(hasStaleFillRaceFinding(retry() + '\n\n' + withdrawal)).toBe(false);
|
|
expect(hasStaleFillRaceFinding(retry() + '\n\n### Assessment of F1\n' + withdrawal)).toBe(false);
|
|
expect(hasStaleFillRaceFinding(retry().replace(' S2 join stale flight', withdrawal + '\n S2 join stale flight'))).toBe(false);
|
|
}
|
|
for (const withdrawal of ['S1 is withdrawn.', 'F1 is rejected.']) {
|
|
expect(hasStaleFillRaceFinding(retry().replace(captured.retry.trace, captured.retry.trace + '\n' + withdrawal))).toBe(false);
|
|
}
|
|
expect(hasStaleFillRaceFinding(retry().split('\n').map(line => '> ' + line).join('\n'))).toBe(false);
|
|
expect(hasStaleFillRaceFinding('````\n' + retry() + '\n````')).toBe(false);
|
|
expect(hasStaleFillRaceFinding(retry().replace('Late fill after write.', 'If a late fill happens after write.'))).toBe(false);
|
|
});
|
|
|
|
test('consistent versions and independent later findings remain valid', () => {
|
|
expect(hasStaleFillRaceFinding(retry().replaceAll('v1', 'v7').replaceAll('v2', 'v8').replaceAll('[key]', '[profileKey]').replaceAll('key#1', 'profileKey#1'))).toBe(true);
|
|
expect(hasStaleFillRaceFinding('## Prior example\nHistorical only.\n\n## Current review\n' + retry().replace('## Findings Registry', '### Findings Registry'))).toBe(true);
|
|
expect(hasStaleFillRaceFinding(retry() + '\n\n### Other finding\nF9 is rejected.')).toBe(true);
|
|
});
|
|
});
|
|
});
|
|
|
|
describe('sdk-reported-coordination-ar', () => {
|
|
const fs = fs_sdk_ordered_schedule_ar;
|
|
const report = fs.readFileSync(new URL('./fixtures/sdk-reported-coordination-ar.md', import.meta.url), 'utf8');
|
|
const paragraph = report.split('\n\n').find(text => text.startsWith('## Proposed wrapper integration'))!.split('\n').slice(1).join('\n');
|
|
const matches = (text = paragraph) => hasStaleFillRaceFinding(text);
|
|
|
|
test('the actual retry independently reports the original coordination violation', () => {
|
|
expect(paragraph).toContain('review found that this violates the read-after-write rule above (F1)');
|
|
expect(paragraph).not.toMatch(/stale|in-flight|race|pending/);
|
|
expect(matches()).toBe(true);
|
|
expect(matches(report)).toBe(true);
|
|
expect(matches(paragraph.replace('proposed no coordination', 'had no coordination'))).toBe(true);
|
|
expect(matches(paragraph.replace('proposed no coordination', 'has no coordination'))).toBe(true);
|
|
expect(matches(paragraph.replace('sketch', 'wrapper'))).toBe(true);
|
|
expect(matches(paragraph.replace('rule above', 'contract'))).toBe(true);
|
|
expect(matches(paragraph.replace(/ and omits[\s\S]*/, '.'))).toBe(true);
|
|
});
|
|
|
|
test('missing or hypothetical premise and conclusion cannot become findings', () => {
|
|
for (const [from, to] of [
|
|
['proposed no coordination', 'proposed coordination'],
|
|
['proposed no coordination', 'may propose no coordination'],
|
|
['review found that this violates', 'review may find that this violates'],
|
|
['review found that this violates', 'review found that this does not violate'],
|
|
['review found that this violates', 'review hypothesized that this violates'],
|
|
['review found that this violates', 'review found that another wrapper violates'],
|
|
['read-after-write rule above', 'formatting rule'],
|
|
['(F1)', '(unknown)'],
|
|
['; the\nreview found', '. Another unrelated finding. The\nreview found'],
|
|
['; the\nreview found', '\n\nThe\nreview found'],
|
|
['; the\nreview found', ' | The\nreview found'],
|
|
]) {
|
|
expect(paragraph).toContain(from);
|
|
expect(matches(paragraph.replace(from, to))).toBe(false);
|
|
}
|
|
});
|
|
|
|
test('source and quoted evidence cannot assert the current violation', () => {
|
|
for (const text of [
|
|
'Source:\n\n' + paragraph,
|
|
'Hypothetical scenario. ' + paragraph,
|
|
'Earlier review:\n\n' + paragraph,
|
|
'## Historical example\n' + paragraph,
|
|
'> ' + paragraph.replaceAll('\n', '\n> '),
|
|
'```text\n' + paragraph + '\n```',
|
|
'~~~text\n' + paragraph + '\n~~~',
|
|
paragraph.replace('original sketch proposed no coordination between a cache fill and a write', '`original sketch proposed no coordination between a cache fill and a write`'),
|
|
paragraph.replace('review found that this violates the read-after-write rule above (F1)', '"review found that this violates the read-after-write rule above (F1)"'),
|
|
]) expect(matches(text)).toBe(false);
|
|
});
|
|
|
|
test('the referenced finding owns its later assessment', () => {
|
|
for (const tail of ['F1 is withdrawn.', 'F1 is "withdrawn".', 'F1 is rejected.', 'This finding is dismissed.', 'No coordination is required.', '| ID | Assessment |\n| F1 | Withdrawn: no coordination is required. |', '| F1 | Withdrawn |', '| F1 | "rejected" |']) {
|
|
expect(matches(paragraph + '\n\n' + tail)).toBe(false);
|
|
}
|
|
expect(matches(paragraph + '\n\nF2 is withdrawn.')).toBe(true);
|
|
expect(matches(paragraph + '\n\n| F2 | Withdrawn |')).toBe(true);
|
|
expect(matches(paragraph + '\n\n## Historical assessment\n| F1 | Withdrawn |')).toBe(true);
|
|
expect(matches('## Earlier material\nSource:\nOld source.\n\n## Current findings\n' + paragraph)).toBe(true);
|
|
});
|
|
});
|
|
|
|
describe('sdk-schedule-continuation-ah', () => {
|
|
const fixture = fixture_sdk_schedule_continuation_ah;
|
|
const frame = fixture.compact;
|
|
function replace(from: string, to: string, input = frame): string {
|
|
expect(input.includes(from)).toBe(true);
|
|
return input.replace(from, to);
|
|
}
|
|
const originalRows = ' S2* | await read ... | write commits, delete(noop) | | - |\n'
|
|
+ ' | resolves V0 → set V0 | | hit → V0 | V0 (30 s) | VIOLATION\n';
|
|
|
|
test('retains both exact public report forms as affirmative original-race findings', () => {
|
|
expect(hasStaleFillRaceFinding(fixture.report)).toBe(true);
|
|
expect(hasStaleFillRaceFinding(frame)).toBe(true);
|
|
expect(fixture.report.includes(frame.trim())).toBe(true);
|
|
});
|
|
|
|
test('binds consistently renamed actors, shared key, versions and finding/schedule IDs', () => {
|
|
const renamed = frame.replace(/\bR1\b/g, 'R7').replace(/\bR2\b/g, 'R8').replace(/\bW\b/g, 'W9')
|
|
.replace(/\bV0\b/g, 'oldValue').replace(/\bV1\b/g, 'freshValue')
|
|
.replace(/\bkey\b/g, 'profile_key').replace(/\bF1\b/g, 'F9').replace(/\bS2\b/g, 'S9');
|
|
expect(hasStaleFillRaceFinding(renamed)).toBe(true);
|
|
expect(hasStaleFillRaceFinding(frame.replace(/→/g, '->'))).toBe(true);
|
|
const unrelated = '## Historical example\nAn unrelated old example.\n\n## Current findings\n\n';
|
|
expect(hasStaleFillRaceFinding(unrelated + frame)).toBe(true);
|
|
});
|
|
|
|
test('amendments, permitted earlier readers and missing continuation do not supply the original race', () => {
|
|
for (const changed of [
|
|
replace(originalRows, ''),
|
|
replace('S2* | await read', 'S2 A1 | await read'),
|
|
replace(originalRows, ' S2* | begins before W, joins | delete + forget | — | — | OK: R1 began before W completed (permitted clause)\n'),
|
|
replace('resolves V0 → set V0', 'resolves V0, slot gone→drop'),
|
|
replace('hit → V0', 'miss→read V1→set'),
|
|
replace('hit → V0', ''),
|
|
replace('VIOLATION\n S2 A1', 'OK (permitted earlier return)\n S2 A1'),
|
|
replace('VIOLATION\n S2 A1', 'VIOLATION\n | already guarded | | | | OK\n S2 A1'),
|
|
]) expect(hasStaleFillRaceFinding(changed)).toBe(false);
|
|
});
|
|
|
|
test('requires the original schedule citation, legend and explicit post-completion boundary', () => {
|
|
for (const changed of [
|
|
replace('Schedule S2 makes', 'Schedule S9 makes'),
|
|
replace('`*` = original sketch.', '`*` = amended sketch.'),
|
|
replace('`*` = original sketch.', ''),
|
|
replace('`*` = original sketch.', 'Hypothetically, `*` = original sketch.'),
|
|
replace('CRITICAL GAP | 1, 2, 4, 5, 6', 'CRITICAL GAP | 1, 2, 5, 6'),
|
|
replace('Violates retained invariant.', 'No defect in the retained invariant.'),
|
|
replace('R2 (begins after W)', 'R2 (begins before W)'),
|
|
replace('after `writeProfile` resolves', 'before `writeProfile` resolves'),
|
|
replace('after `writeProfile` resolves', 'after `writeProfile` begins'),
|
|
replace('after `writeProfile` resolves', 'after `readProfile` resolves'),
|
|
]) expect(hasStaleFillRaceFinding(changed)).toBe(false);
|
|
});
|
|
|
|
test('rejects actor, key, value, invalidation and ordering mismatches', () => {
|
|
for (const changed of [
|
|
replace('R2 (begins after W)', 'R1 (begins after W)'),
|
|
replace('R2 (begins after W)', 'R2 (begins after W9)'),
|
|
replace('`inflight[key]`', '`inflight[other_key]`'),
|
|
replace('| cache[key] | Result', '| cache[other_key] | Result'),
|
|
replace('W (commits V1)', 'W (commits V0)'),
|
|
replace('resolves V0 → set V0', 'resolves V1 → set V0'),
|
|
replace('resolves V0 → set V0', 'resolves V0 → set V1'),
|
|
replace('hit → V0', 'hit → V1'),
|
|
replace('write commits, delete(noop)', 'write begins, delete(noop)'),
|
|
replace('write commits, delete(noop)', 'write commits'),
|
|
replace('await read ...', 'await write ...'),
|
|
replace(originalRows, originalRows.split('\n').slice(0, 2).reverse().join('\n') + '\n'),
|
|
replace('V0 (30 s) | VIOLATION', 'V1 (30 s) | VIOLATION'),
|
|
replace('see V0 for 30 s;', 'see V1 for 30 s;'),
|
|
]) expect(hasStaleFillRaceFinding(changed)).toBe(false);
|
|
});
|
|
|
|
test('quotes, source introductions and withdrawn findings remain negative', () => {
|
|
for (const prefix of ['An unproven hypothesis.', 'Historical example only.', 'The following is a hypothetical example.']) {
|
|
expect(hasStaleFillRaceFinding(prefix + '\n\n' + frame)).toBe(false);
|
|
expect(hasStaleFillRaceFinding(replace('### Async Ordering Record', prefix + '\n\n### Async Ordering Record'))).toBe(false);
|
|
expect(hasStaleFillRaceFinding(replace('### Findings Registry\n', '### Findings Registry\n\n' + prefix))).toBe(false);
|
|
}
|
|
expect(hasStaleFillRaceFinding(frame.split('\n').map(line => '> ' + line).join('\n'))).toBe(false);
|
|
expect(hasStaleFillRaceFinding('````text\n' + frame + '\n````')).toBe(false);
|
|
expect(hasStaleFillRaceFinding(replace('```\n Sched', '```javascript\n Sched'))).toBe(false);
|
|
for (const dismissal of [
|
|
'The original trace is impossible.', 'This schedule is not a bug.',
|
|
'The original race is permitted.', 'The stale fill is accepted.',
|
|
'No coordination is required.',
|
|
]) {
|
|
expect(hasStaleFillRaceFinding(frame + '\n' + dismissal)).toBe(false);
|
|
expect(hasStaleFillRaceFinding(replace('Violates retained invariant.', 'Violates retained invariant. ' + dismissal))).toBe(false);
|
|
}
|
|
});
|
|
|
|
|
|
test('completed prior decision section is independent; spoofed or withdrawn framing is not', () => {
|
|
const close = '### Decision Registry (all auto-resolved to recommended option)\n\n| D1 | A | B |\n\nLake Score: 7/7 recommendations chose the complete option.\n\n';
|
|
expect(hasStaleFillRaceFinding(close + frame)).toBe(true);
|
|
expect(hasStaleFillRaceFinding(close.replace('### Decision Registry (all auto-resolved to recommended option)', '### Historical example') + frame)).toBe(false);
|
|
expect(hasStaleFillRaceFinding(close.replace('Lake Score: 7/7 recommendations chose the complete option.', 'An unproven hypothesis.') + frame)).toBe(false);
|
|
const row = frame.split('\n').find(line => line.startsWith('| F1 |'))!;
|
|
expect(hasStaleFillRaceFinding(replace(row, row + '\n' + row))).toBe(false);
|
|
});
|
|
test('same finding or schedule tail withdrawals remain authoritative', () => {
|
|
for (const tail of ['S2 is impossible.', 'F1 is rejected. The original trace is impossible.', 'F1 is rejected.', 'S2 is withdrawn.']) {
|
|
expect(hasStaleFillRaceFinding(frame + '\n' + tail)).toBe(false);
|
|
}
|
|
expect(hasStaleFillRaceFinding(frame + '\nF2 is rejected. The original trace is impossible.')).toBe(true);
|
|
expect(hasStaleFillRaceFinding(frame + '\nS3 is impossible.')).toBe(true);
|
|
});
|
|
});
|
|
|
|
describe('sdk-stale-table-ad-v3', () => {
|
|
const fixture = fixture_sdk_stale_table_ad_v3;
|
|
const found = hasStaleFillRaceFinding;
|
|
const allowance='Original reader still returns v1 to its own caller (allowed: it began before commit)';
|
|
test('actual table finding distinguishes forbidden later stale reads from the permitted original caller',()=>{
|
|
expect(found(fixture.report)).toBe(true);
|
|
expect(found(fixture.table)).toBe(true);
|
|
expect(fixture.provenance.noRetroactivePass).toBe(true);
|
|
});
|
|
test('the already-started original read may use a version label without changing ownership',()=>{
|
|
for(const token of ['v17','VERSION_A','snapshot-A'])expect(found(fixture.table.replaceAll('v1',token))).toBe(true);
|
|
});
|
|
test('allowance cannot migrate to later readers, a post-commit start, or a cache fill',()=>{
|
|
for(const changed of [
|
|
'Later readers return v1 (allowed: they began after commit)',
|
|
'Original reader still returns v1 to its own caller (allowed: it began after commit)',
|
|
'Original reader still returns v1 to its own caller (allowed: it never began before commit)',
|
|
'Original reader fills the cache with v1 (allowed: it began before commit)',
|
|
'Original reader still returns v1 to later readers (allowed: it began before commit)',
|
|
])expect(found(fixture.table.replace(allowance,changed))).toBe(false);
|
|
});
|
|
test('a permitted original caller cannot hide acceptance of later stale reads or no required fix',()=>{
|
|
for(const suffix of [' This stale-read behavior is accepted.',' No fix is required.',' Later readers may return stale data; this is the accepted consistency model.'])
|
|
expect(found(fixture.table.replace('None against the invariant.','None against the invariant.'+suffix))).toBe(false);
|
|
});
|
|
test('copied table source and absent late-fill evidence cannot provide coverage',()=>{
|
|
expect(found('```text\n'+fixture.table+'\n```')).toBe(false);
|
|
expect(found(fixture.table.split('\n').map(x=>'> '+x).join('\n'))).toBe(false);
|
|
expect(found(fixture.table.split('\n').map(x=>' '+x).join('\n'))).toBe(false);
|
|
const rows=fixture.table.split('\n'),cells=rows[2]!.split('|');
|
|
cells[4]=' There is no stale-fill race; later reads observe the committed value. ';
|
|
rows[2]=cells.join('|');expect(found(rows.join('\n'))).toBe(false);
|
|
});
|
|
|
|
|
|
test('original-caller exception requires asserted chronology for that reader',()=>{
|
|
for(const changed of [
|
|
'Original reader still returns v1 to its own caller (allowed: it may have begun before commit)',
|
|
'Original reader still returns v1 to its own caller (allowed: it did not begin before commit)',
|
|
'Original reader still returns v1 to its own caller (allowed: it began before commit only if the write failed)',
|
|
'Original reader still returns v1 to its own caller (allowed: another reader began before commit)',
|
|
'Original reader still returns v1 to its own caller (allowed: the write began before commit)',
|
|
'If the original reader still returns v1 to its own caller, that is allowed: it began before commit',
|
|
])expect(found(fixture.table.replace(allowance,changed))).toBe(false);
|
|
});
|
|
|
|
test('an original-return allowance cannot erase another allowed stale consequence',()=>{
|
|
for(const changed of [
|
|
allowance+' and stores that v1 in the cache for later readers',
|
|
allowance+'; later readers may reuse this old value and that is allowed',
|
|
allowance+'. New readers may reuse this old value and that is permitted',
|
|
allowance+'. The stale cache refill is acceptable',
|
|
])expect(found(fixture.table.replace(allowance,changed))).toBe(false);
|
|
});
|
|
|
|
test('table rows cannot borrow an ordering defect from another issue or from quoted source',()=>{
|
|
const rows=fixture.table.split('\n'),cells=rows[2]!.split('|');
|
|
const originalFailure=cells[4]!;
|
|
cells[4]=' The original reader receives its pre-commit snapshot; later reads observe the committed version. ';
|
|
const missing=rows.slice(0,2).concat(cells.join('|')).join('\n');
|
|
expect(found(missing)).toBe(false);
|
|
const other=cells.slice();other[1]=' D2 ';other[4]=originalFailure;
|
|
other[5]=' This stale-read behavior is accepted; no fix is required. ';
|
|
expect(found(missing+'\n'+other.join('|'))).toBe(false);
|
|
expect(found('> '+originalFailure+'\n\n'+missing)).toBe(false);
|
|
expect(found('```text\n'+originalFailure+'\n```\n\n'+missing)).toBe(false);
|
|
});
|
|
});
|
|
|
|
// Census 36597762183 slice 3: the final PLAN.md (rebuilt from the captured Edits)
|
|
// traced the race as an arrow-ordered execution in its WR-1 ledger row.
|
|
describe('arrow-ordered stale-fill execution', () => {
|
|
const report = readFileSync(join(import.meta.dir, 'fixtures/ceo-section-loading-36597762183-report.md'), 'utf8');
|
|
const trace = 'R1 miss -> R1 store read (v1) -> W commit v2 -> W cache.delete -> W fulfills -> R1 cache.set(v1) -> R2 (begun after W) hits v1.';
|
|
test('the captured report identifies the seeded race', () => {
|
|
expect(report).toContain(trace);
|
|
expect(hasStaleFillRaceFinding(report)).toBe(true);
|
|
});
|
|
test.each([
|
|
['fill before invalidation', trace.replace('W cache.delete -> W fulfills -> R1 cache.set(v1)', 'R1 cache.set(v1) -> W cache.delete -> W fulfills')],
|
|
['later reader is the filling reader', trace.replace('R2 (begun after W)', 'R1 (begun after W)')],
|
|
['later reader began before the write', trace.replace('begun after W', 'begun before W')],
|
|
['fill stores the committed version', trace.replace('R1 cache.set(v1)', 'R1 cache.set(v2)')],
|
|
['later reader sees the committed version', trace.replace('hits v1.', 'hits v2.')],
|
|
['trace declared impossible', trace + ' This order is impossible here.'],
|
|
])('%s is not the seeded race', (_name, mutated) => {
|
|
expect(hasStaleFillRaceFinding(report.replace(trace, mutated))).toBe(false);
|
|
});
|
|
});
|