mirror of
https://github.com/garrytan/gstack.git
synced 2026-09-09 14:38:59 +02:00
* fix: pin the claude CLI to an exact version in the CI image + tripwire The image installed @anthropic-ai/claude-code UNPINNED and rebuilt weekly 'to pick up CLI updates' — while bun sat carefully pinned at 1.3.13 two RUN lines above. The PTY harness screen-scrapes this CLI's TUI, and that drift broke it three separate times (welcome-screen wedge on 2.1.233, skillify HOME discovery on 2.1.237, guard/freeze hooks on 2.1.162), each debugged as a flake first. Pin 2.1.251 (current latest), bump deliberately via a PR that runs the PTY gate, and enforce with test/ci-image-cli-pin.test.ts: any global npm install in Dockerfile.ci without an exact @X.Y.Z pin fails the free suite. The weekly ci-image cron stays as a cheap tag self-heal. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat: stamp the claude CLI version into every eval-store run record Three harness breakages were traced to claude-CLI TUI drift only after long flake hunts, because no run record said which CLI it actually exercised. EvalCollector now stamps claude_cli_version (claude --version, cached once per process, 'unknown' when the binary is absent) into both partial and finalized records — schema-additive optional field, no SCHEMA_VERSION bump. Correlating a flake wave with a CLI release becomes a grep over ~/.gstack/projects/<slug>/evals/ instead of archaeology. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix: give the spinning-shard kill test load headroom (30s -> 90s) The test spawns and group-kills three real children (one a busy-loop burning a full core) while five sibling shard processes compete for eight vCPUs. Under full-suite load it blew bun's default 30s per-test ceiling at 30,009ms — while passing in isolation in 1.4s — and red the only required lane. Every assertion in it is event-based (statuses, group-kill proof, heartbeat lines); the sole latency claim is the <30s kill-deadline sanity bound, which stays. Explicit 90s headroom, not a weakened oracle. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat: green-by-skip census — skip counts in the classifier, all-skipped labeling in the paid runner bun's 'Ran N tests' line COUNTS skipped tests, so a codex/gemini shard whose every test self-skipped (binary absent on the runner — true of every CI runner today) exits 0, dodges the hollow-shard guard, and reads as coverage in the weekly census. The classifier now parses bun's ' N skip' / ' N pass' recap lines; ShardOutcome carries skippedTests; formatSummary and the fail-closed slices report label an all-skipped pass explicitly: 'all N tests SKIPPED — verified nothing'. Status stays 'passed' (external service availability is host state, not a repo regression) but the census can no longer mistake absence for coverage. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * refactor: extract composite actions for eval-lane setup; surviving lanes gain the fail-fast registry verification 'Fix bun temp' x3, 'Restore deps' x5, 'Seed claude interactive config' x3, and 'Register gstack skills' x3 were byte-near-identical copies across the legacy matrix, the sliced lane, and the periodic lane — and only the MATRIX copy of register-skills carried the 19-line dangling-symlink + frontmatter fail-fast loop written after a silent 'Unknown command' + 35-min-timeout incident. Extract all four into .github/actions/ composites; the register composite carries the verification loop (generalized over the skill list), so the sliced and periodic lanes — the lanes that SURVIVE the matrix deletion — now inherit the check they had silently dropped. Matrix-job inline copies are left untouched: that job is deleted next. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat: delete the legacy 17-row eval matrix — the sliced lane is the only paid lane Every PR paid twice: the hand-enumerated matrix (18 test files, 22.6 min, ~$21 API measured on run 33263204465) ran serialized AHEAD of the strictly superior sliced lane via 'needs: evals' — 35.5 min wall and ~2x paid spend for the same diff. 14 of 17 rows carried no tier:, so periodic Opus benchmarks leaked into every PR (the e2e-plan row alone: 12/12 tests, 21.7 min, $7.28 — the wall-clock bound of ALL of CI). Parity receipt (static, pre-deletion): the sliced lane's gate census (49 files, derived from the runner itself) strictly contains all 18 matrix test files, plus 31 files the matrix never ran. Pure deletion — one revert restores it. The PR comment moved into slices-report (same '## E2E Evals' upsert marker, now sourced from slice artifacts + carrying the fail-closed reconciliation verdict). plan-slices loses the needs edge; the dead workflow-level EVALS_TIER env goes with it. test/evals-workflow-matrix.test.ts (and its KNOWN_MATRIX_GAPS / KNOWN_TIER_UNSET burn-down ratchets — retired: the sliced census makes 'every gate file runs' true by construction) is rewritten as test/evals-workflow-wiring.test.ts: matrix stays deleted, planner/executor/ report tier + slice-count agreement, both surviving lanes on the shared register-skills composite with its fail-fast verification loop, PR comment survival. Expected: PR eval wall 35.5 -> ~13 min, per-PR paid spend ~halved. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix: provider-runner timeouts kill the whole process GROUP; codex/gemini inherit the orphan-drain hardening All three provider runners (claude/codex/gemini) killed only the direct child on timeout: tool subprocesses the CLI spawned survived as orphans holding our pipes open and burning shared API rate (observed: a 600s timeout stretching past 1400s; a stalled run once burned a core for 15 hours). gstack-detach's watchdog had the same shape one level up — killpg SIGTERM, 5s grace, then a direct-child proc.kill() that orphaned grandchildren. Fix: spawn provider children via node:child_process with detached (own process group) and killProcessGroup(SIGKILL) in the timeout handler — runShardChild's proven pattern, EPERM/ESRCH fallbacks included. The codex and gemini copies also gain the reader.cancel() + stderr Promise.race hardening only the claude copy had (they still carried the blocked-drain hang it fixed). gstack-detach's watchdog now group-SIGKILLs after the grace. Regression net: test/session-runner-groupkill.test.ts drives the REAL runSkillTest against a fake claude shim (PATH override) that spawns a grandchild and wedges — the run must classify timeout within budget and leave neither shim nor grandchild alive — plus source pins on all three runners (detached + killProcessGroup, no bare timeout kill, no Bun.spawn reversion). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix: skill-e2e-opus-47 renders SKILL.md fixtures into a mkdtemp — never the live tree mkEvalRoot ran gen-skill-docs with cwd=ROOT, regenerating every in-repo SKILL.md mid-run while concurrent paid shards copyFileSync those same files in their beforeAll (EVALS_JOBS>=4 locally, 2 per CI slice) — a sibling could capture a half-regenerated or opus-rendered SKILL.md, and a timeout before afterAll stranded the whole tree at the wrong model for every later shard. A cross-shard race that could flake ANY concurrent paid test. Render via the --out-dir flag gen-skill-docs grew for exactly this reason (mirrors the repo layout, which is all the fixture reads), read the skill heads from the render dir, delete it, and drop the afterAll restore-regen entirely. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix: claude CLI version resolves in the runner parent, never on a test thread Eng-review finding: getClaudeCliVersion's fallback is a SYNCHRONOUS spawnSync on the same thread that polls concurrent PTY/session tests — the judgePtyState blocking class this overhaul kills elsewhere. The paid runner parent now resolves it once (cached) and stamps GSTACK_CLAUDE_CLI_VERSION into every shard's env; eval-store short-circuits on the env var, and the fallback spawn's budget tightens 10s -> 3s (bounded one-time stall, records 'unknown' on a slow CLI). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * test: wire skippedTests end-to-end through runPaidShard The census unit tests hand-built outcomes and the classifier tests parsed strings; nothing proved a real child's ' N skip' recap flows into outcome.skippedTests and the formatSummary label. A commandFor fake now prints the recap shape and the test asserts the parsed counts, the all-skipped predicate, and the 'verified nothing' label. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix: make the setup composites rerun-safe (codex diff-review hardenings) restore-deps: 'cp -r SRC node_modules' with an existing node_modules NESTS the copy and leaves stale deps active — rm first. register-gstack-skills: 'ln -snf' hard-errors under set -eu when a REAL directory occupies the gstack slot — clear a non-symlink leftover first. CI workspaces are fresh today; a reusable composite must survive dirty reruns. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix: sweep — every sync spawn in the test trees carries a timeout (436 sites, 157 files) spawnSync/execSync/Bun.spawnSync BLOCK the main thread, so bun's in-process per-test timeout can never fire while one waits — a hung child (stdin read, network probe, dead daemon) wedges the whole shard until the runner's external wall-clock SIGKILL. This exact class reached main: free-tests run 33262077256, test/gstack-memory-ingest.test.ts (normally 2.3s) held shard 2 at the 360s wall while its five siblings finished in ~65s. Mechanical sweep in two waves (12 + 4 fan-out agents, every edit verified against its call site): default timeout: 30_000 (matches the free runner's per-test budget), 120_000 for genuinely slow ops (installs, builds, playwright, provider CLIs), helper wrappers fixed ONCE where call sites route through them. Sites that only LOOK like calls (string fixtures, grep needles, comments) were skipped with reasons — the enforcement commit that follows marks them exempt. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * test: sync-spawn timeout tripwire — the wedge class stays extinct Free scanner over all test trees (test/, browse/test/, design/test/, make-pdf/test/, ios-qa, browser-skills): every spawnSync/execSync/ Bun.spawnSync call site must carry a timeout within a 30-line options window, or an explicit '// tripwire-exempt: <reason>' marker. Comment lines are skipped; exemptions are counted and ratcheted shrink-only (ceiling 6 = the 6 string-fixture/grep-needle sites where the pattern is CONTENT, not a call — marked in this commit). A scan-sanity test pins that the scanner still sees >100 real call sites so it can never rot to a vacuous green. Companion to the 436-site sweep in the previous commit. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat: paid-lane flake telemetry — record-level attempts, flaky_retries, report surfacing bun --retry leaves a retried pass INVISIBLE in its output: a fail-then-pass prints the error detail but no (fail) result line and recaps as a clean pass (probed live on 1.3.10). So attempts are recorded where they cannot lie: EvalCollector.addTest stamps a 1-based attempt on same-name re-records (a retried test runs its body again and re-records), finalized runs carry flaky_retries, printSummary warns loudly, and the fail-closed slices report lists every passed-only-on-retry test — recorded and ranked, never blocking and never silent. Cross-model confirmed (codex reached the same don't-parse -the-stream conclusion independently). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat: free-lane flake ledger — retry ON in CI, flaky-passes recorded and uploaded The runner's attribution-gated flaky-retry pass (cap 5, truncation veto) was OFF in the required lane and its FLAKY-PASS evidence was console-only — so a single timing flake red the merge gate while repeat offenders stayed unenumerable. free-tests.yml now sets GSTACK_FREE_RETRY_FLAKY=1 and points GSTACK_FLAKE_LEDGER at runner.temp; every flaky-pass appends a JSONL entry (SINGLE writer: the parent runner — no concurrent-append hazard by construction; fail-open with a loud warning so a broken ledger can never red the lane) and the artifact uploads UNCONDITIONALLY — a flaky-pass run is green, which is exactly when the evidence matters. Wiring pinned by free-tests-workflow-wiring; ledger behavior unit-tested incl. the fail-open path. Matches 2026 industry practice (retry for data, quarantine out of merge-blocking but never out of logging) with the repo's own receipts. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat: eval:flake-rank — the flake-telemetry dial Aggregates per-test series across every finalized eval-store run (shard dirs included) plus the free flake ledger: runs, fails, RETRIED PASSES (the flake signature), avg duration — ranked retries-first. This is the readable dial behind two policies: a flaky pass never blocks a merge but is always ranked here, and the WS16 required-check promotion needs weeks of clean flake-rank, not vibes. --json for machines, --dir for downloaded CI artifacts. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat: two-phase session timeout — silent APIs die at the startup grace, named The single spawn-armed timer charged API queue latency to the work budget: the recurring '0 turns / $0.00 / x3 attempts' failure with four budget-bump receipts (180->300s, 240->360s, 300->420s, 90->300s). Split: startup phase (no NDJSON byte yet) kills EARLY at min(grace, timeout) with the distinct exitReason 'timeout_startup' — an availability verdict, not transcript archaeology — and the work phase arms on the first byte for the REMAINING budget, so total wall never exceeds the timeout (tier envelopes are margin-free: tests pass timeout: CAPTURE_MS and bun-budget the same tier). Local grace 90s (observed queue latency 60-90s), CI floor 300s (TODOS-filed; shared runners queue harder), both pinned by the new grace tests with fake -claude shims covering the late-first-byte and silent-API paths. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix: census integrity — 17 phantom selection keys deleted, reverse invariant added, gitignored dep patterns replaced, local map forks derived The merge-blocking gate census counted tests that could not run. Deleted (critic-verified against both quoted-occurrence and dep-registration liveness): 7 *-prosons-format keys with no declaring test, ship-plan- completion/-verification, review-plan-completion, design-shotgun-path/ session/full, autoplan-core (dead ~10 months), e2e-harness-audit (its namesake is a FREE-suite file), plus 2 dead LLM-judge keys and 2 free-file keys (budget-regression-pty, global-discover) misplaced in the PAID maps. Census: 191 -> 174 keys, gate 86 -> 78 honest. The new reverse invariant in touchfiles.test.ts makes the class structurally impossible: every key must be quoted in a living paid test file OR registered to an existing paid test file via its dep list (the constructed- name binding the 2026-08 self-registration sweep established) — zero exceptions needed today, with a live-file check on any future exception. Also: '.agents/skills/**' dep patterns replaced with the generator (scripts/gen-skill-docs.ts) — .agents/ is gitignored, so those patterns could NEVER match a git diff and review-template edits silently stopped selecting codex/gemini tests; the codex/gemini local touchfile maps are now DERIVED from the canonical map (loud throw if a key vanishes) instead of hand-forked copies that had already drifted. ios-qa-e2e demoted gate -> periodic: its gate declaration was never executable in CI (hardware exclusion only applies at tier=periodic), so every Linux PR planned a hollow shard. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix: routing journeys lose their answer key and end at the routing decision The journey tests exist to catch skill-DESCRIPTION regressions (touchfiles: */SKILL.md.tmpl), but the fixture CLAUDE.md shipped an explicit prompt->skill lookup table — with the answer key in context, a badly regressed frontmatter description still routed correctly, so the tests could not fail on the exact class they select for. The fixture now carries only the generic invoke-skills nudge; the frontmatter carries the routing load. Also capped all 10 journeys at maxTurns 2 / tools [Skill, Read]: only the FIRST Skill call is asserted, so 5 turns of Read/Bash/Glob/Grep was pure spend — roughly halves each journey's cost. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix: retire decided A/B experiments; vendor the pre-cut fixture; ban raw-SHA fixtures Three one-shot decision experiments kept re-running weekly as N=1 stochastic comparisons — flaky by construction with near-zero remaining information: skill-e2e-auq-repetition-cut-ab (its own header: gate "passed pre-landing, approved 2026-08-25"), skill-e2e-preamble-script-ab ("demoted post-Phase-3"), and opus-47's fanout arm-vs-arm (parA >= parB across two SINGLE stochastic runs — a coin flip). Deleted, with their selection keys; the SDK overlay-harness stays as the maintained instrument for the next experiment, and opus-47 keeps its routing-precision cases. verboseSkill() now reads the VENDORED test/fixtures/auq-pre-cut-...-SKILL.md instead of `git show ab66193e^:...` — a branch-local ref that dies on branch prune and already failed on shallow clones. New free tripwire (test/git-ref-fixture-tripwire.test.ts) bans the raw-SHA fixture class outright: quoted SHA:path rev-specs and gitRef-style hex defaults in the test trees fail the suite with the vendor-instead instruction. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix: demote plan-ceo-review-expansion-energy to periodic Opus generator + a subjective 2-axis >=4/5 LLM-judge threshold sat in the MERGE-BLOCKING gate — the exact class its sibling posture tests were demoted for, with a receipt (a +21-line preamble change once flipped the score). CLAUDE.md's own tiering rule: Opus model test -> periodic. The weekly lane keeps the regression signal; merges stop paying a judge- temperament tax. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix: paid shards get per-shard TMPDIR + CHROMIUM_PROFILE isolation and a kill-path cleanup backstop The free runner treats this isolation as MANDATORY (two concurrent shards on one Chromium profile kill each other's browser; shared tmp cross-contaminates) — the paid lane had none of it. Doubly load-bearing here: a shard that hits its 30-min wall is group-SIGKILLed, so per-test afterAll cleanup never runs; the rmSync backstop is the only thing keeping wedged runs from accumulating full git-repo workspaces in the shared tmpdir forever. This is the DAG prerequisite for raising EVALS_JOBS (next commit) — more concurrency on shared state amplifies exactly the shared-tree race class opus-47 exhibited. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat: paid-runner defaults 4x4 -> 8x2 — halve the local gate worst case 39 of 75 skill-e2e files hold exactly ONE test, so within-shard concurrency was dead weight for most shards: 4 jobs x 4 concurrency yielded only ~4-6 real in-flight sessions and a 13-wave local gate worst case (~6.5h). 8 jobs x 2 gives ~10-13 in-flight — under the documented-safe ~15 — and ~7 waves (~3.3h worst case). CI lanes keep their explicit EVALS_JOBS env (2 per slice; 4 for gate-census); this changes local defaults. Rollback trigger: sustained 429 storms in the WS1 telemetry across 2 PR cycles. test/eval-detach-timeout-floor.test.ts recomputed green (the raise LOWERS the worst-case floor). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix: SHA-pin every action in the secrets-bearing eval lanes evals.yml and evals-periodic.yml execute PR-authored code with three provider API keys in env, yet rode mutable action tags (@v7/@v8/@v2/@v4) — while quality-gate.yml, osv-scanner.yml, and dependency-review.yml already model the SHA-pin pattern. All 30 uses sites across both lanes now pin the exact commit (tag noted in a trailing comment); dependabot's github-actions ecosystem keeps them fresh via PRs instead of silent tag moves. Pulled forward from the plan's endgame on the CEO-review + outside- voice agreement: supply-chain pins on secret lanes go first, not last. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix: sweep wave 3 — the execFileSync family gets timeouts (90 sites, 17 files) The tripwire's regex covered spawnSync/execSync/Bun.spawnSync but not execFileSync — an entire blocking sync-spawn API family that could reintroduce the shard-wedge class undetected (ship review army). Same mechanical recipe as waves 1-2: timeout: 30_000 default, 120_000 for slow ops, shared wrappers fixed once, string-needle sites skipped with reasons. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * test: review-army + adversarial test hardening - Tripwire scans execFileSync too (ceiling 8: two more grep-needle string exemptions); merge-introduced timeout-less spawnSync in question-preference-hook fixed — the tripwire caught a site that landed on main AFTER the sweep, on its first day. - gstack-detach gains TWO watchdog kill regression tests: TERM-immune grandchild (the killpg-after-grace escalation) and the leader-dies variant (the pgid-at-spawn fix — the case the first test cannot see). - eval-flake-rank gets its unit suite (final-attempt accounting, artifact exclusion, shard recursion, recency bound). - Groupkill/startup-grace shim markers are per-run unique (pid-suffixed sleep durations): sibling Conductor worktrees run free suites with no machine lock, and fixed markers let one run pgrep/pkill the other's shims — a cross-run flake inside the anti-flake tests. - flake-ledger test pins the project-scoped local default; stale empty section headers in touchfiles-data deleted (they invited entries under deliberately retired categories). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix: adversarial-review runtime fixes across the telemetry + kill paths - session-runner: exit-labeling keys off 'exit', not 'close' — an orphan holding the pipes could relabel a REAL exit (auth failure) as 'timeout_startup' availability noise; the kill path still always group-kills and cancels the reader (labeling and unblocking are separate concerns). Work phase arms on a flag, not firstResponseMs===0 (a same-ms first byte left the startup timer live all run). The CI startup grace is now a real FLOOR (Math.max), matching its name and pinning test. - gstack-detach: pgid captured AT SPAWN (== child pid under start_new_session) — resolving it after the grace raised ESRCH once the leader died on SIGTERM, orphaning TERM-immune grandchildren forever. - test-free-shards: ledger entries carry branch + git_sha (rev-parse split: '--abbrev-ref HEAD HEAD' printed the branch twice and recorded it as the sha); local ledger default is per-PROJECT, not the machine-global tmpdir. - eval-flake-rank: per-LINE ledger parse (one torn JSONL line vanished the whole series), 60-day recency bound (transcript-bearing files are MBs), shared isFinalizedEvalResultFile predicate (the artifact-taxonomy rule lived in three places); eval-store exports the predicate and finalize stops computing flakyRetries twice; paid-shards cleanup uses async rm (a SIGKILLed shard's git-workspace teardown blocked every sibling's stream classification on the parent event loop). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix: CI trust-boundary + fail-closed repairs (adversarial findings) - Token/exec separation restored: slices-report (runs PR-authored code: bun install + the reconcile runner) drops to contents:read; the PR comment moves to a NEW slices-comment job holding the write token with ZERO repo code — no checkout, no bun, only downloaded artifacts + jq/gh. $GITHUB_ENV/BASH_ENV persistence is job-scoped, so the split is the boundary. The matrix-era report job had this property; the consolidation had regressed it. Pinned by the wiring test. - Reconcile exit captured via PIPESTATUS[0] in BOTH lanes: GitHub's default run-step shell has no pipefail, so `$?` after `| tee` was tee's exit — the fail-closed gate was silently fail-open. Wiring test pins it. - PR comment: final-attempt accounting restored the dropped COST accumulation (the dial read $0 forever), flaky passes render as the warning they are (never as failures), and a malformed tests[] artifact skips that file instead of aborting the whole comment under bash -e. - Remaining mutable action tags pinned (free-tests upload-artifact, ci-image checkout/docker trio — the image publisher holds packages:write and feeds the secret-bearing lanes). restore-deps fallback installs --frozen-lockfile; register-gstack-skills validates skill names before its rm -rf. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * chore: bump version and changelog (v1.77.0.0) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs: update project documentation for v1.77.0.0 Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs: cross-model doc-review fixes — flake-ledger env knobs, CI retry-on note, stale version comment Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs: correct CHANGELOG receipt numbers to measured values Gate census keys: 78 -> 77 (bun-imported E2E_TIERS count). Sweep receipt: 586 sites/176 files -> 499 sites/146 files, measured by running this branch's spawnsync-timeout-tripwire against origin/main (exit 1, 499 violations across 146 unique files; green on this branch). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix: slices-comment creates the PR comment via REST — the write-token job has no git context The token/exec split gives slices-comment NO checkout by design, and gh's pr-comment subcommand resolves the repo FROM git — it died with 'not a git repository' on PR #2746's first run (the update-existing PATCH path was already explicit-repo REST and worked). Create now posts through gh api repos/.../issues/N/comments, and the wiring test pins that no git-context-requiring comment call can creep back into the job. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix: startup-grace probes clear CI for local semantics; new probe pins the floor clamp The two shim probes pass explicit 2s/4s graces, but in CI the runner clamps any explicit grace up to the 300s floor (deliberate adversarial-review fix), so 'silent API killed at the grace' died at the 30s work cap instead of 2s — a deterministic red on every CI run, green locally. The probes now pin LOCAL semantics with CI cleared (same save/restore pattern as their PATH shim), and a fourth probe pins the clamp itself: CI=1 + 2s grace + 6s timeout must kill at the 6s cap, still in the startup phase — proof an explicit low grace cannot bypass the floor. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
1082 lines
37 KiB
TypeScript
1082 lines
37 KiB
TypeScript
// Tests for the gen-accessors TS port. Covers:
|
|
//
|
|
// - Parse: lexical marker isolation + invalid declaration diagnostics
|
|
// - Cache: same input → same key; different swift version → different key;
|
|
// different tool rev/build provenance → different key
|
|
// - Schema: stable hash depends only on ordered accessor signatures
|
|
// - Optional: NSNull read/write/restore behavior and Swift type checking
|
|
// - Prune: >30d entries removed, recent kept
|
|
|
|
import { describe, test, expect, beforeEach, afterEach } from 'bun:test';
|
|
import { spawnSync } from 'child_process';
|
|
import { mkdtempSync, rmSync, writeFileSync, existsSync, readFileSync, mkdirSync, utimesSync } from 'fs';
|
|
import { join } from 'path';
|
|
import { tmpdir } from 'os';
|
|
import {
|
|
collectSwiftFiles,
|
|
parseSwift,
|
|
computeCacheKey,
|
|
computeAccessorHash,
|
|
generate,
|
|
pruneCache,
|
|
render,
|
|
AccessorGenerationError,
|
|
type AccessorSpec,
|
|
} from './gen-accessors';
|
|
|
|
let workDir: string;
|
|
|
|
beforeEach(() => {
|
|
workDir = mkdtempSync(join(tmpdir(), 'gen-accessors-test-'));
|
|
});
|
|
|
|
afterEach(() => {
|
|
rmSync(workDir, { recursive: true, force: true });
|
|
});
|
|
|
|
describe('parseSwift — fork regex-failure-mode fixtures', () => {
|
|
test('parses @Observable class with source-marker comments', () => {
|
|
const src = `
|
|
@Observable
|
|
final class AppState {
|
|
// @Snapshotable
|
|
var isLoggedIn: Bool = false
|
|
// @Snapshotable
|
|
var username: String = ""
|
|
var notSnapshotable: Int = 0
|
|
}
|
|
`;
|
|
const specs = parseSwift(src);
|
|
expect(specs).toHaveLength(1);
|
|
expect(specs[0]!.className).toBe('AppState');
|
|
expect(specs[0]!.fields.map(f => f.name)).toEqual(['isLoggedIn', 'username']);
|
|
expect(specs[0]!.fields.find(f => f.name === 'isLoggedIn')!.typeText).toBe('Bool');
|
|
});
|
|
|
|
test('retains legacy @Snapshotable attribute parsing', () => {
|
|
const specs = parseSwift(`
|
|
@Observable
|
|
final class LegacyState {
|
|
@Snapshotable var counter: Int = 0
|
|
}
|
|
`);
|
|
expect(specs).toEqual([{
|
|
className: 'LegacyState',
|
|
fields: [{ name: 'counter', typeText: 'Int' }],
|
|
}]);
|
|
});
|
|
|
|
test('does not treat documentation prose as a source marker', () => {
|
|
const specs = parseSwift(`
|
|
@Observable
|
|
final class PrivateState {
|
|
/// Do not expose this field through @Snapshotable.
|
|
var token: String = "secret"
|
|
}
|
|
`);
|
|
expect(specs).toHaveLength(0);
|
|
});
|
|
|
|
test('does not let a trailing marker comment bleed into the next field', () => {
|
|
const specs = parseSwift(`
|
|
@Observable
|
|
final class PrivateState {
|
|
var oldValue: Int = 0 // @Snapshotable
|
|
var nextValue: Int = 1
|
|
}
|
|
`);
|
|
expect(specs).toHaveLength(0);
|
|
});
|
|
|
|
test('ignores exact-looking markers inside nested block comments', () => {
|
|
const specs = parseSwift(`
|
|
@Observable
|
|
final class PrivateState {
|
|
/*
|
|
/* // @Snapshotable */
|
|
// @Snapshotable
|
|
var leaked: String = "secret"
|
|
*/
|
|
// @Snapshotable
|
|
var visible: Int = 1
|
|
}
|
|
`);
|
|
expect(specs).toEqual([{
|
|
className: 'PrivateState',
|
|
fields: [{ name: 'visible', typeText: 'Int' }],
|
|
}]);
|
|
});
|
|
|
|
test('ignores declarations and markers inside multiline and raw strings', () => {
|
|
const specs = parseSwift(`
|
|
let ordinary = """
|
|
@Observable class FakeA {
|
|
// @Snapshotable
|
|
var leaked: Int = 0
|
|
}
|
|
"""
|
|
let raw = #"""
|
|
@Observable class FakeB { @Snapshotable var leaked: String = "" }
|
|
"""#
|
|
@Observable
|
|
final class RealState {
|
|
// @Snapshotable
|
|
var safe: Bool = true
|
|
}
|
|
`);
|
|
expect(specs).toEqual([{
|
|
className: 'RealState',
|
|
fields: [{ name: 'safe', typeText: 'Bool' }],
|
|
}]);
|
|
});
|
|
|
|
test('ignores marker words in standalone trailing prose', () => {
|
|
const specs = parseSwift(`
|
|
@Observable
|
|
final class PrivateState {
|
|
// @Snapshotable fields are intentionally disabled here.
|
|
var token: String = "secret"
|
|
}
|
|
`);
|
|
expect(specs).toHaveLength(0);
|
|
});
|
|
|
|
test('handles @Snapshotable on multi-line type signatures', () => {
|
|
const src = `
|
|
@Observable
|
|
class Cart {
|
|
@Snapshotable var items:
|
|
[Dictionary<String, [Int]>]
|
|
= []
|
|
var unrelated: Int = 0
|
|
}
|
|
`;
|
|
const specs = parseSwift(src);
|
|
expect(specs).toHaveLength(1);
|
|
expect(specs[0]!.fields).toHaveLength(1);
|
|
expect(specs[0]!.fields[0]!.name).toBe('items');
|
|
expect(specs[0]!.fields[0]!.typeText).toContain('Dictionary');
|
|
});
|
|
|
|
test('handles JSON-compatible generic types in property signatures', () => {
|
|
const src = `
|
|
@Observable
|
|
class Repo {
|
|
@Snapshotable var pages: Dictionary<String, [Optional<Int>]> = [:]
|
|
}
|
|
`;
|
|
const specs = parseSwift(src);
|
|
expect(specs).toHaveLength(1);
|
|
expect(specs[0]!.fields[0]!.typeText).toContain('Dictionary');
|
|
expect(specs[0]!.fields[0]!.typeText).toContain('Optional');
|
|
});
|
|
|
|
test('accepts qualified JSON-native scalar and collection spellings', () => {
|
|
const specs = parseSwift(`
|
|
@Observable
|
|
class Metrics {
|
|
// @Snapshotable
|
|
var ratio: CoreGraphics.CGFloat = 0
|
|
// @Snapshotable
|
|
var counts: Swift.Array<Swift.Int> = []
|
|
}
|
|
`);
|
|
expect(specs[0]!.fields.map((field) => field.typeText)).toEqual([
|
|
'CoreGraphics.CGFloat',
|
|
'Swift.Array<Swift.Int>',
|
|
]);
|
|
});
|
|
|
|
test('rejects custom values that Foundation JSON cannot serialize', () => {
|
|
expect(() => parseSwift(`
|
|
@Observable
|
|
class Repo {
|
|
// @Snapshotable
|
|
var pages: Dictionary<String, [Result<Item, Error>]> = [:]
|
|
}
|
|
`)).toThrow("unsupported snapshot type 'Result<Item,Error>'");
|
|
});
|
|
|
|
test('rejects nested observable types instead of emitting an unqualified reference', () => {
|
|
expect(() => parseSwift(`
|
|
enum Namespace {
|
|
@Observable
|
|
class State {
|
|
// @Snapshotable
|
|
var count: Int = 0
|
|
}
|
|
}
|
|
`)).toThrow('nested @Observable types are not supported');
|
|
});
|
|
|
|
test('ignores nested observable types that expose no snapshot fields', () => {
|
|
expect(parseSwift(`
|
|
enum Namespace {
|
|
@Observable
|
|
class State { var transient: Int = 0 }
|
|
}
|
|
`)).toEqual([]);
|
|
});
|
|
|
|
test('ignores fields without @Snapshotable marker', () => {
|
|
const src = `
|
|
@Observable
|
|
class M {
|
|
var plain: Int = 0
|
|
@State var stateBacked: String = ""
|
|
}
|
|
`;
|
|
const specs = parseSwift(src);
|
|
expect(specs).toHaveLength(0);
|
|
});
|
|
|
|
test('ignores non-@Observable classes', () => {
|
|
const src = `
|
|
class Plain {
|
|
@Snapshotable var should: Int = 0
|
|
}
|
|
`;
|
|
const specs = parseSwift(src);
|
|
expect(specs).toHaveLength(0);
|
|
});
|
|
|
|
test('handles multiple @Observable classes in one file', () => {
|
|
const src = `
|
|
@Observable
|
|
class A {
|
|
@Snapshotable var a: Int = 0
|
|
}
|
|
@Observable
|
|
class B {
|
|
@Snapshotable var b: String = ""
|
|
}
|
|
`;
|
|
const specs = parseSwift(src);
|
|
expect(specs).toHaveLength(2);
|
|
expect(specs.map(s => s.className).sort()).toEqual(['A', 'B']);
|
|
});
|
|
|
|
test('diagnoses fields with computed body braces', () => {
|
|
const src = `
|
|
@Observable
|
|
class M {
|
|
@Snapshotable var snapshotted: Int = 0
|
|
@Snapshotable var computed: Int {
|
|
get { 42 }
|
|
}
|
|
}
|
|
`;
|
|
expect(() => parseSwift(src)).toThrow("field 'computed' must be stored and writable");
|
|
});
|
|
|
|
test.each([
|
|
['let', '// @Snapshotable\n let immutable: Int = 1', 'must be declared var, not let'],
|
|
['private', '// @Snapshotable\n private var secret: String = ""', 'cannot be private'],
|
|
['fileprivate', '// @Snapshotable\n fileprivate var secret: String = ""', 'cannot be private'],
|
|
['private(set)', '// @Snapshotable\n public private(set) var count: Int = 0', 'cannot be private'],
|
|
['fileprivate(set)', '// @Snapshotable\n fileprivate(set) var count: Int = 0', 'cannot be private'],
|
|
['inferred type', '// @Snapshotable\n var inferred = 42', 'requires an explicit type annotation'],
|
|
['multiple bindings', '// @Snapshotable\n var a: Int = 1, b: Int = 2', 'exactly one binding'],
|
|
['static', '// @Snapshotable\n static var shared: Int = 0', 'must be an instance property'],
|
|
['nested Optional', '// @Snapshotable\n var nested: String?? = nil', 'cannot use a nested Optional'],
|
|
['nested Optional in collection', '// @Snapshotable\n var nestedItems: [String??] = []', 'cannot use a nested Optional'],
|
|
['implicitly unwrapped Optional', '// @Snapshotable\n var legacy: String! = nil', 'cannot use an implicitly unwrapped Optional'],
|
|
['custom type', '// @Snapshotable\n var custom: CustomValue = .init()', 'uses unsupported snapshot type'],
|
|
])('diagnoses invalid %s declarations', (_label, declaration, expected) => {
|
|
expect(() => parseSwift(`
|
|
@Observable
|
|
final class InvalidState {
|
|
${declaration}
|
|
}
|
|
`)).toThrow(expected);
|
|
});
|
|
|
|
test('aggregates invalid declaration diagnostics with class and line context', () => {
|
|
try {
|
|
parseSwift(`
|
|
@Observable
|
|
final class InvalidState {
|
|
// @Snapshotable
|
|
let first: Int = 1
|
|
// @Snapshotable
|
|
private var second: String = ""
|
|
}
|
|
`);
|
|
throw new Error('expected parseSwift to fail');
|
|
} catch (error) {
|
|
expect(error).toBeInstanceOf(AccessorGenerationError);
|
|
const generationError = error as AccessorGenerationError;
|
|
expect(generationError.diagnostics).toHaveLength(2);
|
|
expect(generationError.message).toContain('InvalidState (line 4)');
|
|
expect(generationError.message).toContain('InvalidState (line 6)');
|
|
}
|
|
});
|
|
});
|
|
|
|
describe('computeCacheKey', () => {
|
|
test('same source + same versioning = same key', () => {
|
|
const f = join(workDir, 'a.swift');
|
|
writeFileSync(f, '@Observable class A {}');
|
|
const k1 = computeCacheKey({
|
|
swiftFiles: [f],
|
|
swiftVersion: '6.0.0',
|
|
toolGitRev: 'abc123',
|
|
platformTriple: 'darwin-arm64',
|
|
});
|
|
const k2 = computeCacheKey({
|
|
swiftFiles: [f],
|
|
swiftVersion: '6.0.0',
|
|
toolGitRev: 'abc123',
|
|
platformTriple: 'darwin-arm64',
|
|
});
|
|
expect(k1).toBe(k2);
|
|
});
|
|
|
|
test('source modification changes the key', () => {
|
|
const f = join(workDir, 'a.swift');
|
|
writeFileSync(f, '@Observable class A {}');
|
|
const k1 = computeCacheKey({
|
|
swiftFiles: [f],
|
|
swiftVersion: '6.0.0',
|
|
toolGitRev: 'abc123',
|
|
platformTriple: 'darwin-arm64',
|
|
});
|
|
writeFileSync(f, '@Observable class A { @Snapshotable var x: Int = 0 }');
|
|
const k2 = computeCacheKey({
|
|
swiftFiles: [f],
|
|
swiftVersion: '6.0.0',
|
|
toolGitRev: 'abc123',
|
|
platformTriple: 'darwin-arm64',
|
|
});
|
|
expect(k1).not.toBe(k2);
|
|
});
|
|
|
|
test('swift version change invalidates the key (codex catch)', () => {
|
|
const f = join(workDir, 'a.swift');
|
|
writeFileSync(f, '@Observable class A {}');
|
|
const k1 = computeCacheKey({
|
|
swiftFiles: [f],
|
|
swiftVersion: '5.9.0',
|
|
toolGitRev: 'abc',
|
|
platformTriple: 'darwin-arm64',
|
|
});
|
|
const k2 = computeCacheKey({
|
|
swiftFiles: [f],
|
|
swiftVersion: '6.0.0',
|
|
toolGitRev: 'abc',
|
|
platformTriple: 'darwin-arm64',
|
|
});
|
|
expect(k1).not.toBe(k2);
|
|
});
|
|
|
|
test('generator git rev change invalidates the key (codex catch)', () => {
|
|
const f = join(workDir, 'a.swift');
|
|
writeFileSync(f, '@Observable class A {}');
|
|
const k1 = computeCacheKey({
|
|
swiftFiles: [f],
|
|
swiftVersion: '6.0.0',
|
|
toolGitRev: 'abc123',
|
|
platformTriple: 'darwin-arm64',
|
|
});
|
|
const k2 = computeCacheKey({
|
|
swiftFiles: [f],
|
|
swiftVersion: '6.0.0',
|
|
toolGitRev: 'def456',
|
|
platformTriple: 'darwin-arm64',
|
|
});
|
|
expect(k1).not.toBe(k2);
|
|
});
|
|
|
|
test('platform triple change invalidates the key', () => {
|
|
const f = join(workDir, 'a.swift');
|
|
writeFileSync(f, '@Observable class A {}');
|
|
const k1 = computeCacheKey({
|
|
swiftFiles: [f],
|
|
swiftVersion: '6.0.0',
|
|
toolGitRev: 'abc',
|
|
platformTriple: 'darwin-arm64',
|
|
});
|
|
const k2 = computeCacheKey({
|
|
swiftFiles: [f],
|
|
swiftVersion: '6.0.0',
|
|
toolGitRev: 'abc',
|
|
platformTriple: 'darwin-x86_64',
|
|
});
|
|
expect(k1).not.toBe(k2);
|
|
});
|
|
|
|
test('adding/removing files invalidates the key', () => {
|
|
const f1 = join(workDir, 'a.swift');
|
|
const f2 = join(workDir, 'b.swift');
|
|
writeFileSync(f1, '@Observable class A {}');
|
|
writeFileSync(f2, '@Observable class B {}');
|
|
const k1 = computeCacheKey({
|
|
swiftFiles: [f1],
|
|
swiftVersion: '6.0.0',
|
|
toolGitRev: 'a',
|
|
platformTriple: 'd-arm64',
|
|
});
|
|
const k2 = computeCacheKey({
|
|
swiftFiles: [f1, f2],
|
|
swiftVersion: '6.0.0',
|
|
toolGitRev: 'a',
|
|
platformTriple: 'd-arm64',
|
|
});
|
|
expect(k1).not.toBe(k2);
|
|
});
|
|
|
|
test('app build provenance invalidates the cache key', () => {
|
|
const f = join(workDir, 'a.swift');
|
|
writeFileSync(f, '@Observable class A {}');
|
|
const shared = {
|
|
swiftFiles: [f],
|
|
swiftVersion: '6.0.0',
|
|
toolGitRev: 'abc',
|
|
platformTriple: 'darwin-arm64',
|
|
};
|
|
expect(computeCacheKey({ ...shared, buildId: '100' })).not.toBe(
|
|
computeCacheKey({ ...shared, buildId: '101' }),
|
|
);
|
|
});
|
|
|
|
test('equivalent source content does not depend on its absolute checkout path', () => {
|
|
const firstDir = join(workDir, 'checkout-a');
|
|
const secondDir = join(workDir, 'checkout-b');
|
|
mkdirSync(firstDir);
|
|
mkdirSync(secondDir);
|
|
const first = join(firstDir, 'State.swift');
|
|
const second = join(secondDir, 'State.swift');
|
|
writeFileSync(first, '@Observable class A {}');
|
|
writeFileSync(second, '@Observable class A {}');
|
|
const versioning = { swiftVersion: '6', toolGitRev: 't', platformTriple: 'p' };
|
|
expect(computeCacheKey({ swiftFiles: [first], ...versioning })).toBe(
|
|
computeCacheKey({ swiftFiles: [second], ...versioning }),
|
|
);
|
|
});
|
|
});
|
|
|
|
describe('computeAccessorHash', () => {
|
|
const base: AccessorSpec[] = [{
|
|
className: 'AppState',
|
|
fields: [
|
|
{ name: 'count', typeText: 'Int' },
|
|
{ name: 'nickname', typeText: 'String?' },
|
|
],
|
|
}];
|
|
|
|
test('is deterministic for the same ordered accessor signatures', () => {
|
|
expect(computeAccessorHash(base)).toBe(computeAccessorHash(structuredClone(base)));
|
|
});
|
|
|
|
test('changes when field order, name, or type changes', () => {
|
|
const reordered: AccessorSpec[] = [{
|
|
className: 'AppState',
|
|
fields: [...base[0]!.fields].reverse(),
|
|
}];
|
|
const renamed: AccessorSpec[] = [{
|
|
className: 'AppState',
|
|
fields: [{ name: 'total', typeText: 'Int' }, base[0]!.fields[1]!],
|
|
}];
|
|
const retyped: AccessorSpec[] = [{
|
|
className: 'AppState',
|
|
fields: [{ name: 'count', typeText: 'Int64' }, base[0]!.fields[1]!],
|
|
}];
|
|
const hash = computeAccessorHash(base);
|
|
expect(computeAccessorHash(reordered)).not.toBe(hash);
|
|
expect(computeAccessorHash(renamed)).not.toBe(hash);
|
|
expect(computeAccessorHash(retyped)).not.toBe(hash);
|
|
});
|
|
});
|
|
|
|
describe('generate', () => {
|
|
test('first run writes StateAccessor.swift and populates cache', () => {
|
|
const inputDir = join(workDir, 'src');
|
|
mkdirSync(inputDir);
|
|
writeFileSync(join(inputDir, 'state.swift'), `
|
|
@Observable
|
|
class AppState {
|
|
@Snapshotable var x: Int = 0
|
|
}
|
|
`);
|
|
const cacheRoot = join(workDir, 'cache');
|
|
const r = generate({
|
|
inputDir,
|
|
cacheRoot,
|
|
swiftVersion: '6.0.0',
|
|
toolGitRev: 'test',
|
|
platformTriple: 'darwin-arm64',
|
|
});
|
|
expect(r.cacheHit).toBe(false);
|
|
expect(r.specs).toHaveLength(1);
|
|
expect(r.specs[0]!.className).toBe('AppState');
|
|
expect(existsSync(r.outputPath)).toBe(true);
|
|
expect(existsSync(join(cacheRoot, r.cacheKey, 'StateAccessor.swift'))).toBe(true);
|
|
});
|
|
|
|
test('second run with same inputs hits the cache', () => {
|
|
const inputDir = join(workDir, 'src');
|
|
mkdirSync(inputDir);
|
|
writeFileSync(join(inputDir, 'state.swift'), '@Observable class A { @Snapshotable var x: Int = 0 }');
|
|
const cacheRoot = join(workDir, 'cache');
|
|
const r1 = generate({ inputDir, cacheRoot, swiftVersion: '6', toolGitRev: 't', platformTriple: 'p' });
|
|
const r2 = generate({ inputDir, cacheRoot, swiftVersion: '6', toolGitRev: 't', platformTriple: 'p' });
|
|
expect(r1.cacheHit).toBe(false);
|
|
expect(r2.cacheHit).toBe(true);
|
|
expect(r1.cacheKey).toBe(r2.cacheKey);
|
|
});
|
|
|
|
test('custom nested output subtree cannot poison the second-run cache key', () => {
|
|
const inputDir = join(workDir, 'src');
|
|
const outputDir = join(inputDir, 'generated', 'accessors');
|
|
mkdirSync(inputDir);
|
|
writeFileSync(join(inputDir, 'state.swift'), '@Observable class A { @Snapshotable var x: Int = 0 }');
|
|
const cacheRoot = join(workDir, 'cache');
|
|
|
|
const r1 = generate({
|
|
inputDir,
|
|
outputDir,
|
|
cacheRoot,
|
|
swiftVersion: '6',
|
|
toolGitRev: 't',
|
|
platformTriple: 'p',
|
|
});
|
|
writeFileSync(join(outputDir, 'Ignored.swift'), '@Observable class Poison { @Snapshotable var bad: Int = 1 }');
|
|
const r2 = generate({
|
|
inputDir,
|
|
outputDir,
|
|
cacheRoot,
|
|
swiftVersion: '6',
|
|
toolGitRev: 't',
|
|
platformTriple: 'p',
|
|
});
|
|
|
|
expect(r2.cacheHit).toBe(true);
|
|
expect(r2.cacheKey).toBe(r1.cacheKey);
|
|
});
|
|
|
|
test('modifying source invalidates the cache', () => {
|
|
const inputDir = join(workDir, 'src');
|
|
mkdirSync(inputDir);
|
|
const file = join(inputDir, 'state.swift');
|
|
writeFileSync(file, '@Observable class A { @Snapshotable var x: Int = 0 }');
|
|
const cacheRoot = join(workDir, 'cache');
|
|
const r1 = generate({ inputDir, cacheRoot, swiftVersion: '6', toolGitRev: 't', platformTriple: 'p' });
|
|
writeFileSync(file, '@Observable class A { @Snapshotable var y: String = "" }');
|
|
const r2 = generate({ inputDir, cacheRoot, swiftVersion: '6', toolGitRev: 't', platformTriple: 'p' });
|
|
expect(r1.cacheKey).not.toBe(r2.cacheKey);
|
|
expect(r2.cacheHit).toBe(false);
|
|
});
|
|
|
|
test('unmarked source churn changes cache identity but not schema identity', () => {
|
|
const inputDir = join(workDir, 'src');
|
|
mkdirSync(inputDir);
|
|
const state = join(inputDir, 'state.swift');
|
|
const unrelated = join(inputDir, 'unrelated.swift');
|
|
writeFileSync(state, '@Observable class A { // @Snapshotable\n var x: Int = 0\n }');
|
|
writeFileSync(unrelated, 'struct Unrelated { let a = 1 }');
|
|
const cacheRoot = join(workDir, 'cache');
|
|
const options = { inputDir, cacheRoot, swiftVersion: '6', toolGitRev: 't', platformTriple: 'p' };
|
|
const first = generate(options);
|
|
writeFileSync(unrelated, 'struct Unrelated { let a = 2 }');
|
|
const second = generate(options);
|
|
expect(second.cacheKey).not.toBe(first.cacheKey);
|
|
expect(second.accessorHash).toBe(first.accessorHash);
|
|
});
|
|
|
|
test('buildId changes cannot reuse generated fallback provenance', () => {
|
|
const inputDir = join(workDir, 'src');
|
|
mkdirSync(inputDir);
|
|
writeFileSync(join(inputDir, 'state.swift'), '@Observable class A { @Snapshotable var x: Int = 0 }');
|
|
const cacheRoot = join(workDir, 'cache');
|
|
const shared = { inputDir, cacheRoot, swiftVersion: '6', toolGitRev: 't', platformTriple: 'p' };
|
|
const first = generate({ ...shared, buildId: 'build-100' });
|
|
const second = generate({ ...shared, buildId: 'build-101' });
|
|
expect(second.cacheKey).not.toBe(first.cacheKey);
|
|
expect(second.cacheHit).toBe(false);
|
|
expect(readFileSync(second.outputPath, 'utf8')).toContain('?? "build-101"');
|
|
});
|
|
|
|
test('CLI exits 4 with actionable diagnostics and no generated output', () => {
|
|
const inputDir = join(workDir, 'src');
|
|
const outputDir = join(workDir, 'generated');
|
|
mkdirSync(inputDir);
|
|
writeFileSync(join(inputDir, 'state.swift'), `
|
|
@Observable final class InvalidState {
|
|
// @Snapshotable
|
|
private let secret = "nope"
|
|
}
|
|
`);
|
|
const result = spawnSync('bun', [
|
|
join(import.meta.dir, 'gen-accessors.ts'),
|
|
'--input', inputDir,
|
|
'--output', outputDir,
|
|
], {
|
|
encoding: 'utf8',
|
|
timeout: 30_000,
|
|
env: { ...process.env, GSTACK_IOS_CACHE_ROOT: join(workDir, 'cache') },
|
|
});
|
|
expect(result.status).toBe(4);
|
|
// `let` is diagnosed first; either way the class/property is named and
|
|
// generation never emits an accessor that will fail in xcodebuild.
|
|
expect(result.stderr).toContain('InvalidState');
|
|
expect(result.stderr).toContain("field 'secret' must be declared var, not let");
|
|
expect(result.stderr).not.toContain('AccessorGenerationError:');
|
|
expect(existsSync(join(outputDir, 'StateAccessor.swift'))).toBe(false);
|
|
});
|
|
});
|
|
|
|
describe('pruneCache', () => {
|
|
test('removes entries older than 30d, keeps recent', () => {
|
|
const cacheRoot = join(workDir, 'cache');
|
|
mkdirSync(cacheRoot, { recursive: true });
|
|
const old = join(cacheRoot, 'old-key');
|
|
const fresh = join(cacheRoot, 'fresh-key');
|
|
mkdirSync(old);
|
|
mkdirSync(fresh);
|
|
writeFileSync(join(old, 'StateAccessor.swift'), '// old');
|
|
writeFileSync(join(fresh, 'StateAccessor.swift'), '// fresh');
|
|
|
|
// Backdate the old dir by 60 days.
|
|
const sixtyDaysAgo = (Date.now() - 60 * 24 * 60 * 60 * 1000) / 1000;
|
|
utimesSync(old, sixtyDaysAgo, sixtyDaysAgo);
|
|
|
|
const { pruned } = pruneCache(cacheRoot, 30);
|
|
expect(pruned).toHaveLength(1);
|
|
expect(pruned[0]).toBe(old);
|
|
expect(existsSync(old)).toBe(false);
|
|
expect(existsSync(fresh)).toBe(true);
|
|
});
|
|
|
|
test('no-op on empty cache dir', () => {
|
|
const { pruned } = pruneCache(join(workDir, 'nope'), 30);
|
|
expect(pruned).toHaveLength(0);
|
|
});
|
|
});
|
|
|
|
describe('render', () => {
|
|
test('emits valid-looking Swift for one class with two fields', () => {
|
|
const specs: AccessorSpec[] = [{
|
|
className: 'AppState',
|
|
fields: [{ name: 'a', typeText: 'Int' }, { name: 'b', typeText: 'String' }],
|
|
}];
|
|
const out = render(specs, 'build-1.2.3', 'hash-abc');
|
|
expect(out).toContain('enum AppStateAccessor');
|
|
expect(out).not.toContain('public enum AppStateAccessor');
|
|
expect(out).toContain('static func register(_ state: AppState)');
|
|
expect(out).not.toContain('public static func register(_ state: AppState)');
|
|
expect(out).toContain('key: "a"');
|
|
expect(out).toContain('key: "b"');
|
|
expect(out).toContain('return .missingKey("a")');
|
|
expect(out).toContain('return .typeMismatch("b")');
|
|
expect(out).toContain('guard let restored0 = Self.decodeSnapshotValue(raw0, as: Int.self)');
|
|
expect(out).toContain('guard let restored1 = Self.decodeSnapshotValue(raw1, as: String.self)');
|
|
expect(out).toContain('atomicRestore: { keys, apply in');
|
|
expect(out).toContain('if apply {');
|
|
expect(out).toContain('state.a = restored0');
|
|
expect(out).toContain('state.b = restored1');
|
|
expect(out.indexOf('state.a = restored0')).toBeGreaterThan(out.indexOf('guard let restored1'));
|
|
expect(out).toContain('guard let typed = Self.decodeSnapshotValue(value, as: Int.self) else { return false }');
|
|
expect(out).toContain('state.a = typed');
|
|
expect(out).toContain('return true');
|
|
expect(out).not.toContain('atomicRestore: { _ in .ok }');
|
|
expect(out).not.toContain('write: { _ in false }');
|
|
expect(out).toContain('CFBundleShortVersionString');
|
|
expect(out).toContain('CFBundleVersion');
|
|
expect(out).toContain('return shortVersion ?? bundleVersion ?? "build-1.2.3"');
|
|
expect(out).toContain('accessorHash: "hash-abc"');
|
|
expect(out).toContain('import DebugBridgeCore');
|
|
expect(out).not.toContain('import DebugBridge\n');
|
|
expect(out).toContain('#if DEBUG');
|
|
expect(out).toContain('#endif');
|
|
});
|
|
|
|
test('emits explicit NSNull round-trip handling for Optional fields', () => {
|
|
const out = render([{
|
|
className: 'AppState',
|
|
fields: [
|
|
{ name: 'nickname', typeText: 'String?' },
|
|
{ name: 'selection', typeText: 'Optional<Int>' },
|
|
],
|
|
}], 'build', 'schema');
|
|
|
|
expect(out).toContain('let restored0: String?');
|
|
expect(out).toContain('if raw0 is NSNull');
|
|
expect(out).toContain('else if let typed = Self.decodeSnapshotValue(raw0, as: String.self)');
|
|
expect(out).toContain('let restored1: Optional<Int>');
|
|
expect(out).toContain('else if let typed = Self.decodeSnapshotValue(raw1, as: Int.self)');
|
|
expect(out).toContain('guard let value = state.nickname else { return NSNull() }');
|
|
expect(out).toContain('if value is NSNull');
|
|
expect(out).toContain('state.nickname = nil');
|
|
expect(out).not.toContain('value as? String?');
|
|
});
|
|
|
|
test('rejects duplicate snapshot keys across observable models', () => {
|
|
expect(() => render([
|
|
{ className: 'FirstState', fields: [{ name: 'count', typeText: 'Int' }] },
|
|
{ className: 'SecondState', fields: [{ name: 'count', typeText: 'Int' }] },
|
|
], 'build', 'schema')).toThrow("snapshot key 'count' is declared by both FirstState and SecondState");
|
|
});
|
|
|
|
test('typechecks beside an internal @Observable app state using a comment marker', () => {
|
|
if (spawnSync('swiftc', ['--version'], { encoding: 'utf8', timeout: 30_000 }).status !== 0) return;
|
|
|
|
const coreSource = join(workDir, 'DebugBridgeCore.swift');
|
|
const coreModule = join(workDir, 'DebugBridgeCore.swiftmodule');
|
|
writeFileSync(coreSource, `
|
|
public typealias JSONDict = [String: Any]
|
|
|
|
@MainActor
|
|
public final class StateServer {
|
|
public static let shared = StateServer()
|
|
public enum RestoreResult {
|
|
case ok
|
|
case missingKey(String)
|
|
case typeMismatch(String)
|
|
}
|
|
private init() {}
|
|
public func register(
|
|
buildId: String,
|
|
accessorHash: String,
|
|
atomicRestore: @escaping (JSONDict, Bool) -> RestoreResult
|
|
) {}
|
|
public func registerAccessor(
|
|
key: String,
|
|
type: String,
|
|
read: @escaping () -> Any?,
|
|
write: @escaping (Any) -> Bool
|
|
) {}
|
|
}
|
|
`);
|
|
const emitModule = spawnSync('swiftc', [
|
|
'-emit-module',
|
|
'-parse-as-library',
|
|
'-module-name', 'DebugBridgeCore',
|
|
coreSource,
|
|
'-emit-module-path', coreModule,
|
|
], { encoding: 'utf8', timeout: 120_000 });
|
|
if (emitModule.status !== 0) {
|
|
throw new Error(`failed to build DebugBridgeCore test stub:\n${emitModule.stderr}`);
|
|
}
|
|
|
|
const appSource = join(workDir, 'AppState.swift');
|
|
writeFileSync(appSource, `import Observation
|
|
|
|
@Observable
|
|
final class AppState {
|
|
// @Snapshotable
|
|
var counter: Int = 0
|
|
}
|
|
|
|
${render([{
|
|
className: 'AppState',
|
|
fields: [{ name: 'counter', typeText: 'Int' }],
|
|
}], 'build-test', 'hash-test')}`);
|
|
const typecheck = spawnSync('swiftc', [
|
|
'-typecheck',
|
|
'-D', 'DEBUG',
|
|
'-I', workDir,
|
|
appSource,
|
|
], { encoding: 'utf8', timeout: 120_000 });
|
|
if (typecheck.status !== 0) {
|
|
throw new Error(`generated accessor failed Swift type checking:\n${typecheck.stderr}`);
|
|
}
|
|
});
|
|
|
|
test('strict JSON typing and cross-model validate-before-apply restore run correctly', () => {
|
|
if (process.platform !== 'darwin') return;
|
|
if (spawnSync('swiftc', ['--version'], { encoding: 'utf8', timeout: 30_000 }).status !== 0) return;
|
|
|
|
const coreSource = join(workDir, 'DebugBridgeCore.swift');
|
|
const coreModule = join(workDir, 'DebugBridgeCore.swiftmodule');
|
|
const coreLibrary = join(workDir, 'libDebugBridgeCore.dylib');
|
|
writeFileSync(coreSource, `
|
|
public typealias JSONDict = [String: Any]
|
|
|
|
@MainActor
|
|
public final class StateServer {
|
|
public typealias Restore = (JSONDict, Bool) -> RestoreResult
|
|
public enum RestoreResult { case ok, missingKey(String), typeMismatch(String) }
|
|
public static let shared = StateServer()
|
|
public var restores: [Restore] = []
|
|
public var reads: [String: () -> Any?] = [:]
|
|
public var writes: [String: (Any) -> Bool] = [:]
|
|
private init() {}
|
|
public func register(buildId: String, accessorHash: String, atomicRestore: @escaping Restore) {
|
|
restores.append(atomicRestore)
|
|
}
|
|
public func registerAccessor(
|
|
key: String,
|
|
type: String,
|
|
read: @escaping () -> Any?,
|
|
write: @escaping (Any) -> Bool
|
|
) {
|
|
reads[key] = read
|
|
writes[key] = write
|
|
}
|
|
public func restoreAll(_ keys: JSONDict) -> RestoreResult {
|
|
for restore in restores {
|
|
let result = restore(keys, false)
|
|
guard case .ok = result else { return result }
|
|
}
|
|
for restore in restores {
|
|
let result = restore(keys, true)
|
|
guard case .ok = result else { return result }
|
|
}
|
|
return .ok
|
|
}
|
|
}
|
|
`);
|
|
const emitCore = spawnSync('swiftc', [
|
|
'-emit-library', '-emit-module', '-parse-as-library',
|
|
'-module-name', 'DebugBridgeCore', coreSource,
|
|
'-emit-module-path', coreModule,
|
|
'-o', coreLibrary,
|
|
], { encoding: 'utf8', timeout: 120_000 });
|
|
if (emitCore.status !== 0) throw new Error(`failed to build runtime stub:\n${emitCore.stderr}`);
|
|
|
|
const appSource = join(workDir, 'OptionalRoundTrip.swift');
|
|
writeFileSync(appSource, `
|
|
import Foundation
|
|
import Observation
|
|
import DebugBridgeCore
|
|
|
|
@Observable
|
|
final class AppState {
|
|
// @Snapshotable
|
|
var nickname: String? = nil
|
|
// @Snapshotable
|
|
var count: Int = 1
|
|
}
|
|
|
|
@Observable
|
|
final class FeatureState {
|
|
// @Snapshotable
|
|
var enabled: Bool = false
|
|
}
|
|
|
|
${render([
|
|
{
|
|
className: 'AppState',
|
|
fields: [
|
|
{ name: 'nickname', typeText: 'String?' },
|
|
{ name: 'count', typeText: 'Int' },
|
|
],
|
|
},
|
|
{
|
|
className: 'FeatureState',
|
|
fields: [{ name: 'enabled', typeText: 'Bool' }],
|
|
},
|
|
], 'fallback-build', 'schema-hash')}
|
|
|
|
@main
|
|
struct Runner {
|
|
@MainActor static func main() {
|
|
let state = AppState()
|
|
let feature = FeatureState()
|
|
AppStateAccessor.register(state)
|
|
FeatureStateAccessor.register(feature)
|
|
guard StateServer.shared.reads["nickname"]?() is NSNull else { fatalError("nil read") }
|
|
guard StateServer.shared.writes["nickname"]?("Ada") == true, state.nickname == "Ada" else {
|
|
fatalError("optional write")
|
|
}
|
|
guard StateServer.shared.writes["nickname"]?(NSNull()) == true, state.nickname == nil else {
|
|
fatalError("null write")
|
|
}
|
|
guard StateServer.shared.writes["count"]?(true) == false, state.count == 1 else {
|
|
fatalError("boolean must not coerce to integer")
|
|
}
|
|
guard StateServer.shared.writes["enabled"]?(1) == false, feature.enabled == false else {
|
|
fatalError("integer must not coerce to boolean")
|
|
}
|
|
let valid = try! JSONSerialization.jsonObject(
|
|
with: Data(#"{"nickname":"Grace","count":7,"enabled":true}"#.utf8)
|
|
) as! JSONDict
|
|
switch StateServer.shared.restoreAll(valid) {
|
|
case .ok: break
|
|
default: fatalError("valid restore")
|
|
}
|
|
guard state.nickname == "Grace", state.count == 7, feature.enabled else { fatalError("restore values") }
|
|
switch StateServer.shared.restoreAll(["nickname": NSNull(), "count": 8, "enabled": false]) {
|
|
case .ok: break
|
|
default: fatalError("null restore")
|
|
}
|
|
guard state.nickname == nil, state.count == 8, !feature.enabled else { fatalError("null restore values") }
|
|
state.nickname = "unchanged"
|
|
state.count = 9
|
|
feature.enabled = false
|
|
switch StateServer.shared.restoreAll(["nickname": "would-partially-apply", "count": 10, "enabled": 1]) {
|
|
case .typeMismatch("enabled"): break
|
|
default: fatalError("expected mismatch")
|
|
}
|
|
guard state.nickname == "unchanged", state.count == 9, !feature.enabled else { fatalError("cross-model partial mutation") }
|
|
}
|
|
}
|
|
`);
|
|
const executable = join(workDir, 'optional-round-trip');
|
|
const compile = spawnSync('swiftc', [
|
|
'-D', 'DEBUG', '-I', workDir, '-L', workDir, '-lDebugBridgeCore',
|
|
'-parse-as-library', appSource, '-o', executable,
|
|
], { encoding: 'utf8', timeout: 120_000 });
|
|
if (compile.status !== 0) throw new Error(`generated Optional accessor failed compilation:\n${compile.stderr}`);
|
|
const run = spawnSync(executable, [], {
|
|
encoding: 'utf8',
|
|
timeout: 30_000,
|
|
env: { ...process.env, DYLD_LIBRARY_PATH: workDir },
|
|
});
|
|
if (run.status !== 0) throw new Error(`generated Optional accessor failed at runtime:\n${run.stderr}`);
|
|
});
|
|
});
|
|
|
|
describe('SwiftSyntax generator parity', () => {
|
|
test('isolates canonical markers and rejects inaccessible fields', () => {
|
|
if (process.platform !== 'darwin') return;
|
|
if (spawnSync('swift', ['--version'], { encoding: 'utf8', timeout: 30_000 }).status !== 0) return;
|
|
|
|
const packageDir = join(import.meta.dir, 'gen-accessors-tool');
|
|
const inputDir = join(workDir, 'swift-syntax-input');
|
|
const outputDir = join(workDir, 'swift-syntax-output');
|
|
const cacheRoot = join(workDir, 'swift-syntax-cache');
|
|
mkdirSync(inputDir);
|
|
writeFileSync(join(inputDir, 'State.swift'), `
|
|
import Observation
|
|
@Observable
|
|
final class ToolState {
|
|
let documentation = """
|
|
// @Snapshotable
|
|
var stringLeak: String = "no"
|
|
"""
|
|
/* // @Snapshotable */
|
|
var blockLeak: Int = 1
|
|
var trailingLeak: Int = 2 // @Snapshotable
|
|
// prose about @Snapshotable must not opt in
|
|
var proseLeak: Int = 3
|
|
// @Snapshotable
|
|
var nickname: String? = nil
|
|
// @Snapshotable
|
|
var names:
|
|
[String]
|
|
= []
|
|
// @Snapshotable
|
|
var count: Int = 0
|
|
// @Snapshotable
|
|
var enabled: Bool = false
|
|
}
|
|
`);
|
|
const env = {
|
|
...process.env,
|
|
GSTACK_IOS_CACHE_ROOT: cacheRoot,
|
|
APP_BUILD_ID: 'syntax-test-build',
|
|
GEN_ACCESSORS_REV: 'syntax-test-v5',
|
|
};
|
|
const run = spawnSync('swift', [
|
|
'run', '--package-path', packageDir, 'gen-accessors',
|
|
'--input', inputDir, '--output', outputDir,
|
|
], { encoding: 'utf8', env, timeout: 180_000 });
|
|
if (run.status !== 0) throw new Error(`SwiftSyntax generator failed:\n${run.stderr}`);
|
|
const output = readFileSync(join(outputDir, 'StateAccessor.swift'), 'utf8');
|
|
expect(output).toContain('key: "nickname"');
|
|
expect(output).toContain('key: "names"');
|
|
expect(output).toContain('key: "count"');
|
|
expect(output).toContain('key: "enabled"');
|
|
expect(output).toContain('_GStackDebugBridgeSnapshotJSON.decode(raw0');
|
|
expect(output).toContain('atomicRestore: { keys, apply in');
|
|
expect(output).toContain('if apply {');
|
|
expect(output).not.toMatch(/raw\d+ as\?/);
|
|
expect(output).toContain('CFBundleShortVersionString');
|
|
expect(output).toContain(`accessorHash: "${computeAccessorHash([{
|
|
className: 'ToolState',
|
|
fields: [
|
|
{ name: 'nickname', typeText: 'String?' },
|
|
{ name: 'names', typeText: '[String]' },
|
|
{ name: 'count', typeText: 'Int' },
|
|
{ name: 'enabled', typeText: 'Bool' },
|
|
],
|
|
}])}"`);
|
|
expect(output).not.toContain('key: "stringLeak"');
|
|
expect(output).not.toContain('key: "blockLeak"');
|
|
expect(output).not.toContain('key: "trailingLeak"');
|
|
expect(output).not.toContain('key: "proseLeak"');
|
|
|
|
const invalidInput = join(workDir, 'swift-syntax-invalid');
|
|
const invalidOutput = join(workDir, 'swift-syntax-invalid-output');
|
|
mkdirSync(invalidInput);
|
|
writeFileSync(join(invalidInput, 'Invalid.swift'), `
|
|
import Observation
|
|
@Observable
|
|
final class InvalidState {
|
|
// @Snapshotable
|
|
public private(set) var token: String = "secret"
|
|
// @Snapshotable
|
|
let immutable: Int = 1
|
|
// @Snapshotable
|
|
var inferred = 2
|
|
// @Snapshotable
|
|
var legacy: String! = nil
|
|
// @Snapshotable
|
|
var custom: Date = .now
|
|
}
|
|
|
|
enum Namespace {
|
|
@Observable
|
|
final class NestedState {
|
|
// @Snapshotable
|
|
var nestedValue: Int = 0
|
|
}
|
|
}
|
|
|
|
@Observable
|
|
final class FirstState {
|
|
// @Snapshotable
|
|
var shared: String = "first"
|
|
}
|
|
|
|
@Observable
|
|
final class DuplicateState {
|
|
// @Snapshotable
|
|
var shared: String = "duplicate"
|
|
}
|
|
`);
|
|
const invalid = spawnSync('swift', [
|
|
'run', '--package-path', packageDir, 'gen-accessors',
|
|
'--input', invalidInput, '--output', invalidOutput,
|
|
], { encoding: 'utf8', env, timeout: 180_000 });
|
|
expect(invalid.status).toBe(4);
|
|
expect(invalid.stderr).toContain('InvalidState.token cannot be private');
|
|
expect(invalid.stderr).toContain('InvalidState.immutable must be declared var, not let');
|
|
expect(invalid.stderr).toContain('InvalidState.inferred requires an explicit type annotation');
|
|
expect(invalid.stderr).toContain('InvalidState.legacy cannot use an implicitly unwrapped Optional type');
|
|
expect(invalid.stderr).toContain("InvalidState.custom uses unsupported non-JSON snapshot type 'Date'");
|
|
expect(invalid.stderr).toContain('nested @Observable class NestedState');
|
|
expect(invalid.stderr).toContain("snapshot key 'shared' is declared by both FirstState and DuplicateState");
|
|
expect(existsSync(join(invalidOutput, 'StateAccessor.swift'))).toBe(false);
|
|
}, 180_000);
|
|
});
|
|
|
|
describe('collectSwiftFiles', () => {
|
|
test('walks subdirectories and finds all .swift files sorted', () => {
|
|
const a = join(workDir, 'a.swift');
|
|
const sub = join(workDir, 'sub');
|
|
mkdirSync(sub);
|
|
const b = join(sub, 'b.swift');
|
|
const c = join(workDir, 'c.txt');
|
|
writeFileSync(a, 'a');
|
|
writeFileSync(b, 'b');
|
|
writeFileSync(c, 'c');
|
|
const files = collectSwiftFiles(workDir);
|
|
expect(files.sort()).toEqual([a, b].sort());
|
|
});
|
|
|
|
test('excludes the configured output subtree and every StateAccessor.swift', () => {
|
|
const source = join(workDir, 'App.swift');
|
|
const staleAccessor = join(workDir, 'old', 'StateAccessor.swift');
|
|
const outputDir = join(workDir, 'custom-output');
|
|
mkdirSync(join(workDir, 'old'));
|
|
mkdirSync(outputDir);
|
|
writeFileSync(source, 'struct App {}');
|
|
writeFileSync(staleAccessor, '// stale generated output');
|
|
writeFileSync(join(outputDir, 'Poison.swift'), '// must not be scanned');
|
|
|
|
expect(collectSwiftFiles(workDir, { outputDir })).toEqual([source]);
|
|
});
|
|
});
|