The version gate caught a live queue collision (its whole job); same
MINOR bump level, next free slot per bin/gstack-next-version.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The sliced lane's first run (PR #2721) did its job: the planner and
report worked, the manifest governed, and every failure had a name.
Three were fixable on the spot:
- executor + gate-census checkouts get fetch-depth: 0 — files with
SELF-derived selection (the LLM-judge map, routing) walk git at
module load, and selection is deliberately fail-closed on git errors,
so the shallow checkout crashed those shards ('ambiguous argument
main...HEAD'). The manifest still governs WHICH shards run.
- landscape --toc gate: the exact toBe(3) landscape-page count was
font-metric-dependent (3 on Amazon Linux, 2 on ubuntu CI — the same
disease the file's own page-index comment warns about). Now a
comparative invariant: --toc must not CHANGE the landscape count vs
a baseline render.
- paid-run-manifest parse test builds its manifest under EVALS_ALL so
it never walks git (proven with GIT_DIR=/nonexistent).
Remaining first-run failures are newly-exposed rot in gate files that
had never executed in CI (skillify D1 refusal, session-intelligence
context-restore, one tpa-apple-ban retry flake) — being probed
separately; they are the lane WORKING, not the lane failing.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Version + release notes for the audit-and-overhaul branch: every
silently-skipping or never-running test class fixed and tripwired, the
free suite duration-packed with the serial mutator shard dissolved, the
paid lane re-platformed onto the sharded runner (planner/slices/
fail-closed report, parity phase), the weekly all-periodic coverage
contract, eval-budget timeout tiers, and 95 new coverage tests.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
- bun run test: duration-packed shards + --record-durations; the
trailing serial tree-mutating shard no longer exists
- two-tier system: the sliced CI lanes (one engine local+CI), the
weekly all-periodic coverage contract + exclusions, the gate census
- periodic detach timeout 32400 → 37800
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
LPT-packed free suite + --record-durations, the emptied TREE_MUTATING
mechanism, the sharded paid runner as the single selection engine,
CI planner/executor/report with the fail-closed report and hollow-shard
guard, the weekly coverage contract + exclusions policy, and the
eval-budgets timeout tiers.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
evals-periodic.yml re-platforms onto the sharded runner: planner
manifest → 6 executor slices → FAIL-CLOSED report. This IS the coverage
contract: all ~70 periodic-tier files weekly (EVALS_ALL=1), killing the
silent-rot class where a hard-coded 9-file matrix left ~57 files
running NOWHERE (the autoplan E2E rotted invisibly for months).
- test/helpers/periodic-exclude-data.ts: reasoned exclusions in their
OWN literals file (deliberately not touchfiles-data — map-diff
evaluates old versions of that file standalone). Every entry carries
reason + tracking with a re-entry condition; the runner surfaces each
exclusion per run; policy test pins real-file + non-empty fields.
Initial: ship-idempotency + brain-privacy-gate (documented-red,
never green) and skill-e2e-ios (manual hardware). The TODOS 'sidebar
E2E trio' turned out already deleted — only tombstone tests remain.
- gate-census job: weekly EVALS_ALL gate-tier run — PR lanes are
diff-billed, so without this the full gate census might never execute
anywhere; with the hollow-shard guard it is a census-health check
(exit 0 + zero executed tests fails), not just a test run.
- failure notification is a concrete gh issue UPSERT (one tracking
issue, commented per red week — never issue-per-week spam), with
issues:write scoped to the report job.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The parity-phase re-platform: evals.yml gains a second, sliced lane
driven by scripts/test-paid-shards.ts — the SAME engine local
eval:bg:gate uses, so CI and local share one selection engine.
- plan-slices: ONE planner (fetch-depth 0 — the only job needing
history) emits the manifest; selection fails open to run-all, never
per-slice (the divergence class is structurally dead)
- eval-slices: 6-way matrix consuming the manifest; PTY seed +
skill-registration steps run unconditionally (idempotent — a sliced
lane cannot key them on suite names); aggregate spawn budget
6 x EVALS_JOBS=2 x EVALS_CONCURRENCY=2 = 24 lane-wide (the matrix's
40-way per row queued session startup behind 39 siblings — the
timeout-flake family root); slice results + spooled shard logs
uploaded as artifacts
- slices-report: reconciles slice artifacts against the manifest
FAIL-CLOSED via --report — a slice whose artifact never landed, or a
planned shard nobody reported, is a failure, not an absence
- sequenced needs: evals so provider concurrency never doubles while
both lanes coexist; the matrix + its ratchets are deleted after
demonstrated parity (intersection + expected-additions comparison)
- workflow_dispatch gains evals_all (default true) for parity runs and
post-merge smokes — a dispatch can never silently select zero
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
One PLANNER computes diff selection + the slice plan ONCE and writes a
manifest (--emit-plan <path> --slices K); K executors consume it
(--plan <path> --slice i), never self-selecting, and write slice-result
artifacts; a REPORT reconciles results against the manifest (--report
<dir>) fail-closed: a slice whose artifact never landed is a FAILURE,
a planned shard nobody reported fails, wrong-slice/duplicate/cross-tier
results fail. Kills per-slice selector divergence and hollow-lane
aggregation at the root.
- hollow-shard guard: under EVALS_ALL, exit 0 with ZERO executed tests
(bun's 'Ran N tests' now captured by the classifier — additive) is
'passed-empty' and fails the run; selective runs keep it 'passed'
with one warning (in-file diff/tier self-skips are legitimate there);
unknown counts are never guessed hollow
- retry parity: --retry 1 default + RETRY_OVERRIDES literals for the
three files whose old matrix rows earned retries: 2 (stale entries
pinned against disk)
- live smoke: gate plan = 48 shards across 6 slices; report mode exits
1 on a fabricated missing slice, 0 when complete
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Recorded via --record-durations on a quiescent tree: 479s serial
total, p50 92ms / p90 1.8s / max 31.4s — the top-heavy cost shape LPT
packing exists for. A hint, not a contract: refresh opportunistically
with bun run test:free --record-durations.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
codex render: porcelain unchanged AND out-dir gstack-ship/SKILL.md
byte-identical to a fresh in-place render (+openai.yaml presence);
--host all render: exit 0, porcelain unchanged, claude + .agents +
.factory + llms.txt + openclaw docs all present in the out-dir.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Zero mutators remain (all eight render into out-dirs now), so the four
ratchet READERS (parity caps, size budgets, carve parity/ordering) get
a quiet tree by construction in any shard and rejoin the parallel
phase. The ~35-40s serial tail on every full-suite run is gone. The
mechanism stays: a future test that genuinely must write shared
artifacts in place earns an entry with a reason and is serialized
again; the census pin still fails on renamed keys.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Two renders into two separate out-dirs, EVERY file diffed byte-for-byte
(claude-only and --host all; normalization only for each dir's own
sanctioned section-base repoint; presence-sanity lists guard against a
vacuous empty-dir pass) — strictly stronger than the old in-place
double-regen that sampled 5 files. TREE_MUTATING entry deleted.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The full-catalog smoke no longer rewrites all 71 SKILL.md then
regenerates to restore (with its 'CRITICAL: failed to restore' prayer
path) — it renders into a mkdtemp and additionally asserts tracked
ship/SKILL.md is byte-unchanged. TREE_MUTATING entry deleted.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
regenAndSnapshot renders --host claude --out-dir <mkdtemp> (+
--respect-detection) and snapshots probes from the out-dir. The
git-restore machinery is deleted outright — it restored only
PROBE_FILES of the 71 files each call wrote, so a stale tree kept the
other 68 dirty (the partial-restore bug), and its 'no output-path arg'
comment had been false since --out-dir landed. TREE_MUTATING entry
deleted.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Its goldens were 'produced by gen-skill-docs.test.ts' with a
when-missing beforeAll fallback that wrote the live tree — an
inter-test ordering dependency the serial shard hid. It now renders
codex+factory UNCONDITIONALLY into its own out-dir and reads goldens
only from there (the Claude golden deliberately keeps reading tracked
ship/SKILL.md — a read; out-dir claude renders repoint section-base
paths by design). TREE_MUTATING entry deleted.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
gen-skill-docs.test.ts's 15 in-place generator spawns now render into
mkdtemp out-dirs (gitignored-artifact reads repointed; the handshake
scan's silent console.warn degrade became a hard assertion); its
tracked-tree reads (freshness dry-run, SKILL.md content pins) stay
reads. catalog-trim needed no change beyond the earlier main() guard —
its import is now side-effect-free (pinned by the import-purity test).
Both TREE_MUTATING entries deleted in this commit, per the transition
rule: an entry leaves in the same commit as the file's last in-place
write.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
- watchdog.test: the 20s blind wait for one production parent-watchdog
tick becomes BROWSE_PARENT_WATCHDOG_INTERVAL_MS=250 (new env knob in
server.ts, NaN-safe, production default unchanged) + polls for the
boot line and the tick's stay-alive log — strictly stronger (the old
form never proved a tick observed the parent death). 24s → 3.6s.
- stop-dead-daemon / terminal-agent-owner-watchdog: the 300s/30s
stand-in child lifetimes become stdin-EOF-bound — the child can never
self-exit mid-test on a slow runner (spurious-failure class) and
self-reaps instantly if the test dies (no 300s orphans). Node-compat
stdin APIs (owner-watchdog runs on the Windows lane).
- browser-skill-commands: the sleeper fixture's 30s self-time becomes
8s (no stdin pipe exists in runToFiles) — far above the 1s product
timeout it must outlive, below the test ceiling, so a timeout-kill
regression fails on clean assertions instead of an opaque bun
timeout; added: stdout must NOT contain 'done'.
45/45 green across the four files + server tripwires.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The sharded runner computed diff selection once, then each of its 48-73
children recomputed it at module load — including, on touchfiles-diff
branches, a per-child bun subprocess evaluating the old data file (20s
timeout each). The parent now serializes {version, selected, reason} as
EVALS_SELECTION_JSON into the shard env; e2e-helpers adopts it at load.
Fail-open preserved: any parse/shape violation → ONE stderr warning +
local recompute; absent env → silent local compute (non-sharded
entrypoints unchanged). Drift test pins parent→child round-trip to
identical selection decisions plus the malformed/absent cases.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
- runPaidShard no longer buffers whole 30-min stream-json streams in
RAM (x concurrent jobs): every byte tees to a per-shard log file
(slug-named, path printed at START for mid-run inspection and on the
FAILED terminal line); failures print a 64KiB tail read back from
disk; passing shards stay quiet (the file is the record) — the free
runner's proven contract. Classification unchanged: the strict
classifier still sees every byte first.
- the ~35 duplicated spawn/group-kill/wall-timer/finally-reap lines
move into runShardChild in test-strict-output.ts (detached-per-
platform spawn, signal forwarding, SIGKILL group kill at the wall,
drain-before-verdict); designed so the free runner can migrate later
- expectedFiles drift fixed toward ENFORCEMENT: the injected-command
exemption is gone — a fake command exiting 0 without bun's terminal
summary now reads FAILED (pinned: silent-pass → failed)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
- eval CLI family (eval-list/compare/summary + eval-select smoke): the
primary interface to eval results had no tests; isolation via a fake
gstack-slug under a mkdtemp HOME (the scripts' real resolution path —
they do NOT honor GSTACK_EVAL_DIR; only EvalCollector does). Pinned
current behavior: eval-list does NOT exclude _partial runs (documented
improvement candidate)
- slop-diff (runs on every /review + quality-gate): fixture git repo +
first-on-PATH npx stub (never downloads real slop-scan); no-diff
early exit, missing-scanner fallback, fingerprint line-insensitivity,
merge-base worktree scan
- bin/gstack-code-intelligence CLI arg surface (lib was covered, the
284-line CLI wasn't): select/consent/suggest/index/search gating;
pinned: --help routes to usage failure exit 1 (no handler)
- browse media-extract: the page.evaluate callback exercised in-process
against a mock DOM (no exports added) — lazy-src fallback chain,
HLS/DASH detection, bg-image url() parsing, 500-element cap
- browse session-cookie-store: factory contract (cookieName/ttlMs/
maxSessions eviction, cross-store isolation, mint→validate
round-trip); store is in-memory — no fs cases exist
- lib/version-source direct unit tests (gstack-version-bump.test.ts
spawns the bin, never imports the lib): parse/format/cmp/bump
coercion, npm 4→3 translation, #2501 mangled-JSON regression class
All hermetic (mkdtemp homes, runBin child isolation); windows curation
correctly partitions the six.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
69 files, both shapes (trailing bun-test budgets and runner
timeout/timeoutMs options), ROUND-UP ONLY so nothing that passed can
start failing: 75 → JUDGE_MS, 137 → CAPTURE_MS, 74 → CAPTURE_LONG_MS,
9 → PTY_MS, 3 → PTY_LONG_MS. Raw >=60s literal count in the paid scope:
395 → 97, of which 51 are non-timeout noise (fixture dates, run IDs)
and 46 are enumerated justified holds (comment-carrying calibrated
budgets, poll-loop constants, utility spawn waits, and the seven
physical-ceiling 1_500_000 sites). The eval-budgets policy ratchet
keeps the residue from regrowing.
Known collapse: where an inner runner budget and its enclosing test
budget now share a tier, the old stagger is gone — an overrun surfaces
as a bun test timeout instead of a graceful runner timeout
(diagnosability trade, not a correctness one).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
~36 free test files each carry a near-identical local run() (spawnSync
+ utf-8 + {status, stdout, stderr}) differing only in env composition,
cwd, and timeout. runBin absorbs the invariant core; options carry the
variance (gstackHome sets BOTH GSTACK_HOME and GSTACK_STATE_DIR — the
config-precedence trap several locals rediscovered independently; home
for $HOME-anchored bins; input/trim/timeout/maxBuffer). Free-test-only
by design so it never becomes a de facto global touchfile. Migration of
the 36 call sites lands separately (mechanical batches).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Five named tiers (JUDGE 120s / CAPTURE 300s / CAPTURE_LONG 600s /
PTY 900s / PTY_LONG 1200s) replace hand-ratcheted sprawl (46x300s,
46x120s, 44x360s, 44x180s, 27x240s, 19x150s, 13x420s, 12x600s...),
much of it inflated to paper over the old 40-way in-shard concurrency
that the sharded runner's 1-file-per-shard model kills. Policy test
pins: every tier fits the shard wall minus 120s overhead (the
structural fix for budgets-above-the-wall fiction), tiers stay ordered,
and no paid literal exceeds PTY_LONG x1.25 — oversized tests get split,
not budgeted past the wall.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
'bun run test' silently appended up to two 120s npx slop-scan runs plus
a git worktree add/remove after the suite (2>/dev/null || true) —
invisible in the documented '~90-100s' timing and pure friction in the
pre-commit loop. Decoupling is not coverage removal: quality-gate.yml
now runs slop:diff on every PR (advisory, matching its in-repo 'never
blocking' contract), and /review already invokes it explicitly.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Hash sharding balances file COUNTS (1.15x spread) but not cost — the
Playwright-launching files landed 4/3/4/1/2/1 across 6 shards, giving a
measured 28s–97s shard spread and ~40s of idle tail on every run.
Full-suite mode now packs by recorded per-file durations
(longest-processing-time-first) when the committed seed
scripts/free-test-durations.json exists.
- ONE store, no overlay: the seed is refreshed occasionally via the new
--record-durations mode (each file timed in its own child — exact,
and immune to bun's stream buffering, where silent passers print no
header to timestamp); GSTACK_FREE_TEST_DURATIONS overrides the path
for experiments; CI never records
- seed is a hint: missing → silent hash-shard fallback; corrupt (bad
merge) → one warning + fallback; unknown files → 75th-percentile
pessimism so a surprise long-runner can't recreate the tail
- packed shards get duration-aware walls (max(base, predicted x 3)) —
LPT decouples count from cost BY DESIGN, so the 5s/file heuristic
would undersize a shard holding few expensive files
- one log line per shard (files + predicted seconds) so packing
regressions are diagnosable from any run log
- the --shard CI-matrix path is untouched: stable hash indices are its
contract
- successor note in-code: bun >=1.3.14 ships native --timings/--shard
LPT — swap this packer when the repo unpins 1.3.13
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The self-registration sweep made these eight files' gate-tier keys
visible to the census for the first time — their gate tests run in NO
CI lane today (pre-existing hole, newly measurable). Ratcheted into
KNOWN_MATRIX_GAPS with the burn-down note: the paid-lane re-platform
runs every gate file by construction and retires this ratchet class.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
129-of-177 keys omitted their own test file, so editing only a test's
prompt or assertions selected NOTHING — the changed test never ran on
the change that changed it. 135 keys self-registered (110 E2E + 25
LLM-judge), resolved by strict declaration evidence (testName:/
testIfSelected/judge call sites), with skill-name false positives
excluded.
e2e-tier-alignment's warn-only branch for unregistered files is now a
hard failure with a 4-entry KNOWN_UNREGISTERED ratchet (template-
literal testNames, fail-open-safe) + a burn-down test so the set only
shrinks. Selection sanity: a one-file diff on skill-e2e-qa-workflow now
selects its 4 tests (was 0); skill-llm-eval 0 → 25.
Known follow-ups (filed): 15 E2E + 2 judge PHANTOM keys select tests
that exist nowhere; codex-e2e-plan-format's testIfSelected names have
no map keys (run-all only).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
--out-dir was Claude-host-only (gen-skill-docs.ts:842), which forced
the codex/factory-regenerating tests (gen-skill-docs, skill-validation,
host-config) to mutate the live tree — the reason they sit in the
TREE_MUTATING serial shard. The flag now mirrors ALL outputs into the
out-dir: external-host trees (.agents/.factory/... via
processExternalHost), external section files, openclaw docs, and
gstack/llms.txt (a catalog-mode render must never rewrite the tracked
index). OUTPUTS ONLY — inputs (templates, sections/, host configs) are
always read from ROOT, so an empty out-dir can never feed the render.
rewriteSectionBase stays Claude-only (external hosts have their own
path grammar).
Proofs: in-place --host all is byte-identical (tree clean);
--host all --out-dir <mkdtemp> renders the full multi-host tree with
ROOT untouched; gen-skill-docs-out-dir tests + 415/415
gen-skill-docs.test.ts green (bin/dev-setup's claude rendering
byte-compat).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The generator's whole body executed at module load, so any import of it
(test/gen-skill-docs.test.ts pulls assertSinglePreamble via require();
test/catalog-trim.test.ts imports helpers) regenerated all 71 SKILL.md
in place — the root cause of half the TREE_MUTATING serial-shard
entries (hazard class #2532). The body now lives in an exported
main(): number behind if (import.meta.main).
Semantics preserved exactly: failure exits are immediate (matching the
old top-level process.exit), success leaves the event loop to drain so
the llms.txt fire-and-forget IIFE finishes its write, and the module
stays synchronous/require()-able. Proofs: byte-identical --host all
output (git status clean), --dry-run stale-tree still exits 1 (the
skill-docs freshness lane depends on it), and the new
test/gen-skill-docs-import-purity.test.ts pins load-time purity via a
subprocess probe (mtime-based, so a dirty worktree can't false-fail).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
1,700,000ms (28.3 min) exceeded every wall these tests run inside: the
25-min CI job timeout and the 1800s sharded-runner wall (which also
leaves --retry 1 zero room for a second attempt). Budget above the wall
is fiction, not headroom — a test that actually used it produced a
job-level kill (no bun summary, no artifact) instead of a clean
per-test timeout. No recorded p95 exists for this family (they are
being retiered to periodic in the re-platform wave); the trim stops at
the physical ceiling rather than guessing lower. Final policy lands in
the Wave-2 eval-budgets constants module.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Six paid PTY tests wrote their expected plan artifact to a FIXED shared
/tmp path ('/tmp/gstack-test-plan-<mode>.md') and rmSync'd it in
finally — under --retry 1, EVALS_JOBS>1, or two concurrent worktrees, a
sibling's cleanup deletes this run's artifact and the D19 'agent did
not produce expected plan file' assertion fires spuriously. Each test
now mkdtemps its own dir, interpolates the unique path into the agent
prompt (fixture-sourced prompts get a replaceAll + drift guard that
throws if the fixture's literal ever moves), and cleans up its own dir.
The 18 cwd:-into-the-live-repo sites were audited: all deliberate
(skill registry + hermetic pre-trusted dir, in-repo gen renders, git
history reads, slug resolution) — each now carries a
'// LIVE-REPO CWD: <reason>' comment so the next audit can tell
deliberate from accidental.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Closes the Windows Free Tests red: recent lane failures showed a
platform-unguarded POSIX mode-bit assertion ('Expected: 493' — a
symlink-skip test) from PR-branch variants; the KNOWN_WINDOWS_SAFE
force-include reason ('mode-bitmask hits are POSIX-branch only') did
not hold for that shape, and main had neither the guard nor the
behavior.
- product: lstat first; a symlinked dir gets a warning and a skip on
both platforms — chmod AND icacls dereference the link, so
restricting through a symlink hardens an unvetted target (and
/inheritance:r could lock out its real owner). All callers already
treat hardening as best-effort (try/catch).
- test: the symlink regression test, platform-aware — symlinkSync in
the house try/catch skip pattern (Windows runners without Developer
Mode can't create symlinks), mode-bit assertion guarded off win32,
behavior assertions (no throw, warning text, target readable)
everywhere; POSIX still proves the skip (0o755 unchanged, not 0o700)
- KNOWN_WINDOWS_SAFE reason updated to the now-true premise
20/20 pass on Linux.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
carve-section-loading, codex-e2e-plan-format,
codex-e2e-recommendation-substance, and llm-judge-recommendation gated
on EVALS/tier (free suite loads them as describe.skip) but their names
fell outside PAID_TEST_GLOBS, so no paid lane ever selected them — net
execution zero, forever. The existing matrix tripwire filtered on
isPaidTestFile() first, so it was blind to exactly this class (the same
bug that hid the pre-split monolith's gate tests for ~8 releases).
- PAID_TEST_GLOBS: codex-e2e* + skill-llm-eval* wildcards (replacing
exact names) + llm-judge-recommendation + carve-section-loading;
package.json's six test-script glob lists mirrored
- codex-e2e-plan-format gains the explicit periodic tier gate its
siblings carry (external-service rule) — without it the sharded
runner's no-guard default would spawn Codex in the gate tier per PR
- eval:bg:periodic --timeout 32400→37800: the census growth pushed the
periodic worst case to 35910s; the old value had 270s of headroom
BEFORE this change and would now kill healthy runs mid-flight
- new test/paid-orphan-tripwire.test.ts: any EVALS/tier-gated test file
outside the globs fails the free suite (reasoned SCANNER_EXEMPT for
the gate helpers + meta-tests) — the class-killer
- paid-shards pins updated: the four orphans now assert INSIDE the
census
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
extension-sender-auth's two privileged-message denial tests (content
script + missing sender.url — the extension's security boundary) and
snapshot's three skips were quarantined 'pre-existing' failures. Root
cause: machine-local state on the quarantining dev machines — the test
and gate code are byte-identical between the quarantining commit
(410b4928) and HEAD, and all five pass deterministically on a clean
checkout (68/68 across both files, multiple runs). No assertions
weakened, no product changes.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
skill-e2e-spec-execute (600s budget) and skill-llm-eval-spec (300s)
reported PASS on every periodic run while asserting nothing. Deleting
them would remove the periodic-tier selector surface they exist to
register (diff-based selection for spec/ changes), so they become
test.todo — reported as todo/skip, never pass — with the v1.1
implementation specs kept in-file.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
callJudge inlined GSTACK_EVAL_MODEL_JUDGE || sonnet, silently ignoring
the global GSTACK_EVAL_MODEL override every other eval call site honors
via lib/eval-model.ts. New 'judge' kind in DEFAULTS (sonnet — the D1a
pin-on-regressors calibration stands; model CHOICE unchanged) and
callJudge resolves through it: explicit arg > GSTACK_EVAL_MODEL_JUDGE >
GSTACK_EVAL_MODEL > default.
429 handling upgraded from one fixed 1s retry (reliably lost races at
CI concurrency) to three jittered exponential retries (~1s/4s/16s),
honoring the server's retry-after when present.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The hand-maintained 39-name SKILL_GLOBS list had drifted to 39 of 54
SKILL.md.tmpl on disk. No live gap today (none of the 15 unlisted
skills is interactive), but the next interactive skill would have
landed unguarded with zero signal. The audit now walks top-level dirs
for SKILL.md.tmpl (statSync so symlinked dirs like connect-chrome
count), so new skills are in scope the commit they appear.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Column-0 assignment of GSTACK_HOME / GSTACK_STATE_ROOT in any tracked
*.test.ts fails with the file:line and the fix (beforeAll + afterAll
restore). Kills the cross-file env-leak class the previous commit
swept.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Five files assigned process.env.GSTACK_HOME at module scope. Shard
processes evaluate sibling modules before running their tests, so the
assignment leaked into every other file in the shard — the damage was
already visible in defensive workarounds (relink.test.ts:28 'fresh
install test saw a neighbor's skill_prefix'; cdp-e2e's own comment
documents a sibling's temp dir baked into artifacts).
Pattern: save original, assign in beforeAll, restore in afterAll
(cdp-e2e already restored but still assigned at load — its window now
matches the others). GSTACK_TELEMETRY_OFF and GSTACK_PROJECT_SLUG get
the same treatment where they rode along. Victim files' defenses stay
in place (cheap insurance).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
- timeout-minutes on the 6 remaining unbounded jobs (actionlint 5,
skill-docs 10, version-gate 10, make-pdf-gate 15, pr-title-sync 5,
evals build-image 15) — a hung step sat on GitHub's 360-min default
- right-size measured-over-long timeouts: dependency-review 10→5,
windows-setup-e2e 15→10
- dependency-review: 2-core runner (28s API call on an 8-core box) and
drop .github/workflows/** from its trigger paths (workflow edits have
no dependencies to review)
- windows caches gain restore-keys: a lockfile bump paid the 26s/43s
restore for a guaranteed cold miss
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
fetch-depth:0 cost 74 of the job's 92 seconds; the three gates it feeds
take ~12s combined. Shallow checkout + exact-SHA fetches for the diff's
base/head (an exact-SHA fetch, not a guessed depth — long-lived
branches and merge queues still resolve), with a --deepen fallback for
push events whose 'before' is unusable. timeout right-sized 20→10 min.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
- package.json out of the trigger paths: the tag hash deliberately
excludes it (version bumps every ship), so every merge rebuilt and
re-pushed the IDENTICAL tag (~2m26s for zero content change);
patches/** added (it IS a tag input)
- manifest existence check (mirrors evals.yml): tag already exists →
skip the build
- concurrency group: two rapid main pushes raced pushing the same
:latest/:buildcache tags
- cron staggered 06:00→04:00 Monday: it shared the exact minute with
evals-periodic, which could race a half-pushed tag or duplicate the
build
- timeout-minutes: 30 (was unbounded → 360-min default for a hung
docker build)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
evals.yml, evals-periodic.yml, and ci-image.yml each compute the CI
image tag from hashFiles('.github/docker/Dockerfile.ci', 'bun.lock',
'patches/**') — synced by comment only (TODOS.md 'CI three-way
image-tag drift'). If one input list drifts, that workflow computes a
different tag for the same content: eval lanes silently rebuild the
image every run, or ci-image prebuilds a tag nobody looks up. The test
extracts each tag-computation site and fails on any mismatch.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>