mirror of
https://github.com/garrytan/gstack.git
synced 2026-08-31 10:20:42 +02:00
9e869b2ba30bafa2b06ab319bf0b5abdb67e9f8b
426
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
9e869b2ba3 |
docs: TESTING_INTERNALS covers the 2026-08 runner overhaul
LPT-packed free suite + --record-durations, the emptied TREE_MUTATING mechanism, the sharded paid runner as the single selection engine, CI planner/executor/report with the fail-closed report and hollow-shard guard, the weekly coverage contract + exclusions policy, and the eval-budgets timeout tiers. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
9f2ee58d38 |
feat(ci): weekly periodic lane runs EVERY periodic test + gate census backstop
evals-periodic.yml re-platforms onto the sharded runner: planner manifest → 6 executor slices → FAIL-CLOSED report. This IS the coverage contract: all ~70 periodic-tier files weekly (EVALS_ALL=1), killing the silent-rot class where a hard-coded 9-file matrix left ~57 files running NOWHERE (the autoplan E2E rotted invisibly for months). - test/helpers/periodic-exclude-data.ts: reasoned exclusions in their OWN literals file (deliberately not touchfiles-data — map-diff evaluates old versions of that file standalone). Every entry carries reason + tracking with a re-entry condition; the runner surfaces each exclusion per run; policy test pins real-file + non-empty fields. Initial: ship-idempotency + brain-privacy-gate (documented-red, never green) and skill-e2e-ios (manual hardware). The TODOS 'sidebar E2E trio' turned out already deleted — only tombstone tests remain. - gate-census job: weekly EVALS_ALL gate-tier run — PR lanes are diff-billed, so without this the full gate census might never execute anywhere; with the hollow-shard guard it is a census-health check (exit 0 + zero executed tests fails), not just a test run. - failure notification is a concrete gh issue UPSERT (one tracking issue, commented per red week — never issue-per-week spam), with issues:write scoped to the report job. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
056bc61a26 |
feat(ci): sliced paid lane (planner -> 6 executors -> fail-closed report)
The parity-phase re-platform: evals.yml gains a second, sliced lane driven by scripts/test-paid-shards.ts — the SAME engine local eval:bg:gate uses, so CI and local share one selection engine. - plan-slices: ONE planner (fetch-depth 0 — the only job needing history) emits the manifest; selection fails open to run-all, never per-slice (the divergence class is structurally dead) - eval-slices: 6-way matrix consuming the manifest; PTY seed + skill-registration steps run unconditionally (idempotent — a sliced lane cannot key them on suite names); aggregate spawn budget 6 x EVALS_JOBS=2 x EVALS_CONCURRENCY=2 = 24 lane-wide (the matrix's 40-way per row queued session startup behind 39 siblings — the timeout-flake family root); slice results + spooled shard logs uploaded as artifacts - slices-report: reconciles slice artifacts against the manifest FAIL-CLOSED via --report — a slice whose artifact never landed, or a planned shard nobody reported, is a failure, not an absence - sequenced needs: evals so provider concurrency never doubles while both lanes coexist; the matrix + its ratchets are deleted after demonstrated parity (intersection + expected-additions comparison) - workflow_dispatch gains evals_all (default true) for parity runs and post-merge smokes — a dispatch can never silently select zero Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
f7ef0cc0ce |
feat(evals): planner/executor/report modes — the CI re-platform surface
One PLANNER computes diff selection + the slice plan ONCE and writes a manifest (--emit-plan <path> --slices K); K executors consume it (--plan <path> --slice i), never self-selecting, and write slice-result artifacts; a REPORT reconciles results against the manifest (--report <dir>) fail-closed: a slice whose artifact never landed is a FAILURE, a planned shard nobody reported fails, wrong-slice/duplicate/cross-tier results fail. Kills per-slice selector divergence and hollow-lane aggregation at the root. - hollow-shard guard: under EVALS_ALL, exit 0 with ZERO executed tests (bun's 'Ran N tests' now captured by the classifier — additive) is 'passed-empty' and fails the run; selective runs keep it 'passed' with one warning (in-file diff/tier self-skips are legitimate there); unknown counts are never guessed hollow - retry parity: --retry 1 default + RETRY_OVERRIDES literals for the three files whose old matrix rows earned retries: 2 (stale entries pinned against disk) - live smoke: gate plan = 48 shards across 6 slices; report mode exits 1 on a fabricated missing slice, 0 when complete Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
c7512da67c |
feat(test): commit the initial free-test durations seed (496 files)
Recorded via --record-durations on a quiescent tree: 479s serial total, p50 92ms / p90 1.8s / max 31.4s — the top-heavy cost shape LPT packing exists for. A hint, not a contract: refresh opportunistically with bun run test:free --record-durations. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
13ec721a18 |
test(gen): out-dir byte-identity + tree-clean pins for external hosts
codex render: porcelain unchanged AND out-dir gstack-ship/SKILL.md byte-identical to a fresh in-place render (+openai.yaml presence); --host all render: exit 0, porcelain unchanged, claude + .agents + .factory + llms.txt + openclaw docs all present in the out-dir. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
dadb0a4cb1 |
feat(test): the serial tree-mutating shard dissolves — TREE_MUTATING is empty
Zero mutators remain (all eight render into out-dirs now), so the four ratchet READERS (parity caps, size budgets, carve parity/ordering) get a quiet tree by construction in any shard and rejoin the parallel phase. The ~35-40s serial tail on every full-suite run is gone. The mechanism stays: a future test that genuinely must write shared artifacts in place earns an entry with a reason and is serialized again; the census pin still fails on renamed keys. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
9887f4c769 |
fix(test): spec-template-sync compares an out-dir render, not an in-place one
TREE_MUTATING entry deleted. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
dfb5b61889 |
fix(test): idempotency proof strengthens to two-out-dir recursive diff
Two renders into two separate out-dirs, EVERY file diffed byte-for-byte (claude-only and --host all; normalization only for each dir's own sanctioned section-base repoint; presence-sanity lists guard against a vacuous empty-dir pass) — strictly stronger than the old in-place double-regen that sampled 5 files. TREE_MUTATING entry deleted. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
1d945e74b9 |
fix(test): catalog-mode-full renders to out-dir; restore machinery deleted
The full-catalog smoke no longer rewrites all 71 SKILL.md then regenerates to restore (with its 'CRITICAL: failed to restore' prayer path) — it renders into a mkdtemp and additionally asserts tracked ship/SKILL.md is byte-unchanged. TREE_MUTATING entry deleted. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
39a0b61aaa |
fix(test): gbrain-detection-override drops mutate-then-git-restore
regenAndSnapshot renders --host claude --out-dir <mkdtemp> (+ --respect-detection) and snapshots probes from the out-dir. The git-restore machinery is deleted outright — it restored only PROBE_FILES of the 71 files each call wrote, so a stale tree kept the other 68 dirty (the partial-restore bug), and its 'no output-path arg' comment had been false since --out-dir landed. TREE_MUTATING entry deleted. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
5c525c9aca |
fix(test): host-config self-provisions goldens (ordering dependency severed)
Its goldens were 'produced by gen-skill-docs.test.ts' with a when-missing beforeAll fallback that wrote the live tree — an inter-test ordering dependency the serial shard hid. It now renders codex+factory UNCONDITIONALLY into its own out-dir and reads goldens only from there (the Claude golden deliberately keeps reading tracked ship/SKILL.md — a read; out-dir claude renders repoint section-base paths by design). TREE_MUTATING entry deleted. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
4409ca392f |
fix(test): skill-validation renders codex host into an out-dir
Its 3 in-place --host codex regeneration sites collapse into one module-level --out-dir render; assertions untouched. TREE_MUTATING entry deleted. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
1f6da056fe |
fix(test): gen-skill-docs + catalog-trim leave the serial mutator shard
gen-skill-docs.test.ts's 15 in-place generator spawns now render into mkdtemp out-dirs (gitignored-artifact reads repointed; the handshake scan's silent console.warn degrade became a hard assertion); its tracked-tree reads (freshness dry-run, SKILL.md content pins) stay reads. catalog-trim needed no change beyond the earlier main() guard — its import is now side-effect-free (pinned by the import-purity test). Both TREE_MUTATING entries deleted in this commit, per the transition rule: an entry leaves in the same commit as the file's last in-place write. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
5c081a355a |
fix(test): kill the four worst fixed sleeps (300s/30s/30s/20s)
- watchdog.test: the 20s blind wait for one production parent-watchdog tick becomes BROWSE_PARENT_WATCHDOG_INTERVAL_MS=250 (new env knob in server.ts, NaN-safe, production default unchanged) + polls for the boot line and the tick's stay-alive log — strictly stronger (the old form never proved a tick observed the parent death). 24s → 3.6s. - stop-dead-daemon / terminal-agent-owner-watchdog: the 300s/30s stand-in child lifetimes become stdin-EOF-bound — the child can never self-exit mid-test on a slow runner (spurious-failure class) and self-reaps instantly if the test dies (no 300s orphans). Node-compat stdin APIs (owner-watchdog runs on the Windows lane). - browser-skill-commands: the sleeper fixture's 30s self-time becomes 8s (no stdin pipe exists in runToFiles) — far above the 1s product timeout it must outlive, below the test ceiling, so a timeout-kill regression fails on clean assertions instead of an opaque bun timeout; added: stdout must NOT contain 'done'. 45/45 green across the four files + server tripwires. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
73fe950fbf |
feat(evals): parent-computed selection propagates to shard children
The sharded runner computed diff selection once, then each of its 48-73
children recomputed it at module load — including, on touchfiles-diff
branches, a per-child bun subprocess evaluating the old data file (20s
timeout each). The parent now serializes {version, selected, reason} as
EVALS_SELECTION_JSON into the shard env; e2e-helpers adopts it at load.
Fail-open preserved: any parse/shape violation → ONE stderr warning +
local recompute; absent env → silent local compute (non-sharded
entrypoints unchanged). Drift test pins parent→child round-trip to
identical selection decisions plus the malformed/absent cases.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
|
||
|
|
05bca51961 |
refactor(evals): paid shards spool to disk + shared runShardChild lifecycle
- runPaidShard no longer buffers whole 30-min stream-json streams in RAM (x concurrent jobs): every byte tees to a per-shard log file (slug-named, path printed at START for mid-run inspection and on the FAILED terminal line); failures print a 64KiB tail read back from disk; passing shards stay quiet (the file is the record) — the free runner's proven contract. Classification unchanged: the strict classifier still sees every byte first. - the ~35 duplicated spawn/group-kill/wall-timer/finally-reap lines move into runShardChild in test-strict-output.ts (detached-per- platform spawn, signal forwarding, SIGKILL group kill at the wall, drain-before-verdict); designed so the free runner can migrate later - expectedFiles drift fixed toward ENFORCEMENT: the injected-command exemption is gone — a fake command exiting 0 without bun's terminal summary now reads FAILED (pinned: silent-pass → failed) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
6ef8aaba65 |
refactor(test): first runBin migration batch (3 of ~36 run() duplicates)
explain-level-config, benchmark-cli, evidence move onto the shared helper; each file's remaining special-case spawnSync sites (raw-buffer probes, env-scrub probes) stay put deliberately. 55/55 green. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
1055561cae |
test: coverage fill — 95 tests for six zero-coverage surfaces
- eval CLI family (eval-list/compare/summary + eval-select smoke): the primary interface to eval results had no tests; isolation via a fake gstack-slug under a mkdtemp HOME (the scripts' real resolution path — they do NOT honor GSTACK_EVAL_DIR; only EvalCollector does). Pinned current behavior: eval-list does NOT exclude _partial runs (documented improvement candidate) - slop-diff (runs on every /review + quality-gate): fixture git repo + first-on-PATH npx stub (never downloads real slop-scan); no-diff early exit, missing-scanner fallback, fingerprint line-insensitivity, merge-base worktree scan - bin/gstack-code-intelligence CLI arg surface (lib was covered, the 284-line CLI wasn't): select/consent/suggest/index/search gating; pinned: --help routes to usage failure exit 1 (no handler) - browse media-extract: the page.evaluate callback exercised in-process against a mock DOM (no exports added) — lazy-src fallback chain, HLS/DASH detection, bg-image url() parsing, 500-element cap - browse session-cookie-store: factory contract (cookieName/ttlMs/ maxSessions eviction, cross-store isolation, mint→validate round-trip); store is in-memory — no fs cases exist - lib/version-source direct unit tests (gstack-version-bump.test.ts spawns the bin, never imports the lib): parse/format/cmp/bump coercion, npm 4→3 translation, #2501 mangled-JSON regression class All hermetic (mkdtemp homes, runBin child isolation); windows curation correctly partitions the six. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
6841183c35 |
refactor(test): mechanical sweep — 298 paid-test timeouts onto eval-budget tiers
69 files, both shapes (trailing bun-test budgets and runner timeout/timeoutMs options), ROUND-UP ONLY so nothing that passed can start failing: 75 → JUDGE_MS, 137 → CAPTURE_MS, 74 → CAPTURE_LONG_MS, 9 → PTY_MS, 3 → PTY_LONG_MS. Raw >=60s literal count in the paid scope: 395 → 97, of which 51 are non-timeout noise (fixture dates, run IDs) and 46 are enumerated justified holds (comment-carrying calibrated budgets, poll-loop constants, utility spawn waits, and the seven physical-ceiling 1_500_000 sites). The eval-budgets policy ratchet keeps the residue from regrowing. Known collapse: where an inner runner budget and its enclosing test budget now share a tier, the old stagger is gone — an overrun surfaces as a bun test timeout instead of a graceful runner timeout (diagnosability trade, not a correctness one). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
74ae357e0f |
fix(test): runBin trim assertion — trim shapes stream ends, not interior
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
d87e73fa55 |
feat(test): shared runBin helper for bin-script unit tests
~36 free test files each carry a near-identical local run() (spawnSync
+ utf-8 + {status, stdout, stderr}) differing only in env composition,
cwd, and timeout. runBin absorbs the invariant core; options carry the
variance (gstackHome sets BOTH GSTACK_HOME and GSTACK_STATE_DIR — the
config-precedence trap several locals rediscovered independently; home
for $HOME-anchored bins; input/trim/timeout/maxBuffer). Free-test-only
by design so it never becomes a de facto global touchfile. Migration of
the 36 call sites lands separately (mechanical batches).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
|
||
|
|
bc81d39013 |
feat(test): eval-budgets timeout tiers + fit/ceiling policy test
Five named tiers (JUDGE 120s / CAPTURE 300s / CAPTURE_LONG 600s / PTY 900s / PTY_LONG 1200s) replace hand-ratcheted sprawl (46x300s, 46x120s, 44x360s, 44x180s, 27x240s, 19x150s, 13x420s, 12x600s...), much of it inflated to paper over the old 40-way in-shard concurrency that the sharded runner's 1-file-per-shard model kills. Policy test pins: every tier fits the shard wall minus 120s overhead (the structural fix for budgets-above-the-wall fiction), tiers stay ordered, and no paid literal exceeds PTY_LONG x1.25 — oversized tests get split, not budgeted past the wall. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
2e693a5918 |
fix(test): decouple slop:diff from bun run test; quality-gate runs it per PR
'bun run test' silently appended up to two 120s npx slop-scan runs plus a git worktree add/remove after the suite (2>/dev/null || true) — invisible in the documented '~90-100s' timing and pure friction in the pre-commit loop. Decoupling is not coverage removal: quality-gate.yml now runs slop:diff on every PR (advisory, matching its in-repo 'never blocking' contract), and /review already invokes it explicitly. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
b4cc808ba1 |
feat(test): duration-aware LPT shard packing for the free suite
Hash sharding balances file COUNTS (1.15x spread) but not cost — the Playwright-launching files landed 4/3/4/1/2/1 across 6 shards, giving a measured 28s–97s shard spread and ~40s of idle tail on every run. Full-suite mode now packs by recorded per-file durations (longest-processing-time-first) when the committed seed scripts/free-test-durations.json exists. - ONE store, no overlay: the seed is refreshed occasionally via the new --record-durations mode (each file timed in its own child — exact, and immune to bun's stream buffering, where silent passers print no header to timestamp); GSTACK_FREE_TEST_DURATIONS overrides the path for experiments; CI never records - seed is a hint: missing → silent hash-shard fallback; corrupt (bad merge) → one warning + fallback; unknown files → 75th-percentile pessimism so a surprise long-runner can't recreate the tail - packed shards get duration-aware walls (max(base, predicted x 3)) — LPT decouples count from cost BY DESIGN, so the 5s/file heuristic would undersize a shard holding few expensive files - one log line per shard (files + predicted seconds) so packing regressions are diagnosable from any run log - the --shard CI-matrix path is untouched: stable hash indices are its contract - successor note in-code: bun >=1.3.14 ships native --timings/--shard LPT — swap this packer when the repo unpins 1.3.13 Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
1de75acc27 |
test(evals): ratchet the 8 newly-visible gate-matrix gaps
The self-registration sweep made these eight files' gate-tier keys visible to the census for the first time — their gate tests run in NO CI lane today (pre-existing hole, newly measurable). Ratcheted into KNOWN_MATRIX_GAPS with the burn-down note: the paid-lane re-platform runs every gate file by construction and retires this ratchet class. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
2ec4dcf105 |
fix(evals): every E2E key's dep list names its own declaring test file
129-of-177 keys omitted their own test file, so editing only a test's prompt or assertions selected NOTHING — the changed test never ran on the change that changed it. 135 keys self-registered (110 E2E + 25 LLM-judge), resolved by strict declaration evidence (testName:/ testIfSelected/judge call sites), with skill-name false positives excluded. e2e-tier-alignment's warn-only branch for unregistered files is now a hard failure with a 4-entry KNOWN_UNREGISTERED ratchet (template- literal testNames, fail-open-safe) + a burn-down test so the set only shrinks. Selection sanity: a one-file diff on skill-e2e-qa-workflow now selects its 4 tests (was 0); skill-llm-eval 0 → 25. Known follow-ups (filed): 15 E2E + 2 judge PHANTOM keys select tests that exist nowhere; codex-e2e-plan-format's testIfSelected names have no map keys (run-all only). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
95b779f7cd |
feat(gen): --out-dir renders every host, outputs-only
--out-dir was Claude-host-only (gen-skill-docs.ts:842), which forced the codex/factory-regenerating tests (gen-skill-docs, skill-validation, host-config) to mutate the live tree — the reason they sit in the TREE_MUTATING serial shard. The flag now mirrors ALL outputs into the out-dir: external-host trees (.agents/.factory/... via processExternalHost), external section files, openclaw docs, and gstack/llms.txt (a catalog-mode render must never rewrite the tracked index). OUTPUTS ONLY — inputs (templates, sections/, host configs) are always read from ROOT, so an empty out-dir can never feed the render. rewriteSectionBase stays Claude-only (external hosts have their own path grammar). Proofs: in-place --host all is byte-identical (tree clean); --host all --out-dir <mkdtemp> renders the full multi-host tree with ROOT untouched; gen-skill-docs-out-dir tests + 415/415 gen-skill-docs.test.ts green (bin/dev-setup's claude rendering byte-compat). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
57501f5d6f |
refactor(gen): main() guard — importing gen-skill-docs no longer regenerates the tree
The generator's whole body executed at module load, so any import of it (test/gen-skill-docs.test.ts pulls assertSinglePreamble via require(); test/catalog-trim.test.ts imports helpers) regenerated all 71 SKILL.md in place — the root cause of half the TREE_MUTATING serial-shard entries (hazard class #2532). The body now lives in an exported main(): number behind if (import.meta.main). Semantics preserved exactly: failure exits are immediate (matching the old top-level process.exit), success leaves the event loop to drain so the llms.txt fire-and-forget IIFE finishes its write, and the module stays synchronous/require()-able. Proofs: byte-identical --host all output (git status clean), --dry-run stale-tree still exits 1 (the skill-docs freshness lane depends on it), and the new test/gen-skill-docs-import-purity.test.ts pins load-time purity via a subprocess probe (mtime-based, so a dirty worktree can't false-fail). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
40a292e1df |
fix(test): trim the seven over-wall 1700s timeouts to the 1500s physical ceiling
1,700,000ms (28.3 min) exceeded every wall these tests run inside: the 25-min CI job timeout and the 1800s sharded-runner wall (which also leaves --retry 1 zero room for a second attempt). Budget above the wall is fiction, not headroom — a test that actually used it produced a job-level kill (no bun summary, no artifact) instead of a clean per-test timeout. No recorded p95 exists for this family (they are being retiered to periodic in the re-platform wave); the trim stops at the physical ceiling rather than guessing lower. Final policy lands in the Wave-2 eval-budgets constants module. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
ae5ffefb9a |
fix(test): unique tmp dirs for plan artifacts + audited live-repo cwd sites
Six paid PTY tests wrote their expected plan artifact to a FIXED shared
/tmp path ('/tmp/gstack-test-plan-<mode>.md') and rmSync'd it in
finally — under --retry 1, EVALS_JOBS>1, or two concurrent worktrees, a
sibling's cleanup deletes this run's artifact and the D19 'agent did
not produce expected plan file' assertion fires spuriously. Each test
now mkdtemps its own dir, interpolates the unique path into the agent
prompt (fixture-sourced prompts get a replaceAll + drift guard that
throws if the fixture's literal ever moves), and cleans up its own dir.
The 18 cwd:-into-the-live-repo sites were audited: all deliberate
(skill registry + hermetic pre-trusted dir, in-repo gen renders, git
history reads, slug resolution) — each now carries a
'// LIVE-REPO CWD: <reason>' comment so the next audit can tell
deliberate from accidental.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
|
||
|
|
0d2f703f28 |
fix(browse): restrictDirectoryPermissions warns and skips symlinked dirs
Closes the Windows Free Tests red: recent lane failures showed a
platform-unguarded POSIX mode-bit assertion ('Expected: 493' — a
symlink-skip test) from PR-branch variants; the KNOWN_WINDOWS_SAFE
force-include reason ('mode-bitmask hits are POSIX-branch only') did
not hold for that shape, and main had neither the guard nor the
behavior.
- product: lstat first; a symlinked dir gets a warning and a skip on
both platforms — chmod AND icacls dereference the link, so
restricting through a symlink hardens an unvetted target (and
/inheritance:r could lock out its real owner). All callers already
treat hardening as best-effort (try/catch).
- test: the symlink regression test, platform-aware — symlinkSync in
the house try/catch skip pattern (Windows runners without Developer
Mode can't create symlinks), mode-bit assertion guarded off win32,
behavior assertions (no throw, warning text, target readable)
everywhere; POSIX still proves the skip (0o755 unchanged, not 0o700)
- KNOWN_WINDOWS_SAFE reason updated to the now-true premise
20/20 pass on Linux.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
|
||
|
|
72a5246aae |
fix(evals): activate the 4 paid test files that could never run anywhere
carve-section-loading, codex-e2e-plan-format, codex-e2e-recommendation-substance, and llm-judge-recommendation gated on EVALS/tier (free suite loads them as describe.skip) but their names fell outside PAID_TEST_GLOBS, so no paid lane ever selected them — net execution zero, forever. The existing matrix tripwire filtered on isPaidTestFile() first, so it was blind to exactly this class (the same bug that hid the pre-split monolith's gate tests for ~8 releases). - PAID_TEST_GLOBS: codex-e2e* + skill-llm-eval* wildcards (replacing exact names) + llm-judge-recommendation + carve-section-loading; package.json's six test-script glob lists mirrored - codex-e2e-plan-format gains the explicit periodic tier gate its siblings carry (external-service rule) — without it the sharded runner's no-guard default would spawn Codex in the gate tier per PR - eval:bg:periodic --timeout 32400→37800: the census growth pushed the periodic worst case to 35910s; the old value had 270s of headroom BEFORE this change and would now kill healthy runs mid-flight - new test/paid-orphan-tripwire.test.ts: any EVALS/tier-gated test file outside the globs fails the free suite (reasoned SCANNER_EXEMPT for the gate helpers + meta-tests) — the class-killer - paid-shards pins updated: the four orphans now assert INSIDE the census Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
442a46f989 |
fix(test): reactivate 5 quarantined browse tests (2 security)
extension-sender-auth's two privileged-message denial tests (content
script + missing sender.url — the extension's security boundary) and
snapshot's three skips were quarantined 'pre-existing' failures. Root
cause: machine-local state on the quarantining dev machines — the test
and gate code are byte-identical between the quarantining commit
(
|
||
|
|
9eaf15564c |
fix(test): the two expect(true) paid stubs become test.todo
skill-e2e-spec-execute (600s budget) and skill-llm-eval-spec (300s) reported PASS on every periodic run while asserting nothing. Deleting them would remove the periodic-tier selector surface they exist to register (diff-based selection for spec/ changes), so they become test.todo — reported as todo/skip, never pass — with the v1.1 implementation specs kept in-file. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
e9643e131f |
fix(evals): judges honor the eval-model resolution chain + real 429 backoff
callJudge inlined GSTACK_EVAL_MODEL_JUDGE || sonnet, silently ignoring the global GSTACK_EVAL_MODEL override every other eval call site honors via lib/eval-model.ts. New 'judge' kind in DEFAULTS (sonnet — the D1a pin-on-regressors calibration stands; model CHOICE unchanged) and callJudge resolves through it: explicit arg > GSTACK_EVAL_MODEL_JUDGE > GSTACK_EVAL_MODEL > default. 429 handling upgraded from one fixed 1s retry (reliably lost races at CI concurrency) to three jittered exponential retries (~1s/4s/16s), honoring the server's retry-after when present. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
63f8829578 |
fix(test): e2e-harness-audit derives its skill census from disk
The hand-maintained 39-name SKILL_GLOBS list had drifted to 39 of 54 SKILL.md.tmpl on disk. No live gap today (none of the 15 unlisted skills is interactive), but the next interactive skill would have landed unguarded with zero signal. The audit now walks top-level dirs for SKILL.md.tmpl (statSync so symlinked dirs like connect-chrome count), so new skills are in scope the commit they appear. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
241935adce |
test: tripwire against module-scope GSTACK_HOME assignments
Column-0 assignment of GSTACK_HOME / GSTACK_STATE_ROOT in any tracked *.test.ts fails with the file:line and the fix (beforeAll + afterAll restore). Kills the cross-file env-leak class the previous commit swept. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
c252f00b4c |
fix(test): scope GSTACK_HOME to each file's execution window
Five files assigned process.env.GSTACK_HOME at module scope. Shard processes evaluate sibling modules before running their tests, so the assignment leaked into every other file in the shard — the damage was already visible in defensive workarounds (relink.test.ts:28 'fresh install test saw a neighbor's skill_prefix'; cdp-e2e's own comment documents a sibling's temp dir baked into artifacts). Pattern: save original, assign in beforeAll, restore in afterAll (cdp-e2e already restored but still assigned at load — its window now matches the others). GSTACK_TELEMETRY_OFF and GSTACK_PROJECT_SLUG get the same treatment where they rode along. Victim files' defenses stay in place (cheap insurance). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
05a45ddb89 |
fix(ci): small-lane batch — timeouts, right-sizing, windows cache warm-start
- timeout-minutes on the 6 remaining unbounded jobs (actionlint 5, skill-docs 10, version-gate 10, make-pdf-gate 15, pr-title-sync 5, evals build-image 15) — a hung step sat on GitHub's 360-min default - right-size measured-over-long timeouts: dependency-review 10→5, windows-setup-e2e 15→10 - dependency-review: 2-core runner (28s API call on an 8-core box) and drop .github/workflows/** from its trigger paths (workflow edits have no dependencies to review) - windows caches gain restore-keys: a lockfile bump paid the 26s/43s restore for a guaranteed cold miss Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
5cdc02f555 |
fix(ci): quality-gate drops the 74s full-history checkout
fetch-depth:0 cost 74 of the job's 92 seconds; the three gates it feeds take ~12s combined. Shallow checkout + exact-SHA fetches for the diff's base/head (an exact-SHA fetch, not a guessed depth — long-lived branches and merge queues still resolve), with a --deepen fallback for push events whose 'before' is unusable. timeout right-sized 20→10 min. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
adaad18124 |
fix(ci): ci-image stops rebuilding the identical image every ship
- package.json out of the trigger paths: the tag hash deliberately excludes it (version bumps every ship), so every merge rebuilt and re-pushed the IDENTICAL tag (~2m26s for zero content change); patches/** added (it IS a tag input) - manifest existence check (mirrors evals.yml): tag already exists → skip the build - concurrency group: two rapid main pushes raced pushing the same :latest/:buildcache tags - cron staggered 06:00→04:00 Monday: it shared the exact minute with evals-periodic, which could race a half-pushed tag or duplicate the build - timeout-minutes: 30 (was unbounded → 360-min default for a hung docker build) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
c8722b243d |
test(ci): bind the three-way image-tag hashFiles() expressions
evals.yml, evals-periodic.yml, and ci-image.yml each compute the CI
image tag from hashFiles('.github/docker/Dockerfile.ci', 'bun.lock',
'patches/**') — synced by comment only (TODOS.md 'CI three-way
image-tag drift'). If one input list drifts, that workflow computes a
different tag for the same content: eval lanes silently rebuild the
image every run, or ci-image prebuilds a tag nobody looks up. The test
extracts each tag-computation site and fails on any mismatch.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
|
||
|
|
ace904d40a |
fix(ci): one bun version everywhere + drift tripwire
Lanes disagreed four ways: 1.3.13 (free-tests, windows, Dockerfile.ci), latest (quality-gate, make-pdf-gate), unpinned (skill-docs, version-gate — setup-bun installs latest), 1.3.10 (.gitlab-ci.yml). Different Bun versions change the runner output shapes the strict classifiers regex-match, spawn semantics, and shell parsing — a lane on a different Bun tests a different product; Dockerfile.ci's own comment records this class biting once already (silent 1.3.13/1.3.14 drift). All surfaces pinned to 1.3.13; test/bun-version-drift.test.ts scans every workflow setup-bun stanza + Dockerfile.ci + .gitlab-ci.yml and fails on any mismatch or unpinned stanza. skill-docs also gains --frozen-lockfile (was bare bun install). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
e2904be7a4 |
fix(ci): least-privilege permissions + fork-safe concurrency keys
- evals.yml / evals-periodic.yml evals jobs: explicit contents:read + packages:read (container-image pull) and persist-credentials:false — the jobs that execute PR-authored code with three provider API keys ran on the repo-default token grant with the token written into .git/config - permissions blocks for the 4 workflows that had none (skill-docs, make-pdf-gate, windows-free-tests, windows-setup-e2e) - fork-safe concurrency keys: actionlint, skill-docs, make-pdf-gate, windows-setup-e2e switch from head_ref to PR-number keying — a bare branch name carries no fork prefix, so same-name branches from two forks shared one group and cancelled each other's runs Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
9fbd0700ff |
fix(ci): kill the three zero-test eval jobs (hollow green)
- delete the vestigial e2e-codex / e2e-gemini matrix rows: both files
are whole-file periodic-tier, so with no row tier: they ran ZERO
tests and reported green on every PR (~2 min of runner each, pure
false confidence; the periodic lane owns those suites)
- e2e-pty-plan-smoke gains tier: gate — its two files are whole-file
describeE2ETier('gate'), so the job burned ~7 min of container setup
then skipped every describe
- KNOWN_TIER_UNSET burned down to empty; the ratchet stays armed so a
future row/file tier mismatch fails the suite instead of shipping
hollow green
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
|
||
|
|
29d94a505d |
fix(ci): free-tests lane actually runs the make-pdf e2e gates
The 9 make-pdf/test/e2e gate tests probe make-pdf/dist/pdf, browse/dist/browse, and the diagram-render bundle, then self-skip when absent. The required free-tests lane never built any of them, so the gates silently skipped on Linux for their entire life (verified: 9 of 14 skip, exit 0). make-pdf-gate.yml's justification for deleting its Linux leg claimed the free lane covered this — it didn't. - new build:gates script: exactly the three artifacts the gates probe (full bun run build compiles five binaries; ~60-90s tax on the only required check is not warranted) - free-tests.yml: build:gates step + poppler-utils + fonts-noto-color-emoji (fonts must precede the first browse daemon launch — Chromium snapshots fontconfig at startup; verified live: a warm daemon renders tofu, a fresh one embeds NotoColorEmoji) - make-pdf/test/e2e/ci-prereqs.test.ts: GSTACK_EXPECT_BINARIES=1 (set by the workflow) inverts the skip polarity in CI — dropping the build step or poppler fails the lane instead of re-opening the silent-skip hole Pre-flight: all 9 gates green on Linux locally. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
394db326f2 |
v1.71.0.0 feat: token-load reduction — preamble runtime scripts, gated onboarding, 20 skill carves, CLAUDE.md trim (#2691)
* feat(gen): strip gen-time-only frontmatter keys from Claude renders
interactive + benefits-from are read from the .tmpl by buildContext at
generation time; no runtime, host, or test reader consumes them from the
generated SKILL.md (e2e-harness-audit reads .tmpl; benefits-from tests
assert rendered prose). gbrain: stays (bin/gstack-brain-context-load reads
it from the installed render); hooks: stays (Claude Code host wires
PreToolUse from it).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* chore(gen): regenerate SKILL.md — dead frontmatter keys removed
Mechanical regen after hosts/claude.ts stripFields change.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* feat(test): context-budget ratchet — CI ceilings on always-on + eager token ledgers
New free test grades the two ledgers nothing else guards: the full-frontmatter
always-on catalog (aggregate) and per-skill eager tokens (SKILL.md +
forced-read refs), via checkBudget from lib/context-bill.ts. Ceilings live in
test/fixtures/context-budget.json with x1.05/x1.10 headroom; regenerate with
bun test/helpers/capture-context-budget.ts. New skills fail until consciously
budgeted; removed skills fail until the fixture is refreshed; reductions
ratchet the ceilings down so wins lock in.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* docs(todos): file output-template carve wave + plan-ceo doctrine revisit; mark preamble-carve P3 in flight
Two follow-ups deferred from the approved token-reduction program (CEO review
'NOT in scope' list), filed with full context per TODOS format. The existing
P3 preamble-carve entry gets a status update pointing at the program that
supersedes it.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(test): review findings — Windows path normalization, full totals rebuild, ratchet coverage
Pre-landing review (5 specialists) found one critical: the ratchet test runs
in the curated Windows lane, where path.relative yields backslash skill names
that miss the test/ filter and mismatch every POSIX fixture key. Names are now
normalized once in buildRatchetBill (toPosixName) and the fixture filter is
tightened to test/fixtures/. All eight Bill.totals fields are rebuilt from the
filtered list (no fixture-polluted perInvocation/totalMd numbers for future
consumers). New coverage: Windows-separator normalization pins, a
captureContextBudget round-trip against tree-a (headroom math exact), a
stripFields regression pin (interactive/benefits-from absent from renders,
hooks/gbrain preserved), and the ceilings test no longer double-reports
stale-fixture entries.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(test): adversarial findings — stable root key, symlink-alias dedupe, fixture-shape guard
Adversarial review (Claude subagent) verified the fixture's root-skill key was
the capture machine's checkout dirname: any non-gstack-named clone (every
Conductor worktree) failed the free suite, and the documented re-run-the-capture
recovery baked the local dirname into the committed fixture — silent corruption
through the tool's own protocol. The root skill is now pinned to ROOT_SKILL_KEY
('gstack', its frontmatter name). Symlink aliases are realpath-deduped (census
precedent): connect-chrome no longer gets its own ceiling, so Windows checkouts
that materialize the symlink as a plain file can't fail the stale-ceiling
set-equality test. New guards: fixture-shape validation (a string alwaysOnTotal
can no longer silently disable the ceiling), a mutation pin that the filter
shrinks the always-on ledger vs the raw bill, an alwaysOnTotal violation test
(the branch was load-bearing with only under-budget coverage), and an atomic
temp+rename fixture write. Fixture regenerated: 59 ceilings, alwaysOnTotal 6344.
Deferred with a TODO: anchoring transformFrontmatter's denylist strip to the
frontmatter block (latent, zero live collisions, pre-existing path).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* chore: bump version and changelog (v1.69.1.0)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* docs: update project documentation for v1.69.1.0
CLAUDE.md: Token ceiling section documents the context-budget ratchet as
the third guard (test file, fixture, new-skill budgeting, capture command).
CONTRIBUTING.md: Tier 1 guard list gains a Context-budget ratchet bullet;
the Adding-a-new-skill checklist gains the budget-capture step.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* docs: pin exact guard semantics for the context-budget ratchet in CLAUDE.md
Doc-review finding: "a third enforced ceiling" undercounted the guard
family (skill-size-budget floors and parity ratios also watch these
ledgers, relatively). Rephrased to match the ratchet test's own header:
absolute ceilings vs relative floors/ratios.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* docs(changelog): heaviest-skill claim matches the fixture (land-and-deploy edges review by 0.2%)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* feat(bin): gstack-skill-start + gstack-skill-end — the preamble runtime, consolidated
Absorbs the ~13KB of bash every tier-2+ SKILL.md inlined twice over (bootstrap
fence + artifacts-sync fence) and the skill-end telemetry/sync fences. Same
KEY: value STATUS-line contract the prose interprets, plus SKILL_START_PROTO
handshake (OV5), SESSION_ID/TEL_START echoes, GSTACK_HOME-normalized state
paths (EOV7), --parent-pid session identity (EOV5: $PPID inside the script is
the ephemeral tool-call shell), OV4 sanitization of passthrough output, and a
receipted daily artifacts pull (_receipted_git, brain-sync class, fail-closed).
Per-line || true error style throughout (F3) — a mid-script failure never drops
later STATUS lines.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* feat(gen): preamble resolvers emit a script invocation fence instead of inline bash
generate-preamble-bash: ~6.3KB fence -> 4-line gstack-skill-start invocation
(quoted-tilde pitfall handled: leading ~ interpolates through $HOME; env-var
hosts keep $GSTACK_BIN) + degraded-mode prose (F1/EOV8: safe defaults, consent
gates deferred-never-lost; OV5: proto rule). generate-brain-sync-block: ~6.8KB
bash -> interpretation prose + the privacy stop-gate (stays inline until
Phase 2's gated emission). generate-completion-status: telemetry fence -> one
gstack-skill-end call with SESSION_ID/TEL_START handoff.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* chore(gen): regenerate all skills + golden fixtures — inline preamble bash removed
Mechanical regen after the resolver change: −12,628 lines across 52 renders
(corpus 952K -> 806K render tokens; tier-2 skills −11-13KB each). Golden
per-host ship fixtures refreshed from the fresh claude/codex/factory renders.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* test: skill-start contract suite + preamble A/B eval + touchfiles registration
test/gstack-skill-start.test.ts (11 free tests): STATUS-key contract vs the
prose (F2), per-host fence resolution shapes (E1), proto-first, OV4 marker
sanitization, --parent-pid identity, headless suppression, skill-end duration
math + pending cleanup. test/skill-e2e-preamble-script-ab.test.ts (gate tier,
OV7): inline-bash render (pinned from
|
||
|
|
a3749bfa4b |
v1.70.1.0 fix: ship names the /document-release subagent at every decision point (tripwire + gate E2E) (#2700)
* fix(ship): name the /document-release subagent at every Step 18 decision point The v1.54.0.0 carve moved Step 18 (documentation sync) into ship/sections/pr-body.md and the Claude-host skeleton stopped saying "document-release" anywhere in the workflow body — the dispatch became invisible at exactly the moments an agent decides whether to open the section. Restore visibility at three touchpoints, all subagent-framed (never bare-slash-framed, which would invite an inline Skill invocation that bypasses the fresh-context subagent + JSON contract): - manifest trigger (renders into the section-index row AND the STOP pointer): "dispatching the /document-release subagent to sync docs (Step 18) and then creating or updating the PR/MR (Step 19)" - Step 17 handoff line names Step 18's dispatch explicitly - new hoisted doc-sync invariant beside the PR-title invariant: the dispatch itself is never skipped; only a failed subagent is non-blocking Pin it in carve-guards: 'the /document-release subagent' (all three touchpoints) + 'dispatches the /document-release subagent' (invariant) must stay in the skeleton; the carved imperative 'Dispatch /document-release as a subagent' must stay carved. Skeleton cap 91,600 → 92,300 (measured 91,764; trigger renders twice). Goldens regenerated for all three hosts. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * test: pin the ship→document-release Step 18 wiring with a free tripwire Five substring/structure asserts across the carved section, the Claude skeleton's three touchpoints, the manifest trigger, and the codex/factory goldens (inlined Step 18 ordered before Step 19). Claude-golden asserts deliberately omitted: host-config.test.ts already enforces golden == generated byte-for-byte. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * test: gate-tier E2E proving /ship dispatches the document-release subagent New skill-e2e-ship-docsync: a live agent gets the sliced Step 17→19 tail of the generated ship skeleton in a bare-remote git fixture (Steps 0-16 "done"), under a fake HOME so the STOP pointer and the Step 18 subagent prompt resolve to planted copies, with a stub document-release skill that returns the empty-result JSON contract. Hard assert: an Agent/Task tool-call matching /document-release/i exists in result.toolCalls and precedes any `gh pr create`. Neutral prompt (no STOP-Read priming, no document-release mention — the prompt echoes into the transcript, so asserts read toolCalls only). Hardening from review: throw-on-marker-drift fixture slice; per-test GSTACK_HOME + .redact-prepush-prompted marker (routes Step 17's credential guard to its silent branch — the hermetic GSTACK_HOME pin defeats a HOME-only override); 480s/540s timeouts (nested subagent adds wall clock the 300s sibling never carried); 'timeout' accepted in exitReason only because the dispatch assert is independently hard; whole-file describeE2ETier('gate') composed with diff selection (keeps the file out of the periodic shard census, which sits at its ceiling, and under the hard tier-alignment invariant). Registered as 'ship-docsync' in E2E_TOUCHFILES + E2E_TIERS (gate) in the same commit — touchfiles.test.ts rejects either half landing first. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs: fix stale document-release TODOS entry + three review-deferred items The SHIPPED entry still described the deleted Step 8.5 post-PR cat-delegation design from v0.8.4; replace with the current Step 18 subagent design and its test pins. Add the three P3 items deferred from the v1.69 plan review: dispatch receipt enforcement, land-and-deploy→canary dispatch-pin pattern, and the periodic shard-census boundary. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix: pre-landing review fixes Testing-specialist findings, all mechanical: (1) pin the E2E fixture's git branch (-b main / init.defaultBranch=main) and assert every setup command's exit status so operator git config can't silently corrupt a paid run; (2) tighten the dispatch matcher to Step 18-prompt-specific markers (document-release/SKILL.md | executing the /document-release workflow) so a subagent merely quoting section text can't false-pass the regression assert (verified against recorded burn-in transcripts); (3) replace the subsumed carve-guards anchor with three non-overlapping per-touchpoint anchors (gerund/imperative/3rd-person) so each touchpoint is independently enforced. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix: red-team review fixes Five informational findings: TODOS shard-census arithmetic corrected (census is 67 with one free ungated slot; the SECOND ungated file trips the floor) and version pointer fixed (v0.18.2.0, not v0.18.1.0); the free tripwire now pins the two dispatch-matcher marker strings so a pr-body prompt reword fails the free suite instead of surfacing as a paid-tier mystery; the E2E matcher gains a section-paste exclusion (scaffold strings disqualify) — verified against all recorded runs; the E2E header documents the tierless test:evals invisibility tradeoff. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix: adversarial review fixes Pin the E2E matcher's two EXCLUSION markers in the free tripwire (an unpinned 'Parent processing:' reword would silently deaden the section-paste guard while every test stayed green); add an ordering pin (the hoisted doc-sync invariant must sit above the pr-body STOP pointer — presence-only anchors can't catch drift below it); plant a third cwd-relative pr-body copy inside the fixture repo, gitignored so the agent never tries to commit test scaffolding. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * chore: bump version and changelog (v1.70.1.0) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs: CHANGELOG accuracy fixes from the doc-release review Three factual corrections the Step 18 doc subagent caught in the fresh v1.70.1.0 entry: 5 tripwire tests (not 6), cost floor $0.63 per the cited eval store (not $0.59), and the visibility claim scoped to decision points (the re-run checklist mention survived the carve). Plus the E2E header's stale pending-burn-in note replaced with the observed numbers. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix: raise bun-polyfill subprocess budget to 60s for degraded Windows runners The 50ms-sleep test blew the 20s budget on BOTH bun retry attempts on PR #2700's windows-latest runner (run 32989821401) — sustained AV/runner pressure, not just the documented cold-start. Same flake passed-on-rerun on the prompt-token-load-reduction branch yesterday. Budget only; every assertion still checks exact output. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(ci): run the ship-docsync gate E2E in the evals matrix + silent-skip tripwire The evals.yml matrix is hand-enumerated and the Run step never exported EVALS_TIER, so the new whole-file-gated ship-docsync E2E would have self-skipped even with a row — a hollow green one layer deeper than the documented rehomed-monolith incident. Add the e2e-ship-docsync row with a row-level `tier: gate` property, exported as EVALS_TIER by the Run step (empty = unset for every existing row: all readers are `=== '<tier>'` or truthiness). New free tripwire test/evals-workflow-matrix.test.ts ratchets the class: matrix files must exist; gate-hosting files must have a row; whole-file-gated matrix files must carry a matching row tier; and the burn-down lists enforce their own cleanup. It enumerates the PRE-EXISTING holes found while wiring this (8 gate-hosting files with no row; codex/gemini rows running zero tests; the pty-plan-smoke row hollow since its files adopted describeE2ETier) — tracked in TODOS as the CI gate-lane hollow-coverage burn-down. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
ad8400543c |
v1.69.0.0 fix: the silent-failure wave — 6 fixes, 5 community PRs absorbed, tracker closed with receipts (#2666)
* test(wireup): make gbrain-missing PATH fixture hermetic The gbrain-missing test appended the host PATH (and a hardcoded /opt/homebrew/bin) to the fixture PATH, so on any machine with a real gbrain installed the 'missing' case saw it, exited 0 instead of 2, and could never fail where the bug exists — a false green for a whole machine class. The fixture now keeps only root-owned OS dirs on the child PATH, and a new determinism check plants a host-like gbrain to prove it is unreachable. Absorbed from PR #2615 with authorship preserved; the PR-thread liveness screenshot (docs/images/gstack-pr-liveness-2255.png) is dropped — referenced by nothing in the tree. Fixes #2255 Co-authored-by: CommandCodeBot <noreply@commandcode.ai> * fix(evidence): stop bun's dotenv autoload from reaching the spawned command `bin/gstack-evidence` has a `#!/usr/bin/env bun` shebang, and bun AUTO-LOADS `.env`, `.env.<NODE_ENV>` and `.env.local` from the cwd into `process.env`. The wrapper then spawned the command with no `env` override, so every command run through it inherited those variables — and a repo `.env.local` routinely holds production credentials. Two things go wrong, and the second is worse than the leak: 1. Secrets reach a child that would not otherwise have them. `npm test` run by hand in the same shell sees none of them; the same command through the wrapper sees all of them. 2. THE COMMAND UNDER TEST BEHAVES DIFFERENTLY, so the ledger certifies a run that is not the run CI performs. Observed in a Next.js repo on 2026-08-20: four tests failed 4/4 through the wrapper and passed 5/5 without it, because app code branched on env vars only the wrapper supplied. Nearly an hour went into chasing a "flake" that was the measuring instrument. The wrapper exists to record trustworthy evidence, so silently altering the environment defeats its purpose. The fix builds the child env from `process.env` minus the keys bun injected, and detection is exact rather than heuristic: verified on bun 1.3.11, a dotenv file does NOT override a variable the shell already exported (the shell's value wins). So a key whose live value equals the dotenv file's value was injected by bun, and dropping it restores the environment the user's own shell would have given the command. A key whose live value differs is genuinely the caller's and survives. `BUN_DOTENV_FILES()` mirrors bun's precedence, including that `.env.local` is skipped when NODE_ENV is "test" — scrubbing a key bun never loaded would strip a variable the caller legitimately provided. Escape hatch: GSTACK_EVIDENCE_KEEP_DOTENV=1 keeps the old behaviour. When keys are scrubbed the wrapper warns with the KEY NAMES ONLY, so the diagnostic cannot become the leak it prevents. Tests: 6 cases, mutation-verified — removing `env: spawnEnv` reddens exactly the two leak tests and restoring it gives 30/30. Every leak test asserts the scrub warning fired, because `bun test` sets NODE_ENV=test and the first version of these tests passed vacuously against a `.env.local` bun had never loaded. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Absorbed from PR #2652 with authorship preserved. Wave additions: a doc-comment on the ${VAR}-expansion limitation (bun expands refs, the reader compares raw text — those keys are left in the child env, failing open) and a regression pin for the unreadable-.env fail-open path with a functional DAC-override skip guard. Fixes #2624 * fix(setup): reap dangling skill dirs when the payload is gone cleanup_old_claude_symlinks derived its work list from the payload directory, so when the payload was gone — precisely when orphans exist — the glob matched nothing and the loop never ran; the -f guard also followed symlinks, hiding dangling SKILL.md links even with a payload present. The cleanup now scans the DESTINATION skills dir (-e/-L, so dangling symlinks are visible) and anchors SKILL.md provenance to path segments (gstack/*, */gstack/*, */.gstack/render/claude/*) instead of a bare *gstack* substring that would eat a user skill under ~/tools/gstack-fork/. The Windows real-file arm stays payload-gated: a real file has no provable owner. Absorbed from PR #2634 (2 commits squashed) with authorship preserved. The symmetric cleanup_prefixed_claude_symlinks hole is filed as a TODOS.md residual in this wave. Fixes #2204 * fix(redact): tolerate EEXIST from recursive mkdir in install-prepush-hook on bun/Windows (#2635) fs.mkdirSync(dir, { recursive: true }) is a no-op on an existing directory in Node, but bun on Windows throws EEXIST - crashing hook install on any repo whose .git/hooks already existed, leaving the repo unprotected. Add lib/fs-utils.ts mkdirpSync: swallow EEXIST only when statSync confirms the path is an existing directory; a regular file occupying the path, a stat failure, or any other errno still rethrows. Use it in installPrepushHook(). The regression test emulates the Windows bun fs semantics via a bun --preload fixture, so the exact crash path runs (and fails on the old code) on any platform, including CI Linux. Absorbed from PR #2641 with authorship preserved. Fixes #2635 * fix(bin): route remaining Windows-reachable mkdirSync sites through mkdirpSync Sweep follow-up to #2641's lib/fs-utils.ts helper: bun on Windows throws EEXIST from a recursive mkdir on an existing dir, so every unguarded recursive mkdirSync on a Windows-reachable path is a latent crash. Converted: bin/gstack-decision-log (unguarded, runs on every decision log — the second call on any machine hits the pre-existing projects dir), bin/gstack-evidence logsDir + ledger dir sites, and bin/gstack-redact-prepush's skip-log site (already try-wrapped, so its failure mode was a silent skip-log loss rather than a crash — the fix makes the log survive). The ~15 remaining gbrain/mac-lane sites are deliberately left alone. Regression: fs-utils.test.ts drives gstack-decision-log twice, the second run under the bun-Windows EEXIST preload fixture — the pre-sweep code exits 1 with EEXIST there; verified red against v1.68.3.0. * fix(setup-gbrain): warn about the ZeroEntropy sunset before Sept 4 ZeroEntropy was acquired by Notion and sunsets its hosted API on September 4, 2026. A gbrain configured with the zeroentropyai embedding recipe keeps importing pages after that date but embedding silently fails — pages land structurally with no semantic search, this repo's tracker P1 (TODOS.md NEXT PRIORITY). Nothing in gstack ever recommended ZeroEntropy (the dependency is gbrain-internal), so the gstack side is detection + advisory: the wireup helper warns when ~/.gbrain/config.json names the recipe (fail-open grep — a missing, unreadable, or other-provider config stays silent and never blocks a working setup), the setup-gbrain provider-default comments say never to select the legacy recipe for a new brain, and USING_GBRAIN_WITH_GSTACK.md gains a troubleshooting entry. The gbrain-side provider migration stays open upstream. Refs #2365 * fix(gbrain-source-wireup): first sync targets the registered source, not --repo The wireup registered a federated source by id, then ran 'gbrain sync --repo $WORKTREE' — which resolves against the brain's DEFAULT source and (on gbrain 0.46.x) rewrites that source's local_path anchor to our worktree. Net effect: the user's primary knowledge source silently repointed at the gstack brain worktree while the just-registered source got zero pages, and pages_synced still reported success. The sync now targets the registered id ('gbrain sync --source $id', the same form the repo's own troubleshooting documents). Because the script's stated floor is gbrain >= 0.18.0 and nothing proves --source exists there, support is probed via 'gbrain sync --help' first: an older gbrain keeps the wrong-but-working --repo call with an upgrade warning instead of converting it into a hard failure. The probe sits after the GSTACK_BRAIN_NO_SYNC early-exit and is unreachable in --probe mode. Regression tests (fail on v1.68.3.0): a no-skip sync case asserting the call log shows 'sync --source gstack-brain-<id>' and never 'sync --repo', and an old-gbrain fallback case (fake sync --help without --source) asserting --repo plus the upgrade warning. Fixes #2662 * fix(setup): --host slate exits informatively instead of silently installing nothing slate passed --host validation (added to the accept-list in v1.64.1.0) but never got a dispatch arm, and the all-INSTALL_*-zero fallback lives inside the auto branch — so './setup --host slate' configured nothing and exited 0, a silent no-op strictly worse than the original hard rejection. slate is now an informational arm (per docs/designs/SLATE_HOST.md it is blocked on the host-config refactor; Slate reads .claude/skills as a compatibility fallback, so the arm points at './setup --host claude'), and a defensive guard after the dispatch chain errors loudly (naming the host, the missing arm, and the valid targets, exit 1) if a future host is ever accepted without being wired. Regression tests (fail on v1.68.3.0): a dispatch-arm ratchet asserting every accept-listed install target has a matching dispatch branch — the exact drift class; a registry cross-check deriving both sides from hosts/index.ts and setup's case arms; a behavioral slate probe (exit 0, points at --host claude, never reaches the installer — on unfixed code it fell through into the installer); and a static pin on the guard's shape. Fixes #2361 * fix(make-pdf): resolve the sibling browse binary from execPath, not argv[0] In a bun-compiled binary process.argv[0] is the raw invocation string — often relative ('./pdf', 'pdf') — so dirname(argv[0]) yielded '.' and the sibling candidates (../browse/dist/browse etc.) resolved against the CWD instead of the install dir. Resolution was cwd-dependent: correct-by-luck when the fallbacks rescued it, wrong when a cwd-relative path matched. process.execPath is always the absolute binary path. The resolution step takes an injectable selfPath (defaulted) because under bun test the process path is the bun runtime and the compiled-binary shapes are otherwise unreachable. The issue's other half — pdf setup failing on newtab('about:blank') — was already fixed on main in v1.64.0.0 (browse/src/url-validation.ts exact-match allows about:blank; its comment names this exact smoke). This commit closes what remains. Regression tests (the sibling-via-selfPath case fails on v1.68.3.0 — pre-fix code ignores the seam and either resolves the global install or throws): sibling resolution from an install-shaped tree, and a decoy-browse-DIRECTORY case pinning that a directory never wins resolution. Fixes #2156 * fix(memory-ingest): store the normalized git_remote so unattributed pages hit the policy filter buildTranscriptPage wrote the normalized '_unattributed' sentinel into the page FRONTMATTER but stored the raw resolved remote ('' when unresolvable) on the page object. The policy filter fast-paths !p.git_remote, so under --include-unattributed an explicit '_unattributed → deny' (or read-only) policy never applied to exactly the pages it names — they ingested unpoliced. The stored value now matches the frontmatter. Regression test (fails on v1.68.3.0): seeds the REAL bin/gstack-gbrain-repo-policy store with '_unattributed → deny' through its own set verb, ingests an unresolvable-remote session with --include-unattributed, and asserts nothing reaches gbrain — pre-fix the '' remote bypassed the filter and the import ran. A fake echoing tiers would pass on both sides of the fix; the real helper prints 'none' for unknown keys, so only a genuinely applied deny distinguishes the two. Fixes #2353 * fix(land-and-deploy): MERGED recovery reconciles and reports remote-branch cleanup Step 4's merge commands carry --delete-branch, and the success path tells the user 'The branch has been cleaned up.' When gh exits non-zero AFTER GitHub already merged (routine in worktree layouts: gh's local cleanup runs git checkout <base> and fails), the §4a-postfail MERGED recovery re-established everything EXCEPT the branch deletion — and said nothing about it, so the discrepancy was invisible. The MERGED path now reconciles: git ls-remote --heads distinguishes branch-already-gone (exit 0, empty → 'already cleaned up', idempotent on re-runs) from branch-survived (offer confirm-first deletion, matching the section's worktree posture; -d not -D for any local branch) from check-itself-failed (non-zero exit → 'couldn't verify', skip the offer — never read a failed check as a clean branch). Template + regenerated SKILL.md + test extensions land in one commit (the md-sync assertion goes red otherwise). Regression assertions (fail on v1.68.3.0: no delete-branch reconciliation existed in test/ at all) pin the ls-remote check, the confirm-first delete, and the absent-vs-failed distinction. Fixes #2656 * fix(scripts): stop heredoc bodies deadlocking under Homebrew bash `./setup --help` can hang forever on macOS, printing nothing, with no way to tell it apart from a slow install. Eleven scripts carry the same latent hang, `setup` itself being the one every user hits first. bash 5.2+ delivers a heredoc body of 64KiB or less through a pipe: the forked child writes the entire body before exec, and nothing reads the other end until the command starts. Under macOS pipe-KVA pressure the kernel hands a fresh pipe a 512-byte buffer instead of the usual 16-64KiB, so any body of 512 bytes or more blocks write() permanently. The capacity check bash would need to notice (F_GETPIPE_SZ) is Linux-only, so it never fires here. It is pressure-dependent, which is why it reads as "worked on my machine" — the same script runs fine all day and then wedges. Homebrew bash is what `#!/usr/bin/env bash` resolves to on a Mac with brew on PATH, which is most of them. Apple's /bin/bash 3.2 predates the pipe path and is unaffected, so the bug is invisible to anyone testing with the system shell. The fix is `BASH_COMPAT=50` in each affected script, which restores the pre-5.2 tempfile path: $ bash -c 'probe() { [ -p /dev/stdin ] && echo PIPE || echo TEMPFILE; } probe <<EOF $(printf "x%.0s" $(seq 1 1000)) EOF' PIPE $ BASH_COMPAT=50 bash -c '...same...' TEMPFILE - Not a `#!/bin/bash` shebang swap: that pins the script to whatever bash lives at /bin (3.2 on macOS, absent on some Linux distributions) and is bypassed entirely by `bash script.sh` call sites. The variable survives both. - Not exported, so child processes keep their own compat level. - Placed below any `--help` sed range that reads $0, so usage output is unchanged (verified on all eleven). - Every guarded script is bash-3.2-clean — no associative arrays, case conversion, or mapfile — so compat level 50 costs them nothing. test/heredoc-pipe-deadlock.test.ts scans every tracked shell script for a heredoc body in the 512B-64KiB window and fails without the guard, and proves the mechanism at runtime on bash 5.2+ by asserting the body moves from PIPE to TEMPFILE. On older bash the runtime half is skipped, since the pipe path does not exist there. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Absorbed from PR #2640 with authorship preserved. Wave adaptations: the pipe-probe test skips on minimal-/dev environments without /dev/stdin (it would report OTHER for an unobservable fd), and one caveat verified during review: on bash 4.3/4.4 (e.g. Git Bash), assigning BASH_COMPAT=50 prints a non-fatal 'invalid value' warning to stderr — those bashes are already on tempfiles, so the guard is a no-op there; windows-setup-e2e exercises this empirically. * docs: TODOS.md v1.69 wave close-out Move the slate P4 entry and the ZeroEntropy P1's gstack-side half to Completed (v1.69.0.0); reframe the ZeroEntropy NEXT PRIORITY entry around the remaining gbrain-side work; file the wave's four residuals with rationale — the prefixed-cleanup symmetric conversion, the #2163 legacy-slug checkpoint heal, the invited #2657 --reconcile contribution, and the table-driven setup host dispatch behind the new cross-check ratchet. * chore: bump version and changelog (v1.69.0.0) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: Som Samantray <som.samantray@gmail.com> Co-authored-by: CommandCodeBot <noreply@commandcode.ai> Co-authored-by: Connex Client Access <paul@paulkortman.com> Co-authored-by: y$un_ <forrest.sun527@gmail.com> Co-authored-by: Lockyer <135391289+Lockyer228@users.noreply.github.com> Co-authored-by: Benjamin D. Smith <benjamin.smith@binarysword.com> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |