From 7cb0f2d93a3e4b4afc63d18676bd2303ac499568 Mon Sep 17 00:00:00 2001 From: Garry Tan Date: Mon, 31 Aug 2026 05:41:11 +0000 Subject: [PATCH] chore: bump version and changelog (v1.77.0.0) Co-Authored-By: Claude Fable 5 --- CHANGELOG.md | 56 ++++++++++++++++++++++++++++++++++ TODOS.md | 15 +++++++-- VERSION | 2 +- agents-digest/gstack-AGENTS.md | 2 +- package.json | 2 +- 5 files changed, 72 insertions(+), 5 deletions(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index 7380116dc..9bb974472 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -1,5 +1,61 @@ # Changelog +## [1.77.0.0] - 2026-08-31 + +**Every PR stops paying for evals twice.** +**Flakes are now measured, killed at the root, and fenced.** + +The test infrastructure got its overhaul, wave 1. The legacy 17-row eval matrix that ran serialized AHEAD of the sliced lane on every PR is deleted: one paid lane, its gate census derived from the runner itself, so a new gate test is in the census the moment its file lands. No hand-enumerated rows to drift, and the drift already tried, a new matrix row landed on main mid-branch and the merge resolved to the derived census that covers it by construction. + +The flake war moved from anecdotes to instruments. Every retried pass is now recorded where it cannot hide (bun's own output shows a retry as a clean pass, we probed it), the free lane retries a failing file once, loudly, and appends every flaky pass to a per-project ledger uploaded from CI on green runs. `bun run eval:flake-rank` ranks the series. And the wedge class that hit main, a hung child under a blocking spawnSync that no in-process timeout can interrupt, is extinct: 586 sync-spawn sites across 176 test files now carry timeouts, with a ratcheted tripwire that failed its first real offender the day a timeout-less spawn arrived from a merge. + +### The numbers that matter + +Source: CI runs 33263204465 / 33262077256 (measured 2026-08-29), the repo census (`test/helpers/touchfiles-data.ts`), and the sweep tripwire (`test/spawnsync-timeout-tripwire.test.ts`). + +| Metric | Before | After | Δ | +|---|---|---|---| +| Paid lanes per PR | 2 (serialized) | 1 | eval wall 35.5 → ~13 min (target) | +| Measured duplicate spend per PR | ~$20.94 | $0 | matrix deleted | +| Gate census keys | 86 (6 phantoms) | 78, all provably alive | reverse invariant enforces | +| Sync spawns that can wedge a shard | 586 unbounded | 0 (8 reasoned exemptions) | tripwire-ratcheted | +| Local gate worst case | ~6.5 h (4×4 jobs) | ~3.3 h (8×2) | isolation landed first | +| Retried-pass visibility | invisible | recorded + ranked | `eval:flake-rank` | + +The phantom-census number is the quiet one that matters: six merge-blocking "tests" existed only as map keys. They are deleted, and a key without a living test now fails the free suite. + +What this means for anyone shipping here: PRs get one honest paid verdict faster and cheaper, a flaky pass never blocks your merge but never disappears either, and a test that hangs takes down thirty seconds, not a shard. Run `bun run eval:flake-rank` when you want the flake ledger's verdict. + +### Itemized changes + +#### Added +- **Flake telemetry, end to end.** Eval-store records every retry attempt (`attempt`, `flaky_retries`), the paid report lists passed-only-on-retry tests as warnings, the free lane's flaky-retry pass is ON in CI with a single-writer JSONL ledger (branch + sha attributed, per-project by default) uploaded as an artifact on every run, and `bun run eval:flake-rank` aggregates the series with final-attempt accounting and a 60-day recency bound. +- **Two-phase session timeouts.** A silent API dies at a startup grace (90s local, 300s CI floor, a real `Math.max` floor) with the distinct reason `timeout_startup` instead of burning a 600s work budget into an opaque `0 turns / $0.00` failure; the work budget arms on the first byte and the total wall never grows. +- **Green-by-skip census.** "Ran N tests" counts skips, so a codex/gemini file whose every test self-skipped used to read as coverage. The classifier now parses bun's skip/pass recap and the paid report labels an all-skipped pass "verified nothing." +- **Sync-spawn timeout tripwire** (spawnSync / execSync / execFileSync / Bun.spawnSync, comment-aware, 30-line window, shrink-only exemption ratchet) plus a raw-SHA fixture ban (`git show :path` fixtures must be vendored; the one live offender now reads a committed fixture). +- **Behavioral kill-semantics tests**: a fake-claude shim proves a timed-out session leaves neither the CLI nor its grandchild alive; two gstack-detach watchdog tests cover TERM-immune grandchildren and the leader-dies-first case. +- **CLI version stamping**: every eval run records `claude --version` (resolved once in the runner parent), so the next TUI-drift flake hunt is a grep, not archaeology. + +#### Changed +- **One paid lane.** The legacy evals.yml matrix is deleted (pure deletion, one revert restores it) after a static parity receipt: the sliced lane's 49-file derived census strictly contained the matrix's 18 files. The PR comment moved into the sliced lane with final-attempt accounting and a fail-closed reconciliation verdict. +- **Paid runner defaults 4×4 → 8×2**: ~10-13 real in-flight sessions (under the documented-safe 15) instead of ~4-6; per-shard TMPDIR/Chromium-profile isolation and a kill-path cleanup backstop landed first, deliberately. +- **CI setup deduplicated into four composite actions**; the register-skills composite carries the fail-fast dangling-symlink verification loop that only the deleted matrix copy had, so the surviving lanes inherit it. Rerun-safe, frozen-lockfile fallback, input-validated. +- **Supply chain pinned**: the claude CLI in the CI image is an exact version (bumps ride PRs that run the PTY gate, ending the weekly-latest drift that broke the harness three times), and every action in the secrets-bearing and image-publishing workflows is SHA-pinned. +- **Routing journeys lost their answer key**: the fixture no longer ships a prompt→skill lookup table, so a regressed skill description can actually fail the test again, at roughly half the previous per-journey cost (2 turns, [Skill, Read] only). +- Decided A/B experiments retired (auq-repetition-cut, preamble-script, the opus-47 single-run fanout comparison): one-shot questions answered months ago no longer re-run weekly as coin flips. `plan-ceo-review-expansion-energy` and `ios-qa-e2e` moved to the periodic tier with reasons. + +#### Fixed +- **A fail-open reconcile gate**: GitHub's default run-step shell has no pipefail, so the fail-closed report's exit was read from `tee` (always 0) in both paid lanes. Now `PIPESTATUS[0]`, pinned by a wiring test. +- **A write-token trust boundary**: the job that executes PR-authored code no longer holds the PR-comment write token; commenting moved to a job that runs zero repo code. +- **Provider-runner orphans**: timeouts kill the whole process group (claude, codex, gemini, and gstack-detach's watchdog with the group id captured at spawn), and the codex/gemini runners inherited the orphan-drain hardening only the claude copy had. The observed 600s-timeout-stretching-past-1400s class is gone, with a regression net. +- **Selection integrity**: 17 phantom selection keys deleted (a reverse invariant now requires every key to name a living test), gitignored `.agents/**` dep patterns that could never match a git diff replaced with their generators, and the codex/gemini local touchfile forks now derive from the canonical map. +- **A cross-shard SKILL.md race**: the opus-47 eval regenerated the live tree's skill files mid-run; it now renders into a scratch dir via `--out-dir`. + +#### For contributors +- `bun run eval:flake-rank` (with `--json`, `--dir`, `--since-days`) is the promotion-clock dial; the WS16 required-check decision reads it. +- The overhaul plan (16 workstreams, reviewed by CEO + eng passes with two cross-model outside voices) continues: budget-aware shard walls, PTY readiness events, free-suite splits, judge determinism, and required-check promotion are the next waves. + + ## [1.76.0.0] - 2026-08-31 **Ship's doc-sync now survives Conductor.** diff --git a/TODOS.md b/TODOS.md index 607769cf4..3accde79a 100644 --- a/TODOS.md +++ b/TODOS.md @@ -567,7 +567,7 @@ duration-packed free shards, the sharded paid runner as the CI engine coverage contract + gate census, eval-budget timeout tiers, and the coverage fill. Remaining, in rough priority order: -- **DONE (v1.76 test-infra wave) — Delete the legacy evals.yml matrix after +- **DONE (v1.77.0.0 test-infra wave 1) — Delete the legacy evals.yml matrix after parity.** Deleted as a pure-deletion commit (one revert restores it) after a static parity receipt: sliced gate census (49 files) ⊇ matrix files (18), 31 files of extra coverage. `needs: evals` edge dropped, PR comment moved @@ -813,7 +813,18 @@ SKILL.md untouched). `bun test` is green again. ## Scope-gate follow-ups (filed via /plan-eng-review on the plan-mode auto-select-B change) -### P2: SDK eval budgets charge API-queue latency to the work budget — pick a structural fix +### DONE (v1.77.0.0) — SDK eval budgets charge API-queue latency to the work budget + +**Shipped shape:** the two-phase timer landed WITHOUT the codemod this entry +feared: the total wall stays <= timeout (work phase = remainder after first +byte), so every outer/inner bun-timeout relationship is untouched; a silent +API now dies EARLY at the startup grace (90s local / 300s CI floor, enforced +Math.max) with the distinct reason 'timeout_startup'. Option (b)'s 300s CI +floor is in (test/session-runner-startup-grace.test.ts pins it). The +budget-EXTENSION variant (work budget = full timeout from first byte, which +DOES need the tier/wall reshape) remains wave-2 scope in the overhaul plan. + +Original entry follows for context: **What:** `runSkillTest`'s single `setTimeout(timeout)` arms at spawn, so session startup AND the model's first-completion queue time are charged against the diff --git a/VERSION b/VERSION index 84e596762..6734beb17 100644 --- a/VERSION +++ b/VERSION @@ -1 +1 @@ -1.76.0.0 +1.77.0.0 diff --git a/agents-digest/gstack-AGENTS.md b/agents-digest/gstack-AGENTS.md index ccdd0618c..b73164ced 100644 --- a/agents-digest/gstack-AGENTS.md +++ b/agents-digest/gstack-AGENTS.md @@ -1,4 +1,4 @@ -# gstack digest v1.76.0.0 — regenerate/re-copy after upgrading gstack +# gstack digest v1.77.0.0 — regenerate/re-copy after upgrading gstack Behavioral rules from gstack (https://github.com/garrytan/gstack), compressed for agent hosts without a full skill install. The full skills add workflows, diff --git a/package.json b/package.json index 99f7774c7..863c5d659 100644 --- a/package.json +++ b/package.json @@ -1,6 +1,6 @@ { "name": "gstack", - "version": "1.76.0", + "version": "1.77.0", "description": "Garry's Stack — Claude Code skills + fast headless browser. One repo, one install, entire AI engineering workflow.", "license": "MIT", "type": "module",