mirror of
https://github.com/garrytan/gstack.git
synced 2026-08-20 21:17:19 +02:00
* test(helpers): shared skill-census helper with three explicit counts
physicalSkillFiles (symlinked dirs included, root router included),
authoredSkills (realpath-deduped, router excluded), registryEntries
(what ./setup registers: unique frontmatter names + _gstack-command).
One counting authority for the hermetic seeder, context-bill ground
truth, and the catalog-budget test — connect-chrome's dir symlink and
the root router otherwise produce three subtly different hand-rolled
censuses. Ported-wave foundation (C11).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(evals): stop the harness grading itself
findPreviousRun excluded only the file being written, by name, so every
suite compared against _partial-e2e.json — the current run's own
accumulator, relabelled with the current tier just before each flush.
That is why every block read '+$0.00, +0s, Stable run, no regressions.'
This harness has never been able to detect a regression, and reassuring
output that cannot fail is worse than none. In-progress runs are now
excluded by role, and a run with nothing to compare against says NO
BASELINE instead of claiming stability.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
(cherry picked from commit f3140b5245221fff7fb9411c7ec07c2ca11587b5)
* refactor(evals): shared partial-run predicate + finalized-run lookup
isPartialEval(data, filename) is the one place that decides what counts
as an in-progress accumulator (the _partial flag OR a _partial-prefixed
filename), and findLatestFinalizedRun(evalDir, tier) is the one place
that finds the newest real run — scanning the eval dir plus one level of
shards/<slug>/ subdirs, where the sharded paid runner points each
shard's collector. skill-budget-regression.test.ts's hand-rolled
findLatestRun (flag-blind: a flagged-but-renamed accumulator passed its
name check) is replaced by the shared helper.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
(cherry picked from commit b55fcf6966366fd21a8cdc46de61aab6e1b1d100)
* feat(evals): register shipped skills for hermetic PTY children
Hermetic children get a config dir that deliberately seeds no skills —
right for children that install their own, fatal for the PTY family that
TYPES /office-hours or /plan-ceo-review: claude rejects the command as
Unknown before any model turn, so the plan-family gate smokes measure
nothing. hermeticSkillsConfigDir() is a second, opt-in config dir under
the same runRoot that mirrors ./setup's registration exactly (real dir
per registry name, SKILL.md + sections/ symlinks, frontmatter-name
resolution, _gstack-command root alias), driven by the shared
skill-census so connect-chrome's dir symlink collapses the same way
setup's idempotent overwrite does.
Ported from fork commit 03c4eca2, tree walk rewritten for the upstream
layout (top-level <skill>/SKILL.md dirs, no skills/ tree). Unit tests
are new: seed shape, census parity, symlink resolution, connect-chrome
collapse, idempotence, no-API-key seed.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
(cherry picked from commit 93dae6107b30ce453a07c2d342b60262bba6ce0b)
* feat(evals): seedSkills opt-in for PTY slash-command tests + tripwire
Wire ClaudePtyOptions.seedSkills through launchClaudePty: when set (and
hermetic, and no per-test CLAUDE_CONFIG_DIR override), the child gets
hermeticSkillsConfigDir() so typed /skill slash commands resolve instead
of dying as Unknown command before any model turn. Opted in at the three
runPlanSkill* helpers and the four direct-launch slash-command tests
(plan-design-with-ui, plan-ceo-mode-routing, autoplan-chain,
ship-idempotency).
New static tripwire (test/pty-skill-seeding-wiring.test.ts): any test
file that sends a slash command over the PTY must route through a
runPlanSkill* helper or pass seedSkills: true — an unseeded slash-command
test spends money and measures nothing. hermetic-wiring.test.ts now
blesses the repo-tree seeding path explicitly (config dir under runRoot,
symlinks into the repo checkout, never operator ~/.claude).
The CI "Register gstack skills for PTY smoke" step keeps a keep-me note:
container cross-mount symlinks defeat the TUI scanner and HOME is not
hermeticized, so the real-file copies there must survive this change.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
(cherry picked from commit 63c52269daaffb833b3105ea9b4b99be6df8fec7)
* refactor(evals): single shared paid-test-set module
test/helpers/paid-test-set.ts is now the one definition of which test
files are paid (the exact globs package.json's test:gate expands).
scripts/test-free-shards.ts derives its free/paid exclusion from it
instead of a private regex list, dropping the dead
browse/test/security-review-fullstack.test.ts pattern (file no longer
exists). The sharded paid runner derives its enumeration from the same
module, so a file added to one list can no longer silently miss the
other.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
(cherry picked from commit a7f36479a6a1f3656452370f5883371f3cb65623)
* feat(evals): env-driven lazy eval dir + shard-aware store and tooling
Importing eval-store no longer spawns the gstack-slug subprocess: the
module-level DEFAULT_EVAL_DIR constant is now a memoized defaultEvalDir()
resolved at collector construction. Resolution order: explicit
constructor arg, then GSTACK_EVAL_DIR, then slug detection — so the
sharded paid runner can point each shard child at its own
<evalDir>/shards/<slug>/ dir with plain env, no --preload.
Runs collected under a shards/ subdir record their slug in the eval
JSON (EvalResult.shard). findPreviousRun scans one shards/<slug>/ level
and prefers same-slug priors, so each shard baselines against its own
history instead of whichever shard flushed last. eval:list,
eval:summary, and eval:compare enumerate the same one level of shard
subdirs; eval:compare's no-arg mode also stops picking an in-progress
accumulator as the after-run.
eval-watch stays flat (documented follow-up): it tails a single dir for
live progress and gains nothing from per-shard baselines until the
runner emits a merged stream.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
(cherry picked from commit e1f53f7d9c7fe6b65877d843f2e25bd2e2d12ffd)
* feat(evals): sharded paid tier runner
scripts/test-paid-shards.ts runs the gate/periodic tier one Bun process
per test file, with an EXTERNAL wall-clock timeout that SIGKILLs the
shard's detached process group and an aggregate that distinguishes
passed / failed / timed-out / never-started — partial execution can no
longer read as a pass. Bun's native --shard/--isolate covers none of
this: no process-group kill (hung claude/codex PTY grandchildren
survive in-process isolation), no never-started taxonomy, no per-shard
env. Each shard child gets GSTACK_EVAL_DIR=<evalDir>/shards/<slug>/
(slug = test filename sans extension, stable across runs) so shard
baselines compare against their own prior runs.
Output classification lives in scripts/test-strict-output.ts (strict
exit-code derivation, incremental fail-line classifier, child signal
forwarding) so the runner and any future strict bun-test wrapper share
one implementation. Enumeration derives from the shared paid-test-set
module; tier exclusion fires only on an explicit whole-file
EVALS_TIER === '<other>' guard.
package.json gains test:gate:sharded / test:periodic:sharded, and
eval:bg:gate / eval:bg:periodic now run the sharded scripts with detach
timeouts sized to the worst case (gate: 49 shards x 30min / 4 jobs ~
6.2h -> 25200s; periodic: 59 -> 28800s).
test/paid-shards.test.ts pins enumeration, tier classification, and the
kill-and-continue property with a real busy-loop shard.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
(cherry picked from commit 5e76bd5931836257f896cedfe4e93912cb759c70)
* feat(security): hash-chained egress receipt ledger (core)
Port lib/egress-receipt from the v2 fork as TypeScript: writeReceipt
(sync, fail-closed via typed EGRESS_RECEIPT_FAILED), best-effort
writeOutcome, readLedger/listReceipts/verifyLedger, GSTACK_HOME ->
GSTACK_STATE_DIR -> ~/.gstack resolution, 0600 ledger under a 0700
security dir, and an mkdir spin lock (2.5s budget) with documented
>10s-mtime stale-lock reclaim.
Changes vs the fork:
- lastRawLine tail-reads the final 4KB instead of loading the whole
ledger, so appends stay O(1) as the file grows.
- WARN-at-size: past 25MB writeReceipt emits one self-explanatory
stderr warning per process (what the ledger is, how to inspect it,
rotation TODO); verifyLedger gains a sizeWarning field. Rotation
TODO carries the chain-genesis sketch (new generation's first record
embeds the prior file's tail hash).
bin/gstack-egress-receipt is a bun script bridging shell callers:
write|outcome subcommands, exit 3 + EGRESS_RECEIPT_FAILED on stderr on
failure; --no-payload records sha256:null for git-class ops.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
(cherry picked from commit 619726a3d77d987a2e50151a5727b3faaaf5fc6a)
* chore(bin): delete dead brain-consumer/reader scripts
bin/gstack-brain-consumer and bin/gstack-brain-reader are byte-identical
dead scripts that POST the repo URL + a Bearer token to a /ingest-repo
endpoint gbrain removed (docs/gbrain-sync.md already documents the
removal in past tense). No live references remain; CHANGELOG mentions
are historical.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
(cherry picked from commit 254ddc69fc5a0270fcc973e36b6a81766d835d2d)
* feat(security): shared shell receipt helpers
bin/gstack-egress-lib.sh (sourced library, gstack-gbrain-lib.sh
precedent) provides _receipted_curl and _receipted_git: write the
egress receipt BEFORE the send via gstack-egress-receipt, hand curl the
SAME payload file via --data-binary @file so the receipt hash matches
the wire bytes exactly, then append a best-effort outcome. Per-call
fail policy: 'closed' refuses the send (return 3, problem/cause/fix
message on stderr) and 'open' warns and proceeds. Payload temp files
are consumed immediately per call — no EXIT traps, since callers like
gstack-telemetry-sync own their own EXIT trap and a sourced trap would
clobber it.
Tested end-to-end against a local Bun.serve listener: receipt sha256
equals the sha256 of the bytes the listener received, fail-closed
refusal never touches the network and carries the problem/cause/fix
stderr shape, fail-open warns and proceeds.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
(cherry picked from commit 6d067dce2d4c8815dec98be551763c85a3671357)
* feat(security): receipt core shell sinks
Wire the three core bash egress sinks through gstack-egress-lib.sh:
- gstack-telemetry-sync: the batch POST now writes the payload to a
temp file, receipts those exact bytes fail-closed, and hands curl the
SAME file. On refusal nothing is sent and the cursor does not
advance, so the batch stays buffered for the next run. The HTTP
status is recorded as the receipt outcome.
- gstack-update-check: fail-open receipts (warn + proceed) on the
Supabase ping POST, both VERSION curls (via a local
_receipted_version_fetch helper that skips non-network schemes), and
git ls-remote. The ping receipt is written inside the backgrounded
subshell, so it can never block the script's exit.
- gstack-brain-sync: fail-closed git-class receipts. The push receipt
is written BEFORE the commit consumes the queue, so a refused receipt
leaves the queue intact and the next run retries the whole drain
(pinned by a new queue-intact-on-refusal test, including the
problem/cause/fix refusal message shape). The retry-path fetch and
retry push carry their own fail-closed receipts.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
(cherry picked from commit 3c60f699acceaf1c92a218874711e05fc17dca5d)
* feat(security): receipt TS module sinks + tunnel
writeReceipt (fail-closed, sha256:null — a subprocess or SDK owns the
wire bytes) before every TS-module network-bearing operation:
- bin/gstack-gbrain-sync.ts: before the gbrain code walk that ships
repo content to the user's gbrain DB (may be remote Postgres). A
refused receipt fails the stage with status refused-egress-receipt.
- bin/gstack-memory-ingest.ts: before the gbrain batch import of
transcript pages. A refused receipt returns a system_error verdict
without spawning the import.
- browse/src/server.ts: before both ngrok.forward call sites (start-up
BROWSE_TUNNEL=1 path and the /tunnel/start endpoint). A receipt
failure lands in the existing catch that tears the tunnel listener
back down and refuses the start.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
(cherry picked from commit 5677d618a48fcd0ae2b068bf868781d90f809cb5)
* feat(design): receipted fetch for OpenAI calls
design/src/receipted-fetch.ts wraps every api.openai.com call: a
content-free egress receipt (sink design-openai, sha256 of the JSON
body — hash only, never the body) is written BEFORE the send. Polarity
is FAIL-OPEN: user-facing generation must not die because an audit log
hiccuped, so a receipt failure warns on stderr and the call proceeds.
Streams pass through untouched (response bodies returned as-is;
non-string request bodies receipted as sha256:null rather than drained
to hash).
All ten call sites converted with per-command payload classes:
generate, variants (injected fetchFn passes through), iterate (both
threaded and fresh paths), evolve (image + screenshot analysis), check,
diff, design-to-code, memory.
Unit-tested with injected fetch: receipt-before-send ordering, stream
passthrough, and fail-open on an unwritable ledger.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
(cherry picked from commit c0e5ff6639414ac2fd98e8ac3affb51401746b55)
* feat(security): receipt admin scripts + user git-ops (zero exceptions)
Wire the remaining shell egress through gstack-egress-lib.sh:
- gstack-gbrain-mcp-verify: both JSON-RPC probe POSTs (initialize +
tools/list) receipted fail-closed via payload files (hash == wire
bytes). A refused receipt lands in the NETWORK class — no send.
- gstack-security-dashboard / gstack-community-dashboard: the
community-pulse GETs receipted fail-open (read-only stats must not
break over an audit hiccup).
- gstack-gbrain-supabase-provision: api_call receipted fail-closed.
Each retry attempt hands the helper a fresh copy of the body file
(the helper consumes its payload). The receipt hashes the request
body only — the PAT never reaches the ledger or any log. Refusal
exits 8 without retrying.
- git-class sha256:null receipts, fail-open: gstack-artifacts-init
(ls-remote, initial push, fetch/pull recovery, retry push),
gstack-brain-restore (staging clone, existing-repo fetch),
gstack-session-update (self-update pull).
gstack-team-init needs no wiring: every git clone in it is inside an
echoed instruction string, not an executed command.
The lib now self-locates with shell builtins only (no dirname), so
sourcing works under the whitelist-PATH test harnesses.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
(cherry picked from commit b8c5e2055b21ab72878b3e46f8047782ee65a11c)
* test(security): egress wiring tripwire + polarity contract
Static-grep tripwire pinning the egress-receipt wiring (threat model in
the header: the ledger is forensic observability of ATTEMPTED egress,
not an exfiltration control):
- Per-sink assertions: every wired TS module imports egress-receipt and
calls writeReceipt; every wired shell sink sources
gstack-egress-lib.sh with each network op under a receipt;
ngrok-proximity check for server.ts; every design api.openai.com call
routes through receiptedFetch.
- Absence assertions: the dead brain-consumer/reader scripts stay
deleted (lstat, so a dangling symlink also fails).
- Polarity table pinned as data (fail-closed: brain-sync,
memory-ingest, gbrain-sync, telemetry-sync, ngrok, mcp-verify,
supabase-provision; fail-open: design-openai, update-check,
dashboards, git-class user ops, context-bill --exact) plus per-file
polarity spot-checks.
- NEW-SINK SCANNER with zero KNOWN_UNWIRED: sweeps bin/, lib/,
scripts/, design/src, browse/src for curl, absolute-URL fetch(, and
git remote ops (never local rev-parse/get-url; heredoc bodies and
message strings excluded) and requires every hit to be receipted or
in a REASONED exemption list where each entry carries its why.
Preamble-generated skill prose documented out-of-scope in the header.
- Shebang tripwire: no bin/gstack-* file may carry a node shebang.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
(cherry picked from commit ff69ceeafaf9c017d539b6ad77ff8f95b680b979)
* feat(cli): gstack-egress reader
bin/gstack-egress (bun) — the auditor's view of the receipts ledger:
- list: one row per receipt (what gstack ATTEMPTED to send), with
--since/--host/--sink filters and --json.
- verify: recompute the hash chain; exit 3 on tamper naming the first
broken line; prints the sizeWarning when the ledger passes 25MB.
- grants: what CAN leave, built on the upstream config keys only
(telemetry, artifacts_sync_mode, redact_repo_visibility,
redact_prepush_hook via gstack-config get) — each grant names its
file, key, and the exact revoke command.
CLI smoke tests spawn the real bin against a temp GSTACK_HOME,
including a broken-chain fixture asserting exit 3.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
(cherry picked from commit 9e24eca0f1069fea2ea69e7df4e9b256e93d59a3)
* feat(cli): context-bill — token bill-of-materials (stripped port)
lib/context-bill.ts, ported from the v2 fork and STRIPPED to the tiers
this repo's skills can exercise: ALWAYS-ON (per-skill frontmatter bytes
with dead-key and foreign-host-file flags), EAGER (SKILL.md + any
forced 'for every invocation' references), on-disk totals, --diff,
--budget, and --exact with the calibration table. The fork's
CONDITIONAL/TRANSITIVE/LAZY/FAST-PATH parsers understand only its
dispatcher layout and were dropped; the tier fields stay in the report
shape (empty/zero/null) so re-adding a parser is additive.
TOKEN_DIVISORS and their provenance docblock kept; --help notes
recalibration via --exact's calibration block.
Three upstream fixes over the fork:
(a) findSkillDirs treats the walk ROOT as a container — the repo root's
router SKILL.md is billed AND its children are walked (the fork
short-circuited and billed one skill); walkMd skips node_modules
and dot-directories.
(b) installed-tree layout: subdirs that are their own repo checkout
(a gstack/ clone inside ~/.claude/skills, detected by .git) are
skipped, and directory symlinks (connect-chrome) are followed with
a container-recursion cycle guard.
(c) ROUTER_KEYS widened to the upstream frontmatter contract {name,
description, version, allowed-tools, triggers, preamble-tier}.
--exact writes an egress receipt (sink context-bill-exact, host
api.anthropic.com) BEFORE any count_tokens POST; if the receipt cannot
be written the run degrades to the offline estimate with a warning —
nothing is sent unrecorded. bin/gstack-context-bill is the bun shim.
Tests: fixture-tree ledgers, the three fixes, --diff/--budget exit
codes, --exact with injected fetch (envelope subtraction, receipt
ordering, fail-open degradation), CLI smoke test, and ground truth
against THIS repo via test/helpers/skill-census.ts.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
(cherry picked from commit 675c19876b87ec927b555f5f64c7f93130b3de90)
* test(catalog): aggregate discovery-surface budget with ratchet protocol
Every host loads every skill's frontmatter name + description at
discovery, every session. applyCatalogTrim in scripts/gen-skill-docs.ts
shapes each description and the 160KB per-file warn covers body size,
but nothing capped the aggregate frontmatter — the catalog could grow
one reasonable-looking description at a time. This test is that
enforcement layer.
Measures the catalog via test/helpers/skill-census.ts authoredSkills
(symlink-deduped, root router counted separately as the _gstack-command
alias line item): 53 skills + router = 4,420 bytes = 1,105
token-equivalents today, asserted <= 1,150 (~4% headroom). Per-skill
sub-cap of 260 bytes (largest today: design-consultation at 229), plus
a non-empty-description check.
Failure messages are self-service ratchets: they print the new total,
the delta, and the update protocol (bump the constant AND the
derivation comment in the same commit; trim instead of grow for
existing descriptions). Parser handles folded block scalars
(description: >-) for fork parity; import-free by design so it
survives generator refactors.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
(cherry picked from commit c106fb36f768181b80c257e5cff1cde4f435f9c0)
* fix(browse): extension token bootstrap moves to pinned-origin POST; /health carries no token
GET /health is now liveness/status only in every mode — both token
carve-outs (headed-mode disjunct AND chrome-extension:// Origin
disjunct) are removed. Token bootstrap is POST /extension-token on the
local listener: the Origin header must be exactly
chrome-extension://<GSTACK_EXTENSION_ID> and the Host header's hostname
must parse to 127.0.0.1 or localhost (parsed via new URL, never literal
equality — Host arrives as '127.0.0.1:34567'). Wrong origin/host → 403
with no detail. The tunnel surface 404s the endpoint (not in
TUNNEL_PATHS, verified by test).
The extension ID is pinned by a new "key" field (RSA public key) in
extension/manifest.json; browse/scripts/extension-id.ts reproduces the
ID derivation (first 16 bytes of SHA-256 of the DER public key, hex
mapped 0-9a-f → a-p). The private key is not committed anywhere —
unpacked/baked-in loads only need the public key.
Extension side: background.js bootstraps and refreshes the token via
POST /extension-token (403 → disconnected state); sidepanel.js direct
connect path does the same; sidepanel-terminal.js's dead /health token
fallback (read AUTH_TOKEN/authToken keys the server never sent,
hardcoded port) is replaced with the window.gstackAuthToken path.
MIGRATION NOTE: the manifest key pins the extension ID, so existing
installs' side-panel local state (saved port, snoozes) resets once —
explained in-product via a one-time notice (flag
gstack_id_migrated_v162). After upgrading the server, restart the
browser so the old service worker stops polling for a token GET /health
no longer serves.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
(cherry picked from commit e9a0b6847a2d17fe6656a4686b4efd0c8380eb09)
* docs: correct stale compiled-binaries claim; file three egress/eval follow-ups
CLAUDE.md's compiled-binaries section claimed browse/dist binaries are
tracked by git and appear as modified in git status — false since
64d5a3e4 (v0.11.16.0) untracked them, and actively harmful: it trained
agents to ignore dist binaries in git status. The section now states
the truth (untracked + gitignored; a dist binary in git status means
someone force-added it) and covers make-pdf/dist too.
TODOS.md gains the three follow-ups filed by the v1.62 port-wave
reviews: ledger rotation with chain-genesis records, launch-nonce
token bootstrap, and eval-watch shard-awareness.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix: pre-landing review fixes for the v2 port wave
Review army (checklist + 5 specialists) + coverage/plan audits on the
assembled branch. Genuine correctness/security/hygiene fixes:
- test-paid-shards: strictTestExitCode now receives expectedFiles on the
real bun path, so a shard that runs fewer files than planned (harness
crash, nothing loaded) with exit 0 is no longer recorded 'passed' — the
invisible-non-execution class the runner exists to kill. Pinned by the
new test/strict-output.test.ts (also covers the chunk-boundary classifier).
- test-paid-shards: EVALS_TIER env is validated (gate|periodic) like the
--tier flag, so a typo can't self-skip every test and exit 0 green.
- package.json: test:periodic:sharded sets EVALS_ALL=1, restoring the
full-tier semantics the pre-shard script had (CI already set it; local
eval:bg:periodic silently under-measured without it).
- brain-sync.test: run() pins HOME to the temp home so gstack-artifacts-init
stops writing/clobbering the operator's real ~/.gstack-artifacts-remote.txt
every free-suite run; afterEach now also scrubs the current filename.
- egress-receipt: cap each receipt field at 512B so a serialized line always
fits the 4KB tail-read window — a longer line would make the next append
hash a truncated prior line and verifyLedger report a permanent false
TAMPER. warnLedgerSize short-circuits before statSync once fired (append
hot path).
- gstack-egress: import.meta.dir (Windows-safe) instead of new URL().pathname
so grants doesn't silently report defaults on Windows; strip control chars
from ledger-derived fields on render so a crafted receipt can't spoof the
auditor's view.
- extension/background.js + CLAUDE.md: renumber the identity-pin migration
refs v1.62 -> v1.63 (main claimed 1.62.0.0; this wave queue-advances).
- egress-receipt-wiring: pin lib/context-bill.ts unconditionally (both land
together now); drop the dead RunShardsOptions.tier field.
All fix-affected test files green; gate failures triaged as external-env
(codex/gemini CLI drift) or pre-existing (hermetic-canary fails identically
on base). Deferred polish tracked in the PR body + decision store.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* chore: bump version and changelog (v1.63.0.0)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* docs: file TODO to harden plan-design-with-ui PTY detection
The v1.63 seedSkills change made this gate test execute for the first
time; it reliably times out because its terminal scraper can't parse the
(correctly-rendered) scope-gate AskUserQuestion out of a spinner-mangled
PTY buffer. Shipped skill behavior is correct — test-harness limitation.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* chore: re-slot release as v1.62.1.0 (PATCH per user)
Main claimed 1.62.0.0 while the wave was in flight; the user chose the
PATCH slot over queue-advancing MINOR. Renumbers the identity-pin
migration notice (now version-free flag name so a re-slot never orphans
an already-set flag), the CLAUDE.md /health note, the CHANGELOG heading,
and the TODOS section titles.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(test): stop hard-requiring the literal ok) case label in gbrain-refresh guards
The extractor grepped for 'ok)' but the case label grew to
ok|timeout|thin-client) (#1964, #2051), so the whole file errored on
import — the free suite's only red for months. The extractor now matches
any label starting with ok and its alternations; all 7 guard assertions
run again.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(test): hermetic-canary probes with ${VAR:-} so nounset shells can't fail success
The probe echoed bare $CONDUCTOR_WORKSPACE_PATH — when scrubbing WORKS
the var is unset, and under a nounset shell the echo errors, failing the
canary exactly when isolation succeeds. Defaulted expansions assert
identically under any shell. Fails identically on base; fixed here.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(evals): absorb codex/gemini CLI drift; external-service tests go periodic-tier
- codex exec gains --skip-git-repo-check: newer CLIs refuse exec in an
untrusted non-git dir (our temp skill dirs) — empirically verified.
- gemini: --skip-trust was removed in gemini-cli 0.34 (argv parse error);
dropped from the session runner and the benchmark adapter. A present-
but-unusable CLI (deprecated individual code-assist auth path) now
classifies as SKIP, not a false adapter failure; the benchmark live
smoke skips on auth/rate_limit error codes (environmental) while still
failing on timeout/unknown (the drift classes it exists to catch).
- codex-e2e, gemini-e2e, and benchmark-providers gain the canonical
whole-file EVALS_TIER === 'periodic' guard per CLAUDE.md tiering rule 3
(external service -> periodic) — the sharded gate runner now excludes
all three (gate: 45 -> 42 shards).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(evals): parse single-logical-line AskUserQuestions in the PTY runner
When the PTY reflows a boxed AUQ, ALL options land on ONE logical line
after stripAnsi — parseNumberedOptions parsed one option per line, found
only '1.', and the >=2 check failed forever while the correct question
sat on screen (plan-design-with-ui timed out this way twice, with the
rendered scope-gate AUQ visible in both failure buffers). The cursor
line is now parsed as a stream of ascending N. tokens; DEC cursor-
visibility residue is stripped before matching; plan-design-with-ui's
budgets grow to fit observed ~6min preamble+thinking latency. Pinned by
test/pty-auq-single-line.test.ts using the real failure buffers; all 142
existing parser-consumer unit tests still green.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* chore: restore v1.63.0.0 (MINOR — user-confirmed final slot)
The wave ships new capability (egress receipts + two CLIs, sharded paid
runner, hermetic skill seeding) at ~8K lines — MINOR scale per the
scale-aware bump rules. Supersedes the brief v1.62.1.0 re-slot; the
version-free migration flag means no state churn from the renumber.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* test: spell out AskUserQuestion in the PTY single-line fixture
Rename test/pty-auq-single-line.test.ts to
test/pty-askuserquestion-single-line.test.ts and expand the AUQ
abbreviation in identifiers and comments. House style writes
AskUserQuestion in full in filenames, identifiers, and comments.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* docs: sync every doc surface with the v1.63 release
/document-release audit (4-lane, all claims verified against branch code):
- README: gstack-egress + gstack-context-bill rows in the standalone-binaries
table; Privacy & Telemetry gains the receipted-egress bullet (attempted-
egress framing per the shipped threat model).
- ARCHITECTURE: /health is liveness-only, POST /extension-token endpoint row
+ bootstrap mechanics paragraph; new Egress receipt ledger subsection under
Security model; eval persistence covers the sharded runner, GSTACK_EVAL_DIR,
and the finalized-run baseline rule.
- CLAUDE.md: sharded test scripts in Commands; sharded semantics in the
detached-evals section; PTY skill seeding in the hermetic section; egress
invariant block beside the other server-egress invariants; catalog-budget
ceiling beside the 160KB token ceiling; project-tree entries for
lib/egress-receipt.ts, lib/context-bill.ts, scripts/test-paid-shards.ts.
- CONTRIBUTING: seedSkills + live-tree seeding in the hermetic paragraph;
sharded runner in detached runs; catalog-budget in the Tier 1 list.
- BROWSER: extension token bootstrap section, tunnel egress receipts section,
identity-pin migration note in manual install.
- REMOTE_BROWSER_ACCESS: tunnel-start receipt bullet in the security model.
- gbrain docs: /sync-gbrain + brain-sync egress-receipt behavior documented;
dead consumer-token instructions removed (consumer machinery deleted this
release); new fail-closed refusal added to the error catalog.
- CHANGELOG: measured-vs-ceiling catalog numbers, contributor notes for the
external-service tier move and the PTY single-line AskUserQuestion parser,
release date.
- TODOS: /health token-distribution TODO resolved by this release, removed;
port-wave follow-up sections re-labeled to the shipped version.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* docs: sweep drift that predates this release
Surfaced by the /document-release audit; every fix verified against the
current binaries:
- gstack-brain-init was replaced by gstack-artifacts-init in v1.27.0.0
(hard-delete, no compat shim), but README, USING_GBRAIN_WITH_GSTACK,
docs/gbrain-sync.md, and docs/gbrain-sync-errors.md still instructed
users to run it — command-not-found on every follow. Same sweep updates
~/.gstack-brain-remote.txt to the canonical ~/.gstack-artifacts-remote.txt
(legacy name still honored on restore, noted where users copy the file).
- gbrain-sync-errors.md headings re-matched to the literal messages the
binaries print today (the doc's whole value is grep-by-exact-message):
'gstack-artifacts-init: ~/.gstack/ is already a git repo pointing at:',
'Remote not reachable via SSH:', 'Failed to create or find ...'. The
already-a-repo fix now leads with the command's own set-url suggestion.
- docs/gbrain-sync.md 'Under the hood' linked a plan file that does not
exist in the repo; replaced with the decisions themselves.
- SIDEBAR_MESSAGE_FLOW startup timeline: /pty-session responds with
{terminalPort, sessionId, attachToken, leaseExpiresAt} (v1.44 shape,
verified at browse/src/server.ts:1860), not the retired
{terminalPort, ptySessionToken} pair.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* docs: fold the Codex accuracy review of the release docs
Six findings, all verified against source before fixing:
1. 'Every send writes a receipt' overclaimed — fail-open sinks proceed with
a stderr warning when the receipt write fails, so a fail-open send can go
unrecorded (lib/egress-receipt.ts:8-14). Descriptive prose now says so;
the receipted framing keeps 'attempted'.
2. 'Receipts hash the request body' is wrong for subprocess-owned sends —
git pushes record sha256: null (lib/egress-receipt.ts:71).
3. 'grants shows every consent in force' overclaimed — it reports the four
standing config settings (bin/gstack-egress:139-181). Reworded in
README, ARCHITECTURE, and the CHANGELOG entry.
4. 'Zero-exception scanner' vs reality: the new-sink scanner carries a
reasoned SCANNER_EXEMPT list (user-directed fetches, probes, instruction
strings, skill prose). CLAUDE.md now names it.
5. Error-catalog cause/fix for the receipt refusal: the writer mkdirs the
ledger dir itself, so 'missing' isn't a cause and bare chmod fails when
it is absent — cause reworded, fix is mkdir -p && chmod.
6. gbrain-sync first-run steps described the retired binary's behavior:
default repo is gstack-artifacts-$USER, and init PRINTS the gbrain
hookup command (never auto-executes; bin/gstack-artifacts-init:384-419).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Co-authored-by: Sinabina <sinabina@Sinabinas-MacBook-Pro-4.local>
915 lines
33 KiB
TypeScript
915 lines
33 KiB
TypeScript
/**
|
||
* Eval result persistence and comparison.
|
||
*
|
||
* EvalCollector accumulates test results, writes them to
|
||
* ~/.gstack/projects/$SLUG/evals/{version}-{branch}-{tier}-{timestamp}.json,
|
||
* prints a summary table, and auto-compares with the previous run.
|
||
*
|
||
* Comparison functions are exported for reuse by the eval:compare CLI.
|
||
*/
|
||
|
||
import * as fs from 'fs';
|
||
import * as path from 'path';
|
||
import * as os from 'os';
|
||
import { spawnSync } from 'child_process';
|
||
|
||
const SCHEMA_VERSION = 1;
|
||
const LEGACY_EVAL_DIR = path.join(os.homedir(), '.gstack-dev', 'evals');
|
||
|
||
/**
|
||
* Detect project-scoped eval dir via gstack-slug.
|
||
* Falls back to legacy ~/.gstack-dev/evals/ if slug detection fails.
|
||
*/
|
||
export function getProjectEvalDir(): string {
|
||
try {
|
||
// Try repo-local gstack-slug first, then global install
|
||
const localSlug = spawnSync('bash', ['-c', '.claude/skills/gstack/bin/gstack-slug 2>/dev/null || ~/.claude/skills/gstack/bin/gstack-slug 2>/dev/null'], {
|
||
stdio: 'pipe', timeout: 3000,
|
||
});
|
||
const output = localSlug.stdout?.toString().trim();
|
||
if (output) {
|
||
const slugMatch = output.match(/^SLUG=(.+)$/m);
|
||
if (slugMatch && slugMatch[1]) {
|
||
const dir = path.join(os.homedir(), '.gstack', 'projects', slugMatch[1], 'evals');
|
||
fs.mkdirSync(dir, { recursive: true });
|
||
return dir;
|
||
}
|
||
}
|
||
} catch { /* fall through */ }
|
||
return LEGACY_EVAL_DIR;
|
||
}
|
||
|
||
/**
|
||
* Lazy + memoized so importing this module never spawns the gstack-slug
|
||
* subprocess. Callers that pass an explicit dir or set GSTACK_EVAL_DIR
|
||
* (the sharded paid runner does, per shard) never pay for slug detection.
|
||
*/
|
||
let memoizedDefaultEvalDir: string | null = null;
|
||
function defaultEvalDir(): string {
|
||
if (memoizedDefaultEvalDir === null) memoizedDefaultEvalDir = getProjectEvalDir();
|
||
return memoizedDefaultEvalDir;
|
||
}
|
||
|
||
// --- Interfaces ---
|
||
|
||
export interface EvalTestEntry {
|
||
name: string;
|
||
suite: string;
|
||
tier: 'e2e' | 'llm-judge';
|
||
passed: boolean;
|
||
duration_ms: number;
|
||
cost_usd: number;
|
||
|
||
// E2E
|
||
transcript?: any[];
|
||
prompt?: string;
|
||
output?: string;
|
||
turns_used?: number;
|
||
browse_errors?: string[];
|
||
|
||
// LLM judge
|
||
judge_scores?: Record<string, number>;
|
||
judge_reasoning?: string;
|
||
|
||
// Machine-readable diagnostics
|
||
exit_reason?: string; // 'success' | 'timeout' | 'error_max_turns' | 'exit_code_N'
|
||
timeout_at_turn?: number; // which turn was active when timeout hit
|
||
last_tool_call?: string; // e.g. "Write(review-output.md)"
|
||
|
||
// Model + timing diagnostics (added for Sonnet/Opus split)
|
||
model?: string; // e.g. 'claude-sonnet-4-6' or 'claude-opus-4-7'
|
||
first_response_ms?: number; // time from spawn to first NDJSON line
|
||
max_inter_turn_ms?: number; // peak latency between consecutive tool calls
|
||
|
||
// Outcome eval
|
||
detection_rate?: number;
|
||
false_positives?: number;
|
||
evidence_quality?: number;
|
||
detected_bugs?: string[];
|
||
missed_bugs?: string[];
|
||
|
||
error?: string;
|
||
|
||
// Worktree harvest data
|
||
harvest?: {
|
||
filesChanged: number;
|
||
patchPath: string;
|
||
isDuplicate: boolean;
|
||
};
|
||
}
|
||
|
||
export interface EvalResult {
|
||
schema_version: number;
|
||
version: string;
|
||
branch: string;
|
||
git_sha: string;
|
||
timestamp: string;
|
||
hostname: string;
|
||
tier: 'e2e' | 'llm-judge';
|
||
total_tests: number;
|
||
passed: number;
|
||
failed: number;
|
||
total_cost_usd: number;
|
||
total_duration_ms: number;
|
||
wall_clock_ms?: number; // wall-clock from collector creation to finalization (shows parallelism)
|
||
tests: EvalTestEntry[];
|
||
/** Shard slug when the run was collected under <evalDir>/shards/<slug>/. */
|
||
shard?: string;
|
||
_partial?: boolean; // true for incremental saves, absent in final
|
||
}
|
||
|
||
export interface TestDelta {
|
||
name: string;
|
||
before: { passed: boolean; cost_usd: number; turns_used?: number; duration_ms?: number;
|
||
detection_rate?: number; tool_summary?: Record<string, number> };
|
||
after: { passed: boolean; cost_usd: number; turns_used?: number; duration_ms?: number;
|
||
detection_rate?: number; tool_summary?: Record<string, number> };
|
||
status_change: 'improved' | 'regressed' | 'unchanged';
|
||
}
|
||
|
||
export interface ComparisonResult {
|
||
before_file: string;
|
||
after_file: string;
|
||
before_branch: string;
|
||
after_branch: string;
|
||
before_timestamp: string;
|
||
after_timestamp: string;
|
||
deltas: TestDelta[];
|
||
total_cost_delta: number;
|
||
total_duration_delta: number;
|
||
improved: number;
|
||
regressed: number;
|
||
unchanged: number;
|
||
tool_count_before: number;
|
||
tool_count_after: number;
|
||
/** After-tests that had a same-named entry in the before run. 0 = nothing was
|
||
* actually compared, so no stability claim is warranted. */
|
||
matched?: number;
|
||
}
|
||
|
||
// --- Shared helpers ---
|
||
|
||
/**
|
||
* Is this eval file an in-progress accumulator rather than a finalized run?
|
||
*
|
||
* True on either signal: the `_partial` flag inside the JSON (the authoritative
|
||
* role marker) OR a filename starting with `_partial` (catches accumulators
|
||
* whose body predates the flag, and flagged files that were renamed keep being
|
||
* caught by the flag). Every baseline lookup must exclude these — an
|
||
* accumulator carries the current run's tier, branch, and freshest timestamp,
|
||
* so treating it as a baseline makes the run compare against itself.
|
||
*/
|
||
export function isPartialEval(data: unknown, filename: string): boolean {
|
||
if (path.basename(filename).startsWith('_partial')) return true;
|
||
return Boolean((data as { _partial?: unknown } | null)?._partial);
|
||
}
|
||
|
||
/**
|
||
* List eval JSON files in `evalDir` plus one level of `<evalDir>/shards/<slug>/`
|
||
* subdirectories (where the sharded paid runner points each shard's collector).
|
||
* Returns absolute paths. Missing dirs yield [].
|
||
*/
|
||
export function listEvalJsonFiles(evalDir: string): string[] {
|
||
const jsonIn = (dir: string): string[] => {
|
||
let names: string[];
|
||
try {
|
||
names = fs.readdirSync(dir);
|
||
} catch {
|
||
return [];
|
||
}
|
||
return names.filter(f => f.endsWith('.json')).map(f => path.join(dir, f));
|
||
};
|
||
|
||
const files = jsonIn(evalDir);
|
||
const shardsRoot = path.join(evalDir, 'shards');
|
||
let shardDirs: fs.Dirent[];
|
||
try {
|
||
shardDirs = fs.readdirSync(shardsRoot, { withFileTypes: true });
|
||
} catch {
|
||
return files;
|
||
}
|
||
for (const entry of shardDirs) {
|
||
if (!entry.isDirectory()) continue;
|
||
files.push(...jsonIn(path.join(shardsRoot, entry.name)));
|
||
}
|
||
return files;
|
||
}
|
||
|
||
/**
|
||
* Shard slug for an eval dir: when the dir is directly under a `shards/`
|
||
* directory (the sharded paid runner's per-shard GSTACK_EVAL_DIR layout),
|
||
* the dir name is the slug; otherwise null.
|
||
*/
|
||
export function shardSlugOfEvalDir(evalDir: string): string | null {
|
||
const normalized = path.resolve(evalDir);
|
||
return path.basename(path.dirname(normalized)) === 'shards' ? path.basename(normalized) : null;
|
||
}
|
||
|
||
/**
|
||
* Find the most recent finalized (non-partial) eval file for a tier, scanning
|
||
* `evalDir` and one level of `shards/<slug>/` subdirs. Shared by the budget
|
||
* regression gate and any tooling that needs "the latest real run".
|
||
*/
|
||
export function findLatestFinalizedRun(
|
||
evalDir: string,
|
||
tier: 'e2e' | 'llm-judge',
|
||
): { filepath: string; result: EvalResult } | null {
|
||
let latest: { filepath: string; result: EvalResult; timestamp: string } | null = null;
|
||
for (const filepath of listEvalJsonFiles(evalDir)) {
|
||
let data: EvalResult;
|
||
try {
|
||
data = JSON.parse(fs.readFileSync(filepath, 'utf-8')) as EvalResult;
|
||
} catch { continue; }
|
||
if (isPartialEval(data, filepath)) continue;
|
||
if (data.tier !== tier) continue;
|
||
const timestamp = data.timestamp ?? '';
|
||
if (!latest || timestamp.localeCompare(latest.timestamp) > 0) {
|
||
latest = { filepath, result: data, timestamp };
|
||
}
|
||
}
|
||
return latest ? { filepath: latest.filepath, result: latest.result } : null;
|
||
}
|
||
|
||
/**
|
||
* Determine if a planted-bug eval passed based on judge results vs ground truth thresholds.
|
||
* Centralizes the pass/fail logic so all planted-bug tests use the same criteria.
|
||
*/
|
||
export function judgePassed(
|
||
judgeResult: { detection_rate: number; false_positives: number; evidence_quality: number },
|
||
groundTruth: { minimum_detection: number; max_false_positives: number },
|
||
): boolean {
|
||
return judgeResult.detection_rate >= groundTruth.minimum_detection
|
||
&& judgeResult.false_positives <= groundTruth.max_false_positives
|
||
&& judgeResult.evidence_quality >= 2;
|
||
}
|
||
|
||
// --- Comparison functions (exported for eval:compare CLI) ---
|
||
|
||
/**
|
||
* Extract tool call counts from a transcript.
|
||
* Returns e.g. { Bash: 8, Read: 3, Write: 1 }.
|
||
*/
|
||
export function extractToolSummary(transcript: any[]): Record<string, number> {
|
||
const counts: Record<string, number> = {};
|
||
for (const event of transcript) {
|
||
if (event.type === 'assistant') {
|
||
const content = event.message?.content || [];
|
||
for (const item of content) {
|
||
if (item.type === 'tool_use') {
|
||
const name = item.name || 'unknown';
|
||
counts[name] = (counts[name] || 0) + 1;
|
||
}
|
||
}
|
||
}
|
||
}
|
||
return counts;
|
||
}
|
||
|
||
/**
|
||
* Find the most recent prior COMPLETED eval file for comparison.
|
||
* Scans the eval dir plus one level of `shards/<slug>/` subdirs. Prefers
|
||
* same shard slug (a shard's own history over another shard's or the flat
|
||
* dir's), then same branch, then falls back to anything.
|
||
*
|
||
* In-progress accumulators (`_partial: true`, written by savePartial after every
|
||
* test) are never candidates: the current run's own partial carries the current
|
||
* tier + branch and the freshest timestamp, so including it made every run
|
||
* compare against itself and report "no regressions" unconditionally. The
|
||
* exclusion is by role (the `_partial` flag), not by filename.
|
||
*/
|
||
export function findPreviousRun(
|
||
evalDir: string,
|
||
tier: string,
|
||
branch: string,
|
||
excludeFile: string,
|
||
): string | null {
|
||
// Parse top-level fields from each file (cheap — no full tests array needed)
|
||
const entries: Array<{ file: string; branch: string; timestamp: string; shard: string | null }> = [];
|
||
for (const fullPath of listEvalJsonFiles(evalDir)) {
|
||
if (path.resolve(fullPath) === path.resolve(excludeFile)) continue;
|
||
try {
|
||
const raw = fs.readFileSync(fullPath, 'utf-8');
|
||
// Quick parse — only grab the fields we need
|
||
const data = JSON.parse(raw);
|
||
if (isPartialEval(data, fullPath)) continue; // in-progress run, not a baseline
|
||
if (data.tier !== tier) continue;
|
||
entries.push({
|
||
file: fullPath,
|
||
branch: data.branch || '',
|
||
timestamp: data.timestamp || '',
|
||
shard: data.shard || shardSlugOfEvalDir(path.dirname(fullPath)),
|
||
});
|
||
} catch { continue; }
|
||
}
|
||
|
||
if (entries.length === 0) return null;
|
||
|
||
// Sort by timestamp descending
|
||
entries.sort((a, b) => b.timestamp.localeCompare(a.timestamp));
|
||
|
||
// Prefer same shard slug (null = the flat dir), then same branch, then any.
|
||
const targetShard = shardSlugOfEvalDir(path.dirname(excludeFile));
|
||
const preferences: Array<(e: typeof entries[number]) => boolean> = [
|
||
e => e.shard === targetShard && e.branch === branch,
|
||
e => e.shard === targetShard,
|
||
e => e.branch === branch,
|
||
];
|
||
for (const matches of preferences) {
|
||
const hit = entries.find(matches);
|
||
if (hit) return hit.file;
|
||
}
|
||
return entries[0].file;
|
||
}
|
||
|
||
/**
|
||
* Compare two eval results. Matches tests by name.
|
||
*/
|
||
export function compareEvalResults(
|
||
before: EvalResult,
|
||
after: EvalResult,
|
||
beforeFile: string,
|
||
afterFile: string,
|
||
): ComparisonResult {
|
||
const deltas: TestDelta[] = [];
|
||
let improved = 0, regressed = 0, unchanged = 0;
|
||
let toolCountBefore = 0, toolCountAfter = 0;
|
||
let matched = 0;
|
||
|
||
// Index before tests by name
|
||
const beforeMap = new Map<string, EvalTestEntry>();
|
||
for (const t of before.tests) {
|
||
beforeMap.set(t.name, t);
|
||
}
|
||
|
||
// Walk after tests, match by name
|
||
for (const afterTest of after.tests) {
|
||
const beforeTest = beforeMap.get(afterTest.name);
|
||
const beforeToolSummary = beforeTest?.transcript ? extractToolSummary(beforeTest.transcript) : {};
|
||
const afterToolSummary = afterTest.transcript ? extractToolSummary(afterTest.transcript) : {};
|
||
|
||
const beforeToolCount = Object.values(beforeToolSummary).reduce((a, b) => a + b, 0);
|
||
const afterToolCount = Object.values(afterToolSummary).reduce((a, b) => a + b, 0);
|
||
toolCountBefore += beforeToolCount;
|
||
toolCountAfter += afterToolCount;
|
||
|
||
let statusChange: TestDelta['status_change'] = 'unchanged';
|
||
if (beforeTest) {
|
||
matched++;
|
||
if (!beforeTest.passed && afterTest.passed) { statusChange = 'improved'; improved++; }
|
||
else if (beforeTest.passed && !afterTest.passed) { statusChange = 'regressed'; regressed++; }
|
||
else { unchanged++; }
|
||
} else {
|
||
// New test — treat as unchanged (no prior data)
|
||
unchanged++;
|
||
}
|
||
|
||
deltas.push({
|
||
name: afterTest.name,
|
||
before: {
|
||
passed: beforeTest?.passed ?? false,
|
||
cost_usd: beforeTest?.cost_usd ?? 0,
|
||
turns_used: beforeTest?.turns_used,
|
||
duration_ms: beforeTest?.duration_ms,
|
||
detection_rate: beforeTest?.detection_rate,
|
||
tool_summary: beforeToolSummary,
|
||
},
|
||
after: {
|
||
passed: afterTest.passed,
|
||
cost_usd: afterTest.cost_usd,
|
||
turns_used: afterTest.turns_used,
|
||
duration_ms: afterTest.duration_ms,
|
||
detection_rate: afterTest.detection_rate,
|
||
tool_summary: afterToolSummary,
|
||
},
|
||
status_change: statusChange,
|
||
});
|
||
|
||
beforeMap.delete(afterTest.name);
|
||
}
|
||
|
||
// Tests that were in before but not in after (removed tests)
|
||
for (const [name, beforeTest] of beforeMap) {
|
||
const beforeToolSummary = beforeTest.transcript ? extractToolSummary(beforeTest.transcript) : {};
|
||
const beforeToolCount = Object.values(beforeToolSummary).reduce((a, b) => a + b, 0);
|
||
toolCountBefore += beforeToolCount;
|
||
unchanged++;
|
||
deltas.push({
|
||
name: `${name} (removed)`,
|
||
before: {
|
||
passed: beforeTest.passed,
|
||
cost_usd: beforeTest.cost_usd,
|
||
turns_used: beforeTest.turns_used,
|
||
duration_ms: beforeTest.duration_ms,
|
||
detection_rate: beforeTest.detection_rate,
|
||
tool_summary: beforeToolSummary,
|
||
},
|
||
after: { passed: false, cost_usd: 0, tool_summary: {} },
|
||
status_change: 'unchanged',
|
||
});
|
||
}
|
||
|
||
return {
|
||
before_file: beforeFile,
|
||
after_file: afterFile,
|
||
before_branch: before.branch,
|
||
after_branch: after.branch,
|
||
before_timestamp: before.timestamp,
|
||
after_timestamp: after.timestamp,
|
||
deltas,
|
||
total_cost_delta: after.total_cost_usd - before.total_cost_usd,
|
||
total_duration_delta: after.total_duration_ms - before.total_duration_ms,
|
||
improved,
|
||
regressed,
|
||
unchanged,
|
||
tool_count_before: toolCountBefore,
|
||
tool_count_after: toolCountAfter,
|
||
matched,
|
||
};
|
||
}
|
||
|
||
/**
|
||
* Format a ComparisonResult as a readable string.
|
||
*/
|
||
export function formatComparison(c: ComparisonResult): string {
|
||
const lines: string[] = [];
|
||
const ts = c.before_timestamp ? c.before_timestamp.replace('T', ' ').slice(0, 16) : 'unknown';
|
||
lines.push(`\nvs previous: ${c.before_branch}/${c.deltas.length ? 'eval' : ''} (${ts})`);
|
||
lines.push('─'.repeat(70));
|
||
|
||
// Per-test deltas
|
||
for (const d of c.deltas) {
|
||
const arrow = d.status_change === 'improved' ? '↑' : d.status_change === 'regressed' ? '↓' : '=';
|
||
const beforeStatus = d.before.passed ? 'PASS' : 'FAIL';
|
||
const afterStatus = d.after.passed ? 'PASS' : 'FAIL';
|
||
|
||
// Turns delta
|
||
let turnsDelta = '';
|
||
if (d.before.turns_used !== undefined && d.after.turns_used !== undefined) {
|
||
const td = d.after.turns_used - d.before.turns_used;
|
||
turnsDelta = ` ${d.before.turns_used}→${d.after.turns_used}t`;
|
||
if (td !== 0) turnsDelta += `(${td > 0 ? '+' : ''}${td})`;
|
||
} else if (d.after.turns_used !== undefined) {
|
||
turnsDelta = ` ${d.after.turns_used}t`;
|
||
}
|
||
|
||
// Duration delta
|
||
let durDelta = '';
|
||
if (d.before.duration_ms !== undefined && d.after.duration_ms !== undefined) {
|
||
const bs = Math.round(d.before.duration_ms / 1000);
|
||
const as = Math.round(d.after.duration_ms / 1000);
|
||
const dd = as - bs;
|
||
durDelta = ` ${bs}→${as}s`;
|
||
if (dd !== 0) durDelta += `(${dd > 0 ? '+' : ''}${dd})`;
|
||
} else if (d.after.duration_ms !== undefined) {
|
||
durDelta = ` ${Math.round(d.after.duration_ms / 1000)}s`;
|
||
}
|
||
|
||
let detail = '';
|
||
if (d.before.detection_rate !== undefined || d.after.detection_rate !== undefined) {
|
||
detail = ` ${d.before.detection_rate ?? '?'}→${d.after.detection_rate ?? '?'} det`;
|
||
} else {
|
||
const costBefore = d.before.cost_usd.toFixed(2);
|
||
const costAfter = d.after.cost_usd.toFixed(2);
|
||
detail = ` $${costBefore}→$${costAfter}`;
|
||
}
|
||
|
||
const name = d.name.length > 30 ? d.name.slice(0, 27) + '...' : d.name.padEnd(30);
|
||
lines.push(` ${name} ${beforeStatus.padEnd(5)} → ${afterStatus.padEnd(5)} ${arrow}${detail}${turnsDelta}${durDelta}`);
|
||
}
|
||
|
||
lines.push('─'.repeat(70));
|
||
|
||
// Totals
|
||
const parts: string[] = [];
|
||
if (c.improved > 0) parts.push(`${c.improved} improved`);
|
||
if (c.regressed > 0) parts.push(`${c.regressed} regressed`);
|
||
if (c.unchanged > 0) parts.push(`${c.unchanged} unchanged`);
|
||
lines.push(` Status: ${parts.join(', ')}`);
|
||
|
||
const costSign = c.total_cost_delta >= 0 ? '+' : '';
|
||
lines.push(` Cost: ${costSign}$${c.total_cost_delta.toFixed(2)}`);
|
||
|
||
const durDelta = Math.round(c.total_duration_delta / 1000);
|
||
const durSign = durDelta >= 0 ? '+' : '';
|
||
lines.push(` Duration: ${durSign}${durDelta}s`);
|
||
|
||
const toolDelta = c.tool_count_after - c.tool_count_before;
|
||
const toolSign = toolDelta >= 0 ? '+' : '';
|
||
lines.push(` Tool calls: ${c.tool_count_before} → ${c.tool_count_after} (${toolSign}${toolDelta})`);
|
||
|
||
// Tool breakdown (show tools that changed)
|
||
const allTools = new Set<string>();
|
||
for (const d of c.deltas) {
|
||
for (const t of Object.keys(d.before.tool_summary || {})) allTools.add(t);
|
||
for (const t of Object.keys(d.after.tool_summary || {})) allTools.add(t);
|
||
}
|
||
|
||
if (allTools.size > 0) {
|
||
// Aggregate tool counts across all tests
|
||
const totalBefore: Record<string, number> = {};
|
||
const totalAfter: Record<string, number> = {};
|
||
for (const d of c.deltas) {
|
||
for (const [t, n] of Object.entries(d.before.tool_summary || {})) {
|
||
totalBefore[t] = (totalBefore[t] || 0) + n;
|
||
}
|
||
for (const [t, n] of Object.entries(d.after.tool_summary || {})) {
|
||
totalAfter[t] = (totalAfter[t] || 0) + n;
|
||
}
|
||
}
|
||
|
||
for (const tool of [...allTools].sort()) {
|
||
const b = totalBefore[tool] || 0;
|
||
const a = totalAfter[tool] || 0;
|
||
if (b !== a) {
|
||
const d = a - b;
|
||
lines.push(` ${tool}: ${b} → ${a} (${d >= 0 ? '+' : ''}${d})`);
|
||
}
|
||
}
|
||
}
|
||
|
||
// Commentary — interpret what the deltas mean
|
||
const commentary = generateCommentary(c);
|
||
if (commentary.length > 0) {
|
||
lines.push('');
|
||
lines.push(' Takeaway:');
|
||
for (const line of commentary) {
|
||
lines.push(` ${line}`);
|
||
}
|
||
}
|
||
|
||
return lines.join('\n');
|
||
}
|
||
|
||
/**
|
||
* Generate human-readable commentary interpreting comparison deltas.
|
||
* Pure function — analyzes the numbers and explains what they mean.
|
||
*/
|
||
export function generateCommentary(c: ComparisonResult): string[] {
|
||
const notes: string[] = [];
|
||
|
||
// 1. Regressions are the most important signal — call them out first
|
||
const regressions = c.deltas.filter(d => d.status_change === 'regressed');
|
||
if (regressions.length > 0) {
|
||
for (const d of regressions) {
|
||
notes.push(`REGRESSION: "${d.name}" was passing, now fails. Investigate immediately.`);
|
||
}
|
||
}
|
||
|
||
// 2. Improvements
|
||
const improvements = c.deltas.filter(d => d.status_change === 'improved');
|
||
for (const d of improvements) {
|
||
notes.push(`Fixed: "${d.name}" now passes.`);
|
||
}
|
||
|
||
// 3. Per-test efficiency changes (only for unchanged-status tests — regressions/improvements are already noted)
|
||
const stable = c.deltas.filter(d => d.status_change === 'unchanged' && d.after.passed);
|
||
for (const d of stable) {
|
||
const insights: string[] = [];
|
||
|
||
// Turns
|
||
if (d.before.turns_used !== undefined && d.after.turns_used !== undefined && d.before.turns_used > 0) {
|
||
const turnsDelta = d.after.turns_used - d.before.turns_used;
|
||
const turnsPct = Math.round((turnsDelta / d.before.turns_used) * 100);
|
||
if (Math.abs(turnsPct) >= 20 && Math.abs(turnsDelta) >= 2) {
|
||
if (turnsDelta < 0) {
|
||
insights.push(`${Math.abs(turnsDelta)} fewer turns (${Math.abs(turnsPct)}% more efficient)`);
|
||
} else {
|
||
insights.push(`${turnsDelta} more turns (${turnsPct}% less efficient)`);
|
||
}
|
||
}
|
||
}
|
||
|
||
// Duration
|
||
if (d.before.duration_ms !== undefined && d.after.duration_ms !== undefined && d.before.duration_ms > 0) {
|
||
const durDelta = d.after.duration_ms - d.before.duration_ms;
|
||
const durPct = Math.round((durDelta / d.before.duration_ms) * 100);
|
||
if (Math.abs(durPct) >= 20 && Math.abs(durDelta) >= 5000) {
|
||
if (durDelta < 0) {
|
||
insights.push(`${Math.round(Math.abs(durDelta) / 1000)}s faster`);
|
||
} else {
|
||
insights.push(`${Math.round(durDelta / 1000)}s slower`);
|
||
}
|
||
}
|
||
}
|
||
|
||
// Detection rate
|
||
if (d.before.detection_rate !== undefined && d.after.detection_rate !== undefined) {
|
||
const detDelta = d.after.detection_rate - d.before.detection_rate;
|
||
if (detDelta !== 0) {
|
||
if (detDelta > 0) {
|
||
insights.push(`detecting ${detDelta} more bug${detDelta > 1 ? 's' : ''}`);
|
||
} else {
|
||
insights.push(`detecting ${Math.abs(detDelta)} fewer bug${Math.abs(detDelta) > 1 ? 's' : ''} — check prompt quality`);
|
||
}
|
||
}
|
||
}
|
||
|
||
// Cost
|
||
if (d.before.cost_usd > 0) {
|
||
const costDelta = d.after.cost_usd - d.before.cost_usd;
|
||
const costPct = Math.round((costDelta / d.before.cost_usd) * 100);
|
||
if (Math.abs(costPct) >= 30 && Math.abs(costDelta) >= 0.05) {
|
||
if (costDelta < 0) {
|
||
insights.push(`${Math.abs(costPct)}% cheaper`);
|
||
} else {
|
||
insights.push(`${costPct}% more expensive`);
|
||
}
|
||
}
|
||
}
|
||
|
||
if (insights.length > 0) {
|
||
notes.push(`"${d.name}": ${insights.join(', ')}.`);
|
||
}
|
||
}
|
||
|
||
// 4. No baseline — say so. A run with nothing to compare against must never
|
||
// read as "stable"; silence or a false all-clear is worse than no output.
|
||
if (c.matched === 0 && c.deltas.length > 0) {
|
||
notes.push(
|
||
`NO BASELINE: none of these ${c.deltas.length} test(s) appear in ${path.basename(c.before_file)}. ` +
|
||
'Nothing was compared, so this run says nothing about regressions.',
|
||
);
|
||
return notes;
|
||
}
|
||
|
||
// 5. Overall summary
|
||
if (c.deltas.length >= 3 && regressions.length === 0) {
|
||
const overallParts: string[] = [];
|
||
|
||
// Total cost
|
||
const totalBefore = c.deltas.reduce((s, d) => s + d.before.cost_usd, 0);
|
||
if (totalBefore > 0) {
|
||
const costPct = Math.round((c.total_cost_delta / totalBefore) * 100);
|
||
if (Math.abs(costPct) >= 10) {
|
||
overallParts.push(`${Math.abs(costPct)}% ${costPct < 0 ? 'cheaper' : 'more expensive'} overall`);
|
||
}
|
||
}
|
||
|
||
// Total duration
|
||
const totalDurBefore = c.deltas.reduce((s, d) => s + (d.before.duration_ms || 0), 0);
|
||
if (totalDurBefore > 0) {
|
||
const durPct = Math.round((c.total_duration_delta / totalDurBefore) * 100);
|
||
if (Math.abs(durPct) >= 10) {
|
||
overallParts.push(`${Math.abs(durPct)}% ${durPct < 0 ? 'faster' : 'slower'}`);
|
||
}
|
||
}
|
||
|
||
// Total turns
|
||
const turnsBefore = c.deltas.reduce((s, d) => s + (d.before.turns_used || 0), 0);
|
||
const turnsAfter = c.deltas.reduce((s, d) => s + (d.after.turns_used || 0), 0);
|
||
if (turnsBefore > 0) {
|
||
const turnsPct = Math.round(((turnsAfter - turnsBefore) / turnsBefore) * 100);
|
||
if (Math.abs(turnsPct) >= 10) {
|
||
overallParts.push(`${Math.abs(turnsPct)}% ${turnsPct < 0 ? 'fewer' : 'more'} turns`);
|
||
}
|
||
}
|
||
|
||
if (overallParts.length > 0) {
|
||
notes.push(`Overall: ${overallParts.join(', ')}. ${regressions.length === 0 ? 'No regressions.' : ''}`);
|
||
} else if (regressions.length === 0) {
|
||
notes.push('Stable run — no significant efficiency changes, no regressions.');
|
||
}
|
||
}
|
||
|
||
return notes;
|
||
}
|
||
|
||
// --- Budget regression assertion ---
|
||
|
||
export interface BudgetRegression {
|
||
testName: string;
|
||
metric: 'tools' | 'turns';
|
||
before: number;
|
||
after: number;
|
||
ratio: number;
|
||
}
|
||
|
||
/**
|
||
* Compute budget regressions: tests where tool calls or turns grew by more
|
||
* than `ratioCap` between two runs. Pure function — caller decides how to
|
||
* surface the result. Used by test/skill-budget-regression.test.ts and any
|
||
* future ship gate.
|
||
*
|
||
* `ratioCap` defaults to 2.0 (>2× growth is a regression). Override via
|
||
* `GSTACK_BUDGET_RATIO` env var. New tests with no prior data are skipped.
|
||
*/
|
||
export function findBudgetRegressions(
|
||
comparison: ComparisonResult,
|
||
opts?: { ratioCap?: number; minPriorTools?: number; minPriorTurns?: number },
|
||
): BudgetRegression[] {
|
||
const envRatio = Number(process.env.GSTACK_BUDGET_RATIO);
|
||
const cap = opts?.ratioCap ?? (Number.isFinite(envRatio) && envRatio > 0 ? envRatio : 2.0);
|
||
// Floors avoid noise on tiny numbers (1 → 3 tools is 3× but meaningless).
|
||
const minPriorTools = opts?.minPriorTools ?? 5;
|
||
const minPriorTurns = opts?.minPriorTurns ?? 3;
|
||
const out: BudgetRegression[] = [];
|
||
for (const d of comparison.deltas) {
|
||
const beforeTools = Object.values(d.before.tool_summary ?? {}).reduce((a, b) => a + b, 0);
|
||
const afterTools = Object.values(d.after.tool_summary ?? {}).reduce((a, b) => a + b, 0);
|
||
const beforeTurns = d.before.turns_used ?? 0;
|
||
const afterTurns = d.after.turns_used ?? 0;
|
||
if (beforeTools >= minPriorTools && afterTools / beforeTools > cap) {
|
||
out.push({ testName: d.name, metric: 'tools', before: beforeTools, after: afterTools, ratio: afterTools / beforeTools });
|
||
}
|
||
if (beforeTurns >= minPriorTurns && afterTurns / beforeTurns > cap) {
|
||
out.push({ testName: d.name, metric: 'turns', before: beforeTurns, after: afterTurns, ratio: afterTurns / beforeTurns });
|
||
}
|
||
}
|
||
return out;
|
||
}
|
||
|
||
/**
|
||
* Throw if any test in the comparison exceeds the budget cap. Convenience
|
||
* wrapper around findBudgetRegressions for use in test assertions.
|
||
*/
|
||
export function assertNoBudgetRegression(
|
||
comparison: ComparisonResult,
|
||
opts?: { ratioCap?: number; minPriorTools?: number; minPriorTurns?: number },
|
||
): void {
|
||
const regressions = findBudgetRegressions(comparison, opts);
|
||
if (regressions.length === 0) return;
|
||
const cap = opts?.ratioCap ?? (Number(process.env.GSTACK_BUDGET_RATIO) || 2.0);
|
||
const lines = regressions.map(
|
||
r => ` "${r.testName}" ${r.metric}: ${r.before} → ${r.after} (${r.ratio.toFixed(2)}× > ${cap.toFixed(2)}× cap)`,
|
||
);
|
||
throw new Error(
|
||
`Budget regression: ${regressions.length} test(s) exceeded ${cap.toFixed(2)}× prior usage:\n` +
|
||
lines.join('\n') +
|
||
`\n(Override per run: GSTACK_BUDGET_RATIO=<n>. ${comparison.before_file} vs ${comparison.after_file})`,
|
||
);
|
||
}
|
||
|
||
// --- EvalCollector ---
|
||
|
||
function getGitInfo(): { branch: string; sha: string } {
|
||
try {
|
||
const branch = spawnSync('git', ['rev-parse', '--abbrev-ref', 'HEAD'], { stdio: 'pipe', timeout: 5000 });
|
||
const sha = spawnSync('git', ['rev-parse', '--short', 'HEAD'], { stdio: 'pipe', timeout: 5000 });
|
||
return {
|
||
branch: branch.stdout?.toString().trim() || 'unknown',
|
||
sha: sha.stdout?.toString().trim() || 'unknown',
|
||
};
|
||
} catch {
|
||
return { branch: 'unknown', sha: 'unknown' };
|
||
}
|
||
}
|
||
|
||
function getVersion(): string {
|
||
try {
|
||
const pkgPath = path.resolve(__dirname, '..', '..', 'package.json');
|
||
const pkg = JSON.parse(fs.readFileSync(pkgPath, 'utf-8'));
|
||
return pkg.version || 'unknown';
|
||
} catch {
|
||
return 'unknown';
|
||
}
|
||
}
|
||
|
||
export class EvalCollector {
|
||
private tier: 'e2e' | 'llm-judge';
|
||
private tests: EvalTestEntry[] = [];
|
||
private finalized = false;
|
||
private evalDir: string;
|
||
private shard: string | null;
|
||
private createdAt = Date.now();
|
||
|
||
constructor(tier: 'e2e' | 'llm-judge', evalDir?: string) {
|
||
this.tier = tier;
|
||
this.evalDir = evalDir || process.env.GSTACK_EVAL_DIR || defaultEvalDir();
|
||
this.shard = shardSlugOfEvalDir(this.evalDir);
|
||
}
|
||
|
||
addTest(entry: EvalTestEntry): void {
|
||
this.tests.push(entry);
|
||
this.savePartial();
|
||
}
|
||
|
||
/** Write incremental results after each test. Atomic write, non-fatal. */
|
||
savePartial(): void {
|
||
try {
|
||
const git = getGitInfo();
|
||
const version = getVersion();
|
||
const totalCost = this.tests.reduce((s, t) => s + t.cost_usd, 0);
|
||
const totalDuration = this.tests.reduce((s, t) => s + t.duration_ms, 0);
|
||
const passed = this.tests.filter(t => t.passed).length;
|
||
|
||
const partial: EvalResult = {
|
||
schema_version: SCHEMA_VERSION,
|
||
version,
|
||
branch: git.branch,
|
||
git_sha: git.sha,
|
||
timestamp: new Date().toISOString(),
|
||
hostname: os.hostname(),
|
||
tier: this.tier,
|
||
total_tests: this.tests.length,
|
||
passed,
|
||
failed: this.tests.length - passed,
|
||
total_cost_usd: Math.round(totalCost * 100) / 100,
|
||
total_duration_ms: totalDuration,
|
||
tests: this.tests,
|
||
...(this.shard ? { shard: this.shard } : {}),
|
||
_partial: true,
|
||
};
|
||
|
||
fs.mkdirSync(this.evalDir, { recursive: true });
|
||
const partialPath = path.join(this.evalDir, '_partial-e2e.json');
|
||
const tmp = partialPath + '.tmp';
|
||
fs.writeFileSync(tmp, JSON.stringify(partial, null, 2) + '\n');
|
||
fs.renameSync(tmp, partialPath);
|
||
} catch { /* non-fatal — partial saves are best-effort */ }
|
||
}
|
||
|
||
async finalize(): Promise<string> {
|
||
if (this.finalized) return '';
|
||
this.finalized = true;
|
||
|
||
const git = getGitInfo();
|
||
const version = getVersion();
|
||
const timestamp = new Date().toISOString();
|
||
const totalCost = this.tests.reduce((s, t) => s + t.cost_usd, 0);
|
||
const totalDuration = this.tests.reduce((s, t) => s + t.duration_ms, 0);
|
||
const passed = this.tests.filter(t => t.passed).length;
|
||
|
||
const result: EvalResult = {
|
||
schema_version: SCHEMA_VERSION,
|
||
version,
|
||
branch: git.branch,
|
||
git_sha: git.sha,
|
||
timestamp,
|
||
hostname: os.hostname(),
|
||
tier: this.tier,
|
||
total_tests: this.tests.length,
|
||
passed,
|
||
failed: this.tests.length - passed,
|
||
total_cost_usd: Math.round(totalCost * 100) / 100,
|
||
total_duration_ms: totalDuration,
|
||
wall_clock_ms: Date.now() - this.createdAt,
|
||
tests: this.tests,
|
||
...(this.shard ? { shard: this.shard } : {}),
|
||
};
|
||
|
||
// Write eval file
|
||
fs.mkdirSync(this.evalDir, { recursive: true });
|
||
const dateStr = timestamp.replace(/[:.]/g, '').replace('T', '-').slice(0, 15);
|
||
const safeBranch = git.branch.replace(/[^a-zA-Z0-9._-]/g, '-');
|
||
const filename = `${version}-${safeBranch}-${this.tier}-${dateStr}.json`;
|
||
const filepath = path.join(this.evalDir, filename);
|
||
fs.writeFileSync(filepath, JSON.stringify(result, null, 2) + '\n');
|
||
|
||
// Print summary table
|
||
this.printSummary(result, filepath, git);
|
||
|
||
// Auto-compare with previous run
|
||
try {
|
||
const prevFile = findPreviousRun(this.evalDir, this.tier, git.branch, filepath);
|
||
if (prevFile) {
|
||
const prevResult: EvalResult = JSON.parse(fs.readFileSync(prevFile, 'utf-8'));
|
||
const comparison = compareEvalResults(prevResult, result, prevFile, filepath);
|
||
process.stderr.write(formatComparison(comparison) + '\n');
|
||
} else {
|
||
process.stderr.write(
|
||
`\nNO BASELINE: no completed prior ${this.tier} run found in ${this.evalDir}` +
|
||
' (the in-progress accumulator is not a baseline). Nothing compared —' +
|
||
' this run says nothing about regressions.\n',
|
||
);
|
||
}
|
||
} catch (err: any) {
|
||
process.stderr.write(`\nCompare error: ${err.message}\n`);
|
||
}
|
||
|
||
return filepath;
|
||
}
|
||
|
||
private printSummary(result: EvalResult, filepath: string, git: { branch: string; sha: string }): void {
|
||
const lines: string[] = [];
|
||
lines.push('');
|
||
lines.push(`Eval Results — v${result.version} @ ${git.branch} (${git.sha}) — ${this.tier}`);
|
||
lines.push('═'.repeat(70));
|
||
|
||
for (const t of this.tests) {
|
||
const status = t.passed ? ' PASS ' : ' FAIL ';
|
||
const cost = `$${t.cost_usd.toFixed(2)}`;
|
||
const dur = t.duration_ms ? `${Math.round(t.duration_ms / 1000)}s` : '';
|
||
const turns = t.turns_used !== undefined ? `${t.turns_used}t` : '';
|
||
|
||
let detail = '';
|
||
if (t.detection_rate !== undefined) {
|
||
detail = `${t.detection_rate}/${(t.detected_bugs?.length || 0) + (t.missed_bugs?.length || 0)} det`;
|
||
} else if (t.judge_scores) {
|
||
const scores = Object.entries(t.judge_scores).map(([k, v]) => `${k[0]}:${v}`).join(' ');
|
||
detail = scores;
|
||
}
|
||
|
||
const name = t.name.length > 35 ? t.name.slice(0, 32) + '...' : t.name.padEnd(35);
|
||
lines.push(` ${name} ${status} ${cost.padStart(6)} ${turns.padStart(4)} ${dur.padStart(5)} ${detail}`);
|
||
}
|
||
|
||
lines.push('─'.repeat(70));
|
||
const totalCost = `$${result.total_cost_usd.toFixed(2)}`;
|
||
const totalDur = `${Math.round(result.total_duration_ms / 1000)}s`;
|
||
lines.push(` Total: ${result.passed}/${result.total_tests} passed${' '.repeat(20)}${totalCost.padStart(6)} ${totalDur}`);
|
||
lines.push(`Saved: ${filepath}`);
|
||
|
||
process.stderr.write(lines.join('\n') + '\n');
|
||
}
|
||
}
|