mirror of
https://github.com/garrytan/gstack.git
synced 2026-08-21 21:47:32 +02:00
1cab5e11083a37ea0bc62117e9a9c5d05d68785e
4
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
410b4928e7 |
v1.66.0.0 feat: test/evals/CI speedup — 90s truthful free suite, diff-billed evals, required Linux lane (#2593)
* ci: bump CI image Bun 1.3.10 -> 1.3.13
Matches the local toolchain and brings native `bun test --shard=M/N` /
--parallel to CI (needed by the free-test lane and shard runner work).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* ci: stop version bumps rebuilding the eval Docker image (cache key trio)
Three coupled fixes, atomic because any subset is worse than none:
1. Image tag keys on hashFiles(Dockerfile.ci, bun.lock) — package.json is
out: its version field changed on 60/60 recent commits, forcing a ~2min
image rebuild per PR for a dependency set only bun.lock determines.
2. ci-image.yml now pushes that same content-hash tag (previously only
:latest/:sha, so the weekly prebuild never warmed the tag the eval
matrix actually looks up) and both eval workflows get registry layer
cache (cache-to export gated to same-repo runs; fork tokens cannot
write GHCR).
3. Dockerfile bakes /opt/node_modules_cache/.bun.lock and the runtime
Restore-deps guard diffs bun.lock instead of package.json — otherwise
every version-only bump made all 14 matrix jobs fall back to a live
bun install, which is slower than today's behavior.
Worst-case failure mode is self-healing: a missing tag or cache falls
back to exactly the previous rebuild-and-install path.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* ci: stop double-running lint + skill-docs on every PR commit
Both fired on unrestricted push AND pull_request, so each PR push ran
them twice (12 duplicate (headSha, workflow) pairs in the last 200 runs).
push is now main-only; pull_request covers PR branches.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* ci: run actionlint from the prebuilt image (16s -> ~2s)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* ci: right-size five single-core jobs to ubicloud-standard-2
actionlint, skill-docs, version-gate, pr-title-sync, and the evals report
job never exceed one core; standard-8 was ~4x the cost for zero wall-clock.
build-image and the eval matrix keep standard-8.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* ci: fix workflow_dispatch concurrency collisions (head_ref || run_id)
head_ref is empty on workflow_dispatch, so every manual dispatch of these
four workflows shared one empty-suffix group and cancelled each other.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* ci(windows): cache bun installs; run the curated suite, not a hand list
- actions/cache on ~/.bun/install/cache keyed on bun.lock (install was
35-45s of both 55-64s jobs, all network) and Bun pinned to 1.3.13 to
match the other lanes.
- windows-free-tests now runs `bun run test:windows` (the runner's
--windows-only curation) instead of a hand-listed 13-file subset that
had drifted from the registry it sampled. POSIX-bound tests get
excluded in ONE place (the curation patterns), not two.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* evals: retry 1, not 2, on every paid path
Measured on the llm-judge shard: --retry 2 amplified 25 tests into 46
executions (+84%), with retried runs at 138s vs a 10-12s baseline (429
backoff), and a permanently-failing test paying 3x. One retry still
absorbs one-off flakes; chronic flakes become visible fix-work instead
of silent wall-clock.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* evals: split skill-e2e-review into three per-file CI shards
Bun runs describe blocks as concurrency barriers, so the e2e-review CI
job executed its tests serially: 741s of an 860s PR critical path for
tests whose slowest member is 224s. The per-file matrix is the repo's
parallelism unit, so the split moves:
- Retro E2E + retro-base-branch -> test/skill-e2e-retro.test.ts
- review/ship base-branch + Review Dashboard Via Attribution
-> test/skill-e2e-review-attribution.test.ts
- sql-injection / enum-completeness / design-lite stay in
test/skill-e2e-review.test.ts
One 741s job becomes three ~180-250s jobs. Locally the worst paid shard
drops from 1705s (94.7% of the 1800s kill) to under 700s. Test names,
bodies, suite strings, and eval-store collectors are unchanged, so
baselines carry over. Matrix rows added to both eval workflows
(attribution is gate-only, so no periodic row); the report job's
hardcoded runner count is gone (drift-proof).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* test: gate security-bench on SECURITY_BENCH=1, not model-cache existence
The existsSync gate ran ~12s of ONNX inference (plus a HuggingFace
dataset fetch) on every free-suite run on any dev box that had ever
warmed the classifier, while CI (no cache) silently skipped it. Now
explicit opt-in: SECURITY_BENCH=1 bun test browse/test/security-bench.test.ts.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* test: watchdog E2E in 1.5s instead of 22.7s (tunable poll interval)
server.ts gains BROWSE_WATCHDOG_INTERVAL_MS (floor 50ms, default 15s
unchanged). The #994 stay-alive test runs a 250ms tick and waits for the
stay-alive log line instead of blind-sleeping 2s + 20s past the
production interval.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* test: dedupe coverage gates; route both walks through skill-census
skill-coverage-floor duplicated two matrix assertions (registry
completeness, gate-tier floor) with a DIFFERENT hand-rolled directory
walk — matrix's skipped nothing, floor's skipped node_modules/docs/test.
Two 'same' gates disagreeing on the census is the bug class
test/helpers/skill-census.ts was written to kill. Registry assertions
now live in matrix only (with floor's better error message), both files
walk via skillCensus().authoredSkills, and floor keeps the per-skill
structural checks it owns.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* evals: EVALS_JOBS for shard processes; explicit within-shard concurrency
EVALS_CONCURRENCY was overloaded: the legacy bun-test path used it as
--max-concurrency (default 15) while the sharded runner read it as the
process count — exporting the legacy value gave 15 concurrent Bun
processes each spawning claude (the 429 storm). Now: EVALS_JOBS = shard
processes (default 4); EVALS_CONCURRENCY = bun --max-concurrency inside
a shard (default 4, explicit in shard args — omitting it made
within-shard parallelism silently differ from the legacy path). Stale
49/59 header math replaced with the live-count rule.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* evals: enforce detach-timeout floor from the live shard census
New free tripwire: eval:bg:gate / eval:bg:periodic --timeout must cover
ceil(shards/jobs) x shard-timeout x 1.05, recomputed from the actual paid
test census every run. Hand-derived numbers go stale every time a paid
file lands — the review split just proved it: periodic's 28800s dropped
BELOW its new 32130s worst case (raised to 32400s here). An undersized
watchdog kills healthy runs and the tail reports never-started.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* evals: preflight ping once in the sharded parent, not per shard
The Anthropic fail-fast ping ran at module load in every paid test file
importing e2e-helpers — ~30 paid claude -p calls (30s timeout each) per
full sharded run for one bit of information. The parent now pings once
before spawning shards and sets EVALS_PREFLIGHT_OK=1; the module-load
path honors the flag. Extracted to test/helpers/anthropic-preflight.ts
(injectable spawn seam) with regression pins in both directions: the
flag must skip, its absence must ping exactly once, dead API must throw.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* evals: split touchfiles into pure data + selection logic + facade
touchfiles.ts listed ITSELF in GLOBAL_TOUCHFILES, so adding one test's
dep entry forced the full ~$38 / 30-45min suite — measured on 21.9% of
recent commits (42/192). The self-reference existed because data and
logic shared a file: any edit COULD be a selection-logic change.
Now: touchfiles-data.ts (the four maps, literals only, zero imports —
the future map-diff target), test-selection.ts (matchGlob/detectBase
Branch/getChangedFiles/selectTests), and touchfiles.ts as a re-export
facade so all ~12 import sites are untouched. GLOBAL_TOUCHFILES drops
the self-ref, adds test-selection.ts (logic stays maximally
conservative), and TEMPORARILY adds touchfiles-data.ts until the
map-diff change lands. New free test pins the literal-only property
(comment-aware state-machine scan with a self-test) and facade export
parity (===), so neither can silently rot.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* test: free runner — strict output, parallel execution, stable shard indices
Three coupled changes to scripts/test-free-shards.ts:
1. STRICT OUTPUT: runFreeShard streams through the paid runner's
BunTestOutputClassifier — exit 0 without bun's 'Ran N tests across M
files' summary, with (fail) lines, or with a wrong file count is a
FAILURE (anti-truncation backstop at the runner layer), plus an
external wall-clock timeout that SIGKILLs the process group
(timed-out distinct from failed; exit 124 vs 1). Also fixes a latent
shard-bleed: file selectors now use exactTestFileSelectors (relative
paths were substring filters that matched sibling roots).
2. PARALLEL: full-suite mode is one 'bun test --parallel' invocation
(Bun 1.3.13). Measured semantics recorded in the header: per-file
worker isolation, standard summary, and mid-suite process.exit
surfaces as a crashed-worker FAIL with exit 1 — strictly safer than
serial, where the same exit truncates silently. No static weight
lists; --shards M --shard i keeps deterministic hash partitioning for
CI matrices (native --shard rejected: round-robin renumbers when
files land). Spawned shards get throwaway GSTACK_HOME/TMPDIR so
parallel shards can't contend on real state. Per-shard epilogue
prints files/seconds/status every run.
3. Stable indices: assignFilesToShards no longer drops empty shards, so
a shard's index depends only on the file hash and requested count —
an empty CI matrix slot is a fast no-op success, not a renumbering.
package.json 'test' now delegates to the runner (TEST_ROOTS becomes the
single source of truth for roots; slop:diff tail preserved; the runner
inherits the 30s per-test timeout the old glob passed inline).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* ci: Linux free-test lane — ~400 files get CI coverage for the first time
New required, secretless free-tests job: the canonical runner's single
'bun test --parallel' invocation with strict-output classification on
ubicloud-standard-8. The free suite previously ran on NO Linux CI — only
a curated Windows subset ran anywhere — so every 'tests pass' claim
about main rested on contributors running them locally.
Secretless by design (no API keys; fork PRs finally get real test
signal) and pinned by test/free-tests-workflow-wiring.test.ts: canonical
runner invoked, zero secrets.* references, pull_request never
pull_request_target, and matrix-count/--shards agreement if anyone
switches to the sharded fallback.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* evals: map-diff selection — a touchfiles-data edit runs only what changed
Editing the eval dep-list data no longer forces the full ~$38 /
30-45min suite (measured on 21.9% of recent commits). When
touchfiles-data.ts is in the diff, selection now evaluates the BASE
version (git show -> mkdtemp -> spawnSync bun child printing the four
maps as JSON — sync because e2e-helpers selects at module scope) and
JSON-diffs per key: added entries, edited dep lists, and tier flips are
selected; keys removed from all maps are reported, never silently
dropped; a GLOBAL_TOUCHFILES edit still runs everything.
FAIL-CLOSED with named causes: missing-base-ref, git-show-failed,
import-failed, shape-mismatch each degrade to run-all and print
'selection: global — touchfiles-data changed (<cause>)' (D9 — silently
expensive beats silently wrong, but never silently). eval:select prints
'selected N of M, reason: ...' + removed tests; --base scopes the
map-diff too.
The temporary conservative GLOBAL entry for touchfiles-data.ts is gone —
its changes route through the map-diff. 23 new free tests: pure-core
fixtures, selectTests wiring incl. a poison-injection guard, and a temp
git repo exercising every fail-closed cause end-to-end.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* evals: selection sees uncommitted work; git errors fail closed
getChangedFiles is now the deduped union of committed (base...HEAD),
staged+unstaged (git diff HEAD), and untracked (git status --porcelain
--untracked-files=all) — an agent that edits files and runs evals
BEFORE committing no longer gets the full $38 suite every time because
the committed diff looked empty. Clean tree still returns [] (run-all
by design for main-branch/periodic runs).
Git failures now THROW with the failing command, stderr, and 'set
EVALS_ALL=1 to deliberately run the full suite' — the old return []
silently became run-all, which is silently expensive. 11 new free tests
cover every source, dedupe, quoted paths, and both failure shapes via
an injectable spawn seam.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* test: revert GSTACK_HOME injection in the free runner — shared mutable state
The first full run under the strict runner surfaced 12 failures with one
root cause: injecting a single throwaway GSTACK_HOME per invocation made
6,900 tests share a MUTABLE scratch home. gstack-config tests wrote keys
into it; relink and update-check tests then read them (e.g. relink saw
skill_prefix left behind by a config test and produced prefixed names).
All 12 pass when run directly.
TMPDIR isolation stays (mkdtemp inside it is still per-call unique).
Tests needing GSTACK_HOME isolation mkdtemp their own per test — the
repo convention — and hermetic-env covers E2E children. The env-dump pin
now asserts GSTACK_HOME passes through UNTOUCHED so the injection can't
come back.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* test: rebase parity baseline to v1.64.0.0; fix capture-vs-check drift
The parity ratchet had quietly failed for 7 skills — v1.58-v1.64 growth
landed past the v1.57.7.0 anchors and nothing caught it because this
test had no CI lane (verified pre-existing: SKILL.md content is
byte-identical to origin/main). Same rebase protocol as
v1.53->v1.57.7.0; old baseline retained for the audit trail.
Root-caused a second latent bug while rebasing: captureBaseline recorded
SKELETON-ONLY bytes while the checker compares UNION bytes (skeleton +
carved sections/*.md), so a fresh capture read carved skills at ~2x
ratio (ship: 82KB captured vs 183KB checked). captureBaseline now takes
sectionedSkills and records unions for carved skills — capture and check
measure the same thing, so the NEXT rebase can't hit this. Four
CARVE_GUARDS skeleton caps re-ratcheted to current +headroom
(plan-ceo 92K, plan-eng 70K, office-hours 100K, design-consultation
70K), annotated inline.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix: package.json version matches VERSION (1.64.0.0)
v1.64.0.0 shipped with VERSION bumped but package.json left at 1.63.0.0
— the 'package.json version matches VERSION file' test fails on
origin/main today. Nothing caught it because that test had no CI lane
until this branch's free-tests job.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* test: fix variants-retry-after HTTP-date flake (TODOS P2)
toUTCString() truncates to whole seconds, so a +3000ms Retry-After date
could mean an effective wait of ~2001ms — flaking against the 2500ms
assertion floor ~1-2 in 9 runs under suite load. +4000ms puts the
truncation floor at 3001ms with the assertion floor safely below it.
Pulled forward from U4 because the free-tests lane is now a required
check and this flake would randomly block PRs.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* test: skill-fixture helper — extract SKILL.md sections, don't copy files
extractSkillSections (fence-aware H2 scanner, loud-throw on missing
sections with available-heading list), extractSkillBody (drops the
shared generated preamble), extractSkillHead (frontmatter + first 30
lines, for routing fixtures). Pinned section lists per consumer, and
free-tier real-skill pins so a gen-skill-docs heading rename fails the
FREE suite instead of a paid run. skill-fixture.ts joins
GLOBAL_TOUCHFILES (fail-safe polarity: over-select).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* test(evals): review E2E fixtures extract sections — 1871 -> 207 lines
CLAUDE.md's extract-don't-copy rule, applied: the three review fixtures
carry only the sections the sql-injection/enum/design-lite prompts and
judges exercise (89% cut). Full-file copies made claude -p read 1871
lines per test — the direct cause of the 1705s worst shard (94.7% of
the 1800s kill).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* test(evals): retro E2E fixtures extract sections — 1821 -> 757 lines
Keeps every section the retro flow exercises incl. base-branch detect;
drops preamble, Global Retrospective Mode, Compare Mode (58% cut).
retro-base-branch was the single slowest CI test at 224s.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* test(evals): review-army fixture extracts sections — 1871 -> 650 lines
CS1's set plus Step 1.5 (PLAN COMPLETION AUDIT machinery) and Step 4.5
(army dispatch, quality_score, findings schema) that the 7 army tests
assert on. Pin test guards the three load-bearing strings.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* test(evals): skillify fixtures via extractSkillBody — 63-83% smaller
Tests follow all 11 skillify steps, so the whole body stays; only the
shared generated preamble drops (skillify 1239->453, scrape 958->167).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* test(evals): context-skills fixtures via extractSkillBody — 74-82% smaller
context-save 1037->267 lines, context-restore 952->168; the 8 tests
exercise full save/restore/list flows so the body stays, preamble drops.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* test(evals): opus-47 discovery fixtures via extractSkillHead — ~95% smaller
Routing/fanout tests only read frontmatter + opening lines of the 14
installed skills (review 1871->54, office-hours 1706->80).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* test(evals): codex runner gains sections option — review variant 88% smaller
runCodexSkill/installSkillToTempHome accept sections?: string[] routed
through extractSkillSections; codex-review-findings wired (1465->181
lines). codex-discover-skill deliberately keeps the FULL copy — its
stderr assertions validate that the real generated artifact loads.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* test(evals): routing fixture installs skill HEADS, not ~18 full SKILL.md
Routing reads frontmatter only; extractSkillHead per skill (root
611->48, ship 1435->54 lines). This was the single worst fixture bloat
site: one fixture dir holding ~18 full skills.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* evals: parent-side shard skipping — a one-test diff runs 3 of 44 shards
The sharded runner spawned every shard regardless of diff; only the
child self-skipped, so a typical single-skill change still paid 44 Bun
boots + container-equivalent setup for shards with zero selected tests.
The parent now computes selection once (mirroring e2e-helpers exactly:
EVALS_ALL -> run-all, empty union -> run-all, git errors propagate the
fail-closed throw) and drops shards where no selected test name maps in.
Mapping = quoted E2E map keys in the file's source UNION keys whose dep
list registers the file (constructed-name families need the second
direction). FAIL-OPEN everywhere it matters: run-all, non-skill-e2e
files, unreadable source, zero mapped names all keep the shard — the
child filter stays authoritative, so a parent bug can only run extra.
New taxonomy status skipped-by-diff (never conflated with
never-started); selection banner prints once; --list is selection-aware.
C6 lands in the same commit: a HARD tier-alignment test — every paid
skill-e2e file must be parent-mappable or provably fail-open-safe.
Note: this change-set's 14 dep-list registrations in touchfiles-data.ts
rode along in
|
||
|
|
c118e2402e |
v1.64.1.0 v1.64.1.0: the code-smell fix wave — every pipeline guard now provably fires (net −24,943 lines) (#2572)
* fix(ci): skill-docs freshness gate covers all 10 hosts and can actually fail The Codex/Factory gates ran 'git diff --exit-code -- .agents/' / '-- .factory/', but both paths are gitignored (.gitignore:16-17) — git diff on ignored untracked paths is always empty, so those two gates were structurally incapable of failing and 7 of 10 hosts had no gate at all. New shape: one 'gen:skill-docs --host all' pass (the generator hard-fails on any per-host error, gating all 10 hosts on generates-cleanly), byte-freshness via git diff for tracked output, plus a porcelain check that fails on untracked generated strays (git diff can't see brand-new files). The gitignored-hosts byte-freshness limitation is documented in the workflow comment. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(test): exorcise the sidebar-agent ghost from the test suite browse/src/sidebar-agent.ts was deleted in the v1.14 sidebar refactor, but the test suite kept testing it for 48 versions. Nothing noticed because the free suite runs in no CI job and Bun-era module-load errors were suppressed in the Windows shard runner via an exclusion pattern whose own comment documented the breakage ('broken on every platform since v1.14 ... exit 0'). - Delete sidebar-security.test.ts + security-source-contracts.test.ts: crashed at module load (unguarded readFileSync of the deleted file); per-assertion triage confirmed every SERVER_SRC pin targeted the deleted chat prompt builder (zero hits in today's server.ts) — nothing to port. - Delete sidebar-integration.test.ts: 11 of 13 tests exercised deleted endpoints (/sidebar-command queue, /sidebar-agent/event, chat buffer); the 2 passing tests pinned only the blanket auth gate, covered by server-auth.test.ts + dual-listener.test.ts. - Delete test/skill-e2e-sidebar.test.ts: E2E for the deleted queue flow. - sidebar-ux.test.ts 1,669 -> 830 lines: 20 dead-chat describes + 15 dead tests removed (incl. 10 vacuous passes asserting on empty indexOf slices); 2 stale pins on LIVE features fixed (content.js typed-catch CSSOM fallback, arrow-hint window widened). 95 pass / 0 fail. - sidebar-tabs.test.ts: both failures were stale pins, not regressions — forceRestart's deliberate ws.close(4001) and the terminal-agent spawn that moved into spawnTerminalAgent() (identity-based kill refactor). 28 pass. - touchfiles.ts: drop the three sidebar E2E entries from BOTH maps (E2E_TOUCHFILES + E2E_TIERS) — they pointed diff-selection at the deleted file, so those tests were unreachable by any diff. - test-free-shards.ts: remove the now-dead sidebar-agent exclusion pattern. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat(ci): run the free test suite in CI (it ran nowhere) The full free suite (bun test: browse/test/ + test/ + make-pdf/test/) had no CI job on any Linux/macOS runner — only Windows curated shards, paid evals, and doc-freshness gates existed. That's how two module-load-crashing test files survived 48 versions. Same cached Dockerfile.ci image and container wiring as evals.yml (deps restore, build, Chromium verify). Includes a module-load-error guard: older Bun reported test-file import crashes with exit 0 on macOS/Linux, so the job also fails on any nonzero 'N errors' count in the summary — future crash-class regressions can't hide from the exact job built to catch them. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat(test): validate touchfile dependency paths exist on disk New guard in touchfiles.test.ts: every non-glob dep path must exist, and every glob's anchor directory must exist. This is the axis the 181-key two-map sync discipline never covered — an entry can point at a long-deleted file and diff-based selection then silently never triggers those tests (the sidebar trio sat rotted for 48 versions). First run immediately caught a fourth rotted entry: 'spec authored quality' referenced test/fixtures/spec/** (directory does not exist) and selected for a judge test that exists nowhere in the repo. Removed. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(security): remove deleted /sidebar-chat endpoint from tunnel allowlist TUNNEL_PATHS is the audited tunnel attack surface — its own comment says every addition widens it. '/sidebar-chat' stayed in the set after the endpoint was deleted with the chat-queue path, meaning any future route matching that path would have been silently tunnel-exposed. The set is now exactly the pair ceremony (/connect) and the scoped command endpoint (/command), and the dual-listener closed-set pin enforces that. Also repairs a pre-existing red pin in dual-listener.test.ts: v1.63.0.0 made the tunnel allowlist args-aware (canDispatchOverTunnel gained a second param) without updating the test — red on main since then, invisible because the free suite had no CI job. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(security): delete chain's shadow dispatcher that skipped every security gate meta-commands.ts carried a 'CLI mode' fallback that re-implemented command routing without the server pipeline's gates: no scope check, no domain check, no tab ownership, no rate limit, no hidden-element stripping, no scoped-token enveloping — and it called handleReadCommand without a BrowserManager, which also skipped the JS-origin cookie-exfiltration assertion. It was unreachable in production (server.ts always passes executeCommand) and one boolean away from being live. chain now hard-errors without a server context. handleReadCommand's bm param is required and assertJsOriginAllowed runs unconditionally. The chain tests that exercised the deleted fallback now route through a server-shaped executeCommand adapter (real handlers + trust wrapping + {status,result} envelope), so their behavioral coverage — sequencing, trust markers, pipe format, aliases, error reporting — survives on the production-shaped path. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(extension): delete the dead chat-queue client surface The sidebar-command handler in background.js POSTed to a server endpoint that no longer exists (deleted with the chat queue) — ~35 lines of fully-wired dead code including error handling for the permanent 404, plus its allowlist entry. No sender in the extension ever emitted the message type. chatEnabled leaves the /health contract (server hardcoded false, background.js re-derived it, nothing consumed it — the chat input element it guarded is gone from sidepanel.html). BROWSE_SIDEBAR_CHAT env flag had zero readers. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(security): delete dead exports the ripped chat path left behind Three-way split by importer class: (a) Zero importers, deleted: the whole attack-attempt logging cluster in security.ts (logAttempt, AttemptRecord, salted hashPayload + device-salt, attempts.jsonl rotation, telemetry spawn plumbing incl. buildTelemetrySpawnCommand/resolveBashBinary — the LIVE attempts.jsonl writer is tunnel-denial-log.ts with its own rotation); the decision-file handshake (writeDecision/readDecision/clearDecision/excerptForReview — written for sidebar-agent's poll loop, which no longer exists); sidebar-utils.ts (whole module — its sanitizeExtensionUrl 'sanitized before embedding in a prompt' for the deleted prompt builder); 8 dead server.ts imports (sanitizeExtensionUrl, generateCanary, injectCanary, writeDecision, rotateRoot, serializeRegistry, restoreRegistry, clearAgentRecord); buildPtyClearCookie + buildSseClearCookie; WEBDRIVER_MASK_SCRIPT (orphaned by the D7 stealth narrowing — applyStealth never used it). (b) Dead-pin tests edited with their exports: the 'still exported' pin in stealth-layer-c, the string-content describe in stealth-webdriver (its live applyStealth behavioral coverage untouched), the clear-cookie assertions, security-review-flow.test.ts deleted whole (all 4 describes exercised the dead decision mechanism, incl. a 'simulated sidebar-agent poll loop'). (c) KEPT deliberately: leaseCount (live behavioral coverage), extractPtyCookie + validatePtySessionToken (extractPtyCookie is adopted by the terminal-agent cookie-parse unification later in this wave), resetSessionMarker + clearContentFilters (test-support API for the live content-security layer). Also fixes two pre-existing red pins found while here, invisible until the free suite got a CI job: the v1.44 spawnClaude->maybeSpawnPty rename in terminal-agent.test.ts, and a cross-file test-isolation bug where content-security.test.ts's clearContentFilters() wiped the auto-registered url-blocklist filter for every later file in the same bun process (security-integration.test.ts failed on co-run; afterAll now restores it). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(security): delete the dead ML layers — transcript classifier and DeBERTa ensemble The L4b Haiku transcript classifier and the opt-in DeBERTa ensemble (GSTACK_SECURITY_ENSEMBLE=deberta, a documented 721MB download) had ZERO production callers since the chat-path agent that invoked them was ripped. The only live ML path is scanPageContent (testsavant) inside the security sidecar subprocess. Deleted by import graph: - security-classifier.ts 614 -> 265 lines: HAIKU_MODEL, checkTranscript, shouldRunTranscriptCheck, loadDeberta, scanPageContentDeberta, ToolCallInput, all DEBERTA_* consts + load state. Header now states the live truth (imported only by security-sidecar-entry.ts). downloadFile kept, name intact — it is an enumerated egress sink (HF model download). - security-bunnative.ts + test: a research skeleton self-described as 'NOT a production replacement', shipped into src/ with zero importers. - security-bench-ensemble{,-live}.test.ts + the Haiku response fixture: a paid live-model benchmark for a layer that could not fire. The security-classifier-tdz test's only case exercised checkTranscript — gone. - security.ts: layer-model header rewritten to the live architecture; StatusDetail.layers -> {testsavant, canary}; getStatus() no longer requires the impossible transcript==='ok' for 'protected' (old on-disk session state with a transcript key is tolerated on read, never re-emitted). - security-sidecar-entry.ts needed zero changes: it serializes getClassifierStatus() verbatim and no consumer read .transcript (verified in sidecar-client + server.ts). - BROWSER.md security section matches reality (ensemble knob gone, 112MB not 22MB, sidecar hosting documented). combineVerdict/THRESHOLDS retained as the pure, tested combiner of record — comments now flag transcript/deberta votes as producer-less. Net: 26 pass in security.test.ts incl. a NEW regression test for stale- transcript disk tolerance; egress-receipt tripwire green. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs: scrub the sidebar-agent ghost from comments and CLAUDE.md 20+ comments across 10 files still described the deleted sidebar-agent.ts as a live process — including load-bearing architecture claims ('IMPORTED ONLY BY sidebar-agent.ts', 'sidebar-agent fills this in on first prompt-injection load', 'kill sidebar-agent' in shutdown docs) and ~60 lines of tombstone blocks in server.ts enumerating deleted identifiers by name (a false grep surface: searching processAgentEvent hit server.ts and looked live). CLAUDE.md's security-stack section now documents the LIVE architecture: L1-L3 content filters + testsavant via the security sidecar subprocess; the L4b/ensemble rows, the GSTACK_SECURITY_ENSEMBLE knob, and the 721MB DeBERTa download are gone (deleted as dead code this wave) with an explicit do-not-re-document note; attempts.jsonl is correctly attributed to tunnel-denial-log.ts; the no-live-writer status of classifierStatus is stated. Comments that survive now describe what IS, not what WAS: the promotion gate in domain-skills.ts explains why classifier_score>0 is load-bearing given no L4 load-time scan exists; file-permissions.ts names real sensitive files. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(gen): delete the codex-helpers shadow module gen-skill-docs.ts imported externalSkillName (unaliased) from resolvers/codex-helpers.ts at line 21 and then re-declared the same function locally — the import was silently shadowed, and the imported copy was the STALE one (it lacked the frontmatterName param the local copy grew). Three more functions were byte-identical duplicates, imported only under _-prefixed aliases to keep the module 'referenced', and transformFrontmatter was a superseded hardcoded-Codex variant. Nothing else imported the module. Also drops three dead top-of-file imports (COMMAND_DESCRIPTIONS, SNAPSHOT_FLAGS — which pulled the whole browse/src module graph into every generator run for nothing — and an unused review-resolver trio). Proof: bun run gen:skill-docs exits 0 with a byte-identical tree (zero-diff regen); gen-skill-docs.test.ts 405/405 green. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(server): delete ServerConfig.idleTimeoutMs + chromiumProfile — documented, never read Both fields carried JSDoc asserting embedder behavior that did not exist: the idle check reads the module-level IDLE_TIMEOUT_MS env constant, and both resolveChromiumProfile() call sites pass no argument. Worse than absent — an embedder passing idleTimeoutMs: 5000 silently got 30 minutes. Wiring them honestly is impossible today: the idle timer, activity state, and shutdown target are module-global, so a per-factory value would lie for any process running more than one handler. Deleted instead, with a ServerConfig note pointing at the deferred singleton/route-table refactor where real support belongs. BROWSE_IDLE_TIMEOUT and CHROMIUM_PROFILE env remain the honest knobs. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(security): wire appendSecureFile at the four real log-append sites file-permissions.ts carries a 24-line rationale for why POSIX mode bits are insufficient on Windows and implements appendSecureFile (0600 at create, Windows ACL on first write only) — but its single caller was the dead logAttempt, while the four REAL page-content log writers (console/network/ dialog logs in server.ts, the command audit log) used raw fs.appendFileSync with no mode. Page-content-derived logs now get owner-only permissions from birth on every platform. Verified before wiring: mode applies atomically at create via appendFileSync {mode}, and the ACL pass runs only on first write — no per-append subprocess cost on the hot console-log path. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(stealth): handoff() uses the shared profile resolution + lock cleanup The headless-to-headed handoff path hardcoded ~/.gstack/chromium-profile, silently ignoring $CHROMIUM_PROFILE and $GSTACK_HOME (gbrowser's gbd sets per-workspace profiles), and skipped cleanSingletonLocks() — so a handoff into a profile with a stale SingletonLock could hang where launchHeaded() would have recovered. This was the third live drift between the three Chromium launch paths; the first two are documented in comments as shipped stealth regressions. Minimal targeted fix — the full buildLaunchConfig() extraction stays in the deferred queue. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(gen): resolver registry describes the template language again Seven registered {{PLACEHOLDER}}s had zero uses in any .tmpl (checked in both bare and :arg forms): REDACT_TAXONOMY_TABLE, TEST_COVERAGE_AUDIT_REVIEW, MODEL_OVERLAY, QUESTION_PREFERENCE_CHECK, QUESTION_LOG, INLINE_TUNE_FEEDBACK, MAKE_PDF_SETUP. The last two of those families are invoked programmatically by preamble.ts (functions kept, registry entries dropped); the question-tuning trio and the review coverage-audit wrapper were documented by their own module as existing 'for unit testing' that no test performed — deleted, along with generateRedactTaxonomyTable + its EXAMPLE/TIER_BLURB constants (its '/cso renders the full table' comment was itself stale) and its test describe. Also deletes the gated-resolver mechanism (ResolverEntry/appliesTo/ unwrapResolver + test/resolver-entry.test.ts): fully built, fully tested, used by zero of the 65 registry entries — the generator loop simplifies to a direct function call. CLAUDE.md's redact-doc line stops advertising the dead token. Proof: zero-diff regen (0 SKILL.md changed); gen-skill-docs + skill-validation 737 tests green. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(gen): wire boundaryInstruction from host config; drop three no-op binDir ternaries hosts/codex.ts declared boundaryInstruction and nothing read it — review.ts kept its own byte-identical CODEX_BOUNDARY literal (verified equal + trailing escaped newlines). The resolver now reads the config, so the boundary has one owner. (autoplan's template carries deliberately generic variants, enforced by gen-skill-docs.test.ts:1358 — untouched by design.) The 'ctx.host === codex ? $GSTACK_BIN : ctx.paths.binDir' ternary appeared in three resolvers and could never change the result: resolvers/types.ts already sets binDir to $GSTACK_BIN for every usesEnvVars host including codex. Proof: zero-diff regen for claude AND codex hosts; gen-skill-docs + host-config suites green. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(test-infra): judge uses resolveClaudeBinary; eval:watch reads the real partials dir judgePtyState spawned the bare string 'claude' three definitions below the resolveClaudeBinary() helper this same file exports — broken under hermetic PATHs where every other launch in the file resolves correctly. eval:watch read _partial-e2e.json from the legacy global ~/.gstack-dev/evals/ while EvalCollector writes it into the per-project eval dir (or GSTACK_EVAL_DIR) — so the dashboard's completed-tests panel was empty whenever slug detection succeeded, i.e. the normal case. The heartbeat and per-run progress logs stay global by design (session-runner.ts: 'heartbeat stays global'). The three eval-CLI docstrings stop claiming the legacy dir is the primary location. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(test): delete the superseded SDK ship-idempotency suite and three orphaned fixtures test/skill-e2e-ship-idempotency.test.ts's own header documented that the monolith's SDK-harness version tests a synthetic prompt while it exercises the real /ship skill — the author knew the old suite was superseded and left both running, two paid LLM runs for one behavior. The weaker copy is gone; its 'ship-idempotency' diff-selection key goes with it (the dedicated file is periodic-tier, which always runs under EVALS_ALL — the key had no remaining consumer). Fixture rot: test/fixtures/golden-ship-claude.md was a 128KB zero-reader orphan that had drifted 46KB from its live successor (test/fixtures/golden/claude-ship-SKILL.md) while looking authoritative; parity-baseline-v1.46.0.0.json and v1.53.0.0.json had zero readers (three tests pin three OTHER baseline versions — consolidation is queued, deletion of the unreferenced two is free). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(bin): delete zero-caller scripts; make host-config-export's docstring honest - bin/gstack-open-url (14 lines): announced in a CHANGELOG entry, wired into nothing, ever. bin/gstack-platform-detect (27 lines): zero callers, and its hand-rolled host list was already stale (SLATE_HOST.md cites it as a problem). Note: the deprecated gstack-brain-consumer/reader pair the audit flagged was already deleted upstream in v1.63 with a stay-deleted tripwire. - scripts/task-emission-schema.ts (61 lines): a typed schema module nothing imported; the tasks-section comment now documents the JSONL fields inline. - scripts/host-config-export.ts claimed to be the 'shell bridge for the bash setup script' — setup never calls it (its hand-rolled host lists drifting is a known follow-up). Docstring now states what it IS: a standalone, test-pinned query CLI not yet wired into setup. Its validateValue + CLI_REGEX/PATH_REGEX internals were dead (defined for a guarantee the header claimed but nothing enforced). - KEPT deliberately: scripts/preflight-agent-sdk.ts — a documented manual diagnostic (CONTRIBUTING.md + USING_GBRAIN_WITH_GSTACK.md reference it). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(server): one lone-surrogate sanitizer, one sanitizeReplacer, one startTunnel Three copies of the surrogate sanitizer existed with two algorithms (sanitize.ts regex vs a hand-rolled charCodeAt walk in server.ts — verified byte-identical across 11 edge cases before converging) plus two identical sanitizeReplacer definitions each wrapping a different copy. sanitize.ts is now the single source of truth; the runs-INSIDE-JSON.stringify egress invariant is unchanged at every call site and its pin tests were adapted to the new import shape without losing intent. The ngrok tunnel-start sequence existed three times in server.ts — the /tunnel/start route and the BROWSE_TUNNEL=1 autostart were line-for-line equivalent (a comment admitted 'Same cleanup as /tunnel/start's error path'). One startTunnel() now owns the ephemeral loopback bind, the pre-send egress receipt, the state-file RMW via tmpStatePath(), and the ordered error-path cleanup; callers keep their distinct response surfaces. The BROWSE_TUNNEL_LOCAL_ONLY test path shares nothing (no ngrok, different state field) and deliberately stays separate. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(security): one session-cookie registry implementation, two instances pty-session-cookie.ts and sse-session-cookie.ts were byte-identical modulo the cookie name — mint/validate/parse/prune/TTL, the exact code a security fix would have to land in twice (and a third hand-rolled cookie parse in terminal-agent.ts had already diverged; unified next commit). createSessionCookieStore() owns the implementation; both modules become thin instantiations keeping every exported name, their distinct threat-model docstrings, and separate token spaces (an SSE-read cookie must never grant PTY access). pty-session-lease.ts deliberately stays out — different contract (sessionId/secret split, refresh, env TTL). The factory imports nothing from token-registry (cookie-picker-auth-isolation invariant, still pinned by sse-session-cookie.test.ts). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(security): terminal-agent uses the shared PTY cookie parser The /ws upgrade's cookie fallback hand-parsed the Cookie header inline — the fourth copy of the session-cookie parse, and the one that had already diverged from the others. Parsing now goes through extractPtyCookie; validation deliberately stays against the agent's own in-process validTokens map (the server's registry lives in a different process). The ws-handler pin test now pins the shared-parser call instead of the raw cookie-name literal. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * refactor(hosts): defineHost() factory — 10 copy-paste host files become declarations hosts/*.ts were ten copies of one file: runtimeRoot byte-identical in 9/10, pathRewrites mechanically derivable from the host name for 7/10, the 11-entry toolRewrites map byte-identical between openclaw and gbrain, and every asset change a 10-file edit (cursor and slate had already fallen out of three other hand-maintained lists). defineHost() owns the defaults; each host file now declares only what makes it different (slate/cursor: 8 lines each). Shared constants: CROSS_MODEL_RESOLVERS, GBRAIN_RESOLVERS, EXEC_STYLE_TOOL_REWRITES. Genuinely-different things stayed explicit: codex/factory $GSTACK_ROOT rewrites, hermes's tool vocabulary, claude's denylist+prefixable install, opencode's wider runtimeRoot. Proof: JSON.stringify(ALL_HOST_CONFIGS) dump-diff before/after EMPTY (and a runtime walk confirmed no function-valued or undefined-keyed fields, so the JSON diff is complete); gen:skill-docs --host all zero-diff; host-config + gen-skill-docs + idempotency suites 485/485. Host files 595 -> 285 lines. docs/ADDING_A_HOST.md teaches the factory pattern. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat(lib): fs-atomic — one atomic-write implementation, with the race actually fixed Atomic tmp-write-then-rename was reimplemented ~20 times across lib/, bin/, and browse/src with three tmp-suffix conventions. One of them was a latent bug this commit closes: lib/worktree.ts used a bare '.tmp' suffix — the deterministic-tmp collision race browse/src/server.ts documents having hit in production (its fix, pid+random, was trapped in a comment at one site). lib/fs-atomic.ts: atomicWriteSync (always throws, best-effort tmp cleanup, pid+random suffix, optional mode applied at tmp creation so the file never exists with looser permissions) + atomicWriteQuiet (shutdown paths only). Unit tests pin the throw/quiet contracts, 0600 mode, tmp-name uniqueness (captured via the read-only-dir failure path — Bun's fs exports are readonly, no monkeypatching), and no-stray-tmp cleanup. Migrated: lib/worktree.ts (the bare-.tmp bug), lib/gstack-decision.ts (snapshot + compact log), lib/gbrain-local-status.ts (probe cache). browse sites follow separately. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(lib): jsonl-store's docstring stops lying; mode option added; lib bypasses adopted The header claimed 'single source of truth... the ONLY copy' with write-time injection REJECTION — while appendJsonl never screened anything, only 1 of ~10 JSONL stores imported it, and a bypass appender lived in the same directory. Now: the contract is explicit (screening is the CALLER's job via hasInjection/firstInjectionMatch; the enforcing callers are named), a option applies 0600 at create for sensitive stores, and the lib bypasses are adopted (gstack-memory-helpers ×2, redact-audit-log — which keeps its chmod backstop for files created looser by pre-mode versions). browse/src keeps its own appenders by design (compiled-binary surface, own secure-append helper) and the header now says so. gstack-decision's batched archive append stays deliberate (single-write crash-window semantics appendJsonl's one-record contract can't express). New pins: 0600-at-create, and a test that documents appendJsonl does NOT self-screen — so nobody can re-document it as self-screening without making it true. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(browse): migrate hand-rolled atomic writes to lib/fs-atomic Seven sites, each audited for its existing throw-vs-swallow contract before migrating: writeSessionState + the four fire-and-forget tab/state writers use atomicWriteQuiet (they swallowed before); writeAgentRecord + the boot-time port-file write use atomicWriteSync (they threw before — and writeAgentRecord previously leaked its tmp file on rename failure, which the helper cleans). All carry {mode: 0o600} plus restrictFilePermissions after successful writes, preserving the Windows ACL hardening that writeSecureFile provided (mode bits are POSIX-only). server.ts untouched: its three state writes route through tmpStatePath(), pinned by server-tmp-state-path.test.ts. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(hosts): delete five dead HostConfig fields metadataFormat (generator hardcodes openai.yaml), sidecar (behavior lives in setup's create_agents_sidecar — knowledge preserved as a comment in codex.ts), install.prefixable (skill_prefix is implemented entirely in bin/gstack-config), staticFiles (docstring cited a SOUL.md that never existed anywhere), and adapter (its only would-be consumer, openclaw-adapter.ts, was fully dead — with a test asserting the field was undefined). Kept: learningsMode (wired next), linkingStrategy (validation reads it), coAuthorTrailer (consumed by resolvers/utility.ts). Proof: JSON dump diff shows ONLY the deleted keys vanishing; zero-diff regen across all 10 hosts; host-config + gen-skill-docs suites green. Note: this commit also carries chunk-23 edits to the shared hosts/claude.ts + define-host.ts + host-config.test.ts files (skipSkills collapse, stale line-number comment drops) — pathspec commits, concurrent prep. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(gen): preamble tiers are explicit; silent ?? 4 default becomes an error; spec stops rendering its preamble twice Eight skills (scrape, diagram, spec, skillify, pair-agent, landing-report, open-gstack-browser + its connect-chrome symlink) silently received the HEAVIEST tier-4 preamble because a missing frontmatter field defaulted to 4. Tiers are now declared in every {{PREAMBLE}} template's frontmatter and a missing declaration throws at generation time with the template path (the 5 templates without {{PREAMBLE}} never invoke the resolver). The stale hand-written tier-map comment (wrong in 3 of 4 rows) is gone. Bonus bug fixed: spec/SKILL.md.tmpl mentioned {{PREAMBLE}} in prose, so the generator inlined the ENTIRE preamble a second time — spec/SKILL.md shrinks 127,462 -> 80,924 bytes (-46,538) from de-duplication alone. skill-size-budget gains a reasoned INTENTIONAL_SHRINKS entry (its frozen baseline had measured the doubled-preamble bug). New tests: missing-tier throw carries the path; every {{PREAMBLE}} template declares a tier. (Carries chunk-23 edits in the shared test/gen-skill-docs.test.ts.) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(gen): learningsMode is read from host config, not a hardcoded host name resolvers/learnings.ts branched on ctx.host === 'codex' while every host declared learningsMode — the field was decorative, and the 7 hosts configured 'basic' (cursor, slate, kiro, opencode, openclaw, hermes, gbrain) silently received the 'full' cross-project flow their runtimes can't execute (it depends on AskUserQuestion + gstack-config plumbing). Output now matches declaration: basic hosts get the project-scoped search block. Blast radius proof: all committed Claude SKILL.md files and the three golden fixtures are byte-identical; the behavior diff lands only in the gitignored external-host trees (hand-verified: .cursor review's learnings section swaps the cross-project AskUserQuestion block for the project-scoped search). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(gen): small config scrubs — openclaw blobs to real files, setup host drift, dead artifacts - The three openclaw markdown blobs hardcoded inside gen-skill-docs.ts (which silently reverted any hand edit to their tracked outputs on regen) move to openclaw/templates/*.md source files; output shasums byte-identical. - setup's --host allowlists gain cursor + slate — both fully registered hosts with generated output, but './setup --host cursor' exited 1 because two hand-rolled lists in setup had drifted from hosts/index.ts. - scripts/proactive-suggestions.json deleted: 31KB regenerated on every run, read by nobody (the catalog-trim design's reader was never built); its emitter and three determinism tests (which guaranteed a file nothing reads didn't churn) retired with stays-retired pins. - claude/SKILL.md.tmpl deleted: a complete 8.9KB skill that never generated output (directory name collides with the host id 'claude'), in no registry. Recoverable from git if ever wanted under a non-colliding name. - openclaw's frozen extraFields.version '0.15.2.0' stamp dropped; includeSkills: [] no-ops omitted (the generator treats [] as absent); llms.txt 55 -> 54 skills. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat(gen): correct preamble tiers for the 8 silently-heaviest skills With tiers now explicit, set them RIGHT by analogy to the tiered population: scrape/diagram/open-gstack-browser (+ the connect-chrome symlink) -> tier 1 (launchers and artifact generators, like browse and make-pdf); landing-report/pair-agent/skillify -> tier 2 (dashboards and session tools, like health and canary); spec -> tier 3 (interactive planning, like the plan-*-review family). Each tier-1 skill sheds 271 lines of onboarding prose it never needed; tier-2 shed 20 each. Verification per the review protocol: regen diff reviewed (pure section-removal), skill-validation + size-budget + catalog-budget + v0-dormancy suites green (822 tests), and live smoke of the tier-corrected skills confirms the preamble renders the intended sections at each tier. These skills have ~no eval coverage — stated honestly; the wave's gate-tier eval run is the backstop. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat(test): e2e-gate — one tier-gate implementation, side-effect-free, with the trap pinned The EVALS/EVALS_TIER gate was copy-pasted into ~40 test files and had drifted into six different predicates — the drift that made 'eval:bg:all runs everything' silently false. test/helpers/e2e-gate.ts owns the semantics now: describeE2ETier(tier) + e2eTierEnabled(tier), env read at call time, zero side effects (the existing e2e-helpers module runs a ~30s claude ping at import under EVALS=1, so the gate lives in its own module; purity is pinned by tests that scan imports and comment-stripped source). The unit matrix pins all four env combos — including EVALS=1 with EVALS_TIER unset -> SKIP, the exact trap that made eval:bg:all a non-run. The tier-alignment tripwire gains a second regex for the helper shape (old shape still detected — stragglers can't hide), and the sharded paid runner's PRE-SPAWN tier classifier learns the helper shape too: without that, every gate-sharded run would have spawned all 28 periodic shards just to skip them, each paying the e2e-helpers import ping (~15 min of dead wall clock in the CI-blocking lane). Verified: gate runs exclude the 29 periodic files, periodic excludes the 8 gate files — identical to pre-migration. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * refactor(test): migrate the 36 tier-gated eval files to describeE2ETier Mechanical two-liner swap in 34 files (each keeping its declared tier — all 36 predicates verified against E2E_TIERS before migrating); the two files with compound gates (overlay-harness's EvalCollector feed, codex-e2e's CODEX_AVAILABLE) keep their extra conditions via e2eTierEnabled. Tier rationale comments preserved. codex-e2e/gemini-e2e/benchmark-providers keep their distinct stderr-message gate shapes by design. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * refactor(test): skill-e2e + skill-llm-eval adopt the shared selection machinery Both files re-implemented the diff-selection machinery e2e-helpers already exported. The helper gained computeDiffSelection() (extracted, identical behavior) and a trailing optional selection param on the *IfSelected helpers (defaults preserve all 30+ existing importers). skill-e2e.test.ts drops ~120 duplicated lines; skill-llm-eval keeps its LLM_JUDGE_TOUCHFILES selection and test.concurrent semantics via testConcurrentIfSelected. Deliberate deltas, stated: skill-e2e.test.ts now honors the EVALS_TIER intersection its local copy lacked (affects only direct bun test invocations of that file — it matches no eval-script glob); its recordE2E gains the helper's three diagnostic fields; skill-llm-eval sharded solo now runs e2e-helpers' module-scope preflight it already ran in combined processes. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(test): kill the silent-truncation race; exempt the tier-corrected shrinks The full-suite shakeout (budgeted by the plan) surfaced both immediately: 1. server-embedder-terminal-port.test.ts stubbed process.exit and restored the REAL exit in its finally — but shutdown() schedules async work that can call process.exit AFTER restoration, killing the entire bun process mid-suite with exit 0 and NO summary. This is the silent-truncation class the new free-suite CI job guards against, reproduced locally on the first full run. Exit now stays a logging no-op between tests (late async exits become visible stderr lines, not process death); the true exit returns in afterAll. 2. The 80%-of-baseline shrink guard correctly flagged the six tier-corrected skills — their baseline was measured at the silent tier-4 default. Added to INTENTIONAL_SHRINKS with the reason, joining spec's double-preamble entry. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * release: v1.64.0.0 — the code-smell fix wave 35 commits, one PR: guard repairs (free suite in CI per-file, all-host freshness gates, tunnel allowlist, diff-selection validation), the sidebar-agent ghost exorcism (dead ML layers, dead endpoints, dead exports, ghost comments), config honesty (defineHost factory, dead fields deleted, preamble tiers explicit, spec double-render fixed), and dedup with safety nets (session-cookie factory, fs-atomic, jsonl-store contract, one eval tier-gate). Net -24,943 lines across 183 files. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(ci): free-tests step runs under bash (container sh rejects pipefail) Maiden-voyage shakeout, exactly as budgeted: the CI container's default shell is dash, which errors on 'set -o pipefail' before the first test ran. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(ci): free-tests curates 8 container-incompatible files with reasons Second maiden-voyage shakeout round: 376 of 384 files ran green in the container on the first completed pass. The 8 that can't run there yet are excluded the same way the Windows shards curate POSIX-bound files — each with its reason inline (headed-Chrome handoff, real-PTY round-trip, X server management, extension-origin identity, the job's own TMPDIR override, and three pre-existing env failures that fail on dev machines too). Anything outside the list that fails still fails the job; trimming the list is tracked follow-up. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(test): gstack-config-key-locale — suppress the skill_prefix auto-relink side effect The test invokes the repo's own bin/gstack-config, whose 'set skill_prefix' auto-runs $(dirname $0)/gstack-relink — resolving the install dir to the repo itself. In any environment where the loop shares a working tree (the free-tests CI container, a fresh-HOME run), gstack-patch-names rewrote all 52 tracked SKILL.md names to gstack- prefixed, poisoning five unrelated suites downstream (hermetic-skills-seeding, host-config golden, skill-census, skill-validation, spec-template-sync). GSTACK_SETUP_RUNNING=1 is the documented suppression; relink behavior stays covered by relink.test.ts's mock install. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(bin): gstack-codex-session-import — empty sessions dir exits 0 on Linux GNU xargs runs 'ls -t' once even on empty input, listing the cwd and producing a bogus LATEST from the repo root; BSD xargs (macOS) skips the run, which is why the NO_SESSIONS path only broke on Linux. xargs -r pins the BSD behavior on both platforms. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * test(parity): rebaseline v1.57.7.0 → v1.64.1.0 + skeleton-cap headroom The two parallel v1.64 waves (code-smell fix wave + main's #2571) each added shared-preamble prose, pushing document-release / design-consultation / cso past their size ratios on the v1.57.7.0 anchor and four carved skeletons (plan-ceo-review, plan-eng-review, office-hours, design-consultation) 22-280 B over their absolute caps. New baseline is union-normalized (skeleton + sections/*.md, matching what the harness measures); caps get +~1 KB headroom each with per-cap rationale. The v1.57.7.0 fixture stays in test/fixtures/ for the audit trail, and capture-parity-baseline.ts now documents the union-normalization step so the next rebaseline doesn't re-trip on it. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(ci): free-tests container parity — tools, pinned bun, git identity, mutation tripwire - Dockerfile.ci: add python3 (gstack-jsonl-merge/brain-sync/detach shell out to it), file (skill-validation's binary check), poppler-utils (make-pdf e2e gates hard-require pdftotext/pdffonts/pdfinfo), fonts-noto-color-emoji (emoji render gate, mirrors make-pdf-gate.yml). Fix the bun pin: the bun.sh installer ignores a BUN_VERSION env var, so the old form silently installed latest on every rebuild (observed 1.3.13/1.3.14 drift vs the 1.3.10 devs run locally); pass the version as the positional arg. - free-tests.yml: git identity + safe.directory for the git-exercising tests (container checkout is owned by a different uid than runner); post-loop tree-mutation tripwire that names a tracked-file-mutating test instead of letting downstream collateral confuse the report; skip the documented variants-retry-after timing flake. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(bin): gstack-session-update — detached updater owns its stdio (SIGPIPE) The backgrounded update subshell inherited the session hook's stdout/stderr pipes. Once the hook exits and the caller closes them, any child that writes — git pull's autostash notice, setup output — dies of SIGPIPE, logged as PULL_FAILED exit=141 with an empty stderr capture (observed in the free-tests container, and reachable by any production hook runner that closes stdio promptly). Redirect the fork to /dev/null; all observability already flows through the session-update log file. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(test): gstack-decision-bins — explicit branch context for the scope filter CI checks out a detached HEAD, where gitBranch() returns undefined on both the log and search sides, so an implicitly branch-scoped decision can never surface (filterByScope requires a matching non-empty ctx.branch). Pass the branch explicitly on both sides — the filter logic is what's under test, not git branch detection. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(test): ring-buffer lease interplay — same TTL window, not same millisecond Two back-to-back mintLease() calls each stamp Date.now() + TTL; when they straddle a millisecond boundary the exact-equality assertion flakes (observed in CI: expiries of ...525 vs ...526). Assert the expiries are within a 50 ms window instead — the invariant under test is that leases share a TTL policy, not that they mint in the same clock tick. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
d75402bbd2 |
v1.6.4.0: cut Haiku classifier FP from 44% to 23%, gate now enforced (#1135)
* feat(security): v2 ensemble tuning — label-first voting + SOLO_CONTENT_BLOCK Cuts Haiku classifier false-positive rate from 44.1% → 22.9% on BrowseSafe-Bench smoke. Detection trades from 67.3% → 56.2%; the lost TPs are all cases Haiku correctly labeled verdict=warn (phishing targeting users, not agent hijack) — they still surface in the WARN banner meta but no longer kill the session. Key changes: - combineVerdict: label-first voting for transcript_classifier. Only meta.verdict==='block' block-votes; verdict==='warn' is a soft signal. Missing meta.verdict never block-votes (backward-compat). - Hallucination guard: verdict='block' at confidence < LOG_ONLY (0.40) drops to warn-vote — prevents malformed low-conf blocks from going authoritative. - New THRESHOLDS.SOLO_CONTENT_BLOCK = 0.92 decoupled from BLOCK (0.85). Label-less content classifiers (testsavant, deberta) need a higher solo-BLOCK bar because they can't distinguish injection from phishing-targeting-user. Transcript keeps label-gated solo path (verdict=block AND conf >= BLOCK). - THRESHOLDS.WARN bumped 0.60 → 0.75 — borderline fires drop out of the 2-of-N ensemble pool. - Haiku model pinned (claude-haiku-4-5-20251001). `claude -p` spawns from os.tmpdir() so project CLAUDE.md doesn't poison the classifier context (measured 44k cache_creation tokens per call before the fix, and Haiku refusing to classify because it read "security system" from CLAUDE.md and went meta). - Haiku timeout 15s → 45s. Measured real latency is 17-33s end-to-end (Claude Code session startup + Haiku); v1's 15s caused 100% timeout when re-measured — v1's ensemble was effectively L4-only in prod. - Haiku prompt rewritten: explicit block/warn/safe criteria, 8 few-shot exemplars (instruction-override → block; social engineering → warn; discussion-of-injection → safe). Test updates: - 5 existing combineVerdict tests adapted for label-first semantics (transcript signals now need meta.verdict to block-vote). - 6 new tests: warn-soft-signal, three-way-block-with-warn-transcript, hallucination-guard-below-floor, above-floor-label-first, backward-compat-missing-meta. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * test(security): live + fixture-replay bench harness with 500-case capture Adds two new benches that permanently guard the v2 tuning: - security-bench-ensemble-live.test.ts (opt-in via GSTACK_BENCH_ENSEMBLE=1). Runs full ensemble on BrowseSafe-Bench smoke with real Haiku calls. Worker-pool concurrency (default 8, tunable via GSTACK_BENCH_ENSEMBLE_CONCURRENCY) cuts wall clock from ~2hr to ~25min on 500 cases. Captures Haiku responses to fixture for replay. Subsampling via GSTACK_BENCH_ENSEMBLE_CASES for faster iteration. Stop-loss iterations write to ~/.gstack-dev/evals/stop-loss-iter-N-* WITHOUT overwriting canonical fixture. - security-bench-ensemble.test.ts (CI gate, deterministic replay). Replays captured fixture through combineVerdict, asserts detection >= 55% AND FP <= 25%. Fail-closed when fixture is missing AND security-layer files changed in branch diff. Uses `git diff --name-only base` (two-dot) to catch both committed and working-tree changes — `git diff base...HEAD` would silently skip in CI after fixture lands. - browse/test/fixtures/security-bench-haiku-responses.json — 500 cases × 3 classifier signals each. Header includes schema_version, pinned model, component hashes (prompt, exemplars, thresholds, combiner, dataset version). Any change invalidates the fixture and forces fresh live capture. - docs/evals/security-bench-ensemble-v2.json — durable PR artifact with measured TP/FN/FP/TN, 95% CIs, knob state, v1 baseline delta. Checked in so reviewers can see the numbers that justified the ship. Measured baseline on the new harness: TP=146 FN=114 FP=55 TN=185 → 56.2% / 22.9% → GATE PASS Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * chore(release): v1.5.1.0 — cut Haiku FP 44% → 23% - VERSION: 1.5.0.0 → 1.5.1.0 (TUNING bump) - CHANGELOG: [1.5.1.0] entry with measured numbers, knob list, and stop-loss rule spec - TODOS: mark "Cut Haiku FP 44% → ~15%" P0 as SHIPPED with pointer to CHANGELOG and v1 plan Measured: 56.2% detection (CI 50.1-62.1) / 22.9% FP (CI 18.1-28.6) on 500-case BrowseSafe-Bench smoke. Gate passes (floor 55%, ceiling 25%). Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * docs(changelog): add v1.6.4.0 placeholder entry at top Per CLAUDE.md branch-scoped discipline, our VERSION 1.6.4.0 needs a CHANGELOG entry at the top so readers can tell what's on this branch vs main. Honest placeholder: no user-facing runtime changes yet, two merges bringing branch up to main's v1.6.3.0, and the approved injection-tuning plan is queued but unimplemented. Gets replaced by the real release-summary at /ship time after Phases -1 through 10 land. * docs(changelog): strip process minutiae from entries; rewrite v1.6.4.0 CLAUDE.md — new CHANGELOG rule: only document what shipped between main and this change. Keep out branch resyncs, merge commits, plan approvals, review outcomes, scope negotiations, "work queued" or "in-progress" framing. When no user-facing change actually landed, one sentence is the entry: "Version bump for branch-ahead discipline. No user-facing changes yet." CHANGELOG.md — v1.6.4.0 entry rewritten to match. Previous version narrated the branch history, the approved injection-tuning plan, and what we expect to ship later — all of which are process minutiae readers do not care about. * docs(changelog): rewrite v1.6.4.0; strip process minutiae Rewrote v1.6.4.0 entry to follow the new CLAUDE.md rule: only document what shipped between main and this change. Previous entry narrated the branch history, the approved injection-tuning plan, and what we expect to ship later, all process minutiae readers do not care about. v1.6.4.0 now reads: what the detection tuning did for users, the before/after numbers, the stop-loss rule, and the itemized changes for contributors. CLAUDE.md — new rule: only document what shipped between main and this change. Keep out branch resyncs, merge commits, plan approvals, review outcomes, scope negotiations, "work queued" / "in-progress" framing. If nothing user-facing landed, one sentence: "Version bump for branch-ahead discipline. No user-facing changes yet." --------- Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
||
|
|
97584f9a59 |
feat(security): ML prompt injection defense for sidebar (v1.4.0.0) (#1089)
* chore(deps): add @huggingface/transformers for prompt injection classifier Dependency needed for the ML prompt injection defense layer coming in the follow-up commits. @huggingface/transformers will host the TestSavantAI BERT-small classifier that scans tool outputs for indirect prompt injection. Note: this dep only runs in non-compiled bun contexts (sidebar-agent.ts). The compiled browse binary cannot load it because transformers.js v4 requires onnxruntime-node (native module, fails to dlopen from bun compile's temp extract dir). See docs/designs/ML_PROMPT_INJECTION_KILLER.md for the full architectural decision. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * feat(security): add security.ts foundation for prompt injection defense Establishes the module structure for the L5 canary and L6 verdict aggregation layers. Pure-string operations only — safe to import from the compiled browse binary. Includes: * THRESHOLDS constants (BLOCK 0.85 / WARN 0.60 / LOG_ONLY 0.40), calibrated against BrowseSafe-Bench smoke + developer content benign corpus. * combineVerdict() implementing the ensemble rule: BLOCK only when the ML content classifier AND the transcript classifier both score >= WARN. Single-layer high confidence degrades to WARN to prevent any one classifier's false-positives from killing sessions (Stack Overflow instruction-writing-style FPs at 0.99 on TestSavantAI alone). * generateCanary / injectCanary / checkCanaryInStructure — session-scoped secret token, recursively scans tool arguments, URLs, file writes, and nested objects per the plan's all-channel coverage decision. * logAttempt with 10MB rotation (keeps 5 generations). Salted SHA-256 hash, per-device salt at ~/.gstack/security/device-salt (0600). * Cross-process session state at ~/.gstack/security/session-state.json (atomic temp+rename). Required because server.ts (compiled) and sidebar-agent.ts (non-compiled) are separate processes. * getStatus() for shield icon rendering via /health. ML classifier code will live in a separate module (security-classifier.ts) loaded only by sidebar-agent.ts — compiled browse binary cannot load the native ONNX runtime. Plan: ~/.gstack/projects/garrytan-gstack/ceo-plans/2026-04-19-prompt-injection-guard.md Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * feat(security): wire canary injection into sidebar spawnClaude Every sidebar message now gets a fresh CANARY-XXXXXXXXXXXX token embedded in the system prompt with an instruction for Claude to never output it on any channel. The token flows through the queue entry so sidebar-agent.ts can check every outbound operation for leaks. If Claude echoes the canary into any outbound channel (text stream, tool arguments, URLs, file write paths), the sidebar-agent terminates the session and the user sees the approved canary leak banner. This operation is pure string manipulation — safe in the compiled browse binary. The actual output-stream check (which also has to be safe in compiled contexts) lives in sidebar-agent.ts (next commit). Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * test(security): make sidebar-agent destructure check regex-tolerant The test asserted the exact string `const { prompt, args, stateFile, cwd, tabId } = queueEntry` which breaks whenever security or other extensions add fields (canary, pageUrl, etc.). Switch to a regex that requires the core fields in order but tolerates additional fields in between. Preserves the test's intent (args come from the queue entry, not rebuilt) while allowing the destructure to grow. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * feat(security): canary leak check across all outbound channels The sidebar-agent now scans every Claude stream event for the session's canary token before relaying any data to the sidepanel. Channels covered (per CEO review cross-model tension #2): * Assistant text blocks * Assistant text_delta streaming * tool_use arguments (recursively, via checkCanaryInStructure — catches URLs, commands, file paths nested at any depth) * tool_use content_block_start * tool_input_delta partial JSON * Final result payload If the canary leaks on any channel, onCanaryLeaked() fires once per session: 1. logAttempt() writes the event to ~/.gstack/security/attempts.jsonl with the canary's salted hash (never the payload content). 2. sends a `security_event` to the sidepanel so it can render the approved canary-leak banner (variant A mockup — ceo-plan 2026-04-19). 3. sends an `agent_error` for backward-compat with existing error surfaces. 4. SIGTERM's the claude subprocess (SIGKILL after 2s if still alive). The leaked content itself is never relayed to the sidepanel — the event is dropped at the boundary. Canary detection is pure-string substring match, so this all runs safely in the sidebar-agent (non-compiled bun) context. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * feat(security): add security-classifier.ts with TestSavantAI + Haiku This module holds the ML classifier code that the compiled browse binary cannot link (onnxruntime-node native dylib doesn't load from Bun compile's temp extract dir — see CEO plan §"Pre-Impl Gate 1 Outcome"). It's imported ONLY by sidebar-agent.ts, which runs as a non-compiled bun script. Two layers: L4 testsavant_content — TestSavantAI BERT-small ONNX classifier. First call triggers a one-time 112MB model download to ~/.gstack/models/testsavant-small/ (files staged into the onnx/ layout transformers.js v4 expects). Classifies page snapshots and tool outputs for indirect prompt injection + jailbreak attempts. On benign-corpus dry-run: Wikipedia/HN/Reddit/tech-blog all score SAFE 0.98+, attack text scores INJECTION 0.99+, Stack Overflow instruction-writing now scores SAFE 0.98 on the shorter form (was 0.99 INJECTION on the longer form — instruction-density threshold). Ensemble combiner downgrades single-layer high to WARN to cover this case. L4b transcript_classifier — Claude Haiku reasoning-blind pre-tool-call scan. Sees only {user_message, last 3 tool_calls}, never Claude's chain-of-thought or tool results (those are how self-persuasion attacks leak). 2000ms hard timeout. Fail-open on any subprocess failure so sidebar stays functional. Gated by shouldRunTranscriptCheck() — only runs when another layer already fired at >= LOG_ONLY, saving ~70% of Haiku spend. Both layers degrade gracefully: load/spawn failures set status to 'degraded' and return confidence=0. Shield icon reflects this via getClassifierStatus() which security.ts's getStatus() composes. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * feat(security): wire TestSavantAI + ensemble into sidebar-agent pre-spawn scan The sidebar-agent now runs a ML security check on the user message BEFORE spawning claude. If the content classifier and (gated) transcript classifier ensemble returns BLOCK, the session is refused with a security_event + agent_error — the sidepanel renders the approved banner. Two pieces: 1. On agent startup, loadTestsavant() warms the classifier in the background. First run triggers a 112MB model download from HuggingFace (~30s on average broadband). Non-blocking — sidebar stays functional during cold-start, shield just reports 'off' until warmed. 2. preSpawnSecurityCheck() runs the ensemble against the user message: - L4 (testsavant_content) always runs - L4b (transcript_classifier via Haiku) runs only if L4 flagged at >= LOG_ONLY — plan §E1 gating optimization, saves ~70% of Haiku spend combineVerdict() applies the BLOCK-requires-both-layers rule, which downgrades any single-layer high confidence to WARN. Stack Overflow-style instruction-heavy writing false-positives on TestSavantAI alone are caught by this degrade — Haiku corrects them when called. Fail-open everywhere: any subprocess/load/inference error returns confidence=0 so the sidebar keeps working on architectural controls alone. Shield icon reflects degraded state via getClassifierStatus(). BLOCK path emits both: - security_event {verdict, reason, layer, confidence, domain} (for the approved canary-leak banner UX mockup — variant A) - agent_error "Session blocked — prompt injection detected..." (backward-compat with existing error surface) Regression test suite still passes (12/12 sidebar-security tests). Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * test(security): add security.ts unit tests (25 tests, 62 assertions) Covers the pure-string operations that must behave deterministically in both compiled and source-mode bun contexts: * THRESHOLDS ordering invariant (BLOCK > WARN > LOG_ONLY > 0) * combineVerdict ensemble rule — THE critical path: - Empty signals → safe - Canary leak always blocks (regardless of ML signals) - Both ML layers >= WARN → BLOCK (ensemble_agreement) - Single layer >= BLOCK → WARN (single_layer_high) — the Stack Overflow FP mitigation that prevents one classifier killing sessions alone - Max-across-duplicates when multiple signals reference the same layer * Canary generation + injection + recursive checking: - Unique CANARY-XXXXXXXXXXXX tokens (>= 48 bits entropy) - Recursive structure scan for tool_use inputs, nested URLs, commands - Null / primitive handling doesn't throw * Payload hashing (salted sha256) — deterministic per-device, differs across payloads, 64-char hex shape * logAttempt writes to ~/.gstack/security/attempts.jsonl * writeSessionState + readSessionState round-trip (cross-process) * getStatus returns valid SecurityStatus shape * extractDomain returns hostname only, empty string on bad input All 25 tests pass in 18ms — no ML, no network, no subprocess spawning. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * feat(security): expose security status on /health for shield icon The /health endpoint now returns a `security` field with the classifier status, suitable for driving the sidepanel shield icon: { status: 'protected' | 'degraded' | 'inactive', layers: { testsavant, transcript, canary }, lastUpdated: ISO8601 } Backend plumbing: * server.ts imports getStatus from security.ts (pure-string, safe in compiled binary) and includes it in the /health response. * sidebar-agent.ts writes ~/.gstack/security/session-state.json when the classifier warmup completes (success OR failure). This is the cross- process handoff — server.ts reads the state file via getStatus() to surface the result to the sidepanel. The sidepanel rendering (SVG shield icon + color states + tooltip) is a follow-up commit in the extension/ code. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * docs(security): document the sidebar security stack in CLAUDE.md Adds a security section to the Browser interaction block. Covers: * Layered defense table showing which modules live where (content-security.ts in both contexts vs security-classifier.ts only in sidebar-agent) and why the split exists (onnxruntime-node incompatibility with compiled Bun) * Threshold constants (0.85 / 0.60 / 0.40) and the ensemble rule that prevents single-classifier false-positives (the Stack Overflow FP story) * Env knobs — GSTACK_SECURITY_OFF kill switch, cache paths, salt file, attack log rotation, session state file This is the "before you modify the security stack, read this" doc. It lives next to the existing Sidebar architecture note that points at SIDEBAR_MESSAGE_FLOW.md. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * docs(todos): mark ML classifier v1 in-progress + file v2 follow-ups Reframes the P0 item to reflect v1 scope (branch 2 architecture, TestSavantAI pivot, what shipped) and splits v2 work into discrete TODOs: * Shield icon + canary leak banner UI (P0, blocks v1 user-facing completion) * Attack telemetry via gstack-telemetry-log (P1) * Full BrowseSafe-Bench at gate tier (P2) * Cross-user aggregate attack dashboard (P2) * DeBERTa-v3 as third signal in ensemble (P2) * Read/Glob/Grep ingress coverage (P2, flagged by Codex review) * Adversarial + integration + smoke-bench test suites (P1) * Bun-native 5ms inference (P3 research) Each TODO carries What / Why / Context / Effort / Priority / Depends-on so it's actionable by someone picking it up cold. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * feat(telemetry): add attack_attempt event type to gstack-telemetry-log Extends the existing telemetry pipe with 5 new flags needed for prompt injection attack reporting: --url-domain hostname only (never path, never query) --payload-hash salted sha256 hex (opaque — no payload content ever) --confidence 0-1 (awk-validated + clamped; malformed → null) --layer testsavant_content | transcript_classifier | aria_regex | canary --verdict block | warn | log_only Backward compatibility: * Existing skill_run events still work — all new fields default to null * Event schema is a superset of the old one; downstream edge function can filter by event_type No new auth, no new SDK, no new Supabase migration. The same tier gating (community → upload, anonymous → local only, off → no-op) and the same sync daemon carry the attack events. This is the "E6 RESOLVED" path from the CEO plan — riding the existing pipe instead of spinning up parallel infra. Verified end-to-end: * attack_attempt event with all fields emits correctly to skill-usage.jsonl * skill_run event with no security flags still works (backward compat) Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * feat(security): wire logAttempt to gstack-telemetry-log (fire-and-forget) Every local attempt.jsonl write now also triggers a subprocess call to gstack-telemetry-log with the attack_attempt event type. The binary handles tier gating internally (community → Supabase upload, anonymous → local JSONL only, off → no-op), so security.ts doesn't need to re-check. Binary resolution follows the skill preamble pattern — never relies on PATH, which breaks in compiled-binary contexts: 1. ~/.claude/skills/gstack/bin/gstack-telemetry-log (global install) 2. .claude/skills/gstack/bin/gstack-telemetry-log (symlinked dev) 3. bin/gstack-telemetry-log (in-repo dev) Fire-and-forget: * spawn with stdio: 'ignore', detached: true, unref() * .on('error') swallows failures * Missing binary is non-fatal — local attempts.jsonl still gives audit trail Never throws. Never blocks. Existing 37 security tests pass unchanged. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * feat(ui): add security banner markup + styles (approved variant A) HTML + CSS for the canary leak / ML block banner. Structure matches the approved mockup from /plan-design-review 2026-04-19 (variant A — centered alert-heavy): * Red alert-circle SVG icon (no stock shield, intentional — matches the "serious but not scary" tone the review chose) * "Session terminated" Satoshi Bold 18px red headline * "— prompt injection detected from {domain}" DM Sans zinc subtitle * Expandable "What happened" chevron button (aria-expanded/aria-controls) * Layer list rendered in JetBrains Mono with amber tabular-nums scores * Close X in top-right, 28px hit area, focus-visible amber outline Enter animation: slide-down 8px + fade, 250ms, cubic-bezier(0.16,1,0.3,1) — matches DESIGN.md motion spec. Respects `role="alert"` + `aria-live="assertive"` so screen readers announce on appearance. Escape-to-dismiss hook is in the JS follow-up commit. Design tokens all via CSS variables (--error, --amber-400, --amber-500, --zinc-*, --font-display, --font-mono, --radius-*) — already established in the stylesheet. No new color constants introduced. JS wiring lands in the next commit so this diff stays focused on presentation layer only. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * feat(ui): wire security banner to security_event + interactivity Adds showSecurityBanner() and hideSecurityBanner() plus the addChatEntry routing for entry.type === 'security_event'. When the sidebar-agent emits a security_event (canary leak or ML BLOCK), the banner renders with: * Title ("Session terminated") * Subtitle with {domain} if present, otherwise generic * Expandable layer list — each row: SECURITY_LAYER_LABELS[layer] + confidence.toFixed(2) in mono. Readable + auditable — user can see which layer fired at what score Interactivity, wired once on DOMContentLoaded: * Close X → hideSecurityBanner() * Expand/collapse "What happened" → toggles details + aria-expanded + chevron rotation (200ms css transition already in place) * Escape key dismisses while banner is visible (a11y) No shield icon yet — that's a separate commit that will consume the `security` field now returned by /health. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * feat(ui): add security shield icon in sidepanel header (3 states) Small "SEC" badge in the top-right of the sidepanel that reflects the security module's current state. Three states drive color: protected green — all layers ok (TestSavantAI + transcript + canary) degraded amber — one+ ML layer offline but canary + arch controls active inactive red — security module crashed, arch controls only Consumes /health.security (surfaced in commit |