Merge origin/main (v1.64.1.0 code-smell wave) into test-evals-ci-speedup

Both sides shipped overlapping test-infra work in parallel; resolutions
compose intent rather than picking sides:

- free-tests.yml (both added): keep this branch's lane (canonical
  strict-parallel runner, secretless, plain runner, ~2min) over main's
  per-file-serial container loop (45min budget, hand-curated skip list,
  needs GITHUB_TOKEN); ported main's git safe.directory insight.
- Dockerfile.ci Bun install: main discovered the installer IGNORES the
  BUN_VERSION env var (the old form silently installed latest) — main's
  arg-form mechanism + this branch's 1.3.13 target.
- parity baseline: both sides rebased after hitting the same silent
  drift; adopted main's v1.64.1.0 union-normalized fixture and dropped
  this branch's interim v1.64.0.0 capture. carve-guards caps: main's
  tighter re-ratchets win (all four).
- touchfiles: kept this branch's three-file facade split; ported main's
  pure-data removals (dead sidebar-agent entries, spec judge entry,
  ship-idempotency) into touchfiles-data.ts.
- ship-idempotency SDK variant: main deliberately removed it as
  redundant with the real-PTY test; adopted — dropped this branch's
  rehomed copy and its periodic matrix row (the zombie-monolith deletion
  stands; coverage-audit + triage rehomes verified untouched by main).
- e2e-tier-alignment: taught the new parent-mapper hard check main's
  consolidated describeE2ETier()/e2eTierEnabled() self-gate shapes (the
  helper's header names this file as a required recognizer).
- browse/test/compare-board.test.ts: quarantined behind
  GSTACK_COMPARE_BOARD_TESTS=1 — all 16 tests fail identically on
  origin/main solo on dev machines (blame protocol receipts in-file);
  main's own CI lane skip-lists it. An always-red file would block every
  PR now that free-tests is a required check.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
This commit is contained in:
Garry Tan
2026-08-15 10:29:25 -07:00
co-authored by Claude Fable 5
192 changed files with 2371 additions and 22872 deletions
+7 -8
View File
@@ -113,16 +113,16 @@ describe('BrowseSafe-Bench smoke (200 cases)', () => {
// subtle social engineering ("URGENT: verify your account within 24h")
// which the bench is dense with.
//
// The production ensemble (L4 content + L4b transcript classifier
// agreement) filters FPs in practice. This bench measures L4 alone.
// Production runs L4 alone (testsavant via the security sidecar) — the
// Haiku transcript and DeBERTa ensemble layers were removed with their
// host process. This bench measures exactly the live layer.
//
// Tracked for improvement:
// * DeBERTa-v3 ensemble (TODOS "TestSavantAI ensemble" P2 item)
// * Calibration against real attack-log data from production
let tp = 0, fp = 0, tn = 0, fn = 0;
// intentionally 0.6 — L4-only bench pinned to v1 WARN for historical
// continuity. The ensemble bench (security-bench-ensemble.test.ts) uses
// THRESHOLDS.WARN from security.ts (0.75 in v1.5.2.0+).
// continuity (production THRESHOLDS.WARN in security.ts is 0.75 in
// v1.5.2.0+).
const WARN = 0.6;
for (const row of rows) {
const signal = await scanPageContent(row.content);
@@ -142,9 +142,8 @@ describe('BrowseSafe-Bench smoke (200 cases)', () => {
console.log(`[browsesafe-bench] False-positive rate: ${(fpRate * 100).toFixed(1)}% (v1 baseline — ensemble filters in prod)`);
// V1 sanity gates — does the classifier provide ANY signal?
// These are intentionally loose. Quality gates arrive when the DeBERTa
// ensemble lands (P2 TODO) and we can measure the 2-of-3 agreement
// rate against this same bench.
// These are intentionally loose: L4 alone is a signal source, not a
// verdict — combineVerdict + the L1-L3 layers own the final decision.
expect(tp).toBeGreaterThan(0); // classifier fires on some attacks
expect(tn).toBeGreaterThan(0); // classifier is not stuck-on
expect(tp + fp).toBeGreaterThan(0); // classifier fires at all