mirror of
https://github.com/garrytan/gstack.git
synced 2026-09-23 05:10:50 +02:00
fix(security): delete the dead ML layers — transcript classifier and DeBERTa ensemble
The L4b Haiku transcript classifier and the opt-in DeBERTa ensemble
(GSTACK_SECURITY_ENSEMBLE=deberta, a documented 721MB download) had ZERO
production callers since the chat-path agent that invoked them was ripped.
The only live ML path is scanPageContent (testsavant) inside the security
sidecar subprocess. Deleted by import graph:
- security-classifier.ts 614 -> 265 lines: HAIKU_MODEL, checkTranscript,
shouldRunTranscriptCheck, loadDeberta, scanPageContentDeberta, ToolCallInput,
all DEBERTA_* consts + load state. Header now states the live truth
(imported only by security-sidecar-entry.ts). downloadFile kept, name
intact — it is an enumerated egress sink (HF model download).
- security-bunnative.ts + test: a research skeleton self-described as 'NOT a
production replacement', shipped into src/ with zero importers.
- security-bench-ensemble{,-live}.test.ts + the Haiku response fixture: a
paid live-model benchmark for a layer that could not fire. The
security-classifier-tdz test's only case exercised checkTranscript — gone.
- security.ts: layer-model header rewritten to the live architecture;
StatusDetail.layers -> {testsavant, canary}; getStatus() no longer requires
the impossible transcript==='ok' for 'protected' (old on-disk session state
with a transcript key is tolerated on read, never re-emitted).
- security-sidecar-entry.ts needed zero changes: it serializes
getClassifierStatus() verbatim and no consumer read .transcript (verified
in sidecar-client + server.ts).
- BROWSER.md security section matches reality (ensemble knob gone, 112MB not
22MB, sidecar hosting documented). combineVerdict/THRESHOLDS retained as
the pure, tested combiner of record — comments now flag transcript/deberta
votes as producer-less.
Net: 26 pass in security.test.ts incl. a NEW regression test for stale-
transcript disk tolerance; egress-receipt tripwire green.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
This commit is contained in:
co-authored by
Claude Fable 5
parent
ef186cccb3
commit
4328748136
@@ -109,16 +109,16 @@ describe('BrowseSafe-Bench smoke (200 cases)', () => {
|
||||
// subtle social engineering ("URGENT: verify your account within 24h")
|
||||
// which the bench is dense with.
|
||||
//
|
||||
// The production ensemble (L4 content + L4b transcript classifier
|
||||
// agreement) filters FPs in practice. This bench measures L4 alone.
|
||||
// Production runs L4 alone (testsavant via the security sidecar) — the
|
||||
// Haiku transcript and DeBERTa ensemble layers were removed with their
|
||||
// host process. This bench measures exactly the live layer.
|
||||
//
|
||||
// Tracked for improvement:
|
||||
// * DeBERTa-v3 ensemble (TODOS "TestSavantAI ensemble" P2 item)
|
||||
// * Calibration against real attack-log data from production
|
||||
let tp = 0, fp = 0, tn = 0, fn = 0;
|
||||
// intentionally 0.6 — L4-only bench pinned to v1 WARN for historical
|
||||
// continuity. The ensemble bench (security-bench-ensemble.test.ts) uses
|
||||
// THRESHOLDS.WARN from security.ts (0.75 in v1.5.2.0+).
|
||||
// continuity (production THRESHOLDS.WARN in security.ts is 0.75 in
|
||||
// v1.5.2.0+).
|
||||
const WARN = 0.6;
|
||||
for (const row of rows) {
|
||||
const signal = await scanPageContent(row.content);
|
||||
@@ -138,9 +138,8 @@ describe('BrowseSafe-Bench smoke (200 cases)', () => {
|
||||
console.log(`[browsesafe-bench] False-positive rate: ${(fpRate * 100).toFixed(1)}% (v1 baseline — ensemble filters in prod)`);
|
||||
|
||||
// V1 sanity gates — does the classifier provide ANY signal?
|
||||
// These are intentionally loose. Quality gates arrive when the DeBERTa
|
||||
// ensemble lands (P2 TODO) and we can measure the 2-of-3 agreement
|
||||
// rate against this same bench.
|
||||
// These are intentionally loose: L4 alone is a signal source, not a
|
||||
// verdict — combineVerdict + the L1-L3 layers own the final decision.
|
||||
expect(tp).toBeGreaterThan(0); // classifier fires on some attacks
|
||||
expect(tn).toBeGreaterThan(0); // classifier is not stuck-on
|
||||
expect(tp + fp).toBeGreaterThan(0); // classifier fires at all
|
||||
|
||||
Reference in New Issue
Block a user