NeuroSploit · assurance benchmark · 2026-09-20

Does TypeSafe make the run better?

Two identical NeuroSploit engagements against the same vulnerable target — one plain, one with TypeSafe System One (Jev) as a calibrated confirmation layer. Same model, same focus, same 13 seeded vulnerabilities. Only the --typesafe flag differs.

target NimbusCart (BenchMarkBurpAT) · localhost:3000 model claude-opus-4-8 (subscription) recon 2 · vote-n 1 · max-agents 15 ground truth 13 targets
Recall — no TypeSafe
10/13
16 findings · 32m12s
Recall — with TypeSafe
9/13
18 findings · 26m53s
Union coverage
11/13
the two runs together
TypeSafe recalibrated
9
findings, calibrated confidence

Head to head

The recall is a tie inside the noise; the real difference is shape. TypeSafe was faster, surfaced two real findings the plain run missed, and pulled inflated severities down toward what the evidence actually demonstrated — at the cost of being conservative enough to drop two scenarios and under-rate one genuine critical.

MetricA — no TypeSafeB — TypeSafe
Seeded targets hit10 / 139 / 13
Total findings reported1618
Findings beyond the 13 targets69 (2 real: config leak, no-lockout)
Wall-clock time32m 12s26m 53s
Criticals reported52 (recalibrated)
Belief-gate holds (POMDP)3—
Assurance P1–P5all presentall present
Model API cost$0 subscription$0 + TypeSafe ≪ $5

Per-scenario coverage

Each seeded vulnerability, and whether each run confirmed it. Neither run reached the second-order SQLi or the CRLF header injection — the two that need a multi-step chain the single-vote config didn't pursue.

Scenario
A
B·TS
web_sqli_login_bypassSQLi
✓
✓
web_sqli_union_searchSQLi
✓
✕
web_sqli_blind_booleanSQLi
✓
✓
web_sqli_blind_timeSQLi
✓
✕
web_sqli_second_orderSQLi
✕
✕
web_xss_reflected_searchXSS
✓
✓
web_xss_stored_reviewXSS
✓
✓
web_xss_svg_uploadXSS
✓
✓
web_xss_dom_redirectXSS
✓
✓
web_idor_invoiceIDOR
✕
✓
api_bola_ordersBOLA
✓
✓
web_open_redirect_loginRedirect
✓
✓
web_crlf_header_goCRLF
✕
✕

Severity shape

The clearest effect of TypeSafe: the severity distribution flattens. The plain run stacks five Criticals; the calibrated run keeps two and pushes the rest down to where the demonstrated-impact evidence puts them.

CriticalHighMediumLowInfo

A — no TypeSafe · 16

Critical5
High4
Medium1
Low2
Info4

B — TypeSafe · 18

Critical2
High4
Medium3
Low4
Info5

What calibration actually did

The same BOLA, two severities

Both runs found the object-level auth flaw on GET /api/v2/users/:id — a customer token reads any user's full record, including the admin's plaintext password. The plain run rated it Critical (9.1) on the class. TypeSafe, grading against the demonstrated-impact receipts and its calibrated judgment, rated it Low.

A — class-graded
Critical 9.1
BOLA + excessive data exposure
B — evidence-graded
Low
same finding, impact receipts weighted

Why it fired: the severity is graded from the structured evidence slot (the recorded request/response exchange), not the agent's prose. This finding proved the dump in its narrative and claims ledger but left evidence_data null — so the demonstrated-impact rung saw no machine-readable C/I/A receipt, and the calibrated grader dropped the impact metrics to None, collapsing 9.1 → Low. The proof existed; it just wasn't in the slot the grader reads.

This is the honest edge: calibration removes inflated Criticals (good — most scanners over-rate by class), but a receipt in the wrong slot gets under-rated. It is a dial toward defensibility, not a correctness oracle — the operator still owns the final severity, and the fix is to make agents populate evidence_data for impact, not to loosen the grader.

Takeaways

Method & honesty. Both runs: NeuroSploit v4.0.0, claude-opus-4-8 via subscription, black-box, recon intensity 2, single-model vote (vote-n 1), same natural-language focus naming the 13 endpoints, no pre-baked solver — the LLM discovered and confirmed everything live. Recall is scored by class + endpoint keyword match against the target's ground-truth list, so a match is coverage, not a graded proof. Confounders: the two runs are single samples, not averages; an earlier TypeSafe run collapsed to zero when the subscription hit a session limit mid-run (a real harness gap, since fixed — session-limit stdout now parks the run instead of burning agents); vote-n 1 means no cross-model agreement in either arm. Treat this as one honest data point on one target, not a leaderboard. Not measured here: multi-sample variance, higher vote-n, and TypeSafe's agent-pruning effect on a broader agent set.