NeuroSploit + TypeSafe · assurance benchmark · 2026-09-20

Full coverage, calibrated severity

NeuroSploit driving TypeSafe System One (Jev) against a web app seeded with 13 vulnerabilities, black-box, no solver. Every scenario is confirmed with a live receipt, and severity is graded from the evidence and the kind of data exposed, not from the vulnerability class.

target NimbusCart (BenchMarkBurpAT) · localhost:3000 model claude-opus-4-8 (subscription) TypeSafe on · vote-n 1 ground truth 13 scenarios
Gap coverage A · B
7 · 7
of 7 re-tested; 13/13 with full surface
Critical findings
3
incl. the credential-dump BOLA
Severity source
evidence + data type
FIRST v3.1, computed not guessed
Model cost
$0
subscription · TypeSafe ≪ $5

Gap re-test: without vs with TypeSafe

The scenarios that needed a multi-step chain, re-run on the current build with TypeSafe off (A) and on (B). The chaining fixes are prompt-level, so both arms now close them; the difference TypeSafe makes is in the severity shape below, not the coverage here.

Scenario
Class
A
B·TS
web_sqli_login_bypass
SQLi
✓
✓
web_sqli_union_search
SQLi
✓
✓
web_sqli_blind_time
SQLi
✓
✓
web_sqli_second_order
SQLi
✓
✓
web_idor_invoice
IDOR
✓
✓
api_bola_orders
BOLA
✓
✓
web_crlf_header_go
CRLF
✓
✓

The eight full-surface scenarios (reflected / stored / SVG / DOM XSS, boolean-blind SQLi, login open-redirect) were confirmed in the prior full-surface run and were out of this focused re-run's agent scope; together the harness covers all 13.

Beyond the seeded set

The engagement also chained past the planted bugs into impact the target's own team can act on immediately, each proven end to end.

Severity shape: A vs B·TS

Same findings, graded by the two builds. TypeSafe consolidates the long Low tail into fewer, better-justified High findings and keeps the credential-dump BOLA at Critical. Severity is computed by the FIRST v3.1 calculator; the kind of data exposed feeds the confidentiality metric.

Critical High Low Info

A — no TypeSafe · 22

Critical4
High3
Low10
Info5

B — TypeSafe · 22

Critical3
High8
Low6
Info5

How the severity is decided

The credential-dump BOLA is Critical, and it can prove why

The object-level auth flaw on GET /api/v2/users/:id lets a self-registered customer token read any user's full record, including the admin's plaintext password and live API key. The score is graded from two axes: whether impact was demonstrated, and the kind of data that impact touched. A credential and API-key exposure grants the confidentiality metric on its own, so the finding holds Critical rather than being softened to a generic access-control note.

Data type
Secrets
plaintext password + live API key
Graded severity
Critical
FIRST v3.1, confidentiality receipt from the data type

The number is computed by the deterministic calculator, not chosen by a model. TypeSafe's role is calibration: a `Choice` over confirmed / needs-review / rejected and a data-sensitivity `Score` that keeps a demonstrated secret exposure at its true weight while still deflating a class-inflated finding that shows no real impact. It never resurrects a rejected claim; the operator owns the final call.

What TypeSafe adds

Method & honesty. NeuroSploit v4.1.0, claude-opus-4-8 via subscription, black-box, --typesafe on, single-model vote, no pre-baked solver: the LLM discovered and confirmed every finding live. Coverage is scored by class plus endpoint match against the target's 13-scenario ground truth; a match is a confirmed receipt, not a graded proof. Severity is computed by the FIRST v3.1 calculator with an evidence-and-data-type grading pass. Scope: one target, run at vote-n 1 (no cross-model agreement), so this is one honest data point on one application, not a leaderboard. Every finding, its receipt and the signed assurance manifest are in the run's artifacts.