Two identical NeuroSploit engagements against the same vulnerable target — one plain,
one with TypeSafe System One (Jev) as a calibrated confirmation layer. Same model, same focus,
same 13 seeded vulnerabilities. Only the --typesafe flag differs.
The recall is a tie inside the noise; the real difference is shape. TypeSafe was faster, surfaced two real findings the plain run missed, and pulled inflated severities down toward what the evidence actually demonstrated — at the cost of being conservative enough to drop two scenarios and under-rate one genuine critical.
| Metric | A — no TypeSafe | B — TypeSafe |
|---|---|---|
| Seeded targets hit | 10 / 13 | 9 / 13 |
| Total findings reported | 16 | 18 |
| Findings beyond the 13 targets | 6 | 9 (2 real: config leak, no-lockout) |
| Wall-clock time | 32m 12s | 26m 53s |
| Criticals reported | 5 | 2 (recalibrated) |
| Belief-gate holds (POMDP) | 3 | — |
| Assurance P1–P5 | all present | all present |
| Model API cost | $0 subscription | $0 + TypeSafe ≪ $5 |
Each seeded vulnerability, and whether each run confirmed it. Neither run reached the second-order SQLi or the CRLF header injection — the two that need a multi-step chain the single-vote config didn't pursue.
The clearest effect of TypeSafe: the severity distribution flattens. The plain run stacks five Criticals; the calibrated run keeps two and pushes the rest down to where the demonstrated-impact evidence puts them.
Both runs found the object-level auth flaw on GET /api/v2/users/:id — a customer token
reads any user's full record, including the admin's plaintext password. The plain run rated it
Critical (9.1) on the class. TypeSafe, grading against the demonstrated-impact receipts and its
calibrated judgment, rated it Low.
Why it fired: the severity is graded from the structured evidence
slot (the recorded request/response exchange), not the agent's prose. This finding proved the dump in its
narrative and claims ledger but left evidence_data null — so the demonstrated-impact rung saw no
machine-readable C/I/A receipt, and the calibrated grader dropped the impact metrics to None,
collapsing 9.1 → Low. The proof existed; it just wasn't in the slot the grader reads.
This is the honest edge: calibration removes inflated Criticals (good — most scanners over-rate by class),
but a receipt in the wrong slot gets under-rated. It is a dial toward defensibility, not a correctness
oracle — the operator still owns the final severity, and the fix is to make agents populate
evidence_data for impact, not to loosen the grader.
config.json API-key exposure (CWE-200) and a no-lockout brute-force (CWE-307) the plain run never reported — and it caught web_idor_invoice, which the plain run missed.union_search and blind_time, and under-rated the credential-dump BOLA. A confirmation layer that demands receipts will sometimes discard a real thing it couldn't re-prove in-budget.claude-opus-4-8 via subscription,
black-box, recon intensity 2, single-model vote (vote-n 1), same natural-language focus naming
the 13 endpoints, no pre-baked solver — the LLM discovered and confirmed everything live. Recall is scored by
class + endpoint keyword match against the target's ground-truth list, so a match is coverage, not a graded
proof. Confounders: the two runs are single samples, not averages; an earlier TypeSafe run collapsed to
zero when the subscription hit a session limit mid-run (a real harness gap, since fixed — session-limit stdout
now parks the run instead of burning agents); vote-n 1 means no cross-model agreement in either arm. Treat this
as one honest data point on one target, not a leaderboard. Not measured here: multi-sample variance,
higher vote-n, and TypeSafe's agent-pruning effect on a broader agent set.