Files
NeuroSploit/BENCHMARK.md
T
CyberSecurityUPandClaude Opus 5 fce86522ca feat(chain,skills): close benchmark misses — CRLF-on-Location, second-order precondition; condense BENCHMARK
The 13-target benchmark left 3 misses. Root-caused and fixed the two that were
coverage gaps (the third was single-run variance, already handled by the
session-limit fix):

- CRLF header injection (web_crlf_header_go): the agent confirmed the open
  redirect on /go?url= and stopped; the CRLF payload was never generated. The
  open_redirect skill now tests %0d%0a header injection on the SAME param, and
  CHAIN_DOCTRINE says a param landing in a Location header must also be tested
  for response splitting. chain.rs: CWE-113/93/644 now provide capabilities;
  attack_graph maps their kill-chain stage.
- Second-order SQLi (web_sqli_second_order): the sink was behind /admin, which
  the customer account could not reach. CHAIN_DOCTRINE now teaches the
  precondition pattern (store the payload, trigger from every identity, escalate
  first if the trigger page needs a role you lack, else report as a chained
  lead). chain.rs: CWE-564 requires PrivilegedContext so it chains after privesc.

BENCHMARK.md: added the TypeSafe calibrated-adjudication row; dropped the
"genuinely ahead" prose (the table is the summary); condensed the rest
188 -> 89 lines; refreshed scale (27 validators, 47 modules, 383 tests).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-20 12:42:36 -03:00

4.8 KiB

NeuroSploit vs. the open-source AI pentest agents

A rough benchmark, written honestly. Last updated 14 September 2026.

This is a capability comparison, not a scored competition. Nobody in this space has published a head-to-head on a shared target set, so anyone claiming a rank order — including this document — is comparing designs, not results. Where NeuroSploit is behind, it says so. The table is the summary; the prose below is only the honest caveats.

The tools compared: Strix (Apache 2.0), Shannon (AGPLv3, Keygraph), Penligent (commercial SaaS), PentAGI and PentestGPT (open source), with XBOW as the commercial reference point.


The short version

Strix Shannon Penligent NeuroSploit
Language Python Node + Docker SaaS Rust (+ Node web console)
Black-box ✅ ⚠️ needs source ✅ ✅
White-box ✅ SAST+DAST ✅ core design ⚠️ ✅ + grey-box
Browser validation ✅ built-in ✅ ✅ ✅ Playwright, XSS proven by execution
Intercepting proxy ✅ Caido — ✅ Burp ✅ own interceptor + Burp/Caido/ZAP/mitmproxy
Container isolation ✅ ✅ ephemeral Docker ✅ ✅ Kali docker/podman (no host net, no socket)
Exploit-only reporting ✅ "working PoCs" ✅ "no exploit, no report" ✅ ⚠️ different rule — see below
CVSS tag on the finding not scored ✅ ✅ evidence-graded, computed not guessed
Multi-model adversarial vote — — — ✅
Signed authorization (capability tokens) — — — ✅
Hash-chained audit trail — — — ✅
OT/SCADA/ICS safety policy — — — ✅
Internal network / AD attack graph — — — ✅
Self-hosted OOB channel (blind SSRF/XXE/RCE) via tools — ✅ Burp ✅ own DNS+HTTP listeners
Fail-closed egress (VPN/bastion/tunnel) — — — ✅
WAF-aware inference (block ≠ "not vulnerable") — — — ✅
Calibrated adjudication (TypeSafe System One) — — — ✅ evidence-graded, data-type aware
PoC re-validation (re-run, demote what's gone) — — — ✅
Compliance mapping (PCI-DSS/HIPAA/SOC 2) SOC2/ISO/PCI report shapes — ✅ ✅ control-level, disclaimer enforced
Deterministic per-CWE validators — — — ✅ 27 classes
FAIR loss quantification — — — ✅
Provenance / watermarking — — — ✅
Published benchmark results dir exists, empty — marketing ❌ none, including this one
Stars / adoption growing ~40k commercial small

Where NeuroSploit is behind — honestly

  • Container isolation is young. It runs commands in a Kali docker/podman container (no host net, no socket, no-new-privileges), but wiring every agent-authored command through it is still partial.
  • TLS interception delegates to the tools. The own interceptor records HTTP fully and tunnels HTTPS honestly; decrypted HTTPS chains to Burp/Caido/ZAP.
  • No cross-tool benchmark. The only run published here is with/without TypeSafe on one target. This document is not evidence of comparative performance.
  • Adoption. Shannon has ~40k stars and a company; sharp edges get found by users.
  • Exploit-dev ergonomics. Strix's interactive Python PoC sandbox is nicer than agent-authored scripts.

So: Strix or NeuroSploit?

Different halves of the problem. Strix optimises finding things and is the more finished product to hand someone today. NeuroSploit optimises being able to defend what you reported: signed scope, per-action audit, a recomputable CVSS, findings neither silently dropped nor inflated, OT rules in code, and now TypeSafe calibrated adjudication. Better in front of a client's legal and compliance team; still closing the isolation and cross-tool-benchmark gaps.

Current scale

Agents / skills 446 (255 vulnerability, plus recon, code, infra, AI, chains, meta)
Deterministic validators 27 CWE classes with evidence preconditions
Rust modules 47
Rust LOC ~24k
Tests 383, all passing

Next, to make this a real benchmark

  1. Ephemeral container execution (closes the largest gap).
  2. Own the request stream — a real intercepting proxy.
  3. Run all four tools against a fixed target set (Juice Shop, WebGoat, a deliberately vulnerable API, one real authorized scope) and publish: true positives, false positives, time, and cost per finding.
  4. Publish the CVSS deltas — where the evidence-graded score differs from the by-class score, and which one the target's own team agreed with.