mirror of
https://github.com/CyberSecurityUP/NeuroSploit.git
synced 2026-09-29 04:21:44 +02:00
feat(chain,skills): close benchmark misses — CRLF-on-Location, second-order precondition; condense BENCHMARK
The 13-target benchmark left 3 misses. Root-caused and fixed the two that were coverage gaps (the third was single-run variance, already handled by the session-limit fix): - CRLF header injection (web_crlf_header_go): the agent confirmed the open redirect on /go?url= and stopped; the CRLF payload was never generated. The open_redirect skill now tests %0d%0a header injection on the SAME param, and CHAIN_DOCTRINE says a param landing in a Location header must also be tested for response splitting. chain.rs: CWE-113/93/644 now provide capabilities; attack_graph maps their kill-chain stage. - Second-order SQLi (web_sqli_second_order): the sink was behind /admin, which the customer account could not reach. CHAIN_DOCTRINE now teaches the precondition pattern (store the payload, trigger from every identity, escalate first if the trigger page needs a role you lack, else report as a chained lead). chain.rs: CWE-564 requires PrivilegedContext so it chains after privesc. BENCHMARK.md: added the TypeSafe calibrated-adjudication row; dropped the "genuinely ahead" prose (the table is the summary); condensed the rest 188 -> 89 lines; refreshed scale (27 validators, 47 modules, 383 tests). Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
co-authored by
Claude Opus 5
parent
dff2e3c0f0
commit
fce86522ca
+20
-119
@@ -5,7 +5,7 @@
|
||||
This is a capability comparison, not a scored competition. Nobody in this
|
||||
space has published a head-to-head on a shared target set, so anyone claiming
|
||||
a rank order — including this document — is comparing designs, not results.
|
||||
Where NeuroSploit is behind, it says so.
|
||||
Where NeuroSploit is behind, it says so. The table is the summary; the prose below is only the honest caveats.
|
||||
|
||||
The tools compared: [Strix](https://github.com/usestrix/strix) (Apache 2.0),
|
||||
[Shannon](https://github.com/KeygraphHQ/shannon) (AGPLv3, Keygraph),
|
||||
@@ -36,6 +36,7 @@ The tools compared: [Strix](https://github.com/usestrix/strix) (Apache 2.0),
|
||||
| Self-hosted OOB channel (blind SSRF/XXE/RCE) | via tools | — | ✅ Burp | ✅ own DNS+HTTP listeners |
|
||||
| Fail-closed egress (VPN/bastion/tunnel) | — | — | — | ✅ |
|
||||
| WAF-aware inference (block ≠ "not vulnerable") | — | — | — | ✅ |
|
||||
| Calibrated adjudication (TypeSafe System One) | — | — | — | ✅ evidence-graded, data-type aware |
|
||||
| PoC re-validation (re-run, demote what's gone) | — | — | — | ✅ |
|
||||
| Compliance mapping (PCI-DSS/HIPAA/SOC 2) | SOC2/ISO/PCI report shapes | — | ✅ | ✅ control-level, disclaimer enforced |
|
||||
| Deterministic per-CWE validators | — | — | — | ✅ 27 classes |
|
||||
@@ -46,136 +47,36 @@ The tools compared: [Strix](https://github.com/usestrix/strix) (Apache 2.0),
|
||||
|
||||
---
|
||||
|
||||
## Where NeuroSploit is genuinely ahead
|
||||
|
||||
**1. Evidence is a first-class object, not a field on a finding.**
|
||||
Every claim carries an evidence ledger (`E01`, `E02`, …), and a claim's
|
||||
asserted status can never outrun its citations. A finding whose impact loses
|
||||
its evidence is not deleted — it is *rewritten* down to the mechanic that
|
||||
survived, and only rejected if nothing security-relevant is left:
|
||||
|
||||
```rust
|
||||
if remove_unproven_impact(f).still_security_relevant() { retain_and_rewrite() }
|
||||
else { reject() }
|
||||
```
|
||||
|
||||
Strix and Shannon both take the simpler rule — no exploit, no report. That is
|
||||
a good rule and it produces clean reports, but it throws away the middle
|
||||
ground, and the middle ground is where most real engagements live: a
|
||||
rate-limit failure you measured but could not chain, a credential path you
|
||||
proved up to the authenticated surface. NeuroSploit keeps those, downgraded
|
||||
and labelled, instead of discarding them or inflating them.
|
||||
|
||||
**2. CVSS is computed, not asked for.**
|
||||
The model proposes metrics and must point each one at evidence; a
|
||||
deterministic calculator produces the number; a demonstrated-impact ladder
|
||||
caps it (*reached* < *read data* < *wrote data* < *RCE* < *crossed systems*).
|
||||
So SQL injection without extraction lands Medium/High and the same class with
|
||||
a sensitive table read lands High/Critical — by class it would be Critical
|
||||
every time, which is how scanners produce reports nobody believes.
|
||||
|
||||
**3. Authorization is enforced in code, not in a prompt.**
|
||||
Scope is a signed capability token (HMAC, expiry, max action, risk ceiling)
|
||||
that acts as a ceiling nothing in-session can widen — a bug we found and fixed
|
||||
when `/inscope` managed to widen scope past its own grant. Every action lands
|
||||
in a hash-chained audit log. No other tool on this list has an answer for
|
||||
"prove the agent stayed inside what the client authorized" beyond "we told it
|
||||
to".
|
||||
|
||||
**4. OT/SCADA/ICS is modelled, not banned.**
|
||||
`effective_risk = action_risk + asset_criticality + protocol_risk +
|
||||
privilege_level + blast_radius`, scaled by environment. The OT profile forbids
|
||||
write/disruptive *action kinds* and specific industrial function codes
|
||||
(Modbus 5/6/8/15/16/22/23/43, S7 0x28/0x29, DNP3 13/14/18) while still
|
||||
allowing the reads OT findings actually come from. Calibrating that took a
|
||||
real correction: our first ceiling refused a plain read of a critical PLC,
|
||||
which would have made the whole profile useless.
|
||||
|
||||
**5. Internal network and AD as a graph.**
|
||||
The layered taxonomy (Asset → Exposure → Weakness → Credential → Privilege →
|
||||
Movement → Crown Jewel, with business impact, detection and remediation on the
|
||||
**edges**) plus the credential→identity→permission→machine loop. The output
|
||||
that matters is `choke_points()`: the single edge whose removal cuts the most
|
||||
value to crown jewels. A CVSS-sorted list of 40 findings cannot answer "what
|
||||
do we fix first"; this can. The web-focused tools do not attempt this at all.
|
||||
|
||||
**6. Provenance.** Per-build fingerprint, `JOASNSCOPE` sigil on every canary,
|
||||
signed run manifests, and a structural signature that survives rewording but
|
||||
not a changed result set. Nobody else on this list can tell you whether a
|
||||
report that came back to them is theirs.
|
||||
|
||||
**7. Resilience.** Model fallback, pause on quota exhaustion with every
|
||||
finding kept, resume on a different backend, and "report from where it
|
||||
stopped". Long engagements die of token exhaustion more often than of bugs.
|
||||
|
||||
---
|
||||
|
||||
## Where NeuroSploit is behind — honestly
|
||||
|
||||
**1. Container isolation is new and shallow.** NeuroSploit now runs commands in
|
||||
a Kali docker/podman container (no host network, no mounted socket,
|
||||
`no-new-privileges`), which closes the headline gap — but Strix and Shannon
|
||||
have run this way from day one and have found the sharp edges. Ours is young.
|
||||
And wiring *every* agent-authored command through the container (versus the
|
||||
harness's own tool commands) is still partial.
|
||||
|
||||
**2. TLS interception delegates to the tools.** The own interceptor records
|
||||
plaintext HTTP fully and tunnels HTTPS honestly (host, timing, byte counts) —
|
||||
for decrypted HTTPS it chains to Burp/Caido/ZAP/mitmproxy, which own the CA
|
||||
machinery. That is a deliberate honesty split, not a full re-implementation of
|
||||
what those tools do.
|
||||
|
||||
**3. Nobody has run it against a benchmark.** Strix has an empty `benchmarks/`
|
||||
directory, Shannon publishes none, and neither does this project. Until
|
||||
NeuroSploit is run against something like a Juice Shop / DVWA / OWASP
|
||||
Benchmark suite alongside the others, every claim in the "ahead" section above
|
||||
is an argument about design. **This document is not evidence of performance.**
|
||||
|
||||
**4. Adoption.** Shannon has roughly 40k stars and a company behind it. Most
|
||||
of the sharp edges in a security tool are found by other people using it.
|
||||
|
||||
**5. Exploit-development ergonomics.** Strix's Python sandbox for writing PoCs
|
||||
interactively is better developer experience than our agent-authored scripts.
|
||||
|
||||
**6. Compliance report templates.** Strix advertises SOC 2 / ISO 27001 / PCI
|
||||
DSS report shapes. Ours is one (good) template.
|
||||
|
||||
---
|
||||
- **Container isolation is young.** It runs commands in a Kali docker/podman
|
||||
container (no host net, no socket, `no-new-privileges`), but wiring *every*
|
||||
agent-authored command through it is still partial.
|
||||
- **TLS interception delegates to the tools.** The own interceptor records HTTP
|
||||
fully and tunnels HTTPS honestly; decrypted HTTPS chains to Burp/Caido/ZAP.
|
||||
- **No cross-tool benchmark.** The only run published here is with/without
|
||||
TypeSafe on one target. This document is not evidence of comparative performance.
|
||||
- **Adoption.** Shannon has ~40k stars and a company; sharp edges get found by users.
|
||||
- **Exploit-dev ergonomics.** Strix's interactive Python PoC sandbox is nicer than agent-authored scripts.
|
||||
|
||||
## So: Strix or NeuroSploit?
|
||||
|
||||
**If you want a well-packaged autonomous scanner today**, with container
|
||||
isolation, a proxy, a Python exploit sandbox and compliance report templates —
|
||||
Strix is the more finished product, and its team is shipping.
|
||||
|
||||
**If the engagement has to withstand scrutiny** — a signed scope you can prove
|
||||
you stayed inside, an audit trail per action, a CVSS number someone can
|
||||
recompute from the evidence, findings that were not silently dropped or
|
||||
silently inflated, and OT rules that are enforced by code — NeuroSploit is
|
||||
built for that and Strix is not attempting it.
|
||||
|
||||
They are aimed at different halves of the problem. Strix optimises *finding
|
||||
things*; NeuroSploit optimises *being able to defend what you reported*. A
|
||||
harness that finds ten bugs and cannot show its work is not obviously better
|
||||
than one that finds six and can.
|
||||
|
||||
The honest summary: **Strix is the better tool to hand someone today;
|
||||
NeuroSploit is the better tool to put in front of a client's legal and
|
||||
compliance team.** Closing the isolation and proxy gaps, then publishing a
|
||||
real benchmark run, is what would make that a comparison of results instead of
|
||||
a comparison of intentions.
|
||||
|
||||
---
|
||||
Different halves of the problem. Strix optimises *finding things* and is the
|
||||
more finished product to hand someone today. NeuroSploit optimises *being able
|
||||
to defend what you reported*: signed scope, per-action audit, a recomputable
|
||||
CVSS, findings neither silently dropped nor inflated, OT rules in code, and now
|
||||
TypeSafe calibrated adjudication. Better in front of a client's legal and
|
||||
compliance team; still closing the isolation and cross-tool-benchmark gaps.
|
||||
|
||||
## Current scale
|
||||
|
||||
| | |
|
||||
|---|---|
|
||||
| Agents / skills | 446 (255 vulnerability, plus recon, code, infra, AI, chains, meta) |
|
||||
| Deterministic validators | 22 CWE classes with evidence preconditions |
|
||||
| Rust modules | 37 |
|
||||
| Deterministic validators | 27 CWE classes with evidence preconditions |
|
||||
| Rust modules | 47 |
|
||||
| Rust LOC | ~24k |
|
||||
| Tests | 296, all passing |
|
||||
| Tests | 383, all passing |
|
||||
|
||||
## Next, to make this a real benchmark
|
||||
|
||||
|
||||
Reference in New Issue
Block a user