From fce86522cad5f37f873560ddfdcaecdfffe56ff0 Mon Sep 17 00:00:00 2001 From: CyberSecurityUP Date: Sun, 20 Sep 2026 12:42:36 -0300 Subject: [PATCH] =?UTF-8?q?feat(chain,skills):=20close=20benchmark=20misse?= =?UTF-8?q?s=20=E2=80=94=20CRLF-on-Location,=20second-order=20precondition?= =?UTF-8?q?;=20condense=20BENCHMARK?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit The 13-target benchmark left 3 misses. Root-caused and fixed the two that were coverage gaps (the third was single-run variance, already handled by the session-limit fix): - CRLF header injection (web_crlf_header_go): the agent confirmed the open redirect on /go?url= and stopped; the CRLF payload was never generated. The open_redirect skill now tests %0d%0a header injection on the SAME param, and CHAIN_DOCTRINE says a param landing in a Location header must also be tested for response splitting. chain.rs: CWE-113/93/644 now provide capabilities; attack_graph maps their kill-chain stage. - Second-order SQLi (web_sqli_second_order): the sink was behind /admin, which the customer account could not reach. CHAIN_DOCTRINE now teaches the precondition pattern (store the payload, trigger from every identity, escalate first if the trigger page needs a role you lack, else report as a chained lead). chain.rs: CWE-564 requires PrivilegedContext so it chains after privesc. BENCHMARK.md: added the TypeSafe calibrated-adjudication row; dropped the "genuinely ahead" prose (the table is the summary); condensed the rest 188 -> 89 lines; refreshed scale (27 validators, 47 modules, 383 tests). Co-Authored-By: Claude Opus 5 (1M context) --- BENCHMARK.md | 139 +++--------------- agents_md/vulns/open_redirect.md | 7 + .../crates/harness/src/attack_graph.rs | 3 + neurosploit-rs/crates/harness/src/chain.rs | 6 + neurosploit-rs/crates/harness/src/pipeline.rs | 2 + 5 files changed, 38 insertions(+), 119 deletions(-) diff --git a/BENCHMARK.md b/BENCHMARK.md index d33855b..def136f 100644 --- a/BENCHMARK.md +++ b/BENCHMARK.md @@ -5,7 +5,7 @@ This is a capability comparison, not a scored competition. Nobody in this space has published a head-to-head on a shared target set, so anyone claiming a rank order — including this document — is comparing designs, not results. -Where NeuroSploit is behind, it says so. +Where NeuroSploit is behind, it says so. The table is the summary; the prose below is only the honest caveats. The tools compared: [Strix](https://github.com/usestrix/strix) (Apache 2.0), [Shannon](https://github.com/KeygraphHQ/shannon) (AGPLv3, Keygraph), @@ -36,6 +36,7 @@ The tools compared: [Strix](https://github.com/usestrix/strix) (Apache 2.0), | Self-hosted OOB channel (blind SSRF/XXE/RCE) | via tools | — | ✅ Burp | ✅ own DNS+HTTP listeners | | Fail-closed egress (VPN/bastion/tunnel) | — | — | — | ✅ | | WAF-aware inference (block ≠ "not vulnerable") | — | — | — | ✅ | +| Calibrated adjudication (TypeSafe System One) | — | — | — | ✅ evidence-graded, data-type aware | | PoC re-validation (re-run, demote what's gone) | — | — | — | ✅ | | Compliance mapping (PCI-DSS/HIPAA/SOC 2) | SOC2/ISO/PCI report shapes | — | ✅ | ✅ control-level, disclaimer enforced | | Deterministic per-CWE validators | — | — | — | ✅ 27 classes | @@ -46,136 +47,36 @@ The tools compared: [Strix](https://github.com/usestrix/strix) (Apache 2.0), --- -## Where NeuroSploit is genuinely ahead - -**1. Evidence is a first-class object, not a field on a finding.** -Every claim carries an evidence ledger (`E01`, `E02`, …), and a claim's -asserted status can never outrun its citations. A finding whose impact loses -its evidence is not deleted — it is *rewritten* down to the mechanic that -survived, and only rejected if nothing security-relevant is left: - -```rust -if remove_unproven_impact(f).still_security_relevant() { retain_and_rewrite() } -else { reject() } -``` - -Strix and Shannon both take the simpler rule — no exploit, no report. That is -a good rule and it produces clean reports, but it throws away the middle -ground, and the middle ground is where most real engagements live: a -rate-limit failure you measured but could not chain, a credential path you -proved up to the authenticated surface. NeuroSploit keeps those, downgraded -and labelled, instead of discarding them or inflating them. - -**2. CVSS is computed, not asked for.** -The model proposes metrics and must point each one at evidence; a -deterministic calculator produces the number; a demonstrated-impact ladder -caps it (*reached* < *read data* < *wrote data* < *RCE* < *crossed systems*). -So SQL injection without extraction lands Medium/High and the same class with -a sensitive table read lands High/Critical — by class it would be Critical -every time, which is how scanners produce reports nobody believes. - -**3. Authorization is enforced in code, not in a prompt.** -Scope is a signed capability token (HMAC, expiry, max action, risk ceiling) -that acts as a ceiling nothing in-session can widen — a bug we found and fixed -when `/inscope` managed to widen scope past its own grant. Every action lands -in a hash-chained audit log. No other tool on this list has an answer for -"prove the agent stayed inside what the client authorized" beyond "we told it -to". - -**4. OT/SCADA/ICS is modelled, not banned.** -`effective_risk = action_risk + asset_criticality + protocol_risk + -privilege_level + blast_radius`, scaled by environment. The OT profile forbids -write/disruptive *action kinds* and specific industrial function codes -(Modbus 5/6/8/15/16/22/23/43, S7 0x28/0x29, DNP3 13/14/18) while still -allowing the reads OT findings actually come from. Calibrating that took a -real correction: our first ceiling refused a plain read of a critical PLC, -which would have made the whole profile useless. - -**5. Internal network and AD as a graph.** -The layered taxonomy (Asset → Exposure → Weakness → Credential → Privilege → -Movement → Crown Jewel, with business impact, detection and remediation on the -**edges**) plus the credential→identity→permission→machine loop. The output -that matters is `choke_points()`: the single edge whose removal cuts the most -value to crown jewels. A CVSS-sorted list of 40 findings cannot answer "what -do we fix first"; this can. The web-focused tools do not attempt this at all. - -**6. Provenance.** Per-build fingerprint, `JOASNSCOPE` sigil on every canary, -signed run manifests, and a structural signature that survives rewording but -not a changed result set. Nobody else on this list can tell you whether a -report that came back to them is theirs. - -**7. Resilience.** Model fallback, pause on quota exhaustion with every -finding kept, resume on a different backend, and "report from where it -stopped". Long engagements die of token exhaustion more often than of bugs. - ---- - ## Where NeuroSploit is behind — honestly -**1. Container isolation is new and shallow.** NeuroSploit now runs commands in -a Kali docker/podman container (no host network, no mounted socket, -`no-new-privileges`), which closes the headline gap — but Strix and Shannon -have run this way from day one and have found the sharp edges. Ours is young. -And wiring *every* agent-authored command through the container (versus the -harness's own tool commands) is still partial. - -**2. TLS interception delegates to the tools.** The own interceptor records -plaintext HTTP fully and tunnels HTTPS honestly (host, timing, byte counts) — -for decrypted HTTPS it chains to Burp/Caido/ZAP/mitmproxy, which own the CA -machinery. That is a deliberate honesty split, not a full re-implementation of -what those tools do. - -**3. Nobody has run it against a benchmark.** Strix has an empty `benchmarks/` -directory, Shannon publishes none, and neither does this project. Until -NeuroSploit is run against something like a Juice Shop / DVWA / OWASP -Benchmark suite alongside the others, every claim in the "ahead" section above -is an argument about design. **This document is not evidence of performance.** - -**4. Adoption.** Shannon has roughly 40k stars and a company behind it. Most -of the sharp edges in a security tool are found by other people using it. - -**5. Exploit-development ergonomics.** Strix's Python sandbox for writing PoCs -interactively is better developer experience than our agent-authored scripts. - -**6. Compliance report templates.** Strix advertises SOC 2 / ISO 27001 / PCI -DSS report shapes. Ours is one (good) template. - ---- +- **Container isolation is young.** It runs commands in a Kali docker/podman + container (no host net, no socket, `no-new-privileges`), but wiring *every* + agent-authored command through it is still partial. +- **TLS interception delegates to the tools.** The own interceptor records HTTP + fully and tunnels HTTPS honestly; decrypted HTTPS chains to Burp/Caido/ZAP. +- **No cross-tool benchmark.** The only run published here is with/without + TypeSafe on one target. This document is not evidence of comparative performance. +- **Adoption.** Shannon has ~40k stars and a company; sharp edges get found by users. +- **Exploit-dev ergonomics.** Strix's interactive Python PoC sandbox is nicer than agent-authored scripts. ## So: Strix or NeuroSploit? -**If you want a well-packaged autonomous scanner today**, with container -isolation, a proxy, a Python exploit sandbox and compliance report templates — -Strix is the more finished product, and its team is shipping. - -**If the engagement has to withstand scrutiny** — a signed scope you can prove -you stayed inside, an audit trail per action, a CVSS number someone can -recompute from the evidence, findings that were not silently dropped or -silently inflated, and OT rules that are enforced by code — NeuroSploit is -built for that and Strix is not attempting it. - -They are aimed at different halves of the problem. Strix optimises *finding -things*; NeuroSploit optimises *being able to defend what you reported*. A -harness that finds ten bugs and cannot show its work is not obviously better -than one that finds six and can. - -The honest summary: **Strix is the better tool to hand someone today; -NeuroSploit is the better tool to put in front of a client's legal and -compliance team.** Closing the isolation and proxy gaps, then publishing a -real benchmark run, is what would make that a comparison of results instead of -a comparison of intentions. - ---- +Different halves of the problem. Strix optimises *finding things* and is the +more finished product to hand someone today. NeuroSploit optimises *being able +to defend what you reported*: signed scope, per-action audit, a recomputable +CVSS, findings neither silently dropped nor inflated, OT rules in code, and now +TypeSafe calibrated adjudication. Better in front of a client's legal and +compliance team; still closing the isolation and cross-tool-benchmark gaps. ## Current scale | | | |---|---| | Agents / skills | 446 (255 vulnerability, plus recon, code, infra, AI, chains, meta) | -| Deterministic validators | 22 CWE classes with evidence preconditions | -| Rust modules | 37 | +| Deterministic validators | 27 CWE classes with evidence preconditions | +| Rust modules | 47 | | Rust LOC | ~24k | -| Tests | 296, all passing | +| Tests | 383, all passing | ## Next, to make this a real benchmark diff --git a/agents_md/vulns/open_redirect.md b/agents_md/vulns/open_redirect.md index 25d1e0a..a577ce4 100644 --- a/agents_md/vulns/open_redirect.md +++ b/agents_md/vulns/open_redirect.md @@ -20,6 +20,13 @@ You are testing **{target}** for Open Redirect vulnerabilities. - Follow the redirect chain manually - Check if Location header points to external domain - Verify the browser actually navigates to evil.com +### 3b. Same param, test CRLF / header injection +A parameter that lands in the `Location` header is also a response-splitting +sink. On the SAME parameter, try: +- `/go?url=/%0d%0aX-Injected:%20pwned` — look for `X-Injected: pwned` as a real response header +- `/go?url=/%0d%0aSet-Cookie:%20session=attacker` — a planted cookie header +- If the marker appears as a HEADER (not the body), that is CRLF injection (CWE-113), report it IN ADDITION to the open redirect. Never stop at the redirect. + ### 4. Chain with Other Vulns - OAuth token theft via redirect_uri manipulation - Phishing: redirect from trusted domain to fake login diff --git a/neurosploit-rs/crates/harness/src/attack_graph.rs b/neurosploit-rs/crates/harness/src/attack_graph.rs index 5870c29..dddb839 100644 --- a/neurosploit-rs/crates/harness/src/attack_graph.rs +++ b/neurosploit-rs/crates/harness/src/attack_graph.rs @@ -45,6 +45,9 @@ fn map_cwe(cwe: &str) -> (&'static str, &'static str, &'static str) { // Session fixation. 384 => ("A07:2021-Auth-Failures", "T1539", "credential-access"), 601 => ("A01:2021-Broken-Access-Control", "T1566", "initial-access"), + 113 | 93 => ("A03:2021-Injection", "T1557", "initial-access"), + 644 => ("A03:2021-Injection", "T1557", "initial-access"), + 564 => ("A03:2021-Injection", "T1190", "execution"), 352 => ("A01:2021-Broken-Access-Control", "T1189", "execution"), 434 => ("A04:2021-Insecure-Design", "T1505.003", "execution"), 1321 | 915 => ("A08:2021-Software-Data-Integrity", "T1059", "execution"), diff --git a/neurosploit-rs/crates/harness/src/chain.rs b/neurosploit-rs/crates/harness/src/chain.rs index d61234c..b438c51 100644 --- a/neurosploit-rs/crates/harness/src/chain.rs +++ b/neurosploit-rs/crates/harness/src/chain.rs @@ -90,6 +90,9 @@ pub fn provides(f: &Finding) -> Vec { } } 209 | 532 | 538 | 540 | 548 | 693 | 1021 => vec![Capability::InternalKnowledge], + // CRLF / response splitting / host-header: header control feeds cache + // poisoning and redirect abuse downstream. + 113 | 93 | 644 => vec![Capability::InternalKnowledge, Capability::SessionMaterial], // Missing throttling turns any guess into an unlimited one. 307 | 770 | 799 | 400 => vec![Capability::UnlimitedAttempts], // Password policy. @@ -130,6 +133,9 @@ pub fn requires(f: &Finding) -> Vec { 614 | 1004 | 1275 => vec![Capability::SessionMaterial], // Escalation needs a foothold. 269 | 250 | 668 => vec![Capability::PrivilegedContext], + // Second-order SQLi: the stored payload only fires on the (often + // privileged) trigger page, so it needs that context to be reached. + 564 => vec![Capability::PrivilegedContext], _ => vec![], } } diff --git a/neurosploit-rs/crates/harness/src/pipeline.rs b/neurosploit-rs/crates/harness/src/pipeline.rs index e4b27c6..0c03dbe 100644 --- a/neurosploit-rs/crates/harness/src/pipeline.rs +++ b/neurosploit-rs/crates/harness/src/pipeline.rs @@ -582,6 +582,8 @@ const CHAIN_DOCTRINE: &str = "CHAIN THE FOOTHOLD (pivot to deeper, provable impa · XXE → SSRF/file read → creds; deserialization/SSTI → RCE via a gadget/template sink; prove exec with a marker.\n\ · IDOR/BOLA/mass-assignment → account/tenant takeover or role escalation (`role=admin`); open-redirect/XSS/CORS → token/session theft → ATO.\n\ · Exposed `.git`/backup/`.env`/secrets → reconstruct source & keys → auth to internal APIs, cloud, DB; default/leaked creds → domain/service compromise.\n\ + · A param that lands in a redirect/`Location` header → ALSO test CRLF/header injection on the SAME param (`%0d%0aX-Injected: pwned`, `%0d%0aSet-Cookie:`): an open redirect and response splitting share the sink, so never stop at the redirect.\n\ +- Second-order & preconditions: a payload you STORE (profile/bio/name/review/filename) may only fire on a DIFFERENT page, often a privileged one (e.g. an admin search). Plant the payload, then TRIGGER it from every identity you hold; if the trigger page needs a role you lack, FIRST look for a privesc/IDOR/mass-assign to reach it, and if none exists, report the second-order as a CHAINED lead (payload stored + trigger located, blocked only by authorization) rather than dropping it.\n\ - Reuse loot relentlessly: every credential/JWT/cookie/API key/host you obtain is input to the next step — carry it forward across modules and try it everywhere it might be accepted.\n\ - Understand the BUSINESS & LOGIC: reason about what the app is FOR (payments, orders, tenancy, KYC, entitlements) and chain toward business impact — payment/price/coupon abuse, cross-tenant data access, entitlement/limit bypass, workflow/state-machine skips (skip approval/verification steps), race conditions on balance/stock. These compound: each finding updates your model of the app for the next probe.\n\ - Stop at proof: demonstrate the impact with the SMALLEST safe step and report the CHAIN end-to-end; never destroy, overwrite, encrypt, mass-exfiltrate, or DoS to 'prove' it.\n\n";