feat: PoC validator, Kali sandbox, intercept proxy, compliance, +8 validators

Closes the three benchmark gaps and adds the two the user asked for.

poc.rs — re-runs each finding's recorded proof and sorts it into reproduced /
changed / gone / unverifiable. The last two are kept apart deliberately: a PoC
that could not be tested (out of scope now, state-changing, nothing recorded)
is never reported as one that failed. Never re-runs a mutating request to
"confirm" it. Can only lower a finding's standing, never raise it. Wired as a
run pass (--revalidate-poc) and a subcommand (neurosploit poc <run> --apply).

proxy.rs — own recording forward proxy (HTTP in full; HTTPS tunnelled with
honest metadata, no fake CA) that chains upstream to Burp / Caido / ZAP /
mitmproxy. A bare tool routes straight through it; own+tool records here and
forwards for full TLS interception. Flows -> flows.jsonl, distinct hosts become
passive-discovery leads. Harness and agent child commands share one route.

sandbox.rs — Kali docker/podman container: no host network, no mounted socket,
no-new-privileges, workdir mounted, proxy/transport env inherited. A missing
runtime is an explicit error, never a silent fallback to host execution — the
whole point being to keep attack payloads off the operator's host. Subcommands
sandbox up|exec|install|down.

compliance.rs — maps confirmed findings onto PCI-DSS v4.0, HIPAA Security Rule
and SOC 2 controls. Phrased as "bears on control X", never "compliant/non-
compliant"; the disclaimer is rendered on top and absence of a finding is never
presented as compliance. Report section + `neurosploit compliance <run>`.

validation.rs — 8 new deterministic validators (19 -> 27 classes): verbose
errors/stack traces (CWE-209), cleartext/HSTS (319), CRLF response splitting
(113), dangerous HTTP methods (650), GraphQL introspection, exposed backup
files (530), Host header injection (644), cacheable private responses (525).
Each names exactly what it saw and rejects the classic false positives (a
block page echoing a payload, the SPA served under a bogus path, a copyright
year mistaken for a code).

All wired through RunConfig, the CLI (global --intercept/--sandbox; run-level
--revalidate-poc/--compliance) and the web console's Tooling & assurance block.
328 tests.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
CyberSecurityUPandClaude Opus 5 committed 2026-09-14 10:26:06 -03:00
1 parent 64d6efa8c3
commit f1fb6b8bc7
16 files changed
+2528 -12

No files matched your search

+16 -10
View File
@@ -24,8 +24,8 @@ The tools compared: [Strix](https://github.com/usestrix/strix) (Apache 2.0),
| Black-box | ✅ | ⚠️ needs source | ✅ | ✅ |
| White-box | ✅ SAST+DAST | ✅ core design | ⚠️ | ✅ + grey-box |
| Browser validation | ✅ built-in | ✅ | ✅ | ✅ Playwright, XSS proven by execution |
| Intercepting proxy | ✅ Caido | — | ✅ Burp | ⚠️ upstream proxy only |
| Container isolation | ✅ | ✅ ephemeral Docker | ✅ | ❌ **runs on the host** |
| Intercepting proxy | ✅ Caido | — | ✅ Burp | ✅ own interceptor + Burp/Caido/ZAP/mitmproxy |
| Container isolation | ✅ | ✅ ephemeral Docker | ✅ | ✅ Kali docker/podman (no host net, no socket) |
| Exploit-only reporting | ✅ "working PoCs" | ✅ "no exploit, no report" | ✅ | ⚠️ **different rule — see below** |
| CVSS | tag on the finding | not scored | ✅ | ✅ **evidence-graded, computed not guessed** |
| Multi-model adversarial vote | — | — | — | ✅ |
@@ -36,6 +36,9 @@ The tools compared: [Strix](https://github.com/usestrix/strix) (Apache 2.0),
| Self-hosted OOB channel (blind SSRF/XXE/RCE) | via tools | — | ✅ Burp | ✅ own DNS+HTTP listeners |
| Fail-closed egress (VPN/bastion/tunnel) | — | — | — | ✅ |
| WAF-aware inference (block ≠ "not vulnerable") | — | — | — | ✅ |
| PoC re-validation (re-run, demote what's gone) | — | — | — | ✅ |
| Compliance mapping (PCI-DSS/HIPAA/SOC 2) | SOC2/ISO/PCI report shapes | — | ✅ | ✅ control-level, disclaimer enforced |
| Deterministic per-CWE validators | — | — | — | ✅ 27 classes |
| FAIR loss quantification | — | — | — | ✅ |
| Provenance / watermarking | — | — | — | ✅ |
| Published benchmark results | dir exists, empty | — | marketing | ❌ **none, including this one** |
@@ -109,15 +112,18 @@ stopped". Long engagements die of token exhaustion more often than of bugs.
## Where NeuroSploit is behind — honestly
**1. No container isolation.** Strix and Shannon run each scan in an ephemeral
container. NeuroSploit runs on the operator's host. For a tool that executes
attacker-supplied-shaped payloads this is the largest single gap in the
comparison, and the next thing worth building.
**1. Container isolation is new and shallow.** NeuroSploit now runs commands in
a Kali docker/podman container (no host network, no mounted socket,
`no-new-privileges`), which closes the headline gap — but Strix and Shannon
have run this way from day one and have found the sharp edges. Ours is young.
And wiring *every* agent-authored command through the container (versus the
harness's own tool commands) is still partial.
**2. No real intercepting proxy.** Strix ships Caido integration; Penligent
drives Burp. NeuroSploit can route through an upstream proxy — and now through
a VPN, bastion, or Cloudflare tunnel, fail-closed — but it does not own the
request/response stream, which limits replay fidelity and passive discovery.
**2. TLS interception delegates to the tools.** The own interceptor records
plaintext HTTP fully and tunnels HTTPS honestly (host, timing, byte counts) —
for decrypted HTTPS it chains to Burp/Caido/ZAP/mitmproxy, which own the CA
machinery. That is a deliberate honesty split, not a full re-implementation of
what those tools do.
**3. Nobody has run it against a benchmark.** Strix has an empty `benchmarks/`
directory, Shannon publishes none, and neither does this project. Until