The 13-target benchmark left 3 misses. Root-caused and fixed the two that were
coverage gaps (the third was single-run variance, already handled by the
session-limit fix):
- CRLF header injection (web_crlf_header_go): the agent confirmed the open
redirect on /go?url= and stopped; the CRLF payload was never generated. The
open_redirect skill now tests %0d%0a header injection on the SAME param, and
CHAIN_DOCTRINE says a param landing in a Location header must also be tested
for response splitting. chain.rs: CWE-113/93/644 now provide capabilities;
attack_graph maps their kill-chain stage.
- Second-order SQLi (web_sqli_second_order): the sink was behind /admin, which
the customer account could not reach. CHAIN_DOCTRINE now teaches the
precondition pattern (store the payload, trigger from every identity, escalate
first if the trigger page needs a role you lack, else report as a chained
lead). chain.rs: CWE-564 requires PrivilegedContext so it chains after privesc.
BENCHMARK.md: added the TypeSafe calibrated-adjudication row; dropped the
"genuinely ahead" prose (the table is the summary); condensed the rest
188 -> 89 lines; refreshed scale (27 validators, 47 modules, 383 tests).
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Closes the three benchmark gaps and adds the two the user asked for.
poc.rs — re-runs each finding's recorded proof and sorts it into reproduced /
changed / gone / unverifiable. The last two are kept apart deliberately: a PoC
that could not be tested (out of scope now, state-changing, nothing recorded)
is never reported as one that failed. Never re-runs a mutating request to
"confirm" it. Can only lower a finding's standing, never raise it. Wired as a
run pass (--revalidate-poc) and a subcommand (neurosploit poc <run> --apply).
proxy.rs — own recording forward proxy (HTTP in full; HTTPS tunnelled with
honest metadata, no fake CA) that chains upstream to Burp / Caido / ZAP /
mitmproxy. A bare tool routes straight through it; own+tool records here and
forwards for full TLS interception. Flows -> flows.jsonl, distinct hosts become
passive-discovery leads. Harness and agent child commands share one route.
sandbox.rs — Kali docker/podman container: no host network, no mounted socket,
no-new-privileges, workdir mounted, proxy/transport env inherited. A missing
runtime is an explicit error, never a silent fallback to host execution — the
whole point being to keep attack payloads off the operator's host. Subcommands
sandbox up|exec|install|down.
compliance.rs — maps confirmed findings onto PCI-DSS v4.0, HIPAA Security Rule
and SOC 2 controls. Phrased as "bears on control X", never "compliant/non-
compliant"; the disclaimer is rendered on top and absence of a finding is never
presented as compliance. Report section + `neurosploit compliance <run>`.
validation.rs — 8 new deterministic validators (19 -> 27 classes): verbose
errors/stack traces (CWE-209), cleartext/HSTS (319), CRLF response splitting
(113), dangerous HTTP methods (650), GraphQL introspection, exposed backup
files (530), Host header injection (644), cacheable private responses (525).
Each names exactly what it saw and rejects the classic false positives (a
block page echoing a payload, the SPA served under a bogus path, a copyright
year mistaken for a code).
All wired through RunConfig, the CLI (global --intercept/--sandbox; run-level
--revalidate-poc/--compliance) and the web console's Tooling & assurance block.
328 tests.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
A WAF breaks inference in both directions and agents make both mistakes:
a 403 from Cloudflare read as "tested, not vulnerable" (the expensive one —
the app may be wide open and simply never reached), and a block page that
echoes the payload read as reflection (the embarrassing one).
classify() answers one question: did the application see this request?
Proxy markers and enforcement markers are separate lists, because cf-ray is
on every response Cloudflare proxies — treating that as a block would
discard every finding on every CDN-fronted site, including the ordinary
authorization 403s that are often the finding itself.
Coverage::summary() says how many probes actually reached the application,
so a clean result on a WAF-fronted target cannot be read as a clean app.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
BENCHMARK.md is a capability comparison, not a scored result — and it says
so. It names the three places NeuroSploit is genuinely behind (no container
isolation, no intercepting proxy, no benchmark anyone has actually run) as
plainly as the places it is ahead.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>