A WAF breaks inference in both directions and agents make both mistakes:
a 403 from Cloudflare read as "tested, not vulnerable" (the expensive one —
the app may be wide open and simply never reached), and a block page that
echoes the payload read as reflection (the embarrassing one).
classify() answers one question: did the application see this request?
Proxy markers and enforcement markers are separate lists, because cf-ray is
on every response Cloudflare proxies — treating that as a block would
discard every finding on every CDN-fronted site, including the ordinary
authorization 403s that are often the finding itself.
Coverage::summary() says how many probes actually reached the application,
so a clean result on a WAF-fronted target cannot be read as a clean app.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
BENCHMARK.md is a capability comparison, not a scored result — and it says
so. It names the three places NeuroSploit is genuinely behind (no container
isolation, no intercepting proxy, no benchmark anyone has actually run) as
plainly as the places it is ahead.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>