feat(waf): tell the edge apart from the application

A WAF breaks inference in both directions and agents make both mistakes:
a 403 from Cloudflare read as "tested, not vulnerable" (the expensive one —
the app may be wide open and simply never reached), and a block page that
echoes the payload read as reflection (the embarrassing one).

classify() answers one question: did the application see this request?
Proxy markers and enforcement markers are separate lists, because cf-ray is
on every response Cloudflare proxies — treating that as a block would
discard every finding on every CDN-fronted site, including the ordinary
authorization 403s that are often the finding itself.

Coverage::summary() says how many probes actually reached the application,
so a clean result on a WAF-fronted target cannot be read as a clean app.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
CyberSecurityUP
2026-09-14 01:39:20 -03:00
co-authored by Claude Opus 5
parent 8aa776665a
commit 64d6efa8c3
4 changed files with 476 additions and 6 deletions
+8 -5
View File
@@ -33,6 +33,9 @@ The tools compared: [Strix](https://github.com/usestrix/strix) (Apache 2.0),
| Hash-chained audit trail | — | — | — | ✅ |
| OT/SCADA/ICS safety policy | — | — | — | ✅ |
| Internal network / AD attack graph | — | — | — | ✅ |
| Self-hosted OOB channel (blind SSRF/XXE/RCE) | via tools | — | ✅ Burp | ✅ own DNS+HTTP listeners |
| Fail-closed egress (VPN/bastion/tunnel) | — | — | — | ✅ |
| WAF-aware inference (block ≠ "not vulnerable") | — | — | — | ✅ |
| FAIR loss quantification | — | — | — | ✅ |
| Provenance / watermarking | — | — | — | ✅ |
| Published benchmark results | dir exists, empty | — | marketing | ❌ **none, including this one** |
@@ -112,9 +115,9 @@ attacker-supplied-shaped payloads this is the largest single gap in the
comparison, and the next thing worth building.
**2. No real intercepting proxy.** Strix ships Caido integration; Penligent
drives Burp. NeuroSploit can route through an upstream proxy but does not own
the request/response stream, which limits replay fidelity and passive
discovery.
drives Burp. NeuroSploit can route through an upstream proxy — and now through
a VPN, bastion, or Cloudflare tunnel, fail-closed — but it does not own the
request/response stream, which limits replay fidelity and passive discovery.
**3. Nobody has run it against a benchmark.** Strix has an empty `benchmarks/`
directory, Shannon publishes none, and neither does this project. Until
@@ -164,9 +167,9 @@ a comparison of intentions.
|---|---|
| Agents / skills | 446 (255 vulnerability, plus recon, code, infra, AI, chains, meta) |
| Deterministic validators | 22 CWE classes with evidence preconditions |
| Rust modules | 32 |
| Rust modules | 37 |
| Rust LOC | ~24k |
| Tests | 260, all passing |
| Tests | 296, all passing |
## Next, to make this a real benchmark
+1
View File
@@ -40,6 +40,7 @@ pub mod transport;
pub mod types;
pub mod uncertainty;
pub mod validation;
pub mod waf;
pub use agents::{Agent, Library};
pub use models::{
@@ -371,6 +371,7 @@ fn engagement_ops(cfg: &RunConfig) -> String {
};
let oob = oob_ops(cfg);
let sms = sms_ops(cfg);
let waf = WAF_OPS;
format!(
"ENGAGEMENT OPS — TEST ACCOUNTS & VAULT:\n\
- CREDENTIAL VAULT: whenever you create a test account or generate any credential, APPEND one JSON line to \
@@ -385,7 +386,7 @@ fn engagement_ops(cfg: &RunConfig) -> String {
explicit about which findings needed a login. In black-box, record in `how`/evidence exactly what you did \
to create the user.\n\
- {temp}\n\
{oob}{sms}\n"
{oob}{sms}{waf}\n"
)
}
@@ -408,6 +409,13 @@ fn oob_ops(cfg: &RunConfig) -> String {
)
}
/// What an agent must do when the edge answers instead of the application.
///
/// Both failure directions are named explicitly, because agents make both: a
/// 403 from a WAF read as "not vulnerable" (the expensive one), and a block
/// page echoing the payload read as reflection (the embarrassing one).
const WAF_OPS: &str = "- WAF / EDGE: if a response came from a CDN or WAF rather than the application (vendor headers plus block-page wording, a challenge, or HTTP 429), the application NEVER SAW your request. Never record that as 'tested, not vulnerable' — record that the control could not be reached, and say which probes were blocked. A payload echoed back by a block page is the EDGE reflecting it, not the application: it is not XSS evidence. On a 429 or a browser challenge, pace the requests and retry — that is not a verdict on the payload.\n";
/// Inbound SMS instructions, when a number is configured.
fn sms_ops(cfg: &RunConfig) -> String {
match cfg.sms.as_deref().filter(|s| !s.trim().is_empty()) {
+458
View File
@@ -0,0 +1,458 @@
//! WAF awareness — telling the edge apart from the application.
//!
//! When a WAF sits in front of a target, every response an agent reads may
//! have been written by the edge rather than by the application. That breaks
//! inference in two opposite directions at once, and both are common:
//!
//! ```text
//! payload → 403 from Cloudflare → "not vulnerable" ← false NEGATIVE
//! payload → block page echoing it → "payload reflected!" ← false POSITIVE
//! ```
//!
//! The first is the expensive one. A WAF blocking a probe says nothing about
//! the code behind it: the application may be wide open and simply never
//! reached. Reporting "SQL injection tested, not vulnerable" on the strength of
//! a 403 from the edge is a statement about the WAF, not the app — and it is
//! the statement that gets a real bug missed.
//!
//! The second is the embarrassing one. A block page frequently includes the
//! offending payload ("Your request contained: `<script>…`"), so a harness
//! looking for reflection finds it — on a page served by the CDN, under a
//! different content type, with no application involvement at all.
//!
//! So this module answers one question, [`classify`]: **did the application
//! see this request?** Everything else follows from the answer.
use serde::{Deserialize, Serialize};
/// Who wrote the response.
#[derive(Debug, Clone, Copy, PartialEq, Eq, Serialize, Deserialize)]
#[serde(rename_all = "kebab-case")]
pub enum Origin {
/// The application answered. Inference is valid.
Application,
/// The edge answered instead. The application never saw the request, so
/// nothing about the application can be concluded from this response.
EdgeBlocked,
/// The edge answered because of rate limiting or a challenge, not a rule
/// match. Also not the application — but it means "slow down", not
/// "this payload is interesting".
EdgeThrottled,
}
impl Origin {
pub fn as_str(self) -> &'static str {
match self {
Origin::Application => "application",
Origin::EdgeBlocked => "edge-blocked",
Origin::EdgeThrottled => "edge-throttled",
}
}
/// May a finding be drawn from this response?
pub fn supports_a_finding(self) -> bool {
self == Origin::Application
}
}
/// Which product, when it names itself.
#[derive(Debug, Clone, PartialEq, Eq, Serialize, Deserialize)]
pub struct Vendor {
pub name: String,
/// What gave it away — kept because a report that says "WAF detected" and
/// cannot say how is not checkable.
pub signal: String,
}
/// Vendor markers, split by what they actually tell you.
///
/// The distinction is the whole reason this is not one list. `cf-ray` says
/// "Cloudflare proxied this" and appears on every response from half the
/// internet, including the ordinary ones — it identifies the vendor and
/// nothing else. `cf-mitigated` says "Cloudflare acted on this request". Only
/// the second kind may turn a response into a block on its own; treating the
/// first that way would discard every finding on every CDN-fronted site.
struct Signature {
name: &'static str,
/// Present on normal traffic too. Identifies, never accuses.
proxy: &'static [&'static str],
/// Present only when the edge intervened.
enforcement: &'static [&'static str],
}
const SIGNATURES: &[Signature] = &[
Signature { name: "Cloudflare", proxy: &["cf-ray", "__cf_bm", "server: cloudflare"], enforcement: &["cf-mitigated", "attention required!", "error 1020", "cf-chl"] },
Signature { name: "AWS WAF", proxy: &["x-amz-cf-id", "awselb"], enforcement: &["x-amzn-waf", "x-amzn-errortype: accessdenied"] },
Signature { name: "Akamai", proxy: &["akamaighost", "x-akamai"], enforcement: &["reference #", "access denied", "akamai error"] },
Signature { name: "Imperva/Incapsula", proxy: &["incap_ses", "visid_incap", "x-iinfo"], enforcement: &["incapsula incident", "request unsuccessful"] },
Signature { name: "F5 BIG-IP ASM", proxy: &["bigipserver"], enforcement: &["x-waf-event", "the requested url was rejected"] },
Signature { name: "Sucuri", proxy: &["x-sucuri-id"], enforcement: &["sucuri website firewall", "access denied - sucuri"] },
Signature { name: "Fastly", proxy: &["x-fastly", "fastly-io"], enforcement: &["fastly error: security"] },
Signature { name: "Azure Front Door", proxy: &["x-azure-ref", "x-msedge-ref"], enforcement: &["the request is blocked"] },
Signature { name: "ModSecurity", proxy: &[], enforcement: &["mod_security", "modsecurity", "not acceptable!"] },
Signature { name: "Wordfence", proxy: &[], enforcement: &["wordfence", "generated by wordfence"] },
];
/// What the edge did with one request.
#[derive(Debug, Clone, Serialize, Deserialize)]
pub struct Verdict {
pub origin: Origin,
#[serde(skip_serializing_if = "Option::is_none")]
pub vendor: Option<Vendor>,
/// Plain-language reason, for the evidence trail.
pub reason: String,
}
impl Verdict {
fn app(reason: &str) -> Verdict {
Verdict { origin: Origin::Application, vendor: None, reason: reason.into() }
}
}
/// Decide who answered.
///
/// Deliberately conservative in one direction: when the signals are ambiguous
/// the answer is `Application`, because treating a real application response as
/// edge noise would discard true findings. A WAF verdict requires a vendor
/// signature *or* an unmistakable block shape — not merely a 403, which
/// applications return all the time for ordinary authorization failures.
pub fn classify(status: u16, headers: &[(String, String)], body: &str) -> Verdict {
let hay: String = headers
.iter()
.map(|(k, v)| format!("{}: {}", k.to_lowercase(), v.to_lowercase()))
.collect::<Vec<_>>()
.join("\n");
let body_low = body.chars().take(4096).collect::<String>().to_lowercase();
let find = |marks: &'static [&'static str]| -> Option<(&'static str, String)> {
marks.iter().find_map(|m| {
if hay.contains(m) {
Some((*m, format!("header contains `{m}`")))
} else if body_low.contains(m) {
Some((*m, format!("body contains `{m}`")))
} else {
None
}
})
};
// Enforcement markers win: a response carrying both says the edge acted.
let mut vendor: Option<Vendor> = None;
let mut enforced = false;
for sig in SIGNATURES {
if let Some((_, signal)) = find(sig.enforcement) {
vendor = Some(Vendor { name: sig.name.into(), signal });
enforced = true;
break;
}
if vendor.is_none() {
if let Some((_, signal)) = find(sig.proxy) {
vendor = Some(Vendor { name: sig.name.into(), signal });
}
}
}
// Rate limiting and challenges are their own thing: the request was not
// judged on its content, so retrying the same payload later is valid.
if status == 429 || body_low.contains("rate limit") || body_low.contains("too many requests") {
return Verdict {
origin: Origin::EdgeThrottled,
vendor,
reason: format!("HTTP {status} with rate-limit signals — slow down, this says nothing about the payload"),
};
}
if status == 503 && (body_low.contains("checking your browser") || body_low.contains("challenge") || body_low.contains("just a moment")) {
return Verdict {
origin: Origin::EdgeThrottled,
vendor,
reason: "interstitial challenge — a browser check, not a verdict on the request".into(),
};
}
let block_phrases = [
"attention required",
"request blocked",
"access denied",
"the requested url was rejected",
"your request has been blocked",
"malicious",
"web application firewall",
"security policy",
"forbidden by rule",
"error 1020",
"reference #",
];
let says_blocked = block_phrases.iter().any(|p| body_low.contains(p));
// A proxy signature alone is NOT a block. What makes a response the edge's
// rather than the application's is an enforcement marker, or block-page
// wording on a status that fits — never the mere presence of a CDN.
let blocked_status = matches!(status, 403 | 406 | 419);
if enforced || (says_blocked && blocked_status) {
let reason = match &vendor {
Some(v) if enforced => format!(
"HTTP {status}: {} acted on this request ({}) — the application never saw it",
v.name, v.signal
),
Some(v) => format!("HTTP {status} with block-page wording, {} in front ({})", v.name, v.signal),
None => format!("HTTP {status} with block-page wording but no vendor signature — an unidentified filter answered"),
};
return Verdict { origin: Origin::EdgeBlocked, vendor, reason };
}
Verdict { vendor, ..Verdict::app("application response") }
}
/// What a blocked probe means for the test that produced it.
///
/// This is the correction that matters. An edge block turns a *negative* result
/// into **no result** — the test did not run. Recording it as "tested, not
/// vulnerable" is how a WAF-fronted target ends up with a clean report over an
/// exploitable application.
pub fn reinterpret(verdict: &Verdict, payload_reflected: bool) -> Interpretation {
match verdict.origin {
Origin::Application => Interpretation {
conclusive: true,
reflection_is_meaningful: payload_reflected,
note: String::new(),
},
Origin::EdgeThrottled => Interpretation {
conclusive: false,
reflection_is_meaningful: false,
note: "throttled at the edge — retry with pacing; this is not a result".into(),
},
Origin::EdgeBlocked => Interpretation {
conclusive: false,
// A block page that echoes the payload is a reflection on the CDN's
// page, not the application's. This is the false positive.
reflection_is_meaningful: false,
note: format!(
"blocked at the edge{} — the application was NOT tested. Do not report 'not vulnerable'; report that the \
control could not be reached, and note that any payload echoed here was echoed by the block page.",
verdict.vendor.as_ref().map(|v| format!(" by {}", v.name)).unwrap_or_default()
),
},
}
}
/// What the caller should do with a response.
#[derive(Debug, Clone, PartialEq, Eq, Serialize, Deserialize)]
pub struct Interpretation {
/// Did the test actually produce a result about the application?
pub conclusive: bool,
/// Is an observed reflection evidence about the application?
pub reflection_is_meaningful: bool,
pub note: String,
}
/// Classify a recorded exchange.
///
/// The convenience the rest of the harness actually calls: replayed evidence
/// already carries status, headers and body, so the question "did the
/// application see this?" can be asked of any exchange without re-requesting.
pub fn classify_exchange(ex: &crate::validation::Exchange) -> Verdict {
let headers: Vec<(String, String)> = ex.headers.iter().map(|(k, v)| (k.clone(), v.clone())).collect();
classify(ex.status, &headers, &ex.body)
}
/// Summary of a batch of probes, for the report's coverage section.
///
/// A run against a WAF-fronted target should say how much of its testing
/// actually reached the application. "47 payloads sent" means nothing if 44
/// died at the edge, and the client deserves to know which number they are
/// looking at.
#[derive(Debug, Clone, Default, Serialize, Deserialize)]
pub struct Coverage {
pub total: usize,
pub reached_application: usize,
pub blocked: usize,
pub throttled: usize,
#[serde(skip_serializing_if = "Option::is_none")]
pub vendor: Option<String>,
}
impl Coverage {
pub fn of(verdicts: &[Verdict]) -> Coverage {
Coverage {
total: verdicts.len(),
reached_application: verdicts.iter().filter(|v| v.origin == Origin::Application).count(),
blocked: verdicts.iter().filter(|v| v.origin == Origin::EdgeBlocked).count(),
throttled: verdicts.iter().filter(|v| v.origin == Origin::EdgeThrottled).count(),
vendor: verdicts.iter().find_map(|v| v.vendor.as_ref().map(|x| x.name.clone())),
}
}
/// Fraction of probes the application actually saw.
pub fn ratio(&self) -> f64 {
if self.total == 0 {
return 1.0;
}
self.reached_application as f64 / self.total as f64
}
/// Is the run's negative space meaningful at all?
pub fn negative_results_are_meaningful(&self) -> bool {
self.ratio() >= 0.5
}
pub fn summary(&self) -> String {
if self.total == 0 {
return "no probes classified".into();
}
let vendor = self.vendor.clone().unwrap_or_else(|| "an edge filter".into());
format!(
"{}/{} probes reached the application ({:.0}%); {} blocked by {}, {} throttled.{}",
self.reached_application,
self.total,
self.ratio() * 100.0,
self.blocked,
vendor,
self.throttled,
if self.negative_results_are_meaningful() {
""
} else {
" Most testing never reached the application — the absence of findings here is NOT evidence the application is secure."
}
)
}
}
#[cfg(test)]
mod tests {
use super::*;
fn h(pairs: &[(&str, &str)]) -> Vec<(String, String)> {
pairs.iter().map(|(k, v)| (k.to_string(), v.to_string())).collect()
}
#[test]
fn an_ordinary_cloudflare_fronted_response_is_still_the_application() {
// cf-ray is on every response Cloudflare proxies. If that alone read as
// "blocked", every finding on half the internet would be discarded.
let v = classify(200, &h(&[("cf-ray", "8a1b2c3d4e5f"), ("server", "cloudflare")]), "<html>Welcome back, Ana</html>");
assert_eq!(v.origin, Origin::Application);
assert_eq!(v.vendor.as_ref().map(|x| x.name.as_str()), Some("Cloudflare"));
assert!(v.origin.supports_a_finding());
}
#[test]
fn a_cloudflare_fronted_application_403_is_still_the_application() {
// Real sites behind Cloudflare carry cf-ray on EVERY response. An
// ordinary authorization 403 must not be read as a WAF block, or the
// access-control findings — the ones worth having — all disappear.
let v = classify(
403,
&h(&[("cf-ray", "8a1b2c"), ("server", "cloudflare"), ("content-type", "application/json")]),
r#"{"error":"forbidden","detail":"admin role required"}"#,
);
assert_eq!(v.origin, Origin::Application, "reason: {}", v.reason);
assert_eq!(v.vendor.as_ref().map(|x| x.name.as_str()), Some("Cloudflare"));
}
#[test]
fn an_enforcement_header_is_enough_on_its_own() {
// cf-mitigated only appears when Cloudflare acted, so it does not need
// corroborating wording — and a JSON-shaped body must not hide it.
let v = classify(403, &h(&[("cf-ray", "8a1b"), ("cf-mitigated", "challenge")]), "{}");
assert_eq!(v.origin, Origin::EdgeBlocked);
assert!(v.reason.contains("acted on this request"));
}
#[test]
fn an_application_403_is_not_a_waf_block() {
// Applications return 403 for ordinary authorization failures all day.
let v = classify(403, &h(&[("content-type", "application/json")]), r#"{"error":"forbidden","detail":"role required"}"#);
assert_eq!(v.origin, Origin::Application, "a plain 403 is an app decision, and often the finding itself");
}
#[test]
fn a_real_block_is_recognised_and_turns_a_negative_into_no_result() {
let v = classify(
403,
&h(&[("cf-ray", "8a1b"), ("server", "cloudflare")]),
"<html><h1>Attention Required! | Cloudflare</h1><p>You have been blocked.</p></html>",
);
assert_eq!(v.origin, Origin::EdgeBlocked);
let i = reinterpret(&v, false);
assert!(!i.conclusive, "a blocked probe is not a test of the application");
assert!(i.note.contains("NOT tested"));
assert!(i.note.contains("Cloudflare"));
}
#[test]
fn a_block_page_echoing_the_payload_is_not_reflection() {
// The classic false positive: the CDN's block page quotes the payload
// back, and a reflection check finds it.
let body = "<h1>Request Blocked</h1><p>Your request contained: &lt;script&gt;alert(1)&lt;/script&gt;</p>";
let v = classify(403, &h(&[("x-amzn-waf-action", "block")]), body);
assert_eq!(v.origin, Origin::EdgeBlocked);
let i = reinterpret(&v, true);
assert!(!i.reflection_is_meaningful, "a payload echoed by the edge says nothing about the app");
assert!(!i.conclusive);
}
#[test]
fn throttling_is_kept_apart_from_blocking() {
let v = classify(429, &h(&[("retry-after", "60")]), "Too Many Requests");
assert_eq!(v.origin, Origin::EdgeThrottled);
let i = reinterpret(&v, false);
assert!(!i.conclusive);
assert!(i.note.contains("retry"), "throttling means slow down, not 'this payload is interesting'");
let challenge = classify(503, &h(&[("server", "cloudflare")]), "Just a moment... checking your browser");
assert_eq!(challenge.origin, Origin::EdgeThrottled);
}
#[test]
fn an_unbranded_filter_still_counts_when_it_talks_like_one() {
let v = classify(406, &h(&[]), "Not Acceptable! An appropriate representation ... mod_security");
assert_eq!(v.origin, Origin::EdgeBlocked);
assert_eq!(v.vendor.as_ref().map(|x| x.name.as_str()), Some("ModSecurity"));
}
#[test]
fn coverage_says_whether_the_absence_of_findings_means_anything() {
let blocked = Verdict {
origin: Origin::EdgeBlocked,
vendor: Some(Vendor { name: "Cloudflare".into(), signal: "cf-ray".into() }),
reason: String::new(),
};
let app = Verdict { origin: Origin::Application, vendor: None, reason: String::new() };
// 44 of 47 probes died at the edge: the clean result is about the WAF.
let mut mostly_blocked = vec![blocked.clone(); 44];
mostly_blocked.extend(vec![app.clone(); 3]);
let c = Coverage::of(&mostly_blocked);
assert!(!c.negative_results_are_meaningful());
assert!(c.summary().contains("NOT evidence"));
assert_eq!(c.vendor.as_deref(), Some("Cloudflare"));
// Testing that mostly got through says something real.
let mut mostly_through = vec![app; 40];
mostly_through.extend(vec![blocked; 7]);
let c2 = Coverage::of(&mostly_through);
assert!(c2.negative_results_are_meaningful());
assert!(!c2.summary().contains("NOT evidence"));
}
#[test]
fn a_recorded_exchange_classifies_the_same_way_a_live_one_does() {
let mut ex = crate::validation::Exchange {
status: 403,
body: "<h1>Attention Required! | Cloudflare</h1>".into(),
..Default::default()
};
ex.headers.insert("cf-ray".into(), "8a1b".into());
assert_eq!(classify_exchange(&ex).origin, Origin::EdgeBlocked);
ex.status = 200;
ex.body = "<html>dashboard</html>".into();
assert_eq!(classify_exchange(&ex).origin, Origin::Application);
}
#[test]
fn an_empty_batch_does_not_claim_zero_coverage() {
let c = Coverage::of(&[]);
assert_eq!(c.ratio(), 1.0, "no probes is not the same as every probe blocked");
assert!(c.negative_results_are_meaningful());
}
}