diff --git a/README.md b/README.md index 06866b8..e9b34dd 100755 --- a/README.md +++ b/README.md @@ -250,6 +250,13 @@ Zero npm dependencies (Node built-ins only). log tab grows a prompt box (`❭`) to send `/status`, `/stop`, `/continue`, or a plain-language instruction mid-run — same REPL described in [§6](TUTORIAL.md#6-the-interactive-repl). `host` / `aitest` / `skills` stay one-shot (their onboarding menu can't be scripted over piped stdin). +- **Dashboard** — coverage (engagements, targets, agents run), findings by severity, most + frequent weaknesses, and an **annualized loss exposure computed with FAIR** + (Loss Event Frequency × Loss Magnitude): frequency from each finding's exploitability and + validation confidence, magnitude from assumptions that are shown on screen and editable. + Reported as a min / most-likely / max range, never a single number. +- **Run history in folders** — runs group into one folder per target with a filter box, instead + of one flat list that grows forever. - **Terminal dock** — `Ctrl+\`` (or `❭_` in the sidebar) opens a real terminal, xterm.js over an unstripped stdout stream, so the harness renders with its own colour and panels. Its header switches the terminal between a standalone REPL session and the engagement currently running, @@ -262,6 +269,42 @@ Zero npm dependencies (Node built-ins only). Full API reference: **[web/API.md](web/API.md)** · quick start: **[web/README.md](web/README.md)**. +### Knowledge: memory + attack knowledge graph + +Every model call starts with an empty context window, so without somewhere to put what a run +learned, the harness re-derives the same stack, the same endpoints and the same dead ends every +time. Two stores fix that, both under `.neurosploit/` in the project directory: + +- **Layered memory** (`/memory`, `/forget`) — four tiers by scope, not importance: *working* + (one run), *engagement* (one target), *technique* (one agent/CWE), *reusable* (generalized). + Promotion is evidence-gated: a claim repeated within a run becomes engagement knowledge, one + confirmed across runs becomes technique knowledge, and one that held on **two different + targets** is generalized into a reusable lesson with the host-specific tokens stripped. Recall + is scored (term overlap × past success × recency) and injected into recon/exploit prompts as + leads to verify — never as assertions. +- **Attack knowledge graph** (`/graph`, `graph.json`) — typed entities (asset, endpoint, + weakness, technique, finding, account, credential, impact) joined by typed, weighted, + provenance-carrying edges, accumulated across runs. It answers what a finding list can't: + ranked attack paths, which endpoint accumulated the most weaknesses, and the *frontier* — + entities observed but never proven, i.e. where chaining should look next. Chain edges the + harness derived itself are marked `inferred` and drawn dashed in the web console. Secrets + never enter the graph; they stay in the vault. + +### Keeping a run going + +- **Command rectification** — a mistyped command is corrected (`/staus` → `/status`), completed + (`/onb` → `/onboard`), or reported as ambiguous, never guessed at. Arguments too: a bare host + gets its scheme, an out-of-range count is clamped *with a note*, a near-miss model id is + matched against the live catalog. +- **Automatic backend fallback** — when every configured model is quota-exhausted or its token + is dead, the pool switches to whatever else this machine can reach (an installed CLI + subscription, or a provider whose API key is in the environment) and keeps going. It only + parks the run when nothing at all is available. +- **Resume where it stopped** — findings are checkpointed live, so an interrupted run is + recovered on the next start and `/continue` carries them forward. Non-interactive sessions + (the web console drives the REPL over a pipe) resume automatically, since no one is there to + type it; set `NEUROSPLOIT_AUTO_RESUME=1` to get the same at a terminal. + --- ## 🔌 Integrations (GitHub · GitLab · Jira) diff --git a/neurosploit-rs/app/src/main.rs b/neurosploit-rs/app/src/main.rs index 5c236b6..098fad9 100644 --- a/neurosploit-rs/app/src/main.rs +++ b/neurosploit-rs/app/src/main.rs @@ -1,5 +1,6 @@ //! NeuroSploit v4.0.0 — interactive harness + CLI (`run` / `whitebox` / `agents` / `models`). +mod rectify; mod repl; mod tui; diff --git a/neurosploit-rs/app/src/rectify.rs b/neurosploit-rs/app/src/rectify.rs new file mode 100644 index 0000000..215c732 --- /dev/null +++ b/neurosploit-rs/app/src/rectify.rs @@ -0,0 +1,326 @@ +//! Command rectification for the REPL. +//! +//! A mistyped command used to cost the operator a full round trip: `unknown +//! command '/staus' — try /help`, then reading the help, then retyping. During a +//! live run that is the worst possible moment to lose your place. This module +//! turns a typo into either the command that was obviously meant, or a short +//! list of what was probably meant — never a silent guess. +//! +//! Three rules, in order, and the order is the point: +//! +//! 1. **Accepted as typed** wins over everything. The dispatch accepts ~70 +//! literals including aliases (`/q`, `/url`, `/log`), so correction must only +//! ever see input the dispatch would have rejected — otherwise `/url` gets +//! "corrected" to `/ua` and a working command starts doing something else. +//! 2. **Unique prefix**: `/stat` completes to `/status` when nothing else starts +//! that way. This is what Tab would have done. +//! 3. **Edit distance** with transpositions (`/staus`, `/sttaus` → `/status`), +//! with the budget scaled to word length, and only when one candidate is +//! strictly closer than the runner-up. A tie is ambiguity, and ambiguity is +//! reported, not resolved — running the wrong command against a live target +//! is worse than asking. +//! +//! Arguments get the same treatment where a mistake has one obvious reading: a +//! bare host is a URL missing its scheme, and an out-of-range count is a clamp +//! with a note, not a silent default. + +/// What to do with an input command. +#[derive(Debug, Clone, PartialEq, Eq)] +pub enum Fix { + /// The dispatch accepts it as typed. + Accepted, + /// Unambiguously a typo for `to`; `note` explains the substitution. + Corrected { to: String, note: String }, + /// Several equally plausible commands — the operator has to pick. + Ambiguous(Vec), + /// Nothing close enough; `Vec` holds any weak suggestions (possibly empty). + Unknown(Vec), +} + +/// Damerau-Levenshtein (optimal string alignment) distance. +/// +/// Plain Levenshtein scores a transposition as two edits, which is exactly the +/// typo a fast typist makes most (`/sttaus`); counting it as one is what lets a +/// tight budget still catch it. +pub fn distance(a: &str, b: &str) -> usize { + let a: Vec = a.chars().collect(); + let b: Vec = b.chars().collect(); + let (n, m) = (a.len(), b.len()); + if n == 0 { + return m; + } + if m == 0 { + return n; + } + let mut d = vec![vec![0usize; m + 1]; n + 1]; + for (i, row) in d.iter_mut().enumerate().take(n + 1) { + row[0] = i; + } + for j in 0..=m { + d[0][j] = j; + } + for i in 1..=n { + for j in 1..=m { + let cost = usize::from(a[i - 1] != b[j - 1]); + d[i][j] = (d[i - 1][j] + 1).min(d[i][j - 1] + 1).min(d[i - 1][j - 1] + cost); + if i > 1 && j > 1 && a[i - 1] == b[j - 2] && a[i - 2] == b[j - 1] { + d[i][j] = d[i][j].min(d[i - 2][j - 2] + 1); + } + } + } + d[n][m] +} + +/// Edit budget for a word of this length. Two edits on `/ua` would reach half +/// the command list, so short commands get a tighter budget than long ones. +fn budget(len: usize) -> usize { + match len { + 0..=3 => 0, + 4..=5 => 1, + _ => 2, + } +} + +/// Decide what a typed command should become. `accepted` is every literal the +/// dispatch handles, aliases included. +pub fn rectify_command(input: &str, accepted: &[&str]) -> Fix { + let raw = input.trim(); + if raw.is_empty() { + return Fix::Unknown(vec![]); + } + let lower = raw.to_lowercase(); + if accepted.iter().any(|c| *c == lower) { + return Fix::Accepted; + } + + // A command typed without its slash (`status`) is a command, not prose — + // prose does not collide with the dispatch table. + let slashed = if lower.starts_with('/') { lower.clone() } else { format!("/{lower}") }; + if !lower.starts_with('/') && accepted.iter().any(|c| *c == slashed) { + return Fix::Corrected { to: slashed.clone(), note: format!("read '{raw}' as '{slashed}'") }; + } + + // Unique prefix — what Tab completion would have produced. + if slashed.len() >= 3 { + let pre: Vec<&str> = accepted.iter().copied().filter(|c| c.starts_with(&slashed)).collect(); + if pre.len() == 1 { + return Fix::Corrected { to: pre[0].to_string(), note: format!("completed '{raw}' → '{}'", pre[0]) }; + } + if pre.len() > 1 { + let mut v: Vec = pre.iter().map(|s| s.to_string()).collect(); + v.sort(); + v.dedup(); + return Fix::Ambiguous(v); + } + } + + let mut scored: Vec<(usize, &str)> = accepted.iter().map(|c| (distance(&slashed, c), *c)).collect(); + scored.sort_by(|a, b| a.0.cmp(&b.0).then_with(|| a.1.cmp(b.1))); + let budget = budget(slashed.len()); + let best = scored.first().copied(); + + if let Some((d0, c0)) = best { + if d0 <= budget { + let runner_up = scored.iter().skip(1).find(|(_, c)| *c != c0).map(|(d, _)| *d).unwrap_or(usize::MAX); + if d0 < runner_up { + return Fix::Corrected { to: c0.to_string(), note: format!("corrected '{raw}' → '{c0}'") }; + } + let tied: Vec = scored.iter().filter(|(d, _)| *d == d0).map(|(_, c)| c.to_string()).collect(); + return Fix::Ambiguous(tied); + } + } + // Nothing within budget: offer the nearest few as a hint, not a correction. + let hints: Vec = scored.iter().filter(|(d, _)| *d <= budget + 2).take(3).map(|(_, c)| c.to_string()).collect(); + Fix::Unknown(hints) +} + +/// Normalize a target the way an operator meant it: add the missing scheme, fix +/// a mistyped one, and drop trailing punctuation a shell or a paste left behind. +/// Returns `None` when the input is already fine. +pub fn rectify_url(input: &str) -> Option { + let raw = input.trim(); + if raw.is_empty() { + return None; + } + let mut s = raw.trim_end_matches(['.', ',', ';', ')', '\'', '"']).to_string(); + let mut changed = s != raw; + + // Common near-misses of the scheme, including the single-slash paste. + for (bad, good) in [ + ("htp://", "http://"), + ("htttp://", "http://"), + ("htps://", "https://"), + ("htpps://", "https://"), + ("httpss://", "https://"), + ("hhttp://", "http://"), + ("http:/", "http://"), + ("https:/", "https://"), + ] { + if s.starts_with(bad) && !s.starts_with(good) { + s = format!("{good}{}", &s[bad.len()..]); + changed = true; + break; + } + } + + if !s.contains("://") { + // A local path is a repo, not a URL — leave it for /repo to handle. + if s.starts_with('/') || s.starts_with("./") || s.starts_with("~") { + return None; + } + s = format!("https://{s}"); + changed = true; + } + if changed { + Some(s) + } else { + None + } +} + +/// Parse a count, clamped into range. Returns the value and an optional note +/// explaining what was changed — an out-of-range number is a typo worth +/// reporting, and silently falling back to the old value hides it. +pub fn rectify_count(input: &str, min: usize, max: usize, current: usize) -> (usize, Option) { + let t = input.trim(); + if t.is_empty() { + return (current, None); + } + // Tolerate "3x", "3 votes", "v3" — the digits are the intent. + let digits: String = t.chars().filter(|c| c.is_ascii_digit()).collect(); + let Ok(n) = digits.parse::() else { + return (current, Some(format!("'{t}' is not a number — keeping {current}"))); + }; + if n < min { + (min, Some(format!("{n} is below the minimum — using {min}"))) + } else if n > max { + (max, Some(format!("{n} is above the maximum — using {max}"))) + } else if digits != t { + (n, Some(format!("read '{t}' as {n}"))) + } else { + (n, None) + } +} + +/// Nearest `provider:model` in the catalog, for `/model` typos. Only returns a +/// candidate when it is close enough to be the same identifier mistyped. +pub fn nearest_model(input: &str, catalog: &[String]) -> Option { + let q = input.trim().to_lowercase(); + if q.is_empty() || catalog.iter().any(|m| m.to_lowercase() == q) { + return None; + } + let mut best: Option<(usize, &String)> = None; + for m in catalog { + let d = distance(&q, &m.to_lowercase()); + if best.map(|(bd, _)| d < bd).unwrap_or(true) { + best = Some((d, m)); + } + } + // Scale with the identifier's length: `gpt-5.4` and `gpt-5.1` differ by one + // character and are different models, so the budget has to stay tight. + best.filter(|(d, m)| *d <= (m.len() / 6).clamp(1, 3)).map(|(_, m)| m.clone()) +} + +#[cfg(test)] +mod tests { + use super::*; + + const ACCEPTED: &[&str] = &[ + "/help", "/status", "/stop", "/show", "/sub", "/run", "/runs", "/report", "/results", + "/target", "/ua", "/url", "/model", "/models", "/mcp", "/only", "/onboard", "/q", "/quit", + "/log", "/logs", "/votes", "/recon", "/repo", + ]; + + #[test] + fn an_accepted_alias_is_never_rewritten() { + // The regression this whole ordering exists to prevent: /url is a real + // alias and must not be "corrected" to the nearby /ua. + for c in ["/url", "/ua", "/q", "/log", "/models"] { + assert_eq!(rectify_command(c, ACCEPTED), Fix::Accepted, "{c}"); + } + } + + #[test] + fn a_transposition_is_one_edit_away() { + assert_eq!(distance("/staus", "/status"), 1, "a dropped character"); + assert_eq!(distance("/sttaus", "/status"), 1, "a swapped pair is one edit, not two"); + assert_eq!(distance("/status", "/statsu"), 1); + match rectify_command("/staus", ACCEPTED) { + Fix::Corrected { to, .. } => assert_eq!(to, "/status"), + other => panic!("expected a correction, got {other:?}"), + } + } + + #[test] + fn a_unique_prefix_completes() { + match rectify_command("/onb", ACCEPTED) { + Fix::Corrected { to, .. } => assert_eq!(to, "/onboard"), + other => panic!("expected completion, got {other:?}"), + } + } + + #[test] + fn a_shared_prefix_asks_instead_of_guessing() { + match rectify_command("/ru", ACCEPTED) { + Fix::Ambiguous(v) => assert_eq!(v, vec!["/run".to_string(), "/runs".to_string()]), + other => panic!("expected ambiguity, got {other:?}"), + } + } + + #[test] + fn a_missing_slash_is_read_as_the_command() { + match rectify_command("status", ACCEPTED) { + Fix::Corrected { to, .. } => assert_eq!(to, "/status"), + other => panic!("expected /status, got {other:?}"), + } + } + + #[test] + fn nonsense_is_not_forced_onto_a_command() { + match rectify_command("/zzzzzzzz", ACCEPTED) { + Fix::Unknown(_) => {} + other => panic!("expected unknown, got {other:?}"), + } + } + + #[test] + fn short_commands_get_no_edit_budget() { + // With a budget, "/ub" would land on "/ua" or "/sub" — both wrong, and + // both a command that changes how the engagement runs. + match rectify_command("/ub", ACCEPTED) { + Fix::Unknown(_) => {} + other => panic!("expected unknown for a 3-char typo, got {other:?}"), + } + } + + #[test] + fn urls_get_the_scheme_they_were_missing() { + assert_eq!(rectify_url("example.com").as_deref(), Some("https://example.com")); + assert_eq!(rectify_url("htp://example.com").as_deref(), Some("http://example.com")); + assert_eq!(rectify_url("https:/example.com").as_deref(), Some("https://example.com")); + assert_eq!(rectify_url("https://example.com/x,").as_deref(), Some("https://example.com/x")); + assert_eq!(rectify_url("https://example.com"), None, "a correct URL is left alone"); + assert_eq!(rectify_url("/opt/src/repo"), None, "a local path is not a URL"); + } + + #[test] + fn counts_clamp_and_say_so() { + assert_eq!(rectify_count("3", 1, 9, 3), (3, None)); + let (v, note) = rectify_count("99", 1, 9, 3); + assert_eq!(v, 9); + assert!(note.unwrap().contains("above the maximum")); + let (v, note) = rectify_count("banana", 1, 9, 4); + assert_eq!(v, 4); + assert!(note.unwrap().contains("not a number")); + let (v, note) = rectify_count("5 votes", 1, 9, 3); + assert_eq!(v, 5); + assert!(note.is_some()); + } + + #[test] + fn a_near_model_id_is_offered_but_a_sibling_version_is_not() { + let catalog: Vec = vec!["openai:gpt-5.4".into(), "anthropic:claude-opus-5".into()]; + assert_eq!(nearest_model("openai:gpt5.4", &catalog).as_deref(), Some("openai:gpt-5.4")); + assert_eq!(nearest_model("openai:gpt-5.4", &catalog), None, "an exact id needs no fix"); + } +} diff --git a/neurosploit-rs/app/src/repl.rs b/neurosploit-rs/app/src/repl.rs index c28f29f..f9272a3 100644 --- a/neurosploit-rs/app/src/repl.rs +++ b/neurosploit-rs/app/src/repl.rs @@ -139,12 +139,32 @@ struct LiveCheckpoint { commands: Vec, } +/// Every literal the dispatch below accepts, aliases included. +/// +/// [`COMMANDS`] is the *discoverable* subset offered by Tab completion; this is +/// the full set, and command rectification needs the full set: correcting input +/// the dispatch would have accepted (`/url`, `/q`, `/log`) into some +/// near-neighbour would break working commands. A test keeps the two in sync. +pub(crate) const ACCEPTED: &[&str] = &[ + "/?", "/agents", "/attach", "/auth", "/burp", "/chain", "/changed", "/clear", "/config", + "/context", "/continue", "/creds", "/diff", "/exclude", "/exit", "/expand", "/feed", + "/finding", "/findings", "/focus", "/forget", "/full", "/go", "/goal", "/graph", "/help", + "/history", "/idle", "/instructions", "/integration", "/integrations", "/key", "/log", + "/logs", "/mcp", "/memory", "/model", "/models", "/objective", "/objectives", "/offline", + "/onboard", "/only", "/oos", "/outofscope", "/providers", "/proxy", "/q", "/quit", "/recon", + "/repo", "/report", "/results", "/resume", "/retest", "/revalidate", "/run", "/runs", + "/scope", "/scope-out", "/show", "/status", "/stop", "/sub", "/subscription", "/target", + "/temp-email", "/tempmail", "/theme", "/timeout", "/ua", "/url", "/useragent", "/validate", + "/votes", +]; + /// All slash-commands, for Tab completion. const COMMANDS: &[&str] = &[ "/help", "/onboard", "/show", "/config", "/providers", "/model", "/key", "/sub", "/target", "/repo", "/auth", "/creds", "/focus", "/objective", "/scope-out", "/attach", "/context", "/mcp", "/offline", "/votes", "/chain", "/recon", "/tempmail", "/timeout", "/proxy", "/burp", "/ua", "/agents", "/only", "/theme", "/clear", "/run", "/stop", "/continue", "/runs", "/results", "/report", - "/status", "/logs", "/diff", "/retest", "/validate", "/finding", "/expand", "/integrations", "/quit", + "/status", "/logs", "/diff", "/retest", "/validate", "/finding", "/expand", "/integrations", + "/memory", "/forget", "/graph", "/quit", ]; /// rustyline helper: Tab-completes `/commands` and `@filesystem-paths`, @@ -397,6 +417,8 @@ pub async fn repl(base: &Path) -> anyhow::Result<()> { // A recovered interrupted run, carried in memory so `/continue` can relaunch // the engagement on the same target with these findings folded forward. let mut resumable: Option<(String, Vec)> = None; + // Set when a recovered run should continue without waiting for a human. + let mut auto_resume = false; // Recover an interrupted run (REPL was quit/crashed mid-engagement): its // live findings were checkpointed to disk — fold them into /runs so // /results, /finding and /report still work. @@ -411,8 +433,19 @@ pub async fn repl(base: &Path) -> anyhow::Result<()> { save_runs(base, &h); println!(" \x1b[1;33m↻ recovered interrupted run on {} — {} finding(s) saved as run #{}\x1b[0m (/results {id} · /report {id})", cp.target, cp.findings.len(), id); - println!(" \x1b[36m ↳ /continue to keep testing this target — the {} finding(s) carry forward\x1b[0m", cp.findings.len()); resumable = Some((cp.target.clone(), cp.findings.clone())); + // Resume by itself where nobody is watching: the web console drives + // this REPL over a pipe, and a run that stops there waits forever + // for a `/continue` no one will type. An interactive operator keeps + // the choice — relaunching an engagement spends tokens, and at a + // real terminal there is someone to decide. + auto_resume = !std::io::stdin().is_terminal() + || std::env::var("NEUROSPLOIT_AUTO_RESUME").map(|v| v == "1" || v == "true").unwrap_or(false); + if auto_resume { + println!(" \x1b[36m ↳ resuming automatically — the {} finding(s) carry forward\x1b[0m", cp.findings.len()); + } else { + println!(" \x1b[36m ↳ /continue to keep testing this target — the {} finding(s) carry forward\x1b[0m", cp.findings.len()); + } } clear_checkpoint(); } @@ -420,6 +453,12 @@ pub async fn repl(base: &Path) -> anyhow::Result<()> { let mut reader = Reader::new(base); let mut active: Option = None; let mut queue: Vec = Vec::new(); // remaining targets for a multi-target /run + // Commands to run before reading from the user — how an auto-resumed run + // re-enters the normal dispatch instead of duplicating /continue's logic. + let mut pending: Vec = Vec::new(); + if auto_resume && resumable.is_some() { + pending.push("/continue".into()); + } // First-launch onboarding: pick scope (web/infra/cloud/ai/skills) → box → setup. if s.target.is_none() && s.repo.is_none() && std::io::stdin().is_terminal() { onboarding(&mut s); @@ -434,7 +473,14 @@ pub async fn repl(base: &Path) -> anyhow::Result<()> { active = start_background(base, &s, &mut reader, history.clone(), Some(&next), vec![]).await; } println!("{}", context_prompt(&s)); // dim context line above the prompt - let Some(line) = reader.read(PROMPT) else { println!("\n bye."); break }; + let line = if pending.is_empty() { + let Some(l) = reader.read(PROMPT) else { println!("\n bye."); break }; + l + } else { + let l = pending.remove(0); + println!("{PROMPT}{l}"); + l + }; // Ctrl-C → confirm before doing anything drastic (don't lose a live run). if line == CTRL_C { let run_active = active.as_ref().map(|a| !a.done.load(Ordering::Relaxed)).unwrap_or(false); @@ -481,7 +527,30 @@ pub async fn repl(base: &Path) -> anyhow::Result<()> { None => continue, } }; - let (cmd, arg) = (cmd.as_str(), arg.as_str()); + // Rectify before dispatch, so the match below only ever sees a command + // it handles. A typo mid-run costs an operator their place in the + // output; correcting the obvious ones — and asking about the rest — + // keeps a slip from becoming a round trip through /help. + let cmd_owned = match crate::rectify::rectify_command(&cmd, ACCEPTED) { + crate::rectify::Fix::Accepted => cmd.clone(), + crate::rectify::Fix::Corrected { to, note } => { + println!(" \x1b[2m↻ {note}\x1b[0m"); + to + } + crate::rectify::Fix::Ambiguous(v) => { + println!(" '{cmd}' matches {} commands: {}", v.len(), v.join(" ")); + continue; + } + crate::rectify::Fix::Unknown(hints) => { + if hints.is_empty() { + println!(" unknown command '{cmd}' — /help lists them all"); + } else { + println!(" unknown command '{cmd}' — did you mean {}?", hints.join(", ")); + } + continue; + } + }; + let (cmd, arg) = (cmd_owned.as_str(), arg.as_str()); match cmd { "/help" | "/?" => help(), "/show" | "/config" => show(&s), @@ -496,7 +565,17 @@ pub async fn repl(base: &Path) -> anyhow::Result<()> { if arg.is_empty() { pick_models(&mut s); } else { - s.models = arg.split([',', ' ']).filter(|x| !x.is_empty()).map(String::from).collect(); + // A model id is long and easy to fumble; an unrecognized one + // otherwise fails much later, inside the run. + let catalog: Vec = harness::providers().iter() + .flat_map(|p| p.models.iter().map(move |m| format!("{}:{}", p.key, m))) + .collect(); + s.models = arg.split([',', ' ']).filter(|x| !x.is_empty()).map(|x| { + match crate::rectify::nearest_model(x, &catalog) { + Some(fixed) => { println!(" \x1b[2m↻ corrected '{x}' → '{fixed}'\x1b[0m"); fixed } + None => x.to_string(), + } + }).collect(); println!(" models: {}", s.models.join(", ")); } // If a run is paused on exhaustion, queue the newly-chosen models @@ -518,9 +597,11 @@ pub async fn repl(base: &Path) -> anyhow::Result<()> { if arg.is_empty() { println!(" target: {}", s.target.clone().unwrap_or_else(|| "(none) — set with /target , clear with /target clear".into())); } else if arg == "clear" { s.target = None; println!(" target cleared"); } else { - // Accept one URL or a comma-separated list; normalize each. + // Accept one URL or a comma-separated list; normalize each — + // a missing scheme, a mistyped one (`htp://`, `https:/`) or + // a trailing comma from a paste all resolve to one reading. let ts: Vec = arg.split(',').map(|x| x.trim()).filter(|x| !x.is_empty()) - .map(|x| if x.starts_with("http") { x.to_string() } else { format!("https://{x}") }) + .map(|x| crate::rectify::rectify_url(x).unwrap_or_else(|| x.to_string())) .collect(); s.target = Some(ts.join(",")); if ts.len() > 1 { println!(" targets ({}): {}", ts.len(), ts.join(", ")); println!(" \x1b[2m/run tests them sequentially, one report each\x1b[0m"); } @@ -640,7 +721,14 @@ pub async fn repl(base: &Path) -> anyhow::Result<()> { "/mcp" => { s.mcp = !matches!(arg, "off" | "false" | "0" | "no"); println!(" Playwright MCP: {}", onoff(s.mcp)); } "/offline" => { s.offline = !matches!(arg, "off" | "false" | "0" | "no"); println!(" offline: {}", onoff(s.offline)); } "/integrations" | "/integration" => integrations_cmd(arg), - "/votes" => { s.vote_n = arg.parse().unwrap_or(s.vote_n); println!(" votes: {}", s.vote_n); } + "/votes" => { + // Out of range used to fall back to the current value in + // silence, so `/votes 30` looked applied and wasn't. + let (n, note) = crate::rectify::rectify_count(arg, 1, 9, s.vote_n); + if let Some(note) = note { println!(" \x1b[2m↻ {note}\x1b[0m"); } + s.vote_n = n; + println!(" votes: {}", s.vote_n); + } "/chain" => { if arg.is_empty() { println!(" attack-chain depth: {} (0 disables) — set with /chain ", s.chain_depth); } else { s.chain_depth = arg.parse().unwrap_or(s.chain_depth); println!(" attack-chain depth: {}", s.chain_depth); } @@ -925,7 +1013,23 @@ pub async fn repl(base: &Path) -> anyhow::Result<()> { } save_session(&s); println!(" session saved → {} · bye.", proj_dir().display()); break; } - other => println!(" unknown command '{other}' — try /help"), + "/memory" => memory_cmd(&s, arg), + "/forget" => { + if arg.trim().is_empty() { + println!(" usage: /forget — drops every memory whose text contains it"); + } else { + let mut mem = harness::memory::Memory::open(proj_dir().join("memory")); + let n = mem.forget(arg.trim()); + println!(" forgot {n} memo(s) matching '{}'", arg.trim()); + } + } + "/graph" => { + let g = harness::knowledge_graph::KnowledgeGraph::load(proj_dir().join("graph.json")); + print!("{}", g.summary()); + } + // Rectification only forwards commands listed in ACCEPTED, so + // reaching here means ACCEPTED lists something this match forgot. + other => println!(" '{other}' is listed but not implemented — please report this."), } } Ok(()) @@ -1348,6 +1452,47 @@ fn merge_findings(prior: Vec, mut fresh: Vec) -> Vec fresh } +/// `/memory` — inspect what the harness has learned, or search it. +/// +/// The four tiers are shown separately because they mean different things: an +/// engagement memo is about *this* target, a reusable one is a lesson that +/// already held on two of them. Collapsing them into one list would hide the +/// distinction that makes the promotion ladder worth having. +fn memory_cmd(s: &Session, arg: &str) { + let mem = harness::memory::Memory::open(proj_dir().join("memory")); + let (w, e, t, r) = mem.counts(); + let q = arg.trim(); + if q.is_empty() { + println!(" ┌ memory · working {w} · engagement {e} · technique {t} · reusable {r}"); + let recent = mem.dump(); + if recent.is_empty() { + println!(" │ (nothing learned yet — memory fills in as runs finish)"); + } + for m in recent.iter().take(12) { + println!(" │ [{:<10} {:>3}%] {}", m.tier.as_str(), (m.confidence * 100.0) as u32, trunc(&m.text, 92)); + } + if recent.len() > 12 { + println!(" │ … {} more · /memory to search", recent.len() - 12); + } + println!(" └ /forget removes matching memos"); + return; + } + let hits = mem.recall(&harness::memory::Query { + text: q.to_string(), + target: s.target.clone().unwrap_or_default(), + limit: 15, + ..Default::default() + }); + if hits.is_empty() { + println!(" no memory matches '{q}'"); + return; + } + println!(" ── {} match(es) for '{q}' ──", hits.len()); + for h in hits { + println!(" [{:.2}] \x1b[2m{:<10}\x1b[0m {}", h.score, h.memo.tier.as_str(), trunc(&h.memo.text, 96)); + } +} + /// Project-local store: `/.neurosploit/` so each project keeps its own /// session, run history and command history (resume on reopen). No DB needed — /// it's structured state, not semantic search. @@ -1760,6 +1905,11 @@ fn help() { h("/retest [n]", "re-verify a past run's findings (re-runs the test)"); h("/validate [n]", "false-positive validate a recovered/past run (no re-test)"); + println!("\n \x1b[2mKNOWLEDGE\x1b[0m"); + h("/memory [text]", "what the harness learned (working·engagement·technique·reusable); search with text"); + h("/forget ", "drop every memory whose text matches"); + h("/graph", "attack knowledge graph: entities, top attack paths, unproven frontier"); + println!("\n \x1b[2mINTEGRATIONS\x1b[0m"); h("/integrations", "show · enable/disable github|gitlab|jira · setup "); @@ -2287,4 +2437,23 @@ mod nl_tests { assert_eq!(parse_intent_fast("recon 4 em example.com").0.recon, Some(4)); } + /// Rectification forwards only what ACCEPTED lists, so anything offered by + /// Tab completion but missing from ACCEPTED would become unreachable — the + /// user would type a real command and be told it doesn't exist. + #[test] + fn every_completable_command_is_accepted_by_the_dispatch() { + let missing: Vec<&&str> = COMMANDS.iter().filter(|c| !ACCEPTED.contains(c)).collect(); + assert!(missing.is_empty(), "completed but not dispatchable: {missing:?}"); + } + + #[test] + fn accepted_commands_survive_rectification_untouched() { + for c in ACCEPTED { + assert_eq!( + crate::rectify::rectify_command(c, ACCEPTED), + crate::rectify::Fix::Accepted, + "{c} must reach the dispatch as typed" + ); + } + } } diff --git a/neurosploit-rs/crates/harness/src/knowledge_graph.rs b/neurosploit-rs/crates/harness/src/knowledge_graph.rs new file mode 100644 index 0000000..0037a75 --- /dev/null +++ b/neurosploit-rs/crates/harness/src/knowledge_graph.rs @@ -0,0 +1,602 @@ +//! The attack knowledge graph: what was learned about a target, as a graph. +//! +//! [`crate::attack_graph`] maps a finding to OWASP/MITRE/stage and draws it. +//! That is a *per-run view*. This module is the durable structure underneath: +//! typed entities (asset, endpoint, weakness, technique, finding, account, +//! credential, impact) joined by typed, weighted, provenance-carrying edges, so +//! the harness can answer questions a flat finding list cannot — +//! +//! - which endpoint accumulated the most distinct weaknesses across runs; +//! - which credential a finding actually yielded, and what that credential then +//! unlocked; +//! - what paths run from the asset to an impact node, ranked by how likely and +//! how damaging they are; +//! - what the frontier is: entities we have observed but never proved anything +//! about — the natural next targets for chaining. +//! +//! ## Inferred edges are marked as inferred +//! +//! Agents only sometimes populate `chains_from`. Without it a "chain" view +//! degenerates into a fan of unconnected findings, so this module also *infers* +//! progression edges between kill-chain stages. Those carry `inferred: true` and +//! a lower probability, and every renderer draws them differently, because an +//! inferred edge is a hypothesis about an attack path — presenting it as a +//! proven one would be the graph lying about its own evidence. + +use crate::types::Finding; +use serde::{Deserialize, Serialize}; +use std::collections::{BTreeMap, BTreeSet}; +use std::path::Path; + +#[derive(Clone, Copy, Debug, PartialEq, Eq, PartialOrd, Ord, Serialize, Deserialize)] +#[serde(rename_all = "kebab-case")] +pub enum NodeKind { + Asset, + Endpoint, + Tech, + Weakness, + Technique, + Finding, + Account, + Credential, + Impact, +} + +impl NodeKind { + pub fn as_str(&self) -> &'static str { + match self { + NodeKind::Asset => "asset", + NodeKind::Endpoint => "endpoint", + NodeKind::Tech => "tech", + NodeKind::Weakness => "weakness", + NodeKind::Technique => "technique", + NodeKind::Finding => "finding", + NodeKind::Account => "account", + NodeKind::Credential => "credential", + NodeKind::Impact => "impact", + } + } +} + +#[derive(Clone, Copy, Debug, PartialEq, Eq, PartialOrd, Ord, Serialize, Deserialize)] +#[serde(rename_all = "kebab-case")] +pub enum EdgeKind { + /// asset → endpoint + Exposes, + /// asset → tech + Runs, + /// endpoint → weakness + Vulnerable, + /// finding → weakness (this finding proves that weakness) + Proves, + /// finding → endpoint (where it was proven) + ObservedOn, + /// finding → technique (MITRE) + Uses, + /// finding → finding (attack path) + Chains, + /// finding → account/credential + Grants, + /// finding → impact + Leads, +} + +impl EdgeKind { + pub fn as_str(&self) -> &'static str { + match self { + EdgeKind::Exposes => "exposes", + EdgeKind::Runs => "runs", + EdgeKind::Vulnerable => "vulnerable", + EdgeKind::Proves => "proves", + EdgeKind::ObservedOn => "observed-on", + EdgeKind::Uses => "uses", + EdgeKind::Chains => "chains", + EdgeKind::Grants => "grants", + EdgeKind::Leads => "leads", + } + } +} + +#[derive(Clone, Debug, Serialize, Deserialize)] +pub struct Node { + pub id: String, + pub kind: NodeKind, + pub label: String, + #[serde(default)] + pub meta: BTreeMap, + /// Run ids that touched this node — provenance, and the "how often" signal. + #[serde(default)] + pub runs: BTreeSet, + #[serde(default)] + pub first_seen: u64, + #[serde(default)] + pub last_seen: u64, +} + +#[derive(Clone, Debug, Serialize, Deserialize)] +pub struct Edge { + pub from: String, + pub to: String, + pub kind: EdgeKind, + /// Confidence that the relation holds, 0..1. + #[serde(default)] + pub p: f64, + /// True when the harness derived this edge rather than an agent asserting it. + #[serde(default)] + pub inferred: bool, + #[serde(default)] + pub runs: BTreeSet, +} + +#[derive(Default, Clone, Serialize, Deserialize)] +pub struct KnowledgeGraph { + pub nodes: BTreeMap, + pub edges: Vec, +} + +fn now() -> u64 { + std::time::SystemTime::now() + .duration_since(std::time::UNIX_EPOCH) + .map(|d| d.as_secs()) + .unwrap_or(0) +} + +fn sev_weight(sev: &str) -> f64 { + match sev.to_lowercase().as_str() { + s if s.starts_with("crit") => 1.0, + s if s.starts_with("high") => 0.75, + s if s.starts_with("med") => 0.5, + s if s.starts_with("low") => 0.25, + _ => 0.1, + } +} + +/// Kill-chain progression order. An attack moves down this list; an edge that +/// would move *up* it is not progression and is never inferred. +pub const STAGES: &[&str] = &[ + "recon", + "discovery", + "initial-access", + "execution", + "persistence", + "privesc", + "credential-access", + "lateral", + "collection", + "exfil", + "impact", +]; + +pub fn stage_rank(s: &str) -> usize { + STAGES.iter().position(|x| *x == s).unwrap_or(STAGES.len()) +} + +/// Strip the query string and normalize the host so two spellings of one URL +/// collapse into a single endpoint node. +fn endpoint_key(url: &str) -> String { + let u = url.trim(); + let no_scheme = u.split_once("://").map(|(_, r)| r).unwrap_or(u); + let path = no_scheme.split(['?', '#']).next().unwrap_or(no_scheme); + let path = path.trim_end_matches('/'); + let path = path.trim_start_matches("www."); + if path.is_empty() { + no_scheme.to_string() + } else { + path.to_lowercase() + } +} + +impl KnowledgeGraph { + pub fn new() -> Self { + Self::default() + } + + pub fn load(path: impl AsRef) -> Self { + std::fs::read_to_string(path) + .ok() + .and_then(|s| serde_json::from_str(&s).ok()) + .unwrap_or_default() + } + + pub fn save(&self, path: impl AsRef) { + let path = path.as_ref(); + if let Some(parent) = path.parent() { + let _ = std::fs::create_dir_all(parent); + } + if let Ok(j) = serde_json::to_string_pretty(self) { + let _ = std::fs::write(path, j); + } + } + + pub fn to_json(&self) -> String { + serde_json::to_string_pretty(self).unwrap_or_else(|_| "{}".into()) + } + + fn upsert(&mut self, id: &str, kind: NodeKind, label: &str, run: &str) -> String { + let ts = now(); + let n = self.nodes.entry(id.to_string()).or_insert_with(|| Node { + id: id.to_string(), + kind, + label: label.to_string(), + meta: BTreeMap::new(), + runs: BTreeSet::new(), + first_seen: ts, + last_seen: ts, + }); + n.last_seen = ts; + if !run.is_empty() { + n.runs.insert(run.to_string()); + } + if n.label.is_empty() { + n.label = label.to_string(); + } + id.to_string() + } + + fn meta(&mut self, id: &str, k: &str, v: &str) { + if v.is_empty() { + return; + } + if let Some(n) = self.nodes.get_mut(id) { + n.meta.insert(k.to_string(), v.to_string()); + } + } + + /// Add or reinforce an edge. Re-observing an edge raises its probability + /// toward certainty rather than appending a duplicate, and an edge first + /// inferred but later asserted by an agent stops being marked inferred. + pub fn link(&mut self, from: &str, to: &str, kind: EdgeKind, p: f64, inferred: bool, run: &str) { + if from == to || !self.nodes.contains_key(from) || !self.nodes.contains_key(to) { + return; + } + if let Some(e) = self.edges.iter_mut().find(|e| e.from == from && e.to == to && e.kind == kind) { + e.p = (e.p + 0.3 * (p.max(e.p) - e.p)).clamp(0.0, 0.99); + e.inferred = e.inferred && inferred; + if !run.is_empty() { + e.runs.insert(run.to_string()); + } + return; + } + let mut runs = BTreeSet::new(); + if !run.is_empty() { + runs.insert(run.to_string()); + } + self.edges.push(Edge { from: from.into(), to: to.into(), kind, p: p.clamp(0.0, 0.99), inferred, runs }); + } + + /// Fold one run's findings into the graph. + pub fn ingest(&mut self, target: &str, run: &str, findings: &[Finding]) { + let akey = crate::memory::engagement_key(target); + let asset = self.upsert(&format!("asset:{akey}"), NodeKind::Asset, if akey.is_empty() { target } else { &akey }, run); + + for f in findings { + let fid = self.upsert( + &format!("find:{}:{}", run, f.id), + NodeKind::Finding, + if f.title.is_empty() { &f.id } else { &f.title }, + run, + ); + self.meta(&fid, "severity", &f.severity); + self.meta(&fid, "cwe", &f.cwe); + self.meta(&fid, "stage", &f.stage); + self.meta(&fid, "owasp", &f.owasp); + self.meta(&fid, "mitre", &f.mitre); + self.meta(&fid, "exploitability", &f.exploitability); + self.meta(&fid, "agent", &f.agent); + self.meta(&fid, "endpoint", &f.endpoint); + self.meta(&fid, "confidence", &format!("{:.2}", f.confidence)); + self.meta(&fid, "review_status", &f.review_status); + self.meta(&fid, "finding_id", &f.id); + + if !f.endpoint.is_empty() { + let ek = endpoint_key(&f.endpoint); + let ep = self.upsert(&format!("ep:{ek}"), NodeKind::Endpoint, &ek, run); + self.link(&asset, &ep, EdgeKind::Exposes, 0.9, false, run); + self.link(&fid, &ep, EdgeKind::ObservedOn, 0.95, false, run); + if !f.cwe.is_empty() { + let w = self.upsert(&format!("cwe:{}", f.cwe), NodeKind::Weakness, &f.cwe, run); + self.link(&ep, &w, EdgeKind::Vulnerable, f.confidence.max(0.5), false, run); + } + } + if !f.cwe.is_empty() { + let w = self.upsert(&format!("cwe:{}", f.cwe), NodeKind::Weakness, &f.cwe, run); + self.link(&fid, &w, EdgeKind::Proves, f.confidence.max(0.5), false, run); + } + if !f.mitre.is_empty() { + let t = self.upsert(&format!("att:{}", f.mitre), NodeKind::Technique, &f.mitre, run); + self.link(&fid, &t, EdgeKind::Uses, 0.9, false, run); + } + if !f.account.is_empty() { + let a = self.upsert(&format!("acct:{}", f.account), NodeKind::Account, &f.account, run); + self.link(&fid, &a, EdgeKind::Grants, 0.9, false, run); + // The secret itself never enters the graph — the graph is an + // artifact that gets shared; the vault is where secrets live. + if !f.secret.is_empty() { + let c = self.upsert(&format!("cred:{}", f.account), NodeKind::Credential, "credential (vaulted)", run); + self.link(&a, &c, EdgeKind::Grants, 0.9, false, run); + } + } + if sev_weight(&f.severity) >= 0.75 || f.stage == "impact" { + let label = if f.business_impact.is_empty() { f.impact.clone() } else { f.business_impact.clone() }; + let label: String = label.split_whitespace().take(12).collect::>().join(" "); + if !label.is_empty() { + let i = self.upsert(&format!("impact:{}:{}", run, f.id), NodeKind::Impact, &label, run); + self.link(&fid, &i, EdgeKind::Leads, sev_weight(&f.severity), false, run); + } + } + } + + // Asserted chains first: they are evidence. + for f in findings { + for src in &f.chains_from { + let a = format!("find:{}:{}", run, src); + let b = format!("find:{}:{}", run, f.id); + self.link(&a, &b, EdgeKind::Chains, 0.9, false, run); + } + } + self.infer_chains(run, findings); + } + + /// Connect consecutive kill-chain stages when the agents asserted nothing. + /// + /// Only forward moves, only between *adjacent populated* stages, and only + /// from the strongest finding of the earlier stage — a full cross-product + /// would draw a plausible-looking web that encodes no information at all. + fn infer_chains(&mut self, run: &str, findings: &[Finding]) { + let asserted: usize = findings.iter().map(|f| f.chains_from.len()).sum(); + if asserted > 0 || findings.len() < 2 { + return; + } + let mut by_stage: BTreeMap> = BTreeMap::new(); + for f in findings { + by_stage.entry(stage_rank(&f.stage)).or_default().push(f); + } + let ranks: Vec = by_stage.keys().copied().collect(); + for w in ranks.windows(2) { + let (Some(a), Some(b)) = (by_stage.get(&w[0]), by_stage.get(&w[1])) else { continue }; + let best = a + .iter() + .max_by(|x, y| { + (sev_weight(&x.severity) * x.confidence) + .partial_cmp(&(sev_weight(&y.severity) * y.confidence)) + .unwrap_or(std::cmp::Ordering::Equal) + }) + .copied(); + let Some(src) = best else { continue }; + for dst in b { + let from = format!("find:{}:{}", run, src.id); + let to = format!("find:{}:{}", run, dst.id); + self.link(&from, &to, EdgeKind::Chains, 0.35, true, run); + } + } + } + + pub fn neighbors(&self, id: &str) -> Vec<&Edge> { + self.edges.iter().filter(|e| e.from == id).collect() + } + + /// Entities observed but never proved: endpoints with no finding on them, + /// accounts nothing was done with. These are where chaining should look + /// next, and the reason the graph is worth keeping between runs. + pub fn frontier(&self) -> Vec<&Node> { + let proven: BTreeSet<&str> = self + .edges + .iter() + .filter(|e| matches!(e.kind, EdgeKind::ObservedOn | EdgeKind::Proves)) + .map(|e| e.to.as_str()) + .collect(); + let mut v: Vec<&Node> = self + .nodes + .values() + .filter(|n| matches!(n.kind, NodeKind::Endpoint | NodeKind::Account | NodeKind::Credential)) + .filter(|n| !proven.contains(n.id.as_str())) + .collect(); + v.sort_by(|a, b| b.runs.len().cmp(&a.runs.len()).then_with(|| a.id.cmp(&b.id))); + v + } + + /// Ranked attack paths: chains of findings ordered by kill-chain stage, + /// scored by severity × edge probability. Returns node-id paths, longest and + /// most damaging first. + pub fn paths(&self, max: usize) -> Vec<(Vec, f64)> { + let findings: Vec<&Node> = self.nodes.values().filter(|n| n.kind == NodeKind::Finding).collect(); + let has_parent: BTreeSet<&str> = self + .edges + .iter() + .filter(|e| e.kind == EdgeKind::Chains) + .map(|e| e.to.as_str()) + .collect(); + let roots: Vec<&Node> = findings.iter().copied().filter(|n| !has_parent.contains(n.id.as_str())).collect(); + + let mut out: Vec<(Vec, f64)> = Vec::new(); + for r in roots { + let mut stack = vec![(vec![r.id.clone()], self.node_score(&r.id))]; + while let Some((path, score)) = stack.pop() { + let last = path.last().cloned().unwrap_or_default(); + let next: Vec<&Edge> = self + .edges + .iter() + .filter(|e| e.kind == EdgeKind::Chains && e.from == last && !path.contains(&e.to)) + .collect(); + if next.is_empty() { + out.push((path, score)); + continue; + } + for e in next { + let mut p = path.clone(); + p.push(e.to.clone()); + // Depth is capped: cycles are already excluded, but a long + // inferred tail is noise, not a deeper attack. + if p.len() > 12 { + out.push((p, score)); + continue; + } + let s = score + self.node_score(&e.to) * e.p; + stack.push((p, s)); + } + } + } + out.sort_by(|a, b| { + b.0.len() + .cmp(&a.0.len()) + .then_with(|| b.1.partial_cmp(&a.1).unwrap_or(std::cmp::Ordering::Equal)) + }); + out.truncate(if max == 0 { 5 } else { max }); + out + } + + fn node_score(&self, id: &str) -> f64 { + self.nodes + .get(id) + .map(|n| { + let sev = n.meta.get("severity").map(|s| sev_weight(s)).unwrap_or(0.1); + let conf: f64 = n.meta.get("confidence").and_then(|c| c.parse().ok()).unwrap_or(0.5); + sev * conf.clamp(0.2, 1.0) + }) + .unwrap_or(0.0) + } + + /// One line per kind, then the top attack paths — the `/graph` view. + pub fn summary(&self) -> String { + if self.nodes.is_empty() { + return " (knowledge graph empty — run an engagement first)".into(); + } + let mut by_kind: BTreeMap<&str, usize> = BTreeMap::new(); + for n in self.nodes.values() { + *by_kind.entry(n.kind.as_str()).or_insert(0) += 1; + } + let mut s = String::from(" ┌ knowledge graph\n"); + for (k, v) in &by_kind { + s.push_str(&format!(" │ {:<12} {}\n", k, v)); + } + let inferred = self.edges.iter().filter(|e| e.inferred).count(); + s.push_str(&format!(" │ {:<12} {} ({} inferred)\n", "edges", self.edges.len(), inferred)); + let paths = self.paths(3); + if paths.iter().any(|(p, _)| p.len() > 1) { + s.push_str(" │\n │ top attack paths\n"); + for (p, score) in paths.iter().filter(|(p, _)| p.len() > 1) { + let labels: Vec<&str> = p + .iter() + .filter_map(|id| self.nodes.get(id)) + .map(|n| n.label.as_str()) + .collect(); + s.push_str(&format!(" │ [{score:.2}] {}\n", labels.join(" → "))); + } + } + let fr = self.frontier(); + if !fr.is_empty() { + s.push_str(&format!(" │\n │ frontier ({} unproven): {}\n", fr.len(), + fr.iter().take(4).map(|n| n.label.as_str()).collect::>().join(", "))); + } + s.push_str(" └\n"); + s + } +} + +#[cfg(test)] +mod tests { + use super::*; + + fn f(id: &str, sev: &str, cwe: &str, stage: &str, endpoint: &str) -> Finding { + Finding { + id: id.into(), + title: format!("finding {id}"), + severity: sev.into(), + cwe: cwe.into(), + stage: stage.into(), + endpoint: endpoint.into(), + confidence: 0.9, + ..Default::default() + } + } + + #[test] + fn one_endpoint_node_however_the_url_was_written() { + let mut g = KnowledgeGraph::new(); + g.ingest( + "https://ex.com", + "r1", + &[ + f("a", "High", "CWE-89", "initial-access", "https://ex.com/login.aspx?id=1"), + f("b", "Low", "CWE-200", "recon", "http://www.ex.com/login.aspx/"), + ], + ); + let eps: Vec<&Node> = g.nodes.values().filter(|n| n.kind == NodeKind::Endpoint).collect(); + assert_eq!(eps.len(), 1, "got {:?}", eps.iter().map(|n| &n.id).collect::>()); + } + + #[test] + fn asserted_chains_win_and_nothing_is_inferred_alongside_them() { + let mut g = KnowledgeGraph::new(); + let mut b = f("b", "High", "CWE-89", "execution", "https://ex.com/x"); + b.chains_from = vec!["a".into()]; + g.ingest("https://ex.com", "r1", &[f("a", "Medium", "CWE-200", "recon", "https://ex.com/"), b]); + let chains: Vec<&Edge> = g.edges.iter().filter(|e| e.kind == EdgeKind::Chains).collect(); + assert_eq!(chains.len(), 1); + assert!(!chains[0].inferred, "an agent-asserted chain must not be marked inferred"); + } + + #[test] + fn inferred_chains_only_move_forward_through_the_kill_chain() { + let mut g = KnowledgeGraph::new(); + g.ingest( + "https://ex.com", + "r1", + &[ + f("a", "Medium", "CWE-200", "recon", "https://ex.com/"), + f("b", "Critical", "CWE-89", "initial-access", "https://ex.com/login"), + ], + ); + let chains: Vec<&Edge> = g.edges.iter().filter(|e| e.kind == EdgeKind::Chains).collect(); + assert_eq!(chains.len(), 1); + assert!(chains[0].inferred, "a derived edge must say so"); + assert!(chains[0].from.ends_with(":a") && chains[0].to.ends_with(":b"), "recon must precede initial-access"); + assert!(chains[0].p < 0.5, "an inferred edge must carry less weight than an asserted one"); + } + + #[test] + fn a_secret_never_lands_in_the_graph() { + let mut g = KnowledgeGraph::new(); + let mut x = f("a", "High", "CWE-287", "credential-access", "https://ex.com/register"); + x.account = "user1".into(); + x.secret = "hunter2-super-secret".into(); + g.ingest("https://ex.com", "r1", &[x]); + let json = g.to_json(); + assert!(!json.contains("hunter2"), "the vault holds secrets, the graph does not"); + assert!(json.contains("credential (vaulted)")); + } + + #[test] + fn paths_rank_the_longest_most_severe_chain_first() { + let mut g = KnowledgeGraph::new(); + let mut b = f("b", "High", "CWE-89", "initial-access", "https://ex.com/login"); + b.chains_from = vec!["a".into()]; + let mut c = f("c", "Critical", "CWE-78", "execution", "https://ex.com/exec"); + c.chains_from = vec!["b".into()]; + g.ingest("https://ex.com", "r1", &[f("a", "Low", "CWE-200", "recon", "https://ex.com/"), b, c]); + let paths = g.paths(3); + assert_eq!(paths[0].0.len(), 3, "the three-step chain must rank above any single node"); + } + + #[test] + fn re_ingesting_the_same_run_does_not_duplicate_edges() { + let mut g = KnowledgeGraph::new(); + let fs = [f("a", "High", "CWE-89", "initial-access", "https://ex.com/login")]; + g.ingest("https://ex.com", "r1", &fs); + let n = g.edges.len(); + g.ingest("https://ex.com", "r1", &fs); + assert_eq!(g.edges.len(), n); + } + + #[test] + fn the_frontier_lists_endpoints_nothing_was_proven_on() { + let mut g = KnowledgeGraph::new(); + g.ingest("https://ex.com", "r1", &[f("a", "High", "CWE-89", "initial-access", "https://ex.com/login")]); + // An endpoint learned by recon, with no finding attached to it. + let asset = "asset:ex.com".to_string(); + g.upsert("ep:ex.com/admin", NodeKind::Endpoint, "ex.com/admin", "r1"); + g.link(&asset, "ep:ex.com/admin", EdgeKind::Exposes, 0.8, false, "r1"); + let fr: Vec<&str> = g.frontier().iter().map(|n| n.id.as_str()).collect(); + assert_eq!(fr, vec!["ep:ex.com/admin"]); + } +} diff --git a/neurosploit-rs/crates/harness/src/lib.rs b/neurosploit-rs/crates/harness/src/lib.rs index 00bc199..d73c501 100644 --- a/neurosploit-rs/crates/harness/src/lib.rs +++ b/neurosploit-rs/crates/harness/src/lib.rs @@ -13,6 +13,8 @@ pub mod creds; pub mod grounding; pub mod hygiene; pub mod integrations; +pub mod knowledge_graph; +pub mod memory; pub mod pomdp; pub mod models; pub mod pipeline; @@ -29,5 +31,7 @@ pub use models::{ }; pub use pipeline::{run_greybox, run_host, run_whitebox, RunOutput}; pub use pipeline::run; +pub use knowledge_graph::{EdgeKind, KnowledgeGraph, NodeKind}; +pub use memory::{Memory, Query as MemoryQuery, Tier as MemoryTier}; pub use pool::{ModelPool, Task}; pub use types::{Finding, RunConfig}; diff --git a/neurosploit-rs/crates/harness/src/memory.rs b/neurosploit-rs/crates/harness/src/memory.rs new file mode 100644 index 0000000..3157e92 --- /dev/null +++ b/neurosploit-rs/crates/harness/src/memory.rs @@ -0,0 +1,707 @@ +//! Layered memory for the harness. +//! +//! An engagement is a long-running investigation, but every agent call starts +//! from a blank context window. Without a place to put what was learned, the +//! same facts get re-derived every round — the harness re-probes an endpoint it +//! already fingerprinted, re-tries a payload shape that already failed, and +//! forgets across runs entirely. The RL weights in [`crate::rl`] remember *which +//! agent* pays off; they cannot remember *what was true*. +//! +//! Four tiers, separated by what they are scoped to and how long they survive — +//! not by importance: +//! +//! | tier | scope | lives | example | +//! |------|-------|-------|---------| +//! | [`Tier::Working`] | one run | until the run ends | "`/admin` returned 302 to `/login`" | +//! | [`Tier::Engagement`] | one target | forever, decaying | "this host runs IIS 8.5 / ASP.NET 2.0" | +//! | [`Tier::Technique`] | one technique/agent | forever, decaying | "`sqli_error` lands on `.aspx` id params" | +//! | [`Tier::Reusable`] | nothing (generalized) | forever | "ASP.NET verbose errors leak the ViewState key" | +//! +//! Promotion is evidence-gated and moves *up* the table: a working memo repeated +//! within a run becomes engagement knowledge; engagement knowledge confirmed on +//! a second run becomes technique knowledge; a technique memo that holds on two +//! **different targets** is generalized into reusable knowledge with the +//! target-specific tokens stripped. Nothing is promoted on a single observation, +//! because one observation is exactly how a hallucination looks. +//! +//! Recall is scored, not exhaustive: prompts have a budget, so [`Memory::recall`] +//! ranks by term overlap, how often the memo preceded a real finding, and +//! recency, then [`Memory::prompt_block`] renders the top few as plain lines. + +use serde::{Deserialize, Serialize}; +use std::collections::{BTreeMap, HashMap}; +use std::path::{Path, PathBuf}; + +/// Which memory a memo belongs to. See the module docs for the scoping rules. +#[derive(Clone, Copy, Debug, PartialEq, Eq, PartialOrd, Ord, Serialize, Deserialize)] +#[serde(rename_all = "kebab-case")] +pub enum Tier { + Working, + Engagement, + Technique, + Reusable, +} + +impl Tier { + pub fn as_str(&self) -> &'static str { + match self { + Tier::Working => "working", + Tier::Engagement => "engagement", + Tier::Technique => "technique", + Tier::Reusable => "reusable", + } + } +} + +/// One remembered fact. Deliberately a *sentence*, not a struct of fields: the +/// consumer is a language model, and the thing that has to survive the round +/// trip is the claim, not a schema. +#[derive(Clone, Debug, Serialize, Deserialize)] +pub struct Memo { + pub id: String, + pub tier: Tier, + /// Scope key — target key for engagement, technique id for technique, + /// empty for reusable. + #[serde(default)] + pub key: String, + pub text: String, + #[serde(default)] + pub tags: Vec, + #[serde(default)] + pub source_run: String, + #[serde(default)] + pub target: String, + /// Belief that the claim holds, 0..1. + #[serde(default)] + pub confidence: f64, + /// Distinct runs that observed this. + #[serde(default)] + pub observations: u32, + /// Times this memo was fed into a prompt. + #[serde(default)] + pub uses: u32, + /// Times a run that recalled this memo went on to produce a finding. + #[serde(default)] + pub wins: u32, + #[serde(default)] + pub created: u64, + #[serde(default)] + pub updated: u64, +} + +impl Memo { + /// Fraction of recalls that preceded a finding. Unused memos sit at the + /// neutral 0.5 rather than 0 — never having been tried is not evidence of + /// being wrong, and starting them at zero would bury them forever. + pub fn success_rate(&self) -> f64 { + if self.uses == 0 { + 0.5 + } else { + self.wins as f64 / self.uses as f64 + } + } +} + +fn now() -> u64 { + std::time::SystemTime::now() + .duration_since(std::time::UNIX_EPOCH) + .map(|d| d.as_secs()) + .unwrap_or(0) +} + +/// Lowercased alphanumeric terms of length ≥ 3, deduped. Used for both indexing +/// and query matching so a memo and a query are compared the same way. +pub fn terms(s: &str) -> Vec { + let mut out: Vec = Vec::new(); + for raw in s.split(|c: char| !c.is_alphanumeric() && c != '-' && c != '_' && c != '.') { + let t = raw.trim_matches(|c: char| c == '.' || c == '-' || c == '_').to_lowercase(); + if t.len() >= 3 && !out.contains(&t) { + out.push(t); + } + } + out +} + +/// Stable identity of a claim: same normalized wording = same memo, so repeating +/// an observation reinforces it instead of duplicating it. +fn fingerprint(tier: Tier, key: &str, text: &str) -> String { + let norm: String = text + .to_lowercase() + .chars() + .filter(|c| c.is_alphanumeric() || c.is_whitespace()) + .collect::() + .split_whitespace() + .collect::>() + .join(" "); + let mut h: u64 = 0xcbf2_9ce4_8422_2325; + for b in format!("{}|{}|{}", tier.as_str(), key, norm).bytes() { + h ^= b as u64; + h = h.wrapping_mul(0x1000_0000_01b3); + } + format!("{}-{:016x}", tier.as_str(), h) +} + +/// Normalize a target into a stable engagement key: scheme, port, path, `www.` +/// and case all drop out, so `https://WWW.Example.com:443/login` and +/// `http://example.com/` are one engagement and not two. +pub fn engagement_key(target: &str) -> String { + let t = target.trim().to_lowercase(); + let t = t.split_once("://").map(|(_, rest)| rest).unwrap_or(&t); + let t = t.split(['/', '?', '#']).next().unwrap_or(t); + let t = t.rsplit_once(':').map(|(h, p)| if p.chars().all(|c| c.is_ascii_digit()) { h } else { t }).unwrap_or(t); + t.trim_start_matches("www.").trim().to_string() +} + +/// What to recall for. +#[derive(Debug, Clone, Default)] +pub struct Query { + /// Free text — agent prompt, objective, endpoint, whatever is at hand. + pub text: String, + /// Restrict engagement recall to this target (empty = any). + pub target: String, + /// Restrict technique recall to these ids (empty = any). + pub techniques: Vec, + /// Tiers to search. Empty means all but [`Tier::Working`]. + pub tiers: Vec, + pub limit: usize, +} + +/// A scored recall hit. +#[derive(Debug, Clone)] +pub struct Hit { + pub memo: Memo, + pub score: f64, +} + +/// The four-tier store. Persisted under `/` as one file per tier; working +/// memory is written too, so a crashed run can be resumed with its scratchpad +/// intact instead of restarting cold. +#[derive(Default)] +pub struct Memory { + dir: Option, + working: Vec, + engagement: BTreeMap>, + technique: BTreeMap>, + reusable: Vec, + /// Fingerprints seen this run — the promotion gate for working → engagement. + seen_this_run: HashMap, + /// Memo ids injected into prompts during this run, so a run that lands a + /// finding can credit what it was told beforehand. + recalled: Vec, +} + +/// Process-wide store for the current project. +/// +/// Recall happens while prompts are built and reinforcement happens when the +/// run finishes — far apart in the call graph, with the async pipeline in +/// between. Two independently opened handles would each hold a stale copy and +/// the last one to save would silently discard the other's counters, so the +/// process shares one. The directory is bound on first call; later calls return +/// that same store regardless of the path passed, which is correct because one +/// CLI process serves one project. +pub fn shared(dir: impl AsRef) -> &'static std::sync::Mutex { + static STORE: std::sync::OnceLock> = std::sync::OnceLock::new(); + STORE.get_or_init(|| std::sync::Mutex::new(Memory::open(dir))) +} + +/// Working memory is a scratchpad, not a log: past this many memos the oldest +/// go, because a run that emits thousands of lines would otherwise turn recall +/// into a scan of its own noise. +const WORKING_CAP: usize = 400; +/// Confidence floor below which a never-useful memo is pruned on save. +const PRUNE_BELOW: f64 = 0.15; + +impl Memory { + /// In-memory only — used by tests and by callers with no project dir. + pub fn ephemeral() -> Memory { + Memory::default() + } + + /// Open (or create) the store under `dir`, e.g. `.neurosploit/memory`. + pub fn open(dir: impl AsRef) -> Memory { + let dir = dir.as_ref().to_path_buf(); + let _ = std::fs::create_dir_all(&dir); + let read = |name: &str| -> Option { std::fs::read_to_string(dir.join(name)).ok() }; + Memory { + working: read("working.json").and_then(|s| serde_json::from_str(&s).ok()).unwrap_or_default(), + engagement: read("engagement.json").and_then(|s| serde_json::from_str(&s).ok()).unwrap_or_default(), + technique: read("technique.json").and_then(|s| serde_json::from_str(&s).ok()).unwrap_or_default(), + reusable: read("reusable.json").and_then(|s| serde_json::from_str(&s).ok()).unwrap_or_default(), + dir: Some(dir), + seen_this_run: HashMap::new(), + recalled: Vec::new(), + } + } + + pub fn counts(&self) -> (usize, usize, usize, usize) { + ( + self.working.len(), + self.engagement.values().map(|v| v.len()).sum(), + self.technique.values().map(|v| v.len()).sum(), + self.reusable.len(), + ) + } + + fn bucket_mut(&mut self, tier: Tier, key: &str) -> &mut Vec { + match tier { + Tier::Working => &mut self.working, + Tier::Engagement => self.engagement.entry(key.to_string()).or_default(), + Tier::Technique => self.technique.entry(key.to_string()).or_default(), + Tier::Reusable => &mut self.reusable, + } + } + + /// Record a claim. Re-recording the same claim reinforces it (confidence + /// rises toward 1, observation count grows) instead of adding a duplicate, + /// which is what makes "seen twice" a meaningful promotion signal. + pub fn remember(&mut self, tier: Tier, key: &str, text: &str, tags: &[&str], target: &str, run: &str, confidence: f64) -> String { + let text = text.trim(); + if text.is_empty() { + return String::new(); + } + let id = fingerprint(tier, key, text); + *self.seen_this_run.entry(id.clone()).or_insert(0) += 1; + let ts = now(); + let bucket = self.bucket_mut(tier, key); + if let Some(m) = bucket.iter_mut().find(|m| m.id == id) { + // Bounded reinforcement: each repeat closes 35% of the remaining gap + // to certainty, so a claim asymptotically approaches — but never + // reaches — "known", which is the honest shape for an observation. + m.confidence = (m.confidence + 0.35 * (1.0 - m.confidence)).clamp(0.0, 0.99); + m.observations += 1; + m.updated = ts; + if m.source_run != run && !run.is_empty() { + m.source_run = run.to_string(); + } + for t in tags { + if !m.tags.iter().any(|x| x == t) { + m.tags.push((*t).to_string()); + } + } + return id; + } + bucket.push(Memo { + id: id.clone(), + tier, + key: key.to_string(), + text: text.to_string(), + tags: tags.iter().map(|s| s.to_string()).collect(), + source_run: run.to_string(), + target: target.to_string(), + confidence: confidence.clamp(0.0, 0.99), + observations: 1, + uses: 0, + wins: 0, + created: ts, + updated: ts, + }); + if tier == Tier::Working && self.working.len() > WORKING_CAP { + let drop = self.working.len() - WORKING_CAP; + self.working.drain(0..drop); + } + id + } + + /// Convenience: note something learned about the target during this run. + pub fn note(&mut self, target: &str, run: &str, text: &str, tags: &[&str]) -> String { + self.remember(Tier::Working, &engagement_key(target), text, tags, target, run, 0.5) + } + + /// Rank memos against a query. Scoring blends three signals that answer + /// three different questions: overlap ("is this about what I'm doing?"), + /// success rate ("did acting on it ever pay off?") and recency ("is it + /// still likely to be true?"). Confidence gates the whole thing, so a + /// once-observed guess cannot outrank a repeatedly confirmed fact. + pub fn recall(&self, q: &Query) -> Vec { + let qterms = terms(&q.text); + let tkey = engagement_key(&q.target); + let tiers: Vec = if q.tiers.is_empty() { + vec![Tier::Engagement, Tier::Technique, Tier::Reusable] + } else { + q.tiers.clone() + }; + let ts = now(); + let mut hits: Vec = Vec::new(); + + let mut consider = |m: &Memo| { + let mterms = terms(&format!("{} {}", m.text, m.tags.join(" "))); + let overlap = if qterms.is_empty() || mterms.is_empty() { + 0.0 + } else { + let inter = qterms.iter().filter(|t| mterms.contains(t)).count() as f64; + inter / (qterms.len() as f64).sqrt().max(1.0) / (mterms.len() as f64).sqrt().max(1.0) + }; + // Half-life of 30 days: a fingerprint from last week is worth more + // than one from last quarter, but never worthless. + let age_days = (ts.saturating_sub(m.updated)) as f64 / 86_400.0; + let recency = 0.5f64.powf(age_days / 30.0); + let score = m.confidence * (0.55 * overlap.min(1.0) + 0.25 * m.success_rate() + 0.20 * recency); + if score > 0.0 { + hits.push(Hit { memo: m.clone(), score }); + } + }; + + for tier in tiers { + match tier { + Tier::Working => self.working.iter().for_each(&mut consider), + Tier::Engagement => { + for (k, v) in &self.engagement { + if tkey.is_empty() || *k == tkey { + v.iter().for_each(&mut consider); + } + } + } + Tier::Technique => { + for (k, v) in &self.technique { + if q.techniques.is_empty() || q.techniques.iter().any(|t| t == k) { + v.iter().for_each(&mut consider); + } + } + } + Tier::Reusable => self.reusable.iter().for_each(&mut consider), + } + } + hits.sort_by(|a, b| b.score.partial_cmp(&a.score).unwrap_or(std::cmp::Ordering::Equal)); + let limit = if q.limit == 0 { 8 } else { q.limit }; + hits.truncate(limit); + hits + } + + /// Render recalled memos as a prompt section, and mark them used so their + /// success rate can be scored against what the run finds. Returns an empty + /// string when nothing is worth injecting — an empty "what you know" header + /// is worse than none, it invites the model to invent the contents. + pub fn prompt_block(&mut self, q: &Query) -> String { + let hits = self.recall(q); + if hits.is_empty() { + return String::new(); + } + let ids: Vec = hits.iter().map(|h| h.memo.id.clone()).collect(); + self.mark_used(&ids); + for id in &ids { + if !self.recalled.contains(id) { + self.recalled.push(id.clone()); + } + } + let mut out = String::from("## What NeuroSploit already knows (prior engagements)\n\nTreat as leads, not facts — verify before reporting.\n"); + for h in &hits { + out.push_str(&format!( + "- [{} · {:.0}%] {}\n", + h.memo.tier.as_str(), + h.memo.confidence * 100.0, + h.memo.text + )); + } + out + } + + fn all_mut(&mut self) -> impl Iterator { + self.working + .iter_mut() + .chain(self.engagement.values_mut().flatten()) + .chain(self.technique.values_mut().flatten()) + .chain(self.reusable.iter_mut()) + } + + pub fn mark_used(&mut self, ids: &[String]) { + for m in self.all_mut() { + if ids.iter().any(|i| *i == m.id) { + m.uses += 1; + } + } + } + + /// Credit every memo recalled during a run that produced findings. This is + /// what turns recall from a guess into a measurement over time. + pub fn mark_win(&mut self, ids: &[String]) { + for m in self.all_mut() { + if ids.iter().any(|i| *i == m.id) { + m.wins += 1; + m.confidence = (m.confidence + 0.1).min(0.99); + } + } + } + + /// Credit everything recalled during this run. Call it only when the run + /// actually produced findings — that is the whole signal. + pub fn credit_recalled(&mut self) { + let ids = std::mem::take(&mut self.recalled); + self.mark_win(&ids); + } + + /// Promote what this run proved, then persist. Called once at the end of a + /// run; `run` is the run id and `target` the engagement's target. + /// + /// Every step needs *independent* evidence, so nothing here can be triggered + /// twice by one loud observation: + /// - working → engagement: the claim recurred within the run; + /// - engagement → technique: it also carries a technique tag and has been + /// observed in more than one run; + /// - technique → reusable: it held on two different targets, and the + /// generalized copy has target-specific tokens stripped. + pub fn consolidate(&mut self, target: &str, run: &str) -> (usize, usize, usize) { + let key = engagement_key(target); + let mut to_engagement: Vec = Vec::new(); + for m in &self.working { + if self.seen_this_run.get(&m.id).copied().unwrap_or(0) >= 2 || m.observations >= 2 { + to_engagement.push(m.clone()); + } + } + let promoted_e = to_engagement.len(); + for m in to_engagement { + let tags: Vec<&str> = m.tags.iter().map(|s| s.as_str()).collect(); + self.remember(Tier::Engagement, &key, &m.text, &tags, target, run, m.confidence.max(0.55)); + } + self.working.clear(); + + // engagement → technique + let mut to_technique: Vec<(String, Memo)> = Vec::new(); + for memos in self.engagement.values() { + for m in memos { + if m.observations < 2 { + continue; + } + if let Some(t) = m.tags.iter().find(|t| t.starts_with("technique:") || t.starts_with("cwe:") || t.starts_with("agent:")) { + to_technique.push((t.clone(), m.clone())); + } + } + } + let promoted_t = to_technique.len(); + for (tech, m) in to_technique { + let tags: Vec<&str> = m.tags.iter().map(|s| s.as_str()).collect(); + self.remember(Tier::Technique, &tech, &m.text, &tags, target, run, m.confidence); + } + + // technique → reusable: needs two distinct targets. + let mut to_reusable: Vec = Vec::new(); + for memos in self.technique.values() { + let mut by_text: HashMap> = HashMap::new(); + for m in memos { + by_text.entry(generalize(&m.text)).or_default().push(m); + } + for (gen, group) in by_text { + let mut targets: Vec<&str> = group.iter().map(|m| m.target.as_str()).filter(|t| !t.is_empty()).collect(); + targets.sort_unstable(); + targets.dedup(); + if targets.len() >= 2 { + let best = group.iter().map(|m| m.confidence).fold(0.0f64, f64::max); + let mut m = (*group[0]).clone(); + m.text = gen; + m.confidence = best; + to_reusable.push(m); + } + } + } + let promoted_r = to_reusable.len(); + for m in to_reusable { + let tags: Vec<&str> = m.tags.iter().map(|s| s.as_str()).collect(); + self.remember(Tier::Reusable, "", &m.text, &tags, "", run, m.confidence); + } + + self.seen_this_run.clear(); + self.save(); + (promoted_e, promoted_t, promoted_r) + } + + /// Drop memos that were never useful and have decayed — otherwise a store + /// that only grows eventually recalls noise as readily as knowledge. + pub fn decay(&mut self, factor: f64) { + let f = factor.clamp(0.5, 1.0); + for m in self.all_mut() { + if m.wins == 0 { + m.confidence *= f; + } + } + let keep = |m: &Memo| m.confidence >= PRUNE_BELOW || m.wins > 0; + self.working.retain(keep); + self.reusable.retain(keep); + for v in self.engagement.values_mut() { + v.retain(keep); + } + for v in self.technique.values_mut() { + v.retain(keep); + } + self.engagement.retain(|_, v| !v.is_empty()); + self.technique.retain(|_, v| !v.is_empty()); + } + + /// Forget by substring across every tier. Returns how many went. + pub fn forget(&mut self, needle: &str) -> usize { + let n = needle.to_lowercase(); + if n.is_empty() { + return 0; + } + let before = self.counts(); + let drop = |m: &Memo| !(m.text.to_lowercase().contains(&n) || m.id == needle || m.key.to_lowercase() == n); + self.working.retain(drop); + self.reusable.retain(drop); + for v in self.engagement.values_mut() { + v.retain(drop); + } + for v in self.technique.values_mut() { + v.retain(drop); + } + self.engagement.retain(|_, v| !v.is_empty()); + self.technique.retain(|_, v| !v.is_empty()); + let after = self.counts(); + self.save(); + (before.0 + before.1 + before.2 + before.3) - (after.0 + after.1 + after.2 + after.3) + } + + pub fn save(&self) { + let Some(dir) = &self.dir else { return }; + let _ = std::fs::create_dir_all(dir); + let put = |name: &str, v: String| { + let _ = std::fs::write(dir.join(name), v); + }; + if let Ok(j) = serde_json::to_string_pretty(&self.working) { + put("working.json", j); + } + if let Ok(j) = serde_json::to_string_pretty(&self.engagement) { + put("engagement.json", j); + } + if let Ok(j) = serde_json::to_string_pretty(&self.technique) { + put("technique.json", j); + } + if let Ok(j) = serde_json::to_string_pretty(&self.reusable) { + put("reusable.json", j); + } + } + + /// Everything, newest first — for `/memory` in the REPL and the web console. + pub fn dump(&self) -> Vec { + let mut all: Vec = self + .working + .iter() + .chain(self.engagement.values().flatten()) + .chain(self.technique.values().flatten()) + .chain(self.reusable.iter()) + .cloned() + .collect(); + all.sort_by(|a, b| b.updated.cmp(&a.updated)); + all + } +} + +/// Strip target-specific tokens so a technique memo can be stated about the +/// class of system rather than the host it was first seen on. A claim that +/// still names one host is not a general lesson. +fn generalize(text: &str) -> String { + let mut out = String::with_capacity(text.len()); + for word in text.split_whitespace() { + let w = word.trim_matches(|c: char| c == ',' || c == ';'); + let looks_like_host = w.contains("://") + || (w.contains('.') && w.split('.').count() >= 3 && !w.ends_with('.')) + || w.chars().filter(|c| *c == '.').count() >= 3; + if looks_like_host { + out.push_str(""); + } else { + out.push_str(word); + } + out.push(' '); + } + out.trim().to_string() +} + +#[cfg(test)] +mod tests { + use super::*; + + #[test] + fn one_engagement_key_per_host_however_the_url_was_written() { + assert_eq!(engagement_key("https://WWW.Example.com:443/login?x=1"), "example.com"); + assert_eq!(engagement_key("http://example.com/"), "example.com"); + assert_eq!(engagement_key("10.0.0.7"), "10.0.0.7"); + } + + #[test] + fn repeating_a_claim_reinforces_it_instead_of_duplicating() { + let mut m = Memory::ephemeral(); + m.note("http://t.test", "run1", "/admin returns 302 to /login", &["endpoint"]); + m.note("http://t.test", "run1", "/admin returns 302 to /login", &["endpoint"]); + assert_eq!(m.counts().0, 1); + let memo = &m.working[0]; + assert_eq!(memo.observations, 2); + assert!(memo.confidence > 0.5, "second observation must raise confidence"); + } + + #[test] + fn a_claim_seen_once_is_not_promoted_but_a_repeat_is() { + let mut m = Memory::ephemeral(); + m.note("http://t.test", "run1", "seen once only", &[]); + m.note("http://t.test", "run1", "seen twice here", &[]); + m.note("http://t.test", "run1", "seen twice here", &[]); + let (to_eng, _, _) = m.consolidate("http://t.test", "run1"); + assert_eq!(to_eng, 1); + let eng = &m.engagement["t.test"]; + assert_eq!(eng.len(), 1); + assert_eq!(eng[0].text, "seen twice here"); + assert!(m.working.is_empty(), "working memory is cleared once consolidated"); + } + + #[test] + fn technique_knowledge_generalizes_only_after_a_second_target() { + let mut m = Memory::ephemeral(); + let claim = "verbose ASP.NET errors on https://a.example.com/x leak the stack trace"; + m.remember(Tier::Technique, "cwe:209", claim, &["cwe:209"], "https://a.example.com", "r1", 0.8); + assert_eq!(m.consolidate("https://a.example.com", "r1").2, 0, "one target is not a general lesson"); + + let claim2 = "verbose ASP.NET errors on https://b.other.org/y leak the stack trace"; + m.remember(Tier::Technique, "cwe:209", claim2, &["cwe:209"], "https://b.other.org", "r2", 0.8); + assert!(m.consolidate("https://b.other.org", "r2").2 >= 1); + assert!( + m.reusable.iter().any(|r| r.text.contains("")), + "the reusable copy must not name a specific host: {:?}", + m.reusable.iter().map(|r| &r.text).collect::>() + ); + } + + #[test] + fn recall_prefers_the_memo_that_matches_the_question() { + let mut m = Memory::ephemeral(); + m.remember(Tier::Engagement, "t.test", "login.aspx is vulnerable to SQL injection in tbUsername", &["cwe:89"], "http://t.test", "r1", 0.9); + m.remember(Tier::Engagement, "t.test", "the site serves a robots.txt with two entries", &["recon"], "http://t.test", "r1", 0.9); + let hits = m.recall(&Query { text: "sql injection on login".into(), target: "http://t.test".into(), limit: 1, ..Default::default() }); + assert_eq!(hits.len(), 1); + assert!(hits[0].memo.text.contains("SQL injection")); + } + + #[test] + fn an_empty_recall_injects_no_prompt_section() { + let mut m = Memory::ephemeral(); + assert_eq!(m.prompt_block(&Query { text: "anything".into(), ..Default::default() }), ""); + } + + #[test] + fn wins_raise_a_memo_above_an_equally_relevant_one() { + let mut m = Memory::ephemeral(); + let a = m.remember(Tier::Reusable, "", "idor on numeric order ids", &[], "", "r1", 0.8); + m.remember(Tier::Reusable, "", "idor on numeric invoice ids", &[], "", "r1", 0.8); + m.mark_used(&[a.clone()]); + m.mark_win(&[a.clone()]); + let hits = m.recall(&Query { text: "idor numeric ids".into(), limit: 2, ..Default::default() }); + assert_eq!(hits[0].memo.id, a, "the memo with a win must rank first"); + } + + #[test] + fn decay_drops_stale_never_useful_memos_and_keeps_proven_ones() { + let mut m = Memory::ephemeral(); + let keep = m.remember(Tier::Reusable, "", "proven lesson", &[], "", "r1", 0.5); + m.remember(Tier::Reusable, "", "never useful", &[], "", "r1", 0.2); + m.mark_win(&[keep.clone()]); + for _ in 0..6 { + m.decay(0.7); + } + assert!(m.reusable.iter().any(|x| x.id == keep)); + assert!(!m.reusable.iter().any(|x| x.text == "never useful")); + } + + #[test] + fn forget_removes_matching_memos_from_every_tier() { + let mut m = Memory::ephemeral(); + m.note("http://t.test", "r1", "secret token abc123 in page source", &[]); + m.remember(Tier::Reusable, "", "secret token patterns leak in source maps", &[], "", "r1", 0.6); + assert_eq!(m.forget("secret token"), 2); + assert_eq!(m.counts(), (0, 0, 0, 0)); + } +} diff --git a/neurosploit-rs/crates/harness/src/pipeline.rs b/neurosploit-rs/crates/harness/src/pipeline.rs index ba8f096..af1ff90 100644 --- a/neurosploit-rs/crates/harness/src/pipeline.rs +++ b/neurosploit-rs/crates/harness/src/pipeline.rs @@ -49,12 +49,67 @@ fn operator_directives(cfg: &RunConfig) -> String { if let Some(auth) = cfg.auth.as_deref().filter(|x| !x.trim().is_empty()) { s.push_str(&format!("AUTHENTICATION — test as an authenticated user; send this with each request: {auth}\n")); } + let recalled = memory_directives(cfg); + if !recalled.is_empty() { + s.push_str(&recalled); + } if !s.is_empty() { s.push('\n'); } s } +/// Where this project's durable state lives. The app already points +/// `vault_dir` at `/.neurosploit/vault`, so its parent is the project +/// store; a caller that set neither falls back to the run's own workdir, which +/// keeps a one-off run from writing into an unrelated directory. +pub(crate) fn proj_store(cfg: &RunConfig) -> PathBuf { + if let Some(v) = cfg.vault_dir.as_deref() { + if let Some(parent) = Path::new(v).parent() { + return parent.to_path_buf(); + } + } + cfg.workdir + .as_deref() + .map(PathBuf::from) + .unwrap_or_else(|| PathBuf::from(".neurosploit")) +} + +/// The run id used as provenance in the graph and memory: the workdir basename +/// (`ns--`), which is also what the report and the web console use. +pub(crate) fn run_id(cfg: &RunConfig) -> String { + cfg.workdir + .as_deref() + .and_then(|d| Path::new(d).file_name()) + .map(|n| n.to_string_lossy().to_string()) + .unwrap_or_default() +} + +/// Prior knowledge about this target, injected into recon/exploit prompts. +/// +/// An engagement is usually not the first look at a host, but every model call +/// starts blank — without this the harness re-derives the same stack, the same +/// endpoints and the same dead ends on every run. Recall is scored and capped +/// (see [`crate::memory`]) so the block stays a handful of lines, and it is +/// explicitly framed as leads to verify: prior belief must not become an +/// assertion the model is willing to report. +fn memory_directives(cfg: &RunConfig) -> String { + let dir = proj_store(cfg).join("memory"); + let q = crate::memory::Query { + text: format!( + "{} {} {}", + cfg.target, + cfg.objective.clone().unwrap_or_default(), + cfg.instructions.clone().unwrap_or_default() + ), + target: cfg.target.clone(), + limit: 6, + ..Default::default() + }; + let Ok(mut mem) = crate::memory::shared(&dir).lock() else { return String::new() }; + mem.prompt_block(&q) +} + /// Tool-usage doctrine prepended to recon/exploit prompts so the agent knows /// exactly what it may use. Best run on Kali Linux (or the Kali Docker image), /// where these tools are preinstalled. @@ -485,6 +540,12 @@ pub async fn run(cfg: RunConfig, lib: &Library, pool: &ModelPool, tx: Sender = ranked.into_iter().take(cap).collect(); let _ = tx.send(format!("selected {} specialist agents (RL-ranked)", selected.len())).await; let _ = tx.send("offline: no exploitation performed (provide API keys or --subscription to run live)".into()).await; + // Recon still learned something about the target even with no + // exploitation, and that is exactly the kind of knowledge the next run + // should not have to re-derive. + for n in absorb(&cfg, &recon, &[]) { + let _ = tx.send(n).await; + } let artifacts = persist(&cfg, &recon, "", &[]); return RunOutput { target: cfg.target.clone(), workdir: cfg.workdir.clone().unwrap_or_default(), findings: vec![], agents_ran: selected.iter().map(|a| a.name.clone()).collect(), candidates: 0, recon, artifacts }; } @@ -1394,6 +1455,13 @@ async fn finish(cfg: RunConfig, _lib: &Library, recon: String, transcript: Strin let _ = tx.send("RL rewards updated".into()).await; } + // Durable knowledge. Everything above this point is about *this* run; these + // two stores are what makes the next one start from further along. + let notes = absorb(&cfg, &recon, &findings); + for n in notes { + let _ = tx.send(n).await; + } + let artifacts = persist(&cfg, &recon, &transcript, &findings); if !artifacts.is_empty() { let _ = tx.send(format!("notify: evidence saved → {}", cfg.workdir.clone().unwrap_or_default())).await; @@ -1419,6 +1487,102 @@ async fn finish(cfg: RunConfig, _lib: &Library, recon: String, transcript: Strin } } +/// Fold a finished run into the attack knowledge graph and the layered memory. +/// +/// Returns the lines to report to the operator. It is deliberately synchronous +/// and returns its messages instead of sending them: the memory store is behind +/// a `std::sync::Mutex`, and holding that guard across an `.await` would make +/// the pipeline future non-`Send`. +fn absorb(cfg: &RunConfig, recon: &str, findings: &[Finding]) -> Vec { + let mut out = Vec::new(); + let rid = run_id(cfg); + let store = proj_store(cfg); + + // Graph: the project-wide one accumulates across runs, and each run keeps + // its own copy so the report and the web console can draw just this run. + let proj_graph = store.join("graph.json"); + let mut kg = crate::knowledge_graph::KnowledgeGraph::load(&proj_graph); + kg.ingest(&cfg.target, &rid, findings); + kg.save(&proj_graph); + if let Some(dir) = cfg.workdir.as_deref() { + let mut run_kg = crate::knowledge_graph::KnowledgeGraph::new(); + run_kg.ingest(&cfg.target, &rid, findings); + run_kg.save(Path::new(dir).join("graph.json")); + } + let paths = kg.paths(1); + let depth = paths.first().map(|(p, _)| p.len()).unwrap_or(0); + out.push(format!( + "knowledge graph: {} node(s), {} edge(s), longest attack path {} step(s) → graph.json", + kg.nodes.len(), + kg.edges.len(), + depth + )); + + // Memory: what was proven is engagement knowledge immediately (a validated + // finding is evidence, not a guess); everything else has to earn promotion. + let key = crate::memory::engagement_key(&cfg.target); + let Ok(mut mem) = crate::memory::shared(store.join("memory")).lock() else { return out }; + + for f in findings { + let where_ = if f.endpoint.is_empty() { cfg.target.as_str() } else { f.endpoint.as_str() }; + let text = format!("{} — {} on {} [{} · {}]", f.title, f.severity, where_, f.cwe, f.stage); + let mut tags: Vec = vec![format!("agent:{}", f.agent)]; + if !f.cwe.is_empty() { + tags.push(format!("cwe:{}", f.cwe)); + } + if !f.mitre.is_empty() { + tags.push(format!("technique:{}", f.mitre)); + } + if !f.stage.is_empty() { + tags.push(f.stage.clone()); + } + let refs: Vec<&str> = tags.iter().map(|s| s.as_str()).collect(); + mem.remember( + crate::memory::Tier::Engagement, + &key, + &text, + &refs, + &cfg.target, + &rid, + f.confidence.clamp(0.4, 0.95), + ); + } + + // Recon facts go to working memory: one sighting of an endpoint is a lead, + // and only a repeat earns a place in the engagement's knowledge. + if let Ok(v) = serde_json::from_str::(recon) { + for k in ["endpoints", "apis", "hosts", "subdomains"] { + if let Some(arr) = v.get(k).and_then(|x| x.as_array()) { + for item in arr.iter().take(25) { + let s = item.as_str().map(|s| s.to_string()).unwrap_or_else(|| item.to_string()); + let s = s.trim_matches('"').trim(); + if s.len() > 2 { + mem.note(&cfg.target, &rid, &format!("{k}: {s}"), &["recon", k]); + } + } + } + } + if let Some(tech) = v.get("tech").and_then(|x| x.as_array()) { + let list: Vec = tech.iter().filter_map(|t| t.as_str().map(|s| s.to_string())).collect(); + if !list.is_empty() { + mem.note(&cfg.target, &rid, &format!("stack: {}", list.join(", ")), &["recon", "tech"]); + } + } + } + + // Recall only earns credit when the run it informed actually found something. + if !findings.is_empty() { + mem.credit_recalled(); + } + mem.decay(0.98); + let (e, t, r) = mem.consolidate(&cfg.target, &rid); + let (w, eng, tech, reuse) = mem.counts(); + out.push(format!( + "memory: +{e} engagement, +{t} technique, +{r} reusable (now {w}/{eng}/{tech}/{reuse}) → memory/" + )); + out +} + /// Write recon/exploit/findings/report as json+md for downstream reuse. fn persist(cfg: &RunConfig, recon: &str, transcript: &str, findings: &[Finding]) -> Vec { let Some(dir) = &cfg.workdir else { return vec![] }; diff --git a/neurosploit-rs/crates/harness/src/pool.rs b/neurosploit-rs/crates/harness/src/pool.rs index 49647f5..101983b 100644 --- a/neurosploit-rs/crates/harness/src/pool.rs +++ b/neurosploit-rs/crates/harness/src/pool.rs @@ -86,6 +86,9 @@ pub struct ModelPool { /// When this exceeds `AUTH_FAIL_THRESHOLD`, the pool auto-pauses instead of /// burning through the remaining agents on a dead token. consecutive_auth_fails: Arc, + /// Backends already tried as an automatic fallback, so a failing one is not + /// retried in a loop. + tried_auto: Arc>>, } impl ModelPool { @@ -119,6 +122,7 @@ impl ModelPool { resume: Arc::new(Notify::new()), fallback: Arc::new(Mutex::new(Vec::new())), consecutive_auth_fails: Arc::new(std::sync::atomic::AtomicUsize::new(0)), + tried_auto: Arc::new(Mutex::new(Vec::new())), } } @@ -325,9 +329,28 @@ impl ModelPool { } } } - // Every candidate failed. Park the run (keeping all state) so the user - // can fix auth or wait for quota renewal, then /continue. + // Every configured candidate failed. Before parking the run and + // waiting for a human, use whatever else this machine can actually + // reach — another logged-in CLI subscription, or a provider whose + // API key is in the environment. A run that stops because one + // provider ran out of quota, on a box with three other usable + // backends, is a run that stopped for no reason. if (auth_failed || exhausted) && !self.is_cancelled() { + if let Some(alt) = self.next_auto_fallback(&order) { + if let Some(tx) = self.progress() { + let _ = tx.send(format!( + "notify: ⇄ {} unavailable — falling back to {}:{} and continuing.", + order.first().map(|m| m.provider.clone()).unwrap_or_default(), + alt.provider, alt.model + )).await; + } + if let Ok(mut fb) = self.fallback.lock() { + fb.insert(0, alt.clone()); + } + self.reset_auth_fails(); + continue; + } + // Nothing else is reachable — now a human really is required. self.park_exhausted(&last, auth_failed).await; continue; } @@ -335,6 +358,48 @@ impl ModelPool { } } + /// A backend this machine can use right now that is not already in `tried` + /// and not already a candidate. + /// + /// Two sources, in this order: a subscription CLI that is installed (the + /// operator already logged into it, and it costs no API key), then any + /// provider whose API key is present in the environment. Each is offered + /// once — a backend that also fails is recorded so the loop cannot spin. + pub fn next_auto_fallback(&self, current: &[ModelRef]) -> Option { + let mut tried = self.tried_auto.lock().ok()?; + let known = |p: &str, m: &str, tried: &Vec| { + current.iter().any(|c| c.provider == p && c.model == m) || tried.iter().any(|t| t == &format!("{p}:{m}")) + }; + let installed = crate::models::installed_cli_backends(); + for pr in crate::models::providers() { + if pr.kind != "cli" { + continue; + } + let Some(bin) = crate::models::cli_binary_for(pr.key) else { continue }; + if !installed.contains(&bin) { + continue; + } + let Some(model) = pr.models.first() else { continue }; + if known(pr.key, model, &tried) { + continue; + } + tried.push(format!("{}:{}", pr.key, model)); + return Some(ModelRef { provider: pr.key.to_string(), model: (*model).to_string() }); + } + for pr in crate::models::providers() { + if std::env::var(pr.env_key).ok().filter(|v| !v.trim().is_empty()).is_none() { + continue; + } + let Some(model) = pr.models.first() else { continue }; + if known(pr.key, model, &tried) { + continue; + } + tried.push(format!("{}:{}", pr.key, model)); + return Some(ModelRef { provider: pr.key.to_string(), model: (*model).to_string() }); + } + None + } + /// Reorder candidates for a task. With a single-model panel this is a no-op. pub fn route(&self, task: Task) -> Vec { let mut order = self.candidates.clone(); @@ -457,6 +522,42 @@ pub fn quorum_confirmed(severity: &str, yes: usize, total: usize) -> bool { #[cfg(test)] mod verdict_tests { + /// Whatever this machine happens to have installed, the automatic fallback + /// must never re-offer a model already in the panel and never offer the + /// same one twice — either turns "keep going" into a spin. + #[test] + fn auto_fallback_never_repeats_itself_or_the_current_panel() { + let current = vec![ModelRef::parse("anthropic:claude-opus-4-8")]; + let pool = ModelPool::new(current.clone(), 1); + let mut seen: Vec = Vec::new(); + for _ in 0..8 { + let Some(m) = pool.next_auto_fallback(¤t) else { break }; + let id = format!("{}:{}", m.provider, m.model); + assert!( + !(m.provider == "anthropic" && m.model == "claude-opus-4-8"), + "offered the model that just failed" + ); + assert!(!seen.contains(&id), "offered {id} twice"); + seen.push(id); + } + } + + /// A provider whose key is in the environment is reachable, so it must be + /// offered before the run parks and waits for a human. + #[test] + fn a_provider_with_a_key_in_the_environment_is_offered() { + std::env::set_var("DEEPSEEK_API_KEY", "test-key-for-fallback"); + let current = vec![ModelRef::parse("anthropic:claude-opus-4-8")]; + let pool = ModelPool::new(current.clone(), 1); + let mut found = false; + for _ in 0..30 { + let Some(m) = pool.next_auto_fallback(¤t) else { break }; + if m.provider == "deepseek" { found = true; break; } + } + std::env::remove_var("DEEPSEEK_API_KEY"); + assert!(found, "a provider with a usable API key must be reachable as a fallback"); + } + use super::*; #[test] fn parses_json_and_prose() { diff --git a/web/public/app.js b/web/public/app.js index 5d2197a..726358f 100644 --- a/web/public/app.js +++ b/web/public/app.js @@ -534,6 +534,7 @@ function attachLiveJob(id, target, name, pinnedAgents) { show($('#wizardView'), false); show($('#detailView'), false); + show($('#dashView'), false); show($('#liveView'), true); $('#liveTarget').textContent = name || target || '—'; $('#liveTargetSub').textContent = name ? target : ''; @@ -772,9 +773,9 @@ function leaveLiveJob() { state.currentJob = null; termSyncTargets(); } -$('#btnBackToBoard').addEventListener('click', () => { leaveLiveJob(); show($('#liveView'), false); show($('#wizardView'), true); }); -$('#btnDetailBack').addEventListener('click', () => { clearInterval(state.detailPoll); show($('#detailView'), false); show($('#wizardView'), true); }); -$('#btnNewEngagement').addEventListener('click', () => { leaveLiveJob(); clearInterval(state.detailPoll); show($('#detailView'), false); show($('#liveView'), false); show($('#wizardView'), true); }); +$('#btnBackToBoard').addEventListener('click', () => { leaveLiveJob(); show($('#liveView'), false); show($('#dashView'), false); show($('#wizardView'), true); }); +$('#btnDetailBack').addEventListener('click', () => { clearInterval(state.detailPoll); show($('#detailView'), false); show($('#dashView'), false); show($('#wizardView'), true); }); +$('#btnNewEngagement').addEventListener('click', () => { leaveLiveJob(); clearInterval(state.detailPoll); show($('#detailView'), false); show($('#liveView'), false); show($('#dashView'), false); show($('#wizardView'), true); }); // --------------------------------------------------------------------------- // Generative Attack Path Chaining @@ -860,7 +861,39 @@ function openFindingModal(f, pocs, runId) { $('#btnCloseFinding').addEventListener('click', () => show($('#findingModal'), false)); $('#findingModal').addEventListener('click', (e) => { if (e.target.id === 'findingModal') show($('#findingModal'), false); }); -const KILL_CHAIN_STAGES = ['recon', 'initial-access', 'execution', 'privesc', 'lateral', 'exfil', 'impact']; +// --------------------------------------------------------------------------- +// Generative Attack Path Chaining +// +// The graph answers one question: how does an attacker get from the target to +// impact? Three things it must not do, each of which the first version did: +// +// 1. **Drop findings.** Stages were matched against a hardcoded list of seven, +// so anything the harness emitted outside it (`credential-access`, +// `discovery`, `persistence`, …) silently vanished — 5 of 27 findings on a +// real run. The stage list now mirrors `knowledge_graph::STAGES`, and any +// unknown stage still gets its own column rather than being discarded. +// 2. **Blur into unreadable boxes.** Titles were cut at 22 characters, so a +// column read "SQL Injection Authent…" six times. Nodes now wrap onto two +// lines and carry CWE / technique / exploitability. +// 3. **Present a guess as evidence.** Agents only sometimes fill `chains_from`. +// Without it every node fanned off the root, which looks like a chain and +// is not one. Inferred progression edges are drawn dashed, counted +// separately in the toolbar, and can be hidden. +// +// When a run wrote `graph.json` (the harness's own knowledge graph), its edges +// are used verbatim — including which ones it inferred. Older runs fall back to +// deriving the same shape client-side, so the view degrades rather than empties. +// --------------------------------------------------------------------------- + +// Mirrors knowledge_graph::STAGES on the Rust side. Order = attack progression. +const KILL_CHAIN_STAGES = [ + 'recon', 'discovery', 'initial-access', 'execution', 'persistence', + 'privesc', 'credential-access', 'lateral', 'collection', 'exfil', 'impact', +]; +const stageRank = (s) => { + const i = KILL_CHAIN_STAGES.indexOf(s); + return i === -1 ? KILL_CHAIN_STAGES.length : i; +}; // Same severity tokens the rest of the console uses — the graph canvas // follows the light/dark theme instead of a fixed dark palette. @@ -877,101 +910,350 @@ function nodeIcon(f) { return '⚠'; } -// Generative Attack Path Chaining — a real node graph (root = target, one -// node per confirmed finding, edges from chains_from when the harness set -// it, else fanned from root) instead of flat cards, so a single finding -// still reads as a graph and not an empty list. -function renderAttackPath(container, findings, target) { - if (!findings.length) { +/// Greedy wrap into at most `lines` lines of `max` chars, ellipsizing the tail. +function wrapLabel(s, max, lines) { + const words = String(s || '').split(/\s+/).filter(Boolean); + const out = []; + let cur = ''; + for (const w of words) { + const next = cur ? `${cur} ${w}` : w; + if (next.length <= max) { cur = next; continue; } + if (out.length === lines - 1) { cur = `${next.slice(0, max - 1)}…`; break; } + out.push(cur || w.slice(0, max)); + cur = cur ? w : ''; + } + if (cur) out.push(cur); + return out.slice(0, lines); +} + +/// Chain edges between findings, and where they came from. +/// Returns `{ edges: [{from, to, inferred}], source }` with indices into +/// `findings`, so the caller can tell the operator what it is looking at. +function chainEdges(findings, graph) { + const byId = new Map(findings.map((f, i) => [f.id, i])); + + // 1. The harness's own graph, when the run wrote one. + if (graph?.edges?.length) { + const nodeToFinding = new Map(); + for (const [id, n] of Object.entries(graph.nodes || {})) { + const fid = n.meta?.finding_id; + if (n.kind === 'finding' && fid !== undefined && byId.has(fid)) nodeToFinding.set(id, byId.get(fid)); + } + const edges = []; + for (const e of graph.edges) { + if (e.kind !== 'chains') continue; + const a = nodeToFinding.get(e.from), b = nodeToFinding.get(e.to); + if (a !== undefined && b !== undefined && a !== b) edges.push({ from: a, to: b, inferred: !!e.inferred }); + } + if (edges.length) return { edges, source: edges.every((e) => e.inferred) ? 'graph-inferred' : 'graph' }; + } + + // 2. Edges the agents asserted on the findings themselves. + const asserted = []; + findings.forEach((f, i) => { + for (const src of f.chains_from || []) { + const a = byId.get(src); + if (a !== undefined && a !== i) asserted.push({ from: a, to: i, inferred: false }); + } + }); + if (asserted.length) return { edges: asserted, source: 'asserted' }; + + // 3. Derive progression the same way the harness does: forward only, between + // adjacent populated stages, from the strongest finding of the earlier one. + // A full cross-product would look richer and mean nothing. + const byStage = new Map(); + findings.forEach((f, i) => { + const r = stageRank(f.stage || ''); + if (!byStage.has(r)) byStage.set(r, []); + byStage.get(r).push(i); + }); + const ranks = [...byStage.keys()].sort((a, b) => a - b); + const weight = (i) => (4 - sevRank(findings[i].severity)) * (findings[i].confidence || 0.5); + const edges = []; + for (let k = 0; k + 1 < ranks.length; k++) { + const from = byStage.get(ranks[k]).slice().sort((a, b) => weight(b) - weight(a))[0]; + for (const to of byStage.get(ranks[k + 1])) edges.push({ from, to, inferred: true }); + } + return { edges, source: edges.length ? 'derived' : 'none' }; +} + +const AP_SEV_FILTERS = ['all', 'critical', 'high', 'medium', 'low']; + +function renderAttackPath(container, allFindings, target, graph) { + const view = (container.__ap = container.__ap || { k: 1, tx: 0, ty: 0, sev: 'all', hideInferred: false, fitted: false }); + + if (!allFindings.length) { container.innerHTML = '
The attack path builds automatically as findings chain together — nothing confirmed yet.
'; return; } - const byId = new Map(findings.filter((f) => f.id).map((f) => [f.id, f])); - const hasStages = findings.some((f) => f.stage); - let groups; - if (hasStages) { - groups = KILL_CHAIN_STAGES - .map((stage) => ({ label: stage.replace('-', ' '), items: findings.filter((f) => (f.stage || '') === stage) })) - .filter((g) => g.items.length); - const other = findings.filter((f) => !f.stage); - if (other.length) groups.push({ label: 'unstaged', items: other }); - } else { - groups = [{ label: 'confirmed findings', items: findings }]; + + const minRank = view.sev === 'all' ? 99 : sevRank(view.sev); + const findings = view.sev === 'all' ? allFindings : allFindings.filter((f) => sevRank(f.severity) <= minRank); + if (!findings.length) { + container.innerHTML = `
No ${esc(view.sev)}-or-higher finding to chain.
`; + container.querySelector('[data-ap-reset]')?.addEventListener('click', () => { view.sev = 'all'; renderAttackPath(container, allFindings, target, graph); }); + return; } - // Layout: root at column 0; each kill-chain stage is its own column. - const COL_W = 210, ROW_H = 78, NODE_W = 176, NODE_H = 54, PAD = 40; - const rootX = PAD, rootY = PAD + (Math.max(...groups.map((g) => g.items.length)) * ROW_H) / 2; - const nodes = [{ id: '__root', x: rootX, y: rootY, root: true, label: target || 'target' }]; - const nodeById = new Map(); // finding.id -> node (for chains_from edges) - groups.forEach((g, ci) => { - const colX = PAD + NODE_W / 2 + (ci + 1) * COL_W; - const colH = g.items.length * ROW_H; - const offsetY = rootY - colH / 2 + ROW_H / 2; - g.items.forEach((f, ri) => { - const node = { id: f.id || `${ci}-${ri}`, x: colX, y: offsetY + ri * ROW_H, finding: f, stageLabel: g.label }; - nodes.push(node); - if (f.id) nodeById.set(f.id, node); - }); + const model = chainEdges(findings, graph); + const edges = view.hideInferred ? model.edges.filter((e) => !e.inferred) : model.edges; + + // ---- columns: one per stage actually present, in progression order -------- + const groups = new Map(); + findings.forEach((f, i) => { + const key = (f.stage || '').trim() || 'unstaged'; + if (!groups.has(key)) groups.set(key, []); + groups.get(key).push(i); }); - - const edges = []; - for (const n of nodes) { - if (n.root) continue; - const parents = (n.finding.chains_from || []).map((cid) => nodeById.get(cid)).filter(Boolean); - if (parents.length) parents.forEach((p) => edges.push([p, n])); - else edges.push([nodes[0], n]); + const cols = [...groups.entries()].sort((a, b) => { + const ra = a[0] === 'unstaged' ? 999 : stageRank(a[0]); + const rb = b[0] === 'unstaged' ? 999 : stageRank(b[0]); + return ra - rb || a[0].localeCompare(b[0]); + }); + for (const [, idxs] of cols) { + idxs.sort((a, b) => sevRank(findings[a].severity) - sevRank(findings[b].severity) || (findings[b].confidence || 0) - (findings[a].confidence || 0)); } - const width = PAD * 2 + NODE_W + (groups.length) * COL_W; - const height = Math.max(...nodes.map((n) => n.y)) + NODE_H + PAD; + const NODE_W = 236, NODE_H = 66, COL_GAP = 88, ROW_GAP = 14, PAD = 28, HEAD_H = 34, ROOT_W = 132; + const rowH = NODE_H + ROW_GAP; + const tallest = Math.max(...cols.map(([, v]) => v.length)); + const bodyH = tallest * rowH; + const pos = new Map(); // finding index -> {x, y} + cols.forEach(([, idxs], ci) => { + const x = PAD + ROOT_W + ci * (NODE_W + COL_GAP); + const colH = idxs.length * rowH; + const top = PAD + HEAD_H + (bodyH - colH) / 2; + idxs.forEach((fi, ri) => pos.set(fi, { x, y: top + ri * rowH })); + }); + const width = PAD * 2 + ROOT_W + cols.length * NODE_W + Math.max(0, cols.length - 1) * COL_GAP; + const height = PAD * 2 + HEAD_H + bodyH; + const rootY = PAD + HEAD_H + bodyH / 2; - const edgePath = (a, b) => { - const x1 = a.root ? a.x + 14 : a.x + NODE_W / 2, y1 = a.y; - const x2 = b.x - NODE_W / 2, y2 = b.y; - const midX = (x1 + x2) / 2; - return `M ${x1},${y1} C ${midX},${y1} ${midX},${y2} ${x2},${y2}`; + const hasParent = new Set(edges.map((e) => e.to)); + const roots = findings.map((_, i) => i).filter((i) => !hasParent.has(i)); + + const curve = (x1, y1, x2, y2) => { + const mid = (x1 + x2) / 2; + return `M ${x1},${y1} C ${mid},${y1} ${mid},${y2} ${x2},${y2}`; }; - const nodeSvg = (n) => { - if (n.root) { - return ` - - 🎯 - ${esc(trimMid(n.label, 26))} - `; - } - const f = n.finding; + const edgeSvg = edges.map((e, i) => { + const a = pos.get(e.from), b = pos.get(e.to); + if (!a || !b) return ''; + return ``; + }).join(''); + + const rootEdges = roots.map((i) => { + const b = pos.get(i); + return ``; + }).join(''); + + const nodeSvg = findings.map((f, i) => { + const p = pos.get(i); + if (!p) return ''; const color = canvasColor(f.severity); - const x = n.x - NODE_W / 2, y = n.y - NODE_H / 2; - return ` - - ${nodeIcon(f)} - ${esc(trimMid(f.title, 22))} - ${esc((f.mitre || f.owasp || f.cwe || n.stageLabel || '').slice(0, 26))} - + const title = wrapLabel(f.title, 30, 2); + const meta = [f.cwe, f.mitre || f.owasp, f.exploitability].filter(Boolean).join(' · '); + return ` + + + ${nodeIcon(f)} + ${title.map((line, li) => `${esc(line)}`).join('')} + ${esc(meta || f.agent)} + ${f.confidence ? f.confidence.toFixed(2) : ''} `; - }; + }).join(''); + + // Column headers name the stage and count it — unlabelled separator lines + // made the columns unreadable, which defeats a kill-chain layout entirely. + const headSvg = cols.map(([stage, idxs], ci) => { + const x = PAD + ROOT_W + ci * (NODE_W + COL_GAP); + return ` + + ${esc(stage.replace(/-/g, ' '))} + ${idxs.length} + `; + }).join(''); + + const inferredCount = model.edges.filter((e) => e.inferred).length; + const provenance = { + graph: 'chain edges from this run\'s knowledge graph', + 'graph-inferred': 'no chain was asserted — progression inferred by the harness', + asserted: 'chain edges asserted by the agents', + derived: 'no chain was asserted — progression inferred from kill-chain stages', + none: 'no chain links', + }[model.source]; container.innerHTML = ` - ${!hasStages ? '
No kill-chain stage data yet — shown as a flat graph from the target.
' : ''} -
- - ${groups.map((g, ci) => ``).join('')} - ${edges.map(([a, b]) => ``).join('')} - ${nodes.map(nodeSvg).join('')} - +
+
+ ${findings.length} finding(s) · ${cols.length} stage(s) · ${edges.length} link(s)${inferredCount ? ` (${inferredCount} inferred)` : ''} · ${roots.length} entry point(s) +
+
+ + +
+ + + +
- `; - container.querySelectorAll('.ap-node-g').forEach((g) => g.addEventListener('click', () => { - const f = findings[Number(g.dataset.idx)]; - const isLive = container.id === 'liveAttackPath'; - const pocs = isLive ? (state.currentJob?.pocs || []) : (state.detailPocs || []); - const runId = isLive ? state.currentJob?.runId : state.currentDetailId; - if (f) openFindingModal(f, pocs, runId); - })); -} +
${esc(provenance)} — inferred links are hypotheses, drawn dashed.
+
+ + + + + + + + ${headSvg} + ${rootEdges} + ${edgeSvg} + + + 🎯 + target + ${esc(trimMid(target || '', 14))} + + ${nodeSvg} + + +
drag to pan · scroll to zoom · click a node for the finding
+
+
+ ${['critical', 'high', 'medium', 'low', 'info'].map((s) => `${s}`).join('')} + asserted chain + inferred +
`; + const svg = container.querySelector('#apSvg'); + const pan = container.querySelector('#apPan'); + const wrap = container.querySelector('#apWrap'); + + const apply = () => pan.setAttribute('transform', `translate(${view.tx},${view.ty}) scale(${view.k})`); + const fit = () => { + const box = wrap.getBoundingClientRect(); + // The panel is rendered while its tab is still hidden, so the box measures + // zero and a "fit" there would lock in a garbage scale. Report the failure + // so the caller can try again once the tab is actually on screen. + if (!box.width || !box.height) return false; + view.k = Math.min(box.width / width, box.height / height, 1); + view.tx = (box.width - width * view.k) / 2; + view.ty = (box.height - height * view.k) / 2; + apply(); + return true; + }; + apply(); + // Auto-fit once, and only once it can actually measure: re-fitting on every + // live finding would yank the canvas out from under someone mid-inspection. + if (!view.fitted) { + const tryFit = () => { if (fit()) view.fitted = true; }; + requestAnimationFrame(tryFit); + if (window.ResizeObserver) { + const ro = new ResizeObserver(() => { if (view.fitted) ro.disconnect(); else tryFit(); }); + ro.observe(wrap); + } + } + + container.querySelector('[data-ap-fit]').addEventListener('click', fit); + container.querySelectorAll('[data-ap-zoom]').forEach((b) => b.addEventListener('click', () => { + const box = wrap.getBoundingClientRect(); + const factor = Number(b.dataset.apZoom) > 0 ? 1.2 : 1 / 1.2; + const cx = box.width / 2, cy = box.height / 2; + view.tx = cx - (cx - view.tx) * factor; + view.ty = cy - (cy - view.ty) * factor; + view.k = Math.max(0.15, Math.min(3, view.k * factor)); + apply(); + })); + container.querySelector('#apHideInferred').addEventListener('change', (e) => { + view.hideInferred = e.target.checked; + renderAttackPath(container, allFindings, target, graph); + }); + container.querySelector('#apSev').addEventListener('change', (e) => { + view.sev = e.target.value; + renderAttackPath(container, allFindings, target, graph); + }); + + wrap.addEventListener('wheel', (e) => { + e.preventDefault(); + const box = wrap.getBoundingClientRect(); + const mx = e.clientX - box.left, my = e.clientY - box.top; + const factor = e.deltaY < 0 ? 1.12 : 1 / 1.12; + view.tx = mx - (mx - view.tx) * factor; + view.ty = my - (my - view.ty) * factor; + view.k = Math.max(0.15, Math.min(3, view.k * factor)); + apply(); + }, { passive: false }); + + let dragging = false, sx = 0, sy = 0, moved = 0; + wrap.addEventListener('pointerdown', (e) => { + dragging = true; moved = 0; sx = e.clientX - view.tx; sy = e.clientY - view.ty; + wrap.setPointerCapture(e.pointerId); + wrap.classList.add('dragging'); + }); + wrap.addEventListener('pointermove', (e) => { + if (!dragging) return; + moved += Math.abs(e.movementX) + Math.abs(e.movementY); + view.tx = e.clientX - sx; view.ty = e.clientY - sy; + apply(); + }); + const endDrag = (e) => { dragging = false; wrap.classList.remove('dragging'); if (e?.pointerId !== undefined) { try { wrap.releasePointerCapture(e.pointerId); } catch { /* already released */ } } }; + wrap.addEventListener('pointerup', endDrag); + wrap.addEventListener('pointercancel', endDrag); + + // Highlight the whole path through a node, both directions — the question a + // reader has in front of a graph is "what led here, and where does it go". + const up = new Map(), down = new Map(); + edges.forEach((e) => { + if (!down.has(e.from)) down.set(e.from, []); + down.get(e.from).push(e.to); + if (!up.has(e.to)) up.set(e.to, []); + up.get(e.to).push(e.from); + }); + const reach = (start, map) => { + const seen = new Set(), stack = [start]; + while (stack.length) { + const n = stack.pop(); + for (const m of map.get(n) || []) if (!seen.has(m)) { seen.add(m); stack.push(m); } + } + return seen; + }; + const focusOn = (idx) => { + const set = new Set([idx, ...reach(idx, up), ...reach(idx, down)]); + svg.classList.add('has-focus'); + container.querySelectorAll('.ap-node-g').forEach((g) => g.classList.toggle('focus', set.has(Number(g.dataset.idx)))); + container.querySelectorAll('.ap-edge').forEach((p) => { + const f = Number(p.dataset.from), t = Number(p.dataset.to); + const r = Number(p.dataset.rootTo); + p.classList.toggle('focus', (set.has(f) && set.has(t)) || (!Number.isNaN(r) && r === idx)); + }); + }; + const clearFocus = () => { + svg.classList.remove('has-focus'); + container.querySelectorAll('.focus').forEach((el) => el.classList.remove('focus')); + }; + + container.querySelectorAll('.ap-node-g').forEach((g) => { + const idx = Number(g.dataset.idx); + g.addEventListener('mouseenter', () => focusOn(idx)); + g.addEventListener('focus', () => focusOn(idx)); + g.addEventListener('mouseleave', clearFocus); + g.addEventListener('blur', clearFocus); + const open = () => { + const isLive = container.id === 'liveAttackPath'; + const pocs = isLive ? (state.currentJob?.pocs || []) : (state.detailPocs || []); + const runId = isLive ? state.currentJob?.runId : state.currentDetailId; + openFindingModal(findings[idx], pocs, runId); + }; + // A pan that ends on a node is not a click on it. + g.addEventListener('click', () => { if (moved < 6) open(); }); + g.addEventListener('keydown', (e) => { if (e.key === 'Enter' || e.key === ' ') { e.preventDefault(); open(); } }); + }); +} function trimMid(s, n) { s = String(s || ''); return s.length > n ? s.slice(0, n - 1) + '…' : s; @@ -997,34 +1279,80 @@ function stepClassFor(phase, step) { return 'pending'; } +/// Group runs into target folders. Twelve runs of three hosts was a flat list +/// of twelve near-identical rows; the host is what an operator actually scans +/// for, so it becomes the folder and the runs live inside it. +function runFolders(runs) { + const folders = new Map(); + for (const r of runs) { + const key = engagementKey(r.target || r.id); + if (!folders.has(key)) folders.set(key, { key, items: [], ts: 0, findings: 0, severities: {} }); + const f = folders.get(key); + f.items.push(r); + f.ts = Math.max(f.ts, r.ts || 0); + f.findings += r.findings || 0; + for (const [k, n] of Object.entries(r.severities || {})) f.severities[k] = (f.severities[k] || 0) + n; + } + for (const f of folders.values()) f.items.sort((a, b) => (b.ts || 0) - (a.ts || 0)); + return [...folders.values()].sort((a, b) => b.ts - a.ts); +} + +/// Same normalization the harness uses for its engagement key, so a folder here +/// and an engagement in the harness's memory mean the same thing. +function engagementKey(target) { + const t = String(target || '').trim().toLowerCase(); + const noScheme = t.includes('://') ? t.split('://')[1] : t; + const host = noScheme.split(/[/?#]/)[0].replace(/:\d+$/, '').replace(/^www\./, ''); + return host || t || 'unknown'; +} + +function worstSeverity(severities) { + return SEV_ORDER.find((s) => Object.entries(severities || {}).some(([k, n]) => n && SEV_ORDER[sevRank(k)] === s)); +} + +const SB_OPEN_KEY = 'ns-sb-open'; +function openFolders() { + try { return new Set(JSON.parse(localStorage.getItem(SB_OPEN_KEY) || '[]')); } catch { return new Set(); } +} +function setFolderOpen(key, on) { + const s = openFolders(); + if (on) s.add(key); else s.delete(key); + localStorage.setItem(SB_OPEN_KEY, JSON.stringify([...s])); +} + +function runButton(r) { + const btn = document.createElement('button'); + btn.className = 'sb-run' + (state.currentDetailId === r.id ? ' active' : ''); + // Every line here truncates: a long target URL used to run past the + // sidebar's edge and collide with the main pane. + const worst = worstSeverity(r.severities); + btn.innerHTML = ` + ${worst ? `` : ''}${esc(r.name || r.target)} + ${r.name ? esc(r.target) : esc(r.id)} + ${r.findings} finding${r.findings === 1 ? '' : 's'}${esc(timeAgo(r.ts))}`; + btn.title = `${r.name ? r.name + '\n' : ''}${r.target}\n${r.id}${r.ts ? '\n' + new Date(r.ts * 1000).toLocaleString() : ''}`; + btn.addEventListener('click', () => openRun(r)); + return btn; +} + function renderSidebar() { const root = $('#sbGroups'); root.innerHTML = ''; - const running = state.runs.filter((r) => r.state === 'running'); - const completed = state.runs.filter((r) => r.state !== 'running'); - const groups = [{ label: 'Running', items: running }, { label: 'Completed', items: completed }]; + const q = (state.runFilter || '').trim().toLowerCase(); + const match = (r) => !q || `${r.name} ${r.target} ${r.id}`.toLowerCase().includes(q); + const runs = state.runs.filter(match); + const running = runs.filter((r) => r.state === 'running'); + const past = runs.filter((r) => r.state !== 'running'); - for (const g of groups) { + if (running.length) { const wrap = document.createElement('div'); wrap.className = 'sb-group'; - wrap.innerHTML = `
▾${g.label}${g.items.length}
`; + wrap.innerHTML = `
▾Running${running.length}
`; wrap.querySelector('.sb-group-head').addEventListener('click', () => wrap.classList.toggle('collapsed')); const items = wrap.querySelector('.sb-items'); - for (const r of g.items) { - const btn = document.createElement('button'); - btn.className = 'sb-run' + (state.currentDetailId === r.id ? ' active' : ''); - // Every line here truncates: a long target URL used to run past the - // sidebar's edge and collide with the main pane. - const worst = SEV_ORDER.find((s) => Object.entries(r.severities || {}).some(([k, n]) => n && SEV_ORDER[sevRank(k)] === s)); - btn.innerHTML = ` - ${worst ? `` : ''}${esc(r.name || r.target)} - ${r.name ? esc(r.target) : esc(r.id)} - ${r.findings} finding${r.findings === 1 ? '' : 's'}${esc(timeAgo(r.ts))}`; - btn.title = `${r.name ? r.name + '\n' : ''}${r.target}\n${r.id}${r.ts ? '\n' + new Date(r.ts * 1000).toLocaleString() : ''}`; - btn.addEventListener('click', () => openRun(r)); - items.appendChild(btn); - const isThisJob = r.state === 'running' && state.currentJob && r.id === state.currentJob.runId; - if (isThisJob) { + for (const r of running) { + items.appendChild(runButton(r)); + if (state.currentJob && r.id === state.currentJob.runId) { const steps = document.createElement('div'); steps.className = 'sb-steps'; steps.innerHTML = ['recon', 'planning', 'exploiting', 'remediation'].map((s) => @@ -1034,16 +1362,51 @@ function renderSidebar() { } root.appendChild(wrap); } + + const folders = runFolders(past); + if (!folders.length) { + const empty = document.createElement('div'); + empty.className = 'sb-empty'; + empty.textContent = q ? `No run matches “${q}”.` : 'No runs yet.'; + root.appendChild(empty); + return; + } + + const open = openFolders(); + for (const f of folders) { + const holdsActive = f.items.some((r) => r.id === state.currentDetailId); + // A search is a request to see what matched — collapsing the results would + // hide the very thing that was searched for. + const isOpen = !!q || holdsActive || open.has(f.key); + const wrap = document.createElement('div'); + wrap.className = 'sb-folder' + (isOpen ? '' : ' collapsed'); + const worst = worstSeverity(f.severities); + wrap.innerHTML = ` +
+ ▾ + ${worst ? `` : ''} + ${esc(f.key)} + ${f.items.length} +
+
`; + wrap.querySelector('.sb-folder-head').addEventListener('click', () => { + const nowCollapsed = wrap.classList.toggle('collapsed'); + setFolderOpen(f.key, !nowCollapsed); + }); + const items = wrap.querySelector('.sb-items'); + for (const r of f.items) items.appendChild(runButton(r)); + root.appendChild(wrap); + } } function openRun(run) { state.currentDetailId = run.id; if (run.state === 'running' && state.currentJob && run.id === state.currentJob.runId) { - show($('#wizardView'), false); show($('#detailView'), false); show($('#liveView'), true); + show($('#wizardView'), false); show($('#detailView'), false); show($('#dashView'), false); show($('#liveView'), true); renderSidebar(); return; } - show($('#wizardView'), false); show($('#liveView'), false); show($('#detailView'), true); + show($('#wizardView'), false); show($('#liveView'), false); show($('#dashView'), false); show($('#detailView'), true); loadDetail(run.id); renderSidebar(); } @@ -1077,7 +1440,11 @@ async function loadDetail(id) { id, ].filter(Boolean); $('#detailFacts').innerHTML = facts.map((f) => `${esc(f)}`).join(''); - renderAttackPath($('#detailAttackPath'), detail.findings, target); + // The harness writes graph.json per run (knowledge_graph::ingest). When it + // exists, its edges — including which ones it inferred — beat anything the + // browser could re-derive; older runs simply fall back. + const graph = await api(`/api/runs/${encodeURIComponent(id)}/asset/graph.json`).catch(() => null); + renderAttackPath($('#detailAttackPath'), detail.findings, target, graph); const reportLink = $('#detailOpenReport'); if (detail.assets.includes('report.html')) { reportLink.href = `/api/runs/${encodeURIComponent(id)}/asset/report.html`; @@ -1538,6 +1905,347 @@ document.addEventListener('keydown', (e) => { if (!$('#termDock').hidden) termClose(); }); + +// --------------------------------------------------------------------------- +// Dashboard — coverage, findings, and FAIR loss exposure +// +// The console could show one run at a time and nothing about the programme as a +// whole: how much has been tested, what keeps coming back, and what any of it +// is worth in money. The last question is the one a security owner is actually +// asked, and "14 highs" is not an answer to it. +// +// ## FAIR, and why the assumptions are on screen +// +// Risk here follows FAIR (Factor Analysis of Information Risk): annualized loss +// exposure = Loss Event Frequency × Loss Magnitude, where +// +// LEF = Threat Event Frequency × Vulnerability +// +// Both factors are estimates, not measurements. TEF (how often someone tries) +// is derived from the harness's own `exploitability` rating — a trivially +// exploitable bug on an internet-facing app gets attempted constantly, a hard +// one rarely. Vulnerability (how often an attempt succeeds) uses the finding's +// validation confidence, which is exactly what the multi-model vote measured. +// Loss magnitude cannot be derived from a scan at all: it depends on the +// business. So the defaults below are stated openly, shown in the UI, and +// editable — and the result is a RANGE, never a single number, because a point +// estimate of a distribution is the classic way risk quantification lies. +// --------------------------------------------------------------------------- + +const FAIR_DEFAULTS = { + // Threat event frequency: attempts per year, by how easy the harness judged + // the finding to exploit. + tef: { trivial: 12, moderate: 4, hard: 1, unknown: 3 }, + // Loss magnitude in USD per event: [minimum, most likely, maximum]. + // Order-of-magnitude industry defaults — replace with your own loss data. + lm: { + critical: [250000, 1200000, 5000000], + high: [75000, 400000, 1500000], + medium: [15000, 80000, 300000], + low: [2000, 15000, 60000], + info: [0, 1000, 5000], + }, +}; +const FAIR_KEY = 'ns-fair-params'; + +function fairParams() { + try { + const saved = JSON.parse(localStorage.getItem(FAIR_KEY) || 'null'); + if (saved?.tef && saved?.lm) return saved; + } catch { /* fall through to defaults */ } + return structuredClone(FAIR_DEFAULTS); +} +function saveFairParams(p) { localStorage.setItem(FAIR_KEY, JSON.stringify(p)); } + +const money = (n) => { + if (!isFinite(n)) return '—'; + if (n >= 1e9) return `$${(n / 1e9).toFixed(1)}B`; + if (n >= 1e6) return `$${(n / 1e6).toFixed(1)}M`; + if (n >= 1e3) return `$${Math.round(n / 1e3)}K`; + return `$${Math.round(n)}`; +}; + +/// Annualized loss exposure for one finding, as [min, likely, max]. +function fairForFinding(f, params) { + const sev = SEV_ORDER[sevRank(f.severity)]; + const lm = params.lm[sev] || params.lm.info; + const tef = params.tef[(f.exploitability || 'unknown').toLowerCase()] ?? params.tef.unknown; + // A finding flagged for human review is a maybe, not a fact — halving its + // frequency keeps it visible without letting unreviewed leads drive the total. + const reviewFactor = f.reviewStatus === 'needs-review' ? 0.5 : 1; + const vuln = Math.min(0.95, Math.max(0.2, f.confidence || 0.5)); + const lef = tef * vuln * reviewFactor; + return lm.map((m) => lef * m); +} + +/// Heuristic posture score, 0-100. Deliberately simple and fully stated in the +/// UI: an opaque score invites arguing with the number instead of the findings. +/// +/// Subtracting a fixed penalty per finding hit zero after one critical and a +/// handful of highs, which makes the score useless exactly when there is +/// something to track — a programme that fixes half its criticals must be able +/// to see the number move. The saturating form has diminishing returns instead, +/// so it keeps discriminating at any volume and never quite reaches 0. +const SCORE_WEIGHT = { critical: 10, high: 5, medium: 2, low: 0.5, info: 0.1 }; +const SCORE_SCALE = 25; +function riskLoad(findings) { + return findings.reduce((acc, f) => { + const sev = SEV_ORDER[sevRank(f.severity)]; + return acc + (SCORE_WEIGHT[sev] || 0) * Math.min(1, Math.max(0.3, f.confidence || 0.5)); + }, 0); +} +function exposureScore(findings) { + return Math.round(100 / (1 + riskLoad(findings) / SCORE_SCALE)); +} +function scoreBand(score) { + if (score >= 85) return { label: 'low exposure', cls: 'low' }; + if (score >= 65) return { label: 'moderate exposure', cls: 'medium' }; + if (score >= 40) return { label: 'high exposure', cls: 'high' }; + return { label: 'critical exposure', cls: 'critical' }; +} + +/// One horizontal bar row: a label, a proportional fill, a direct value. +/// Values are labeled on every row, so the bar is a second encoding of a number +/// that is already readable — not the only way to read it. +function barRow(label, value, max, cls, title) { + const pct = max > 0 ? Math.max(2, (value / max) * 100) : 0; + return `
+ ${esc(label)} + + ${esc(String(value))} +
`; +} + +function dashRangeCutoff() { + const days = Number($('#dashRange')?.value || 0); + return days > 0 ? Date.now() / 1000 - days * 86400 : 0; +} + +async function renderDashboard() { + const body = $('#dashBody'); + let data; + try { + data = await api('/api/stats'); + } catch (e) { + body.innerHTML = `
Couldn't load stats: ${esc(e.message)}
`; + return; + } + const cutoff = dashRangeCutoff(); + const runs = data.runs.filter((r) => !cutoff || r.ts >= cutoff); + const findings = data.findings.filter((f) => !cutoff || f.ts >= cutoff); + const params = fairParams(); + + if (!runs.length) { + body.innerHTML = '
No runs in this range yet — start an engagement and the dashboard fills in.
'; + return; + } + + const targets = new Set(runs.map((r) => engagementKey(r.target))); + const sevCounts = {}; + for (const f of findings) { + const s = SEV_ORDER[sevRank(f.severity)]; + sevCounts[s] = (sevCounts[s] || 0) + 1; + } + const needsReview = findings.filter((f) => f.reviewStatus === 'needs-review').length; + const score = exposureScore(findings); + const band = scoreBand(score); + + const ale = findings.reduce((acc, f) => { + const [lo, ml, hi] = fairForFinding(f, params); + return [acc[0] + lo, acc[1] + ml, acc[2] + hi]; + }, [0, 0, 0]); + + // Which findings actually drive the exposure — the reason to quantify at all + // is to rank remediation, and the ranking is what gets acted on. + const contributors = findings + .map((f) => ({ f, ml: fairForFinding(f, params)[1] })) + .sort((a, b) => b.ml - a.ml) + .slice(0, 6); + + const byCwe = {}; + for (const f of findings) { + if (!f.cwe) continue; + byCwe[f.cwe] = (byCwe[f.cwe] || 0) + 1; + } + const topCwe = Object.entries(byCwe).sort((a, b) => b[1] - a[1]).slice(0, 6); + + const byTarget = {}; + for (const r of runs) { + const k = engagementKey(r.target); + byTarget[k] = byTarget[k] || { runs: 0, findings: 0, last: 0 }; + byTarget[k].runs++; + byTarget[k].findings += r.findings; + byTarget[k].last = Math.max(byTarget[k].last, r.ts); + } + const targetRows = Object.entries(byTarget).sort((a, b) => b[1].findings - a[1].findings); + + const maxSev = Math.max(1, ...Object.values(sevCounts)); + const maxCwe = Math.max(1, ...topCwe.map(([, n]) => n)); + + body.innerHTML = ` +
+
+
Engagements
+
${runs.length}
+
${targets.size} distinct target(s)
+
+
+
Findings
+
${findings.length}
+
${needsReview} awaiting human review
+
+
+
Agents run
+
${runs.reduce((a, r) => a + (r.agentsRan || 0), 0).toLocaleString('en-US')}
+
across ${runs.length} run(s)
+
+
+
Exposure score
+
${score}/100
+
${esc(band.label)}
+
+
+ +
+
+
+

Annualized loss exposure (FAIR)

+ +
+
+
+
minimum${money(ale[0])}
+
most likely${money(ale[1])}
+
maximum${money(ale[2])}
+
+
+ Loss Event Frequency × Loss Magnitude, summed over ${findings.length} finding(s). + Frequency comes from each finding's exploitability and validation confidence; + magnitude comes from the assumptions you set. A range, not a forecast. + Findings are summed independently, so shared root causes count twice — + read it as an upper bound on annual exposure, not a portfolio model. +
+
+
+
Top contributors (most likely annual loss)
+ ${contributors.map(({ f, ml }) => ` +
+ ${esc(f.severity)} + ${esc(f.title || f.cwe || 'untitled')} + ${money(ml)} +
`).join('') || '
No findings in range.
'} +
+
+ +
+

Findings by severity

+ ${SEV_ORDER.filter((s) => sevCounts[s]).map((s) => + barRow(s, sevCounts[s], maxSev, s, `${sevCounts[s]} ${s} finding(s)`)).join('') + || '
No findings in range.
'} +
+ +
+

Most frequent weaknesses

+ ${topCwe.map(([cwe, n]) => barRow(cwe, n, maxCwe, 'neutral', `${cwe}: ${n} finding(s)`)).join('') + || '
No CWE data in range.
'} +
+ +
+

Targets

+
+ + + + ${targetRows.map(([k, v]) => ` + + + + + `).join('')} + +
TargetRunsFindingsLast tested
${esc(k)}${v.runs}${v.findings}${esc(timeAgo(v.last))}
+
+
+
+ +
Score = 100 / (1 + risk / ${SCORE_SCALE}) where risk = Σ(severity weight × confidence) = ${riskLoad(findings).toFixed(1)}. Weights: critical ${SCORE_WEIGHT.critical}, high ${SCORE_WEIGHT.high}, medium ${SCORE_WEIGHT.medium}, low ${SCORE_WEIGHT.low}, info ${SCORE_WEIGHT.info}. A heuristic for tracking direction over time, not a certification.
+ `; + + $('#btnFairParams').addEventListener('click', () => openFairParams(params)); + $$('#dashBody tbody tr[data-target]').forEach((tr) => tr.addEventListener('click', () => { + state.runFilter = tr.dataset.target; + $('#runFilter').value = tr.dataset.target; + renderSidebar(); + })); +} + +/// The assumptions panel. FAIR without visible inputs is a magic number; with +/// them it is a model the reader can disagree with concretely. +function openFairParams(params) { + const p = structuredClone(params); + const body = $('#dashBody'); + const host = document.createElement('div'); + host.className = 'modal-overlay'; + host.innerHTML = ` + `; + document.body.appendChild(host); + const close = () => host.remove(); + host.addEventListener('click', (e) => { if (e.target === host) close(); }); + host.querySelector('[data-close]').addEventListener('click', close); + host.querySelector('[data-reset]').addEventListener('click', () => { + saveFairParams(structuredClone(FAIR_DEFAULTS)); + close(); + renderDashboard(); + }); + host.querySelector('[data-save]').addEventListener('click', () => { + host.querySelectorAll('[data-tef]').forEach((i) => { p.tef[i.dataset.tef] = Number(i.value) || 0; }); + host.querySelectorAll('[data-lm]').forEach((i) => { p.lm[i.dataset.lm][Number(i.dataset.i)] = Number(i.value) || 0; }); + // A max below the minimum would silently invert the range. + for (const k of Object.keys(p.lm)) p.lm[k].sort((a, b) => a - b); + saveFairParams(p); + close(); + renderDashboard(); + }); + body.scrollTop = body.scrollTop; // keep the page anchored while the modal opens +} + +function showDashboard() { + leaveLiveJob(); + clearInterval(state.detailPoll); + show($('#wizardView'), false); + show($('#liveView'), false); + show($('#detailView'), false); + show($('#dashView'), true); + renderDashboard(); +} +$('#btnDashboard').addEventListener('click', showDashboard); +$('#btnDashRefresh').addEventListener('click', renderDashboard); +$('#dashRange').addEventListener('change', renderDashboard); +$('#runFilter').addEventListener('input', (e) => { state.runFilter = e.target.value; renderSidebar(); }); + // --------------------------------------------------------------------------- // boot // --------------------------------------------------------------------------- diff --git a/web/public/index.html b/web/public/index.html index fcf45ae..4f839c3 100644 --- a/web/public/index.html +++ b/web/public/index.html @@ -22,6 +22,12 @@
+ + +
@@ -37,6 +43,27 @@
+ + +
diff --git a/web/public/style.css b/web/public/style.css index e5e4fa2..a3c4ce2 100644 --- a/web/public/style.css +++ b/web/public/style.css @@ -156,6 +156,97 @@ a { color: var(--accent); text-decoration: none; } .sb-version { font-size: 11px; color: var(--text-faint); font-family: var(--mono); } .sb-bottom-actions { display: flex; gap: var(--sp-1); } +.sb-link { + margin: 0 var(--sp-4) var(--sp-3); padding: var(--sp-2) var(--sp-3); + border: 1px solid transparent; border-radius: var(--radius-sm); + background: transparent; color: var(--text-dim); font-size: 12.5px; text-align: left; +} +.sb-link:hover { background: var(--surface-3); color: var(--text); } + +.sb-search { position: relative; margin: 0 var(--sp-4) var(--sp-3); } +.sb-search input { padding-left: 26px; font-size: 12px; } +.sb-search .search-icon { left: 8px; } + +/* Runs are grouped into one folder per target: twelve rows of near-identical + URLs was a wall to scroll past, and the host is what an operator scans for. */ +.sb-folder { margin-bottom: 2px; } +.sb-folder-head { + display: flex; align-items: center; gap: var(--sp-2); padding: var(--sp-2); + border-radius: var(--radius-sm); cursor: pointer; user-select: none; font-size: 12px; +} +.sb-folder-head:hover { background: var(--surface-3); } +.sb-folder-head .caret { font-size: 11px; line-height: 1; color: var(--text-faint); transition: transform .15s; } +.sb-folder-head .fname { flex: 1; min-width: 0; font-weight: 600; white-space: nowrap; overflow: hidden; text-overflow: ellipsis; } +.sb-folder-head .fmeta { font-family: var(--mono); font-size: 11px; color: var(--text-faint); } +.sb-folder.collapsed .caret { transform: rotate(-90deg); } +.sb-folder.collapsed .sb-items { display: none; } +.sb-folder .sb-items { padding-left: var(--sp-3); border-left: 1px solid var(--border); margin-left: 9px; } +.sb-empty { padding: var(--sp-4) var(--sp-2); font-size: 12px; color: var(--text-faint); } + +/* ============================================================ Dashboard */ + +.dashboard { flex: 1; display: flex; flex-direction: column; overflow: hidden; } +.dashboard[hidden] { display: none; } +.dash-range { width: auto; padding: 6px 8px; font-size: 12.5px; } +.dash-body { flex: 1; overflow-y: auto; padding: var(--sp-5); } + +.stat-row { display: grid; grid-template-columns: repeat(auto-fit, minmax(190px, 1fr)); gap: var(--sp-3); margin-bottom: var(--sp-4); } +.stat-tile { border: 1px solid var(--border); border-radius: var(--radius-md); padding: var(--sp-4); background: var(--surface); } +.stat-k { font-size: 10.5px; text-transform: uppercase; letter-spacing: .05em; color: var(--text-faint); font-weight: 600; } +.stat-v { font-size: 30px; font-weight: 650; line-height: 1.15; margin-top: 2px; font-variant-numeric: tabular-nums; } +.stat-unit { font-size: 14px; color: var(--text-faint); font-weight: 500; } +.stat-sub { font-size: 11.5px; color: var(--text-dim); margin-top: 2px; } +/* The score tile carries a status color, so it also carries a word — the band + label — because color alone is not an encoding. */ +.stat-score.sev-critical { border-color: var(--sev-critical-fg); } +.stat-score.sev-critical .stat-v { color: var(--sev-critical-fg); } +.stat-score.sev-high { border-color: var(--sev-high-fg); } +.stat-score.sev-high .stat-v { color: var(--sev-high-fg); } +.stat-score.sev-medium { border-color: var(--sev-medium-fg); } +.stat-score.sev-medium .stat-v { color: var(--sev-medium-fg); } +.stat-score.sev-low { border-color: var(--sev-low-fg); } +.stat-score.sev-low .stat-v { color: var(--sev-low-fg); } + +.dash-grid { display: grid; grid-template-columns: repeat(auto-fit, minmax(320px, 1fr)); gap: var(--sp-3); } +.dash-card { border: 1px solid var(--border); border-radius: var(--radius-md); padding: var(--sp-4); background: var(--surface); } +.dash-card h3 { margin: 0 0 var(--sp-3); font-size: 12.5px; font-weight: 600; } +.dash-card-head { display: flex; align-items: center; justify-content: space-between; gap: var(--sp-2); margin-bottom: var(--sp-3); } +.dash-card-head h3 { margin: 0; } +.dash-wide { grid-column: 1 / -1; } +.dash-sub { font-size: 10.5px; text-transform: uppercase; letter-spacing: .05em; color: var(--text-faint); font-weight: 600; margin: var(--sp-4) 0 var(--sp-2); } +.dash-foot { margin-top: var(--sp-4); font-size: 11px; color: var(--text-faint); } +.dash-table td.mono { font-family: var(--mono); font-size: 11.5px; } +.dash-table { min-width: 480px; } + +/* Bars: thin marks, value labeled on every row, recessive track. */ +.bar-row { display: grid; grid-template-columns: 90px 1fr 42px; align-items: center; gap: var(--sp-3); padding: 3px 0; font-size: 12px; } +.bar-label { color: var(--text-dim); text-transform: capitalize; white-space: nowrap; overflow: hidden; text-overflow: ellipsis; } +.bar-track { height: 10px; background: var(--surface-3); border-radius: 3px; overflow: hidden; } +.bar-fill { display: block; height: 100%; border-radius: 3px; background: var(--text-faint); } +.bar-critical { background: var(--sev-critical-fg); } +.bar-high { background: var(--sev-high-fg); } +.bar-medium { background: var(--sev-medium-fg); } +.bar-low { background: var(--sev-low-fg); } +.bar-info { background: var(--sev-info-fg); } +.bar-neutral { background: var(--accent); } +.bar-value { font-family: var(--mono); font-size: 11.5px; text-align: right; color: var(--text); } + +.fair-hero { border: 1px solid var(--border); border-radius: var(--radius-sm); padding: var(--sp-4); background: var(--surface-2); } +.fair-range { display: flex; gap: var(--sp-5); flex-wrap: wrap; align-items: baseline; } +.fair-point { display: flex; flex-direction: column; } +.fair-point .k { font-size: 10.5px; text-transform: uppercase; letter-spacing: .05em; color: var(--text-faint); } +.fair-point .v { font-size: 18px; font-weight: 600; font-variant-numeric: tabular-nums; } +.fair-likely .v { font-size: 30px; color: var(--accent); } +.fair-note { font-size: 11.5px; color: var(--text-dim); margin-top: var(--sp-3); } +.contrib-row { display: flex; align-items: center; gap: var(--sp-3); padding: 5px 0; border-bottom: 1px solid var(--border); font-size: 12px; } +.contrib-row:last-child { border-bottom: none; } +.contrib-title { flex: 1; min-width: 0; white-space: nowrap; overflow: hidden; text-overflow: ellipsis; } +.contrib-v { font-family: var(--mono); font-size: 12px; } +.fair-params { display: flex; flex-direction: column; gap: var(--sp-2); } +.fair-param { display: flex; align-items: center; gap: var(--sp-2); font-size: 12px; } +.fair-param > span:first-child { width: 92px; flex: none; text-transform: capitalize; } +.fair-param input { flex: 1; font-family: var(--mono); font-size: 12px; } + /* ============================================================ Main / Topbar */ .main { flex: 1; display: flex; flex-direction: column; min-width: 0; } @@ -341,10 +432,51 @@ textarea { resize: vertical; min-height: 72px; } /* Generative Attack Path Chaining */ .attackpath-empty { font-size: 12.5px; color: var(--text-faint); padding: var(--sp-5); text-align: center; border: 1px dashed var(--border-strong); border-radius: var(--radius-sm); } -.ap-canvas-wrap { border-radius: var(--radius-md); overflow: auto; background: var(--surface-2); border: 1px solid var(--border); } -.ap-canvas { display: block; min-width: 100%; } + +.ap-toolbar { display: flex; align-items: center; gap: var(--sp-3); flex-wrap: wrap; margin-bottom: var(--sp-2); } +.ap-stats { font-size: 12px; color: var(--text-dim); } +.ap-stats b { color: var(--text); font-family: var(--mono); } +.ap-inferred-note { color: var(--text-faint); } +.ap-check { display: flex; align-items: center; gap: 6px; font-size: 12px; color: var(--text-dim); } +.ap-sev { width: auto; padding: 5px 8px; font-size: 12px; } +.ap-zoom { display: flex; gap: 4px; } +.ap-zoom .btn { min-width: 32px; justify-content: center; } +.ap-provenance { font-size: 11.5px; color: var(--text-faint); margin-bottom: var(--sp-2); } + +/* The canvas is a fixed viewport that the graph pans inside — letting the box + grow to the graph's height (1187px on a 27-finding run) meant scrolling the + page blind, with no way to see the shape of the path. */ +.ap-canvas-wrap { position: relative; height: min(60vh, 560px); border-radius: var(--radius-md); overflow: hidden; background: var(--surface-2); border: 1px solid var(--border); cursor: grab; touch-action: none; } +.ap-canvas-wrap.dragging { cursor: grabbing; } +.ap-canvas { display: block; width: 100%; height: 100%; } .ap-canvas text { font-family: var(--sans); } -.ap-node-g:hover rect:first-child { filter: brightness(0.97); } +.ap-hint { position: absolute; right: 8px; bottom: 6px; font-size: 10.5px; color: var(--text-faint); pointer-events: none; } + +.ap-col-line { stroke: var(--border); stroke-width: 1; } +.ap-col-name { font-size: 10.5px; fill: var(--text-faint); text-transform: uppercase; letter-spacing: .06em; font-weight: 600; } +.ap-col-count { font-size: 10.5px; fill: var(--text-faint); font-family: var(--mono); } + +.ap-node { fill: var(--surface); stroke-width: 1.6; } +.ap-root .ap-node { stroke: var(--accent); } +.ap-icon { font-size: 13px; } +.ap-title { font-size: 11.5px; fill: var(--text); font-weight: 600; } +.ap-meta, .ap-conf { font-size: 10px; fill: var(--text-faint); font-family: var(--mono); } +.ap-node-g { cursor: pointer; } +.ap-node-g:hover .ap-node, .ap-node-g:focus-visible .ap-node { filter: brightness(1.04); stroke-width: 2.4; } +.ap-node-g:focus { outline: none; } + +.ap-edge { stroke: var(--border-strong); stroke-width: 1.6; fill: none; } +.ap-edge marker, #apArrow path { fill: var(--border-strong); } +/* An inferred edge is a hypothesis about the path, not evidence of it. */ +.ap-edge.inferred { stroke-dasharray: 5 4; opacity: .55; } +.ap-edge.root-edge { opacity: .35; } +.has-focus .ap-node-g:not(.focus) { opacity: .22; } +.has-focus .ap-edge:not(.focus) { opacity: .1; } +.ap-edge.focus { stroke: var(--accent); stroke-width: 2.2; opacity: 1; } + +.ap-legend { display: flex; gap: var(--sp-4); flex-wrap: wrap; align-items: center; margin-top: var(--sp-2); font-size: 11px; color: var(--text-faint); } +.ap-key { display: flex; align-items: center; gap: 5px; } +.ap-key i { width: 9px; height: 9px; border-radius: 2px; display: inline-block; } /* findings table */ /* Wide tables scroll inside their own box; without this the whole page diff --git a/web/server.js b/web/server.js index 7639f1b..5f93d80 100644 --- a/web/server.js +++ b/web/server.js @@ -391,6 +391,61 @@ async function listRuns() { return runs; } +/// Flat aggregate over every run for the dashboard. +/// +/// Returns per-finding tuples rather than a computed risk number: the FAIR +/// estimate depends on assumptions (contact frequency, loss magnitude per +/// severity) that belong to the operator, not to this server, so the browser +/// computes it from parameters the operator can see and change. +async function stats() { + let ids = []; + try { + ids = (await fsp.readdir(RUNS_DIR)).filter((d) => d.startsWith('ns-')); + } catch { + return { runs: [], findings: [], generated: Date.now() }; + } + const runs = []; + const findings = []; + await Promise.all(ids.map(async (id) => { + const dir = path.join(RUNS_DIR, id); + const [status, fs_] = await Promise.all([ + readJsonSafe(path.join(dir, 'status.json'), {}), + readJsonSafe(path.join(dir, 'findings.json'), []), + ]); + const meta = await readJsonSafe(path.join(dir, 'meta.json'), {}); + const tsMatch = id.match(/^ns-(\d+)-/); + const ts = status.ts || (tsMatch ? Number(tsMatch[1]) : 0); + const target = status.target || meta.target || id.replace(/^ns-\d+-/, ''); + runs.push({ + id, + ts, + name: engagementNames.get(id) || '', + target, + state: status.state || 'unknown', + agentsRan: status.agents_ran || 0, + findings: fs_.length, + }); + for (const f of fs_) { + findings.push({ + runId: id, + target, + ts, + severity: f.severity || 'Info', + cwe: f.cwe || '', + owasp: f.owasp || '', + stage: f.stage || '', + agent: f.agent || '', + title: f.title || '', + exploitability: f.exploitability || '', + confidence: typeof f.confidence === 'number' ? f.confidence : 0, + reviewStatus: f.review_status || '', + }); + } + })); + runs.sort((a, b) => b.ts - a.ts); + return { runs, findings, generated: Date.now() }; +} + async function runDetail(id) { const dir = safeRunDir(id); if (!dir) return null; @@ -962,6 +1017,11 @@ const server = http.createServer(async (req, res) => { return; } + // ---- aggregate stats for the dashboard ---- + if (req.method === 'GET' && p === '/api/stats') { + return sendJson(res, 200, await stats()); + } + if (req.method === 'GET' && p === '/api/meta') { return sendJson(res, 200, { version: '4.0.0', binary: BIN, root: ROOT }); }