feat: attack knowledge graph, layered memory, command rectification, FAIR dashboard

Backend
-------
- knowledge_graph.rs — the durable structure under attack_graph's per-run view:
  typed entities (asset/endpoint/weakness/technique/finding/account/credential/
  impact) joined by typed, weighted, provenance-carrying edges, accumulated
  across runs in .neurosploit/graph.json plus a per-run copy the report and web
  console can draw. Answers what a finding list can't: ranked attack paths, and
  the frontier of entities observed but never proven — where chaining should
  look next. Agents only sometimes fill chains_from, so progression is also
  inferred between adjacent kill-chain stages; those edges are marked inferred,
  weighted lower, and drawn dashed, because presenting a hypothesis as evidence
  is the graph lying about itself. Secrets stay in the vault, never the graph.

- memory.rs — four tiers scoped by lifetime, not importance: working (one run),
  engagement (one target), technique (one agent/CWE), reusable (generalized).
  Promotion is evidence-gated and needs independent evidence at each step: a
  claim repeated within a run becomes engagement knowledge; one confirmed
  across runs becomes technique knowledge; one that held on two DIFFERENT
  targets is generalized into a reusable lesson with host-specific tokens
  stripped. Nothing is promoted on a single observation, which is exactly what
  a hallucination looks like. Recall is scored (overlap × past success ×
  recency) and injected into recon/exploit prompts as leads to verify. Recalled
  memos are credited only when the run they informed actually found something.

- rectify.rs — a mistyped command cost a full round trip through /help, at the
  worst possible moment during a live run. Accepted-as-typed wins over
  everything (so the /url alias is never "corrected" to /ua), then unique
  prefix, then Damerau-Levenshtein with a length-scaled budget, and a tie is
  reported rather than resolved. Arguments too: a bare host gets its scheme, an
  out-of-range count is clamped with a note instead of silently reverting, a
  near-miss model id is matched against the live catalog.

- pool.rs — when every configured model is exhausted or its token is dead, try
  whatever else this machine can actually reach (an installed CLI subscription,
  or a provider whose key is in the environment) before parking. A run that
  stops on a box with three other usable backends stopped for no reason.

- repl.rs — /memory, /forget, /graph; a recovered run resumes by itself where
  nobody is watching (piped stdin — the web console — or NEUROSPLOIT_AUTO_RESUME),
  since a `/continue` prompt there waits forever.

Web
---
- Attack path: the stage list was seven hardcoded values, so findings the
  harness staged outside it were silently dropped — 5 of 27 on a real run.
  Rewritten against the harness's own stage list with unknown stages kept,
  two-line labels (every node used to read "SQL Injection Authent…"), stage
  column headers, pan/zoom/fit, path highlighting, severity filter, and the
  run's graph.json used when present.
- Dashboard: coverage, findings by severity, top weaknesses, and annualized
  loss exposure via FAIR — frequency from exploitability × validation
  confidence, magnitude from assumptions shown on screen and editable, reported
  as a range. The posture score saturates instead of subtracting, so it keeps
  discriminating past the first critical.
- Run history groups into one folder per target with a filter, instead of one
  flat list that grows forever.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BvdGy9XtVWSdXDTa3FFLJv
This commit is contained in:
CyberSecurityUPandClaude Opus 5 committed 2026-09-07 15:35:36 -03:00
1 parent 0ef0ce8d94
commit 9d83cb6e30
13 files changed
+3163 -119

No files matched your search

@@ -0,0 +1,602 @@
//! The attack knowledge graph: what was learned about a target, as a graph.
//!
//! [`crate::attack_graph`] maps a finding to OWASP/MITRE/stage and draws it.
//! That is a *per-run view*. This module is the durable structure underneath:
//! typed entities (asset, endpoint, weakness, technique, finding, account,
//! credential, impact) joined by typed, weighted, provenance-carrying edges, so
//! the harness can answer questions a flat finding list cannot —
//!
//! - which endpoint accumulated the most distinct weaknesses across runs;
//! - which credential a finding actually yielded, and what that credential then
//! unlocked;
//! - what paths run from the asset to an impact node, ranked by how likely and
//! how damaging they are;
//! - what the frontier is: entities we have observed but never proved anything
//! about — the natural next targets for chaining.
//!
//! ## Inferred edges are marked as inferred
//!
//! Agents only sometimes populate `chains_from`. Without it a "chain" view
//! degenerates into a fan of unconnected findings, so this module also *infers*
//! progression edges between kill-chain stages. Those carry `inferred: true` and
//! a lower probability, and every renderer draws them differently, because an
//! inferred edge is a hypothesis about an attack path — presenting it as a
//! proven one would be the graph lying about its own evidence.
use crate::types::Finding;
use serde::{Deserialize, Serialize};
use std::collections::{BTreeMap, BTreeSet};
use std::path::Path;
#[derive(Clone, Copy, Debug, PartialEq, Eq, PartialOrd, Ord, Serialize, Deserialize)]
#[serde(rename_all = "kebab-case")]
pub enum NodeKind {
Asset,
Endpoint,
Tech,
Weakness,
Technique,
Finding,
Account,
Credential,
Impact,
}
impl NodeKind {
pub fn as_str(&self) -> &'static str {
match self {
NodeKind::Asset => "asset",
NodeKind::Endpoint => "endpoint",
NodeKind::Tech => "tech",
NodeKind::Weakness => "weakness",
NodeKind::Technique => "technique",
NodeKind::Finding => "finding",
NodeKind::Account => "account",
NodeKind::Credential => "credential",
NodeKind::Impact => "impact",
}
}
}
#[derive(Clone, Copy, Debug, PartialEq, Eq, PartialOrd, Ord, Serialize, Deserialize)]
#[serde(rename_all = "kebab-case")]
pub enum EdgeKind {
/// asset → endpoint
Exposes,
/// asset → tech
Runs,
/// endpoint → weakness
Vulnerable,
/// finding → weakness (this finding proves that weakness)
Proves,
/// finding → endpoint (where it was proven)
ObservedOn,
/// finding → technique (MITRE)
Uses,
/// finding → finding (attack path)
Chains,
/// finding → account/credential
Grants,
/// finding → impact
Leads,
}
impl EdgeKind {
pub fn as_str(&self) -> &'static str {
match self {
EdgeKind::Exposes => "exposes",
EdgeKind::Runs => "runs",
EdgeKind::Vulnerable => "vulnerable",
EdgeKind::Proves => "proves",
EdgeKind::ObservedOn => "observed-on",
EdgeKind::Uses => "uses",
EdgeKind::Chains => "chains",
EdgeKind::Grants => "grants",
EdgeKind::Leads => "leads",
}
}
}
#[derive(Clone, Debug, Serialize, Deserialize)]
pub struct Node {
pub id: String,
pub kind: NodeKind,
pub label: String,
#[serde(default)]
pub meta: BTreeMap<String, String>,
/// Run ids that touched this node — provenance, and the "how often" signal.
#[serde(default)]
pub runs: BTreeSet<String>,
#[serde(default)]
pub first_seen: u64,
#[serde(default)]
pub last_seen: u64,
}
#[derive(Clone, Debug, Serialize, Deserialize)]
pub struct Edge {
pub from: String,
pub to: String,
pub kind: EdgeKind,
/// Confidence that the relation holds, 0..1.
#[serde(default)]
pub p: f64,
/// True when the harness derived this edge rather than an agent asserting it.
#[serde(default)]
pub inferred: bool,
#[serde(default)]
pub runs: BTreeSet<String>,
}
#[derive(Default, Clone, Serialize, Deserialize)]
pub struct KnowledgeGraph {
pub nodes: BTreeMap<String, Node>,
pub edges: Vec<Edge>,
}
fn now() -> u64 {
std::time::SystemTime::now()
.duration_since(std::time::UNIX_EPOCH)
.map(|d| d.as_secs())
.unwrap_or(0)
}
fn sev_weight(sev: &str) -> f64 {
match sev.to_lowercase().as_str() {
s if s.starts_with("crit") => 1.0,
s if s.starts_with("high") => 0.75,
s if s.starts_with("med") => 0.5,
s if s.starts_with("low") => 0.25,
_ => 0.1,
}
}
/// Kill-chain progression order. An attack moves down this list; an edge that
/// would move *up* it is not progression and is never inferred.
pub const STAGES: &[&str] = &[
"recon",
"discovery",
"initial-access",
"execution",
"persistence",
"privesc",
"credential-access",
"lateral",
"collection",
"exfil",
"impact",
];
pub fn stage_rank(s: &str) -> usize {
STAGES.iter().position(|x| *x == s).unwrap_or(STAGES.len())
}
/// Strip the query string and normalize the host so two spellings of one URL
/// collapse into a single endpoint node.
fn endpoint_key(url: &str) -> String {
let u = url.trim();
let no_scheme = u.split_once("://").map(|(_, r)| r).unwrap_or(u);
let path = no_scheme.split(['?', '#']).next().unwrap_or(no_scheme);
let path = path.trim_end_matches('/');
let path = path.trim_start_matches("www.");
if path.is_empty() {
no_scheme.to_string()
} else {
path.to_lowercase()
}
}
impl KnowledgeGraph {
pub fn new() -> Self {
Self::default()
}
pub fn load(path: impl AsRef<Path>) -> Self {
std::fs::read_to_string(path)
.ok()
.and_then(|s| serde_json::from_str(&s).ok())
.unwrap_or_default()
}
pub fn save(&self, path: impl AsRef<Path>) {
let path = path.as_ref();
if let Some(parent) = path.parent() {
let _ = std::fs::create_dir_all(parent);
}
if let Ok(j) = serde_json::to_string_pretty(self) {
let _ = std::fs::write(path, j);
}
}
pub fn to_json(&self) -> String {
serde_json::to_string_pretty(self).unwrap_or_else(|_| "{}".into())
}
fn upsert(&mut self, id: &str, kind: NodeKind, label: &str, run: &str) -> String {
let ts = now();
let n = self.nodes.entry(id.to_string()).or_insert_with(|| Node {
id: id.to_string(),
kind,
label: label.to_string(),
meta: BTreeMap::new(),
runs: BTreeSet::new(),
first_seen: ts,
last_seen: ts,
});
n.last_seen = ts;
if !run.is_empty() {
n.runs.insert(run.to_string());
}
if n.label.is_empty() {
n.label = label.to_string();
}
id.to_string()
}
fn meta(&mut self, id: &str, k: &str, v: &str) {
if v.is_empty() {
return;
}
if let Some(n) = self.nodes.get_mut(id) {
n.meta.insert(k.to_string(), v.to_string());
}
}
/// Add or reinforce an edge. Re-observing an edge raises its probability
/// toward certainty rather than appending a duplicate, and an edge first
/// inferred but later asserted by an agent stops being marked inferred.
pub fn link(&mut self, from: &str, to: &str, kind: EdgeKind, p: f64, inferred: bool, run: &str) {
if from == to || !self.nodes.contains_key(from) || !self.nodes.contains_key(to) {
return;
}
if let Some(e) = self.edges.iter_mut().find(|e| e.from == from && e.to == to && e.kind == kind) {
e.p = (e.p + 0.3 * (p.max(e.p) - e.p)).clamp(0.0, 0.99);
e.inferred = e.inferred && inferred;
if !run.is_empty() {
e.runs.insert(run.to_string());
}
return;
}
let mut runs = BTreeSet::new();
if !run.is_empty() {
runs.insert(run.to_string());
}
self.edges.push(Edge { from: from.into(), to: to.into(), kind, p: p.clamp(0.0, 0.99), inferred, runs });
}
/// Fold one run's findings into the graph.
pub fn ingest(&mut self, target: &str, run: &str, findings: &[Finding]) {
let akey = crate::memory::engagement_key(target);
let asset = self.upsert(&format!("asset:{akey}"), NodeKind::Asset, if akey.is_empty() { target } else { &akey }, run);
for f in findings {
let fid = self.upsert(
&format!("find:{}:{}", run, f.id),
NodeKind::Finding,
if f.title.is_empty() { &f.id } else { &f.title },
run,
);
self.meta(&fid, "severity", &f.severity);
self.meta(&fid, "cwe", &f.cwe);
self.meta(&fid, "stage", &f.stage);
self.meta(&fid, "owasp", &f.owasp);
self.meta(&fid, "mitre", &f.mitre);
self.meta(&fid, "exploitability", &f.exploitability);
self.meta(&fid, "agent", &f.agent);
self.meta(&fid, "endpoint", &f.endpoint);
self.meta(&fid, "confidence", &format!("{:.2}", f.confidence));
self.meta(&fid, "review_status", &f.review_status);
self.meta(&fid, "finding_id", &f.id);
if !f.endpoint.is_empty() {
let ek = endpoint_key(&f.endpoint);
let ep = self.upsert(&format!("ep:{ek}"), NodeKind::Endpoint, &ek, run);
self.link(&asset, &ep, EdgeKind::Exposes, 0.9, false, run);
self.link(&fid, &ep, EdgeKind::ObservedOn, 0.95, false, run);
if !f.cwe.is_empty() {
let w = self.upsert(&format!("cwe:{}", f.cwe), NodeKind::Weakness, &f.cwe, run);
self.link(&ep, &w, EdgeKind::Vulnerable, f.confidence.max(0.5), false, run);
}
}
if !f.cwe.is_empty() {
let w = self.upsert(&format!("cwe:{}", f.cwe), NodeKind::Weakness, &f.cwe, run);
self.link(&fid, &w, EdgeKind::Proves, f.confidence.max(0.5), false, run);
}
if !f.mitre.is_empty() {
let t = self.upsert(&format!("att:{}", f.mitre), NodeKind::Technique, &f.mitre, run);
self.link(&fid, &t, EdgeKind::Uses, 0.9, false, run);
}
if !f.account.is_empty() {
let a = self.upsert(&format!("acct:{}", f.account), NodeKind::Account, &f.account, run);
self.link(&fid, &a, EdgeKind::Grants, 0.9, false, run);
// The secret itself never enters the graph — the graph is an
// artifact that gets shared; the vault is where secrets live.
if !f.secret.is_empty() {
let c = self.upsert(&format!("cred:{}", f.account), NodeKind::Credential, "credential (vaulted)", run);
self.link(&a, &c, EdgeKind::Grants, 0.9, false, run);
}
}
if sev_weight(&f.severity) >= 0.75 || f.stage == "impact" {
let label = if f.business_impact.is_empty() { f.impact.clone() } else { f.business_impact.clone() };
let label: String = label.split_whitespace().take(12).collect::<Vec<_>>().join(" ");
if !label.is_empty() {
let i = self.upsert(&format!("impact:{}:{}", run, f.id), NodeKind::Impact, &label, run);
self.link(&fid, &i, EdgeKind::Leads, sev_weight(&f.severity), false, run);
}
}
}
// Asserted chains first: they are evidence.
for f in findings {
for src in &f.chains_from {
let a = format!("find:{}:{}", run, src);
let b = format!("find:{}:{}", run, f.id);
self.link(&a, &b, EdgeKind::Chains, 0.9, false, run);
}
}
self.infer_chains(run, findings);
}
/// Connect consecutive kill-chain stages when the agents asserted nothing.
///
/// Only forward moves, only between *adjacent populated* stages, and only
/// from the strongest finding of the earlier stage — a full cross-product
/// would draw a plausible-looking web that encodes no information at all.
fn infer_chains(&mut self, run: &str, findings: &[Finding]) {
let asserted: usize = findings.iter().map(|f| f.chains_from.len()).sum();
if asserted > 0 || findings.len() < 2 {
return;
}
let mut by_stage: BTreeMap<usize, Vec<&Finding>> = BTreeMap::new();
for f in findings {
by_stage.entry(stage_rank(&f.stage)).or_default().push(f);
}
let ranks: Vec<usize> = by_stage.keys().copied().collect();
for w in ranks.windows(2) {
let (Some(a), Some(b)) = (by_stage.get(&w[0]), by_stage.get(&w[1])) else { continue };
let best = a
.iter()
.max_by(|x, y| {
(sev_weight(&x.severity) * x.confidence)
.partial_cmp(&(sev_weight(&y.severity) * y.confidence))
.unwrap_or(std::cmp::Ordering::Equal)
})
.copied();
let Some(src) = best else { continue };
for dst in b {
let from = format!("find:{}:{}", run, src.id);
let to = format!("find:{}:{}", run, dst.id);
self.link(&from, &to, EdgeKind::Chains, 0.35, true, run);
}
}
}
pub fn neighbors(&self, id: &str) -> Vec<&Edge> {
self.edges.iter().filter(|e| e.from == id).collect()
}
/// Entities observed but never proved: endpoints with no finding on them,
/// accounts nothing was done with. These are where chaining should look
/// next, and the reason the graph is worth keeping between runs.
pub fn frontier(&self) -> Vec<&Node> {
let proven: BTreeSet<&str> = self
.edges
.iter()
.filter(|e| matches!(e.kind, EdgeKind::ObservedOn | EdgeKind::Proves))
.map(|e| e.to.as_str())
.collect();
let mut v: Vec<&Node> = self
.nodes
.values()
.filter(|n| matches!(n.kind, NodeKind::Endpoint | NodeKind::Account | NodeKind::Credential))
.filter(|n| !proven.contains(n.id.as_str()))
.collect();
v.sort_by(|a, b| b.runs.len().cmp(&a.runs.len()).then_with(|| a.id.cmp(&b.id)));
v
}
/// Ranked attack paths: chains of findings ordered by kill-chain stage,
/// scored by severity × edge probability. Returns node-id paths, longest and
/// most damaging first.
pub fn paths(&self, max: usize) -> Vec<(Vec<String>, f64)> {
let findings: Vec<&Node> = self.nodes.values().filter(|n| n.kind == NodeKind::Finding).collect();
let has_parent: BTreeSet<&str> = self
.edges
.iter()
.filter(|e| e.kind == EdgeKind::Chains)
.map(|e| e.to.as_str())
.collect();
let roots: Vec<&Node> = findings.iter().copied().filter(|n| !has_parent.contains(n.id.as_str())).collect();
let mut out: Vec<(Vec<String>, f64)> = Vec::new();
for r in roots {
let mut stack = vec![(vec![r.id.clone()], self.node_score(&r.id))];
while let Some((path, score)) = stack.pop() {
let last = path.last().cloned().unwrap_or_default();
let next: Vec<&Edge> = self
.edges
.iter()
.filter(|e| e.kind == EdgeKind::Chains && e.from == last && !path.contains(&e.to))
.collect();
if next.is_empty() {
out.push((path, score));
continue;
}
for e in next {
let mut p = path.clone();
p.push(e.to.clone());
// Depth is capped: cycles are already excluded, but a long
// inferred tail is noise, not a deeper attack.
if p.len() > 12 {
out.push((p, score));
continue;
}
let s = score + self.node_score(&e.to) * e.p;
stack.push((p, s));
}
}
}
out.sort_by(|a, b| {
b.0.len()
.cmp(&a.0.len())
.then_with(|| b.1.partial_cmp(&a.1).unwrap_or(std::cmp::Ordering::Equal))
});
out.truncate(if max == 0 { 5 } else { max });
out
}
fn node_score(&self, id: &str) -> f64 {
self.nodes
.get(id)
.map(|n| {
let sev = n.meta.get("severity").map(|s| sev_weight(s)).unwrap_or(0.1);
let conf: f64 = n.meta.get("confidence").and_then(|c| c.parse().ok()).unwrap_or(0.5);
sev * conf.clamp(0.2, 1.0)
})
.unwrap_or(0.0)
}
/// One line per kind, then the top attack paths — the `/graph` view.
pub fn summary(&self) -> String {
if self.nodes.is_empty() {
return " (knowledge graph empty — run an engagement first)".into();
}
let mut by_kind: BTreeMap<&str, usize> = BTreeMap::new();
for n in self.nodes.values() {
*by_kind.entry(n.kind.as_str()).or_insert(0) += 1;
}
let mut s = String::from(" ┌ knowledge graph\n");
for (k, v) in &by_kind {
s.push_str(&format!(" │ {:<12} {}\n", k, v));
}
let inferred = self.edges.iter().filter(|e| e.inferred).count();
s.push_str(&format!(" │ {:<12} {} ({} inferred)\n", "edges", self.edges.len(), inferred));
let paths = self.paths(3);
if paths.iter().any(|(p, _)| p.len() > 1) {
s.push_str(" │\n │ top attack paths\n");
for (p, score) in paths.iter().filter(|(p, _)| p.len() > 1) {
let labels: Vec<&str> = p
.iter()
.filter_map(|id| self.nodes.get(id))
.map(|n| n.label.as_str())
.collect();
s.push_str(&format!(" │ [{score:.2}] {}\n", labels.join(" → ")));
}
}
let fr = self.frontier();
if !fr.is_empty() {
s.push_str(&format!(" │\n │ frontier ({} unproven): {}\n", fr.len(),
fr.iter().take(4).map(|n| n.label.as_str()).collect::<Vec<_>>().join(", ")));
}
s.push_str(" └\n");
s
}
}
#[cfg(test)]
mod tests {
use super::*;
fn f(id: &str, sev: &str, cwe: &str, stage: &str, endpoint: &str) -> Finding {
Finding {
id: id.into(),
title: format!("finding {id}"),
severity: sev.into(),
cwe: cwe.into(),
stage: stage.into(),
endpoint: endpoint.into(),
confidence: 0.9,
..Default::default()
}
}
#[test]
fn one_endpoint_node_however_the_url_was_written() {
let mut g = KnowledgeGraph::new();
g.ingest(
"https://ex.com",
"r1",
&[
f("a", "High", "CWE-89", "initial-access", "https://ex.com/login.aspx?id=1"),
f("b", "Low", "CWE-200", "recon", "http://www.ex.com/login.aspx/"),
],
);
let eps: Vec<&Node> = g.nodes.values().filter(|n| n.kind == NodeKind::Endpoint).collect();
assert_eq!(eps.len(), 1, "got {:?}", eps.iter().map(|n| &n.id).collect::<Vec<_>>());
}
#[test]
fn asserted_chains_win_and_nothing_is_inferred_alongside_them() {
let mut g = KnowledgeGraph::new();
let mut b = f("b", "High", "CWE-89", "execution", "https://ex.com/x");
b.chains_from = vec!["a".into()];
g.ingest("https://ex.com", "r1", &[f("a", "Medium", "CWE-200", "recon", "https://ex.com/"), b]);
let chains: Vec<&Edge> = g.edges.iter().filter(|e| e.kind == EdgeKind::Chains).collect();
assert_eq!(chains.len(), 1);
assert!(!chains[0].inferred, "an agent-asserted chain must not be marked inferred");
}
#[test]
fn inferred_chains_only_move_forward_through_the_kill_chain() {
let mut g = KnowledgeGraph::new();
g.ingest(
"https://ex.com",
"r1",
&[
f("a", "Medium", "CWE-200", "recon", "https://ex.com/"),
f("b", "Critical", "CWE-89", "initial-access", "https://ex.com/login"),
],
);
let chains: Vec<&Edge> = g.edges.iter().filter(|e| e.kind == EdgeKind::Chains).collect();
assert_eq!(chains.len(), 1);
assert!(chains[0].inferred, "a derived edge must say so");
assert!(chains[0].from.ends_with(":a") && chains[0].to.ends_with(":b"), "recon must precede initial-access");
assert!(chains[0].p < 0.5, "an inferred edge must carry less weight than an asserted one");
}
#[test]
fn a_secret_never_lands_in_the_graph() {
let mut g = KnowledgeGraph::new();
let mut x = f("a", "High", "CWE-287", "credential-access", "https://ex.com/register");
x.account = "user1".into();
x.secret = "hunter2-super-secret".into();
g.ingest("https://ex.com", "r1", &[x]);
let json = g.to_json();
assert!(!json.contains("hunter2"), "the vault holds secrets, the graph does not");
assert!(json.contains("credential (vaulted)"));
}
#[test]
fn paths_rank_the_longest_most_severe_chain_first() {
let mut g = KnowledgeGraph::new();
let mut b = f("b", "High", "CWE-89", "initial-access", "https://ex.com/login");
b.chains_from = vec!["a".into()];
let mut c = f("c", "Critical", "CWE-78", "execution", "https://ex.com/exec");
c.chains_from = vec!["b".into()];
g.ingest("https://ex.com", "r1", &[f("a", "Low", "CWE-200", "recon", "https://ex.com/"), b, c]);
let paths = g.paths(3);
assert_eq!(paths[0].0.len(), 3, "the three-step chain must rank above any single node");
}
#[test]
fn re_ingesting_the_same_run_does_not_duplicate_edges() {
let mut g = KnowledgeGraph::new();
let fs = [f("a", "High", "CWE-89", "initial-access", "https://ex.com/login")];
g.ingest("https://ex.com", "r1", &fs);
let n = g.edges.len();
g.ingest("https://ex.com", "r1", &fs);
assert_eq!(g.edges.len(), n);
}
#[test]
fn the_frontier_lists_endpoints_nothing_was_proven_on() {
let mut g = KnowledgeGraph::new();
g.ingest("https://ex.com", "r1", &[f("a", "High", "CWE-89", "initial-access", "https://ex.com/login")]);
// An endpoint learned by recon, with no finding attached to it.
let asset = "asset:ex.com".to_string();
g.upsert("ep:ex.com/admin", NodeKind::Endpoint, "ex.com/admin", "r1");
g.link(&asset, "ep:ex.com/admin", EdgeKind::Exposes, 0.8, false, "r1");
let fr: Vec<&str> = g.frontier().iter().map(|n| n.id.as_str()).collect();
assert_eq!(fr, vec!["ep:ex.com/admin"]);
}
}
+4
View File
@@ -13,6 +13,8 @@ pub mod creds;
pub mod grounding;
pub mod hygiene;
pub mod integrations;
pub mod knowledge_graph;
pub mod memory;
pub mod pomdp;
pub mod models;
pub mod pipeline;
@@ -29,5 +31,7 @@ pub use models::{
};
pub use pipeline::{run_greybox, run_host, run_whitebox, RunOutput};
pub use pipeline::run;
pub use knowledge_graph::{EdgeKind, KnowledgeGraph, NodeKind};
pub use memory::{Memory, Query as MemoryQuery, Tier as MemoryTier};
pub use pool::{ModelPool, Task};
pub use types::{Finding, RunConfig};
+707
View File
@@ -0,0 +1,707 @@
//! Layered memory for the harness.
//!
//! An engagement is a long-running investigation, but every agent call starts
//! from a blank context window. Without a place to put what was learned, the
//! same facts get re-derived every round — the harness re-probes an endpoint it
//! already fingerprinted, re-tries a payload shape that already failed, and
//! forgets across runs entirely. The RL weights in [`crate::rl`] remember *which
//! agent* pays off; they cannot remember *what was true*.
//!
//! Four tiers, separated by what they are scoped to and how long they survive —
//! not by importance:
//!
//! | tier | scope | lives | example |
//! |------|-------|-------|---------|
//! | [`Tier::Working`] | one run | until the run ends | "`/admin` returned 302 to `/login`" |
//! | [`Tier::Engagement`] | one target | forever, decaying | "this host runs IIS 8.5 / ASP.NET 2.0" |
//! | [`Tier::Technique`] | one technique/agent | forever, decaying | "`sqli_error` lands on `.aspx` id params" |
//! | [`Tier::Reusable`] | nothing (generalized) | forever | "ASP.NET verbose errors leak the ViewState key" |
//!
//! Promotion is evidence-gated and moves *up* the table: a working memo repeated
//! within a run becomes engagement knowledge; engagement knowledge confirmed on
//! a second run becomes technique knowledge; a technique memo that holds on two
//! **different targets** is generalized into reusable knowledge with the
//! target-specific tokens stripped. Nothing is promoted on a single observation,
//! because one observation is exactly how a hallucination looks.
//!
//! Recall is scored, not exhaustive: prompts have a budget, so [`Memory::recall`]
//! ranks by term overlap, how often the memo preceded a real finding, and
//! recency, then [`Memory::prompt_block`] renders the top few as plain lines.
use serde::{Deserialize, Serialize};
use std::collections::{BTreeMap, HashMap};
use std::path::{Path, PathBuf};
/// Which memory a memo belongs to. See the module docs for the scoping rules.
#[derive(Clone, Copy, Debug, PartialEq, Eq, PartialOrd, Ord, Serialize, Deserialize)]
#[serde(rename_all = "kebab-case")]
pub enum Tier {
Working,
Engagement,
Technique,
Reusable,
}
impl Tier {
pub fn as_str(&self) -> &'static str {
match self {
Tier::Working => "working",
Tier::Engagement => "engagement",
Tier::Technique => "technique",
Tier::Reusable => "reusable",
}
}
}
/// One remembered fact. Deliberately a *sentence*, not a struct of fields: the
/// consumer is a language model, and the thing that has to survive the round
/// trip is the claim, not a schema.
#[derive(Clone, Debug, Serialize, Deserialize)]
pub struct Memo {
pub id: String,
pub tier: Tier,
/// Scope key — target key for engagement, technique id for technique,
/// empty for reusable.
#[serde(default)]
pub key: String,
pub text: String,
#[serde(default)]
pub tags: Vec<String>,
#[serde(default)]
pub source_run: String,
#[serde(default)]
pub target: String,
/// Belief that the claim holds, 0..1.
#[serde(default)]
pub confidence: f64,
/// Distinct runs that observed this.
#[serde(default)]
pub observations: u32,
/// Times this memo was fed into a prompt.
#[serde(default)]
pub uses: u32,
/// Times a run that recalled this memo went on to produce a finding.
#[serde(default)]
pub wins: u32,
#[serde(default)]
pub created: u64,
#[serde(default)]
pub updated: u64,
}
impl Memo {
/// Fraction of recalls that preceded a finding. Unused memos sit at the
/// neutral 0.5 rather than 0 — never having been tried is not evidence of
/// being wrong, and starting them at zero would bury them forever.
pub fn success_rate(&self) -> f64 {
if self.uses == 0 {
0.5
} else {
self.wins as f64 / self.uses as f64
}
}
}
fn now() -> u64 {
std::time::SystemTime::now()
.duration_since(std::time::UNIX_EPOCH)
.map(|d| d.as_secs())
.unwrap_or(0)
}
/// Lowercased alphanumeric terms of length ≥ 3, deduped. Used for both indexing
/// and query matching so a memo and a query are compared the same way.
pub fn terms(s: &str) -> Vec<String> {
let mut out: Vec<String> = Vec::new();
for raw in s.split(|c: char| !c.is_alphanumeric() && c != '-' && c != '_' && c != '.') {
let t = raw.trim_matches(|c: char| c == '.' || c == '-' || c == '_').to_lowercase();
if t.len() >= 3 && !out.contains(&t) {
out.push(t);
}
}
out
}
/// Stable identity of a claim: same normalized wording = same memo, so repeating
/// an observation reinforces it instead of duplicating it.
fn fingerprint(tier: Tier, key: &str, text: &str) -> String {
let norm: String = text
.to_lowercase()
.chars()
.filter(|c| c.is_alphanumeric() || c.is_whitespace())
.collect::<String>()
.split_whitespace()
.collect::<Vec<_>>()
.join(" ");
let mut h: u64 = 0xcbf2_9ce4_8422_2325;
for b in format!("{}|{}|{}", tier.as_str(), key, norm).bytes() {
h ^= b as u64;
h = h.wrapping_mul(0x1000_0000_01b3);
}
format!("{}-{:016x}", tier.as_str(), h)
}
/// Normalize a target into a stable engagement key: scheme, port, path, `www.`
/// and case all drop out, so `https://WWW.Example.com:443/login` and
/// `http://example.com/` are one engagement and not two.
pub fn engagement_key(target: &str) -> String {
let t = target.trim().to_lowercase();
let t = t.split_once("://").map(|(_, rest)| rest).unwrap_or(&t);
let t = t.split(['/', '?', '#']).next().unwrap_or(t);
let t = t.rsplit_once(':').map(|(h, p)| if p.chars().all(|c| c.is_ascii_digit()) { h } else { t }).unwrap_or(t);
t.trim_start_matches("www.").trim().to_string()
}
/// What to recall for.
#[derive(Debug, Clone, Default)]
pub struct Query {
/// Free text — agent prompt, objective, endpoint, whatever is at hand.
pub text: String,
/// Restrict engagement recall to this target (empty = any).
pub target: String,
/// Restrict technique recall to these ids (empty = any).
pub techniques: Vec<String>,
/// Tiers to search. Empty means all but [`Tier::Working`].
pub tiers: Vec<Tier>,
pub limit: usize,
}
/// A scored recall hit.
#[derive(Debug, Clone)]
pub struct Hit {
pub memo: Memo,
pub score: f64,
}
/// The four-tier store. Persisted under `<dir>/` as one file per tier; working
/// memory is written too, so a crashed run can be resumed with its scratchpad
/// intact instead of restarting cold.
#[derive(Default)]
pub struct Memory {
dir: Option<PathBuf>,
working: Vec<Memo>,
engagement: BTreeMap<String, Vec<Memo>>,
technique: BTreeMap<String, Vec<Memo>>,
reusable: Vec<Memo>,
/// Fingerprints seen this run — the promotion gate for working → engagement.
seen_this_run: HashMap<String, u32>,
/// Memo ids injected into prompts during this run, so a run that lands a
/// finding can credit what it was told beforehand.
recalled: Vec<String>,
}
/// Process-wide store for the current project.
///
/// Recall happens while prompts are built and reinforcement happens when the
/// run finishes — far apart in the call graph, with the async pipeline in
/// between. Two independently opened handles would each hold a stale copy and
/// the last one to save would silently discard the other's counters, so the
/// process shares one. The directory is bound on first call; later calls return
/// that same store regardless of the path passed, which is correct because one
/// CLI process serves one project.
pub fn shared(dir: impl AsRef<Path>) -> &'static std::sync::Mutex<Memory> {
static STORE: std::sync::OnceLock<std::sync::Mutex<Memory>> = std::sync::OnceLock::new();
STORE.get_or_init(|| std::sync::Mutex::new(Memory::open(dir)))
}
/// Working memory is a scratchpad, not a log: past this many memos the oldest
/// go, because a run that emits thousands of lines would otherwise turn recall
/// into a scan of its own noise.
const WORKING_CAP: usize = 400;
/// Confidence floor below which a never-useful memo is pruned on save.
const PRUNE_BELOW: f64 = 0.15;
impl Memory {
/// In-memory only — used by tests and by callers with no project dir.
pub fn ephemeral() -> Memory {
Memory::default()
}
/// Open (or create) the store under `dir`, e.g. `.neurosploit/memory`.
pub fn open(dir: impl AsRef<Path>) -> Memory {
let dir = dir.as_ref().to_path_buf();
let _ = std::fs::create_dir_all(&dir);
let read = |name: &str| -> Option<String> { std::fs::read_to_string(dir.join(name)).ok() };
Memory {
working: read("working.json").and_then(|s| serde_json::from_str(&s).ok()).unwrap_or_default(),
engagement: read("engagement.json").and_then(|s| serde_json::from_str(&s).ok()).unwrap_or_default(),
technique: read("technique.json").and_then(|s| serde_json::from_str(&s).ok()).unwrap_or_default(),
reusable: read("reusable.json").and_then(|s| serde_json::from_str(&s).ok()).unwrap_or_default(),
dir: Some(dir),
seen_this_run: HashMap::new(),
recalled: Vec::new(),
}
}
pub fn counts(&self) -> (usize, usize, usize, usize) {
(
self.working.len(),
self.engagement.values().map(|v| v.len()).sum(),
self.technique.values().map(|v| v.len()).sum(),
self.reusable.len(),
)
}
fn bucket_mut(&mut self, tier: Tier, key: &str) -> &mut Vec<Memo> {
match tier {
Tier::Working => &mut self.working,
Tier::Engagement => self.engagement.entry(key.to_string()).or_default(),
Tier::Technique => self.technique.entry(key.to_string()).or_default(),
Tier::Reusable => &mut self.reusable,
}
}
/// Record a claim. Re-recording the same claim reinforces it (confidence
/// rises toward 1, observation count grows) instead of adding a duplicate,
/// which is what makes "seen twice" a meaningful promotion signal.
pub fn remember(&mut self, tier: Tier, key: &str, text: &str, tags: &[&str], target: &str, run: &str, confidence: f64) -> String {
let text = text.trim();
if text.is_empty() {
return String::new();
}
let id = fingerprint(tier, key, text);
*self.seen_this_run.entry(id.clone()).or_insert(0) += 1;
let ts = now();
let bucket = self.bucket_mut(tier, key);
if let Some(m) = bucket.iter_mut().find(|m| m.id == id) {
// Bounded reinforcement: each repeat closes 35% of the remaining gap
// to certainty, so a claim asymptotically approaches — but never
// reaches — "known", which is the honest shape for an observation.
m.confidence = (m.confidence + 0.35 * (1.0 - m.confidence)).clamp(0.0, 0.99);
m.observations += 1;
m.updated = ts;
if m.source_run != run && !run.is_empty() {
m.source_run = run.to_string();
}
for t in tags {
if !m.tags.iter().any(|x| x == t) {
m.tags.push((*t).to_string());
}
}
return id;
}
bucket.push(Memo {
id: id.clone(),
tier,
key: key.to_string(),
text: text.to_string(),
tags: tags.iter().map(|s| s.to_string()).collect(),
source_run: run.to_string(),
target: target.to_string(),
confidence: confidence.clamp(0.0, 0.99),
observations: 1,
uses: 0,
wins: 0,
created: ts,
updated: ts,
});
if tier == Tier::Working && self.working.len() > WORKING_CAP {
let drop = self.working.len() - WORKING_CAP;
self.working.drain(0..drop);
}
id
}
/// Convenience: note something learned about the target during this run.
pub fn note(&mut self, target: &str, run: &str, text: &str, tags: &[&str]) -> String {
self.remember(Tier::Working, &engagement_key(target), text, tags, target, run, 0.5)
}
/// Rank memos against a query. Scoring blends three signals that answer
/// three different questions: overlap ("is this about what I'm doing?"),
/// success rate ("did acting on it ever pay off?") and recency ("is it
/// still likely to be true?"). Confidence gates the whole thing, so a
/// once-observed guess cannot outrank a repeatedly confirmed fact.
pub fn recall(&self, q: &Query) -> Vec<Hit> {
let qterms = terms(&q.text);
let tkey = engagement_key(&q.target);
let tiers: Vec<Tier> = if q.tiers.is_empty() {
vec![Tier::Engagement, Tier::Technique, Tier::Reusable]
} else {
q.tiers.clone()
};
let ts = now();
let mut hits: Vec<Hit> = Vec::new();
let mut consider = |m: &Memo| {
let mterms = terms(&format!("{} {}", m.text, m.tags.join(" ")));
let overlap = if qterms.is_empty() || mterms.is_empty() {
0.0
} else {
let inter = qterms.iter().filter(|t| mterms.contains(t)).count() as f64;
inter / (qterms.len() as f64).sqrt().max(1.0) / (mterms.len() as f64).sqrt().max(1.0)
};
// Half-life of 30 days: a fingerprint from last week is worth more
// than one from last quarter, but never worthless.
let age_days = (ts.saturating_sub(m.updated)) as f64 / 86_400.0;
let recency = 0.5f64.powf(age_days / 30.0);
let score = m.confidence * (0.55 * overlap.min(1.0) + 0.25 * m.success_rate() + 0.20 * recency);
if score > 0.0 {
hits.push(Hit { memo: m.clone(), score });
}
};
for tier in tiers {
match tier {
Tier::Working => self.working.iter().for_each(&mut consider),
Tier::Engagement => {
for (k, v) in &self.engagement {
if tkey.is_empty() || *k == tkey {
v.iter().for_each(&mut consider);
}
}
}
Tier::Technique => {
for (k, v) in &self.technique {
if q.techniques.is_empty() || q.techniques.iter().any(|t| t == k) {
v.iter().for_each(&mut consider);
}
}
}
Tier::Reusable => self.reusable.iter().for_each(&mut consider),
}
}
hits.sort_by(|a, b| b.score.partial_cmp(&a.score).unwrap_or(std::cmp::Ordering::Equal));
let limit = if q.limit == 0 { 8 } else { q.limit };
hits.truncate(limit);
hits
}
/// Render recalled memos as a prompt section, and mark them used so their
/// success rate can be scored against what the run finds. Returns an empty
/// string when nothing is worth injecting — an empty "what you know" header
/// is worse than none, it invites the model to invent the contents.
pub fn prompt_block(&mut self, q: &Query) -> String {
let hits = self.recall(q);
if hits.is_empty() {
return String::new();
}
let ids: Vec<String> = hits.iter().map(|h| h.memo.id.clone()).collect();
self.mark_used(&ids);
for id in &ids {
if !self.recalled.contains(id) {
self.recalled.push(id.clone());
}
}
let mut out = String::from("## What NeuroSploit already knows (prior engagements)\n\nTreat as leads, not facts — verify before reporting.\n");
for h in &hits {
out.push_str(&format!(
"- [{} · {:.0}%] {}\n",
h.memo.tier.as_str(),
h.memo.confidence * 100.0,
h.memo.text
));
}
out
}
fn all_mut(&mut self) -> impl Iterator<Item = &mut Memo> {
self.working
.iter_mut()
.chain(self.engagement.values_mut().flatten())
.chain(self.technique.values_mut().flatten())
.chain(self.reusable.iter_mut())
}
pub fn mark_used(&mut self, ids: &[String]) {
for m in self.all_mut() {
if ids.iter().any(|i| *i == m.id) {
m.uses += 1;
}
}
}
/// Credit every memo recalled during a run that produced findings. This is
/// what turns recall from a guess into a measurement over time.
pub fn mark_win(&mut self, ids: &[String]) {
for m in self.all_mut() {
if ids.iter().any(|i| *i == m.id) {
m.wins += 1;
m.confidence = (m.confidence + 0.1).min(0.99);
}
}
}
/// Credit everything recalled during this run. Call it only when the run
/// actually produced findings — that is the whole signal.
pub fn credit_recalled(&mut self) {
let ids = std::mem::take(&mut self.recalled);
self.mark_win(&ids);
}
/// Promote what this run proved, then persist. Called once at the end of a
/// run; `run` is the run id and `target` the engagement's target.
///
/// Every step needs *independent* evidence, so nothing here can be triggered
/// twice by one loud observation:
/// - working → engagement: the claim recurred within the run;
/// - engagement → technique: it also carries a technique tag and has been
/// observed in more than one run;
/// - technique → reusable: it held on two different targets, and the
/// generalized copy has target-specific tokens stripped.
pub fn consolidate(&mut self, target: &str, run: &str) -> (usize, usize, usize) {
let key = engagement_key(target);
let mut to_engagement: Vec<Memo> = Vec::new();
for m in &self.working {
if self.seen_this_run.get(&m.id).copied().unwrap_or(0) >= 2 || m.observations >= 2 {
to_engagement.push(m.clone());
}
}
let promoted_e = to_engagement.len();
for m in to_engagement {
let tags: Vec<&str> = m.tags.iter().map(|s| s.as_str()).collect();
self.remember(Tier::Engagement, &key, &m.text, &tags, target, run, m.confidence.max(0.55));
}
self.working.clear();
// engagement → technique
let mut to_technique: Vec<(String, Memo)> = Vec::new();
for memos in self.engagement.values() {
for m in memos {
if m.observations < 2 {
continue;
}
if let Some(t) = m.tags.iter().find(|t| t.starts_with("technique:") || t.starts_with("cwe:") || t.starts_with("agent:")) {
to_technique.push((t.clone(), m.clone()));
}
}
}
let promoted_t = to_technique.len();
for (tech, m) in to_technique {
let tags: Vec<&str> = m.tags.iter().map(|s| s.as_str()).collect();
self.remember(Tier::Technique, &tech, &m.text, &tags, target, run, m.confidence);
}
// technique → reusable: needs two distinct targets.
let mut to_reusable: Vec<Memo> = Vec::new();
for memos in self.technique.values() {
let mut by_text: HashMap<String, Vec<&Memo>> = HashMap::new();
for m in memos {
by_text.entry(generalize(&m.text)).or_default().push(m);
}
for (gen, group) in by_text {
let mut targets: Vec<&str> = group.iter().map(|m| m.target.as_str()).filter(|t| !t.is_empty()).collect();
targets.sort_unstable();
targets.dedup();
if targets.len() >= 2 {
let best = group.iter().map(|m| m.confidence).fold(0.0f64, f64::max);
let mut m = (*group[0]).clone();
m.text = gen;
m.confidence = best;
to_reusable.push(m);
}
}
}
let promoted_r = to_reusable.len();
for m in to_reusable {
let tags: Vec<&str> = m.tags.iter().map(|s| s.as_str()).collect();
self.remember(Tier::Reusable, "", &m.text, &tags, "", run, m.confidence);
}
self.seen_this_run.clear();
self.save();
(promoted_e, promoted_t, promoted_r)
}
/// Drop memos that were never useful and have decayed — otherwise a store
/// that only grows eventually recalls noise as readily as knowledge.
pub fn decay(&mut self, factor: f64) {
let f = factor.clamp(0.5, 1.0);
for m in self.all_mut() {
if m.wins == 0 {
m.confidence *= f;
}
}
let keep = |m: &Memo| m.confidence >= PRUNE_BELOW || m.wins > 0;
self.working.retain(keep);
self.reusable.retain(keep);
for v in self.engagement.values_mut() {
v.retain(keep);
}
for v in self.technique.values_mut() {
v.retain(keep);
}
self.engagement.retain(|_, v| !v.is_empty());
self.technique.retain(|_, v| !v.is_empty());
}
/// Forget by substring across every tier. Returns how many went.
pub fn forget(&mut self, needle: &str) -> usize {
let n = needle.to_lowercase();
if n.is_empty() {
return 0;
}
let before = self.counts();
let drop = |m: &Memo| !(m.text.to_lowercase().contains(&n) || m.id == needle || m.key.to_lowercase() == n);
self.working.retain(drop);
self.reusable.retain(drop);
for v in self.engagement.values_mut() {
v.retain(drop);
}
for v in self.technique.values_mut() {
v.retain(drop);
}
self.engagement.retain(|_, v| !v.is_empty());
self.technique.retain(|_, v| !v.is_empty());
let after = self.counts();
self.save();
(before.0 + before.1 + before.2 + before.3) - (after.0 + after.1 + after.2 + after.3)
}
pub fn save(&self) {
let Some(dir) = &self.dir else { return };
let _ = std::fs::create_dir_all(dir);
let put = |name: &str, v: String| {
let _ = std::fs::write(dir.join(name), v);
};
if let Ok(j) = serde_json::to_string_pretty(&self.working) {
put("working.json", j);
}
if let Ok(j) = serde_json::to_string_pretty(&self.engagement) {
put("engagement.json", j);
}
if let Ok(j) = serde_json::to_string_pretty(&self.technique) {
put("technique.json", j);
}
if let Ok(j) = serde_json::to_string_pretty(&self.reusable) {
put("reusable.json", j);
}
}
/// Everything, newest first — for `/memory` in the REPL and the web console.
pub fn dump(&self) -> Vec<Memo> {
let mut all: Vec<Memo> = self
.working
.iter()
.chain(self.engagement.values().flatten())
.chain(self.technique.values().flatten())
.chain(self.reusable.iter())
.cloned()
.collect();
all.sort_by(|a, b| b.updated.cmp(&a.updated));
all
}
}
/// Strip target-specific tokens so a technique memo can be stated about the
/// class of system rather than the host it was first seen on. A claim that
/// still names one host is not a general lesson.
fn generalize(text: &str) -> String {
let mut out = String::with_capacity(text.len());
for word in text.split_whitespace() {
let w = word.trim_matches(|c: char| c == ',' || c == ';');
let looks_like_host = w.contains("://")
|| (w.contains('.') && w.split('.').count() >= 3 && !w.ends_with('.'))
|| w.chars().filter(|c| *c == '.').count() >= 3;
if looks_like_host {
out.push_str("<target>");
} else {
out.push_str(word);
}
out.push(' ');
}
out.trim().to_string()
}
#[cfg(test)]
mod tests {
use super::*;
#[test]
fn one_engagement_key_per_host_however_the_url_was_written() {
assert_eq!(engagement_key("https://WWW.Example.com:443/login?x=1"), "example.com");
assert_eq!(engagement_key("http://example.com/"), "example.com");
assert_eq!(engagement_key("10.0.0.7"), "10.0.0.7");
}
#[test]
fn repeating_a_claim_reinforces_it_instead_of_duplicating() {
let mut m = Memory::ephemeral();
m.note("http://t.test", "run1", "/admin returns 302 to /login", &["endpoint"]);
m.note("http://t.test", "run1", "/admin returns 302 to /login", &["endpoint"]);
assert_eq!(m.counts().0, 1);
let memo = &m.working[0];
assert_eq!(memo.observations, 2);
assert!(memo.confidence > 0.5, "second observation must raise confidence");
}
#[test]
fn a_claim_seen_once_is_not_promoted_but_a_repeat_is() {
let mut m = Memory::ephemeral();
m.note("http://t.test", "run1", "seen once only", &[]);
m.note("http://t.test", "run1", "seen twice here", &[]);
m.note("http://t.test", "run1", "seen twice here", &[]);
let (to_eng, _, _) = m.consolidate("http://t.test", "run1");
assert_eq!(to_eng, 1);
let eng = &m.engagement["t.test"];
assert_eq!(eng.len(), 1);
assert_eq!(eng[0].text, "seen twice here");
assert!(m.working.is_empty(), "working memory is cleared once consolidated");
}
#[test]
fn technique_knowledge_generalizes_only_after_a_second_target() {
let mut m = Memory::ephemeral();
let claim = "verbose ASP.NET errors on https://a.example.com/x leak the stack trace";
m.remember(Tier::Technique, "cwe:209", claim, &["cwe:209"], "https://a.example.com", "r1", 0.8);
assert_eq!(m.consolidate("https://a.example.com", "r1").2, 0, "one target is not a general lesson");
let claim2 = "verbose ASP.NET errors on https://b.other.org/y leak the stack trace";
m.remember(Tier::Technique, "cwe:209", claim2, &["cwe:209"], "https://b.other.org", "r2", 0.8);
assert!(m.consolidate("https://b.other.org", "r2").2 >= 1);
assert!(
m.reusable.iter().any(|r| r.text.contains("<target>")),
"the reusable copy must not name a specific host: {:?}",
m.reusable.iter().map(|r| &r.text).collect::<Vec<_>>()
);
}
#[test]
fn recall_prefers_the_memo_that_matches_the_question() {
let mut m = Memory::ephemeral();
m.remember(Tier::Engagement, "t.test", "login.aspx is vulnerable to SQL injection in tbUsername", &["cwe:89"], "http://t.test", "r1", 0.9);
m.remember(Tier::Engagement, "t.test", "the site serves a robots.txt with two entries", &["recon"], "http://t.test", "r1", 0.9);
let hits = m.recall(&Query { text: "sql injection on login".into(), target: "http://t.test".into(), limit: 1, ..Default::default() });
assert_eq!(hits.len(), 1);
assert!(hits[0].memo.text.contains("SQL injection"));
}
#[test]
fn an_empty_recall_injects_no_prompt_section() {
let mut m = Memory::ephemeral();
assert_eq!(m.prompt_block(&Query { text: "anything".into(), ..Default::default() }), "");
}
#[test]
fn wins_raise_a_memo_above_an_equally_relevant_one() {
let mut m = Memory::ephemeral();
let a = m.remember(Tier::Reusable, "", "idor on numeric order ids", &[], "", "r1", 0.8);
m.remember(Tier::Reusable, "", "idor on numeric invoice ids", &[], "", "r1", 0.8);
m.mark_used(&[a.clone()]);
m.mark_win(&[a.clone()]);
let hits = m.recall(&Query { text: "idor numeric ids".into(), limit: 2, ..Default::default() });
assert_eq!(hits[0].memo.id, a, "the memo with a win must rank first");
}
#[test]
fn decay_drops_stale_never_useful_memos_and_keeps_proven_ones() {
let mut m = Memory::ephemeral();
let keep = m.remember(Tier::Reusable, "", "proven lesson", &[], "", "r1", 0.5);
m.remember(Tier::Reusable, "", "never useful", &[], "", "r1", 0.2);
m.mark_win(&[keep.clone()]);
for _ in 0..6 {
m.decay(0.7);
}
assert!(m.reusable.iter().any(|x| x.id == keep));
assert!(!m.reusable.iter().any(|x| x.text == "never useful"));
}
#[test]
fn forget_removes_matching_memos_from_every_tier() {
let mut m = Memory::ephemeral();
m.note("http://t.test", "r1", "secret token abc123 in page source", &[]);
m.remember(Tier::Reusable, "", "secret token patterns leak in source maps", &[], "", "r1", 0.6);
assert_eq!(m.forget("secret token"), 2);
assert_eq!(m.counts(), (0, 0, 0, 0));
}
}
@@ -49,12 +49,67 @@ fn operator_directives(cfg: &RunConfig) -> String {
if let Some(auth) = cfg.auth.as_deref().filter(|x| !x.trim().is_empty()) {
s.push_str(&format!("AUTHENTICATION — test as an authenticated user; send this with each request: {auth}\n"));
}
let recalled = memory_directives(cfg);
if !recalled.is_empty() {
s.push_str(&recalled);
}
if !s.is_empty() {
s.push('\n');
}
s
}
/// Where this project's durable state lives. The app already points
/// `vault_dir` at `<cwd>/.neurosploit/vault`, so its parent is the project
/// store; a caller that set neither falls back to the run's own workdir, which
/// keeps a one-off run from writing into an unrelated directory.
pub(crate) fn proj_store(cfg: &RunConfig) -> PathBuf {
if let Some(v) = cfg.vault_dir.as_deref() {
if let Some(parent) = Path::new(v).parent() {
return parent.to_path_buf();
}
}
cfg.workdir
.as_deref()
.map(PathBuf::from)
.unwrap_or_else(|| PathBuf::from(".neurosploit"))
}
/// The run id used as provenance in the graph and memory: the workdir basename
/// (`ns-<ts>-<target>`), which is also what the report and the web console use.
pub(crate) fn run_id(cfg: &RunConfig) -> String {
cfg.workdir
.as_deref()
.and_then(|d| Path::new(d).file_name())
.map(|n| n.to_string_lossy().to_string())
.unwrap_or_default()
}
/// Prior knowledge about this target, injected into recon/exploit prompts.
///
/// An engagement is usually not the first look at a host, but every model call
/// starts blank — without this the harness re-derives the same stack, the same
/// endpoints and the same dead ends on every run. Recall is scored and capped
/// (see [`crate::memory`]) so the block stays a handful of lines, and it is
/// explicitly framed as leads to verify: prior belief must not become an
/// assertion the model is willing to report.
fn memory_directives(cfg: &RunConfig) -> String {
let dir = proj_store(cfg).join("memory");
let q = crate::memory::Query {
text: format!(
"{} {} {}",
cfg.target,
cfg.objective.clone().unwrap_or_default(),
cfg.instructions.clone().unwrap_or_default()
),
target: cfg.target.clone(),
limit: 6,
..Default::default()
};
let Ok(mut mem) = crate::memory::shared(&dir).lock() else { return String::new() };
mem.prompt_block(&q)
}
/// Tool-usage doctrine prepended to recon/exploit prompts so the agent knows
/// exactly what it may use. Best run on Kali Linux (or the Kali Docker image),
/// where these tools are preinstalled.
@@ -485,6 +540,12 @@ pub async fn run(cfg: RunConfig, lib: &Library, pool: &ModelPool, tx: Sender<Str
let selected: Vec<Agent> = ranked.into_iter().take(cap).collect();
let _ = tx.send(format!("selected {} specialist agents (RL-ranked)", selected.len())).await;
let _ = tx.send("offline: no exploitation performed (provide API keys or --subscription to run live)".into()).await;
// Recon still learned something about the target even with no
// exploitation, and that is exactly the kind of knowledge the next run
// should not have to re-derive.
for n in absorb(&cfg, &recon, &[]) {
let _ = tx.send(n).await;
}
let artifacts = persist(&cfg, &recon, "", &[]);
return RunOutput { target: cfg.target.clone(), workdir: cfg.workdir.clone().unwrap_or_default(), findings: vec![], agents_ran: selected.iter().map(|a| a.name.clone()).collect(), candidates: 0, recon, artifacts };
}
@@ -1394,6 +1455,13 @@ async fn finish(cfg: RunConfig, _lib: &Library, recon: String, transcript: Strin
let _ = tx.send("RL rewards updated".into()).await;
}
// Durable knowledge. Everything above this point is about *this* run; these
// two stores are what makes the next one start from further along.
let notes = absorb(&cfg, &recon, &findings);
for n in notes {
let _ = tx.send(n).await;
}
let artifacts = persist(&cfg, &recon, &transcript, &findings);
if !artifacts.is_empty() {
let _ = tx.send(format!("notify: evidence saved → {}", cfg.workdir.clone().unwrap_or_default())).await;
@@ -1419,6 +1487,102 @@ async fn finish(cfg: RunConfig, _lib: &Library, recon: String, transcript: Strin
}
}
/// Fold a finished run into the attack knowledge graph and the layered memory.
///
/// Returns the lines to report to the operator. It is deliberately synchronous
/// and returns its messages instead of sending them: the memory store is behind
/// a `std::sync::Mutex`, and holding that guard across an `.await` would make
/// the pipeline future non-`Send`.
fn absorb(cfg: &RunConfig, recon: &str, findings: &[Finding]) -> Vec<String> {
let mut out = Vec::new();
let rid = run_id(cfg);
let store = proj_store(cfg);
// Graph: the project-wide one accumulates across runs, and each run keeps
// its own copy so the report and the web console can draw just this run.
let proj_graph = store.join("graph.json");
let mut kg = crate::knowledge_graph::KnowledgeGraph::load(&proj_graph);
kg.ingest(&cfg.target, &rid, findings);
kg.save(&proj_graph);
if let Some(dir) = cfg.workdir.as_deref() {
let mut run_kg = crate::knowledge_graph::KnowledgeGraph::new();
run_kg.ingest(&cfg.target, &rid, findings);
run_kg.save(Path::new(dir).join("graph.json"));
}
let paths = kg.paths(1);
let depth = paths.first().map(|(p, _)| p.len()).unwrap_or(0);
out.push(format!(
"knowledge graph: {} node(s), {} edge(s), longest attack path {} step(s) → graph.json",
kg.nodes.len(),
kg.edges.len(),
depth
));
// Memory: what was proven is engagement knowledge immediately (a validated
// finding is evidence, not a guess); everything else has to earn promotion.
let key = crate::memory::engagement_key(&cfg.target);
let Ok(mut mem) = crate::memory::shared(store.join("memory")).lock() else { return out };
for f in findings {
let where_ = if f.endpoint.is_empty() { cfg.target.as_str() } else { f.endpoint.as_str() };
let text = format!("{} — {} on {} [{} · {}]", f.title, f.severity, where_, f.cwe, f.stage);
let mut tags: Vec<String> = vec![format!("agent:{}", f.agent)];
if !f.cwe.is_empty() {
tags.push(format!("cwe:{}", f.cwe));
}
if !f.mitre.is_empty() {
tags.push(format!("technique:{}", f.mitre));
}
if !f.stage.is_empty() {
tags.push(f.stage.clone());
}
let refs: Vec<&str> = tags.iter().map(|s| s.as_str()).collect();
mem.remember(
crate::memory::Tier::Engagement,
&key,
&text,
&refs,
&cfg.target,
&rid,
f.confidence.clamp(0.4, 0.95),
);
}
// Recon facts go to working memory: one sighting of an endpoint is a lead,
// and only a repeat earns a place in the engagement's knowledge.
if let Ok(v) = serde_json::from_str::<serde_json::Value>(recon) {
for k in ["endpoints", "apis", "hosts", "subdomains"] {
if let Some(arr) = v.get(k).and_then(|x| x.as_array()) {
for item in arr.iter().take(25) {
let s = item.as_str().map(|s| s.to_string()).unwrap_or_else(|| item.to_string());
let s = s.trim_matches('"').trim();
if s.len() > 2 {
mem.note(&cfg.target, &rid, &format!("{k}: {s}"), &["recon", k]);
}
}
}
}
if let Some(tech) = v.get("tech").and_then(|x| x.as_array()) {
let list: Vec<String> = tech.iter().filter_map(|t| t.as_str().map(|s| s.to_string())).collect();
if !list.is_empty() {
mem.note(&cfg.target, &rid, &format!("stack: {}", list.join(", ")), &["recon", "tech"]);
}
}
}
// Recall only earns credit when the run it informed actually found something.
if !findings.is_empty() {
mem.credit_recalled();
}
mem.decay(0.98);
let (e, t, r) = mem.consolidate(&cfg.target, &rid);
let (w, eng, tech, reuse) = mem.counts();
out.push(format!(
"memory: +{e} engagement, +{t} technique, +{r} reusable (now {w}/{eng}/{tech}/{reuse}) → memory/"
));
out
}
/// Write recon/exploit/findings/report as json+md for downstream reuse.
fn persist(cfg: &RunConfig, recon: &str, transcript: &str, findings: &[Finding]) -> Vec<String> {
let Some(dir) = &cfg.workdir else { return vec![] };
+103 -2
View File
@@ -86,6 +86,9 @@ pub struct ModelPool {
/// When this exceeds `AUTH_FAIL_THRESHOLD`, the pool auto-pauses instead of
/// burning through the remaining agents on a dead token.
consecutive_auth_fails: Arc<std::sync::atomic::AtomicUsize>,
/// Backends already tried as an automatic fallback, so a failing one is not
/// retried in a loop.
tried_auto: Arc<Mutex<Vec<String>>>,
}
impl ModelPool {
@@ -119,6 +122,7 @@ impl ModelPool {
resume: Arc::new(Notify::new()),
fallback: Arc::new(Mutex::new(Vec::new())),
consecutive_auth_fails: Arc::new(std::sync::atomic::AtomicUsize::new(0)),
tried_auto: Arc::new(Mutex::new(Vec::new())),
}
}
@@ -325,9 +329,28 @@ impl ModelPool {
}
}
}
// Every candidate failed. Park the run (keeping all state) so the user
// can fix auth or wait for quota renewal, then /continue.
// Every configured candidate failed. Before parking the run and
// waiting for a human, use whatever else this machine can actually
// reach — another logged-in CLI subscription, or a provider whose
// API key is in the environment. A run that stops because one
// provider ran out of quota, on a box with three other usable
// backends, is a run that stopped for no reason.
if (auth_failed || exhausted) && !self.is_cancelled() {
if let Some(alt) = self.next_auto_fallback(&order) {
if let Some(tx) = self.progress() {
let _ = tx.send(format!(
"notify: ⇄ {} unavailable — falling back to {}:{} and continuing.",
order.first().map(|m| m.provider.clone()).unwrap_or_default(),
alt.provider, alt.model
)).await;
}
if let Ok(mut fb) = self.fallback.lock() {
fb.insert(0, alt.clone());
}
self.reset_auth_fails();
continue;
}
// Nothing else is reachable — now a human really is required.
self.park_exhausted(&last, auth_failed).await;
continue;
}
@@ -335,6 +358,48 @@ impl ModelPool {
}
}
/// A backend this machine can use right now that is not already in `tried`
/// and not already a candidate.
///
/// Two sources, in this order: a subscription CLI that is installed (the
/// operator already logged into it, and it costs no API key), then any
/// provider whose API key is present in the environment. Each is offered
/// once — a backend that also fails is recorded so the loop cannot spin.
pub fn next_auto_fallback(&self, current: &[ModelRef]) -> Option<ModelRef> {
let mut tried = self.tried_auto.lock().ok()?;
let known = |p: &str, m: &str, tried: &Vec<String>| {
current.iter().any(|c| c.provider == p && c.model == m) || tried.iter().any(|t| t == &format!("{p}:{m}"))
};
let installed = crate::models::installed_cli_backends();
for pr in crate::models::providers() {
if pr.kind != "cli" {
continue;
}
let Some(bin) = crate::models::cli_binary_for(pr.key) else { continue };
if !installed.contains(&bin) {
continue;
}
let Some(model) = pr.models.first() else { continue };
if known(pr.key, model, &tried) {
continue;
}
tried.push(format!("{}:{}", pr.key, model));
return Some(ModelRef { provider: pr.key.to_string(), model: (*model).to_string() });
}
for pr in crate::models::providers() {
if std::env::var(pr.env_key).ok().filter(|v| !v.trim().is_empty()).is_none() {
continue;
}
let Some(model) = pr.models.first() else { continue };
if known(pr.key, model, &tried) {
continue;
}
tried.push(format!("{}:{}", pr.key, model));
return Some(ModelRef { provider: pr.key.to_string(), model: (*model).to_string() });
}
None
}
/// Reorder candidates for a task. With a single-model panel this is a no-op.
pub fn route(&self, task: Task) -> Vec<ModelRef> {
let mut order = self.candidates.clone();
@@ -457,6 +522,42 @@ pub fn quorum_confirmed(severity: &str, yes: usize, total: usize) -> bool {
#[cfg(test)]
mod verdict_tests {
/// Whatever this machine happens to have installed, the automatic fallback
/// must never re-offer a model already in the panel and never offer the
/// same one twice — either turns "keep going" into a spin.
#[test]
fn auto_fallback_never_repeats_itself_or_the_current_panel() {
let current = vec![ModelRef::parse("anthropic:claude-opus-4-8")];
let pool = ModelPool::new(current.clone(), 1);
let mut seen: Vec<String> = Vec::new();
for _ in 0..8 {
let Some(m) = pool.next_auto_fallback(&current) else { break };
let id = format!("{}:{}", m.provider, m.model);
assert!(
!(m.provider == "anthropic" && m.model == "claude-opus-4-8"),
"offered the model that just failed"
);
assert!(!seen.contains(&id), "offered {id} twice");
seen.push(id);
}
}
/// A provider whose key is in the environment is reachable, so it must be
/// offered before the run parks and waits for a human.
#[test]
fn a_provider_with_a_key_in_the_environment_is_offered() {
std::env::set_var("DEEPSEEK_API_KEY", "test-key-for-fallback");
let current = vec![ModelRef::parse("anthropic:claude-opus-4-8")];
let pool = ModelPool::new(current.clone(), 1);
let mut found = false;
for _ in 0..30 {
let Some(m) = pool.next_auto_fallback(&current) else { break };
if m.provider == "deepseek" { found = true; break; }
}
std::env::remove_var("DEEPSEEK_API_KEY");
assert!(found, "a provider with a usable API key must be reachable as a fallback");
}
use super::*;
#[test]
fn parses_json_and_prose() {