feat: attack knowledge graph, layered memory, command rectification, FAIR dashboard

Backend
-------
- knowledge_graph.rs — the durable structure under attack_graph's per-run view:
  typed entities (asset/endpoint/weakness/technique/finding/account/credential/
  impact) joined by typed, weighted, provenance-carrying edges, accumulated
  across runs in .neurosploit/graph.json plus a per-run copy the report and web
  console can draw. Answers what a finding list can't: ranked attack paths, and
  the frontier of entities observed but never proven — where chaining should
  look next. Agents only sometimes fill chains_from, so progression is also
  inferred between adjacent kill-chain stages; those edges are marked inferred,
  weighted lower, and drawn dashed, because presenting a hypothesis as evidence
  is the graph lying about itself. Secrets stay in the vault, never the graph.

- memory.rs — four tiers scoped by lifetime, not importance: working (one run),
  engagement (one target), technique (one agent/CWE), reusable (generalized).
  Promotion is evidence-gated and needs independent evidence at each step: a
  claim repeated within a run becomes engagement knowledge; one confirmed
  across runs becomes technique knowledge; one that held on two DIFFERENT
  targets is generalized into a reusable lesson with host-specific tokens
  stripped. Nothing is promoted on a single observation, which is exactly what
  a hallucination looks like. Recall is scored (overlap × past success ×
  recency) and injected into recon/exploit prompts as leads to verify. Recalled
  memos are credited only when the run they informed actually found something.

- rectify.rs — a mistyped command cost a full round trip through /help, at the
  worst possible moment during a live run. Accepted-as-typed wins over
  everything (so the /url alias is never "corrected" to /ua), then unique
  prefix, then Damerau-Levenshtein with a length-scaled budget, and a tie is
  reported rather than resolved. Arguments too: a bare host gets its scheme, an
  out-of-range count is clamped with a note instead of silently reverting, a
  near-miss model id is matched against the live catalog.

- pool.rs — when every configured model is exhausted or its token is dead, try
  whatever else this machine can actually reach (an installed CLI subscription,
  or a provider whose key is in the environment) before parking. A run that
  stops on a box with three other usable backends stopped for no reason.

- repl.rs — /memory, /forget, /graph; a recovered run resumes by itself where
  nobody is watching (piped stdin — the web console — or NEUROSPLOIT_AUTO_RESUME),
  since a `/continue` prompt there waits forever.

Web
---
- Attack path: the stage list was seven hardcoded values, so findings the
  harness staged outside it were silently dropped — 5 of 27 on a real run.
  Rewritten against the harness's own stage list with unknown stages kept,
  two-line labels (every node used to read "SQL Injection Authent…"), stage
  column headers, pan/zoom/fit, path highlighting, severity filter, and the
  run's graph.json used when present.
- Dashboard: coverage, findings by severity, top weaknesses, and annualized
  loss exposure via FAIR — frequency from exploitability × validation
  confidence, magnitude from assumptions shown on screen and editable, reported
  as a range. The posture score saturates instead of subtracting, so it keeps
  discriminating past the first critical.
- Run history groups into one folder per target with a filter, instead of one
  flat list that grows forever.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BvdGy9XtVWSdXDTa3FFLJv
This commit is contained in:
CyberSecurityUPandClaude Opus 5 committed 2026-09-07 15:35:36 -03:00
1 parent 0ef0ce8d94
commit 9d83cb6e30
13 files changed
+3163 -119

No files matched your search

+103 -2
View File
@@ -86,6 +86,9 @@ pub struct ModelPool {
/// When this exceeds `AUTH_FAIL_THRESHOLD`, the pool auto-pauses instead of
/// burning through the remaining agents on a dead token.
consecutive_auth_fails: Arc<std::sync::atomic::AtomicUsize>,
/// Backends already tried as an automatic fallback, so a failing one is not
/// retried in a loop.
tried_auto: Arc<Mutex<Vec<String>>>,
}
impl ModelPool {
@@ -119,6 +122,7 @@ impl ModelPool {
resume: Arc::new(Notify::new()),
fallback: Arc::new(Mutex::new(Vec::new())),
consecutive_auth_fails: Arc::new(std::sync::atomic::AtomicUsize::new(0)),
tried_auto: Arc::new(Mutex::new(Vec::new())),
}
}
@@ -325,9 +329,28 @@ impl ModelPool {
}
}
}
// Every candidate failed. Park the run (keeping all state) so the user
// can fix auth or wait for quota renewal, then /continue.
// Every configured candidate failed. Before parking the run and
// waiting for a human, use whatever else this machine can actually
// reach — another logged-in CLI subscription, or a provider whose
// API key is in the environment. A run that stops because one
// provider ran out of quota, on a box with three other usable
// backends, is a run that stopped for no reason.
if (auth_failed || exhausted) && !self.is_cancelled() {
if let Some(alt) = self.next_auto_fallback(&order) {
if let Some(tx) = self.progress() {
let _ = tx.send(format!(
"notify: ⇄ {} unavailable — falling back to {}:{} and continuing.",
order.first().map(|m| m.provider.clone()).unwrap_or_default(),
alt.provider, alt.model
)).await;
}
if let Ok(mut fb) = self.fallback.lock() {
fb.insert(0, alt.clone());
}
self.reset_auth_fails();
continue;
}
// Nothing else is reachable — now a human really is required.
self.park_exhausted(&last, auth_failed).await;
continue;
}
@@ -335,6 +358,48 @@ impl ModelPool {
}
}
/// A backend this machine can use right now that is not already in `tried`
/// and not already a candidate.
///
/// Two sources, in this order: a subscription CLI that is installed (the
/// operator already logged into it, and it costs no API key), then any
/// provider whose API key is present in the environment. Each is offered
/// once — a backend that also fails is recorded so the loop cannot spin.
pub fn next_auto_fallback(&self, current: &[ModelRef]) -> Option<ModelRef> {
let mut tried = self.tried_auto.lock().ok()?;
let known = |p: &str, m: &str, tried: &Vec<String>| {
current.iter().any(|c| c.provider == p && c.model == m) || tried.iter().any(|t| t == &format!("{p}:{m}"))
};
let installed = crate::models::installed_cli_backends();
for pr in crate::models::providers() {
if pr.kind != "cli" {
continue;
}
let Some(bin) = crate::models::cli_binary_for(pr.key) else { continue };
if !installed.contains(&bin) {
continue;
}
let Some(model) = pr.models.first() else { continue };
if known(pr.key, model, &tried) {
continue;
}
tried.push(format!("{}:{}", pr.key, model));
return Some(ModelRef { provider: pr.key.to_string(), model: (*model).to_string() });
}
for pr in crate::models::providers() {
if std::env::var(pr.env_key).ok().filter(|v| !v.trim().is_empty()).is_none() {
continue;
}
let Some(model) = pr.models.first() else { continue };
if known(pr.key, model, &tried) {
continue;
}
tried.push(format!("{}:{}", pr.key, model));
return Some(ModelRef { provider: pr.key.to_string(), model: (*model).to_string() });
}
None
}
/// Reorder candidates for a task. With a single-model panel this is a no-op.
pub fn route(&self, task: Task) -> Vec<ModelRef> {
let mut order = self.candidates.clone();
@@ -457,6 +522,42 @@ pub fn quorum_confirmed(severity: &str, yes: usize, total: usize) -> bool {
#[cfg(test)]
mod verdict_tests {
/// Whatever this machine happens to have installed, the automatic fallback
/// must never re-offer a model already in the panel and never offer the
/// same one twice — either turns "keep going" into a spin.
#[test]
fn auto_fallback_never_repeats_itself_or_the_current_panel() {
let current = vec![ModelRef::parse("anthropic:claude-opus-4-8")];
let pool = ModelPool::new(current.clone(), 1);
let mut seen: Vec<String> = Vec::new();
for _ in 0..8 {
let Some(m) = pool.next_auto_fallback(&current) else { break };
let id = format!("{}:{}", m.provider, m.model);
assert!(
!(m.provider == "anthropic" && m.model == "claude-opus-4-8"),
"offered the model that just failed"
);
assert!(!seen.contains(&id), "offered {id} twice");
seen.push(id);
}
}
/// A provider whose key is in the environment is reachable, so it must be
/// offered before the run parks and waits for a human.
#[test]
fn a_provider_with_a_key_in_the_environment_is_offered() {
std::env::set_var("DEEPSEEK_API_KEY", "test-key-for-fallback");
let current = vec![ModelRef::parse("anthropic:claude-opus-4-8")];
let pool = ModelPool::new(current.clone(), 1);
let mut found = false;
for _ in 0..30 {
let Some(m) = pool.next_auto_fallback(&current) else { break };
if m.provider == "deepseek" { found = true; break; }
}
std::env::remove_var("DEEPSEEK_API_KEY");
assert!(found, "a provider with a usable API key must be reachable as a fallback");
}
use super::*;
#[test]
fn parses_json_and_prose() {