feat: attack knowledge graph, layered memory, command rectification, FAIR dashboard

Backend
-------
- knowledge_graph.rs — the durable structure under attack_graph's per-run view:
  typed entities (asset/endpoint/weakness/technique/finding/account/credential/
  impact) joined by typed, weighted, provenance-carrying edges, accumulated
  across runs in .neurosploit/graph.json plus a per-run copy the report and web
  console can draw. Answers what a finding list can't: ranked attack paths, and
  the frontier of entities observed but never proven — where chaining should
  look next. Agents only sometimes fill chains_from, so progression is also
  inferred between adjacent kill-chain stages; those edges are marked inferred,
  weighted lower, and drawn dashed, because presenting a hypothesis as evidence
  is the graph lying about itself. Secrets stay in the vault, never the graph.

- memory.rs — four tiers scoped by lifetime, not importance: working (one run),
  engagement (one target), technique (one agent/CWE), reusable (generalized).
  Promotion is evidence-gated and needs independent evidence at each step: a
  claim repeated within a run becomes engagement knowledge; one confirmed
  across runs becomes technique knowledge; one that held on two DIFFERENT
  targets is generalized into a reusable lesson with host-specific tokens
  stripped. Nothing is promoted on a single observation, which is exactly what
  a hallucination looks like. Recall is scored (overlap × past success ×
  recency) and injected into recon/exploit prompts as leads to verify. Recalled
  memos are credited only when the run they informed actually found something.

- rectify.rs — a mistyped command cost a full round trip through /help, at the
  worst possible moment during a live run. Accepted-as-typed wins over
  everything (so the /url alias is never "corrected" to /ua), then unique
  prefix, then Damerau-Levenshtein with a length-scaled budget, and a tie is
  reported rather than resolved. Arguments too: a bare host gets its scheme, an
  out-of-range count is clamped with a note instead of silently reverting, a
  near-miss model id is matched against the live catalog.

- pool.rs — when every configured model is exhausted or its token is dead, try
  whatever else this machine can actually reach (an installed CLI subscription,
  or a provider whose key is in the environment) before parking. A run that
  stops on a box with three other usable backends stopped for no reason.

- repl.rs — /memory, /forget, /graph; a recovered run resumes by itself where
  nobody is watching (piped stdin — the web console — or NEUROSPLOIT_AUTO_RESUME),
  since a `/continue` prompt there waits forever.

Web
---
- Attack path: the stage list was seven hardcoded values, so findings the
  harness staged outside it were silently dropped — 5 of 27 on a real run.
  Rewritten against the harness's own stage list with unknown stages kept,
  two-line labels (every node used to read "SQL Injection Authent…"), stage
  column headers, pan/zoom/fit, path highlighting, severity filter, and the
  run's graph.json used when present.
- Dashboard: coverage, findings by severity, top weaknesses, and annualized
  loss exposure via FAIR — frequency from exploitability × validation
  confidence, magnitude from assumptions shown on screen and editable, reported
  as a range. The posture score saturates instead of subtracting, so it keeps
  discriminating past the first critical.
- Run history groups into one folder per target with a filter, instead of one
  flat list that grows forever.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BvdGy9XtVWSdXDTa3FFLJv
This commit is contained in:
CyberSecurityUPandClaude Opus 5 committed 2026-09-07 15:35:36 -03:00
1 parent 0ef0ce8d94
commit 9d83cb6e30
13 files changed
+3163 -119

No files matched your search

+60
View File
@@ -391,6 +391,61 @@ async function listRuns() {
return runs;
}
/// Flat aggregate over every run for the dashboard.
///
/// Returns per-finding tuples rather than a computed risk number: the FAIR
/// estimate depends on assumptions (contact frequency, loss magnitude per
/// severity) that belong to the operator, not to this server, so the browser
/// computes it from parameters the operator can see and change.
async function stats() {
let ids = [];
try {
ids = (await fsp.readdir(RUNS_DIR)).filter((d) => d.startsWith('ns-'));
} catch {
return { runs: [], findings: [], generated: Date.now() };
}
const runs = [];
const findings = [];
await Promise.all(ids.map(async (id) => {
const dir = path.join(RUNS_DIR, id);
const [status, fs_] = await Promise.all([
readJsonSafe(path.join(dir, 'status.json'), {}),
readJsonSafe(path.join(dir, 'findings.json'), []),
]);
const meta = await readJsonSafe(path.join(dir, 'meta.json'), {});
const tsMatch = id.match(/^ns-(\d+)-/);
const ts = status.ts || (tsMatch ? Number(tsMatch[1]) : 0);
const target = status.target || meta.target || id.replace(/^ns-\d+-/, '');
runs.push({
id,
ts,
name: engagementNames.get(id) || '',
target,
state: status.state || 'unknown',
agentsRan: status.agents_ran || 0,
findings: fs_.length,
});
for (const f of fs_) {
findings.push({
runId: id,
target,
ts,
severity: f.severity || 'Info',
cwe: f.cwe || '',
owasp: f.owasp || '',
stage: f.stage || '',
agent: f.agent || '',
title: f.title || '',
exploitability: f.exploitability || '',
confidence: typeof f.confidence === 'number' ? f.confidence : 0,
reviewStatus: f.review_status || '',
});
}
}));
runs.sort((a, b) => b.ts - a.ts);
return { runs, findings, generated: Date.now() };
}
async function runDetail(id) {
const dir = safeRunDir(id);
if (!dir) return null;
@@ -962,6 +1017,11 @@ const server = http.createServer(async (req, res) => {
return;
}
// ---- aggregate stats for the dashboard ----
if (req.method === 'GET' && p === '/api/stats') {
return sendJson(res, 200, await stats());
}
if (req.method === 'GET' && p === '/api/meta') {
return sendJson(res, 200, { version: '4.0.0', binary: BIN, root: ROOT });
}