Backend
-------
- knowledge_graph.rs — the durable structure under attack_graph's per-run view:
typed entities (asset/endpoint/weakness/technique/finding/account/credential/
impact) joined by typed, weighted, provenance-carrying edges, accumulated
across runs in .neurosploit/graph.json plus a per-run copy the report and web
console can draw. Answers what a finding list can't: ranked attack paths, and
the frontier of entities observed but never proven — where chaining should
look next. Agents only sometimes fill chains_from, so progression is also
inferred between adjacent kill-chain stages; those edges are marked inferred,
weighted lower, and drawn dashed, because presenting a hypothesis as evidence
is the graph lying about itself. Secrets stay in the vault, never the graph.
- memory.rs — four tiers scoped by lifetime, not importance: working (one run),
engagement (one target), technique (one agent/CWE), reusable (generalized).
Promotion is evidence-gated and needs independent evidence at each step: a
claim repeated within a run becomes engagement knowledge; one confirmed
across runs becomes technique knowledge; one that held on two DIFFERENT
targets is generalized into a reusable lesson with host-specific tokens
stripped. Nothing is promoted on a single observation, which is exactly what
a hallucination looks like. Recall is scored (overlap × past success ×
recency) and injected into recon/exploit prompts as leads to verify. Recalled
memos are credited only when the run they informed actually found something.
- rectify.rs — a mistyped command cost a full round trip through /help, at the
worst possible moment during a live run. Accepted-as-typed wins over
everything (so the /url alias is never "corrected" to /ua), then unique
prefix, then Damerau-Levenshtein with a length-scaled budget, and a tie is
reported rather than resolved. Arguments too: a bare host gets its scheme, an
out-of-range count is clamped with a note instead of silently reverting, a
near-miss model id is matched against the live catalog.
- pool.rs — when every configured model is exhausted or its token is dead, try
whatever else this machine can actually reach (an installed CLI subscription,
or a provider whose key is in the environment) before parking. A run that
stops on a box with three other usable backends stopped for no reason.
- repl.rs — /memory, /forget, /graph; a recovered run resumes by itself where
nobody is watching (piped stdin — the web console — or NEUROSPLOIT_AUTO_RESUME),
since a `/continue` prompt there waits forever.
Web
---
- Attack path: the stage list was seven hardcoded values, so findings the
harness staged outside it were silently dropped — 5 of 27 on a real run.
Rewritten against the harness's own stage list with unknown stages kept,
two-line labels (every node used to read "SQL Injection Authent…"), stage
column headers, pan/zoom/fit, path highlighting, severity filter, and the
run's graph.json used when present.
- Dashboard: coverage, findings by severity, top weaknesses, and annualized
loss exposure via FAIR — frequency from exploitability × validation
confidence, magnitude from assumptions shown on screen and editable, reported
as a range. The posture score saturates instead of subtracting, so it keeps
discriminating past the first critical.
- Run history groups into one folder per target with a filter, instead of one
flat list that grows forever.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BvdGy9XtVWSdXDTa3FFLJv