mirror of
https://github.com/CyberSecurityUP/NeuroSploit.git
synced 2026-09-30 13:09:36 +02:00
feat: attack knowledge graph, layered memory, command rectification, FAIR dashboard
Backend ------- - knowledge_graph.rs — the durable structure under attack_graph's per-run view: typed entities (asset/endpoint/weakness/technique/finding/account/credential/ impact) joined by typed, weighted, provenance-carrying edges, accumulated across runs in .neurosploit/graph.json plus a per-run copy the report and web console can draw. Answers what a finding list can't: ranked attack paths, and the frontier of entities observed but never proven — where chaining should look next. Agents only sometimes fill chains_from, so progression is also inferred between adjacent kill-chain stages; those edges are marked inferred, weighted lower, and drawn dashed, because presenting a hypothesis as evidence is the graph lying about itself. Secrets stay in the vault, never the graph. - memory.rs — four tiers scoped by lifetime, not importance: working (one run), engagement (one target), technique (one agent/CWE), reusable (generalized). Promotion is evidence-gated and needs independent evidence at each step: a claim repeated within a run becomes engagement knowledge; one confirmed across runs becomes technique knowledge; one that held on two DIFFERENT targets is generalized into a reusable lesson with host-specific tokens stripped. Nothing is promoted on a single observation, which is exactly what a hallucination looks like. Recall is scored (overlap × past success × recency) and injected into recon/exploit prompts as leads to verify. Recalled memos are credited only when the run they informed actually found something. - rectify.rs — a mistyped command cost a full round trip through /help, at the worst possible moment during a live run. Accepted-as-typed wins over everything (so the /url alias is never "corrected" to /ua), then unique prefix, then Damerau-Levenshtein with a length-scaled budget, and a tie is reported rather than resolved. Arguments too: a bare host gets its scheme, an out-of-range count is clamped with a note instead of silently reverting, a near-miss model id is matched against the live catalog. - pool.rs — when every configured model is exhausted or its token is dead, try whatever else this machine can actually reach (an installed CLI subscription, or a provider whose key is in the environment) before parking. A run that stops on a box with three other usable backends stopped for no reason. - repl.rs — /memory, /forget, /graph; a recovered run resumes by itself where nobody is watching (piped stdin — the web console — or NEUROSPLOIT_AUTO_RESUME), since a `/continue` prompt there waits forever. Web --- - Attack path: the stage list was seven hardcoded values, so findings the harness staged outside it were silently dropped — 5 of 27 on a real run. Rewritten against the harness's own stage list with unknown stages kept, two-line labels (every node used to read "SQL Injection Authent…"), stage column headers, pan/zoom/fit, path highlighting, severity filter, and the run's graph.json used when present. - Dashboard: coverage, findings by severity, top weaknesses, and annualized loss exposure via FAIR — frequency from exploitability × validation confidence, magnitude from assumptions shown on screen and editable, reported as a range. The posture score saturates instead of subtracting, so it keeps discriminating past the first critical. - Run history groups into one folder per target with a filter, instead of one flat list that grows forever. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BvdGy9XtVWSdXDTa3FFLJv
This commit is contained in:
1 parent
0ef0ce8d94
commit
9d83cb6e30
13 files changed
+3163
-119
No files matched your search
@@ -250,6 +250,13 @@ Zero npm dependencies (Node built-ins only).
|
||||
log tab grows a prompt box (`❭`) to send `/status`, `/stop`, `/continue`, or a plain-language
|
||||
instruction mid-run — same REPL described in [§6](TUTORIAL.md#6-the-interactive-repl). `host` /
|
||||
`aitest` / `skills` stay one-shot (their onboarding menu can't be scripted over piped stdin).
|
||||
- **Dashboard** — coverage (engagements, targets, agents run), findings by severity, most
|
||||
frequent weaknesses, and an **annualized loss exposure computed with FAIR**
|
||||
(Loss Event Frequency × Loss Magnitude): frequency from each finding's exploitability and
|
||||
validation confidence, magnitude from assumptions that are shown on screen and editable.
|
||||
Reported as a min / most-likely / max range, never a single number.
|
||||
- **Run history in folders** — runs group into one folder per target with a filter box, instead
|
||||
of one flat list that grows forever.
|
||||
- **Terminal dock** — `Ctrl+\`` (or `❭_` in the sidebar) opens a real terminal, xterm.js over an
|
||||
unstripped stdout stream, so the harness renders with its own colour and panels. Its header
|
||||
switches the terminal between a standalone REPL session and the engagement currently running,
|
||||
@@ -262,6 +269,42 @@ Zero npm dependencies (Node built-ins only).
|
||||
|
||||
Full API reference: **[web/API.md](web/API.md)** · quick start: **[web/README.md](web/README.md)**.
|
||||
|
||||
### Knowledge: memory + attack knowledge graph
|
||||
|
||||
Every model call starts with an empty context window, so without somewhere to put what a run
|
||||
learned, the harness re-derives the same stack, the same endpoints and the same dead ends every
|
||||
time. Two stores fix that, both under `.neurosploit/` in the project directory:
|
||||
|
||||
- **Layered memory** (`/memory`, `/forget`) — four tiers by scope, not importance: *working*
|
||||
(one run), *engagement* (one target), *technique* (one agent/CWE), *reusable* (generalized).
|
||||
Promotion is evidence-gated: a claim repeated within a run becomes engagement knowledge, one
|
||||
confirmed across runs becomes technique knowledge, and one that held on **two different
|
||||
targets** is generalized into a reusable lesson with the host-specific tokens stripped. Recall
|
||||
is scored (term overlap × past success × recency) and injected into recon/exploit prompts as
|
||||
leads to verify — never as assertions.
|
||||
- **Attack knowledge graph** (`/graph`, `graph.json`) — typed entities (asset, endpoint,
|
||||
weakness, technique, finding, account, credential, impact) joined by typed, weighted,
|
||||
provenance-carrying edges, accumulated across runs. It answers what a finding list can't:
|
||||
ranked attack paths, which endpoint accumulated the most weaknesses, and the *frontier* —
|
||||
entities observed but never proven, i.e. where chaining should look next. Chain edges the
|
||||
harness derived itself are marked `inferred` and drawn dashed in the web console. Secrets
|
||||
never enter the graph; they stay in the vault.
|
||||
|
||||
### Keeping a run going
|
||||
|
||||
- **Command rectification** — a mistyped command is corrected (`/staus` → `/status`), completed
|
||||
(`/onb` → `/onboard`), or reported as ambiguous, never guessed at. Arguments too: a bare host
|
||||
gets its scheme, an out-of-range count is clamped *with a note*, a near-miss model id is
|
||||
matched against the live catalog.
|
||||
- **Automatic backend fallback** — when every configured model is quota-exhausted or its token
|
||||
is dead, the pool switches to whatever else this machine can reach (an installed CLI
|
||||
subscription, or a provider whose API key is in the environment) and keeps going. It only
|
||||
parks the run when nothing at all is available.
|
||||
- **Resume where it stopped** — findings are checkpointed live, so an interrupted run is
|
||||
recovered on the next start and `/continue` carries them forward. Non-interactive sessions
|
||||
(the web console drives the REPL over a pipe) resume automatically, since no one is there to
|
||||
type it; set `NEUROSPLOIT_AUTO_RESUME=1` to get the same at a terminal.
|
||||
|
||||
---
|
||||
|
||||
## 🔌 Integrations (GitHub · GitLab · Jira)
|
||||
|
||||
Reference in new issue
Block a user