Every finding in the Arena engagement carried votes reading 1/1 while the run
was configured for a 3-model vote. The panel was candidates.take(n), so a run
with one configured model produced a "multi-model adversarial validation" that
was one model agreeing with itself — the single most important thing the
engagement revealed about the harness.
Filling the panel by repeating that model would not fix it: its errors are
correlated with themselves, and three confident repetitions of one mistake are
indistinguishable from a consensus. Anthropic checking Anthropic is not
independent; a second vendor is.
build_panel() now takes at most one model per provider from the configured
candidates, then fills any shortfall from backends this machine can actually
reach (an installed subscription CLI, or a provider whose API key is in the
environment). When nothing else is available the panel stays small and the
yes/total the caller prints tells the truth about it, rather than being padded
to look like a quorum.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Backend
-------
- knowledge_graph.rs — the durable structure under attack_graph's per-run view:
typed entities (asset/endpoint/weakness/technique/finding/account/credential/
impact) joined by typed, weighted, provenance-carrying edges, accumulated
across runs in .neurosploit/graph.json plus a per-run copy the report and web
console can draw. Answers what a finding list can't: ranked attack paths, and
the frontier of entities observed but never proven — where chaining should
look next. Agents only sometimes fill chains_from, so progression is also
inferred between adjacent kill-chain stages; those edges are marked inferred,
weighted lower, and drawn dashed, because presenting a hypothesis as evidence
is the graph lying about itself. Secrets stay in the vault, never the graph.
- memory.rs — four tiers scoped by lifetime, not importance: working (one run),
engagement (one target), technique (one agent/CWE), reusable (generalized).
Promotion is evidence-gated and needs independent evidence at each step: a
claim repeated within a run becomes engagement knowledge; one confirmed
across runs becomes technique knowledge; one that held on two DIFFERENT
targets is generalized into a reusable lesson with host-specific tokens
stripped. Nothing is promoted on a single observation, which is exactly what
a hallucination looks like. Recall is scored (overlap × past success ×
recency) and injected into recon/exploit prompts as leads to verify. Recalled
memos are credited only when the run they informed actually found something.
- rectify.rs — a mistyped command cost a full round trip through /help, at the
worst possible moment during a live run. Accepted-as-typed wins over
everything (so the /url alias is never "corrected" to /ua), then unique
prefix, then Damerau-Levenshtein with a length-scaled budget, and a tie is
reported rather than resolved. Arguments too: a bare host gets its scheme, an
out-of-range count is clamped with a note instead of silently reverting, a
near-miss model id is matched against the live catalog.
- pool.rs — when every configured model is exhausted or its token is dead, try
whatever else this machine can actually reach (an installed CLI subscription,
or a provider whose key is in the environment) before parking. A run that
stops on a box with three other usable backends stopped for no reason.
- repl.rs — /memory, /forget, /graph; a recovered run resumes by itself where
nobody is watching (piped stdin — the web console — or NEUROSPLOIT_AUTO_RESUME),
since a `/continue` prompt there waits forever.
Web
---
- Attack path: the stage list was seven hardcoded values, so findings the
harness staged outside it were silently dropped — 5 of 27 on a real run.
Rewritten against the harness's own stage list with unknown stages kept,
two-line labels (every node used to read "SQL Injection Authent…"), stage
column headers, pan/zoom/fit, path highlighting, severity filter, and the
run's graph.json used when present.
- Dashboard: coverage, findings by severity, top weaknesses, and annualized
loss exposure via FAIR — frequency from exploitability × validation
confidence, magnitude from assumptions shown on screen and editable, reported
as a range. The posture score saturates instead of subtracting, so it keeps
discriminating past the first critical.
- Run history groups into one folder per target with a filter, instead of one
flat list that grows forever.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BvdGy9XtVWSdXDTa3FFLJv
Version bump 3.6.5 -> 3.6.6.
Local & uncensored models
- New `llamacpp:` provider (llama-server, OpenAI-compatible, localhost:8080,
no API key, CPU-only or GPU-offloaded). Override via LLAMACPP_BASE_URL;
model name is the loaded gguf (pass-through). 15 -> 16 providers.
- README: local/uncensored highlight, provider table + key-less note, badges.
Quality
- clippy clean under `-D warnings`: clamp(), sort_by_key(Reverse), struct-literal
init, too_many_arguments allows, scoped await_holding_lock on the REPL blocking
fallback (guard intentionally held across run().await), plus clippy --fix set.
CI
- examples/github-actions/ci.yml: cargo build/test/clippy -D warnings for the
neurosploit-rs workspace (template, kept out of .github/workflows).
Claude-Session: https://claude.ai/code/session_01QDses7zTSa9YF7pPRjphvh
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
- Robust verdict parsing (pool::parse_verdict): whitespace-insensitive, checks
explicit rejection first, counts only explicit confirmations; ambiguous →
Unclear (not confirmed). Replaces the fragile exact-JSON / loose "yes" match.
- Severity-aware quorum (pool::quorum_confirmed): High/Critical now need ≥2
validators AND ≥2/3 agreement (a single vote can no longer confirm a
Critical); lower severities need a strict majority (>half, was ≥half). Single-
model panels fall back to majority so they aren't nuked.
- Adversarial refute pass (REFUTE_SYS): every confirmed High/Critical is
re-examined by a skeptical panel that assumes false-positive; findings that
can't withstand a majority of skeptics are dropped. Survives on infra failure.
- Strengthened VOTE_SYS with an explicit false-positive checklist (reflected-not-
executed, version/banner guesses, self-XSS, error-as-injection, thin evidence,
inflated severity); validator query now also includes impact.
- Unit tests for parse_verdict + quorum_confirmed.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Token/quota exhaustion no longer silently drops agents. When every candidate
model is rate-limited / out of quota, the run PARKS (keeping all state) and
prints "⏸ token/quota exhausted … PAUSED". The user can:
- wait for renewal and /continue (retry same model), or
- /model <provider:model> (or the /model selector) then /continue to switch.
Implemented via ModelPool: is_exhaustion() detection, park_exhausted() that
awaits a resume Notify, and a fallback-model slot tried first on retry. /model
queues the chosen models into a paused run's fallback so a plain /continue
resumes on them.
Findings now survive a crash/quit: each finding is checkpointed live to
.neurosploit/active_run.json; on next launch an interrupted run is recovered
into /runs (a raw report is materialized) so /results, /finding and /report
keep working.
/stop now actually halts immediately on raw/discard: one() races the in-flight
model call against the hard-cancel flag, so the CLI child (kill_on_drop) is
terminated at once instead of finishing its whole command sequence. The
validate path still soft-stops (lets validation run).
Docs: TUTORIAL documents the 3-way /stop, crash recovery and pause/continue;
/help lists /continue and the new behaviors.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
- Streamed Claude events now tagged with the agent label (@name) so every
command/tool/file is attributable to the agent that ran it.
- Token/cost telemetry: parse usage from the stream-json result event; feed shows
per-call in/out/cost and a running total in the run summary.
- Ctrl-C during a run no longer hard-kills: it cancels cooperatively (no new
agents launch, in-flight bounded), then asks "generate report from partial
results? [Y/n]" — discard removes the run dir. Second Ctrl-C aborts.
- pool: cancel handle + is_cancelled; one()/complete_routed/chat_cli carry a label.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Harness:
- ModelPool gains a progress channel (set_progress); chat_cli forwards it.
- New chat_claude_stream: drives Claude Code with --output-format stream-json and
parses the event stream live — assistant text, and tool_use blocks categorized
into tagged events (exec/danger command, read/edit file, net request/browser,
grep/glob tool). 900s bound; clear error surfacing.
- Wired set_progress into run / whitebox / greybox.
REPL renderer (render_line):
- Tagged events render as the conversation feed: tool/command/network as compact
CARDS (tool-runner visual), files/edits/AI text/states as iconized lines.
- Clear "what the AI is doing" states: reconning, planning, testing, validating,
chaining, report, complete — plus a ⚠ DANGEROUS marker for risky commands.
- Untagged harness lines mapped to the same state vocabulary.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
- Task-based model ROUTER (recon/select prefer a fast model; exploit prefers primary; validate uses a different model than the finder)
- ReAct doctrine injected into exploit prompts (Thought→Action→Observation, token-efficient)
- Dedup: unique agents per run + findings deduped by CWE/endpoint/title (highest confidence kept)
- Token economy: recon blob capped for selector + per-agent context
- Configurable MCP: merge user mcp.servers.json into the pipeline's .mcp.json
- +54 white-box/code-analysis agents (NoSQLi, LDAP/XPath, JWT-none, Java/.NET/PHP/Go/Node/Python
specifics, SSTI, ReDoS, deserialization, etc.) → 303 agents total (78 code)
- Credits: Joas A Santos & Red Team Leaders (CLI banner, interactive header, HTML+Typst report)
- README: GitHub stars/forks badges, 60-second quick start, full API config steps, intuitive layout
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The 'recon failed (claude subscription CLI failed: )' was a transient CLI failure
(rate limit / cold start) reported with a blank message and no retry.
- chat_cli: on non-zero exit, surface exit code + stdout (CLI writes the real
reason there, not stderr); treat empty output as an error
- pool.one(): retry up to 3x with backoff for transient failures (both
subscription and API paths)
- with_auth: cap concurrency to 3 on the subscription path — spawning many
parallel CLI processes itself trips provider rate limits
Verified: live subscription run recovers and completes recon → select → exploit
→ vote → artifacts.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>