Both bugs surfaced during a live black-box engagement.
1. extract_findings took the span from the first '[' to the last ']'. One
agent's reply opened with the prose line "[low] Antiforgery cookie missing
Secure flag" and put its real findings in a fenced ```json block further
down — so the span started inside prose, failed to parse, and every finding
that agent had proven was thrown away. Fenced blocks are now tried first
(last one wins: models narrate, then answer), with the span kept only as a
fallback and the trailing-comma salvage preserved.
2. "no findings" was reported as malformed JSON. The guard compared the raw
text to "[]", but models wrap the empty array in a fence, so every honest
negative result was logged as a parse failure — which teaches an operator to
ignore a warning that sometimes means a real one. reported_nothing() now
recognises a bare [], a fenced [], and {"findings": []}.
The second bug made the first one harder to see: the log was already full of
"malformed JSON" warnings for agents that had simply found nothing, so the one
warning that meant a genuine loss looked like more of the same.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
A PDF was only ever produced while a run was finishing. If `typst` was missing
at that moment — or the template improved afterwards — the operator had no way
to get one without re-running the whole engagement against the target.
report::rebuild() regenerates every artifact (md · json · html · pdf) from the
findings already on disk, exposed as `neurosploit rebuild <run-id|dir>` and as
POST /api/runs/:id/report with a "Generate report" button in the run view. The
endpoint shells out to the harness rather than reimplementing report generation
in JavaScript, so there is one implementation instead of two that drift, and it
says plainly when the PDF was skipped for want of `typst` instead of handing
back a link to a file that was never produced.
Also fixes write_all() to pass the run's pocs/ listing into the HTML report, so
a rebuilt report links the scripts each finding cites — the run-time path
already did this and the rebuild path silently did not.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The risk model, grants and hash-chained trail existed as modules nothing
called. Now every engagement runs under them.
Capability
- `neurosploit capability issue|verify` mints and inspects grants.
- `--capability-token` (global, so the REPL takes it too), `--in-scope`,
`--environment`, `--policy` on `run`; verification happens at the command
line, so an invalid grant fails with a readable message instead of halfway
through an engagement.
- The pipeline verifies before anything else and REFUSES to run on a token that
does not verify — proceeding would mean acting on an authorization nobody can
prove was issued. `effective_scope` then applies the grant as a ceiling.
- Web: an Authorization tab carrying the token, extra hosts, environment and
policy profile. The browser decodes the claims for display and says plainly
that it is not verifying them — a "valid" badge from a party without the key
would be the UI vouching for something it cannot check.
A hole the smoke test found: `/inscope evil.test` inside a session under a
grant WIDENED the scope past it — the one thing a capability token exists to
prevent. The run itself would still have been constrained (the pipeline
re-applies the grant), but `/policy` reported a boundary that was not real, and
a tool that misreports its own limits is worse than one with none. Scope
mutations now re-apply the ceiling and name what it refused. Session
authorization also arrives from argv rather than a `/`-command, because a
session that can widen its own grant is not constrained by one.
Audit
- One hash-chained record per action in `<run>/audit.jsonl`, in the specified
shape, covering engagement start/end, validator rejections, findings that
reach the report (with the hash of the evidence behind them) and findings
withheld for being out of scope.
- The run verifies its own chain at the end and says loudly if it is broken.
- `/audit [n]` tails the trail and verifies it; the web offers it as a download
next to the report, so "show me what the tool did" is a link.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Four pieces that together answer "may this action happen, under whose
authority, and can we prove afterwards what we did".
policy.rs — effective_risk per action, exactly as specified:
(action_risk + asset_criticality + protocol_risk + privilege_level
+ blast_radius) × environment_multiplier
Every term is named and kept on the result, so the number can be explained
rather than argued with. Three policies sit on it: SafetyPolicy (ceilings,
approval thresholds, hard prohibitions), ReasoningPolicy (baseline before
payload, bounded hypotheses, evidence before escalation, explicit stop
conditions) and ProofOfImpactPolicy (what a severity must carry before it may
be that severity).
OT/ICS/SCADA is treated as its own regime, not web testing on odd ports.
Industrial protocols authenticate nothing — a Modbus write is the protocol
working as intended, addressed to a device that may be holding a valve — and
scanners crash PLCs by sending unexpected data at line rate. So the OT profile
blocks writes, disruptive actions, fuzzing and exploit payloads outright, caps
the rate at ~1 req/s, and refuses the function codes that stop a CPU (Modbus
5/6/8/15/16/22/23/43, S7 start/stop, DNP3 restart/stop). Safety instrumented
systems are off limits in every profile.
A test caught a calibration error worth keeping: a plain READ of a critical PLC
scores 3.6 on this formula, so the obvious tight ceiling would have refused
exactly the observation OT findings come from. In an industrial environment it
is the KIND of action that is forbidden, not the arithmetic — the ceiling
catches extremes and the low approval threshold makes anything past trivial
observation a human's decision.
capability.rs — HMAC-signed grants: who authorized what, against which hosts,
in which environment, until when. The harness verifies the signature before
reading a single claim (a well-formed token from the wrong key must never get
to influence what the harness believes), refuses expired and not-yet-valid
tokens, and treats the grant as a CEILING: constrain() intersects it with local
configuration, so config can narrow authorization and never widen it. Tokens
carry no secrets — the payload is readable by anyone holding it.
audit.rs — one structured record per action, in the specified shape (timestamp,
agent, hypothesis, action, target, policy_decision, operator, tool, result,
evidence_hash, capability_token). Two things make it worth having: it is
hash-chained, so removing or editing an entry breaks every hash that follows
and verify() says which one; and it records REFUSALS, because a trail
containing only what happened cannot demonstrate restraint. Only the grant's
id is recorded, never the token — the trail gets shared.
Hard kill conditions end a run outright: target unresponsive after our traffic,
sustained 5xx, out-of-scope request, forbidden industrial function code, safety
system addressed, capability expired mid-run, repeated policy violations,
budget exhausted, operator stop. Failures BEFORE the target ever answered do
not count — nothing listening is not the same as knocked over. The OT switch
trips far sooner: a PLC missing two requests already warrants stopping.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
A finding is useful only if the reader can find the problem, see why it
matters, fix it, and reproduce it without trusting us. The report answered the
last one badly and the other three not at all: it printed a payload blob and an
evidence blob, and "payload: ' OR 1=1--" tells a developer nothing about WHERE
to look. A PoC script attached as a file is a black box unless you run it.
Findings are now rendered in the order a reader works through them — where the
problem is, what it means, how to fix it, then the proof — in the HTML report,
the Markdown, the Typst/PDF and the web console's finding modal.
The proof is numbered, pasteable steps: baseline request, attack request, how
to read the result, with the real URL and the real payload. They come from the
agent's `repro_steps` when it recorded them, and are derived from the
endpoint/payload/identity pair otherwise, so every finding carries something
runnable. The generated curl redacts Authorization/Cookie/API-key headers — a
report gets shared, and a live session cookie inside one is a new bug. A PoC
script is now offered as an extra artifact that automates the steps, never as
the proof itself.
Technical evidence is the measured difference, not a paraphrase: baseline vs
attack status, size, timing and delta; how many repeats reproduced it; the
controlled marker and whether a browser or a callback observed it; then each
recorded exchange with the headers that decide a class (Location, Set-Cookie,
Access-Control-*, X-Frame-Options, CSP, Retry-After) and a body excerpt.
Finding gains `location` (the parameter/field/flow step, not just the URL) and
`repro_steps`, and the agent contract now asks for them explicitly, along with
impact tied to this app's data and remediation that names the control rather
than saying "sanitise input".
The web console offers the run's PDF when Typst produced one — and only then,
since a dead download button is worse than none.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Validators decide from recorded artifacts, and the weakest link was who
recorded them: "the payload returned a 500" is still an agent's account of what
happened. Replay produces the part that matters most in practice —
reproducibility — by sending the request again through the harness's own
client, with the scope guard in front of it.
Three properties it is built around:
- every request passes ScopePolicy::check_request before a socket is opened, so
replay cannot be the thing that wanders off-scope while verifying a finding;
- it never mutates: a finding proven with DELETE is not re-proven by deleting
the record again, so non-idempotent verbs are refused and repeats of them are
refused outright;
- bodies are truncated at 96KB and SAY they were truncated — a silently clipped
body makes a length differential meaningless.
enrich() fills in repeats and re-measures a recorded baseline (comparing a
fresh attack against an hour-old baseline attributes ordinary drift to the
payload). It deliberately does NOT synthesize a baseline from an attack
request: removing "the payload" from an arbitrary URL is guesswork, and a
guessed baseline would silently decide the verdict.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Audited at 390×844, 844×390 (phone landscape), 768×1024, 1024×768, 1440×900
and ≥1600px. What was actually broken:
- **The run header collapsed.** `justify-content: space-between` let the two
action buttons take the whole row, so at 390px the title wrapped one
character per line ("Test AspNe / t") and the facts line wrapped one word per
line. It now stacks below 680px, with the title block sized to its content
instead of stretching (the desktop `flex: 1` was what left a 180px void above
the buttons in the stacked layout).
- **The terminal header overflowed** its dock at 390px (557px of content in a
390px box) — it wraps now, and the status text drops out on narrow screens
where the coloured dot already carries it.
- **`100vh` is wrong on mobile.** It measures the viewport without the
collapsing address bar, so the wizard footer and its CTA sit underneath it.
Switched to `dvh` with the `vh` line kept as the fallback.
- **The dock took 82% of a phone in landscape** at its fixed 320px. It now
tracks the viewport (`clamp(180px, 42dvh, 340px)`, tighter still under
500px of height).
- **The off-canvas drawer had no way out but the button that opened it.**
Added a scrim that closes it, Esc, and auto-close when a run is picked —
and it closes itself if the window grows past the breakpoint, which
otherwise left a scrim over a sidebar that was no longer a drawer.
- **The stepper scrolls horizontally on a phone**, so advancing to an
off-screen step looked like nothing happened; the active step is scrolled
into view.
Device-type rules rather than width alone: `pointer: coarse` gets 38-44px hit
targets and 16px inputs (under 16px, iOS zooms the page on focus and breaks the
layout the user is typing into); `prefers-reduced-motion` drops the drawer
slide and the progress animation, which are decoration.
Scrolling stays where it belongs — one scroller per pane (`.wizard-body`,
`.run-body`, `.dash-body`, `.sb-groups`, `.modal-body`, `.term-host`), wide
tables scroll inside `.table-wrap`, and the page itself never scrolls
horizontally at any tested size.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Six classes had deterministic rules; the rest of a run still rested on models
voting. These thirteen cover the classes that produce the most false positives
in AI-driven testing, and each one is written around what *disproves* the
claim, because that is the part a language model skips:
SSTI an expression evaluated server-side whose result was never
sent — the payload echoing its own "result" is rejected
XXE entity content or an OOB callback; a parser error mentioning
entities shows the DTD was read, not that anything resolved
Open redirect 3xx WITH a Location off-site; a rendered link is not a redirect
CORS reflected Origin PLUS credentials; ACAO:* without credentials
exposes only what an anonymous client could already read, and
ACAO:* WITH credentials is refused by browsers anyway
Cookie flags fully decidable from Set-Cookie + scheme
Clickjacking neither X-Frame-Options nor CSP frame-ancestors
Auth bypass protected content with NO credentials sent — a "bypass" whose
request still carried a cookie is rejected, as is a redirect
to login
JWT forged token accepted AND privileged content returned
Rate limiting >= 20 attempts with no 429/Retry-After; five attempts prove
nothing about a limit that was never reached
Session fix. the session id surviving login unchanged
Mass assign. a read-back proving the field persisted — a 200 on the write
means nothing, APIs accept and ignore extra fields routinely
CSRF a cross-origin state change read back; a SameSite session
cookie means a browser would never attach it cross-site
Exposure a real secret/listing signature the baseline lacked; a
soft-404 mirroring the baseline page is rejected
Exchange gains response and request headers, because several of these classes
are decided by a header (Location, Set-Cookie, Access-Control-Allow-*) and the
body alone is not evidence for them.
A test asserts no two validators claim the same CWE — ambiguous ownership would
make routing depend on registration order, which is how a class silently gets
the wrong rule.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Two gaps this closes, both found by reading what the code actually did.
Scope was never enforced
------------------------
`out_of_scope` was rendered into the prompt as "HARD CONSTRAINT — do NOT test…"
and nothing checked it. That is a request to a model, not a control: an agent
that decided a discovered subdomain was interesting, or that followed a
redirect off-target, was free to act and the operator found out by reading the
report.
scope.rs adds a guard in code:
- Hard scope: allowlist of hosts, *.wildcards, IPv4 CIDRs, URL prefixes, with
exclusions that always win. Defaults to the engagement's own target, so
discovery cannot widen authorization — finding a host is not permission to
attack it. An unconfigured policy is closed, not open.
- Soft scope: observe-only zones, destructive verbs (off by default), an
account-creation cap, a rate guard that warns rather than silently dropping
requests (a dropped request reads as "target unreachable"), and payload
classes refused even in scope because they damage the target instead of
demonstrating a bug.
- Enforced at the harness's own chokepoint (probe) and as a post-run audit:
findings proven against an unauthorized host are withheld from the report and
written to out-of-scope-findings.json as an incident to disclose, because
shipping one would launder the mistake.
- REPL: /inscope, /observe, /guardrail, /policy; /scope-out now promotes
host-shaped entries into enforced rules immediately, and says plainly when an
entry is prose the guard cannot enforce.
Validation was models checking models
-------------------------------------
N-model voting plus an adversarial refute pass share the failure mode of the
thing they check — agreement is not evidence, and a confident hallucination
survives a vote by being confident. grounding.rs helps but matches keywords
("http/", "status", "alert(") and cannot tell a real response from a plausible
transcript of one.
validation.rs asks a different question — does the recorded evidence
demonstrate THIS class? — with per-CWE rules and no model in the loop:
SQLi baseline/attack difference that reproduces >= 2x
XSS a browser executed a harness-chosen marker; reflection is not proof
IDOR identity B reads A's resource AND the body matches (a 200 returning a
login page is rejected, which is the classic false positive)
SSRF controlled callback or canary retrieval
LFI controlled marker or a file signature the baseline lacked
RCE a unique nonce in output/callback; reflected input is rejected
Absent evidence is never a pass, and a class with no rule is never
auto-confirmed. NEUROSPLOIT_VALIDATION=advisory (default) rejects
contradictions without demoting voted findings for missing artifacts;
enforcing makes the verdict the status. The evidence contract is injected into
exploit prompts so agents collect the artifacts while they still hold the
target.
Finding gains evidence_data so agents can emit structured artifacts alongside
the finding JSON.
Two bugs the tests caught while writing this: the scope guard treated a SAST
`src/auth.rs:42` endpoint as a host and quarantined valid source findings, and
two canaries minted in the same clock tick came out identical — a marker that
repeats would let a stale token vouch for a new finding.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Backend
-------
- knowledge_graph.rs — the durable structure under attack_graph's per-run view:
typed entities (asset/endpoint/weakness/technique/finding/account/credential/
impact) joined by typed, weighted, provenance-carrying edges, accumulated
across runs in .neurosploit/graph.json plus a per-run copy the report and web
console can draw. Answers what a finding list can't: ranked attack paths, and
the frontier of entities observed but never proven — where chaining should
look next. Agents only sometimes fill chains_from, so progression is also
inferred between adjacent kill-chain stages; those edges are marked inferred,
weighted lower, and drawn dashed, because presenting a hypothesis as evidence
is the graph lying about itself. Secrets stay in the vault, never the graph.
- memory.rs — four tiers scoped by lifetime, not importance: working (one run),
engagement (one target), technique (one agent/CWE), reusable (generalized).
Promotion is evidence-gated and needs independent evidence at each step: a
claim repeated within a run becomes engagement knowledge; one confirmed
across runs becomes technique knowledge; one that held on two DIFFERENT
targets is generalized into a reusable lesson with host-specific tokens
stripped. Nothing is promoted on a single observation, which is exactly what
a hallucination looks like. Recall is scored (overlap × past success ×
recency) and injected into recon/exploit prompts as leads to verify. Recalled
memos are credited only when the run they informed actually found something.
- rectify.rs — a mistyped command cost a full round trip through /help, at the
worst possible moment during a live run. Accepted-as-typed wins over
everything (so the /url alias is never "corrected" to /ua), then unique
prefix, then Damerau-Levenshtein with a length-scaled budget, and a tie is
reported rather than resolved. Arguments too: a bare host gets its scheme, an
out-of-range count is clamped with a note instead of silently reverting, a
near-miss model id is matched against the live catalog.
- pool.rs — when every configured model is exhausted or its token is dead, try
whatever else this machine can actually reach (an installed CLI subscription,
or a provider whose key is in the environment) before parking. A run that
stops on a box with three other usable backends stopped for no reason.
- repl.rs — /memory, /forget, /graph; a recovered run resumes by itself where
nobody is watching (piped stdin — the web console — or NEUROSPLOIT_AUTO_RESUME),
since a `/continue` prompt there waits forever.
Web
---
- Attack path: the stage list was seven hardcoded values, so findings the
harness staged outside it were silently dropped — 5 of 27 on a real run.
Rewritten against the harness's own stage list with unknown stages kept,
two-line labels (every node used to read "SQL Injection Authent…"), stage
column headers, pan/zoom/fit, path highlighting, severity filter, and the
run's graph.json used when present.
- Dashboard: coverage, findings by severity, top weaknesses, and annualized
loss exposure via FAIR — frequency from exploitability × validation
confidence, magnitude from assumptions shown on screen and editable, reported
as a range. The posture score saturates instead of subtracting, so it keeps
discriminating past the first critical.
- Run history groups into one folder per target with a filter, instead of one
flat list that grows forever.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BvdGy9XtVWSdXDTa3FFLJv
Replaces the floating REPL drawer with a docked terminal, and fixes the
usability problems a screenshot audit of the console turned up.
Terminal (the reason for the change):
- The drawer rendered the harness into a <div>, so the server had to strip
ANSI before sending it: colour, the box-drawn /status panel and the banner
all arrived flattened, and long lines rewrapped mid-glyph. The stream is now
sent verbatim and rendered by xterm.js (vendored, nothing fetched at
runtime), decoded with a streaming UTF-8 decoder so a multi-byte character
split across two reads survives.
- The drawer floated bottom-right, directly over "Next →" and "Start
Exploitation" — the wizard's primary buttons. The dock is a flex child of
.main, so opening it shortens the view instead of covering it. Drag its top
edge to resize; the height is remembered.
- The child is spawned over a pipe, not a PTY, so it never echoes: line
editing is local — echo, ←/→, Home/End, history, Tab completion over the
slash commands, Ctrl+C/L/U/K/A/E. Ctrl-C is delivered as SIGINT by the
server, since a raw 0x03 byte over a pipe interrupts nothing.
- A target picker switches the terminal between a standalone REPL session and
the engagement currently running, so mid-run instructions go to the same
process doing the testing.
QA fixes:
- Findings tables sorted by severity (a LOW above a CRITICAL made a 27-row
result unreadable), with sortable headers, a severity summary that doubles
as a filter, a text filter, a sticky header, and horizontal scroll confined
to the table instead of the whole page.
- alert()/prompt() replaced by inline field errors, a custom-lead modal and
toasts — a modal alert hid the very field it was complaining about.
- Lead categories start collapsed (412 leads over ~30 categories); search
auto-expands what it matches and shows per-category hit counts.
- Sidebar rows truncate inside the rail (a long target URL used to spill past
its border), and carry a worst-severity dot, finding count and age.
- Past-run header shows when it ran, how many agents ran, the recon asset,
PoC count and run id — two runs of one target were indistinguishable.
- Off-canvas sidebar below 768px had no way to be opened; added the toggle.
- Long evidence values (cookies, tokens) now wrap instead of running under
the finding modal's edge.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BvdGy9XtVWSdXDTa3FFLJv
- README: web console section notes the wizard now scripts a real REPL
session for run/whitebox/greybey so mid-run prompts work
- TUTORIAL.md: new §8 Web console (wizard steps, REPL script, live view,
links to web/API.md and web/README.md); renumbered §9-17 and §9.1-9.5
- web/README.md: bullet on the REPL-backed run + send-prompt box
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0166e1EFP2JrifYsj5bVugut
Root cause of "can't send prompts while a run streams": /api/exploit
spawned a plain `neurosploit run ...` subprocess, and that CLI path
(run_mode() in main.rs) never reads stdin - it only waits on the task or
Ctrl-C. The ONLY thing in the harness that keeps accepting input while an
engagement streams is the interactive REPL's background-run loop. So:
- New startJobViaRepl(): for mode run/whitebox/greybox, spawns a bare
`neurosploit` REPL session and scripts it via stdin (/target or /repo,
/model, /sub, /mcp, /votes, /chain, /recon, /focus, /objective,
/scope-out, /creds, /only <agents> or /only clear, then /run) instead
of building CLI args. Same underlying pipeline, same tagged output
lines, so all existing parsing (findings/phase/progress/runId) works
unchanged. host/aitest/skills modes stay on the old one-shot
startJob() - they need onboarding's scope picker, an interactive
arrow-key menu that silently skips itself over a piped stdin, so they
can't be scripted this way.
- New POST /api/exploit/:id/input writes a line to the session's stdin -
natural language, /status, /continue, anything the REPL accepts - and
the live run view grows a "send prompt" box (in the Activity log tab)
for it, shown only when the job reports interactive: true.
- Stop, for an interactive job, now sends the REPL's own graceful
'/stop\n1\n' (validate what's found, then report) instead of SIGINT -
the REPL's own input loop has no signal handler, so SIGINT there would
just kill the process outright and skip the report step. Non-
interactive jobs still get SIGINT (run_mode() does catch that).
- 'done' can no longer be process-exit only: an interactive session stays
open after the engagement finishes (for /report, /continue, another
/run), so ingestLine() now also flags done from the same "phase
complete" content signal it already used for the phase field.
Verified end-to-end: started an interactive job, confirmed
`interactive: true` and a captured runId, sent /status and /agents mid-
and post-run over the new /input endpoint (both accepted, session stayed
alive and responsive after completion), and confirmed a non-interactive
run is unaffected.
Also: the missing "Activity log" tab a screenshot showed for a "running"
engagement was the sidebar's detail-view fallback (2 tabs, no log) for a
run whose Job object no longer exists in server memory - it happens when
the Node process gets restarted while a spawned neurosploit child is
still alive underneath it (an orphan from testing across many redeploys
this session, not a code bug); the live view itself always had the tab.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0129WdYHccPsH27k5GGuwijd
/only <agent,...> (app/src/repl.rs): the REPL had no way to pin an exact
agent set the way the CLI's --only flag does - Session gained a `pinned`
field, wired into RunConfig in both start_background() (the live
background-run path) and the blocking run() fallback. Needed so the web
console's exploitation jobs can drive a real interactive REPL session
(for live input while a run streams) without losing lead-pinning, which
only existed as a CLI flag until now. Also usable directly from a
terminal REPL session.
HTML report (crates/harness/src/report.rs, html()): rebuilt to match the
Typst PDF template's design (templates/report.typ) instead of its own
inconsistent styling - violet brand accent, an asset table, a 5-box
executive-summary grid (all severities, zero-count included, matching
Typst's grid exactly), a Vulnerability Summary table, and severity-
left-bordered finding cards with a compact field grid (Criticality /
Status / OWASP-CWE / Confidence / Location / Agent / Auth context) before
Description-Impact / Proof of Concept / Evidence / Remediation - same
field order and labels as the Typst template. Dropped the Mermaid
attack-path/kill-chain section entirely (the web console's live
Generative Attack Path Chaining graph covers that now, interactively).
Also tidied two pre-existing formatting quirks while in there: OWASP/CWE
left a dangling " · " when CWE was empty, and the confidence cell said
"<votes-string> votes" even when the votes field already contained a
compound descriptor like "1/1 · receipt_missing".
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0129WdYHccPsH27k5GGuwijd
- Agent library table was stale (196/12/78/17 = 303 total, missing the
infra/chains/ai categories entirely). Corrected to the real counts
(245/78/30/34/23/13/12 = 435), matching the "MD Agents-435" badge and
neurosploit agents output.
- Provider badge said 16; the harness actually ships 18 (litellm and azure
were missing from the README's provider table and the API-key export
block). Added both.
- Web console section was a 3-line stub written before most of the feature
was built. Expanded to cover the 5-step wizard, the lead board's bulk
select/category toggles, custom-lead-generates-a-real-agent, the live
run view's Generative Attack Path Chaining graph and finding/PoC detail,
the Auth & Keys menu, and F5 persistence - with links to web/API.md and
web/README.md.
- Select all / Clear all buttons in the Leads step toolbar - respects the
current search filter, so filtering to "sql" then Select all only pins
those, not all 412 leads. The per-category master switch (already
select/deselect-all for that category, indeterminate when partial) was
the only bulk control before; this adds the "everything" case.
- '+ Custom lead' now generates an ACTUAL specialist-agent markdown file
(agents_md/vulns/custom_<slug>.md, same format every other agent uses)
via the claude CLI on the operator's Anthropic subscription
(claude-opus-4-8 by default - matches the harness's own default model),
instead of folding free text into --focus. The new lead is immediately
selectable and pinnable via --only like any other agent; verified the
Rust harness's own agent loader picks it up (agent count went 435 -> 436,
neurosploit agents confirmed it).
Two things found and fixed while wiring this up:
- the skip-permissions flag gave the model file/bash tool access, which
made it try to write the file itself and narrate doing so instead of
just returning text. Dropped the flag (pure text completion needs no
tools) and told it explicitly not to use any.
- Even so, defensively strip anything before the first '# ' heading
before saving, in case a model still prepends commentary.
Falls back to the old free-text-focus behavior if generation fails
(claude not installed/logged in, malformed output, timeout) so the
operator's intent isn't lost.
- New "Custom Leads" category, shown first, so generated leads have a
visible home instead of landing in the catch-all "Other" bucket.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0129WdYHccPsH27k5GGuwijd
impact and business_impact often carry identical text (both ending with the
reporter's 'Identified and validated by NeuroSploit...' footer), and the
modal concatenated them verbatim — the boilerplate line rendered twice, and
whenever the two fields matched, so did the whole paragraph.
- Strip the attribution sentence out of impact/business_impact/remediation/
evidence wherever it appears; surface it once, at the bottom of the modal,
instead of embedded per field.
- Skip business_impact entirely when it's identical to impact (the common
case) instead of printing the same paragraph twice.
- Split rendering into codeBlock() (endpoint/payload — monospace, looks like
what it is: a request/curl) and proseBlock() (description/impact/
remediation — a readable paragraph, not a code box).
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0129WdYHccPsH27k5GGuwijd
Node/edge colors now use the same --sev-*-fg / --surface / --text / --border
CSS custom properties as the rest of the console (set via inline style=
attributes, since SVG presentation attributes don't resolve var()) — the
graph reads correctly in light mode instead of always rendering as a
dark canvas.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0129WdYHccPsH27k5GGuwijd
The 'Generative Attack Path Chaining' tab previously showed kill-chain
stages as stacked cards in columns — with 1 finding (the common case
early in a run) it looked like an empty list, nothing like an attack
graph.
Rewritten as an inline SVG node/edge graph on a fixed-dark canvas
(matches attack-graph tools like NodeZero regardless of the app's own
light/dark theme — bright severity colors read better against near-black):
- Root node = the target, always present.
- One node per confirmed finding, positioned in its kill-chain-stage
column (falls back to a single flat column when no finding has a
stage yet).
- Edges: from the finding's chains_from parent when the harness set one,
else fanned directly from root — never invents a specific relationship
that doesn't exist in the data.
- Per-node icon inferred from title/evidence/cwe/stage (key/shield/
person/host/db/impact), severity-colored border + corner tick.
- Nodes are clickable — opens the same finding detail modal as the
findings table (PoC included).
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0129WdYHccPsH27k5GGuwijd
Real front-end bugs found and fixed:
- [hidden] never worked on any element whose class also sets 'display'
(every .btn, .chip, ...): the browser's built-in '[hidden]{display:none}'
rule and an author rule of equal specificity tie, and the later one in the
cascade wins — so 'Next' stayed visible on the Review step alongside
'Start Exploitation', and 'Open report'/'Stop' rendered during 'starting'.
Fixed with a single global '[hidden]{display:none!important}' override.
- Progress bar was functionally correct but easy to miss (thin, 0%-width,
low-contrast track) and gave no feedback while the agent count is still
unknown (recon phase). Added a border for visibility and an indeterminate
sliding-segment state for the 'agents: ?' window.
- A live run watched in the browser was lost on F5 (jumped back to the
wizard) even though the job keeps running server-side. The active job id
now persists in localStorage; on load the app reconnects the SSE stream
(the server replays its full event buffer) instead of losing the view.
New:
- Findings are now clickable — a detail modal shows every Finding field
(CWE/CVSS/OWASP/MITRE/stage/exploitability/confidence/votes/review status/
auth context/account/agent), endpoint+payload, evidence, impact, business
impact, remediation, and chains_from — in both the live run and past-run
detail views.
- PoC surfacing: the finding modal looks up any script the run wrote to
pocs/ that's cited in the finding's evidence (per the harness's own
doctrine — see pipeline.rs change below), fetches and previews it inline,
with a link to open the raw file. Live runs poll for new PoC files every
5s once the run id is known.
- Pinned-leads confirmation: the live run header now states plainly how
many leads were pinned (and their names) or that selection is auto
(recon-driven) — this was previously buried in the scrolling activity log
behind the harness's unconditional 'Loaded 435 agents' library-size line,
which describes the full agent library, not what will actually run.
Harness doctrine (crates/harness/src/pipeline.rs, pocs_line()):
PoC-writing for black-box findings was previously conditioned on 'when an
issue needs a custom multi-step exploit/script' — vague enough that a
straightforward finding (single-request XSS/SQLi/IDOR) often got no PoC
file at all. Now required for every confirmed Medium+ finding, one
standalone .py/.sh script per finding, and explicit about citing the exact
file name in the finding's evidence field (which is what the web UI now
matches on to link a PoC to its finding).
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0129WdYHccPsH27k5GGuwijd
Unchecking one agent in an 8-agent category (7/8 left on) rendered the
category header switch fully unchecked — visually indistinguishable from
'category disabled', even though 7 of 8 agents were still on. The
checkbox's checked state only had two positions; a partial selection had
nowhere to render but off.
Fix: set the master switch's .indeterminate property when 0 < selected <
total, with its own CSS state (grey track, thumb parked halfway) instead
of the on/off track+thumb. Clicking a checkbox out of indeterminate
selects everything, per browser default — unchanged behavior, just an
honest visual for the in-between state.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0129WdYHccPsH27k5GGuwijd
Wizard's Asset step now opens with a required 'Engagement name' field
(validated before advancing or launching). The name isn't a harness/CLI
concept, so it's persisted server-side as runId -> name in
.neurosploit/web-engagement-names.json (keyed off the CLI's own run id,
captured from its 'run id : ns-...' log line) so the sidebar, live run
header, and run detail can label a run by name instead of the raw
target/run-id, surviving a server restart.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0129WdYHccPsH27k5GGuwijd
Full frontend rewrite following a deliberate visual direction (dense
security-operations console — borders over shadows, two radii, one accent,
no gradients/glassmorphism) and fixing real bugs found in review:
- EventSource on the exploit stream never called es.close() on 'done',
so the browser silently reconnected and re-streamed the whole job
(duplicate log lines/findings). Fixed.
- Sidebar 'running' step indicator and openRun() matched ANY running run
instead of the one belonging to the current job (by runId). Fixed.
New:
- 5-step engagement wizard (Asset -> Scope & Auth -> Leads -> Model & Run
-> Review) replacing the single flat board — inspired by the
Discovery/Plan/Exploit/Remediate stage model both a.security and
terra.security use publicly.
- Model is now a real dropdown sourced from /api/providers (mirrors
harness::models::providers()), with an API-key vs. subscription toggle
that disables subscription for API-only providers.
- One Auth & Keys menu: target auth header + named roles (IDOR/BOLA/BFLA
multi-identity testing) materialize into an ephemeral creds.yaml passed
via --creds; per-provider API keys live in server memory only (never on
disk) and are merged into every spawned child's env.
- Generative Attack Path Chaining: findings rendered as kill-chain columns
(recon -> initial-access -> ... -> impact) with chains_from resolved to
parent titles, live in the run view and static in run detail.
- Findings are now a proper table (severity/title/endpoint/CWE/agent/
confidence) instead of stacked cards.
- Explicit light/dark theme toggle persisted in localStorage, defaulting
to light (previously light only won when the OS wasn't in dark mode).
- All UI strings in English.
Backend additions: GET /api/providers, GET/POST/DELETE /api/keys,
ephemeral creds.yaml generation for auth/roles, env override merged into
every exploit-job and REPL child spawn.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0129WdYHccPsH27k5GGuwijd
New web/ app (zero npm deps, Node http built-ins only):
- server.js reads agents_md/ to build a categorized lead board (435 agents
auto-classified into Business Logic / Broken Access Control / Injection /
LLM Application / Auth & Session / SSRF / API / Cloud & Infra / etc.),
reads runs/ for history, and spawns the compiled neurosploit CLI binary
for every exploitation job — structured findings/phase/progress are parsed
from its stdout (finding_json:/phase lines), same signal the TUI uses.
- REPL drawer spawns `neurosploit` with no subcommand (real interactive
session, Reader::Plain over the piped stdin) and streams stdin/stdout —
every /command works exactly as in a terminal, nothing reimplemented.
- SSE endpoints for both job and REPL streams; run/finding/report assets
served under /api/runs/:id/asset/*.
- public/{index,app.js,style.css}: lead board with category toggles + custom
leads + Start Exploitation, live run view (progress/findings/log), run
detail view, REPL drawer — screenshot-inspired layout.
- web/API.md: full endpoint reference. web/README.md: quick start.
Bump version 3.6.9 -> 4.0.0 (Cargo.toml, CLI banners, README/TUTORIAL).
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0129WdYHccPsH27k5GGuwijd
Add two new model providers, both usable via API key or --subscription
(local CLI login, no key):
- opencode: OpenCode Zen gateway (OPENCODE_API_KEY, opencode.ai/zen/v1).
Subscription mode drives the `opencode` CLI (`opencode run --auto`).
Supports the Playwright MCP (--mcp): our .mcp.json is converted to
OpenCode's own config schema and injected via OPENCODE_CONFIG.
- nous: Nous Research / Hermes models (NOUS_API_KEY,
inference-api.nousresearch.com/v1). Subscription mode drives the
`hermes` CLI (NousResearch/hermes-agent) on the user's Nous Portal
OAuth login (`hermes setup --portal`), via `hermes chat -q`. No
CLI-level MCP hook — falls back to Hermes's own built-in toolsets
(web/terminal/computer-use).
Both wired into cli_binary_for, installed_cli_backends, cli_login_status
(prompt passed as argv, not stdin — neither CLI reads stdin for this).
Bump version 3.6.8 -> 3.6.9 across Cargo.toml, README, TUTORIAL, setup.sh,
install.ps1, and in-binary version strings. README/.env.example updated
with the new provider rows and subscription-login table.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HHFAVCHMvRkTy9Wgw7SayG
- Add RECON_TOTAL_BUDGET_SECS (300s) total wall-clock cap across all rounds
- Per-round budget directive in prompt: 30-50 commands max, stop early if enough intel
- Elapsed time check between rounds: skip remaining if budget exhausted
- Remaining time communicated to follow-up rounds for self-pacing
- RELEASE.md updated with recon budget section
Previously: subscription CLI recon ran 150+ commands over 15 min, exploitation never started.
Now: recon caps at 5 min total, then proceeds to agent exploitation.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
- models.rs: detect connection-refused and timeout on local providers
(ollama/litellm/llamacpp), show actionable error instead of raw reqwest
- pipeline.rs: findings with empty evidence skip adversarial vote (which
always rejects per 'default to rejected' prompt) and go straight to
needs-review for human triage
- pipeline.rs: warn when single-model panel + vote_n=1 (same model
validates its own findings = weaker validation)
- Bump version 3.6.7 → 3.6.8
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011739wMqPJJPttTLLX6YoQH
- Preflight: abort a run early with '✗ target unreachable … is DOWN' when the
probe gets no HTTP response, instead of running agents against a dead host;
print '✓ target is UP' otherwise.
- When no --auth/creds are set on a web run, force account_registration_and_forms
to run first so the authenticated surface is always attempted and visible.
- Move the credential vault to <cwd>/.neurosploit/vault/<run-id>.json (persistent
project store) via new RunConfig.vault_dir; header now prints the vault path at
launch. engagement_ops + finish() resolve paths through vault_paths().
- New agent account_registration_and_forms (+1 → 430): analyzes the app's forms
and self-registers a benign test account (curl or Playwright) to reach the
authenticated surface when no creds are given.
- Probe extracts form details (action/method/fields/kind/CSRF) so form analysis is
grounded; shown in the probe summary and recon JSON.
- Hard anti-flood guardrail in SAFETY_DOCTRINE + the agent: at most 2 accounts per
engagement, never loop/script/batch the register endpoint or flood the DB; reuse
the account made; a test needing many sign-ups is a lead, not mass-creation.
- Credential vault: engagement_ops directive tells agents to append created
accounts to <run-dir>/vault.jsonl; finish() consolidates to vault.json, masks
secrets in the report, and adds a 'Test accounts created (DELETE after)' cleanup
finding listing each account and how it was created.
- Finding tagging: new auth_context (authenticated/unauthenticated) and account
fields, rendered per-finding in the HTML report.
- Opt-in disposable email (off by default): /tempmail on + RunConfig.temp_email;
agents may use the free mail.tm API to read a registration confirmation code.
- Tests: parse_forms unit tests; docs updated (README/TUTORIAL/RELEASE), counts 430.
- Add 12 technique/scenario LLM red-team agents (AI category 18 → 30, total 429):
jailbreaks — AdvPrefix, PAIR, TAP, Crescendo, many-shot, persona/DAN,
encoding/obfuscation, refusal-suppression; prompt-injection scenarios — direct,
indirect (RAG/web/email/tool output), goal hijacking, tool/function-call abuse,
system-prompt/secret exfiltration. Each runs an attacker→LLM-judge loop
(baseline refusal → technique across variants → verdict), proving the bypass
with a benign, redacted receipt. Generated by scripts/build_llm_redteam_v365.py.
- Add REDTEAM_DOCTRINE and inject it into run_ai so every AI test follows the
baseline→technique→judge method across scenarios.
- Models: add Claude Opus 5 and Sonnet 5 (Anthropic) and a new Moonshot AI (Kimi)
provider with Kimi K3/K2 (moonshot:kimi-k3, MOONSHOT_API_KEY) — 15 providers.
- Docs: README/TUTORIAL/RELEASE — new AI/LLM red-team engagement mode + section,
model/env-key tables, agent-library counts (429), badges.
Also includes the v3.6.4 grounding fix (#33) landing on main.
The grounding gate ran in empirical mode for every engagement, demoting
white-box (and skills/n8n audit) findings that had passed the n-model vote
because a file:line code citation isn't raw tool output. Grounding is now
mode-aware:
- Symbolic (white-box SAST / skills): a file:line reference into the reviewed
source, or a quote of code present in it, is the receipt — no live target.
- Empirical (black-box / host / AI): evidence must resemble tool output (as before).
- Either (grey-box): a source citation OR a tool receipt grounds a finding.
The symbolic check runs against the reviewed source corpus (not the transcript)
and falls back to a structural file:line + quote check when the corpus is
unavailable. Adds unit tests incl. a regression test for #33.
- /continue (and /resume) now relaunch a recovered interrupted run on the same
target, carrying its findings forward and steering agents to widen coverage /
chain from them instead of re-reporting. Offer shown at launch; a fresh /run
supersedes it. Findings merge (dedup by title+endpoint) across both runs.
- Opening /results, /finding or /report while a run streams no longer corrupts
the terminal: live background output is paused for the picker (still captured
in /logs) and restored on exit, so Ctrl-C in a picker can't take the process
down mid-run.
A missing or un-downloadable recon tool must never block the run. Both the recon
intensity directive and the general tool doctrine now instruct agents to:
- wrap every install in `timeout 90 <install> || echo skip` and run non-interactively
- try each tool install at most once; on failure/no-package/no-network/hang, skip
immediately and fall back to an installed alternative or curl/nc/dig/python3
- never wait on, retry, or block the whole recon for a single tool download
- Drive `codex exec --json` and parse its JSONL event stream into the same
categorized live feed as Claude Code (exec/edit/tool/net/tokens), so recon and
exploitation are visible as each command runs instead of a silent black box.
- Fix the activity feed to keep per-agent tool events (commands, network, files,
findings) and only filter model reasoning + token telemetry, so /logs shows the
real command trail and /status 'last:' is a true sign-of-life.
- Surface failed internal commands as 'exec: (exit N)'; keep Codex auth/rate
detection from stderr.
- /status now shows progress in EVERY phase: a real bar once agents are selected,
otherwise the current pre-exploit phase + counters (cmds, activity lines), plus
a "last:" sign-of-life line (the latest activity) and the actual full findings.
Before, the bar only appeared after agent selection, so a long recon looked
frozen. Findings count now uses the full list.
- New /logs [n] — dump the recent activity feed (recon/tools/findings) of the
running test; useful with non-streaming CLIs (codex) or after scrolling. Backed
by a capped feed ring buffer + last/lines counters in RunLive.
Symptom: with a non-streaming subscription CLI (codex), a long/intense recon
showed nothing in the feed ("phase starting") and the 5-min idle guardrail killed
the run before any agent ran.
- render_compact now SHOWS recon/probe/ai-recon/skills-audit/loaded/running lines
(were dropped) so a long recon no longer looks frozen.
- Idle guardrail reworked: resets on ANY streamed activity (not only new
findings) and only ARMS after exploitation starts (agent launch / vote) — recon
can never trip it. Message: "no activity in N min".
- RunLive.ingest sets phase=recon on recon/probe lines (was stuck at "starting").
`codex exec` in --dangerously-bypass-approvals-and-sandbox mode exits non-zero
when a tool/command it ran internally (curl/nmap/etc.) returned non-zero — even
though it produced a valid final answer. chat_cli treated any non-zero exit as a
hard failure and dropped the output ("recon round 1 failed ... exit 1"). Now, on
non-zero exit WITH usable stdout and no auth/rate/quota keyword, we use the
output; only genuine auth/rate/quota errors (or empty output) fail hard.
Added the OpenAI GPT-5.6 line to the provider pool: gpt-5.6-sol (frontier/default),
gpt-5.6-terra (balanced), gpt-5.6-luna (fast/affordable). Version 3.6.0 -> 3.6.1.
Recon was a single quick model pass — now it's deep and iterative:
- deep_recon(): an initial deep enumeration pass then follow-up EXPANSION rounds
that chase discovered subdomains/hosts/endpoints/params, converging when a
round finds nothing new. Rounds scale with intensity.
- recon_intensity_directive(): tells the agent HOW hard to recon and to INSTALL
the tools it needs (apt/pip/go/npm/cargo) — subfinder/amass/httpx/gau/katana/
gf/arjun/ffuf/nuclei/nmap/dnsx/linkfinder/whatweb/nikto/testssl — chained
(subfinder->httpx->katana/gau->gf->ffuf); covers subdomains, crawl+wayback, JS,
content/param discovery, ports, versions, API, exposures, TLS/headers.
- RunConfig.recon_intensity (default 3) + REPL /recon <1-4> + CLI --recon <1-4>
(1 quick .. 4 exhaustive); shown in /show.
- DECISION_DOCTRINE injected into exploit/grey/chain prompts: analyse responses to
pick the technique; map & connect routes (endpoint output → next endpoint input);
hunt sensitive flows; mine parameters (incl. hidden from JS/source maps) and test
per-param; mock realistic (non-PII) data to reach deeper logic; exploit the
authenticated surface after login and compare roles; build PoCs when a proof
needs an artifact; bypass 401/403/redirect controls.
- REPL /auth now supports multiple named identities (/auth admin <hdr>, /auth user
<hdr>; bare token → Bearer). With >=2 roles the run gets the access-control
directive (IDOR/BOLA/BFLA/privesc, authorized-vs-unauthorized) and tests both.
- +6 decision agents (library 389): param_miner, endpoint_flow_linker,
authenticated_surface_exploit, clickjacking_poc (HTML PoC), csrf_poc (HTML PoC),
access_control_bypass.
- Docs: counts 383->389, RELEASE + /auth help updated.
- setup.sh: downloads the prebuilt release asset for the detected OS/arch (no Rust
needed; latest release auto-resolved), installs binary + agents_md to
~/.neurosploit-app, symlinks into ~/.local/bin, and PERSISTS PATH +
NEUROSPLOIT_BASE into the shell rc (bash/zsh/fish). Falls back to a source build
(NEUROSPLOIT_BUILD=1 to force). Idempotent.
- install.ps1: same for Windows — downloads windows-x64 zip, installs to
%LOCALAPPDATA%\NeuroSploit, sets User PATH + NEUROSPLOIT_BASE (setx), source-build
fallback (incl. arm64).
- find_base(): auto-discovers agents_md/ NEXT TO THE EXECUTABLE (resolves the PATH
symlink via current_exe) and at common install dirs — so `neurosploit` runs from
ANY folder even without the env var. Env override still takes precedence.
Verified: symlinked binary run from /tmp with no env finds all 383 agents.
- /results (interactive, no arg) now ALWAYS opens the run/test picker (target →
vuln → detail, Esc back) instead of jumping straight to the current run's vulns.
The live run (if any) appears at the top, past runs newest-first — so you can
browse every test, not only the active one.
- /validate [n]: re-run false-positive validation (N-model voting + adversarial
refute) on a recovered/past run's findings WITHOUT re-testing the target, then
rewrite that run's findings + report. Backed by new harness::pipeline::revalidate.
Use this after a crash/quit recovered raw findings into /runs.
- Ctrl-C at the prompt now CONFIRMS instead of silently cancelling: with a live
run it offers [s]top&validate / [q]uit(keep findings) / keep-running; otherwise
asks "exit? [y/N]" — so a stray Ctrl-C can't lose a running test.
- tool_doctrine: agents now actively DRIVE the browser on JS/SPA targets — use
the Playwright MCP (render, read live DOM, click client-side routes, watch the
network to find the real API, screenshot proof); when no MCP, use the Playwright
CLI (write+run a small script / npx playwright screenshot) to render and capture
XHR/fetch traffic — complementing curl (which only sees the empty shell).
- probe: detect SPAs (<app-root>, ng-version, near-empty body + linked scripts →
Angular/React/Vue/SPA) and note in recon that the browser is required, so the
SPA agents get selected.
- +8 SPA/API agents (library 383): spa_api_discovery, spa_hidden_admin,
login_sqli_bypass, dom_xss_spa, api_bola_numeric_ids,
register_privilege_mass_assign, jwt_forgery_spa, spa_business_logic.
- Docs: README/RELEASE/TUTORIAL counts + notes.
Why runs came back empty / "MCP didn't execute":
- Not logged in: a subscription CLI that isn't authenticated returns empty
instantly (the Juice Shop symptom — every agent 0 candidates, no tool activity).
Added models::cli_login_status + subscription_preflight(): before a run we check
the primary provider's CLI is installed AND logged in and warn clearly if not
(CLI run_mode + REPL start_background).
- Missing browser: ensure_playwright_mcp now also runs `npx playwright install
chromium` (best-effort; NEUROSPLOIT_SKIP_BROWSER_INSTALL=1 to skip) so the first
browser action doesn't fail/hang.
- Codex MCP was mis-wired (`--config mcp_config_file=` is not a codex key). Now
injects our .mcp.json servers via `-c mcp_servers.<name>.command/.args` TOML
overrides — MCP works on Codex, not only Claude. gemini/grok remain built-in-tools
only (no MCP flag).
- REPL diagnostic: subscription+MCP run with zero tool/browser events warns the
CLI likely isn't logged in / MCP didn't start.
New harness::probe runs a real request/response analysis of the target BEFORE
the model recon and injects the observed facts into recon, so agent-selection
and exploitation decisions are grounded in evidence (robust even when model
recon is weak):
- status & redirect, Server/X-Powered-By/content-type, 6 security headers,
cookie flags (HttpOnly/Secure/SameSite), CORS reflection test (arbitrary
Origin + credentials), tech fingerprint, linked scripts, form count, a 404
baseline for soft-404 differentials, and high-signal paths (/robots.txt,
/.git/config, /.env, /sitemap.xml, /.well-known/security.txt).
- Best-effort (never fatal — degrades to a note on network failure), honors the
identifying User-Agent and the Burp/ZAP proxy. Wired into black-box run() and
greybox recon. A one-line probe summary streams to the live feed.
Attribution (anti-plagiarism), multiple layers:
- Identifying User-Agent on every request (default NeuroSploit/<ver> + an
X-NeuroSploit-Scan header), overridable via /ua or NEUROSPLOIT_UA env; shown
in the run banner. RunConfig.user_agent + Session.user_agent wired through.
- Every finding is stamped "Identified and validated by NeuroSploit …" (in
finish() and the raw-report path) so provenance travels in the finding text,
findings.json and the report.
Multi-role authentication for access-control testing (IDOR/BOLA/BFLA/privesc):
- creds.yaml gains named identity blocks (admin:/user:/victim:/…), each with
jwt | header | cookie | apikey | login+username+password. With >=2 roles the
harness injects a cross-role access-control directive (authorized-vs-unauthorized
proof) and defaults the primary auth to the first role.
Also: /help now lists one command per line (fixes smushed OPTIONS/RUN columns);
/ua command + Session field; docs (README + RELEASE) updated.
The OPTIONS/RUN sections crammed a second command into the description column
(/clear, /quit, /offline, /chain, /theme appeared as loose text), which was
confusing. Every command now has its own aligned row; split /attach+/context and
/diff+/retest; added /results, /finding, /report, /offline, /theme rows; added
/finding and /expand to Tab-completion.
REPL (v3.5.5):
- /timeout <min>: idle guardrail — if no NEW finding lands within the window the
run soft-stops and validates what was found (default 5 min; 0 disables).
- /target accepts a comma-separated list; /run tests them SEQUENTIALLY (a queue
auto-advances to the next target when the current run finishes; one report each).
- /results (no arg, interactive): navigation browser — pick target/run → pick
vulnerability → full detail; Esc steps back a level (vuln → target → session).
- /report (no arg, multiple runs): pick which report to open from a menu.
- /show now shows idle-stop; help updated.
Agent prompts:
- RECON_SYS deepened: crawl + params/headers/cookies, DOWNLOAD & analyze linked
JS (endpoints, hidden params, GraphQL, secrets, sourceMappingURL), fingerprint
exact versions, response-differential analysis; richer JSON schema.
- tool_doctrine adds JS-analysis and request/response-analysis guidance
(linkfinder/gau/katana, header/cookie/timing/length differentials).
The prompt passed to rustyline embedded ANSI escapes AND a newline (dim context
line + colored `neurosploit›`), so rustyline mis-measured the prompt width and
cursor position — typing/backspace/history/cursor got garbled in a real
terminal (fine when piped, which has no line editor).
Now: the dim context line is printed with println!() ABOVE the prompt, the
readline prompt is plain "neurosploit› " (correct width), and the magenta color
is applied via Highlighter::highlight_prompt (display-only, doesn't affect width).
Replaces the single-shot chain_round with attack_chain(): an iterative,
per-foothold pivot engine.
- Each round takes the newest confirmed footholds (best-first, capped) and, for
EACH one, an agent DECIDES which directions to expand — post-exploitation
(loot creds/keys/config/source), credential reuse, horizontal+vertical
privesc, lateral movement to adjacent services/hosts, data exfiltration, and
new attack surface the foothold exposes — proving each step with a receipt.
- LOOT (creds/tokens/hosts/endpoints) discovered in one round is carried forward
and reused by later rounds (parsed from a {"findings":[...],"loot":[...]} reply).
- New findings are validated each round (never pivot off a false positive) and
become the next round's footholds. Loop-until-dry or chain_depth rounds.
- New RunConfig.chain_depth (default 2) + --chain-depth flag on all engagement
commands (0 disables). CHAIN_SYS rewritten for decision/post-ex framing.
- Robust verdict parsing (pool::parse_verdict): whitespace-insensitive, checks
explicit rejection first, counts only explicit confirmations; ambiguous →
Unclear (not confirmed). Replaces the fragile exact-JSON / loose "yes" match.
- Severity-aware quorum (pool::quorum_confirmed): High/Critical now need ≥2
validators AND ≥2/3 agreement (a single vote can no longer confirm a
Critical); lower severities need a strict majority (>half, was ≥half). Single-
model panels fall back to majority so they aren't nuked.
- Adversarial refute pass (REFUTE_SYS): every confirmed High/Critical is
re-examined by a skeptical panel that assumes false-positive; findings that
can't withstand a majority of skeptics are dropped. Survives on infra failure.
- Strengthened VOTE_SYS with an explicit false-positive checklist (reflected-not-
executed, version/banner guesses, self-XSS, error-as-injection, thin evidence,
inflated severity); validator query now also includes impact.
- Unit tests for parse_verdict + quorum_confirmed.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
New harness module `integrations` (+ app commands) wiring NeuroSploit into the
SDLC. Config persists per-project to .neurosploit/integrations.json; secrets are
NEVER stored — only the env-var name is saved, values read from the environment.
GitHub:
- private-repo clone (token injected into the clone URL for whitebox/greybox/tui)
- `neurosploit pr <owner/repo> <n>`: clone the PR head (refs/pull/N/head),
white-box review, optional `--comment` (PR summary) and `--jira` (cards)
- `neurosploit watch <owner/repo> --branch --interval`: re-review on each new commit
GitLab:
- private-repo clone (oauth2 token) for whitebox/greybox (gitlab.com or self-hosted)
Jira:
- `--jira` on any engagement opens one card per finding (REST /issue, basic auth)
Control:
- `/integrations` (REPL): show · enable/disable · setup jira|gitlab|github
- `neurosploit integrations [show|enable|disable] [github|gitlab|jira]` (CLI)
Docs: README "Integrations" section + new TUTORIAL-INTEGRATION.md (per-tool setup,
scopes, recipes, troubleshooting). Version bumped 3.5.2 → 3.5.3.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
GitHub Actions workflow that, on a pushed v* tag (or manual dispatch), builds a
self-contained NeuroSploit (binary + agents_md/) for every OS/arch and uploads
the archives to the matching release. macOS builds are also attached manually.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Resolves the only two open issues that still apply to the Rust build:
- #21 Azure OpenAI: new `azure` provider (OpenAI-compatible). Endpoint comes
from AZURE_OPENAI_ENDPOINT, api-version from AZURE_OPENAI_API_VERSION
(default 2024-10-21); the model name is the Azure deployment; auth uses the
`api-key` header instead of Bearer. Use `--model azure:<deployment>`.
- #25 Gemini key confusion: GEMINI_API_KEY now also accepts GOOGLE_API_KEY
(Google's standard env var) as an alias; local providers (ollama/litellm)
require no key. .env.example documents both.
Kept under the v3.5.2 line (additive provider support).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
`whitebox <arg>`, `greybox --repo <arg>`, `tui --repo`, and the REPL `/repo`
now accept a git URL (https://github.com/owner/repo[.git], git@…, ssh://, *.git)
or an `owner/repo` shorthand. A new resolve_source() shallow-clones it into
<base>/repos/<name> (cached, .gitignored) and reviews it; existing local paths
are used unchanged. Works identically with API-key (--model) and --subscription.
Verified: `neurosploit whitebox https://github.com/digininja/DVWA --offline`
clones DVWA and runs the 78 code agents over 120KB of source.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Distilled from reviewing real AI-pentest output that kept stopping at "exposed"
instead of "exploited". Pure-additive, back-compatible.
Behavior (injected into black/grey/chain exploit prompts via DEPTH_DOCTRINE):
- Exposed → exploited: any info-disclosure / exposed service/WSDL / leaked
credential|token / reachable dev host MUST be used before it's a finding;
otherwise it's a lead, not a confirmed High/Critical.
- Chain across modules: reuse obtained session/JWT/cookie/credential and pivot
to IDOR/privesc/exfil; report the chain, not isolated parts.
- Decode & fingerprint → CVE; audit tokens (alg-confusion/none/kid/JWKS, weak
HS256 secret cracking, lifecycle).
Deterministic post-pass (new crates/harness/src/hygiene.rs, wired into finish()):
- calibrate severity to PROVEN impact — unproven High/Critical (hedged, no
payload, thin evidence) capped to Medium and re-titled "(potential)";
- depth_audit — flag exposures on a host with no real exploit;
- hygiene_summary — advise consolidating hygiene classes repeated across assets.
Unit tests cover calibration + depth audit.
5 new doctrine meta-agents (scripts/build_methodology_v352.py → agents_md/meta/):
exploit_depth_doctrine, finding_chainer, artifact_decoder, token_auditor,
report_calibrator (meta 17→22, total 343→348).
Version bumped 3.5.1 → 3.5.2 across crates/app/installers/docs; RELEASE/README
updated.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
These session/runs/history files are runtime state generated during local
testing; .neurosploit/ is already in .gitignore. Untrack them so the repo
doesn't carry test artifacts.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Token/quota exhaustion no longer silently drops agents. When every candidate
model is rate-limited / out of quota, the run PARKS (keeping all state) and
prints "⏸ token/quota exhausted … PAUSED". The user can:
- wait for renewal and /continue (retry same model), or
- /model <provider:model> (or the /model selector) then /continue to switch.
Implemented via ModelPool: is_exhaustion() detection, park_exhausted() that
awaits a resume Notify, and a fallback-model slot tried first on retry. /model
queues the chosen models into a paused run's fallback so a plain /continue
resumes on them.
Findings now survive a crash/quit: each finding is checkpointed live to
.neurosploit/active_run.json; on next launch an interrupted run is recovered
into /runs (a raw report is materialized) so /results, /finding and /report
keep working.
/stop now actually halts immediately on raw/discard: one() races the in-flight
model call against the hard-cancel flag, so the CLI child (kill_on_drop) is
terminated at once instead of finishing its whole command sequence. The
validate path still soft-stops (lets validation run).
Docs: TUTORIAL documents the 3-way /stop, crash recovery and pause/continue;
/help lists /continue and the new behaviors.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Drop the legacy Python-stack settings (DATABASE_URL, HOST/PORT, RAG,
Kali sandbox, Discord/Telegram/Twilio, feature flags) that no longer
exist in the Rust harness. Keep only the provider API-key env vars the
model pool actually reads, plus the Ollama/LiteLLM base-URL overrides.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
- New `litellm` provider (kind=api). Use `litellm:<model>` — model names pass
through to your gateway. No hardcoded key required (proxy may be open).
- Env-configurable base URL: LITELLM_BASE_URL (default http://localhost:4000/v1),
LITELLM_API_KEY. OLLAMA_BASE_URL override added too.
- TUTORIAL documents the LiteLLM env config.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
CRITICAL BUG: truncate()/source-context slices cut strings by BYTE, panicking on
a multibyte char (e.g. '—'). The panic crashed agent tasks → task.await returned
JoinError → unwrap_or_default() → empty RunOutput. Result: real confirmed findings
(win.ini traversal, HTML injection) were silently lost, workdir was empty, report
missing. Now all string truncation is char-safe (models.rs, pipeline.rs, repl.rs).
Also:
- Background runs: /run now runs in the BACKGROUND via rustyline's ExternalPrinter
— the REPL keeps accepting commands while the engagement streams live. New
/status (live phase + progress bar + findings) and /stop (graceful). Findings
persist to history + report on completion (finalize_run ensures workdir is set
even on abort, fixing "no report file in ").
- Progress bar: agents-done/total with %, shown in /status.
- Severity colors in the live feed (Critical=red…Info=grey); confirmed vote = green.
- /help reformatted into clear aligned sections.
- TUTORIAL: document non-blocking runs, /status progress, /stop, colors.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
- BUG: /auth (and /creds /focus /target /repo) with no argument CLEARED the value
instead of showing it — so typing /auth to view wiped your credential. Now no-arg
prints the current value; clear only with an explicit `clear`.
- /show now also displays API-key status (set/missing) for the selected models'
providers, and a hint of which commands edit config.
- REPL /run prints a clear "▶ RUNNING (prompt returns when done; use tui for live)"
banner before and "◀ back to the NeuroSploit REPL" after, so it's obvious the
REPL didn't disappear during a run.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
- Claude-Code-style @ menu: rustyline CompletionType::List so @path shows a
file/folder selection list (Tab), not inline cycling.
- /diff (/changed): shows new (+) / gone (-) findings between the last two runs.
- /retest [n]: loads a past run's target/repo and seeds a re-verify focus on its
findings → /run to check if they're fixed.
- Both added to Tab-complete and /help.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Chaining:
- agents_md/chains/ (12 multi-stage exploitation playbooks): SQLi→RCE→LPE,
SSRF→AWS-creds, SSRF→RCE, upload→RCE, upload→LFI→RCE→LPE, XSS→ATO, IDOR→ATO,
SSTI→RCE→cloud, default-creds→domain, deserialization→RCE, exposed-git→RCE,
subdomain-takeover→trusted-abuse. Each stage proven by a tool receipt before
advancing; reports chains_from edges.
- Loaded as a `chains` category (→ 329 agents). chain_round now injects the chain
recipes as a menu so the LLM applies proven multi-stage paths.
Persistence (no DB — structured state):
- Per-project `<cwd>/.neurosploit/` holding session.json (config), runs.json
(history), history.txt (readline). REPL resumes target/repo/auth/focus/models
on reopen; saves on /run and /quit.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
- Live findings feed: each candidate is surfaced (✦ possible finding [sev] title
@ endpoint) the moment an agent returns it, not only at the end.
- 🔔 notifications in the feed: evidence saved, phase complete (with severity
breakdown = automatic partial summary). Renderer styles notify/finding tags.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Partial observability is now first-class:
- belief.rs — property-graph world model; nodes (host/service/vuln/exploit/cred)
carry a probability, not a boolean. Bayesian observation updates; per-node
Shannon entropy; mean-uncertainty + recon-frontier. Black-box = diffuse priors
that sharpen with observation; white-box collapses toward deterministic (MDP).
- pomdp.rs — value_of_information(), decide() (recon vs exploit falls out of
belief entropy), and may_assert() — the mathematical anti-hallucination gate:
no exploitability claim while the belief is diffuse (high entropy) → observe first.
- grounding.rs — verification engine, hard rule "no claim without a tool receipt":
empirical grounding for black-box (raw HTTP/OOB/error markers), symbolic for
white-box (file:line into reviewed source). Ungrounded claims demoted + flagged
receipt_missing (feeds future reward shaping).
- pipeline.finish(): grounding gate before reporting + belief-uncertainty readout.
- bump 3.5.0 → 3.5.1; README documents the v3.5.1 belief/grounding architecture
and the infra/bandit/reward roadmap.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
- Streamed Claude events now tagged with the agent label (@name) so every
command/tool/file is attributable to the agent that ran it.
- Token/cost telemetry: parse usage from the stream-json result event; feed shows
per-call in/out/cost and a running total in the run summary.
- Ctrl-C during a run no longer hard-kills: it cancels cooperatively (no new
agents launch, in-flight bounded), then asks "generate report from partial
results? [Y/n]" — discard removes the run dir. Second Ctrl-C aborts.
- pool: cancel handle + is_cancelled; one()/complete_routed/chat_cli carry a label.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
REPL (rustyline Helper):
- Tab autocomplete for /commands and @filesystem-paths.
- @path attach: @file, @folder, @file:LINE / @file:START-END fold scope files /
stack traces into the agent context; /attach <path> and /context to manage.
- Multiline input: end a line with `\` to continue (validator-driven).
- /theme color|mono, /config (=/show); history (↑/↓) persists as before.
- Attachments are merged into the run's instruction context.
Install:
- setup.sh: `curl … | bash` — auto-installs Rust, clones to ~/.neurosploit,
builds release, links neurosploit into ~/.local/bin; idempotent; env-tunable.
README: v3.5.0, 🧠 (back to "neuro"), one-line install section, neurosploit-on-PATH usage.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
- Finding enriched with owasp / mitre / kill-chain stage / exploitability /
business_impact / chains_from (attack-path edges).
- attack_graph module: derive OWASP Top 10 + MITRE ATT&CK technique + kill-chain
stage from CWE (heuristic, no extra model call); render a Mermaid attack-path
flowchart (findings grouped by stage, explicit + implicit edges) and an ASCII
kill chain for the REPL.
- enrich() runs in finish() for every engagement.
- HTML report gains an "Attack Path & Kill Chain" section (Mermaid via CDN, dark)
plus a stage/sev/OWASP/MITRE/exploitability table.
- REPL print_findings shows the ASCII kill-chain + severity summary after a run.
- models: add GPT-5.5, GPT-5.4, GPT-5.4-mini, GPT-5.3-codex, GPT-5.2.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Harness:
- ModelPool gains a progress channel (set_progress); chat_cli forwards it.
- New chat_claude_stream: drives Claude Code with --output-format stream-json and
parses the event stream live — assistant text, and tool_use blocks categorized
into tagged events (exec/danger command, read/edit file, net request/browser,
grep/glob tool). 900s bound; clear error surfacing.
- Wired set_progress into run / whitebox / greybox.
REPL renderer (render_line):
- Tagged events render as the conversation feed: tool/command/network as compact
CARDS (tool-runner visual), files/edits/AI text/states as iconized lines.
- Clear "what the AI is doing" states: reconning, planning, testing, validating,
chaining, report, complete — plus a ⚠ DANGEROUS marker for risky commands.
- Untagged harness lines mapped to the same state vocabulary.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Harness:
- Exploit-chaining round: after validation, chain confirmed findings into deeper
impact (SSRF→metadata, SQLi→dump→reuse, IDOR→ATO, file-read→secrets→RCE),
validate the new findings, merge. Wired into black-box and greybox.
- Latest top models surfaced: claude-opus-4-8, gpt-5.1/gpt-5.1-codex, gemini-3-pro.
REPL:
- Real line editing via rustyline: ↑/↓ command-history recall, Ctrl-A/E/K, paste;
Ctrl-C cancels the line, Ctrl-D exits. Command history persists to
data/repl_history.txt. Graceful plain-stdin fallback when not a TTY.
- /model with no arg → arrow-key multi-select (dialoguer); with arg accepts any
provider:model names.
- /key is model-aware: lists the providers your selected models need (set/missing)
and prompts for the missing keys; /key <prov> <key> still works.
- Run history persists to data/repl_runs.json and reloads across sessions
(/runs lists past + current; /results /report /status by run number).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
- RunOutput exposes `workdir` so the session can locate reports.
- Session now records every run (RunRecord: id, mode, target, workdir, findings).
- New commands:
/runs list runs done this session (mode, target, severity counts)
/results [n] show findings of run n (default last), severity-sorted
/report [n] open the PDF/HTML report (open/xdg-open)
/status [n] print the run's status.json
/offline on|off pipeline self-test toggle (no model calls)
- Each /run prints "saved as run #n" with the quick commands.
- Verified offline: run → /runs → /results → /status all work.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
- harness/creds::login(): performs the real HTTP login (POST/GET form), captures
a session Cookie from Set-Cookie or a Bearer token from the JSON body, with a
soft success check (no hard fail on 302). Redirects not followed so Set-Cookie
is visible.
- apply_creds is now async: direct material (jwt/header/cookie) used as-is; a
`login:` flow is EXECUTED to obtain a live session; on failure, falls back to
instructing the agents to log in themselves.
- --creds + --focus added to `run` (authenticated black-box) too.
- Verified live against a local mock: POST /login → 302 + Set-Cookie captured as
the auth header used on subsequent requests.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
- New GREYBOX mode: review a repo's source AND exploit the running app in one
pipeline — code-review findings become LEADS injected into live exploitation.
CLI: `neurosploit greybox <repo> --url <app> [--creds creds.yaml] [--focus ...]`
REPL: set both /repo and /target → greybox auto-selected.
- Credentials (harness/src/creds.rs, dependency-free YAML subset): jwt / header /
cookie, or an automated `login:` flow. Derives an auth header and/or a
"authenticate first via curl" directive injected into prompts so agents test
authenticated. --creds flag + /creds command + creds.example.yaml.
- RunConfig gains `repo`; run_engagement refactored to a Mode enum (Black/White/Grey).
- Verified offline: greybox loads creds, combines repo+URL, runs pipeline, writes report.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>