New --research mode for whitebox/greybox (REPL /research, web 🔬 checkbox, or
auto-detected from natural-language focus/objective in PT/EN). Steers the source
review to find a NOVEL, CVE-reportable issue instead of a known one:
- WHITEBOX_RESEARCH_DOCTRINE: pin version/commit; research known CVEs/advisories
(SECURITY.md, CHANGELOG, GHSA, NVD, git history) to de-duplicate; patch-diff /
n-day->0-day variant analysis (incomplete fixes, bypasses of a new check,
sibling sinks, reintroductions); strict novelty gate (each finding states
novel-why + checked-against); benign PoC + dynamic confirm on greybox.
- RunConfig.research + is_research_intent(); injected in run_whitebox and the
greybox code-review half.
- 6 research skills (code/): known_cve_dedup, patch_diff_variant,
attack_surface_map, source_to_sink_taint, logic_authz_flaw,
dependency_nday_reachability.
- Methodology modeled on a real AppSec-research workflow (no specifics copied).
479 agents, 421 tests.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Parse per-agent activity from the stream (launching agent / exploit·analyze·test
<name> via <model> -> N candidate(s) / failed) into a status map, surfaced in the
live view as a grid of chips coloured by state (pending/running/found/done/failed)
with a per-agent finding count and a running/done tally. Answers 'what is it
testing right now' at a glance; resets per engagement.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
The REPL-backed path (run/whitebox/greybox) spawns `neurosploit` with NO
subcommand, so only global flags are valid in argv — but authArgs() emitted
run-subcommand flags there (--environment, --policy, --in-scope, --budget,
--compliance, --revalidate-poc, --token-limit, --deep-test-limit, --order,
--sample-per-route, --scope-file). clap aborted on the first one
("unexpected argument '--environment'"), so the engagement died at launch and
the live view sat empty. authArgs now emits only the real global flags with
their global names (--session-environment / --session-in-scope / --session-policy,
plus --capability-token/--transport/--oob-*/--sms/--typesafe/--decision-backend/
--intercept/--sandbox); run-only knobs ride the REPL script or defaults.
Also:
- sidebar: a disk run whose status says "running" but has no live job is shown
as "interrupted", not "running" (no more stale RUNNING entries); brand-new
in-memory jobs are injected so an engagement appears the moment it starts;
new "Interrupted" group; clicking a running/interrupted row attaches the live
stream or offers resume.
- stop: robust now — graceful /stop then SIGTERM/SIGKILL fallback, a second
press escalates, a job whose child already exited is marked done so the UI
stops showing it as running.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Billing concern: a full run with 2-3 voters, deep chaining and exhaustive recon
burns a lot of tokens. --quick is one switch for a fast, cheap pass:
one voter, one chain round, light recon (intensity 1), ≤6 agents, eco budget.
Dropping voting from 3 models to 1 is the biggest saver.
- CLI: --quick on run/whitebox/greybox (apply_quick, applied last so it wins)
- REPL: /quick (aliases /economy /eco), listed in help + completion
- web: ⚡ Quick-mode checkbox in the wizard -> /quick in the REPL script
- README: documented in the flags table
- 420 tests passing
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
The web server's job state was in-memory only, so a node restart/crash lost the
live run (the harness still checkpointed to .neurosploit/active_run.json, but
the console couldn't see or re-attach it).
- persist each job to .neurosploit/web-jobs/<id>.json (snapshot + feed tail +
sanitized launch params; NO api keys or creds contents), throttled, on
findings/phase/done
- loadPersistedJobs() on boot: a job that was live becomes `interrupted`, and
`resumable` when it was a REPL-backed run/whitebox/greybox
- POST /api/exploit/:id/resume: relaunch the REPL (NEUROSPLOIT_AUTO_RESUME=1),
which recovers the on-disk checkpoint and /continue's it, carrying findings
forward; reuses the same job id so the live view resumes streaming
- API-key jobs need the provider key re-entered after a full restart (kept only
in memory) — resume returns a clear 409 saying which; subscription resumes clean
- frontend: boot offers a dismissible "N interrupted run(s) — Resume" banner
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
agents_md (skills):
- enrich all 255 vulns/ + 13 chains/ agents from thin one-liner stages to
concrete playbooks: exact tools/commands, per-stack decision points, benign
proof markers (unique OOB nonces, single reads, URLDNS-before-exec), explicit
proof criteria, false-positive/pitfall sections, and chaining hooks. Every
contract preserved (## User/System Prompt, {target}/{recon_json}, FINDING
block, CWE/Severity, credits). avg 37->53 lines; loader parses all 449.
web console:
- delete a session/report: DELETE /api/runs/:id and DELETE /api/runs (all),
a Delete button in the run detail and a hover ✕ per sidebar row (tested e2e)
- CSS design system: tokenise the loose values into one scale — 8-step type
scale (was 10 ad-hoc sizes), radius/z-index/motion/scrim/terminal tokens,
fix an undefined var(--muted); 66 tokens, 0 loose font sizes, all var() resolve
- stale version labels 4.0.0/4.2.0 -> 4.2.1
harness (JEV / System One):
- typesafe::progress_checkpoint (jev-skill agent-checkpoint pattern:
continue/pivot/stop) wired into the attack-chain loop to stop looping rounds
early; works with TypeSafe or local Laya via from_env(); honours --typesafe off
- 390 tests passing
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
New `mobile` engagement mode: `neurosploit mobile <app.apk|app.ipa|binary>`
reverse-engineers a local artifact with a dedicated `mobile` agent set, all
headless and provisioned on demand (Ghidra analyzeHeadless, MobSF REST/Docker,
Frida, apktool/jadx, radare2).
Twelve original, generic skills (agents_md/mobile/, English): static binary
triage, APK static analysis, IPA static analysis, RASP & anti-tamper mapping,
root/jailbreak detection + bypass, TLS pinning detection + bypass, anti-debug
detection + bypass, obfuscation analysis & deobfuscation, code-integrity /
tamper-check bypass, hardcoded-secrets extraction, insecure local storage, and
mobile network traffic analysis. Findings are proven from the artifact
(decompilation or Frida trace), non-destructively.
- agents.rs: new `mobile` Library category (loaded, counted).
- pipeline.rs: run_mobile() mirroring the host pipeline with a mobile recon and
headless tooling doctrine; exported from the crate.
- CLI: `Cmd::Mobile` + `Mode::Mobile`, wired in main and the TUI.
- README + TUTORIAL document the new test type; engagement-modes badge + table
updated; "New in v4.2.0" note. Version bumped to 4.2.0 across the workspace.
383 tests.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Version bumped to 4.1.0 across the workspace, binaries, web console and Typst
template.
README: new "New in v4.1.0" summary; trimmed the verbose highlight bullets and
the TypeSafe section; removed the anti-plagiarism/provenance section (provenance
stays in the code, just not front-and-centre in the README); TypeSafe promoted
to its own top-level section; agent count 446.
TUTORIAL: new section 17 "Assurance & authorization" covering the target gate,
--scope-file, evidence-graded CVSS, audit anchoring + assurance bundle, sandbox,
intercept proxy, PoC re-validation, compliance mapping, TypeSafe, and the
internal/AD graph + budget governor.
benchmarks/typesafe-2026-09-20/: the with/without TypeSafe measurement —
report.html, scorer, both runs' findings/assurance/meta/logs, and a README.
No secrets committed (env-only during the runs, verified clean).
381 tests.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
TypeSafe cannot BE an LLM agent — System One does not generate text or call
tools. But it can be the decision brain of a code-owned confirmation loop, and
that is what typesafe_agent.rs is: an ADDITIONAL confirmation strategy.
typesafe_agent.rs — for enumerable classes (XSS, SQLi, open-redirect, path
traversal, SSRF, IDOR): code lists candidate payloads, a TypeSafe Choice picks
the next one given what's been tried, the replay engine sends it for real, a
TypeSafe Noul judges the response, loop until confirmed or exhausted. Edge/WAF
answers are refused. Pure parts (class table, payload templating, id-swap, OAST
substitution, query encoding) are unit-tested; the networked loop is integration.
Wired as a pipeline pass that runs ONLY on findings the LLM path left
unconfirmed or in needs-review (the recall lever) — it can raise a finding to
confirmed with a calibrated probability, never downgrades (the deterministic
layer owns that).
--typesafe on|off|auto (global flag) resolves into the env the pipeline reads,
governing adjudication, CVSS re-grade, agent pruning and this loop together.
`off` runs the identical pipeline without TypeSafe; meta.json records
"typesafe": true|false so a with/without pair is a clean A/B measurement. Web
console gets the same toggle.
381 tests.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Hard scoping was already enforced in code (every request passes
ScopePolicy::check_request; exclude beats allowlist; capability token caps
it; out-of-scope findings withheld + audited). What was missing was a way to
author that boundary from a file or the web form instead of only CLI flags.
- scope.rs: ScopePolicy::from_yaml / from_file — a dependency-free parser for
the friendly string format (app.example.com, *.wildcard, CIDR, url-prefix),
the same strings Pattern::parse already takes, NOT the raw serde {kind,value}
shape. Strict in one direction: an unreadable file errors, an empty hard list
authorizes nothing (a safe failure, but the operator's choice, not a typo).
- CLI: --scope-file <yaml>. Loaded before authorization so --in-scope adds to
it and the capability grant still caps it.
- Web: a full Scoping & Guardrails section in the Authorization tab — hard
scope, exclusions, observe-only, destructive-method + account-creation
toggles, max accounts, rate limit, forbidden payloads, notes. The server
materializes a scope YAML and passes --scope-file; notes stay labelled
"guidance, NOT enforced" so prose is never mistaken for a control.
- examples/scope.example.yaml documents the format.
End-to-end verified: web form -> YAML -> Rust loader -> enforced boundary.
332 tests (+4).
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Closes the three benchmark gaps and adds the two the user asked for.
poc.rs — re-runs each finding's recorded proof and sorts it into reproduced /
changed / gone / unverifiable. The last two are kept apart deliberately: a PoC
that could not be tested (out of scope now, state-changing, nothing recorded)
is never reported as one that failed. Never re-runs a mutating request to
"confirm" it. Can only lower a finding's standing, never raise it. Wired as a
run pass (--revalidate-poc) and a subcommand (neurosploit poc <run> --apply).
proxy.rs — own recording forward proxy (HTTP in full; HTTPS tunnelled with
honest metadata, no fake CA) that chains upstream to Burp / Caido / ZAP /
mitmproxy. A bare tool routes straight through it; own+tool records here and
forwards for full TLS interception. Flows -> flows.jsonl, distinct hosts become
passive-discovery leads. Harness and agent child commands share one route.
sandbox.rs — Kali docker/podman container: no host network, no mounted socket,
no-new-privileges, workdir mounted, proxy/transport env inherited. A missing
runtime is an explicit error, never a silent fallback to host execution — the
whole point being to keep attack payloads off the operator's host. Subcommands
sandbox up|exec|install|down.
compliance.rs — maps confirmed findings onto PCI-DSS v4.0, HIPAA Security Rule
and SOC 2 controls. Phrased as "bears on control X", never "compliant/non-
compliant"; the disclaimer is rendered on top and absence of a finding is never
presented as compliance. Report section + `neurosploit compliance <run>`.
validation.rs — 8 new deterministic validators (19 -> 27 classes): verbose
errors/stack traces (CWE-209), cleartext/HSTS (319), CRLF response splitting
(113), dangerous HTTP methods (650), GraphQL introspection, exposed backup
files (530), Host header injection (644), cacheable private responses (525).
Each names exactly what it saw and rejects the classic false positives (a
block page echoing a payload, the SPA served under a bogus path, a copyright
year mistaken for a code).
All wired through RunConfig, the CLI (global --intercept/--sandbox; run-level
--revalidate-poc/--compliance) and the web console's Tooling & assurance block.
328 tests.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
transport.rs — internal engagements happen through a VPN, a bastion or a
tunnel, and the dangerous failure is silent: with the VPN down, 10.20.0.15
is a machine on the operator's own network and the scan succeeds against
the wrong host. So an internal target with no transport is refused before
any traffic leaves, and a transport that is up must prove it (the apparent
source address has to change) rather than be assumed. Supports SOCKS, HTTP
proxy, OpenVPN, SSH bastion (dynamic or single-host forward) and cloudflared.
oob.rs — our own Collaborator, self-hosted by default because callbacks are
engagement data (internal hostnames, resolver addresses, sometimes the
exfiltrated value). HTTP and DNS listeners written on tokio directly, no new
dependency. The two levels of proof are separated in code: an HTTP callback
proves egress, a DNS query proves only that a resolver saw the name — the
overclaim this channel otherwise invites.
inbox.rs — mail.tm and inbound SMS (Twilio or webhook). extract_code() scores
candidates by surrounding text and returns nothing rather than a guess, so a
copyright year never gets submitted as an OTP. A throttling claim requires
delivered messages carrying DISTINCT codes, not HTTP 200s.
Wired through RunConfig, the CLI (global flags, so a session cannot re-route
itself mid-engagement), the REPL and the web console's Authorization tab.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Budget (opt-in, unlimited by default so an un-budgeted run is unchanged):
- crates/harness/src/budget.rs — modes, phase shares, Token Governor
- CLI: --budget/--token-limit/--deep-test-limit/--coverage-first/
--depth-first/--sample-per-route; same controls in the web wizard
- pipeline honours it: vote_n narrows, evidence rounds are capped
Run control parity in the web console:
- /pause in the REPL, backed by a pause gate in the model pool: in-flight
agents finish, then the run holds with every finding kept
- POST /api/exploit/:id/{pause,continue,report} + GET .../log
Provenance (crates/harness/src/provenance.rs):
- JOASNSCOPE sigil leads every canary, so a marker found in a response,
a log or someone else's report extracts whole and names its build
- per-build fingerprint, per-run id, optional per-customer build id
- findings.json stamped with _engine; signed provenance.json manifest
- structural signature survives rewording but not a changed result set
- prompts watermarked at the single pool chokepoint
- `neurosploit provenance show|scan|verify`
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
A PDF was only ever produced while a run was finishing. If `typst` was missing
at that moment — or the template improved afterwards — the operator had no way
to get one without re-running the whole engagement against the target.
report::rebuild() regenerates every artifact (md · json · html · pdf) from the
findings already on disk, exposed as `neurosploit rebuild <run-id|dir>` and as
POST /api/runs/:id/report with a "Generate report" button in the run view. The
endpoint shells out to the harness rather than reimplementing report generation
in JavaScript, so there is one implementation instead of two that drift, and it
says plainly when the PDF was skipped for want of `typst` instead of handing
back a link to a file that was never produced.
Also fixes write_all() to pass the run's pocs/ listing into the HTML report, so
a rebuilt report links the scripts each finding cites — the run-time path
already did this and the rebuild path silently did not.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The risk model, grants and hash-chained trail existed as modules nothing
called. Now every engagement runs under them.
Capability
- `neurosploit capability issue|verify` mints and inspects grants.
- `--capability-token` (global, so the REPL takes it too), `--in-scope`,
`--environment`, `--policy` on `run`; verification happens at the command
line, so an invalid grant fails with a readable message instead of halfway
through an engagement.
- The pipeline verifies before anything else and REFUSES to run on a token that
does not verify — proceeding would mean acting on an authorization nobody can
prove was issued. `effective_scope` then applies the grant as a ceiling.
- Web: an Authorization tab carrying the token, extra hosts, environment and
policy profile. The browser decodes the claims for display and says plainly
that it is not verifying them — a "valid" badge from a party without the key
would be the UI vouching for something it cannot check.
A hole the smoke test found: `/inscope evil.test` inside a session under a
grant WIDENED the scope past it — the one thing a capability token exists to
prevent. The run itself would still have been constrained (the pipeline
re-applies the grant), but `/policy` reported a boundary that was not real, and
a tool that misreports its own limits is worse than one with none. Scope
mutations now re-apply the ceiling and name what it refused. Session
authorization also arrives from argv rather than a `/`-command, because a
session that can widen its own grant is not constrained by one.
Audit
- One hash-chained record per action in `<run>/audit.jsonl`, in the specified
shape, covering engagement start/end, validator rejections, findings that
reach the report (with the hash of the evidence behind them) and findings
withheld for being out of scope.
- The run verifies its own chain at the end and says loudly if it is broken.
- `/audit [n]` tails the trail and verifies it; the web offers it as a download
next to the report, so "show me what the tool did" is a link.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Backend
-------
- knowledge_graph.rs — the durable structure under attack_graph's per-run view:
typed entities (asset/endpoint/weakness/technique/finding/account/credential/
impact) joined by typed, weighted, provenance-carrying edges, accumulated
across runs in .neurosploit/graph.json plus a per-run copy the report and web
console can draw. Answers what a finding list can't: ranked attack paths, and
the frontier of entities observed but never proven — where chaining should
look next. Agents only sometimes fill chains_from, so progression is also
inferred between adjacent kill-chain stages; those edges are marked inferred,
weighted lower, and drawn dashed, because presenting a hypothesis as evidence
is the graph lying about itself. Secrets stay in the vault, never the graph.
- memory.rs — four tiers scoped by lifetime, not importance: working (one run),
engagement (one target), technique (one agent/CWE), reusable (generalized).
Promotion is evidence-gated and needs independent evidence at each step: a
claim repeated within a run becomes engagement knowledge; one confirmed
across runs becomes technique knowledge; one that held on two DIFFERENT
targets is generalized into a reusable lesson with host-specific tokens
stripped. Nothing is promoted on a single observation, which is exactly what
a hallucination looks like. Recall is scored (overlap × past success ×
recency) and injected into recon/exploit prompts as leads to verify. Recalled
memos are credited only when the run they informed actually found something.
- rectify.rs — a mistyped command cost a full round trip through /help, at the
worst possible moment during a live run. Accepted-as-typed wins over
everything (so the /url alias is never "corrected" to /ua), then unique
prefix, then Damerau-Levenshtein with a length-scaled budget, and a tie is
reported rather than resolved. Arguments too: a bare host gets its scheme, an
out-of-range count is clamped with a note instead of silently reverting, a
near-miss model id is matched against the live catalog.
- pool.rs — when every configured model is exhausted or its token is dead, try
whatever else this machine can actually reach (an installed CLI subscription,
or a provider whose key is in the environment) before parking. A run that
stops on a box with three other usable backends stopped for no reason.
- repl.rs — /memory, /forget, /graph; a recovered run resumes by itself where
nobody is watching (piped stdin — the web console — or NEUROSPLOIT_AUTO_RESUME),
since a `/continue` prompt there waits forever.
Web
---
- Attack path: the stage list was seven hardcoded values, so findings the
harness staged outside it were silently dropped — 5 of 27 on a real run.
Rewritten against the harness's own stage list with unknown stages kept,
two-line labels (every node used to read "SQL Injection Authent…"), stage
column headers, pan/zoom/fit, path highlighting, severity filter, and the
run's graph.json used when present.
- Dashboard: coverage, findings by severity, top weaknesses, and annualized
loss exposure via FAIR — frequency from exploitability × validation
confidence, magnitude from assumptions shown on screen and editable, reported
as a range. The posture score saturates instead of subtracting, so it keeps
discriminating past the first critical.
- Run history groups into one folder per target with a filter, instead of one
flat list that grows forever.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BvdGy9XtVWSdXDTa3FFLJv
Replaces the floating REPL drawer with a docked terminal, and fixes the
usability problems a screenshot audit of the console turned up.
Terminal (the reason for the change):
- The drawer rendered the harness into a <div>, so the server had to strip
ANSI before sending it: colour, the box-drawn /status panel and the banner
all arrived flattened, and long lines rewrapped mid-glyph. The stream is now
sent verbatim and rendered by xterm.js (vendored, nothing fetched at
runtime), decoded with a streaming UTF-8 decoder so a multi-byte character
split across two reads survives.
- The drawer floated bottom-right, directly over "Next →" and "Start
Exploitation" — the wizard's primary buttons. The dock is a flex child of
.main, so opening it shortens the view instead of covering it. Drag its top
edge to resize; the height is remembered.
- The child is spawned over a pipe, not a PTY, so it never echoes: line
editing is local — echo, ←/→, Home/End, history, Tab completion over the
slash commands, Ctrl+C/L/U/K/A/E. Ctrl-C is delivered as SIGINT by the
server, since a raw 0x03 byte over a pipe interrupts nothing.
- A target picker switches the terminal between a standalone REPL session and
the engagement currently running, so mid-run instructions go to the same
process doing the testing.
QA fixes:
- Findings tables sorted by severity (a LOW above a CRITICAL made a 27-row
result unreadable), with sortable headers, a severity summary that doubles
as a filter, a text filter, a sticky header, and horizontal scroll confined
to the table instead of the whole page.
- alert()/prompt() replaced by inline field errors, a custom-lead modal and
toasts — a modal alert hid the very field it was complaining about.
- Lead categories start collapsed (412 leads over ~30 categories); search
auto-expands what it matches and shows per-category hit counts.
- Sidebar rows truncate inside the rail (a long target URL used to spill past
its border), and carry a worst-severity dot, finding count and age.
- Past-run header shows when it ran, how many agents ran, the recon asset,
PoC count and run id — two runs of one target were indistinguishable.
- Off-canvas sidebar below 768px had no way to be opened; added the toggle.
- Long evidence values (cookies, tokens) now wrap instead of running under
the finding modal's edge.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BvdGy9XtVWSdXDTa3FFLJv
Root cause of "can't send prompts while a run streams": /api/exploit
spawned a plain `neurosploit run ...` subprocess, and that CLI path
(run_mode() in main.rs) never reads stdin - it only waits on the task or
Ctrl-C. The ONLY thing in the harness that keeps accepting input while an
engagement streams is the interactive REPL's background-run loop. So:
- New startJobViaRepl(): for mode run/whitebox/greybox, spawns a bare
`neurosploit` REPL session and scripts it via stdin (/target or /repo,
/model, /sub, /mcp, /votes, /chain, /recon, /focus, /objective,
/scope-out, /creds, /only <agents> or /only clear, then /run) instead
of building CLI args. Same underlying pipeline, same tagged output
lines, so all existing parsing (findings/phase/progress/runId) works
unchanged. host/aitest/skills modes stay on the old one-shot
startJob() - they need onboarding's scope picker, an interactive
arrow-key menu that silently skips itself over a piped stdin, so they
can't be scripted this way.
- New POST /api/exploit/:id/input writes a line to the session's stdin -
natural language, /status, /continue, anything the REPL accepts - and
the live run view grows a "send prompt" box (in the Activity log tab)
for it, shown only when the job reports interactive: true.
- Stop, for an interactive job, now sends the REPL's own graceful
'/stop\n1\n' (validate what's found, then report) instead of SIGINT -
the REPL's own input loop has no signal handler, so SIGINT there would
just kill the process outright and skip the report step. Non-
interactive jobs still get SIGINT (run_mode() does catch that).
- 'done' can no longer be process-exit only: an interactive session stays
open after the engagement finishes (for /report, /continue, another
/run), so ingestLine() now also flags done from the same "phase
complete" content signal it already used for the phase field.
Verified end-to-end: started an interactive job, confirmed
`interactive: true` and a captured runId, sent /status and /agents mid-
and post-run over the new /input endpoint (both accepted, session stayed
alive and responsive after completion), and confirmed a non-interactive
run is unaffected.
Also: the missing "Activity log" tab a screenshot showed for a "running"
engagement was the sidebar's detail-view fallback (2 tabs, no log) for a
run whose Job object no longer exists in server memory - it happens when
the Node process gets restarted while a spawned neurosploit child is
still alive underneath it (an orphan from testing across many redeploys
this session, not a code bug); the live view itself always had the tab.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0129WdYHccPsH27k5GGuwijd
- Select all / Clear all buttons in the Leads step toolbar - respects the
current search filter, so filtering to "sql" then Select all only pins
those, not all 412 leads. The per-category master switch (already
select/deselect-all for that category, indeterminate when partial) was
the only bulk control before; this adds the "everything" case.
- '+ Custom lead' now generates an ACTUAL specialist-agent markdown file
(agents_md/vulns/custom_<slug>.md, same format every other agent uses)
via the claude CLI on the operator's Anthropic subscription
(claude-opus-4-8 by default - matches the harness's own default model),
instead of folding free text into --focus. The new lead is immediately
selectable and pinnable via --only like any other agent; verified the
Rust harness's own agent loader picks it up (agent count went 435 -> 436,
neurosploit agents confirmed it).
Two things found and fixed while wiring this up:
- the skip-permissions flag gave the model file/bash tool access, which
made it try to write the file itself and narrate doing so instead of
just returning text. Dropped the flag (pure text completion needs no
tools) and told it explicitly not to use any.
- Even so, defensively strip anything before the first '# ' heading
before saving, in case a model still prepends commentary.
Falls back to the old free-text-focus behavior if generation fails
(claude not installed/logged in, malformed output, timeout) so the
operator's intent isn't lost.
- New "Custom Leads" category, shown first, so generated leads have a
visible home instead of landing in the catch-all "Other" bucket.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0129WdYHccPsH27k5GGuwijd
Real front-end bugs found and fixed:
- [hidden] never worked on any element whose class also sets 'display'
(every .btn, .chip, ...): the browser's built-in '[hidden]{display:none}'
rule and an author rule of equal specificity tie, and the later one in the
cascade wins — so 'Next' stayed visible on the Review step alongside
'Start Exploitation', and 'Open report'/'Stop' rendered during 'starting'.
Fixed with a single global '[hidden]{display:none!important}' override.
- Progress bar was functionally correct but easy to miss (thin, 0%-width,
low-contrast track) and gave no feedback while the agent count is still
unknown (recon phase). Added a border for visibility and an indeterminate
sliding-segment state for the 'agents: ?' window.
- A live run watched in the browser was lost on F5 (jumped back to the
wizard) even though the job keeps running server-side. The active job id
now persists in localStorage; on load the app reconnects the SSE stream
(the server replays its full event buffer) instead of losing the view.
New:
- Findings are now clickable — a detail modal shows every Finding field
(CWE/CVSS/OWASP/MITRE/stage/exploitability/confidence/votes/review status/
auth context/account/agent), endpoint+payload, evidence, impact, business
impact, remediation, and chains_from — in both the live run and past-run
detail views.
- PoC surfacing: the finding modal looks up any script the run wrote to
pocs/ that's cited in the finding's evidence (per the harness's own
doctrine — see pipeline.rs change below), fetches and previews it inline,
with a link to open the raw file. Live runs poll for new PoC files every
5s once the run id is known.
- Pinned-leads confirmation: the live run header now states plainly how
many leads were pinned (and their names) or that selection is auto
(recon-driven) — this was previously buried in the scrolling activity log
behind the harness's unconditional 'Loaded 435 agents' library-size line,
which describes the full agent library, not what will actually run.
Harness doctrine (crates/harness/src/pipeline.rs, pocs_line()):
PoC-writing for black-box findings was previously conditioned on 'when an
issue needs a custom multi-step exploit/script' — vague enough that a
straightforward finding (single-request XSS/SQLi/IDOR) often got no PoC
file at all. Now required for every confirmed Medium+ finding, one
standalone .py/.sh script per finding, and explicit about citing the exact
file name in the finding's evidence field (which is what the web UI now
matches on to link a PoC to its finding).
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0129WdYHccPsH27k5GGuwijd
Wizard's Asset step now opens with a required 'Engagement name' field
(validated before advancing or launching). The name isn't a harness/CLI
concept, so it's persisted server-side as runId -> name in
.neurosploit/web-engagement-names.json (keyed off the CLI's own run id,
captured from its 'run id : ns-...' log line) so the sidebar, live run
header, and run detail can label a run by name instead of the raw
target/run-id, surviving a server restart.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0129WdYHccPsH27k5GGuwijd
Full frontend rewrite following a deliberate visual direction (dense
security-operations console — borders over shadows, two radii, one accent,
no gradients/glassmorphism) and fixing real bugs found in review:
- EventSource on the exploit stream never called es.close() on 'done',
so the browser silently reconnected and re-streamed the whole job
(duplicate log lines/findings). Fixed.
- Sidebar 'running' step indicator and openRun() matched ANY running run
instead of the one belonging to the current job (by runId). Fixed.
New:
- 5-step engagement wizard (Asset -> Scope & Auth -> Leads -> Model & Run
-> Review) replacing the single flat board — inspired by the
Discovery/Plan/Exploit/Remediate stage model both a.security and
terra.security use publicly.
- Model is now a real dropdown sourced from /api/providers (mirrors
harness::models::providers()), with an API-key vs. subscription toggle
that disables subscription for API-only providers.
- One Auth & Keys menu: target auth header + named roles (IDOR/BOLA/BFLA
multi-identity testing) materialize into an ephemeral creds.yaml passed
via --creds; per-provider API keys live in server memory only (never on
disk) and are merged into every spawned child's env.
- Generative Attack Path Chaining: findings rendered as kill-chain columns
(recon -> initial-access -> ... -> impact) with chains_from resolved to
parent titles, live in the run view and static in run detail.
- Findings are now a proper table (severity/title/endpoint/CWE/agent/
confidence) instead of stacked cards.
- Explicit light/dark theme toggle persisted in localStorage, defaulting
to light (previously light only won when the OS wasn't in dark mode).
- All UI strings in English.
Backend additions: GET /api/providers, GET/POST/DELETE /api/keys,
ephemeral creds.yaml generation for auth/roles, env override merged into
every exploit-job and REPL child spawn.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0129WdYHccPsH27k5GGuwijd
New web/ app (zero npm deps, Node http built-ins only):
- server.js reads agents_md/ to build a categorized lead board (435 agents
auto-classified into Business Logic / Broken Access Control / Injection /
LLM Application / Auth & Session / SSRF / API / Cloud & Infra / etc.),
reads runs/ for history, and spawns the compiled neurosploit CLI binary
for every exploitation job — structured findings/phase/progress are parsed
from its stdout (finding_json:/phase lines), same signal the TUI uses.
- REPL drawer spawns `neurosploit` with no subcommand (real interactive
session, Reader::Plain over the piped stdin) and streams stdin/stdout —
every /command works exactly as in a terminal, nothing reimplemented.
- SSE endpoints for both job and REPL streams; run/finding/report assets
served under /api/runs/:id/asset/*.
- public/{index,app.js,style.css}: lead board with category toggles + custom
leads + Start Exploitation, live run view (progress/findings/log), run
detail view, REPL drawer — screenshot-inspired layout.
- web/API.md: full endpoint reference. web/README.md: quick start.
Bump version 3.6.9 -> 4.0.0 (Cargo.toml, CLI banners, README/TUTORIAL).
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0129WdYHccPsH27k5GGuwijd