RECON_SYS enriched with the concrete bug-bounty recon arsenal and tool pipelines
(subfinder/amass/assetfinder/crt.sh → httpx liveness → gau/waybackurls/katana/
gospider URL harvest → JS analysis & secret regexes → arjun/x8 params → gf vuln
patterns + qsreplace → ffuf/feroxbuster content discovery → naabu ports → dnsx/
nuclei takeovers → cloud bucket grep → targeted nuclei → chained pipeline),
passive-first and scope-respecting. Shared by the black-box run path, so it
applies identically to the CLI, the REPL and the web console.
Version bumped to 4.2.3 across CLI/clap/web/README/TUTORIAL. 423 tests.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
(1) Preflight now checks EVERY configured model, not just the primary — a 3-model
jury that silently collapses to 1 (others not logged in / no key) is now shown:
each model prints usable/needs-login/needs-key, with a '→ N/M models usable'
summary. Fixes 'I set 3 models and only opus ran'. (qwen is API-only: use
nous:qwen3.8-max for Hermes, or export DASHSCOPE_API_KEY.)
(2) Wildcard engagement (*.domain in scope) now gets a deterministic subdomain
fan-out before recon: enumerate via crt.sh + subfinder/amass (in the Kali sandbox
when --sandbox, else host), keep only IN-SCOPE hosts, probe liveness, drop
soft-404/parked catch-alls, rank (401/403 auth hosts first — flagged for bypass —
then interesting names admin/api/dev/staging), and fold the live list into the
test surface so the run tests the WHOLE authorized domain end-to-end, not just the
seed host. Scope-respecting: one lightweight GET per host, no degradation.
(3) recon_tool() runs a recon tool in the Kali sandbox or on the host — the basis
for tool-powered recon.
423 tests passing.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
A copy-paste full-surface focus string (all web classes, prioritise authed
surface + subdomains, chain to impact, reproducible receipt) and an
objective-vs-focus note; the engagement template now carries the broad
focus/objective in its one-file config.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
- Wildcard target from a scope-file (target: "*.nasa.gov") was probed literally
→ "builder error" / target unreachable. It's now reduced to the apex
(https://nasa.gov) for the seed, while the scope keeps *.nasa.gov so subdomain
enumeration stays authorized. (The /target command already did this; the
engagement-file meta path didn't.)
- Provenance was a OnceLock ("first run wins"), so in the REPL every run after
the first minted markers and the provenance line with the FIRST run's id
(nasa run showing a rockstargames id). Now a RwLock that rebinds per run —
each engagement gets its own id; the build fingerprint stays stable.
- Nous/Hermes: qwen3.8-max / qwen3.8-omni-flash added to the provider list so
`nous:qwen3.8-max` routes qwen through the Hermes portal (model name passes
through `hermes chat -m <model> --provider nous`).
- Black-box `run` now gets a broad DEFAULT objective when none is set: a
comprehensive WEB assessment grounded in OWASP Top 10 / ASVS / WSTG / CWE that
traverses every applicable web vuln class then goes deep — web-only (this path
loads only web vuln agents; mobile/binary are separate modes), so it never
drifts into mobile/exe.
423 tests passing.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
A copy-paste recipe (cross-model jury, /ua browser, /recon 4, scope-file +
authorization, natural-language focus on the high-value surface, /run) plus a
table of high-value knobs and an honest benchmark reality-check, so people run
NeuroSploit more effectively against real targets.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
While a run streams, every printed line makes rustyline redraw the input line.
The high-volume tool echoes (exec/net/read — the ⌘/🌐/📄 lines, dozens/sec during
recon/curl bursts) flooded the terminal so arrow keys and backspace felt stuck
while typing. Those lines are now rate-limited to ~1/200ms in the live feed (they
remain in full via /logs); findings, phase changes, votes and notices still print
immediately. Editing stays responsive during a run.
(Note: the /model-rewrite the user hit was the OLD binary — v4.2.2 keeps the exact
models typed; verified anthropic:claude-opus-4-8 + openai:gpt-5.6-sol are kept.)
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
/show showed only the policy profile ('web') and hid the actual hard-scope
allowlist, the authorization reference, the pinned agents (/class, /only) and
research mode — so an operator couldn't confirm what was really configured.
Now shows: authorized (the enforced hard-scope hosts, flagged when pinned),
guardrails (rate/accounts/destructive/excludes), authz ref, pinned agent list,
and research in opts.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Field feedback: real, reproducible findings (missing HSTS, insecure cookie flags,
internal IP leaked in a header) were being down-rated/dropped by the adversarial
opinion-vote. The user's rule: never discard something real; another person must
be able to reproduce the same finding.
Validation:
- has_http_receipt(): a finding whose proof is a captured HTTP response (status
line / security headers / Set-Cookie / a header the class is proven by) or a
file:line citation is REPRODUCIBLE by definition.
- validate(): a finding unanimously rejected by the vote but carrying such a
receipt is NO LONGER dropped — it is kept as needs-review (capped to what the
receipt alone proves), because "the response lacks HSTS" is a fact, not a
story. grounded_receipt() now also recognizes an in-prose HTTP receipt.
- reproducibility: a kept finding at a URL with no explicit repro steps gets a
minimal pasteable `curl -i` so anyone can reproduce the exact finding.
Also:
- /pocs (aliases /poc /evidence /artifacts): list a run's synthesized/written
PoC + evidence files with their paths (was: "unknown command /pocs").
- version bumped to 4.2.2 across CLI/clap/web; stale v4.2.1/v4.1.0 labels fixed.
423 tests passing.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Field feedback from a live run (only SQLi being hit, auth flows skipped, scanner
UA behind a CDN):
- EXPLORE_DOCTRINE injected into every exploit prompt: the agent's named class is
a starting point, not a cage. It maps what the app actually does and reports
ANY class it can prove — with authentication/identity (login, signup, password
reset, MFA, OAuth/OIDC/SAML, JWT, session) as a first-class target, plus
business-logic/multi-step flows and both client- and back-end surfaces. When
its own class yields nothing, it pivots instead of idling.
- Selection (SELECT_SYS) now covers the surface instead of collapsing into one
family: diverse set, MUST include auth/identity agents when any auth/OAuth/JWT
surface is present, include business-logic/access-control on authed/multi-step
flows, cover client + back-end when both exist.
- UA quality: /ua browser sets a realistic Chrome UA (attribution stays in the
X-NeuroSploit-Scan header) for accuracy behind a WAF/CDN, where a self-declaring
scanner UA gets blocked/challenged and causes false negatives; /ua identify
keeps the transparent NeuroSploit UA. The UA doctrine now tells agents to
compare both early and switch to the browser UA if responses differ.
422 tests passing.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
- /scope-file now reads optional engagement keys from the SAME YAML (target,
models, focus, objective, authorization, classes) so one file defines the
whole engagement, not just scope. examples/scopes/nasa.yaml and
engagement.example.yaml show the keys.
- When importing a scope (/scope-file) or declaring one (/authorize), a target
left over from a previous session that falls OUTSIDE the new scope is reset to
a host inside it (was: silently kept, then denied on /run — the "nothing
changed" confusion). Added scope_seed_target() + in_hard_scope() check.
- UI/help strings are English; the natural-language REPL still accepts input in
any language (the two example lines are now English).
- README + TUTORIAL updated: new REPL commands (/authorize, /scope-file, /class,
/research, /quick, /authorization, /guardrail), a "Scope — three ways" section
with the one-file YAML, version/counts refreshed to 4.2.1 / 480 agents.
422 tests passing.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Hard scope stays the safety boundary (you must say what you're allowed to test),
but setting it is now frictionless for a normal client pentest where authorization
comes from a signed SOW/contract — no bug-bounty program or capability token.
- /authorize <host|*.dom|cidr|url> ... (aliases /grant, /inscope-set): set the
entire authorized scope in one line (multiple entries), pins it so /target
won't re-derive, and seeds the target so /run works immediately. The operator
asserts written authorization for the listed assets; guardrails (rate,
accounts, destructive) remain tunable via /guardrail.
- examples/scopes/engagement.example.yaml — neutral direct-engagement template
(no program framing): fill hard scope from the SOW, guardrails documented as
yours to tune (e.g. allow destructive in a staging env, raise rate for a lab).
The frictionless path already worked (/target x -> authorized against x); this
makes the multi-asset direct engagement a single clear command. 422 tests.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
- /scope-file <path> (aliases /scopefile, /import-scope): import a ready scope
config (hard allowlist + exclusions + guardrails) in the REPL — one step to
"scope set correctly", instead of typing /inscope repeatedly. Pins the scope.
- scope_pinned: once scope is set explicitly (scope-file / /inscope / capability),
/target no longer re-derives the scope from the target, so an imported
allowlist is not clobbered by picking a target.
- examples/scopes/rockstargames.yaml and examples/scopes/nasa.yaml — ready
TEMPLATES scoped to *.rockstargames.com / *.nasa.gov with conservative,
bounty/VDP-safe guardrails (no destructive verbs, no mass accounts, low rate,
forbidden payloads) and a clear "verify the program's current in/out-of-scope
before running" banner. Both parse and enforce; subdomain enumeration happens
inside the wildcard boundary.
422 tests passing.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Three improvements prompted by a real bug-bounty run:
1. /class idor,sqli,xss,ssrf — focus a run on specific vulnerability CLASSES.
Expands each class to the matching library agents (by name/title/CWE) and
pins them, so /run tests exactly those classes and skips recon-based
selection. Friendlier than naming each agent via /only. Known aliases cover
idor/bola, sqli, xss, ssrf, csrf, ssti, xxe, rce, lfi, auth/jwt, graphql,
race, upload, cors, prototype-pollution, smuggling, and more.
2. /authorization <url> (aliases /authz, /program) — declare the engagement's
authorization (e.g. https://hackerone.com/<program>). Recorded and added to
the rules-of-engagement context so the run is framed as the authorized test
it is, which reduces false model refusals on in-scope bounty targets. It does
NOT widen scope — the grant still comes from /target, /scope-file or a
capability — and the RoE context tells agents to keep within the program's
rules (no DoS/mass-account-creation/out-of-scope; report those as leads).
3. Model-refusal detection: when an agent DECLINES a technique (safety/RoE
pushback, e.g. "can't run mass account creation at a production service"),
the harness now says so plainly instead of hiding it as "0 parseable
findings (malformed JSON)". The operator sees WHY an agent found nothing.
422 tests passing.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Two reported bugs:
1) A scope left over from a previous session persisted in the project session and
kept denying every new target (DENY_TARGET_OUTSIDE_GRANT ... scope *.example.com
even after /target zoom.us). With no verified capability, /target now re-derives
the authorized scope from the new target (preserving excludes + guardrails), so
the target you pick is the target you test — same model as `neurosploit run <url>`.
2) A wildcard target (`*.zoom.us`) now authorizes the apex AND all subdomains and
seeds recon with the apex (a literal `*.zoom.us` has no DNS record to probe), so
subdomain enumeration happens inside the wildcard scope. Applied in both the REPL
/target and the one-shot `run` path.
Also: the run banner and clap about showed v4.1.0 — now use CARGO_PKG_VERSION / 4.2.1.
421 tests passing.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
New --research mode for whitebox/greybox (REPL /research, web 🔬 checkbox, or
auto-detected from natural-language focus/objective in PT/EN). Steers the source
review to find a NOVEL, CVE-reportable issue instead of a known one:
- WHITEBOX_RESEARCH_DOCTRINE: pin version/commit; research known CVEs/advisories
(SECURITY.md, CHANGELOG, GHSA, NVD, git history) to de-duplicate; patch-diff /
n-day->0-day variant analysis (incomplete fixes, bypasses of a new check,
sibling sinks, reintroductions); strict novelty gate (each finding states
novel-why + checked-against); benign PoC + dynamic confirm on greybox.
- RunConfig.research + is_research_intent(); injected in run_whitebox and the
greybox code-review half.
- 6 research skills (code/): known_cve_dedup, patch_diff_variant,
attack_surface_map, source_to_sink_taint, logic_authz_flaw,
dependency_nday_reachability.
- Methodology modeled on a real AppSec-research workflow (no specifics copied).
479 agents, 421 tests.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
The REPL banner and HTML-report footer hardcoded v4.2.0 while the crate is 4.2.1.
Both now interpolate env!("CARGO_PKG_VERSION"); web title/sidebar/comments set to
4.2.1. --version already read the crate version.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Adds robust AD pentest coverage spanning the full kill chain (initial access →
enumeration → exploitation → lateral movement → privilege escalation →
persistence → pivoting), with concrete tooling, per-technique decision points,
benign-proof-only guidance, lockout/state awareness, and chaining hooks. All
GENERIC — no lab-specific hosts/IPs/creds/flags; works in any AD environment.
New infra/ skills: ad_recon_enum, ad_bloodhound_paths, ad_llmnr_poisoning,
ad_ntlm_relay, ad_password_spray, ad_kerberos_delegation, ad_adcs_esc,
ad_pth_ptt, ad_coerce_auth, ad_critical_cve (Zerologon/noPac), ad_smb_share_hunt,
ad_laps_gmsa_read, ad_gpo_abuse, ad_dpapi_looting, ad_trust_abuse,
ad_persistence_review, ad_mssql_abuse. Enriched: ad_kerberoasting, ad_asreproasting,
ad_dcsync, ad_acl_privesc, ad_default_creds, windows_priv_esc.
New chains/: chain_ad_web_to_forest_root, chain_ad_rbcd_s4u_to_adcs,
chain_ad_coerce_relay_adcs, chain_ad_kerberoast_to_domain,
chain_ad_mssql_linked_pivot, chain_ad_trust_cross_forest, chain_ad_local_to_domain.
attack_graph: map CWE-294/295/1392/269 to OWASP/MITRE/stage + CVSS bands so AD
findings grade and place in the kill chain correctly. 473 agents, 421 tests.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Parse per-agent activity from the stream (launching agent / exploit·analyze·test
<name> via <model> -> N candidate(s) / failed) into a status map, surfaced in the
live view as a grid of chips coloured by state (pending/running/found/done/failed)
with a per-agent finding count and a running/done tally. Answers 'what is it
testing right now' at a glance; resets per engagement.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
The REPL-backed path (run/whitebox/greybox) spawns `neurosploit` with NO
subcommand, so only global flags are valid in argv — but authArgs() emitted
run-subcommand flags there (--environment, --policy, --in-scope, --budget,
--compliance, --revalidate-poc, --token-limit, --deep-test-limit, --order,
--sample-per-route, --scope-file). clap aborted on the first one
("unexpected argument '--environment'"), so the engagement died at launch and
the live view sat empty. authArgs now emits only the real global flags with
their global names (--session-environment / --session-in-scope / --session-policy,
plus --capability-token/--transport/--oob-*/--sms/--typesafe/--decision-backend/
--intercept/--sandbox); run-only knobs ride the REPL script or defaults.
Also:
- sidebar: a disk run whose status says "running" but has no live job is shown
as "interrupted", not "running" (no more stale RUNNING entries); brand-new
in-memory jobs are injected so an engagement appears the moment it starts;
new "Interrupted" group; clicking a running/interrupted row attaches the live
stream or offers resume.
- stop: robust now — graceful /stop then SIGTERM/SIGKILL fallback, a second
press escalates, a job whose child already exited is marked done so the UI
stops showing it as running.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
The runs/evidence and pocs folders were always empty on the API-key path: there
the model only returns findings JSON and never executes a tool to write files,
so nothing populated them (they were only ever written by the subscription
agentic CLIs). The harness already holds the structured evidence (evidence_data:
baseline/attack/identity exchanges, marker) and the payload/endpoint.
synthesize_pocs_and_evidence() now writes, per finding and without overwriting
anything an agent already produced:
- evidence/<slug>.md — the request/response proof (baseline/attack/identity
pairs, marker, callback/browser flags, notes), or the prose receipt as fallback
- pocs/<slug>.md — a runnable curl repro from the recorded request(s),
incl. the cross-identity pair for BOLA/IDOR; falls back to endpoint+payload
and cites the PoC path back into the finding's evidence. Runs in the main
pipeline after evidence collection; idempotent. 421 tests passing.
Closes#44
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
The one-shot `neurosploit run|whitebox|greybox <target>` path dropped the
pause/resume handles and never read stdin, so when the pool parked on quota/auth
exhaustion it printed "type /continue …" into a void — nothing accepted it and
the process hung forever on the parked task; only Ctrl-C worked. This is exactly
the "can't /continue, it's stuck" a run that exhausts during recon hits.
run_mode now, at a real terminal, reads stdin and accepts:
/continue [provider:model] resume (optionally switching model)
/model provider:model switch provider/model, then resume
Ctrl-C still stops and offers a partial report; it also clears the pause so a
parked task can't swallow the interrupt. Over a pipe (CI/web) stdin is skipped.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
A single out-of-credit event (Anthropic's 400 "credit balance too low", which
is_exhaustion correctly classifies) made every in-flight parallel agent print
its own "PAUSED — /continue" line, flooding the console with identical notices.
Now only the agent that wins the paused false->true transition emits the notice;
the rest park quietly. Resume clears the flag so the next episode notices again.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Billing concern: a full run with 2-3 voters, deep chaining and exhaustive recon
burns a lot of tokens. --quick is one switch for a fast, cheap pass:
one voter, one chain round, light recon (intensity 1), ≤6 agents, eco budget.
Dropping voting from 3 models to 1 is the biggest saver.
- CLI: --quick on run/whitebox/greybox (apply_quick, applied last so it wins)
- REPL: /quick (aliases /economy /eco), listed in help + completion
- web: ⚡ Quick-mode checkbox in the wizard -> /quick in the REPL script
- README: documented in the flags table
- 420 tests passing
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
The web server's job state was in-memory only, so a node restart/crash lost the
live run (the harness still checkpointed to .neurosploit/active_run.json, but
the console couldn't see or re-attach it).
- persist each job to .neurosploit/web-jobs/<id>.json (snapshot + feed tail +
sanitized launch params; NO api keys or creds contents), throttled, on
findings/phase/done
- loadPersistedJobs() on boot: a job that was live becomes `interrupted`, and
`resumable` when it was a REPL-backed run/whitebox/greybox
- POST /api/exploit/:id/resume: relaunch the REPL (NEUROSPLOIT_AUTO_RESUME=1),
which recovers the on-disk checkpoint and /continue's it, carrying findings
forward; reuses the same job id so the live view resumes streaming
- API-key jobs need the provider key re-entered after a full restart (kept only
in memory) — resume returns a clear 409 saying which; subscription resumes clean
- frontend: boot offers a dismissible "N interrupted run(s) — Resume" banner
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix(pipeline): robust LLM JSON extraction (json5 + truncation repair)
Model replies that did not exactly match the expected JSON syntax were
either dropped silently or surfaced as "[extract_findings] ... JSON parse
failed" / "... no JSON array/object found". Both came from the same two
weak stages in extract_findings: a greedy first-'['-to-last-']' span that
captured prose, and a salvage pass that only stripped trailing commas.
Add a shared, string/escape-aware extractor (crates/harness/src/json_extract.rs):
- locate balanced [..]/{..} regions, ignoring brackets inside prose/strings,
preferring fenced blocks (last wins);
- parse leniently: serde_json first, then json5 (trailing commas, comments,
single quotes, unquoted keys);
- repair token-limit truncation by closing the open structure, keeping the
complete findings instead of discarding the whole batch.
Route extract_findings, reported_nothing, extract_chain, parse_string_array
and prosecutor::parse_verdict through it. Make the diagnostic tail()
char-boundary-safe (the old slice could panic on UTF-8). Add regression
tests for single quotes/comments, capitalised ```JSON fences, truncated
arrays and pure prose.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix(models): robust LLM response handling + higher token/timeout limits
Harden the OpenAI-compatible chat client against the empty-content and
parse failures hit with reasoning models (GLM/DeepSeek via OpenRouter)
during whitebox runs:
- Accept message `content` as a string, an array of content parts, or a
`reasoning_content` fallback; surface `finish_reason` and empty-choices
errors instead of an opaque "no content in response".
- Stop masking mid-stream body-read failures as a bogus "EOF while
parsing"; report read timeouts and empty bodies explicitly, and
reassemble SSE-framed responses some gateways return unrequested.
- Raise reasoning-model max_tokens to 32768 and the HTTP timeout to 300s;
both overridable via NEUROSPLOIT_MAX_TOKENS / NEUROSPLOIT_HTTP_TIMEOUT.
Adds unit tests for content extraction, token sizing, and SSE reassembly.
Cargo.lock syncs the json5 entry from the prior extraction commit.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix(pipeline): unwrap findings/selection replies wrapped in an object
The json5 + truncation work made JSON *parsing* robust, but the *shape*
handling after it still dropped data when a model wrapped its answer in an
object instead of returning the bare array we asked for. Most visible on
black-box runs, where a long tool-use turn ends with the model narrating
into a report object.
extract_findings treated any object as ONE finding, so a real batch returned
as `{"findings":[…]}` (or `{"vulnerabilities":[…]}`, …) became a single
title-less "finding", was filtered out, and surfaced to the operator as
"returned text but 0 parseable findings" while the findings were lost. Add
findings_items() to normalise the shape: an array is the list; an object with
a title is one bare finding; otherwise an object wrapping a known findings key
unwraps to that array. reported_nothing() now recognises the same wrapper keys
so an empty `{"vulnerabilities":[]}` reads as an honest negative.
Two more consumers of the same class:
- parse_string_array (agent selection) accepted only a bare array of strings,
so a wrapped `{"agents":[…]}` or elements-as-objects `[{"name":"sqli"}]`
silently fell back to RL ranking. Now unwraps the wrapper and pulls the
string from object elements.
- extract_chain hard-coded the "findings" key for its object branch, dropping
the sibling `loot` under any other wrapper key. Now checks all wrapper keys.
Add regression tests for each shape.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Each dashboard "Top contributors" row is now a button: clicking it opens the
run it belongs to and pops that finding's detail modal (evidence/impact/PoC).
Matches the finding by title, falling back to CWE; if it was recalibrated or
merged, opens the run and says so. Fetches the run detail directly to avoid
racing loadDetail's async fill.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
agents_md (skills):
- enrich all 255 vulns/ + 13 chains/ agents from thin one-liner stages to
concrete playbooks: exact tools/commands, per-stack decision points, benign
proof markers (unique OOB nonces, single reads, URLDNS-before-exec), explicit
proof criteria, false-positive/pitfall sections, and chaining hooks. Every
contract preserved (## User/System Prompt, {target}/{recon_json}, FINDING
block, CWE/Severity, credits). avg 37->53 lines; loader parses all 449.
web console:
- delete a session/report: DELETE /api/runs/:id and DELETE /api/runs (all),
a Delete button in the run detail and a hover ✕ per sidebar row (tested e2e)
- CSS design system: tokenise the loose values into one scale — 8-step type
scale (was 10 ad-hoc sizes), radius/z-index/motion/scrim/terminal tokens,
fix an undefined var(--muted); 66 tokens, 0 loose font sizes, all var() resolve
- stale version labels 4.0.0/4.2.0 -> 4.2.1
harness (JEV / System One):
- typesafe::progress_checkpoint (jev-skill agent-checkpoint pattern:
continue/pivot/stop) wired into the attack-chain loop to stop looping rounds
early; works with TypeSafe or local Laya via from_env(); honours --typesafe off
- 390 tests passing
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
- new sarif module: projects findings to SARIF 2.1.0 (rules deduped by CWE,
security-severity from graded CVSS, endpoint locations, OWASP/MITRE tags)
- report::write_all/rebuild now emit report.sarif alongside md/json/html/pdf
- `neurosploit sarif <run> [--out]` re-emits on demand; exposed over MCP
- assurance: report.sarif added to the known-artifacts manifest
- chaining doctrine: harvest every object identifier (ids/UUIDs/tokens/emails)
into a reference pool and substitute across identities/endpoints — the core
of reliable BOLA/IDOR/mass-assignment discovery
- version 4.2.1; 389 tests passing
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
- Container image scanning: new `container` mode + 4 skills (vuln, secret,
misconfig, SBOM) driving trivy/grype/syft headless, read-only. Scans an OCI
ref / tar / Dockerfile for vulnerable packages (CVE/fixed-in/KEV), exposed
secrets in any layer, Dockerfile+runtime misconfig, and writes an SBOM in
both SPDX and CycloneDX to the run's sbom/. Also exposed as an MCP tool
(neurosploit_container).
- Coverage report: every run writes coverage.md — which agents ran (tested
surface), findings per agent, and the high-value classes NOT covered — so the
reader sees the engagement's reach. Added to the assurance bundle.
- Login-verification evidence: doctrine now requires capturing the login
request/response + a Playwright screenshot and recording success/failure
before authenticated testing.
- HTTP traffic export: `neurosploit traffic <run>` turns the intercepted
flows.jsonl into a traffic.http archive for external tools.
Not ported: Asset Discovery (enterprise-only, skipped per request).
383 tests.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The three fragile heuristics now get a calibrated second opinion when a decision
backend is configured. All go through TypeSafe::from_env(), so they work
identically with hosted TypeSafe or local Laya (--decision-backend), and the run
banner names the active backend. Deterministic behaviour is unchanged when no
backend is set or --typesafe off.
- typesafe.rs: three helpers — same_finding (Noul), response_origin (Choice) and
is_prompt_injection (Noul).
- Dedup grey zone (pipeline finish): a fixed 0.4 Jaccard cannot settle
near-duplicates; merge_grey_zone_dupes asks a calibrated Noul on every
same-endpoint/CWE pair scoring in the 0.25..0.40 band and merges the ones it
calls the same bug.
- WAF origin (poc.rs): header signatures are ambiguous; before dropping a PoC as
edge-answered, the backend gets the deciding vote — only bail if it also judges
p(application) < 0.5, so a real finding is not discarded on a false edge.
- Prompt-injection (pipeline probe): the keyword matcher over-flags legit pages
that merely mention "ignore instructions"; a calibrated Noul confirms real
manipulation before raising the neutralised-injection notice.
383 tests.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Models: xai:grok-4.7 added as the preferred xAI model (API via XAI_API_KEY at
api.x.ai/v1, and subscription via the local `grok` CLI) and to the OpenCode Zen
list. `--model xai:grok-4.7` or `--subscription --model xai:grok-4.7`.
PoC persistence hardened so exploitation artifacts are retrievable after a run:
the POCS doctrine now explicitly requires saving repro scripts, Frida hook
scripts (<slug>.frida.js), custom exploit code/source, compiled PoCs, request
collections and adapted public PoCs into the run's pocs/ folder with a run
command in the header — and the mobile mode now carries that directive too, so
Frida bypass/hook scripts written during APK/IPA analysis are kept and re-runnable.
383 tests.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
New `mobile` engagement mode: `neurosploit mobile <app.apk|app.ipa|binary>`
reverse-engineers a local artifact with a dedicated `mobile` agent set, all
headless and provisioned on demand (Ghidra analyzeHeadless, MobSF REST/Docker,
Frida, apktool/jadx, radare2).
Twelve original, generic skills (agents_md/mobile/, English): static binary
triage, APK static analysis, IPA static analysis, RASP & anti-tamper mapping,
root/jailbreak detection + bypass, TLS pinning detection + bypass, anti-debug
detection + bypass, obfuscation analysis & deobfuscation, code-integrity /
tamper-check bypass, hardcoded-secrets extraction, insecure local storage, and
mobile network traffic analysis. Findings are proven from the artifact
(decompilation or Frida trace), non-destructively.
- agents.rs: new `mobile` Library category (loaded, counted).
- pipeline.rs: run_mobile() mirroring the host pipeline with a mobile recon and
headless tooling doctrine; exported from the crate.
- CLI: `Cmd::Mobile` + `Mode::Mobile`, wired in main and the TUI.
- README + TUTORIAL document the new test type; engagement-modes badge + table
updated; "New in v4.2.0" note. Version bumped to 4.2.0 across the workspace.
383 tests.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
MCP server (app/src/mcp.rs): `neurosploit mcp` speaks Model Context Protocol
over stdio (JSON-RPC 2.0), exposing run / list_runs / findings / report /
rebuild / internal / compliance as tools. Each shells out to the same binary,
so scope, safety and authorization match the CLI. Install with
`claude mcp add neurosploit -- neurosploit mcp`. Handshake, tools/list and a
live call verified. TUTORIAL section 8 + README document setup for Claude Code,
Codex and Cursor.
Tooling doctrine expanded so the agent researches and provisions the BEST tool
for the context instead of being limited to a fixed list:
- context toolboxes (AD: netexec/impacket/bloodhound-python/certipy/kerbrute/
responder/evil-winrm; web recon; cloud; exploitation frameworks incl.
metasploit/msfvenom; cracking) — provision on demand.
- CVE -> PoC sourcing as a core capability: on a fingerprinted version
(WordPress/plugin/CMS/OS package/service) go to searchsploit, Exploit-DB,
GitHub, PacketStorm/Vulners, wpscan; clone/fetch, compile (gcc/go/cargo) and
run the PoC non-destructively, vetted and time-boxed.
- headless-only rule for GUI tools: mobsf (REST/Docker), ghidra analyzeHeadless,
jadx/apktool/frida, radare2 — never require an X display.
383 tests.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Laya (github.com/NandhaKishorM/laya) is the same System One abstraction as
TypeSafe — identical choice/score/noul primitives — but local, open-source
(Apache 2.0) and free. Added it as a swappable backend, entirely additively:
the hosted TypeSafe path is byte-for-byte unchanged (key alone → same endpoint,
model, bearer as before).
- typesafe.rs: endpoint/model/bearer are now instance fields with env overrides
(NEUROSPLOIT_DECISION_ENDPOINT / _MODEL). Defaults are the hosted TypeSafe API.
from_env() now also activates when a local endpoint is configured (no key).
backend_label() names the active backend in the run banner.
- tools/laya_shim.py: a stdlib HTTP shim that loads Laya and exposes the exact
POST /systemone contract the client already speaks. Model downloads on first
use (HF cache); no key; evidence stays on the box.
- CLI: --decision-backend typesafe|laya. `laya` installs laya if missing, starts
the shim, waits for readiness, and points the client at it — all optional,
only when the operator selects it.
383 tests; the hosted TypeSafe behaviour is untouched.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Re-ran the previously-missed scenarios on the current build without TypeSafe (A)
and with (B). Both arms now confirm CRLF-on-Location, second-order SQLi, UNION
SQLi, blind-time, IDOR and BOLA — the chaining/skill fixes are prompt-level, not
TypeSafe-gated. TypeSafe's contribution is the severity shape: it consolidates
A's long Low tail (10) into fewer, better-justified High findings (8 vs 3) and
keeps the credential-dump BOLA at Critical via data-type grading. Artifact grid
restored to A vs B·TS columns; both run arms stored.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Current-build run against the 13-scenario target with TypeSafe on: every seeded
class confirmed with a live receipt, chained beyond the set into full admin
takeover, GraphQL authz bypass, a config secret leak and an authenticated RCE.
The credential-dump BOLA holds Critical because severity is graded on the kind
of data exposed, not the class. report.html + run artifacts + README refreshed;
no secrets committed.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The 13-target benchmark left 3 misses. Root-caused and fixed the two that were
coverage gaps (the third was single-run variance, already handled by the
session-limit fix):
- CRLF header injection (web_crlf_header_go): the agent confirmed the open
redirect on /go?url= and stopped; the CRLF payload was never generated. The
open_redirect skill now tests %0d%0a header injection on the SAME param, and
CHAIN_DOCTRINE says a param landing in a Location header must also be tested
for response splitting. chain.rs: CWE-113/93/644 now provide capabilities;
attack_graph maps their kill-chain stage.
- Second-order SQLi (web_sqli_second_order): the sink was behind /admin, which
the customer account could not reach. CHAIN_DOCTRINE now teaches the
precondition pattern (store the payload, trigger from every identity, escalate
first if the trigger page needs a role you lack, else report as a chained
lead). chain.rs: CWE-564 requires PrivilegedContext so it chains after privesc.
BENCHMARK.md: added the TypeSafe calibrated-adjudication row; dropped the
"genuinely ahead" prose (the table is the summary); condensed the rest
188 -> 89 lines; refreshed scale (27 validators, 47 modules, 383 tests).
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Addresses the benchmark's honest edge (a genuine BOLA credential dump graded
Low because evidence_data was null). Two fixes so criticals like it are not
recalibrated away:
- attack_graph::backfill_evidence — when evidence_data is null but the agent
recorded a proof in prose, copy that text into the structured slot the grader
reads (no fabrication, just relocation). Called first in enrich().
- attack_graph::data_class — classifies the demonstrated data (none/data/
sensitive) by scanning every evidence slot for credential/key/PII/payment
signatures. cvss_graded now grants the confidentiality receipt when sensitive
data was shown, even on a thin structured receipt — the KIND of data is itself
the impact.
- TypeSafe adjudication adds a `data_sensitivity` Score (public → PII → secrets),
carried on Adjudication. The pipeline regrade only strips impact when the
model was unconvinced AND no sensitive data was shown AND data_sensitivity is
low; a demonstrated credential/PII exposure keeps its severity.
articles/ — LinkedIn article (PT, no em-dashes) in Markdown + DOCX: explains
TypeSafe/System One/Jev, NeuroSploit, how to configure TypeSafe, the step-by-step
benchmark, results, the refinements this forced, and offensive-security use cases.
383 tests.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Version bumped to 4.1.0 across the workspace, binaries, web console and Typst
template.
README: new "New in v4.1.0" summary; trimmed the verbose highlight bullets and
the TypeSafe section; removed the anti-plagiarism/provenance section (provenance
stays in the code, just not front-and-centre in the README); TypeSafe promoted
to its own top-level section; agent count 446.
TUTORIAL: new section 17 "Assurance & authorization" covering the target gate,
--scope-file, evidence-graded CVSS, audit anchoring + assurance bundle, sandbox,
intercept proxy, PoC re-validation, compliance mapping, TypeSafe, and the
internal/AD graph + budget governor.
benchmarks/typesafe-2026-09-20/: the with/without TypeSafe measurement —
report.html, scorer, both runs' findings/assurance/meta/logs, and a README.
No secrets committed (env-only during the runs, verified clean).
381 tests.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Subscription CLIs (claude) report a hit session limit as ordinary stdout with a
ZERO exit code — 'You've hit your session limit · resets …'. Left as Ok it
became a 'response' each agent then failed to parse, and the run burned every
remaining agent against a dead session instead of pausing. Now the sentinel is
caught (length-guarded so a real finding mentioning 'rate limit' is not misread)
and surfaced as exhaustion, so the pool parks the run for /continue — the
pause-on-quota path that already existed but this case never reached.
Found during a live benchmark when run B collapsed to 0 findings mid-run.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
127.0.0.1/localhost/::1 always mean this host regardless of any VPN, so the
network-position ambiguity the gate guards against does not apply. Needed to
run against a local benchmark target.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
pomdp.rs was dead (belief.rs was on the path, pomdp wasn't). Now
pomdp::may_assert runs as an anti-hallucination gate over the belief WorldModel:
a finding asserted confirmed while the belief about it is diffuse or weak is
held for review. Advisory — never deletes.
inbox.rs was unused (mail.tm happened via agent prompt instructions). Now when
temp-email is enabled the HARNESS creates the mail.tm inbox via crate::inbox and
hands the agent that exact address, so the inbox is one we own and record rather
than an unrecorded one the agent conjures. Best-effort; falls back to the old
prompt path on failure.
Every module is now on a real runtime path. 381 tests.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
TypeSafe cannot BE an LLM agent — System One does not generate text or call
tools. But it can be the decision brain of a code-owned confirmation loop, and
that is what typesafe_agent.rs is: an ADDITIONAL confirmation strategy.
typesafe_agent.rs — for enumerable classes (XSS, SQLi, open-redirect, path
traversal, SSRF, IDOR): code lists candidate payloads, a TypeSafe Choice picks
the next one given what's been tried, the replay engine sends it for real, a
TypeSafe Noul judges the response, loop until confirmed or exhausted. Edge/WAF
answers are refused. Pure parts (class table, payload templating, id-swap, OAST
substitution, query encoding) are unit-tested; the networked loop is integration.
Wired as a pipeline pass that runs ONLY on findings the LLM path left
unconfirmed or in needs-review (the recall lever) — it can raise a finding to
confirmed with a calibrated probability, never downgrades (the deterministic
layer owns that).
--typesafe on|off|auto (global flag) resolves into the env the pipeline reads,
governing adjudication, CVSS re-grade, agent pruning and this loop together.
`off` runs the identical pipeline without TypeSafe; meta.json records
"typesafe": true|false so a with/without pair is a clean A/B measurement. Web
console gets the same toggle.
381 tests.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Honesty audit found four modules written but not on the runtime path. Two mattered
and are now wired; two are noted.
cvss.rs — was NOT called; finding.cvss came from the old attack_graph ladder.
Now attack_graph::cvss_graded() bridges the class shape + demonstrated rung into
crate::cvss::grade (the FIRST-verbatim v3.1 equation), and enrich() sets
finding.cvss from the demonstrated vector, recording the potential ceiling in
the impact text. The class ladder remains only as a fallback for findings with
no evidence to grade.
waf.rs — the deterministic classifier was NOT run on any real exchange (only
WAF_OPS prompt text reached the agent). Now poc.rs classifies each re-run: a PoC
answered by a WAF/CDN is Unverifiable, not "gone" — closing a false-demotion
where an edge block looked like a fix.
TypeSafe (System One) extended per the build-with docs:
- CVSS via System One: when impact_demonstrated < 0.5, the finding's CVSS is
re-graded with impact receipts stripped — the calibrated judgment, not just
the rung, decides the demonstrated number.
- Agent selection: typesafe_prune_agents() asks one batched request (the
fan-out pattern), a Noul per chosen agent, and drops only those it calibrates
as clearly irrelevant (p < 0.25), never prunes to empty. Additive over the
LLM selection; skipped without a key.
Still shelf-ware, flagged honestly (not wired): inbox.rs (mail.tm/SMS happens
via agent prompt instructions, the Rust client is unused) and pomdp.rs
(redundant — belief.rs is the one on the path).
374 tests.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Integrates TypeSafe's System One model (Jev) as an optional, calibrated
adjudicator — the RLCD (Reinforcement Learning for Calibrated Decisions) tier:
typed judgments with probabilities where the harness needs a number, not prose.
- typesafe.rs: HTTP client for POST /v1/systemone (Bearer TYPESAFE_API_KEY,
model jev-latest), with Choice/Noul/Score primitives, retry on 429/529, and
parsed answers exposing the probability distribution + confidence.
adjudicate() asks a Choice {confirmed/needs-review/rejected} plus an
impact-demonstrated Noul over a finding; calibrated_confidence() folds
demonstrated impact into the number, wants_review() gates a split distribution.
- pipeline: an optional pass (runs when TYPESAFE_API_KEY is set, off with
NEUROSPLOIT_TYPESAFE=off) adjudicates each finding over its EVIDENCE — never
its narrative — refining confidence and the needs-review boundary. Additive:
a deterministic validator still rules; TypeSafe can only lower confidence or
flag for review, never resurrect a rejected claim. Audited per finding.
- env.example + README document it; the web console inherits the key via env.
Where the model stack maps in NeuroSploit: LM/BERT ≈ the deterministic
validators (no model), RLHF chat ≈ the exploit/recon agents, RLVR reasoning ≈
the DeepReasoning budget tier, RLCD ≈ this calibrated adjudication.
373 tests (+5).
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>