A finding is useful only if the reader can find the problem, see why it
matters, fix it, and reproduce it without trusting us. The report answered the
last one badly and the other three not at all: it printed a payload blob and an
evidence blob, and "payload: ' OR 1=1--" tells a developer nothing about WHERE
to look. A PoC script attached as a file is a black box unless you run it.
Findings are now rendered in the order a reader works through them — where the
problem is, what it means, how to fix it, then the proof — in the HTML report,
the Markdown, the Typst/PDF and the web console's finding modal.
The proof is numbered, pasteable steps: baseline request, attack request, how
to read the result, with the real URL and the real payload. They come from the
agent's `repro_steps` when it recorded them, and are derived from the
endpoint/payload/identity pair otherwise, so every finding carries something
runnable. The generated curl redacts Authorization/Cookie/API-key headers — a
report gets shared, and a live session cookie inside one is a new bug. A PoC
script is now offered as an extra artifact that automates the steps, never as
the proof itself.
Technical evidence is the measured difference, not a paraphrase: baseline vs
attack status, size, timing and delta; how many repeats reproduced it; the
controlled marker and whether a browser or a callback observed it; then each
recorded exchange with the headers that decide a class (Location, Set-Cookie,
Access-Control-*, X-Frame-Options, CSP, Retry-After) and a body excerpt.
Finding gains `location` (the parameter/field/flow step, not just the URL) and
`repro_steps`, and the agent contract now asks for them explicitly, along with
impact tied to this app's data and remediation that names the control rather
than saying "sanitise input".
The web console offers the run's PDF when Typst produced one — and only then,
since a dead download button is worse than none.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* feat: embed proof screenshots in reports, correlated to findings
Define a convention that ties each proof image to its vulnerability and
renders it in every report format.
- Finding gains `screenshots: Vec<String>` (paths relative to the run
workdir, e.g. evidence/<finding-id>-1.png).
- Exploit prompt injects an EVIDENCE SCREENSHOTS doctrine: agents save
proof PNGs into the run's absolute evidence/ dir named by a vuln slug,
and list them in the finding JSON `screenshots` array.
- collect_evidence() resolves whatever the agent captured (absolute,
workdir-relative, evidence/, /tmp basename), copies it to a stable
evidence/<finding-id>-N.png, and rewrites the field; unresolved refs
are dropped so a report never embeds a missing image.
- Typst (image()), HTML (<img>) and Markdown (![]) render each finding's
screenshots beside its evidence.
Tests: slugify + collect_evidence resolution/rename; verified a real PDF
compiles with an embedded image.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018BGLy4j5qsqqid6CoovowC
* feat: source-able env.sh to activate neurosploit in the current shell
Add env.sh: `source` it to export NEUROSPLOIT (binary path),
NEUROSPLOIT_BASE (agents base) and prepend the binary dir to PATH —
no reinstall or new terminal needed. Auto-detects the install/repo dir,
honors NEUROSPLOIT_DIR, idempotent. setup.sh now writes a ready env.sh
into the install dir and points users at `source <dir>/env.sh`.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018BGLy4j5qsqqid6CoovowC
---------
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Partial observability is now first-class:
- belief.rs — property-graph world model; nodes (host/service/vuln/exploit/cred)
carry a probability, not a boolean. Bayesian observation updates; per-node
Shannon entropy; mean-uncertainty + recon-frontier. Black-box = diffuse priors
that sharpen with observation; white-box collapses toward deterministic (MDP).
- pomdp.rs — value_of_information(), decide() (recon vs exploit falls out of
belief entropy), and may_assert() — the mathematical anti-hallucination gate:
no exploitability claim while the belief is diffuse (high entropy) → observe first.
- grounding.rs — verification engine, hard rule "no claim without a tool receipt":
empirical grounding for black-box (raw HTTP/OOB/error markers), symbolic for
white-box (file:line into reviewed source). Ungrounded claims demoted + flagged
receipt_missing (feeds future reward shaping).
- pipeline.finish(): grounding gate before reporting + belief-uncertainty readout.
- bump 3.5.0 → 3.5.1; README documents the v3.5.1 belief/grounding architecture
and the infra/bandit/reward roadmap.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
- New persistent interactive session (app/src/repl.rs), launched when run with no args:
banner, model selection, API-key config (/key) or subscription (/sub), then a live
session to set /target, /repo, /auth, and free-text /focus instructions (or just type
them) that STEER which agents run and how.
- Slash-commands: /model /providers /key /sub /target /repo /auth /focus /mcp /votes
/agents /show /run /quit (+ bare text = focus).
- RunConfig gains `instructions` and `auth`:
* instructions bias both LLM agent-selection and the heuristic (focus keywords →
injection/access-control/etc. agents get a strong boost)
* operator directives (focus + auth) injected into recon and exploit prompts so agents
test as an authenticated user and prioritise the requested vuln classes
- bump 3.4.1 → 3.5.0 (CLI, harness, reports, credits)
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
- Task-based model ROUTER (recon/select prefer a fast model; exploit prefers primary; validate uses a different model than the finder)
- ReAct doctrine injected into exploit prompts (Thought→Action→Observation, token-efficient)
- Dedup: unique agents per run + findings deduped by CWE/endpoint/title (highest confidence kept)
- Token economy: recon blob capped for selector + per-agent context
- Configurable MCP: merge user mcp.servers.json into the pipeline's .mcp.json
- +54 white-box/code-analysis agents (NoSQLi, LDAP/XPath, JWT-none, Java/.NET/PHP/Go/Node/Python
specifics, SSTI, ReDoS, deserialization, etc.) → 303 agents total (78 code)
- Credits: Joas A Santos & Red Team Leaders (CLI banner, interactive header, HTML+Typst report)
- README: GitHub stars/forks badges, 60-second quick start, full API config steps, intuitive layout
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>