New `mobile` engagement mode: `neurosploit mobile <app.apk|app.ipa|binary>`
reverse-engineers a local artifact with a dedicated `mobile` agent set, all
headless and provisioned on demand (Ghidra analyzeHeadless, MobSF REST/Docker,
Frida, apktool/jadx, radare2).
Twelve original, generic skills (agents_md/mobile/, English): static binary
triage, APK static analysis, IPA static analysis, RASP & anti-tamper mapping,
root/jailbreak detection + bypass, TLS pinning detection + bypass, anti-debug
detection + bypass, obfuscation analysis & deobfuscation, code-integrity /
tamper-check bypass, hardcoded-secrets extraction, insecure local storage, and
mobile network traffic analysis. Findings are proven from the artifact
(decompilation or Frida trace), non-destructively.
- agents.rs: new `mobile` Library category (loaded, counted).
- pipeline.rs: run_mobile() mirroring the host pipeline with a mobile recon and
headless tooling doctrine; exported from the crate.
- CLI: `Cmd::Mobile` + `Mode::Mobile`, wired in main and the TUI.
- README + TUTORIAL document the new test type; engagement-modes badge + table
updated; "New in v4.2.0" note. Version bumped to 4.2.0 across the workspace.
383 tests.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Version bumped to 4.1.0 across the workspace, binaries, web console and Typst
template.
README: new "New in v4.1.0" summary; trimmed the verbose highlight bullets and
the TypeSafe section; removed the anti-plagiarism/provenance section (provenance
stays in the code, just not front-and-centre in the README); TypeSafe promoted
to its own top-level section; agent count 446.
TUTORIAL: new section 17 "Assurance & authorization" covering the target gate,
--scope-file, evidence-graded CVSS, audit anchoring + assurance bundle, sandbox,
intercept proxy, PoC re-validation, compliance mapping, TypeSafe, and the
internal/AD graph + budget governor.
benchmarks/typesafe-2026-09-20/: the with/without TypeSafe measurement —
report.html, scorer, both runs' findings/assurance/meta/logs, and a README.
No secrets committed (env-only during the runs, verified clean).
381 tests.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Closes the three benchmark gaps and adds the two the user asked for.
poc.rs — re-runs each finding's recorded proof and sorts it into reproduced /
changed / gone / unverifiable. The last two are kept apart deliberately: a PoC
that could not be tested (out of scope now, state-changing, nothing recorded)
is never reported as one that failed. Never re-runs a mutating request to
"confirm" it. Can only lower a finding's standing, never raise it. Wired as a
run pass (--revalidate-poc) and a subcommand (neurosploit poc <run> --apply).
proxy.rs — own recording forward proxy (HTTP in full; HTTPS tunnelled with
honest metadata, no fake CA) that chains upstream to Burp / Caido / ZAP /
mitmproxy. A bare tool routes straight through it; own+tool records here and
forwards for full TLS interception. Flows -> flows.jsonl, distinct hosts become
passive-discovery leads. Harness and agent child commands share one route.
sandbox.rs — Kali docker/podman container: no host network, no mounted socket,
no-new-privileges, workdir mounted, proxy/transport env inherited. A missing
runtime is an explicit error, never a silent fallback to host execution — the
whole point being to keep attack payloads off the operator's host. Subcommands
sandbox up|exec|install|down.
compliance.rs — maps confirmed findings onto PCI-DSS v4.0, HIPAA Security Rule
and SOC 2 controls. Phrased as "bears on control X", never "compliant/non-
compliant"; the disclaimer is rendered on top and absence of a finding is never
presented as compliance. Report section + `neurosploit compliance <run>`.
validation.rs — 8 new deterministic validators (19 -> 27 classes): verbose
errors/stack traces (CWE-209), cleartext/HSTS (319), CRLF response splitting
(113), dangerous HTTP methods (650), GraphQL introspection, exposed backup
files (530), Host header injection (644), cacheable private responses (525).
Each names exactly what it saw and rejects the classic false positives (a
block page echoing a payload, the SPA served under a bogus path, a copyright
year mistaken for a code).
All wired through RunConfig, the CLI (global --intercept/--sandbox; run-level
--revalidate-poc/--compliance) and the web console's Tooling & assurance block.
328 tests.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Budget (opt-in, unlimited by default so an un-budgeted run is unchanged):
- crates/harness/src/budget.rs — modes, phase shares, Token Governor
- CLI: --budget/--token-limit/--deep-test-limit/--coverage-first/
--depth-first/--sample-per-route; same controls in the web wizard
- pipeline honours it: vote_n narrows, evidence rounds are capped
Run control parity in the web console:
- /pause in the REPL, backed by a pause gate in the model pool: in-flight
agents finish, then the run holds with every finding kept
- POST /api/exploit/:id/{pause,continue,report} + GET .../log
Provenance (crates/harness/src/provenance.rs):
- JOASNSCOPE sigil leads every canary, so a marker found in a response,
a log or someone else's report extracts whole and names its build
- per-build fingerprint, per-run id, optional per-customer build id
- findings.json stamped with _engine; signed provenance.json manifest
- structural signature survives rewording but not a changed result set
- prompts watermarked at the single pool chokepoint
- `neurosploit provenance show|scan|verify`
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The Arena engagement produced 24 findings, a graph with 83 edges, and exactly
one chain edge. Two defects, both a step nobody was doing rather than a model
reasoning badly.
chains_from came back empty on every finding. An agent works one vulnerability
and has no view of what the other twelve found, so asking it to link its result
to findings it never saw was asking for something it cannot know. Chaining now
happens after the whole set is visible, on rules about ENABLEMENT: what one
weakness yields that another needs. Account enumeration yields valid
identities; absent throttling turns them into unlimited guesses; a permissive
password policy makes the guessing land. None is severe alone, and that
sequence is how accounts get taken over — on the real data it now reads
CWE-307 <- CWE-204, CWE-208 and CWE-614 <- CWE-319.
The CWE->stage fallback sent 23 of 24 findings to initial-access, so the kill
chain had one populated column and drew a star. Enumeration and side channels
are discovery; missing throttling, password policy, cookie flags and session
fixation are credential-access; hardening headers are recon. The same run now
spreads across credential-access 13, discovery 6, initial-access 5.
Two bugs the tests and the real data caught:
- CWE-614 both yields session material and needs it, so a class chained to
itself: duplicates formed a circular "attack path" from a cookie flag to the
same cookie flag. A weakness class no longer enables itself.
- apply_links only fills an empty chains_from, which is right for asserted
chains and wrong for links written by an older version of these rules — a
report kept the circular link through two rebuilds because nothing was
allowed to touch it. repair() now drops links that cannot be true whoever
wrote them: self-references, same-class links, dangling ids, cross-host
links.
enrich() still only fills empty fields during a run (an agent's judgement
should survive), but a rebuild applies the current mappings via remap_stages —
otherwise a finished run is frozen with whatever taxonomy existed that day.
path_for() gives the per-vulnerability view: what precedes this finding, what
it enables, and the narrative to print beside it.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Driven by the Arena Hockey engagement, where all 24 findings shipped with an
empty CVSS field and the PDF ran payloads off the page edge.
CVSS
- Derived deterministically from what the harness knows: the weakness class
sets the impact shape, the PROVEN exploitability sets attack complexity, and
the auth context sets privileges required. The vector is emitted with the
score, because a score without its vector cannot be checked and an unchecked
score is just a bigger adjective.
- The base equation is the v3.1 specification verbatim, including round-up and
the scope-changed privileges table. Tests anchor it against known values
(9.8 unauthenticated RCE, 10.0 with scope change, 6.1 reflected XSS, 0.0 for
no impact).
- An unknown weakness stays conservative — guessing high impact from a class
nobody mapped is how reports get inflated. An agent-supplied score is never
overwritten.
PDF
- Steps are passed as an ARRAY and rendered as a real numbered list, one command
per box. The previous template flattened them into a single `raw` block, which
rendered five separate commands as one run-on paragraph.
- Finding blocks are breakable, so a long evidence dump flows to the next page
instead of off the bottom of this one.
- `wrappable()` inserts zero-width breaks so encoded payloads wrap. The first
attempt broke prose mid-word ("rota ted", "lockoutOnFailu re=false") by
breaking every N characters regardless of context; it now works per token and
leaves anything that fits on a line exactly as it was.
- rebuild() re-enriches before rendering, so a run that finished before a
mapping existed picks it up instead of reprinting the gap forever.
Over-claimed findings are capped, not deleted
- The engagement rejected "no rate limiting on the password-reset flow" because
the agent claimed inbox flooding and only proved 25 unthrottled requests. The
claim was inflated; the measurement was real, and dropping it hid a genuine
gap. A unanimously rejected finding that still carries a checkable receipt is
now capped to Low and flagged for review, with the validator's reason
attached — the reader gets the fact without the story built on it.
- The agent contract now says impact must be what was MEASURED, and warns that
inflating it costs the whole finding.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
A PDF was only ever produced while a run was finishing. If `typst` was missing
at that moment — or the template improved afterwards — the operator had no way
to get one without re-running the whole engagement against the target.
report::rebuild() regenerates every artifact (md · json · html · pdf) from the
findings already on disk, exposed as `neurosploit rebuild <run-id|dir>` and as
POST /api/runs/:id/report with a "Generate report" button in the run view. The
endpoint shells out to the harness rather than reimplementing report generation
in JavaScript, so there is one implementation instead of two that drift, and it
says plainly when the PDF was skipped for want of `typst` instead of handing
back a link to a file that was never produced.
Also fixes write_all() to pass the run's pocs/ listing into the HTML report, so
a rebuilt report links the scripts each finding cites — the run-time path
already did this and the rebuild path silently did not.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
A finding is useful only if the reader can find the problem, see why it
matters, fix it, and reproduce it without trusting us. The report answered the
last one badly and the other three not at all: it printed a payload blob and an
evidence blob, and "payload: ' OR 1=1--" tells a developer nothing about WHERE
to look. A PoC script attached as a file is a black box unless you run it.
Findings are now rendered in the order a reader works through them — where the
problem is, what it means, how to fix it, then the proof — in the HTML report,
the Markdown, the Typst/PDF and the web console's finding modal.
The proof is numbered, pasteable steps: baseline request, attack request, how
to read the result, with the real URL and the real payload. They come from the
agent's `repro_steps` when it recorded them, and are derived from the
endpoint/payload/identity pair otherwise, so every finding carries something
runnable. The generated curl redacts Authorization/Cookie/API-key headers — a
report gets shared, and a live session cookie inside one is a new bug. A PoC
script is now offered as an extra artifact that automates the steps, never as
the proof itself.
Technical evidence is the measured difference, not a paraphrase: baseline vs
attack status, size, timing and delta; how many repeats reproduced it; the
controlled marker and whether a browser or a callback observed it; then each
recorded exchange with the headers that decide a class (Location, Set-Cookie,
Access-Control-*, X-Frame-Options, CSP, Retry-After) and a body excerpt.
Finding gains `location` (the parameter/field/flow step, not just the URL) and
`repro_steps`, and the agent contract now asks for them explicitly, along with
impact tied to this app's data and remediation that names the control rather
than saying "sanitise input".
The web console offers the run's PDF when Typst produced one — and only then,
since a dead download button is worse than none.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
/only <agent,...> (app/src/repl.rs): the REPL had no way to pin an exact
agent set the way the CLI's --only flag does - Session gained a `pinned`
field, wired into RunConfig in both start_background() (the live
background-run path) and the blocking run() fallback. Needed so the web
console's exploitation jobs can drive a real interactive REPL session
(for live input while a run streams) without losing lead-pinning, which
only existed as a CLI flag until now. Also usable directly from a
terminal REPL session.
HTML report (crates/harness/src/report.rs, html()): rebuilt to match the
Typst PDF template's design (templates/report.typ) instead of its own
inconsistent styling - violet brand accent, an asset table, a 5-box
executive-summary grid (all severities, zero-count included, matching
Typst's grid exactly), a Vulnerability Summary table, and severity-
left-bordered finding cards with a compact field grid (Criticality /
Status / OWASP-CWE / Confidence / Location / Agent / Auth context) before
Description-Impact / Proof of Concept / Evidence / Remediation - same
field order and labels as the Typst template. Dropped the Mermaid
attack-path/kill-chain section entirely (the web console's live
Generative Attack Path Chaining graph covers that now, interactively).
Also tidied two pre-existing formatting quirks while in there: OWASP/CWE
left a dangling " · " when CWE was empty, and the confidence cell said
"<votes-string> votes" even when the votes field already contained a
compound descriptor like "1/1 · receipt_missing".
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0129WdYHccPsH27k5GGuwijd
Version bump 3.6.5 -> 3.6.6.
Local & uncensored models
- New `llamacpp:` provider (llama-server, OpenAI-compatible, localhost:8080,
no API key, CPU-only or GPU-offloaded). Override via LLAMACPP_BASE_URL;
model name is the loaded gguf (pass-through). 15 -> 16 providers.
- README: local/uncensored highlight, provider table + key-less note, badges.
Quality
- clippy clean under `-D warnings`: clamp(), sort_by_key(Reverse), struct-literal
init, too_many_arguments allows, scoped await_holding_lock on the REPL blocking
fallback (guard intentionally held across run().await), plus clippy --fix set.
CI
- examples/github-actions/ci.yml: cargo build/test/clippy -D warnings for the
neurosploit-rs workspace (template, kept out of .github/workflows).
Claude-Session: https://claude.ai/code/session_01QDses7zTSa9YF7pPRjphvh
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
* feat: embed proof screenshots in reports, correlated to findings
Define a convention that ties each proof image to its vulnerability and
renders it in every report format.
- Finding gains `screenshots: Vec<String>` (paths relative to the run
workdir, e.g. evidence/<finding-id>-1.png).
- Exploit prompt injects an EVIDENCE SCREENSHOTS doctrine: agents save
proof PNGs into the run's absolute evidence/ dir named by a vuln slug,
and list them in the finding JSON `screenshots` array.
- collect_evidence() resolves whatever the agent captured (absolute,
workdir-relative, evidence/, /tmp basename), copies it to a stable
evidence/<finding-id>-N.png, and rewrites the field; unresolved refs
are dropped so a report never embeds a missing image.
- Typst (image()), HTML (<img>) and Markdown (![]) render each finding's
screenshots beside its evidence.
Tests: slugify + collect_evidence resolution/rename; verified a real PDF
compiles with an embedded image.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018BGLy4j5qsqqid6CoovowC
* feat: source-able env.sh to activate neurosploit in the current shell
Add env.sh: `source` it to export NEUROSPLOIT (binary path),
NEUROSPLOIT_BASE (agents base) and prepend the binary dir to PATH —
no reinstall or new terminal needed. Auto-detects the install/repo dir,
honors NEUROSPLOIT_DIR, idempotent. setup.sh now writes a ready env.sh
into the install dir and points users at `source <dir>/env.sh`.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018BGLy4j5qsqqid6CoovowC
---------
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
- New agent account_registration_and_forms (+1 → 430): analyzes the app's forms
and self-registers a benign test account (curl or Playwright) to reach the
authenticated surface when no creds are given.
- Probe extracts form details (action/method/fields/kind/CSRF) so form analysis is
grounded; shown in the probe summary and recon JSON.
- Hard anti-flood guardrail in SAFETY_DOCTRINE + the agent: at most 2 accounts per
engagement, never loop/script/batch the register endpoint or flood the DB; reuse
the account made; a test needing many sign-ups is a lead, not mass-creation.
- Credential vault: engagement_ops directive tells agents to append created
accounts to <run-dir>/vault.jsonl; finish() consolidates to vault.json, masks
secrets in the report, and adds a 'Test accounts created (DELETE after)' cleanup
finding listing each account and how it was created.
- Finding tagging: new auth_context (authenticated/unauthenticated) and account
fields, rendered per-finding in the HTML report.
- Opt-in disposable email (off by default): /tempmail on + RunConfig.temp_email;
agents may use the free mail.tm API to read a registration confirmation code.
- Tests: parse_forms unit tests; docs updated (README/TUTORIAL/RELEASE), counts 430.
- Add 12 technique/scenario LLM red-team agents (AI category 18 → 30, total 429):
jailbreaks — AdvPrefix, PAIR, TAP, Crescendo, many-shot, persona/DAN,
encoding/obfuscation, refusal-suppression; prompt-injection scenarios — direct,
indirect (RAG/web/email/tool output), goal hijacking, tool/function-call abuse,
system-prompt/secret exfiltration. Each runs an attacker→LLM-judge loop
(baseline refusal → technique across variants → verdict), proving the bypass
with a benign, redacted receipt. Generated by scripts/build_llm_redteam_v365.py.
- Add REDTEAM_DOCTRINE and inject it into run_ai so every AI test follows the
baseline→technique→judge method across scenarios.
- Models: add Claude Opus 5 and Sonnet 5 (Anthropic) and a new Moonshot AI (Kimi)
provider with Kimi K3/K2 (moonshot:kimi-k3, MOONSHOT_API_KEY) — 15 providers.
- Docs: README/TUTORIAL/RELEASE — new AI/LLM red-team engagement mode + section,
model/env-key tables, agent-library counts (429), badges.
Also includes the v3.6.4 grounding fix (#33) landing on main.
The grounding gate ran in empirical mode for every engagement, demoting
white-box (and skills/n8n audit) findings that had passed the n-model vote
because a file:line code citation isn't raw tool output. Grounding is now
mode-aware:
- Symbolic (white-box SAST / skills): a file:line reference into the reviewed
source, or a quote of code present in it, is the receipt — no live target.
- Empirical (black-box / host / AI): evidence must resemble tool output (as before).
- Either (grey-box): a source citation OR a tool receipt grounds a finding.
The symbolic check runs against the reviewed source corpus (not the transcript)
and falls back to a structural file:line + quote check when the corpus is
unavailable. Adds unit tests incl. a regression test for #33.
- /continue (and /resume) now relaunch a recovered interrupted run on the same
target, carrying its findings forward and steering agents to widen coverage /
chain from them instead of re-reporting. Offer shown at launch; a fresh /run
supersedes it. Findings merge (dedup by title+endpoint) across both runs.
- Opening /results, /finding or /report while a run streams no longer corrupts
the terminal: live background output is paused for the picker (still captured
in /logs) and restored on exit, so Ctrl-C in a picker can't take the process
down mid-run.
- Drive `codex exec --json` and parse its JSONL event stream into the same
categorized live feed as Claude Code (exec/edit/tool/net/tokens), so recon and
exploitation are visible as each command runs instead of a silent black box.
- Fix the activity feed to keep per-agent tool events (commands, network, files,
findings) and only filter model reasoning + token telemetry, so /logs shows the
real command trail and /status 'last:' is a true sign-of-life.
- Surface failed internal commands as 'exec: (exit N)'; keep Codex auth/rate
detection from stderr.
Added the OpenAI GPT-5.6 line to the provider pool: gpt-5.6-sol (frontier/default),
gpt-5.6-terra (balanced), gpt-5.6-luna (fast/affordable). Version 3.6.0 -> 3.6.1.
New harness module `integrations` (+ app commands) wiring NeuroSploit into the
SDLC. Config persists per-project to .neurosploit/integrations.json; secrets are
NEVER stored — only the env-var name is saved, values read from the environment.
GitHub:
- private-repo clone (token injected into the clone URL for whitebox/greybox/tui)
- `neurosploit pr <owner/repo> <n>`: clone the PR head (refs/pull/N/head),
white-box review, optional `--comment` (PR summary) and `--jira` (cards)
- `neurosploit watch <owner/repo> --branch --interval`: re-review on each new commit
GitLab:
- private-repo clone (oauth2 token) for whitebox/greybox (gitlab.com or self-hosted)
Jira:
- `--jira` on any engagement opens one card per finding (REST /issue, basic auth)
Control:
- `/integrations` (REPL): show · enable/disable · setup jira|gitlab|github
- `neurosploit integrations [show|enable|disable] [github|gitlab|jira]` (CLI)
Docs: README "Integrations" section + new TUTORIAL-INTEGRATION.md (per-tool setup,
scopes, recipes, troubleshooting). Version bumped 3.5.2 → 3.5.3.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Distilled from reviewing real AI-pentest output that kept stopping at "exposed"
instead of "exploited". Pure-additive, back-compatible.
Behavior (injected into black/grey/chain exploit prompts via DEPTH_DOCTRINE):
- Exposed → exploited: any info-disclosure / exposed service/WSDL / leaked
credential|token / reachable dev host MUST be used before it's a finding;
otherwise it's a lead, not a confirmed High/Critical.
- Chain across modules: reuse obtained session/JWT/cookie/credential and pivot
to IDOR/privesc/exfil; report the chain, not isolated parts.
- Decode & fingerprint → CVE; audit tokens (alg-confusion/none/kid/JWKS, weak
HS256 secret cracking, lifecycle).
Deterministic post-pass (new crates/harness/src/hygiene.rs, wired into finish()):
- calibrate severity to PROVEN impact — unproven High/Critical (hedged, no
payload, thin evidence) capped to Medium and re-titled "(potential)";
- depth_audit — flag exposures on a host with no real exploit;
- hygiene_summary — advise consolidating hygiene classes repeated across assets.
Unit tests cover calibration + depth audit.
5 new doctrine meta-agents (scripts/build_methodology_v352.py → agents_md/meta/):
exploit_depth_doctrine, finding_chainer, artifact_decoder, token_auditor,
report_calibrator (meta 17→22, total 343→348).
Version bumped 3.5.1 → 3.5.2 across crates/app/installers/docs; RELEASE/README
updated.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Partial observability is now first-class:
- belief.rs — property-graph world model; nodes (host/service/vuln/exploit/cred)
carry a probability, not a boolean. Bayesian observation updates; per-node
Shannon entropy; mean-uncertainty + recon-frontier. Black-box = diffuse priors
that sharpen with observation; white-box collapses toward deterministic (MDP).
- pomdp.rs — value_of_information(), decide() (recon vs exploit falls out of
belief entropy), and may_assert() — the mathematical anti-hallucination gate:
no exploitability claim while the belief is diffuse (high entropy) → observe first.
- grounding.rs — verification engine, hard rule "no claim without a tool receipt":
empirical grounding for black-box (raw HTTP/OOB/error markers), symbolic for
white-box (file:line into reviewed source). Ungrounded claims demoted + flagged
receipt_missing (feeds future reward shaping).
- pipeline.finish(): grounding gate before reporting + belief-uncertainty readout.
- bump 3.5.0 → 3.5.1; README documents the v3.5.1 belief/grounding architecture
and the infra/bandit/reward roadmap.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
- Finding enriched with owasp / mitre / kill-chain stage / exploitability /
business_impact / chains_from (attack-path edges).
- attack_graph module: derive OWASP Top 10 + MITRE ATT&CK technique + kill-chain
stage from CWE (heuristic, no extra model call); render a Mermaid attack-path
flowchart (findings grouped by stage, explicit + implicit edges) and an ASCII
kill chain for the REPL.
- enrich() runs in finish() for every engagement.
- HTML report gains an "Attack Path & Kill Chain" section (Mermaid via CDN, dark)
plus a stage/sev/OWASP/MITRE/exploitability table.
- REPL print_findings shows the ASCII kill-chain + severity summary after a run.
- models: add GPT-5.5, GPT-5.4, GPT-5.4-mini, GPT-5.3-codex, GPT-5.2.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
- New persistent interactive session (app/src/repl.rs), launched when run with no args:
banner, model selection, API-key config (/key) or subscription (/sub), then a live
session to set /target, /repo, /auth, and free-text /focus instructions (or just type
them) that STEER which agents run and how.
- Slash-commands: /model /providers /key /sub /target /repo /auth /focus /mcp /votes
/agents /show /run /quit (+ bare text = focus).
- RunConfig gains `instructions` and `auth`:
* instructions bias both LLM agent-selection and the heuristic (focus keywords →
injection/access-control/etc. agents get a strong boost)
* operator directives (focus + auth) injected into recon and exploit prompts so agents
test as an authenticated user and prioritise the requested vuln classes
- bump 3.4.1 → 3.5.0 (CLI, harness, reports, credits)
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
- Task-based model ROUTER (recon/select prefer a fast model; exploit prefers primary; validate uses a different model than the finder)
- ReAct doctrine injected into exploit prompts (Thought→Action→Observation, token-efficient)
- Dedup: unique agents per run + findings deduped by CWE/endpoint/title (highest confidence kept)
- Token economy: recon blob capped for selector + per-agent context
- Configurable MCP: merge user mcp.servers.json into the pipeline's .mcp.json
- +54 white-box/code-analysis agents (NoSQLi, LDAP/XPath, JWT-none, Java/.NET/PHP/Go/Node/Python
specifics, SSTI, ReDoS, deserialization, etc.) → 303 agents total (78 code)
- Credits: Joas A Santos & Red Team Leaders (CLI banner, interactive header, HTML+Typst report)
- README: GitHub stars/forks badges, 60-second quick start, full API config steps, intuitive layout
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>