diff --git a/README.md b/README.md index 9b17525..1d56e2d 100755 --- a/README.md +++ b/README.md @@ -11,7 +11,7 @@ - + @@ -32,7 +32,7 @@ LLMs** — via **API key** or local **subscription** (Claude Code / Codex / Gemi Grok) — recons the target, **intelligently selects only the agents that match the discovered surface**, runs them in parallel, **chains** findings into deeper impact, and **validates every claim by cross-model voting + tool-receipt -grounding** before reporting. It ships **473 markdown agents** and a **Mission +grounding** before reporting. It ships **479 markdown agents** and a **Mission Control TUI**. ### Engagement modes @@ -62,6 +62,8 @@ Control TUI**. > reliable BOLA / IDOR / mass-assignment discovery. > > Also a deep **Active Directory** suite: 25+ host/infra skills and 7 multi-stage AD chains covering the full kill chain — initial access, enumeration (BloodHound), Kerberoasting/AS-REP, NTLM relay + coercion (PetitPotam/PrinterBug), delegation abuse (unconstrained/constrained/RBCD + S4U), AD CS (ESC1-ESC13), MSSQL linked-server pivoting, DCSync, cross-forest trust abuse (SID history/trust keys), and persistence (detect-and-report). Lockout- and state-aware, benign-proof-only. +> +> And a **vulnerability-research mode** (`whitebox --research` / `greybox --research`, REPL `/research`, or natural language): hand it a source repo and it hunts a NOVEL, CVE-reportable bug — pins the version/commit, researches known CVEs/advisories (SECURITY.md, CHANGELOG, GHSA, NVD, git history) to de-duplicate, does patch-diff variant analysis (incomplete-fix bypasses, sibling sinks, reintroductions), and gates strictly on novelty. 6 research skills. > **New in v4.2.0** — **binary / APK / IPA testing**: a new `mobile` mode analyses > a local artifact with 12 reverse-engineering skills (static binary triage, @@ -87,7 +89,7 @@ Control TUI**. > (`--compliance pci-dss,hipaa,soc2`); an **internal-network / AD attack graph**; > a **reasoning-budget governor** (`--budget`); and **TypeSafe System One** > (`--typesafe on|off|auto`) as a calibrated confirmation + adjudication layer. -> 27 deterministic per-CWE validators, 473 agents. +> 27 deterministic per-CWE validators, 479 agents. - 🧠 **POMDP belief + anti-hallucination gate** — findings aren't booleans; a property-graph belief carries probabilities, and `may_assert` refuses to claim @@ -240,7 +242,7 @@ Zero npm dependencies (Node built-ins only). out-of-scope) → Leads (the 435-agent board below) → Model & Run (provider/model picker, API-key vs. subscription toggle, votes/chain-depth/recon) → Review. Every engagement is named up front, so runs are identifiable in history instead of by raw target string. -- **Lead board** — all 473 agents auto-categorized (Business Logic, Broken Access Control, +- **Lead board** — all 479 agents auto-categorized (Business Logic, Broken Access Control, Injection, LLM Application, Auth & Session, SSRF & Network, Cloud & Infra, …). Toggle a single lead, a whole category (indeterminate when partially selected), or use **Select all / Clear all** — respects the active search filter. Leave everything off to let the harness's own @@ -982,7 +984,7 @@ Every run writes a self-contained folder `runs/ns--/`: A reinforcement-learning reward store (`data/rl_state_rs.json`) biases agent selection on future runs. -## Agent library — `agents_md/` (473) +## Agent library — `agents_md/` (479) | Category | Count | Purpose | |----------|-------|---------| diff --git a/agents_md/code/research_attack_surface_map.md b/agents_md/code/research_attack_surface_map.md new file mode 100644 index 0000000..b2c63d0 --- /dev/null +++ b/agents_md/code/research_attack_surface_map.md @@ -0,0 +1,59 @@ +# Research Attack-Surface Mapper Agent + +## User Prompt +You are reviewing the source code of **{target}** to map its richest research attack surface: every untrusted entrypoint and every dangerous sink, ranked by reachability, so later taint/logic/variant passes dig where novel bugs are most likely. + +**Recon Context:** +{recon_json} + +The relevant source files are provided to you below the methodology. + +**METHODOLOGY:** + +### 1. Pin the version/commit +- `git -C {target} rev-parse HEAD`; record `name@version` from the manifest so the map is tied to a reviewable snapshot + +### 2. Enumerate untrusted ENTRYPOINTS (sources) +- HTTP routes/handlers: `grep -rniE "route|@app\.(get|post|put|delete)|HandleFunc|app\.(get|post)|#\[(get|post|route)" .` +- Request data: query/body/header/cookie/path-param readers +- Deserialization inputs, file uploads/reads, template render inputs +- CLI args, environment variables, config files +- IPC/RPC/message consumers, websockets, cron/webhook callbacks + +### 3. Enumerate dangerous SINKS +- Exec/command: `system`, `exec`, `child_process`, `Command::new`, backticks +- SQL/NoSQL: raw query concatenation, `format!`/f-string into queries +- Deserialization: `pickle.loads`, `yaml.load`, native/Java/PHP unserialize +- File/path: open/read/write with dynamic paths (traversal) +- SSRF: outbound HTTP with user-controlled URL +- Template/eval: `eval`, `render_template_string`, dynamic template +- Reflected output: HTML/JS write without escaping (XSS) + +### 4. Connect sources to sinks and rank by REACHABILITY +- For each sink, ask: is there a source whose data can reach it? How many hops? Any obvious guard in between? +- Rank: directly-reachable-from-request (high) > reachable-via-internal-call (medium) > guarded/config-gated (low) +- Note auth/authz posture of each entrypoint (anonymous vs authenticated) + +### 5. Record the map (not exploits) +- Produce a table: `entrypoint (file:line) | source type | candidate sink (file:line) | hops | guard? | reachability | suggested next pass` + +### 6. Report Format +For each surface entry: +``` +FINDING: +- Title: Attack-Surface entry -> at [file:line] +- Severity: Info +- CWE: CWE-1059 +- Endpoint: [file:line of the entrypoint] +- Vector: [source type -> candidate sink (file:line), hop count] +- Payload: [the grep/command that located it + the exact quoted entrypoint/sink lines] +- Evidence: [exact code quoted at source and sink + version/commit pinned + reachability rank + any guard observed] +- Impact: Prioritization only — flags where novel taint/logic/variant bugs are most likely; no exploit asserted here +- Remediation: N/A (map); route high-reachability pairs to the source-to-sink taint and logic/authz agents +``` +- Save the ranked surface table to `$NEUROSPLOIT_POCS/{target}-attack-surface.md` (static-derived) and cite it. + +## System Prompt +You are a senior AppSec vulnerability researcher performing reconnaissance of a codebase's attack surface for responsible, novelty-gated research. Your job is to MAP, not to exploit: enumerate untrusted entrypoints and dangerous sinks from the PROVIDED code, connect plausible source->sink pairs, and rank them by reachability so downstream passes focus their effort. Pin the version/commit. Report ONLY what you can see in the code — quote the exact entrypoint and sink lines (file:line) as the receipt; never invent routes or sinks not present. Do not assert exploitability, impact, or any live/HTTP result here — that is for the taint and logic agents. Be explicit about hops and any guard you observe between source and sink. If a mapping is uncertain because the snippet is incomplete, mark it as unconfirmed rather than guessing. + +Credits: Joas A Santos & Red Team Leaders. diff --git a/agents_md/code/research_dependency_nday_reachability.md b/agents_md/code/research_dependency_nday_reachability.md new file mode 100644 index 0000000..beff96e --- /dev/null +++ b/agents_md/code/research_dependency_nday_reachability.md @@ -0,0 +1,49 @@ +# Dependency n-day Reachability Researcher Agent + +## User Prompt +You are reviewing the source code of **{target}** to determine whether a known-vulnerable dependency's flaw is actually REACHABLE from this application's own code — a reachable n-day, or a novel misuse — not a blind lockfile match. + +**Recon Context:** +{recon_json} + +The relevant source files are provided to you below the methodology. + +**METHODOLOGY:** + +### 1. Pin exact dependency versions from lockfiles +- Read the lockfiles, not the loose ranges: `Cargo.lock`, `package-lock.json` / `pnpm-lock.yaml` / `yarn.lock`, `poetry.lock` / `requirements*.txt`, `go.sum` / `go.mod`, `composer.lock`, `Gemfile.lock` +- Record each `name@exact-version` with the file:line where it is pinned + +### 2. Map pinned versions to known CVEs +- Run/read advisory tooling if available: `cargo audit`, `npm audit`, `pip-audit`, `osv-scanner -r .`, `govulncheck ./...` +- Cross-check GHSA/NVD/OSV; for each hit note: CVE/GHSA id, vulnerable range, fixed-in, the VULNERABLE SYMBOL/function/API in the dependency + +### 3. Determine REACHABILITY from this app (the decisive step) +- Find where the app imports/calls the dependency: `grep -rnE "use |require\(['\"]|import .*|from |\." .` +- Does the app actually invoke the VULNERABLE symbol/code path, with attacker-influenced input? Trace source -> the dependency call (quote every hop, file:line) +- Distinguish: (a) vulnerable API called with untrusted data = reachable n-day; (b) dependency present but vulnerable path never invoked = not reachable (say so); (c) app uses the dep in a way the advisory did not cover but is still dangerous = novel misuse + +### 4. Prove or refute reachability +- For a reachable hit: quote the app callsite + the untrusted source feeding it + the dependency's vulnerable entry +- For a non-reachable hit: quote the absence (the vulnerable symbol is never imported/called) and mark NOT REACHABLE + +### 5. Report Format +For each CONFIRMED finding: +``` +FINDING: +- Title: Reachable n-day in @ at [file:line] +- Severity: High +- CWE: CWE-1395 +- Endpoint: [file:line of the app callsite invoking the vulnerable dependency API] +- Vector: [untrusted source (file:line) -> app callsite (file:line) -> vulnerable dep symbol] +- Payload: [benign crafted input that would traverse the app into the vulnerable dep path + the call chain] +- Evidence: [lockfile pin quoted (file:line) + app callsite quoted + advisory id/vulnerable-range/fixed-in + reachable=yes + novel: reachable-here or novel-misuse + checked-against: ] +- Impact: [what the dep CVE yields WHEN reached from here: RCE / DoS / path traversal / etc.] +- Remediation: [upgrade to fixed-in version; or remove the reachable call / constrain input before the vulnerable API] +``` +- Write a static-derived PoC (failing unit test driving the app into the vulnerable dep call) to `$NEUROSPLOIT_POCS/{target}-nday-.{ext}` and cite it. Mark it SOURCE-DERIVED. + +## System Prompt +You are a senior AppSec vulnerability researcher specializing in software-supply-chain reachability analysis, doing responsible, novelty-gated research. A lockfile match alone is NOT a finding — your contribution is proving REACHABILITY: that this application actually invokes the vulnerable dependency symbol with attacker-influenced input. Pin exact versions from the lockfiles (quote file:line), map them to real advisories (CVE/GHSA/OSV) with the vulnerable range and fixed-in, then trace from an untrusted source in the app to the dependency's vulnerable entry, quoting every hop. Report a reachable n-day or a novel misuse only; if the vulnerable path is never invoked, explicitly report NOT REACHABLE rather than inflating it. State `novel: ` and `checked-against: `. No speculation and no live/HTTP claims — source-only. If you cannot see whether the vulnerable symbol is called, say reachability is unconfirmed rather than guess. + +Credits: Joas A Santos & Red Team Leaders. diff --git a/agents_md/code/research_known_cve_dedup.md b/agents_md/code/research_known_cve_dedup.md new file mode 100644 index 0000000..634f443 --- /dev/null +++ b/agents_md/code/research_known_cve_dedup.md @@ -0,0 +1,52 @@ +# Known-CVE Baseline & Dedup Researcher Agent + +## User Prompt +You are reviewing the source code of **{target}** to build a de-duplication baseline so that later findings are provably NOVEL and not restatements of already-patched CVEs/advisories. + +**Recon Context:** +{recon_json} + +The relevant source files are provided to you below the methodology. + +**METHODOLOGY:** + +### 1. Pin the exact version/commit under review +- `git -C {target} rev-parse HEAD` and `git -C {target} describe --tags --always` +- Read package manifests for the self-reported version: `grep -rniE "version\s*[=:]" Cargo.toml package.json pyproject.toml setup.py pom.xml composer.json go.mod 2>/dev/null` +- Record the precise `name@version` / `commit` — every downstream finding must cite it. + +### 2. Harvest the project's own security record +- `git -C {target} log --oneline --all | grep -iE "cve|security|vuln|advisory|rce|xss|sqli|ssrf|auth bypass|sanitiz|escape|overflow"` +- Read `SECURITY.md`, `CHANGELOG*`, `HISTORY*`, `RELEASES*`, `docs/security*` +- `git log -p -- SECURITY.md CHANGELOG*` to see when each fix landed vs. the pinned commit + +### 3. Map fixed-in-version against the version under review +- For each advisory, note "fixed in X.Y.Z"; compare to the pinned version +- If the pinned version is AT OR AFTER the fix commit, that CVE is already patched here (not reportable as-is) +- If BEFORE, it is a known n-day — still not a NOVEL finding, flag it for the n-day/variant agents instead + +### 4. Build the external baseline +- Cross-reference the ecosystem: GHSA (GitHub Advisories), NVD/NIST, `cargo audit` / `npm audit` / `pip-audit` / `osv-scanner` output if present +- `grep -rniE "cve-[0-9]{4}-[0-9]+|ghsa-" .` to catch CVE ids already noted in code/comments/tests +- Produce a table: `CVE/GHSA | class | fixed-in | present-in-this-version? | patch commit` + +### 5. Report Format +For each baseline entry (this agent reports the BASELINE, not exploits): +``` +FINDING: +- Title: Known-CVE Baseline & Dedup entry at [file:line] +- Severity: Info +- CWE: CWE-1059 +- Endpoint: [file:line of the manifest/SECURITY.md/commit proving version+fix status] +- Vector: [advisory id -> class -> fixed-in vs pinned version] +- Payload: [the git log/grep command + exact quoted version string or commit hash] +- Evidence: [exact code/commit quoted + version/commit pinned + whether patched-here=yes/no + source: GHSA/NVD/CHANGELOG] +- Impact: Establishes the dedup baseline; marks each class as already-known so novel findings can be isolated +- Remediation: N/A (baseline); flag already-known-unpatched classes to the n-day/variant agents +``` +- Persist the full baseline table to `$NEUROSPLOIT_POCS/{target}-cve-baseline.md` (static-derived) and cite it. + +## System Prompt +You are a senior AppSec vulnerability researcher doing responsible, novelty-gated research. Your sole job in this pass is to build an authoritative KNOWN-ISSUE baseline for the exact pinned version/commit of {target} so that no later finding re-reports an existing CVE as new. Pin the version from manifests and `git rev-parse` before asserting anything. Treat SECURITY.md, CHANGELOG, GHSA, and NVD as ground truth for "already known". Report ONLY what you can prove from the provided files and git metadata — quote the exact version string, commit hash, or advisory line as the receipt (file:line). Never guess a fix status; if the snippet does not show the version or the patch commit, say so. Do not claim any live/HTTP/network result — this is source-only. Classify each known class as patched-here or not, so novel findings can be cleanly separated downstream. + +Credits: Joas A Santos & Red Team Leaders. diff --git a/agents_md/code/research_logic_authz_flaw.md b/agents_md/code/research_logic_authz_flaw.md new file mode 100644 index 0000000..d2ed8ae --- /dev/null +++ b/agents_md/code/research_logic_authz_flaw.md @@ -0,0 +1,56 @@ +# Business-Logic & Broken-Access-Control Researcher Agent + +## User Prompt +You are reviewing the source code of **{target}** for NOVEL business-logic and broken-access-control flaws: missing or incorrect authorization checks, IDOR, tenant/owner confusion, and state-machine skips — with the exact unguarded code path. + +**Recon Context:** +{recon_json} + +The relevant source files are provided to you below the methodology. + +**METHODOLOGY:** + +### 1. Pin version & novelty baseline +- `git -C {target} rev-parse HEAD`; record `name@version` +- Check whether the authz gap is already a known advisory: `git log --oneline | grep -iE "authz|access|idor|permission|tenant|privilege|owner"`, `SECURITY.md`, `CHANGELOG*`, GHSA/NVD + +### 2. Inventory protected resources & the intended policy +- Identify sensitive actions: read/update/delete of records, admin ops, money/credit moves, role changes, multi-step flows (checkout, invite, reset) +- Infer the INTENDED rule from code/tests/docs: who may do what to which object + +### 3. Find the enforcement points (and the holes) +- Locate guards: `grep -rniE "authoriz|is_admin|has_role|current_user|require_(login|auth)|@login_required|permission|owner|tenant|can\(" .` +- For each sensitive handler, verify a guard exists AND is correct: + - Object-level: does it check the object belongs to `current_user`/tenant, or only that the user is logged in? (IDOR / BOLA) + - Function-level: is an admin-only route reachable by a normal role? (missing function-level authz) + - Is the id/owner taken from the REQUEST instead of the session? (horizontal escalation) + +### 4. Hunt logic/state-machine skips +- Steps that can be reordered or skipped: pay-after-ship, verify-after-use, approve-after-execute +- Mass-assignment into privileged fields (`role`, `is_admin`, `balance`, `owner_id`) — `grep -rniE "update\(|assign|from_json|serde\(flatten\)|params\.permit" .` +- Replay/toctou: a check and a use separated so the state can change in between + +### 5. Prove the unguarded path +- Quote the handler and show the MISSING or INCORRECT check (file:line); trace how a lower-privileged or non-owner actor reaches the action +- Show the request-controlled identifier that selects another user's/tenant's object + +### 6. Report Format +For each CONFIRMED finding: +``` +FINDING: +- Title: at [file:line] +- Severity: High +- CWE: CWE-285 +- Endpoint: [file:line of the vulnerable handler] +- Vector: [actor/role -> action -> object; where the required check is absent/incorrect] +- Payload: [the exact request-shaped input (changed id/role/step) + the call chain reaching the action] +- Evidence: [the handler quoted showing the missing/incorrect guard + the correct guard elsewhere for contrast + version/commit pinned + novel: why + checked-against: ] +- Impact: [horizontal/vertical priv-esc, cross-tenant data access, unauthorized state change] +- Remediation: [enforce object-level ownership/tenant check server-side; derive id from session; gate by role; order-enforce the state machine] +``` +- Write a static-derived PoC (a failing unit/integration test asserting a non-owner/low-role reaches the action) to `$NEUROSPLOIT_POCS/{target}-authz-.{ext}` and cite it. Mark it SOURCE-DERIVED. + +## System Prompt +You are a senior AppSec vulnerability researcher specializing in business-logic and broken-access-control flaws, doing responsible, novelty-gated research. Report ONLY issues you can PROVE in the provided code: quote the vulnerable handler and show the authorization check that is MISSING or INCORRECT (file:line), and trace how an under-privileged or non-owner actor reaches the sensitive action. Anchor claims in the code's own intended policy — ideally contrasting a correct guard elsewhere with the missing one here. Pin the version/commit. The flaw must be NOVEL: state `novel: ` and `checked-against: `; do not re-report an already-fixed access-control advisory for this version unless it is a concrete bypass. No speculation and no live/HTTP claims — source-only. If you cannot see the guard (it may live in middleware or a decorator not provided), say the finding is unconfirmed rather than assume it is absent. + +Credits: Joas A Santos & Red Team Leaders. diff --git a/agents_md/code/research_patch_diff_variant.md b/agents_md/code/research_patch_diff_variant.md new file mode 100644 index 0000000..3724989 --- /dev/null +++ b/agents_md/code/research_patch_diff_variant.md @@ -0,0 +1,54 @@ +# Patch-Diff & Variant (n-day -> 0-day) Researcher Agent + +## User Prompt +You are reviewing the source code of **{target}** to study a recent security-fix commit and find what it MISSED — an incomplete fix (a reachable sink the patch left behind, or a bypass of the new check) and the same bug pattern repeated elsewhere in the tree. The goal is a NOVEL variant, not the already-patched CVE. + +**Recon Context:** +{recon_json} + +The relevant source files are provided to you below the methodology. + +**METHODOLOGY:** + +### 1. Find the security-fix commits +- `git -C {target} log --oneline -n 200 | grep -iE "fix|security|cve|sanitiz|escape|validat|bypass|injection|traversal|overflow|auth"` +- For each candidate: `git -C {target} show ` and `git -C {target} log -p ` to read the exact diff +- Note the CVE/GHSA it addresses and the precise lines/function it changed + +### 2. Characterize the fix precisely +- What was the vulnerable pattern (source -> sink)? What control did the patch ADD (a check, an escape, an allowlist, a type guard)? +- Where is that control enforced — one callsite, or every callsite? Centralized or copy-pasted? + +### 3. Hunt the incomplete-fix / bypass (patch variant) +- Did the patch guard ONE entrypoint but leave a sibling reaching the same sink unguarded? `grep -rn "" .` and compare each callsite against the added check +- Can the new check be bypassed? Look for: normalization mismatches (check before decode), case/encoding gaps, missing recursion, allowlist holes, early-return paths, `unsafe`/raw-SQL/`eval` reached around the guard +- Did the fix cover the reported input but not an equivalent one (alternate parser, second deserializer, another file-read)? + +### 4. Hunt the SAME pattern elsewhere (variant across the tree) +- Build the sink signature from the patched code and grep the whole repo for structurally identical uses that were NEVER patched +- `semgrep` a pattern mirroring the vulnerable shape if available; the CODE CITATION is the proof, not the scanner + +### 5. Prove reachability from untrusted input +- Trace a concrete source (route param, request body, CLI arg, file, env, IPC) to the unpatched sink; quote the full path (file:line each hop) +- Confirm the added control does NOT sit on this path + +### 6. Report Format +For each CONFIRMED finding: +``` +FINDING: +- Title: Patch-Bypass/Variant of at [file:line] +- Severity: High +- CWE: CWE-1288 +- Endpoint: [file:line of the unguarded sink] +- Vector: [untrusted source -> hops -> sink, and how it evades the patch's new check] +- Payload: [benign crafted input that reaches the sink around the fix + the exact call chain] +- Evidence: [exact vulnerable lines quoted + the patch commit sha it bypasses/mirrors + version/commit pinned + novel: why this is NOT the fixed CVE + checked-against: ] +- Impact: [concrete technical impact of the variant] +- Remediation: [centralize the check / cover all callsites / fix the normalization-order or allowlist gap] +``` +- Write the static-derived PoC (crafted input + failing unit test asserting the sink is reached) to `$NEUROSPLOIT_POCS/{target}-variant-.{ext}` and cite it. + +## System Prompt +You are a senior AppSec vulnerability researcher specializing in patch-diff and variant analysis (turning an n-day into a novel 0-day). Report ONLY high-confidence findings you can prove in the PROVIDED code and git history: an incomplete fix, a concrete bypass of a patch's new check, or the same bug pattern in an unpatched location. Always start from a real security-fix commit (`git show`/`git log -p`) and pin the version/commit. A finding is valid ONLY if it is materially DIFFERENT from the already-fixed CVE — a plain restatement of the patched bug is forbidden; state `novel: ` and `checked-against: ` every time. Prove a reachable, unsanitized path from untrusted input and quote every hop (file:line). No speculation, no live/HTTP claims — source-only. If the provided snippet does not show the sink, the source, or the patch, say so rather than guess. + +Credits: Joas A Santos & Red Team Leaders. diff --git a/agents_md/code/research_source_to_sink_taint.md b/agents_md/code/research_source_to_sink_taint.md new file mode 100644 index 0000000..1d1813b --- /dev/null +++ b/agents_md/code/research_source_to_sink_taint.md @@ -0,0 +1,52 @@ +# Source-to-Sink Taint Researcher Agent + +## User Prompt +You are reviewing the source code of **{target}** to prove a NOVEL, reachable, unsanitized taint path from an untrusted entrypoint to a dangerous sink (injection / RCE / SSRF / deserialization / path traversal / prototype pollution). + +**Recon Context:** +{recon_json} + +The relevant source files are provided to you below the methodology. + +**METHODOLOGY:** + +### 1. Pin version & confirm novelty first +- `git -C {target} rev-parse HEAD`; record `name@version` +- Before tracing, confirm the class is not already a patched CVE at this version: scan `SECURITY.md`, `CHANGELOG*`, `git log --oneline | grep -iE "cve|sanitiz|injection|ssrf|traversal|deser"`, GHSA/NVD for the ecosystem + +### 2. Identify the SOURCE (untrusted input) +- Request params/body/headers/cookies, path params, uploaded files, env, CLI args, IPC/queue messages, config a lower-trust actor controls +- Quote the exact line where the value enters (file:line) + +### 3. Identify the SINK +- Command exec, raw SQL/NoSQL, deserializer, dynamic file path, outbound URL, template/eval, reflected HTML/JS, object-key assignment (prototype pollution) +- Locate with `grep -rnE "" .`; quote the exact line (file:line) + +### 4. Trace the DATAFLOW end to end +- Follow the value hop by hop: assignments, function params, struct fields, closures, await boundaries +- Quote EVERY hop (file:line). Identify any validation/escaping/allowlist encountered and PROVE it is absent, insufficient, or bypassable on this path (wrong order, partial, wrong charset, missing recursion) + +### 5. Confirm exploitability (static) +- State the concrete attacker-controlled value that reaches the sink unmodified (or modified in an attacker-useful way) +- Explain why existing controls do not stop it; explain what the sink does with it (RCE, data read/write, SSRF, file disclosure, pollution) + +### 6. Report Format +For each CONFIRMED finding: +``` +FINDING: +- Title: source-to-sink at [file:line] +- Severity: Critical +- CWE: CWE-20 +- Endpoint: [file:line of the source] +- Vector: [source (file:line) -> hop (file:line) -> ... -> sink (file:line)] +- Payload: [benign crafted input demonstrating the reachable path + the exact call chain] +- Evidence: [every hop quoted + sink quoted + version/commit pinned + novel: why this path is not an already-fixed CVE + checked-against: ] +- Impact: [concrete: RCE / SQLi data exfil / SSRF to metadata / arbitrary file read-write / prototype pollution -> ...] +- Remediation: [validate/parameterize/escape at the sink; safe loader; allowlist; normalize-before-check] +``` +- Write a static-derived PoC (crafted input + a failing unit test that drives the source and asserts the sink is reached) to `$NEUROSPLOIT_POCS/{target}-taint-.{ext}` and cite it. Mark it SOURCE-DERIVED. + +## System Prompt +You are a senior AppSec vulnerability researcher performing deep source-to-sink taint analysis for responsible, novelty-gated research. Report ONLY a finding you can PROVE in the provided code: an untrusted source, a dangerous sink, and a reachable, unsanitized dataflow connecting them — with EVERY hop quoted (file:line). Pin the version/commit. The finding must be NOVEL: state `novel: ` and `checked-against: `; a path that merely re-walks an already-patched CVE for this version is not reportable unless recast as a concrete bypass/variant. Prove that controls on the path are absent or defeatable — do not assume sanitization you cannot see, and do not assume a control works that you cannot trace. No speculation and no live/HTTP/network claims — reason strictly about the source and git metadata. If any hop, the source, or the sink is missing from the snippet, say the path is unconfirmed rather than guess. + +Credits: Joas A Santos & Red Team Leaders. diff --git a/neurosploit-rs/app/src/main.rs b/neurosploit-rs/app/src/main.rs index bb67ae8..787e227 100644 --- a/neurosploit-rs/app/src/main.rs +++ b/neurosploit-rs/app/src/main.rs @@ -318,6 +318,9 @@ enum Cmd { /// Economy preset for a short, low-cost review (see `run --quick`). #[arg(long)] quick: bool, + /// Vulnerability-research mode: hunt a NOVEL, CVE-reportable bug (known-CVE dedup + patch-diff variant analysis). + #[arg(long)] + research: bool, #[arg(long)] offline: bool, #[arg(long)] @@ -359,6 +362,9 @@ enum Cmd { /// Economy preset for a short, low-cost test (see `run --quick`). #[arg(long)] quick: bool, + /// Vulnerability-research mode (see `whitebox --research`). + #[arg(long)] + research: bool, #[arg(long)] offline: bool, #[arg(long)] @@ -880,7 +886,7 @@ async fn main() -> anyhow::Result<()> { let ig = harness::integrations::Integrations::load(&repl::proj_dir()); post_integrations(&ig, &url, &out, jira, false, None).await; } - Cmd::Whitebox { path, models, max_agents, vote_n, chain_depth, recon, quick, offline, subscription, jira, only, verbose } => { + Cmd::Whitebox { path, models, max_agents, vote_n, chain_depth, recon, quick, research, offline, subscription, jira, only, verbose } => { let path = resolve_source(&base, &path)?; // local path OR github URL/owner/repo let mut cfg = RunConfig::new(&path); cfg.max_agents = max_agents; @@ -891,6 +897,7 @@ async fn main() -> anyhow::Result<()> { cfg.subscription = subscription; cfg.verbose = verbose; cfg.pinned = parse_only(&only); + cfg.research = research; if quick { apply_quick(&mut cfg); } if !models.is_empty() { cfg.models = models; @@ -900,7 +907,7 @@ async fn main() -> anyhow::Result<()> { let ig = harness::integrations::Integrations::load(&repl::proj_dir()); post_integrations(&ig, &path, &out, jira, false, None).await; } - Cmd::Greybox { repo, url, models, creds, focus, max_agents, vote_n, chain_depth, recon, quick, offline, subscription, mcp, only, verbose } => { + Cmd::Greybox { repo, url, models, creds, focus, max_agents, vote_n, chain_depth, recon, quick, research, offline, subscription, mcp, only, verbose } => { let repo = resolve_source(&base, &repo)?; // local path OR github URL/owner/repo let url = if url.starts_with("http") { url } else { format!("https://{url}") }; let mut cfg = RunConfig::new(&url); @@ -914,6 +921,7 @@ async fn main() -> anyhow::Result<()> { cfg.verbose = verbose; cfg.instructions = focus; cfg.pinned = parse_only(&only); + cfg.research = research; if quick { apply_quick(&mut cfg); } if !models.is_empty() { cfg.models = models; diff --git a/neurosploit-rs/app/src/repl.rs b/neurosploit-rs/app/src/repl.rs index d07a703..b959503 100644 --- a/neurosploit-rs/app/src/repl.rs +++ b/neurosploit-rs/app/src/repl.rs @@ -152,7 +152,7 @@ pub(crate) const ACCEPTED: &[&str] = &[ "/history", "/idle", "/inscope", "/instructions", "/integration", "/integrations", "/key", "/log", "/logs", "/mcp", "/memory", "/model", "/models", "/objective", "/objectives", "/observe", "/observe-only", "/offline", - "/onboard", "/only", "/oos", "/outofscope", "/policy", "/providers", "/proxy", "/quick", "/economy", "/eco", "/q", "/quit", "/recon", + "/onboard", "/only", "/oos", "/outofscope", "/policy", "/providers", "/proxy", "/research", "/quick", "/economy", "/eco", "/q", "/quit", "/recon", "/pause", "/repo", "/report", "/results", "/resume", "/retest", "/revalidate", "/run", "/runs", "/scope", "/scope-out", "/show", "/status", "/stop", "/sub", "/subscription", "/target", "/temp-email", "/tempmail", "/theme", "/timeout", "/ua", "/url", "/useragent", "/validate", @@ -163,7 +163,7 @@ pub(crate) const ACCEPTED: &[&str] = &[ const COMMANDS: &[&str] = &[ "/help", "/onboard", "/show", "/config", "/providers", "/model", "/key", "/sub", "/target", "/repo", "/auth", "/creds", "/focus", "/objective", "/scope-out", "/attach", "/context", "/mcp", "/offline", - "/quick", "/economy", "/eco", "/votes", "/chain", "/recon", "/tempmail", "/timeout", "/proxy", "/burp", "/ua", "/agents", "/only", "/theme", "/clear", "/run", "/stop", "/pause", "/continue", "/runs", "/results", "/report", + "/research", "/quick", "/economy", "/eco", "/votes", "/chain", "/recon", "/tempmail", "/timeout", "/proxy", "/burp", "/ua", "/agents", "/only", "/theme", "/clear", "/run", "/stop", "/pause", "/continue", "/runs", "/results", "/report", "/status", "/logs", "/diff", "/retest", "/validate", "/finding", "/expand", "/integrations", "/memory", "/forget", "/graph", "/inscope", "/observe", "/guardrail", "/policy", "/capability", "/audit", "/quit", @@ -268,6 +268,7 @@ struct Session { max_agents: usize, chain_depth: usize, recon_intensity: usize, + research: bool, /// Opt-in disposable email (mail.tm) for register flows needing a confirmation code. temp_email: bool, /// Idle guardrail: stop a run if no NEW finding lands in this many seconds @@ -320,6 +321,7 @@ impl Default for Session { max_agents: 0, chain_depth: 2, recon_intensity: 3, + research: false, temp_email: false, idle_secs: 300, // 5-minute idle guardrail by default proxy: None, @@ -846,6 +848,13 @@ pub async fn repl(base: &Path, auth: SessionAuth) -> anyhow::Result<()> { if arg.is_empty() { println!(" recon intensity: {} ({}) — set with /recon <1-4> [1 quick · 2 standard · 3 deep · 4 exhaustive]", s.recon_intensity, lvl(s.recon_intensity)); } else { s.recon_intensity = arg.parse::().unwrap_or(s.recon_intensity).clamp(1, 4); println!(" recon intensity: {} ({}) — more rounds, more enumeration, auto-installs tools", s.recon_intensity, lvl(s.recon_intensity)); } } + "/research" => { + match arg.trim() { + "on" | "true" | "1" => { s.research = true; println!(" \x1b[1;36m🔬 research mode ON\x1b[0m — whitebox/greybox will hunt a NOVEL, CVE-reportable bug (known-CVE dedup + patch-diff variant analysis)"); } + "off" | "false" | "0" => { s.research = false; println!(" research mode off"); } + _ => println!(" research mode: {} — /research on|off (for whitebox/greybox: find a new CVE, not a known one)", if s.research { "\x1b[36mon\x1b[0m" } else { "\x1b[2moff\x1b[0m" }), + } + } "/quick" | "/economy" | "/eco" => { // Economy preset for a short, low-cost test — the single switch // for "fast and cheap" instead of tuning each knob. The big @@ -1471,6 +1480,7 @@ async fn run(base: &Path, s: &Session, history: &mut Vec) { cfg.vote_n = s.vote_n; cfg.chain_depth = s.chain_depth; cfg.recon_intensity = s.recon_intensity; + cfg.research = s.research; cfg.temp_email = s.temp_email; cfg.proxy = s.proxy.clone(); cfg.user_agent = s.user_agent.clone(); @@ -1560,6 +1570,7 @@ async fn start_background(base: &Path, s: &Session, reader: &mut Reader, cfg.vote_n = s.vote_n; cfg.chain_depth = s.chain_depth; cfg.recon_intensity = s.recon_intensity; + cfg.research = s.research; cfg.temp_email = s.temp_email; cfg.proxy = s.proxy.clone(); cfg.user_agent = s.user_agent.clone(); @@ -2236,6 +2247,7 @@ fn help() { h("/votes ", "number of validator votes per finding"); h("/chain ", "attack-chain depth (post-exploitation pivots; 0 = off)"); h("/recon <1-4>", "recon intensity: 1 quick · 2 standard · 3 deep · 4 exhaustive (installs tools)"); + h("/research", "whitebox/greybox: hunt a NOVEL, CVE-reportable bug (known-CVE dedup + patch-diff variant analysis)"); h("/quick", "economy preset: short, low-cost run (1 voter · 1 chain round · light recon · ≤6 agents)"); h("/tempmail on|off", "opt-in disposable inbox (mail.tm) to read a register confirmation code"); h("/timeout ", "idle guardrail: stop if no new finding in (0 = off)"); diff --git a/neurosploit-rs/crates/harness/src/pipeline.rs b/neurosploit-rs/crates/harness/src/pipeline.rs index beabe72..6f26275 100644 --- a/neurosploit-rs/crates/harness/src/pipeline.rs +++ b/neurosploit-rs/crates/harness/src/pipeline.rs @@ -623,6 +623,34 @@ const WHITEBOX_DOCTRINE: &str = "MODE: WHITE-BOX STATIC SOURCE REVIEW. You are r - Repro PoC (optional but valued): when a finding warrants it, WRITE a proof/repro script to $NEUROSPLOIT_POCS — e.g. the exact malicious input + the request/CLI call that would trigger the sink, or a unit-style harness exercising the vulnerable function — with a header comment (file:line it proves, how to run). Cite the PoC path in the evidence. Mark clearly that it demonstrates the code path (static-derived), not a live hit.\n\ - Calibrate: High/Critical only when the sink is reachable and exploitable from untrusted input; guarded/unreachable code is Low or a lead.\n\n"; +/// Vulnerability-RESEARCH doctrine: turn a source review into a hunt for a +/// NOVEL, CVE-reportable bug. Prepended (after the white-box doctrine) when the +/// operator runs research mode (`--research`, `/research`, or a focus/objective +/// that asks for a new CVE / 0-day / patch bypass). The whole point is NOVELTY: +/// do not re-report a known CVE — use the known ones as a map to find what the +/// fix missed, a variant, or a reintroduction. +const WHITEBOX_RESEARCH_DOCTRINE: &str = "MISSION: VULNERABILITY RESEARCH — find a NOVEL, CVE-reportable issue in this codebase, not a known one.\n\ +- PIN THE VERSION FIRST: read the version (package.json/VERSION/__init__/composer.json/go.mod/tag) and the exact commit. Every claim is against THIS version/commit; note it in evidence.\n\ +- RESEARCH KNOWN CVEs (de-duplicate): before reporting anything, build the set of ALREADY-KNOWN issues for this project+version — read SECURITY.md, CHANGELOG/release notes, the security advisories (GHSA), CVE/NVD, the issue tracker and recent security commits (`git log --oneline`, grep messages for CVE/security/fix/vuln/XSS/RCE/injection). A finding that matches a known CVE for this version is NOT novel — drop it or recast it ONLY as a patch-bypass/variant (below). State which known CVEs you checked against.\n\ +- PATCH-DIFF / N-DAY -> 0-DAY (the highest-yield path): take a recent SECURITY fix (its commit) and study the diff. Ask: did the patch fix the ROOT CAUSE or just one path? Look for (1) incomplete fixes — another reachable sink the patch did not cover, a bypass of the new check (different encoding, type juggling, alternate parser, case/Unicode, second-order input); (2) the SAME bug pattern elsewhere in the tree (variant analysis — grep the fixed sink's shape across the repo); (3) reintroduction in a later commit. A proven bypass of an existing patch IS novel and reportable.\n\ +- SOURCE->SINK with reachability: trace untrusted input (request params, headers, body, deserialized objects, file names, env, IPC, config) to a dangerous sink (SQL/exec/eval/template/path/deserialize/SSRF/XXE/prototype/unsafe-reflection). Only call it a vuln when the path is REACHABLE from an untrusted entrypoint without an effective sanitizer; record the full path `entry -> … -> sink`.\n\ +- NOVELTY GATE (strict): report a finding ONLY if (a) it does not match a known CVE for this version, OR (b) it is a concrete bypass/variant of a patched issue. For each, state explicitly: 'novel: ' and 'checked-against: '. No speculation — high-confidence, evidence-backed only.\n\ +- WRITE A PoC: produce a minimal, SAFE proof (a failing unit test, a crafted input + the exact call reaching the sink, or a request) to $NEUROSPLOIT_POCS and cite it. Where a running instance is available (greybox), CONFIRM the source-derived bug dynamically — but never run destructive payloads.\n\ +- REPORT for disclosure: each finding carries file:line, root-cause analysis, the version/commit, the novelty justification, CVSS vector, PoC path, suggested fix, and the project's disclosure channel (SECURITY.md / security@). Prefer a few solid, novel, reportable bugs over a long list of known or speculative ones.\n\n"; + +/// True when the engagement is asking for vulnerability research / a new CVE, +/// either via the explicit flag or from natural-language focus/objective. +fn is_research_intent(cfg: &RunConfig) -> bool { + if cfg.research { return true; } + let hay = format!("{} {}", + cfg.instructions.clone().unwrap_or_default(), + cfg.objective.clone().unwrap_or_default()).to_lowercase(); + ["new cve", "novel", "0-day", "0day", "zero-day", "zero day", "patch bypass", + "patch-bypass", "variant analysis", "n-day", "nday", "reportable", "cve research", + "vulnerability research", "find a cve", "nova cve", "pesquisa de vuln"] + .iter().any(|k| hay.contains(k)) +} + /// Methodology directions for a modern JS SPA backed by a REST/GraphQL API /// (Angular/React/Vue front + Node/Express-style API — the shape of OWASP Juice /// Shop and many real apps). These are DIRECTIONS on HOW to hunt each vuln class, @@ -1167,6 +1195,10 @@ pub async fn run_whitebox(cfg: RunConfig, lib: &Library, pool: &ModelPool, tx: S let context = collect_repo_context(Path::new(&cfg.target), 200, 120_000); let bytes = context.len(); + let research = is_research_intent(&cfg); + if research { + let _ = tx.send("notify: 🔬 research mode — hunting a NOVEL, CVE-reportable issue (known-CVE dedup + patch-diff variant analysis)".into()).await; + } let _ = tx.send(format!("collected {} bytes of source context", bytes)).await; if bytes == 0 { let _ = tx.send("no readable source found at the given path".into()).await; @@ -1214,7 +1246,12 @@ pub async fn run_whitebox(cfg: RunConfig, lib: &Library, pool: &ModelPool, tx: S ); // Prepend the white-box doctrine so code agents stay in static // source-review mode and never hallucinate live/black-box actions. - let sys = format!("{}{}", WHITEBOX_DOCTRINE, ag.system); + // In research mode, add the novelty/patch-diff hunting doctrine. + let sys = if research { + format!("{}{}{}", WHITEBOX_DOCTRINE, WHITEBOX_RESEARCH_DOCTRINE, ag.system) + } else { + format!("{}{}", WHITEBOX_DOCTRINE, ag.system) + }; match pool.complete_routed(Task::Exploit, &ag.name, &sys, &user).await { Ok((m, text)) => { let f = extract_findings(&text, &ag.name); @@ -1267,6 +1304,10 @@ pub async fn run_greybox(cfg: RunConfig, lib: &Library, pool: &ModelPool, tx: Se // ---- 2. Review the source for leads ------------------------------- let context = collect_repo_context(Path::new(&repo), 200, 90_000); let _ = tx.send(format!("collected {} bytes of source for code review", context.len())).await; + let gb_research = is_research_intent(&cfg); + if gb_research { + let _ = tx.send("notify: 🔬 research mode — code review hunts a novel, CVE-reportable bug, then confirms it live".into()).await; + } let mut rl = cfg.rl_path.as_ref().map(|p| RlState::load(Path::new(p))).unwrap_or_default(); let mut code_leads = String::new(); @@ -1284,7 +1325,10 @@ pub async fn run_greybox(cfg: RunConfig, lib: &Library, pool: &ModelPool, tx: Se where endpoint is file:line.", ag.user.replace("{target}", "the repository").replace("{recon_json}", "{}"), ctx ); - match pool.complete_routed(Task::Select, &ag.name, &ag.system, &user).await { + // Research mode steers the code-review half of greybox too: + // find a novel, reportable bug, then confirm it live. + let sys = if gb_research { format!("{}{}", WHITEBOX_RESEARCH_DOCTRINE, ag.system) } else { ag.system.clone() }; + match pool.complete_routed(Task::Select, &ag.name, &sys, &user).await { Ok((_, text)) => { let f = extract_findings(&text, &ag.name); let _ = txc.send(format!("review {} → {} lead(s)", ag.name, f.len())).await; f } Err(_) => vec![], diff --git a/neurosploit-rs/crates/harness/src/types.rs b/neurosploit-rs/crates/harness/src/types.rs index c12c8a4..78b6f40 100644 --- a/neurosploit-rs/crates/harness/src/types.rs +++ b/neurosploit-rs/crates/harness/src/types.rs @@ -213,6 +213,12 @@ pub struct RunConfig { /// more recon rounds, more active enumeration, and auto-installing tools. #[serde(default = "default_recon")] pub recon_intensity: usize, + /// Vulnerability-research mode: hunt for NOVEL, CVE-reportable issues in a + /// source repo — research known CVEs/advisories to de-duplicate, do + /// patch-diff variant analysis (incomplete-fix bypasses, sibling sinks), + /// and gate strictly on novelty. Steers whitebox/greybox. + #[serde(default)] + pub research: bool, /// Opt-in: when the app requires email confirmation to register, allow the /// agent to use a free disposable-inbox API (mail.tm) to read the code/link. /// Off by default. Account creation is still capped by the safety guardrail. @@ -315,6 +321,7 @@ impl RunConfig { repo: None, pinned: Vec::new(), chain_depth: 2, + research: false, proxy: None, user_agent: None, recon_intensity: 3, diff --git a/web/public/app.js b/web/public/app.js index f8e90c1..a45933c 100644 --- a/web/public/app.js +++ b/web/public/app.js @@ -533,6 +533,7 @@ async function startExploitation() { chainDepth: Number($('#fieldChain').value), recon: Number($('#fieldRecon').value), quick: $('#fieldQuick') ? $('#fieldQuick').checked : false, + research: $('#fieldResearch') ? $('#fieldResearch').checked : false, subscription: state.authMode === 'subscription', mcp: $('#fieldMcp').checked, agents: [...state.selected], diff --git a/web/public/index.html b/web/public/index.html index 705ad35..7420e92 100644 --- a/web/public/index.html +++ b/web/public/index.html @@ -221,6 +221,9 @@
Optional. Left on unlimited, the run behaves exactly as it always has — full depth, no cap.
+
+
Known-CVE de-dup + patch-diff variant analysis. Best with a source repo; reports only genuinely new issues (or a concrete patch bypass).
+
Economy preset: 1 voter, 1 chain round, light recon, ≤6 agents, eco budget. The big token saver. Wins over the settings below.
diff --git a/web/server.js b/web/server.js index 43940ab..a7909f3 100644 --- a/web/server.js +++ b/web/server.js @@ -882,6 +882,7 @@ function sanitizeLaunch(body) { sandbox: !!body.sandbox, typesafe: body.typesafe, quick: !!body.quick, + research: !!body.research, }; } @@ -902,6 +903,7 @@ function buildReplScript(body) { if (body.objective) lines.push(`/objective ${body.objective}`); if (body.outOfScope) lines.push(`/scope-out ${body.outOfScope}`); if (body.creds) lines.push(`/creds ${body.creds}`); + if (body.research) lines.push(`/research on`); lines.push((body.agents || []).length ? `/only ${body.agents.join(',')}` : '/only clear'); // Economy preset last, so it wins over the per-knob settings above. if (body.quick) lines.push('/quick');