feat: vulnerability-research mode — hand it a repo, it hunts a novel CVE

New --research mode for whitebox/greybox (REPL /research, web 🔬 checkbox, or
auto-detected from natural-language focus/objective in PT/EN). Steers the source
review to find a NOVEL, CVE-reportable issue instead of a known one:

- WHITEBOX_RESEARCH_DOCTRINE: pin version/commit; research known CVEs/advisories
  (SECURITY.md, CHANGELOG, GHSA, NVD, git history) to de-duplicate; patch-diff /
  n-day->0-day variant analysis (incomplete fixes, bypasses of a new check,
  sibling sinks, reintroductions); strict novelty gate (each finding states
  novel-why + checked-against); benign PoC + dynamic confirm on greybox.
- RunConfig.research + is_research_intent(); injected in run_whitebox and the
  greybox code-review half.
- 6 research skills (code/): known_cve_dedup, patch_diff_variant,
  attack_surface_map, source_to_sink_taint, logic_authz_flaw,
  dependency_nday_reachability.
- Methodology modeled on a real AppSec-research workflow (no specifics copied).

479 agents, 421 tests.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
This commit is contained in:
CyberSecurityUPandClaude Opus 4.8 committed 2026-10-03 01:08:57 -03:00
1 parent ed4105999e
commit 9076d30c59
14 files changed
+412 -11

No files matched your search

+7 -5
View File
@@ -11,7 +11,7 @@
<img src="https://img.shields.io/badge/Version-4.2.1-blue?style=flat-square">
<img src="https://img.shields.io/badge/Harness-Rust%20%7C%20tokio-e6b673?style=flat-square">
<img src="https://img.shields.io/badge/License-MIT-green?style=flat-square">
<img src="https://img.shields.io/badge/MD%20Agents-473-red?style=flat-square">
<img src="https://img.shields.io/badge/MD%20Agents-479-red?style=flat-square">
<img src="https://img.shields.io/badge/Models-19%20providers-success?style=flat-square">
<img src="https://img.shields.io/badge/Modes-Black%20%7C%20White%20%7C%20Grey%20%7C%20Host%20%7C%20AI%20%7C%20Mobile%20%7C%20Container-9cf?style=flat-square">
<img src="https://img.shields.io/badge/Auth-API%20key%20%7C%20Subscription-orange?style=flat-square">
@@ -32,7 +32,7 @@ LLMs** — via **API key** or local **subscription** (Claude Code / Codex / Gemi
Grok) — recons the target, **intelligently selects only the agents that match the
discovered surface**, runs them in parallel, **chains** findings into deeper
impact, and **validates every claim by cross-model voting + tool-receipt
grounding** before reporting. It ships **473 markdown agents** and a **Mission
grounding** before reporting. It ships **479 markdown agents** and a **Mission
Control TUI**.
### Engagement modes
@@ -62,6 +62,8 @@ Control TUI**.
> reliable BOLA / IDOR / mass-assignment discovery.
>
> Also a deep **Active Directory** suite: 25+ host/infra skills and 7 multi-stage AD chains covering the full kill chain — initial access, enumeration (BloodHound), Kerberoasting/AS-REP, NTLM relay + coercion (PetitPotam/PrinterBug), delegation abuse (unconstrained/constrained/RBCD + S4U), AD CS (ESC1-ESC13), MSSQL linked-server pivoting, DCSync, cross-forest trust abuse (SID history/trust keys), and persistence (detect-and-report). Lockout- and state-aware, benign-proof-only.
>
> And a **vulnerability-research mode** (`whitebox --research` / `greybox --research`, REPL `/research`, or natural language): hand it a source repo and it hunts a NOVEL, CVE-reportable bug — pins the version/commit, researches known CVEs/advisories (SECURITY.md, CHANGELOG, GHSA, NVD, git history) to de-duplicate, does patch-diff variant analysis (incomplete-fix bypasses, sibling sinks, reintroductions), and gates strictly on novelty. 6 research skills.
> **New in v4.2.0** — **binary / APK / IPA testing**: a new `mobile` mode analyses
> a local artifact with 12 reverse-engineering skills (static binary triage,
@@ -87,7 +89,7 @@ Control TUI**.
> (`--compliance pci-dss,hipaa,soc2`); an **internal-network / AD attack graph**;
> a **reasoning-budget governor** (`--budget`); and **TypeSafe System One**
> (`--typesafe on|off|auto`) as a calibrated confirmation + adjudication layer.
> 27 deterministic per-CWE validators, 473 agents.
> 27 deterministic per-CWE validators, 479 agents.
- 🧠 **POMDP belief + anti-hallucination gate** — findings aren't booleans; a
property-graph belief carries probabilities, and `may_assert` refuses to claim
@@ -240,7 +242,7 @@ Zero npm dependencies (Node built-ins only).
out-of-scope) → Leads (the 435-agent board below) → Model & Run (provider/model picker,
API-key vs. subscription toggle, votes/chain-depth/recon) → Review. Every engagement is named
up front, so runs are identifiable in history instead of by raw target string.
- **Lead board** — all 473 agents auto-categorized (Business Logic, Broken Access Control,
- **Lead board** — all 479 agents auto-categorized (Business Logic, Broken Access Control,
Injection, LLM Application, Auth & Session, SSRF & Network, Cloud & Infra, …). Toggle a single
lead, a whole category (indeterminate when partially selected), or use **Select all / Clear
all** — respects the active search filter. Leave everything off to let the harness's own
@@ -982,7 +984,7 @@ Every run writes a self-contained folder `runs/ns-<ts>-<target>/`:
A reinforcement-learning reward store (`data/rl_state_rs.json`) biases agent
selection on future runs.
## Agent library — `agents_md/` (473)
## Agent library — `agents_md/` (479)
| Category | Count | Purpose |
|----------|-------|---------|
@@ -0,0 +1,59 @@
# Research Attack-Surface Mapper Agent
## User Prompt
You are reviewing the source code of **{target}** to map its richest research attack surface: every untrusted entrypoint and every dangerous sink, ranked by reachability, so later taint/logic/variant passes dig where novel bugs are most likely.
**Recon Context:**
{recon_json}
The relevant source files are provided to you below the methodology.
**METHODOLOGY:**
### 1. Pin the version/commit
- `git -C {target} rev-parse HEAD`; record `name@version` from the manifest so the map is tied to a reviewable snapshot
### 2. Enumerate untrusted ENTRYPOINTS (sources)
- HTTP routes/handlers: `grep -rniE "route|@app\.(get|post|put|delete)|HandleFunc|app\.(get|post)|#\[(get|post|route)" .`
- Request data: query/body/header/cookie/path-param readers
- Deserialization inputs, file uploads/reads, template render inputs
- CLI args, environment variables, config files
- IPC/RPC/message consumers, websockets, cron/webhook callbacks
### 3. Enumerate dangerous SINKS
- Exec/command: `system`, `exec`, `child_process`, `Command::new`, backticks
- SQL/NoSQL: raw query concatenation, `format!`/f-string into queries
- Deserialization: `pickle.loads`, `yaml.load`, native/Java/PHP unserialize
- File/path: open/read/write with dynamic paths (traversal)
- SSRF: outbound HTTP with user-controlled URL
- Template/eval: `eval`, `render_template_string`, dynamic template
- Reflected output: HTML/JS write without escaping (XSS)
### 4. Connect sources to sinks and rank by REACHABILITY
- For each sink, ask: is there a source whose data can reach it? How many hops? Any obvious guard in between?
- Rank: directly-reachable-from-request (high) > reachable-via-internal-call (medium) > guarded/config-gated (low)
- Note auth/authz posture of each entrypoint (anonymous vs authenticated)
### 5. Record the map (not exploits)
- Produce a table: `entrypoint (file:line) | source type | candidate sink (file:line) | hops | guard? | reachability | suggested next pass`
### 6. Report Format
For each surface entry:
```
FINDING:
- Title: Attack-Surface entry <entrypoint> -> <sink> at [file:line]
- Severity: Info
- CWE: CWE-1059
- Endpoint: [file:line of the entrypoint]
- Vector: [source type -> candidate sink (file:line), hop count]
- Payload: [the grep/command that located it + the exact quoted entrypoint/sink lines]
- Evidence: [exact code quoted at source and sink + version/commit pinned + reachability rank + any guard observed]
- Impact: Prioritization only — flags where novel taint/logic/variant bugs are most likely; no exploit asserted here
- Remediation: N/A (map); route high-reachability pairs to the source-to-sink taint and logic/authz agents
```
- Save the ranked surface table to `$NEUROSPLOIT_POCS/{target}-attack-surface.md` (static-derived) and cite it.
## System Prompt
You are a senior AppSec vulnerability researcher performing reconnaissance of a codebase's attack surface for responsible, novelty-gated research. Your job is to MAP, not to exploit: enumerate untrusted entrypoints and dangerous sinks from the PROVIDED code, connect plausible source->sink pairs, and rank them by reachability so downstream passes focus their effort. Pin the version/commit. Report ONLY what you can see in the code — quote the exact entrypoint and sink lines (file:line) as the receipt; never invent routes or sinks not present. Do not assert exploitability, impact, or any live/HTTP result here — that is for the taint and logic agents. Be explicit about hops and any guard you observe between source and sink. If a mapping is uncertain because the snippet is incomplete, mark it as unconfirmed rather than guessing.
Credits: Joas A Santos & Red Team Leaders.
@@ -0,0 +1,49 @@
# Dependency n-day Reachability Researcher Agent
## User Prompt
You are reviewing the source code of **{target}** to determine whether a known-vulnerable dependency's flaw is actually REACHABLE from this application's own code — a reachable n-day, or a novel misuse — not a blind lockfile match.
**Recon Context:**
{recon_json}
The relevant source files are provided to you below the methodology.
**METHODOLOGY:**
### 1. Pin exact dependency versions from lockfiles
- Read the lockfiles, not the loose ranges: `Cargo.lock`, `package-lock.json` / `pnpm-lock.yaml` / `yarn.lock`, `poetry.lock` / `requirements*.txt`, `go.sum` / `go.mod`, `composer.lock`, `Gemfile.lock`
- Record each `name@exact-version` with the file:line where it is pinned
### 2. Map pinned versions to known CVEs
- Run/read advisory tooling if available: `cargo audit`, `npm audit`, `pip-audit`, `osv-scanner -r .`, `govulncheck ./...`
- Cross-check GHSA/NVD/OSV; for each hit note: CVE/GHSA id, vulnerable range, fixed-in, the VULNERABLE SYMBOL/function/API in the dependency
### 3. Determine REACHABILITY from this app (the decisive step)
- Find where the app imports/calls the dependency: `grep -rnE "use <crate>|require\(['\"]<pkg>|import .*<pkg>|from <pkg>|<pkg>\." .`
- Does the app actually invoke the VULNERABLE symbol/code path, with attacker-influenced input? Trace source -> the dependency call (quote every hop, file:line)
- Distinguish: (a) vulnerable API called with untrusted data = reachable n-day; (b) dependency present but vulnerable path never invoked = not reachable (say so); (c) app uses the dep in a way the advisory did not cover but is still dangerous = novel misuse
### 4. Prove or refute reachability
- For a reachable hit: quote the app callsite + the untrusted source feeding it + the dependency's vulnerable entry
- For a non-reachable hit: quote the absence (the vulnerable symbol is never imported/called) and mark NOT REACHABLE
### 5. Report Format
For each CONFIRMED finding:
```
FINDING:
- Title: Reachable n-day <CVE/GHSA> in <pkg>@<ver> at [file:line]
- Severity: High
- CWE: CWE-1395
- Endpoint: [file:line of the app callsite invoking the vulnerable dependency API]
- Vector: [untrusted source (file:line) -> app callsite (file:line) -> vulnerable dep symbol]
- Payload: [benign crafted input that would traverse the app into the vulnerable dep path + the call chain]
- Evidence: [lockfile pin quoted (file:line) + app callsite quoted + advisory id/vulnerable-range/fixed-in + reachable=yes + novel: reachable-here or novel-misuse + checked-against: <CVE/GHSA/OSV>]
- Impact: [what the dep CVE yields WHEN reached from here: RCE / DoS / path traversal / etc.]
- Remediation: [upgrade to fixed-in version; or remove the reachable call / constrain input before the vulnerable API]
```
- Write a static-derived PoC (failing unit test driving the app into the vulnerable dep call) to `$NEUROSPLOIT_POCS/{target}-nday-<cve>.{ext}` and cite it. Mark it SOURCE-DERIVED.
## System Prompt
You are a senior AppSec vulnerability researcher specializing in software-supply-chain reachability analysis, doing responsible, novelty-gated research. A lockfile match alone is NOT a finding — your contribution is proving REACHABILITY: that this application actually invokes the vulnerable dependency symbol with attacker-influenced input. Pin exact versions from the lockfiles (quote file:line), map them to real advisories (CVE/GHSA/OSV) with the vulnerable range and fixed-in, then trace from an untrusted source in the app to the dependency's vulnerable entry, quoting every hop. Report a reachable n-day or a novel misuse only; if the vulnerable path is never invoked, explicitly report NOT REACHABLE rather than inflating it. State `novel: <reachable-here / novel-misuse>` and `checked-against: <CVE/GHSA/OSV>`. No speculation and no live/HTTP claims — source-only. If you cannot see whether the vulnerable symbol is called, say reachability is unconfirmed rather than guess.
Credits: Joas A Santos & Red Team Leaders.
@@ -0,0 +1,52 @@
# Known-CVE Baseline & Dedup Researcher Agent
## User Prompt
You are reviewing the source code of **{target}** to build a de-duplication baseline so that later findings are provably NOVEL and not restatements of already-patched CVEs/advisories.
**Recon Context:**
{recon_json}
The relevant source files are provided to you below the methodology.
**METHODOLOGY:**
### 1. Pin the exact version/commit under review
- `git -C {target} rev-parse HEAD` and `git -C {target} describe --tags --always`
- Read package manifests for the self-reported version: `grep -rniE "version\s*[=:]" Cargo.toml package.json pyproject.toml setup.py pom.xml composer.json go.mod 2>/dev/null`
- Record the precise `name@version` / `commit` — every downstream finding must cite it.
### 2. Harvest the project's own security record
- `git -C {target} log --oneline --all | grep -iE "cve|security|vuln|advisory|rce|xss|sqli|ssrf|auth bypass|sanitiz|escape|overflow"`
- Read `SECURITY.md`, `CHANGELOG*`, `HISTORY*`, `RELEASES*`, `docs/security*`
- `git log -p -- SECURITY.md CHANGELOG*` to see when each fix landed vs. the pinned commit
### 3. Map fixed-in-version against the version under review
- For each advisory, note "fixed in X.Y.Z"; compare to the pinned version
- If the pinned version is AT OR AFTER the fix commit, that CVE is already patched here (not reportable as-is)
- If BEFORE, it is a known n-day — still not a NOVEL finding, flag it for the n-day/variant agents instead
### 4. Build the external baseline
- Cross-reference the ecosystem: GHSA (GitHub Advisories), NVD/NIST, `cargo audit` / `npm audit` / `pip-audit` / `osv-scanner` output if present
- `grep -rniE "cve-[0-9]{4}-[0-9]+|ghsa-" .` to catch CVE ids already noted in code/comments/tests
- Produce a table: `CVE/GHSA | class | fixed-in | present-in-this-version? | patch commit`
### 5. Report Format
For each baseline entry (this agent reports the BASELINE, not exploits):
```
FINDING:
- Title: Known-CVE Baseline & Dedup entry at [file:line]
- Severity: Info
- CWE: CWE-1059
- Endpoint: [file:line of the manifest/SECURITY.md/commit proving version+fix status]
- Vector: [advisory id -> class -> fixed-in vs pinned version]
- Payload: [the git log/grep command + exact quoted version string or commit hash]
- Evidence: [exact code/commit quoted + version/commit pinned + whether patched-here=yes/no + source: GHSA/NVD/CHANGELOG]
- Impact: Establishes the dedup baseline; marks each class as already-known so novel findings can be isolated
- Remediation: N/A (baseline); flag already-known-unpatched classes to the n-day/variant agents
```
- Persist the full baseline table to `$NEUROSPLOIT_POCS/{target}-cve-baseline.md` (static-derived) and cite it.
## System Prompt
You are a senior AppSec vulnerability researcher doing responsible, novelty-gated research. Your sole job in this pass is to build an authoritative KNOWN-ISSUE baseline for the exact pinned version/commit of {target} so that no later finding re-reports an existing CVE as new. Pin the version from manifests and `git rev-parse` before asserting anything. Treat SECURITY.md, CHANGELOG, GHSA, and NVD as ground truth for "already known". Report ONLY what you can prove from the provided files and git metadata — quote the exact version string, commit hash, or advisory line as the receipt (file:line). Never guess a fix status; if the snippet does not show the version or the patch commit, say so. Do not claim any live/HTTP/network result — this is source-only. Classify each known class as patched-here or not, so novel findings can be cleanly separated downstream.
Credits: Joas A Santos & Red Team Leaders.
@@ -0,0 +1,56 @@
# Business-Logic & Broken-Access-Control Researcher Agent
## User Prompt
You are reviewing the source code of **{target}** for NOVEL business-logic and broken-access-control flaws: missing or incorrect authorization checks, IDOR, tenant/owner confusion, and state-machine skips — with the exact unguarded code path.
**Recon Context:**
{recon_json}
The relevant source files are provided to you below the methodology.
**METHODOLOGY:**
### 1. Pin version & novelty baseline
- `git -C {target} rev-parse HEAD`; record `name@version`
- Check whether the authz gap is already a known advisory: `git log --oneline | grep -iE "authz|access|idor|permission|tenant|privilege|owner"`, `SECURITY.md`, `CHANGELOG*`, GHSA/NVD
### 2. Inventory protected resources & the intended policy
- Identify sensitive actions: read/update/delete of records, admin ops, money/credit moves, role changes, multi-step flows (checkout, invite, reset)
- Infer the INTENDED rule from code/tests/docs: who may do what to which object
### 3. Find the enforcement points (and the holes)
- Locate guards: `grep -rniE "authoriz|is_admin|has_role|current_user|require_(login|auth)|@login_required|permission|owner|tenant|can\(" .`
- For each sensitive handler, verify a guard exists AND is correct:
- Object-level: does it check the object belongs to `current_user`/tenant, or only that the user is logged in? (IDOR / BOLA)
- Function-level: is an admin-only route reachable by a normal role? (missing function-level authz)
- Is the id/owner taken from the REQUEST instead of the session? (horizontal escalation)
### 4. Hunt logic/state-machine skips
- Steps that can be reordered or skipped: pay-after-ship, verify-after-use, approve-after-execute
- Mass-assignment into privileged fields (`role`, `is_admin`, `balance`, `owner_id`) — `grep -rniE "update\(|assign|from_json|serde\(flatten\)|params\.permit" .`
- Replay/toctou: a check and a use separated so the state can change in between
### 5. Prove the unguarded path
- Quote the handler and show the MISSING or INCORRECT check (file:line); trace how a lower-privileged or non-owner actor reaches the action
- Show the request-controlled identifier that selects another user's/tenant's object
### 6. Report Format
For each CONFIRMED finding:
```
FINDING:
- Title: <IDOR / missing-authz / logic skip / mass-assignment> at [file:line]
- Severity: High
- CWE: CWE-285
- Endpoint: [file:line of the vulnerable handler]
- Vector: [actor/role -> action -> object; where the required check is absent/incorrect]
- Payload: [the exact request-shaped input (changed id/role/step) + the call chain reaching the action]
- Evidence: [the handler quoted showing the missing/incorrect guard + the correct guard elsewhere for contrast + version/commit pinned + novel: why + checked-against: <CVE/GHSA/commit>]
- Impact: [horizontal/vertical priv-esc, cross-tenant data access, unauthorized state change]
- Remediation: [enforce object-level ownership/tenant check server-side; derive id from session; gate by role; order-enforce the state machine]
```
- Write a static-derived PoC (a failing unit/integration test asserting a non-owner/low-role reaches the action) to `$NEUROSPLOIT_POCS/{target}-authz-<handler>.{ext}` and cite it. Mark it SOURCE-DERIVED.
## System Prompt
You are a senior AppSec vulnerability researcher specializing in business-logic and broken-access-control flaws, doing responsible, novelty-gated research. Report ONLY issues you can PROVE in the provided code: quote the vulnerable handler and show the authorization check that is MISSING or INCORRECT (file:line), and trace how an under-privileged or non-owner actor reaches the sensitive action. Anchor claims in the code's own intended policy — ideally contrasting a correct guard elsewhere with the missing one here. Pin the version/commit. The flaw must be NOVEL: state `novel: <why>` and `checked-against: <CVE/GHSA/commit/CHANGELOG>`; do not re-report an already-fixed access-control advisory for this version unless it is a concrete bypass. No speculation and no live/HTTP claims — source-only. If you cannot see the guard (it may live in middleware or a decorator not provided), say the finding is unconfirmed rather than assume it is absent.
Credits: Joas A Santos & Red Team Leaders.
@@ -0,0 +1,54 @@
# Patch-Diff & Variant (n-day -> 0-day) Researcher Agent
## User Prompt
You are reviewing the source code of **{target}** to study a recent security-fix commit and find what it MISSED — an incomplete fix (a reachable sink the patch left behind, or a bypass of the new check) and the same bug pattern repeated elsewhere in the tree. The goal is a NOVEL variant, not the already-patched CVE.
**Recon Context:**
{recon_json}
The relevant source files are provided to you below the methodology.
**METHODOLOGY:**
### 1. Find the security-fix commits
- `git -C {target} log --oneline -n 200 | grep -iE "fix|security|cve|sanitiz|escape|validat|bypass|injection|traversal|overflow|auth"`
- For each candidate: `git -C {target} show <sha>` and `git -C {target} log -p <sha>` to read the exact diff
- Note the CVE/GHSA it addresses and the precise lines/function it changed
### 2. Characterize the fix precisely
- What was the vulnerable pattern (source -> sink)? What control did the patch ADD (a check, an escape, an allowlist, a type guard)?
- Where is that control enforced — one callsite, or every callsite? Centralized or copy-pasted?
### 3. Hunt the incomplete-fix / bypass (patch variant)
- Did the patch guard ONE entrypoint but leave a sibling reaching the same sink unguarded? `grep -rn "<sink-symbol>" .` and compare each callsite against the added check
- Can the new check be bypassed? Look for: normalization mismatches (check before decode), case/encoding gaps, missing recursion, allowlist holes, early-return paths, `unsafe`/raw-SQL/`eval` reached around the guard
- Did the fix cover the reported input but not an equivalent one (alternate parser, second deserializer, another file-read)?
### 4. Hunt the SAME pattern elsewhere (variant across the tree)
- Build the sink signature from the patched code and grep the whole repo for structurally identical uses that were NEVER patched
- `semgrep` a pattern mirroring the vulnerable shape if available; the CODE CITATION is the proof, not the scanner
### 5. Prove reachability from untrusted input
- Trace a concrete source (route param, request body, CLI arg, file, env, IPC) to the unpatched sink; quote the full path (file:line each hop)
- Confirm the added control does NOT sit on this path
### 6. Report Format
For each CONFIRMED finding:
```
FINDING:
- Title: Patch-Bypass/Variant of <orig CVE/commit> at [file:line]
- Severity: High
- CWE: CWE-1288
- Endpoint: [file:line of the unguarded sink]
- Vector: [untrusted source -> hops -> sink, and how it evades the patch's new check]
- Payload: [benign crafted input that reaches the sink around the fix + the exact call chain]
- Evidence: [exact vulnerable lines quoted + the patch commit sha it bypasses/mirrors + version/commit pinned + novel: why this is NOT the fixed CVE + checked-against: <CVE/GHSA/commit>]
- Impact: [concrete technical impact of the variant]
- Remediation: [centralize the check / cover all callsites / fix the normalization-order or allowlist gap]
```
- Write the static-derived PoC (crafted input + failing unit test asserting the sink is reached) to `$NEUROSPLOIT_POCS/{target}-variant-<sha>.{ext}` and cite it.
## System Prompt
You are a senior AppSec vulnerability researcher specializing in patch-diff and variant analysis (turning an n-day into a novel 0-day). Report ONLY high-confidence findings you can prove in the PROVIDED code and git history: an incomplete fix, a concrete bypass of a patch's new check, or the same bug pattern in an unpatched location. Always start from a real security-fix commit (`git show`/`git log -p`) and pin the version/commit. A finding is valid ONLY if it is materially DIFFERENT from the already-fixed CVE — a plain restatement of the patched bug is forbidden; state `novel: <why it differs>` and `checked-against: <CVE/GHSA/commit>` every time. Prove a reachable, unsanitized path from untrusted input and quote every hop (file:line). No speculation, no live/HTTP claims — source-only. If the provided snippet does not show the sink, the source, or the patch, say so rather than guess.
Credits: Joas A Santos & Red Team Leaders.
@@ -0,0 +1,52 @@
# Source-to-Sink Taint Researcher Agent
## User Prompt
You are reviewing the source code of **{target}** to prove a NOVEL, reachable, unsanitized taint path from an untrusted entrypoint to a dangerous sink (injection / RCE / SSRF / deserialization / path traversal / prototype pollution).
**Recon Context:**
{recon_json}
The relevant source files are provided to you below the methodology.
**METHODOLOGY:**
### 1. Pin version & confirm novelty first
- `git -C {target} rev-parse HEAD`; record `name@version`
- Before tracing, confirm the class is not already a patched CVE at this version: scan `SECURITY.md`, `CHANGELOG*`, `git log --oneline | grep -iE "cve|sanitiz|injection|ssrf|traversal|deser"`, GHSA/NVD for the ecosystem
### 2. Identify the SOURCE (untrusted input)
- Request params/body/headers/cookies, path params, uploaded files, env, CLI args, IPC/queue messages, config a lower-trust actor controls
- Quote the exact line where the value enters (file:line)
### 3. Identify the SINK
- Command exec, raw SQL/NoSQL, deserializer, dynamic file path, outbound URL, template/eval, reflected HTML/JS, object-key assignment (prototype pollution)
- Locate with `grep -rnE "<sink-shape>" .`; quote the exact line (file:line)
### 4. Trace the DATAFLOW end to end
- Follow the value hop by hop: assignments, function params, struct fields, closures, await boundaries
- Quote EVERY hop (file:line). Identify any validation/escaping/allowlist encountered and PROVE it is absent, insufficient, or bypassable on this path (wrong order, partial, wrong charset, missing recursion)
### 5. Confirm exploitability (static)
- State the concrete attacker-controlled value that reaches the sink unmodified (or modified in an attacker-useful way)
- Explain why existing controls do not stop it; explain what the sink does with it (RCE, data read/write, SSRF, file disclosure, pollution)
### 6. Report Format
For each CONFIRMED finding:
```
FINDING:
- Title: <class> source-to-sink at [file:line]
- Severity: Critical
- CWE: CWE-20
- Endpoint: [file:line of the source]
- Vector: [source (file:line) -> hop (file:line) -> ... -> sink (file:line)]
- Payload: [benign crafted input demonstrating the reachable path + the exact call chain]
- Evidence: [every hop quoted + sink quoted + version/commit pinned + novel: why this path is not an already-fixed CVE + checked-against: <CVE/GHSA/commit/CHANGELOG>]
- Impact: [concrete: RCE / SQLi data exfil / SSRF to metadata / arbitrary file read-write / prototype pollution -> ...]
- Remediation: [validate/parameterize/escape at the sink; safe loader; allowlist; normalize-before-check]
```
- Write a static-derived PoC (crafted input + a failing unit test that drives the source and asserts the sink is reached) to `$NEUROSPLOIT_POCS/{target}-taint-<class>.{ext}` and cite it. Mark it SOURCE-DERIVED.
## System Prompt
You are a senior AppSec vulnerability researcher performing deep source-to-sink taint analysis for responsible, novelty-gated research. Report ONLY a finding you can PROVE in the provided code: an untrusted source, a dangerous sink, and a reachable, unsanitized dataflow connecting them — with EVERY hop quoted (file:line). Pin the version/commit. The finding must be NOVEL: state `novel: <why>` and `checked-against: <CVE/GHSA/commit/CHANGELOG>`; a path that merely re-walks an already-patched CVE for this version is not reportable unless recast as a concrete bypass/variant. Prove that controls on the path are absent or defeatable — do not assume sanitization you cannot see, and do not assume a control works that you cannot trace. No speculation and no live/HTTP/network claims — reason strictly about the source and git metadata. If any hop, the source, or the sink is missing from the snippet, say the path is unconfirmed rather than guess.
Credits: Joas A Santos & Red Team Leaders.
+10 -2
View File
@@ -318,6 +318,9 @@ enum Cmd {
/// Economy preset for a short, low-cost review (see `run --quick`).
#[arg(long)]
quick: bool,
/// Vulnerability-research mode: hunt a NOVEL, CVE-reportable bug (known-CVE dedup + patch-diff variant analysis).
#[arg(long)]
research: bool,
#[arg(long)]
offline: bool,
#[arg(long)]
@@ -359,6 +362,9 @@ enum Cmd {
/// Economy preset for a short, low-cost test (see `run --quick`).
#[arg(long)]
quick: bool,
/// Vulnerability-research mode (see `whitebox --research`).
#[arg(long)]
research: bool,
#[arg(long)]
offline: bool,
#[arg(long)]
@@ -880,7 +886,7 @@ async fn main() -> anyhow::Result<()> {
let ig = harness::integrations::Integrations::load(&repl::proj_dir());
post_integrations(&ig, &url, &out, jira, false, None).await;
}
Cmd::Whitebox { path, models, max_agents, vote_n, chain_depth, recon, quick, offline, subscription, jira, only, verbose } => {
Cmd::Whitebox { path, models, max_agents, vote_n, chain_depth, recon, quick, research, offline, subscription, jira, only, verbose } => {
let path = resolve_source(&base, &path)?; // local path OR github URL/owner/repo
let mut cfg = RunConfig::new(&path);
cfg.max_agents = max_agents;
@@ -891,6 +897,7 @@ async fn main() -> anyhow::Result<()> {
cfg.subscription = subscription;
cfg.verbose = verbose;
cfg.pinned = parse_only(&only);
cfg.research = research;
if quick { apply_quick(&mut cfg); }
if !models.is_empty() {
cfg.models = models;
@@ -900,7 +907,7 @@ async fn main() -> anyhow::Result<()> {
let ig = harness::integrations::Integrations::load(&repl::proj_dir());
post_integrations(&ig, &path, &out, jira, false, None).await;
}
Cmd::Greybox { repo, url, models, creds, focus, max_agents, vote_n, chain_depth, recon, quick, offline, subscription, mcp, only, verbose } => {
Cmd::Greybox { repo, url, models, creds, focus, max_agents, vote_n, chain_depth, recon, quick, research, offline, subscription, mcp, only, verbose } => {
let repo = resolve_source(&base, &repo)?; // local path OR github URL/owner/repo
let url = if url.starts_with("http") { url } else { format!("https://{url}") };
let mut cfg = RunConfig::new(&url);
@@ -914,6 +921,7 @@ async fn main() -> anyhow::Result<()> {
cfg.verbose = verbose;
cfg.instructions = focus;
cfg.pinned = parse_only(&only);
cfg.research = research;
if quick { apply_quick(&mut cfg); }
if !models.is_empty() {
cfg.models = models;
+14 -2
View File
@@ -152,7 +152,7 @@ pub(crate) const ACCEPTED: &[&str] = &[
"/history", "/idle", "/inscope", "/instructions", "/integration", "/integrations", "/key", "/log",
"/logs", "/mcp", "/memory", "/model", "/models", "/objective", "/objectives", "/observe",
"/observe-only", "/offline",
"/onboard", "/only", "/oos", "/outofscope", "/policy", "/providers", "/proxy", "/quick", "/economy", "/eco", "/q", "/quit", "/recon",
"/onboard", "/only", "/oos", "/outofscope", "/policy", "/providers", "/proxy", "/research", "/quick", "/economy", "/eco", "/q", "/quit", "/recon",
"/pause", "/repo", "/report", "/results", "/resume", "/retest", "/revalidate", "/run", "/runs",
"/scope", "/scope-out", "/show", "/status", "/stop", "/sub", "/subscription", "/target",
"/temp-email", "/tempmail", "/theme", "/timeout", "/ua", "/url", "/useragent", "/validate",
@@ -163,7 +163,7 @@ pub(crate) const ACCEPTED: &[&str] = &[
const COMMANDS: &[&str] = &[
"/help", "/onboard", "/show", "/config", "/providers", "/model", "/key", "/sub", "/target",
"/repo", "/auth", "/creds", "/focus", "/objective", "/scope-out", "/attach", "/context", "/mcp", "/offline",
"/quick", "/economy", "/eco", "/votes", "/chain", "/recon", "/tempmail", "/timeout", "/proxy", "/burp", "/ua", "/agents", "/only", "/theme", "/clear", "/run", "/stop", "/pause", "/continue", "/runs", "/results", "/report",
"/research", "/quick", "/economy", "/eco", "/votes", "/chain", "/recon", "/tempmail", "/timeout", "/proxy", "/burp", "/ua", "/agents", "/only", "/theme", "/clear", "/run", "/stop", "/pause", "/continue", "/runs", "/results", "/report",
"/status", "/logs", "/diff", "/retest", "/validate", "/finding", "/expand", "/integrations",
"/memory", "/forget", "/graph", "/inscope", "/observe", "/guardrail", "/policy",
"/capability", "/audit", "/quit",
@@ -268,6 +268,7 @@ struct Session {
max_agents: usize,
chain_depth: usize,
recon_intensity: usize,
research: bool,
/// Opt-in disposable email (mail.tm) for register flows needing a confirmation code.
temp_email: bool,
/// Idle guardrail: stop a run if no NEW finding lands in this many seconds
@@ -320,6 +321,7 @@ impl Default for Session {
max_agents: 0,
chain_depth: 2,
recon_intensity: 3,
research: false,
temp_email: false,
idle_secs: 300, // 5-minute idle guardrail by default
proxy: None,
@@ -846,6 +848,13 @@ pub async fn repl(base: &Path, auth: SessionAuth) -> anyhow::Result<()> {
if arg.is_empty() { println!(" recon intensity: {} ({}) — set with /recon <1-4> [1 quick · 2 standard · 3 deep · 4 exhaustive]", s.recon_intensity, lvl(s.recon_intensity)); }
else { s.recon_intensity = arg.parse::<usize>().unwrap_or(s.recon_intensity).clamp(1, 4); println!(" recon intensity: {} ({}) — more rounds, more enumeration, auto-installs tools", s.recon_intensity, lvl(s.recon_intensity)); }
}
"/research" => {
match arg.trim() {
"on" | "true" | "1" => { s.research = true; println!(" \x1b[1;36m🔬 research mode ON\x1b[0m — whitebox/greybox will hunt a NOVEL, CVE-reportable bug (known-CVE dedup + patch-diff variant analysis)"); }
"off" | "false" | "0" => { s.research = false; println!(" research mode off"); }
_ => println!(" research mode: {} — /research on|off (for whitebox/greybox: find a new CVE, not a known one)", if s.research { "\x1b[36mon\x1b[0m" } else { "\x1b[2moff\x1b[0m" }),
}
}
"/quick" | "/economy" | "/eco" => {
// Economy preset for a short, low-cost test — the single switch
// for "fast and cheap" instead of tuning each knob. The big
@@ -1471,6 +1480,7 @@ async fn run(base: &Path, s: &Session, history: &mut Vec<RunRecord>) {
cfg.vote_n = s.vote_n;
cfg.chain_depth = s.chain_depth;
cfg.recon_intensity = s.recon_intensity;
cfg.research = s.research;
cfg.temp_email = s.temp_email;
cfg.proxy = s.proxy.clone();
cfg.user_agent = s.user_agent.clone();
@@ -1560,6 +1570,7 @@ async fn start_background(base: &Path, s: &Session, reader: &mut Reader,
cfg.vote_n = s.vote_n;
cfg.chain_depth = s.chain_depth;
cfg.recon_intensity = s.recon_intensity;
cfg.research = s.research;
cfg.temp_email = s.temp_email;
cfg.proxy = s.proxy.clone();
cfg.user_agent = s.user_agent.clone();
@@ -2236,6 +2247,7 @@ fn help() {
h("/votes <n>", "number of validator votes per finding");
h("/chain <n>", "attack-chain depth (post-exploitation pivots; 0 = off)");
h("/recon <1-4>", "recon intensity: 1 quick · 2 standard · 3 deep · 4 exhaustive (installs tools)");
h("/research", "whitebox/greybox: hunt a NOVEL, CVE-reportable bug (known-CVE dedup + patch-diff variant analysis)");
h("/quick", "economy preset: short, low-cost run (1 voter · 1 chain round · light recon · ≤6 agents)");
h("/tempmail on|off", "opt-in disposable inbox (mail.tm) to read a register confirmation code");
h("/timeout <min>", "idle guardrail: stop if no new finding in <min> (0 = off)");
+46 -2
View File
@@ -623,6 +623,34 @@ const WHITEBOX_DOCTRINE: &str = "MODE: WHITE-BOX STATIC SOURCE REVIEW. You are r
- Repro PoC (optional but valued): when a finding warrants it, WRITE a proof/repro script to $NEUROSPLOIT_POCS — e.g. the exact malicious input + the request/CLI call that would trigger the sink, or a unit-style harness exercising the vulnerable function — with a header comment (file:line it proves, how to run). Cite the PoC path in the evidence. Mark clearly that it demonstrates the code path (static-derived), not a live hit.\n\
- Calibrate: High/Critical only when the sink is reachable and exploitable from untrusted input; guarded/unreachable code is Low or a lead.\n\n";
/// Vulnerability-RESEARCH doctrine: turn a source review into a hunt for a
/// NOVEL, CVE-reportable bug. Prepended (after the white-box doctrine) when the
/// operator runs research mode (`--research`, `/research`, or a focus/objective
/// that asks for a new CVE / 0-day / patch bypass). The whole point is NOVELTY:
/// do not re-report a known CVE — use the known ones as a map to find what the
/// fix missed, a variant, or a reintroduction.
const WHITEBOX_RESEARCH_DOCTRINE: &str = "MISSION: VULNERABILITY RESEARCH — find a NOVEL, CVE-reportable issue in this codebase, not a known one.\n\
- PIN THE VERSION FIRST: read the version (package.json/VERSION/__init__/composer.json/go.mod/tag) and the exact commit. Every claim is against THIS version/commit; note it in evidence.\n\
- RESEARCH KNOWN CVEs (de-duplicate): before reporting anything, build the set of ALREADY-KNOWN issues for this project+version — read SECURITY.md, CHANGELOG/release notes, the security advisories (GHSA), CVE/NVD, the issue tracker and recent security commits (`git log --oneline`, grep messages for CVE/security/fix/vuln/XSS/RCE/injection). A finding that matches a known CVE for this version is NOT novel — drop it or recast it ONLY as a patch-bypass/variant (below). State which known CVEs you checked against.\n\
- PATCH-DIFF / N-DAY -> 0-DAY (the highest-yield path): take a recent SECURITY fix (its commit) and study the diff. Ask: did the patch fix the ROOT CAUSE or just one path? Look for (1) incomplete fixes — another reachable sink the patch did not cover, a bypass of the new check (different encoding, type juggling, alternate parser, case/Unicode, second-order input); (2) the SAME bug pattern elsewhere in the tree (variant analysis — grep the fixed sink's shape across the repo); (3) reintroduction in a later commit. A proven bypass of an existing patch IS novel and reportable.\n\
- SOURCE->SINK with reachability: trace untrusted input (request params, headers, body, deserialized objects, file names, env, IPC, config) to a dangerous sink (SQL/exec/eval/template/path/deserialize/SSRF/XXE/prototype/unsafe-reflection). Only call it a vuln when the path is REACHABLE from an untrusted entrypoint without an effective sanitizer; record the full path `entry -> … -> sink`.\n\
- NOVELTY GATE (strict): report a finding ONLY if (a) it does not match a known CVE for this version, OR (b) it is a concrete bypass/variant of a patched issue. For each, state explicitly: 'novel: <why>' and 'checked-against: <CVEs/advisories/commits>'. No speculation — high-confidence, evidence-backed only.\n\
- WRITE A PoC: produce a minimal, SAFE proof (a failing unit test, a crafted input + the exact call reaching the sink, or a request) to $NEUROSPLOIT_POCS and cite it. Where a running instance is available (greybox), CONFIRM the source-derived bug dynamically — but never run destructive payloads.\n\
- REPORT for disclosure: each finding carries file:line, root-cause analysis, the version/commit, the novelty justification, CVSS vector, PoC path, suggested fix, and the project's disclosure channel (SECURITY.md / security@). Prefer a few solid, novel, reportable bugs over a long list of known or speculative ones.\n\n";
/// True when the engagement is asking for vulnerability research / a new CVE,
/// either via the explicit flag or from natural-language focus/objective.
fn is_research_intent(cfg: &RunConfig) -> bool {
if cfg.research { return true; }
let hay = format!("{} {}",
cfg.instructions.clone().unwrap_or_default(),
cfg.objective.clone().unwrap_or_default()).to_lowercase();
["new cve", "novel", "0-day", "0day", "zero-day", "zero day", "patch bypass",
"patch-bypass", "variant analysis", "n-day", "nday", "reportable", "cve research",
"vulnerability research", "find a cve", "nova cve", "pesquisa de vuln"]
.iter().any(|k| hay.contains(k))
}
/// Methodology directions for a modern JS SPA backed by a REST/GraphQL API
/// (Angular/React/Vue front + Node/Express-style API — the shape of OWASP Juice
/// Shop and many real apps). These are DIRECTIONS on HOW to hunt each vuln class,
@@ -1167,6 +1195,10 @@ pub async fn run_whitebox(cfg: RunConfig, lib: &Library, pool: &ModelPool, tx: S
let context = collect_repo_context(Path::new(&cfg.target), 200, 120_000);
let bytes = context.len();
let research = is_research_intent(&cfg);
if research {
let _ = tx.send("notify: 🔬 research mode — hunting a NOVEL, CVE-reportable issue (known-CVE dedup + patch-diff variant analysis)".into()).await;
}
let _ = tx.send(format!("collected {} bytes of source context", bytes)).await;
if bytes == 0 {
let _ = tx.send("no readable source found at the given path".into()).await;
@@ -1214,7 +1246,12 @@ pub async fn run_whitebox(cfg: RunConfig, lib: &Library, pool: &ModelPool, tx: S
);
// Prepend the white-box doctrine so code agents stay in static
// source-review mode and never hallucinate live/black-box actions.
let sys = format!("{}{}", WHITEBOX_DOCTRINE, ag.system);
// In research mode, add the novelty/patch-diff hunting doctrine.
let sys = if research {
format!("{}{}{}", WHITEBOX_DOCTRINE, WHITEBOX_RESEARCH_DOCTRINE, ag.system)
} else {
format!("{}{}", WHITEBOX_DOCTRINE, ag.system)
};
match pool.complete_routed(Task::Exploit, &ag.name, &sys, &user).await {
Ok((m, text)) => {
let f = extract_findings(&text, &ag.name);
@@ -1267,6 +1304,10 @@ pub async fn run_greybox(cfg: RunConfig, lib: &Library, pool: &ModelPool, tx: Se
// ---- 2. Review the source for leads -------------------------------
let context = collect_repo_context(Path::new(&repo), 200, 90_000);
let _ = tx.send(format!("collected {} bytes of source for code review", context.len())).await;
let gb_research = is_research_intent(&cfg);
if gb_research {
let _ = tx.send("notify: 🔬 research mode — code review hunts a novel, CVE-reportable bug, then confirms it live".into()).await;
}
let mut rl = cfg.rl_path.as_ref().map(|p| RlState::load(Path::new(p))).unwrap_or_default();
let mut code_leads = String::new();
@@ -1284,7 +1325,10 @@ pub async fn run_greybox(cfg: RunConfig, lib: &Library, pool: &ModelPool, tx: Se
where endpoint is file:line.",
ag.user.replace("{target}", "the repository").replace("{recon_json}", "{}"), ctx
);
match pool.complete_routed(Task::Select, &ag.name, &ag.system, &user).await {
// Research mode steers the code-review half of greybox too:
// find a novel, reportable bug, then confirm it live.
let sys = if gb_research { format!("{}{}", WHITEBOX_RESEARCH_DOCTRINE, ag.system) } else { ag.system.clone() };
match pool.complete_routed(Task::Select, &ag.name, &sys, &user).await {
Ok((_, text)) => { let f = extract_findings(&text, &ag.name);
let _ = txc.send(format!("review {} → {} lead(s)", ag.name, f.len())).await; f }
Err(_) => vec![],
@@ -213,6 +213,12 @@ pub struct RunConfig {
/// more recon rounds, more active enumeration, and auto-installing tools.
#[serde(default = "default_recon")]
pub recon_intensity: usize,
/// Vulnerability-research mode: hunt for NOVEL, CVE-reportable issues in a
/// source repo — research known CVEs/advisories to de-duplicate, do
/// patch-diff variant analysis (incomplete-fix bypasses, sibling sinks),
/// and gate strictly on novelty. Steers whitebox/greybox.
#[serde(default)]
pub research: bool,
/// Opt-in: when the app requires email confirmation to register, allow the
/// agent to use a free disposable-inbox API (mail.tm) to read the code/link.
/// Off by default. Account creation is still capped by the safety guardrail.
@@ -315,6 +321,7 @@ impl RunConfig {
repo: None,
pinned: Vec::new(),
chain_depth: 2,
research: false,
proxy: None,
user_agent: None,
recon_intensity: 3,
+1
View File
@@ -533,6 +533,7 @@ async function startExploitation() {
chainDepth: Number($('#fieldChain').value),
recon: Number($('#fieldRecon').value),
quick: $('#fieldQuick') ? $('#fieldQuick').checked : false,
research: $('#fieldResearch') ? $('#fieldResearch').checked : false,
subscription: state.authMode === 'subscription',
mcp: $('#fieldMcp').checked,
agents: [...state.selected],
+3
View File
@@ -221,6 +221,9 @@
<div class="section-desc">Optional. Left on <em>unlimited</em>, the run behaves exactly as it always has — full depth, no cap.</div>
</div>
<div class="field-row">
<div class="check-row"><input type="checkbox" id="fieldResearch" /> <label for="fieldResearch"><b>🔬 Research mode</b> — hunt a novel, CVE-reportable bug (whitebox/greybox)</label>
<div class="field-help">Known-CVE de-dup + patch-diff variant analysis. Best with a source repo; reports only genuinely new issues (or a concrete patch bypass).</div>
</div>
<div class="check-row check-quick"><input type="checkbox" id="fieldQuick" /> <label for="fieldQuick"><b>⚡ Quick mode</b> — short, low-cost test</label>
<div class="field-help">Economy preset: 1 voter, 1 chain round, light recon, ≤6 agents, eco budget. The big token saver. Wins over the settings below.</div>
</div>
+2
View File
@@ -882,6 +882,7 @@ function sanitizeLaunch(body) {
sandbox: !!body.sandbox,
typesafe: body.typesafe,
quick: !!body.quick,
research: !!body.research,
};
}
@@ -902,6 +903,7 @@ function buildReplScript(body) {
if (body.objective) lines.push(`/objective ${body.objective}`);
if (body.outOfScope) lines.push(`/scope-out ${body.outOfScope}`);
if (body.creds) lines.push(`/creds ${body.creds}`);
if (body.research) lines.push(`/research on`);
lines.push((body.agents || []).length ? `/only ${body.agents.join(',')}` : '/only clear');
// Economy preset last, so it wins over the per-knob settings above.
if (body.quick) lines.push('/quick');