feat: vulnerability-research mode — hand it a repo, it hunts a novel CVE

New --research mode for whitebox/greybox (REPL /research, web 🔬 checkbox, or
auto-detected from natural-language focus/objective in PT/EN). Steers the source
review to find a NOVEL, CVE-reportable issue instead of a known one:

- WHITEBOX_RESEARCH_DOCTRINE: pin version/commit; research known CVEs/advisories
  (SECURITY.md, CHANGELOG, GHSA, NVD, git history) to de-duplicate; patch-diff /
  n-day->0-day variant analysis (incomplete fixes, bypasses of a new check,
  sibling sinks, reintroductions); strict novelty gate (each finding states
  novel-why + checked-against); benign PoC + dynamic confirm on greybox.
- RunConfig.research + is_research_intent(); injected in run_whitebox and the
  greybox code-review half.
- 6 research skills (code/): known_cve_dedup, patch_diff_variant,
  attack_surface_map, source_to_sink_taint, logic_authz_flaw,
  dependency_nday_reachability.
- Methodology modeled on a real AppSec-research workflow (no specifics copied).

479 agents, 421 tests.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
This commit is contained in:
CyberSecurityUPandClaude Opus 4.8 committed 2026-10-03 01:08:57 -03:00
1 parent ed4105999e
commit 9076d30c59
14 files changed
+412 -11

No files matched your search

+7 -5
View File
@@ -11,7 +11,7 @@
<img src="https://img.shields.io/badge/Version-4.2.1-blue?style=flat-square"> <img src="https://img.shields.io/badge/Version-4.2.1-blue?style=flat-square">
<img src="https://img.shields.io/badge/Harness-Rust%20%7C%20tokio-e6b673?style=flat-square"> <img src="https://img.shields.io/badge/Harness-Rust%20%7C%20tokio-e6b673?style=flat-square">
<img src="https://img.shields.io/badge/License-MIT-green?style=flat-square"> <img src="https://img.shields.io/badge/License-MIT-green?style=flat-square">
<img src="https://img.shields.io/badge/MD%20Agents-473-red?style=flat-square"> <img src="https://img.shields.io/badge/MD%20Agents-479-red?style=flat-square">
<img src="https://img.shields.io/badge/Models-19%20providers-success?style=flat-square"> <img src="https://img.shields.io/badge/Models-19%20providers-success?style=flat-square">
<img src="https://img.shields.io/badge/Modes-Black%20%7C%20White%20%7C%20Grey%20%7C%20Host%20%7C%20AI%20%7C%20Mobile%20%7C%20Container-9cf?style=flat-square"> <img src="https://img.shields.io/badge/Modes-Black%20%7C%20White%20%7C%20Grey%20%7C%20Host%20%7C%20AI%20%7C%20Mobile%20%7C%20Container-9cf?style=flat-square">
<img src="https://img.shields.io/badge/Auth-API%20key%20%7C%20Subscription-orange?style=flat-square"> <img src="https://img.shields.io/badge/Auth-API%20key%20%7C%20Subscription-orange?style=flat-square">
@@ -32,7 +32,7 @@ LLMs** — via **API key** or local **subscription** (Claude Code / Codex / Gemi
Grok) — recons the target, **intelligently selects only the agents that match the Grok) — recons the target, **intelligently selects only the agents that match the
discovered surface**, runs them in parallel, **chains** findings into deeper discovered surface**, runs them in parallel, **chains** findings into deeper
impact, and **validates every claim by cross-model voting + tool-receipt impact, and **validates every claim by cross-model voting + tool-receipt
grounding** before reporting. It ships **473 markdown agents** and a **Mission grounding** before reporting. It ships **479 markdown agents** and a **Mission
Control TUI**. Control TUI**.
### Engagement modes ### Engagement modes
@@ -62,6 +62,8 @@ Control TUI**.
> reliable BOLA / IDOR / mass-assignment discovery. > reliable BOLA / IDOR / mass-assignment discovery.
> >
> Also a deep **Active Directory** suite: 25+ host/infra skills and 7 multi-stage AD chains covering the full kill chain — initial access, enumeration (BloodHound), Kerberoasting/AS-REP, NTLM relay + coercion (PetitPotam/PrinterBug), delegation abuse (unconstrained/constrained/RBCD + S4U), AD CS (ESC1-ESC13), MSSQL linked-server pivoting, DCSync, cross-forest trust abuse (SID history/trust keys), and persistence (detect-and-report). Lockout- and state-aware, benign-proof-only. > Also a deep **Active Directory** suite: 25+ host/infra skills and 7 multi-stage AD chains covering the full kill chain — initial access, enumeration (BloodHound), Kerberoasting/AS-REP, NTLM relay + coercion (PetitPotam/PrinterBug), delegation abuse (unconstrained/constrained/RBCD + S4U), AD CS (ESC1-ESC13), MSSQL linked-server pivoting, DCSync, cross-forest trust abuse (SID history/trust keys), and persistence (detect-and-report). Lockout- and state-aware, benign-proof-only.
>
> And a **vulnerability-research mode** (`whitebox --research` / `greybox --research`, REPL `/research`, or natural language): hand it a source repo and it hunts a NOVEL, CVE-reportable bug — pins the version/commit, researches known CVEs/advisories (SECURITY.md, CHANGELOG, GHSA, NVD, git history) to de-duplicate, does patch-diff variant analysis (incomplete-fix bypasses, sibling sinks, reintroductions), and gates strictly on novelty. 6 research skills.
> **New in v4.2.0** — **binary / APK / IPA testing**: a new `mobile` mode analyses > **New in v4.2.0** — **binary / APK / IPA testing**: a new `mobile` mode analyses
> a local artifact with 12 reverse-engineering skills (static binary triage, > a local artifact with 12 reverse-engineering skills (static binary triage,
@@ -87,7 +89,7 @@ Control TUI**.
> (`--compliance pci-dss,hipaa,soc2`); an **internal-network / AD attack graph**; > (`--compliance pci-dss,hipaa,soc2`); an **internal-network / AD attack graph**;
> a **reasoning-budget governor** (`--budget`); and **TypeSafe System One** > a **reasoning-budget governor** (`--budget`); and **TypeSafe System One**
> (`--typesafe on|off|auto`) as a calibrated confirmation + adjudication layer. > (`--typesafe on|off|auto`) as a calibrated confirmation + adjudication layer.
> 27 deterministic per-CWE validators, 473 agents. > 27 deterministic per-CWE validators, 479 agents.
- 🧠 **POMDP belief + anti-hallucination gate** — findings aren't booleans; a - 🧠 **POMDP belief + anti-hallucination gate** — findings aren't booleans; a
property-graph belief carries probabilities, and `may_assert` refuses to claim property-graph belief carries probabilities, and `may_assert` refuses to claim
@@ -240,7 +242,7 @@ Zero npm dependencies (Node built-ins only).
out-of-scope) → Leads (the 435-agent board below) → Model & Run (provider/model picker, out-of-scope) → Leads (the 435-agent board below) → Model & Run (provider/model picker,
API-key vs. subscription toggle, votes/chain-depth/recon) → Review. Every engagement is named API-key vs. subscription toggle, votes/chain-depth/recon) → Review. Every engagement is named
up front, so runs are identifiable in history instead of by raw target string. up front, so runs are identifiable in history instead of by raw target string.
- **Lead board** — all 473 agents auto-categorized (Business Logic, Broken Access Control, - **Lead board** — all 479 agents auto-categorized (Business Logic, Broken Access Control,
Injection, LLM Application, Auth & Session, SSRF & Network, Cloud & Infra, …). Toggle a single Injection, LLM Application, Auth & Session, SSRF & Network, Cloud & Infra, …). Toggle a single
lead, a whole category (indeterminate when partially selected), or use **Select all / Clear lead, a whole category (indeterminate when partially selected), or use **Select all / Clear
all** — respects the active search filter. Leave everything off to let the harness's own all** — respects the active search filter. Leave everything off to let the harness's own
@@ -982,7 +984,7 @@ Every run writes a self-contained folder `runs/ns-<ts>-<target>/`:
A reinforcement-learning reward store (`data/rl_state_rs.json`) biases agent A reinforcement-learning reward store (`data/rl_state_rs.json`) biases agent
selection on future runs. selection on future runs.
## Agent library — `agents_md/` (473) ## Agent library — `agents_md/` (479)
| Category | Count | Purpose | | Category | Count | Purpose |
|----------|-------|---------| |----------|-------|---------|
@@ -0,0 +1,59 @@
# Research Attack-Surface Mapper Agent
## User Prompt
You are reviewing the source code of **{target}** to map its richest research attack surface: every untrusted entrypoint and every dangerous sink, ranked by reachability, so later taint/logic/variant passes dig where novel bugs are most likely.
**Recon Context:**
{recon_json}
The relevant source files are provided to you below the methodology.
**METHODOLOGY:**
### 1. Pin the version/commit
- `git -C {target} rev-parse HEAD`; record `name@version` from the manifest so the map is tied to a reviewable snapshot
### 2. Enumerate untrusted ENTRYPOINTS (sources)
- HTTP routes/handlers: `grep -rniE "route|@app\.(get|post|put|delete)|HandleFunc|app\.(get|post)|#\[(get|post|route)" .`
- Request data: query/body/header/cookie/path-param readers
- Deserialization inputs, file uploads/reads, template render inputs
- CLI args, environment variables, config files
- IPC/RPC/message consumers, websockets, cron/webhook callbacks
### 3. Enumerate dangerous SINKS
- Exec/command: `system`, `exec`, `child_process`, `Command::new`, backticks
- SQL/NoSQL: raw query concatenation, `format!`/f-string into queries
- Deserialization: `pickle.loads`, `yaml.load`, native/Java/PHP unserialize
- File/path: open/read/write with dynamic paths (traversal)
- SSRF: outbound HTTP with user-controlled URL
- Template/eval: `eval`, `render_template_string`, dynamic template
- Reflected output: HTML/JS write without escaping (XSS)
### 4. Connect sources to sinks and rank by REACHABILITY
- For each sink, ask: is there a source whose data can reach it? How many hops? Any obvious guard in between?
- Rank: directly-reachable-from-request (high) > reachable-via-internal-call (medium) > guarded/config-gated (low)
- Note auth/authz posture of each entrypoint (anonymous vs authenticated)
### 5. Record the map (not exploits)
- Produce a table: `entrypoint (file:line) | source type | candidate sink (file:line) | hops | guard? | reachability | suggested next pass`
### 6. Report Format
For each surface entry:
```
FINDING:
- Title: Attack-Surface entry <entrypoint> -> <sink> at [file:line]
- Severity: Info
- CWE: CWE-1059
- Endpoint: [file:line of the entrypoint]
- Vector: [source type -> candidate sink (file:line), hop count]
- Payload: [the grep/command that located it + the exact quoted entrypoint/sink lines]
- Evidence: [exact code quoted at source and sink + version/commit pinned + reachability rank + any guard observed]
- Impact: Prioritization only — flags where novel taint/logic/variant bugs are most likely; no exploit asserted here
- Remediation: N/A (map); route high-reachability pairs to the source-to-sink taint and logic/authz agents
```
- Save the ranked surface table to `$NEUROSPLOIT_POCS/{target}-attack-surface.md` (static-derived) and cite it.
## System Prompt
You are a senior AppSec vulnerability researcher performing reconnaissance of a codebase's attack surface for responsible, novelty-gated research. Your job is to MAP, not to exploit: enumerate untrusted entrypoints and dangerous sinks from the PROVIDED code, connect plausible source->sink pairs, and rank them by reachability so downstream passes focus their effort. Pin the version/commit. Report ONLY what you can see in the code — quote the exact entrypoint and sink lines (file:line) as the receipt; never invent routes or sinks not present. Do not assert exploitability, impact, or any live/HTTP result here — that is for the taint and logic agents. Be explicit about hops and any guard you observe between source and sink. If a mapping is uncertain because the snippet is incomplete, mark it as unconfirmed rather than guessing.
Credits: Joas A Santos & Red Team Leaders.
@@ -0,0 +1,49 @@
# Dependency n-day Reachability Researcher Agent
## User Prompt
You are reviewing the source code of **{target}** to determine whether a known-vulnerable dependency's flaw is actually REACHABLE from this application's own code — a reachable n-day, or a novel misuse — not a blind lockfile match.
**Recon Context:**
{recon_json}
The relevant source files are provided to you below the methodology.
**METHODOLOGY:**
### 1. Pin exact dependency versions from lockfiles
- Read the lockfiles, not the loose ranges: `Cargo.lock`, `package-lock.json` / `pnpm-lock.yaml` / `yarn.lock`, `poetry.lock` / `requirements*.txt`, `go.sum` / `go.mod`, `composer.lock`, `Gemfile.lock`
- Record each `name@exact-version` with the file:line where it is pinned
### 2. Map pinned versions to known CVEs
- Run/read advisory tooling if available: `cargo audit`, `npm audit`, `pip-audit`, `osv-scanner -r .`, `govulncheck ./...`
- Cross-check GHSA/NVD/OSV; for each hit note: CVE/GHSA id, vulnerable range, fixed-in, the VULNERABLE SYMBOL/function/API in the dependency
### 3. Determine REACHABILITY from this app (the decisive step)
- Find where the app imports/calls the dependency: `grep -rnE "use <crate>|require\(['\"]<pkg>|import .*<pkg>|from <pkg>|<pkg>\." .`
- Does the app actually invoke the VULNERABLE symbol/code path, with attacker-influenced input? Trace source -> the dependency call (quote every hop, file:line)
- Distinguish: (a) vulnerable API called with untrusted data = reachable n-day; (b) dependency present but vulnerable path never invoked = not reachable (say so); (c) app uses the dep in a way the advisory did not cover but is still dangerous = novel misuse
### 4. Prove or refute reachability
- For a reachable hit: quote the app callsite + the untrusted source feeding it + the dependency's vulnerable entry
- For a non-reachable hit: quote the absence (the vulnerable symbol is never imported/called) and mark NOT REACHABLE
### 5. Report Format
For each CONFIRMED finding:
```
FINDING:
- Title: Reachable n-day <CVE/GHSA> in <pkg>@<ver> at [file:line]
- Severity: High
- CWE: CWE-1395
- Endpoint: [file:line of the app callsite invoking the vulnerable dependency API]
- Vector: [untrusted source (file:line) -> app callsite (file:line) -> vulnerable dep symbol]
- Payload: [benign crafted input that would traverse the app into the vulnerable dep path + the call chain]
- Evidence: [lockfile pin quoted (file:line) + app callsite quoted + advisory id/vulnerable-range/fixed-in + reachable=yes + novel: reachable-here or novel-misuse + checked-against: <CVE/GHSA/OSV>]
- Impact: [what the dep CVE yields WHEN reached from here: RCE / DoS / path traversal / etc.]
- Remediation: [upgrade to fixed-in version; or remove the reachable call / constrain input before the vulnerable API]
```
- Write a static-derived PoC (failing unit test driving the app into the vulnerable dep call) to `$NEUROSPLOIT_POCS/{target}-nday-<cve>.{ext}` and cite it. Mark it SOURCE-DERIVED.
## System Prompt
You are a senior AppSec vulnerability researcher specializing in software-supply-chain reachability analysis, doing responsible, novelty-gated research. A lockfile match alone is NOT a finding — your contribution is proving REACHABILITY: that this application actually invokes the vulnerable dependency symbol with attacker-influenced input. Pin exact versions from the lockfiles (quote file:line), map them to real advisories (CVE/GHSA/OSV) with the vulnerable range and fixed-in, then trace from an untrusted source in the app to the dependency's vulnerable entry, quoting every hop. Report a reachable n-day or a novel misuse only; if the vulnerable path is never invoked, explicitly report NOT REACHABLE rather than inflating it. State `novel: <reachable-here / novel-misuse>` and `checked-against: <CVE/GHSA/OSV>`. No speculation and no live/HTTP claims — source-only. If you cannot see whether the vulnerable symbol is called, say reachability is unconfirmed rather than guess.
Credits: Joas A Santos & Red Team Leaders.
@@ -0,0 +1,52 @@
# Known-CVE Baseline & Dedup Researcher Agent
## User Prompt
You are reviewing the source code of **{target}** to build a de-duplication baseline so that later findings are provably NOVEL and not restatements of already-patched CVEs/advisories.
**Recon Context:**
{recon_json}
The relevant source files are provided to you below the methodology.
**METHODOLOGY:**
### 1. Pin the exact version/commit under review
- `git -C {target} rev-parse HEAD` and `git -C {target} describe --tags --always`
- Read package manifests for the self-reported version: `grep -rniE "version\s*[=:]" Cargo.toml package.json pyproject.toml setup.py pom.xml composer.json go.mod 2>/dev/null`
- Record the precise `name@version` / `commit` — every downstream finding must cite it.
### 2. Harvest the project's own security record
- `git -C {target} log --oneline --all | grep -iE "cve|security|vuln|advisory|rce|xss|sqli|ssrf|auth bypass|sanitiz|escape|overflow"`
- Read `SECURITY.md`, `CHANGELOG*`, `HISTORY*`, `RELEASES*`, `docs/security*`
- `git log -p -- SECURITY.md CHANGELOG*` to see when each fix landed vs. the pinned commit
### 3. Map fixed-in-version against the version under review
- For each advisory, note "fixed in X.Y.Z"; compare to the pinned version
- If the pinned version is AT OR AFTER the fix commit, that CVE is already patched here (not reportable as-is)
- If BEFORE, it is a known n-day — still not a NOVEL finding, flag it for the n-day/variant agents instead
### 4. Build the external baseline
- Cross-reference the ecosystem: GHSA (GitHub Advisories), NVD/NIST, `cargo audit` / `npm audit` / `pip-audit` / `osv-scanner` output if present
- `grep -rniE "cve-[0-9]{4}-[0-9]+|ghsa-" .` to catch CVE ids already noted in code/comments/tests
- Produce a table: `CVE/GHSA | class | fixed-in | present-in-this-version? | patch commit`
### 5. Report Format
For each baseline entry (this agent reports the BASELINE, not exploits):
```
FINDING:
- Title: Known-CVE Baseline & Dedup entry at [file:line]
- Severity: Info
- CWE: CWE-1059
- Endpoint: [file:line of the manifest/SECURITY.md/commit proving version+fix status]
- Vector: [advisory id -> class -> fixed-in vs pinned version]
- Payload: [the git log/grep command + exact quoted version string or commit hash]
- Evidence: [exact code/commit quoted + version/commit pinned + whether patched-here=yes/no + source: GHSA/NVD/CHANGELOG]
- Impact: Establishes the dedup baseline; marks each class as already-known so novel findings can be isolated
- Remediation: N/A (baseline); flag already-known-unpatched classes to the n-day/variant agents
```
- Persist the full baseline table to `$NEUROSPLOIT_POCS/{target}-cve-baseline.md` (static-derived) and cite it.
## System Prompt
You are a senior AppSec vulnerability researcher doing responsible, novelty-gated research. Your sole job in this pass is to build an authoritative KNOWN-ISSUE baseline for the exact pinned version/commit of {target} so that no later finding re-reports an existing CVE as new. Pin the version from manifests and `git rev-parse` before asserting anything. Treat SECURITY.md, CHANGELOG, GHSA, and NVD as ground truth for "already known". Report ONLY what you can prove from the provided files and git metadata — quote the exact version string, commit hash, or advisory line as the receipt (file:line). Never guess a fix status; if the snippet does not show the version or the patch commit, say so. Do not claim any live/HTTP/network result — this is source-only. Classify each known class as patched-here or not, so novel findings can be cleanly separated downstream.
Credits: Joas A Santos & Red Team Leaders.
@@ -0,0 +1,56 @@
# Business-Logic & Broken-Access-Control Researcher Agent
## User Prompt
You are reviewing the source code of **{target}** for NOVEL business-logic and broken-access-control flaws: missing or incorrect authorization checks, IDOR, tenant/owner confusion, and state-machine skips — with the exact unguarded code path.
**Recon Context:**
{recon_json}
The relevant source files are provided to you below the methodology.
**METHODOLOGY:**
### 1. Pin version & novelty baseline
- `git -C {target} rev-parse HEAD`; record `name@version`
- Check whether the authz gap is already a known advisory: `git log --oneline | grep -iE "authz|access|idor|permission|tenant|privilege|owner"`, `SECURITY.md`, `CHANGELOG*`, GHSA/NVD
### 2. Inventory protected resources & the intended policy
- Identify sensitive actions: read/update/delete of records, admin ops, money/credit moves, role changes, multi-step flows (checkout, invite, reset)
- Infer the INTENDED rule from code/tests/docs: who may do what to which object
### 3. Find the enforcement points (and the holes)
- Locate guards: `grep -rniE "authoriz|is_admin|has_role|current_user|require_(login|auth)|@login_required|permission|owner|tenant|can\(" .`
- For each sensitive handler, verify a guard exists AND is correct:
- Object-level: does it check the object belongs to `current_user`/tenant, or only that the user is logged in? (IDOR / BOLA)
- Function-level: is an admin-only route reachable by a normal role? (missing function-level authz)
- Is the id/owner taken from the REQUEST instead of the session? (horizontal escalation)
### 4. Hunt logic/state-machine skips
- Steps that can be reordered or skipped: pay-after-ship, verify-after-use, approve-after-execute
- Mass-assignment into privileged fields (`role`, `is_admin`, `balance`, `owner_id`) — `grep -rniE "update\(|assign|from_json|serde\(flatten\)|params\.permit" .`
- Replay/toctou: a check and a use separated so the state can change in between
### 5. Prove the unguarded path
- Quote the handler and show the MISSING or INCORRECT check (file:line); trace how a lower-privileged or non-owner actor reaches the action
- Show the request-controlled identifier that selects another user's/tenant's object
### 6. Report Format
For each CONFIRMED finding:
```
FINDING:
- Title: <IDOR / missing-authz / logic skip / mass-assignment> at [file:line]
- Severity: High
- CWE: CWE-285
- Endpoint: [file:line of the vulnerable handler]
- Vector: [actor/role -> action -> object; where the required check is absent/incorrect]
- Payload: [the exact request-shaped input (changed id/role/step) + the call chain reaching the action]
- Evidence: [the handler quoted showing the missing/incorrect guard + the correct guard elsewhere for contrast + version/commit pinned + novel: why + checked-against: <CVE/GHSA/commit>]
- Impact: [horizontal/vertical priv-esc, cross-tenant data access, unauthorized state change]
- Remediation: [enforce object-level ownership/tenant check server-side; derive id from session; gate by role; order-enforce the state machine]
```
- Write a static-derived PoC (a failing unit/integration test asserting a non-owner/low-role reaches the action) to `$NEUROSPLOIT_POCS/{target}-authz-<handler>.{ext}` and cite it. Mark it SOURCE-DERIVED.
## System Prompt
You are a senior AppSec vulnerability researcher specializing in business-logic and broken-access-control flaws, doing responsible, novelty-gated research. Report ONLY issues you can PROVE in the provided code: quote the vulnerable handler and show the authorization check that is MISSING or INCORRECT (file:line), and trace how an under-privileged or non-owner actor reaches the sensitive action. Anchor claims in the code's own intended policy — ideally contrasting a correct guard elsewhere with the missing one here. Pin the version/commit. The flaw must be NOVEL: state `novel: <why>` and `checked-against: <CVE/GHSA/commit/CHANGELOG>`; do not re-report an already-fixed access-control advisory for this version unless it is a concrete bypass. No speculation and no live/HTTP claims — source-only. If you cannot see the guard (it may live in middleware or a decorator not provided), say the finding is unconfirmed rather than assume it is absent.
Credits: Joas A Santos & Red Team Leaders.
@@ -0,0 +1,54 @@
# Patch-Diff & Variant (n-day -> 0-day) Researcher Agent
## User Prompt
You are reviewing the source code of **{target}** to study a recent security-fix commit and find what it MISSED — an incomplete fix (a reachable sink the patch left behind, or a bypass of the new check) and the same bug pattern repeated elsewhere in the tree. The goal is a NOVEL variant, not the already-patched CVE.
**Recon Context:**
{recon_json}
The relevant source files are provided to you below the methodology.
**METHODOLOGY:**
### 1. Find the security-fix commits
- `git -C {target} log --oneline -n 200 | grep -iE "fix|security|cve|sanitiz|escape|validat|bypass|injection|traversal|overflow|auth"`
- For each candidate: `git -C {target} show <sha>` and `git -C {target} log -p <sha>` to read the exact diff
- Note the CVE/GHSA it addresses and the precise lines/function it changed
### 2. Characterize the fix precisely
- What was the vulnerable pattern (source -> sink)? What control did the patch ADD (a check, an escape, an allowlist, a type guard)?
- Where is that control enforced — one callsite, or every callsite? Centralized or copy-pasted?
### 3. Hunt the incomplete-fix / bypass (patch variant)
- Did the patch guard ONE entrypoint but leave a sibling reaching the same sink unguarded? `grep -rn "<sink-symbol>" .` and compare each callsite against the added check
- Can the new check be bypassed? Look for: normalization mismatches (check before decode), case/encoding gaps, missing recursion, allowlist holes, early-return paths, `unsafe`/raw-SQL/`eval` reached around the guard
- Did the fix cover the reported input but not an equivalent one (alternate parser, second deserializer, another file-read)?
### 4. Hunt the SAME pattern elsewhere (variant across the tree)
- Build the sink signature from the patched code and grep the whole repo for structurally identical uses that were NEVER patched
- `semgrep` a pattern mirroring the vulnerable shape if available; the CODE CITATION is the proof, not the scanner
### 5. Prove reachability from untrusted input
- Trace a concrete source (route param, request body, CLI arg, file, env, IPC) to the unpatched sink; quote the full path (file:line each hop)
- Confirm the added control does NOT sit on this path
### 6. Report Format
For each CONFIRMED finding:
```
FINDING:
- Title: Patch-Bypass/Variant of <orig CVE/commit> at [file:line]
- Severity: High
- CWE: CWE-1288
- Endpoint: [file:line of the unguarded sink]
- Vector: [untrusted source -> hops -> sink, and how it evades the patch's new check]
- Payload: [benign crafted input that reaches the sink around the fix + the exact call chain]
- Evidence: [exact vulnerable lines quoted + the patch commit sha it bypasses/mirrors + version/commit pinned + novel: why this is NOT the fixed CVE + checked-against: <CVE/GHSA/commit>]
- Impact: [concrete technical impact of the variant]
- Remediation: [centralize the check / cover all callsites / fix the normalization-order or allowlist gap]
```
- Write the static-derived PoC (crafted input + failing unit test asserting the sink is reached) to `$NEUROSPLOIT_POCS/{target}-variant-<sha>.{ext}` and cite it.
## System Prompt
You are a senior AppSec vulnerability researcher specializing in patch-diff and variant analysis (turning an n-day into a novel 0-day). Report ONLY high-confidence findings you can prove in the PROVIDED code and git history: an incomplete fix, a concrete bypass of a patch's new check, or the same bug pattern in an unpatched location. Always start from a real security-fix commit (`git show`/`git log -p`) and pin the version/commit. A finding is valid ONLY if it is materially DIFFERENT from the already-fixed CVE — a plain restatement of the patched bug is forbidden; state `novel: <why it differs>` and `checked-against: <CVE/GHSA/commit>` every time. Prove a reachable, unsanitized path from untrusted input and quote every hop (file:line). No speculation, no live/HTTP claims — source-only. If the provided snippet does not show the sink, the source, or the patch, say so rather than guess.
Credits: Joas A Santos & Red Team Leaders.
@@ -0,0 +1,52 @@
# Source-to-Sink Taint Researcher Agent
## User Prompt
You are reviewing the source code of **{target}** to prove a NOVEL, reachable, unsanitized taint path from an untrusted entrypoint to a dangerous sink (injection / RCE / SSRF / deserialization / path traversal / prototype pollution).
**Recon Context:**
{recon_json}
The relevant source files are provided to you below the methodology.
**METHODOLOGY:**
### 1. Pin version & confirm novelty first
- `git -C {target} rev-parse HEAD`; record `name@version`
- Before tracing, confirm the class is not already a patched CVE at this version: scan `SECURITY.md`, `CHANGELOG*`, `git log --oneline | grep -iE "cve|sanitiz|injection|ssrf|traversal|deser"`, GHSA/NVD for the ecosystem
### 2. Identify the SOURCE (untrusted input)
- Request params/body/headers/cookies, path params, uploaded files, env, CLI args, IPC/queue messages, config a lower-trust actor controls
- Quote the exact line where the value enters (file:line)
### 3. Identify the SINK
- Command exec, raw SQL/NoSQL, deserializer, dynamic file path, outbound URL, template/eval, reflected HTML/JS, object-key assignment (prototype pollution)
- Locate with `grep -rnE "<sink-shape>" .`; quote the exact line (file:line)
### 4. Trace the DATAFLOW end to end
- Follow the value hop by hop: assignments, function params, struct fields, closures, await boundaries
- Quote EVERY hop (file:line). Identify any validation/escaping/allowlist encountered and PROVE it is absent, insufficient, or bypassable on this path (wrong order, partial, wrong charset, missing recursion)
### 5. Confirm exploitability (static)
- State the concrete attacker-controlled value that reaches the sink unmodified (or modified in an attacker-useful way)
- Explain why existing controls do not stop it; explain what the sink does with it (RCE, data read/write, SSRF, file disclosure, pollution)
### 6. Report Format
For each CONFIRMED finding:
```
FINDING:
- Title: <class> source-to-sink at [file:line]
- Severity: Critical
- CWE: CWE-20
- Endpoint: [file:line of the source]
- Vector: [source (file:line) -> hop (file:line) -> ... -> sink (file:line)]
- Payload: [benign crafted input demonstrating the reachable path + the exact call chain]
- Evidence: [every hop quoted + sink quoted + version/commit pinned + novel: why this path is not an already-fixed CVE + checked-against: <CVE/GHSA/commit/CHANGELOG>]
- Impact: [concrete: RCE / SQLi data exfil / SSRF to metadata / arbitrary file read-write / prototype pollution -> ...]
- Remediation: [validate/parameterize/escape at the sink; safe loader; allowlist; normalize-before-check]
```
- Write a static-derived PoC (crafted input + a failing unit test that drives the source and asserts the sink is reached) to `$NEUROSPLOIT_POCS/{target}-taint-<class>.{ext}` and cite it. Mark it SOURCE-DERIVED.
## System Prompt
You are a senior AppSec vulnerability researcher performing deep source-to-sink taint analysis for responsible, novelty-gated research. Report ONLY a finding you can PROVE in the provided code: an untrusted source, a dangerous sink, and a reachable, unsanitized dataflow connecting them — with EVERY hop quoted (file:line). Pin the version/commit. The finding must be NOVEL: state `novel: <why>` and `checked-against: <CVE/GHSA/commit/CHANGELOG>`; a path that merely re-walks an already-patched CVE for this version is not reportable unless recast as a concrete bypass/variant. Prove that controls on the path are absent or defeatable — do not assume sanitization you cannot see, and do not assume a control works that you cannot trace. No speculation and no live/HTTP/network claims — reason strictly about the source and git metadata. If any hop, the source, or the sink is missing from the snippet, say the path is unconfirmed rather than guess.
Credits: Joas A Santos & Red Team Leaders.
+10 -2
View File
@@ -318,6 +318,9 @@ enum Cmd {
/// Economy preset for a short, low-cost review (see `run --quick`). /// Economy preset for a short, low-cost review (see `run --quick`).
#[arg(long)] #[arg(long)]
quick: bool, quick: bool,
/// Vulnerability-research mode: hunt a NOVEL, CVE-reportable bug (known-CVE dedup + patch-diff variant analysis).
#[arg(long)]
research: bool,
#[arg(long)] #[arg(long)]
offline: bool, offline: bool,
#[arg(long)] #[arg(long)]
@@ -359,6 +362,9 @@ enum Cmd {
/// Economy preset for a short, low-cost test (see `run --quick`). /// Economy preset for a short, low-cost test (see `run --quick`).
#[arg(long)] #[arg(long)]
quick: bool, quick: bool,
/// Vulnerability-research mode (see `whitebox --research`).
#[arg(long)]
research: bool,
#[arg(long)] #[arg(long)]
offline: bool, offline: bool,
#[arg(long)] #[arg(long)]
@@ -880,7 +886,7 @@ async fn main() -> anyhow::Result<()> {
let ig = harness::integrations::Integrations::load(&repl::proj_dir()); let ig = harness::integrations::Integrations::load(&repl::proj_dir());
post_integrations(&ig, &url, &out, jira, false, None).await; post_integrations(&ig, &url, &out, jira, false, None).await;
} }
Cmd::Whitebox { path, models, max_agents, vote_n, chain_depth, recon, quick, offline, subscription, jira, only, verbose } => { Cmd::Whitebox { path, models, max_agents, vote_n, chain_depth, recon, quick, research, offline, subscription, jira, only, verbose } => {
let path = resolve_source(&base, &path)?; // local path OR github URL/owner/repo let path = resolve_source(&base, &path)?; // local path OR github URL/owner/repo
let mut cfg = RunConfig::new(&path); let mut cfg = RunConfig::new(&path);
cfg.max_agents = max_agents; cfg.max_agents = max_agents;
@@ -891,6 +897,7 @@ async fn main() -> anyhow::Result<()> {
cfg.subscription = subscription; cfg.subscription = subscription;
cfg.verbose = verbose; cfg.verbose = verbose;
cfg.pinned = parse_only(&only); cfg.pinned = parse_only(&only);
cfg.research = research;
if quick { apply_quick(&mut cfg); } if quick { apply_quick(&mut cfg); }
if !models.is_empty() { if !models.is_empty() {
cfg.models = models; cfg.models = models;
@@ -900,7 +907,7 @@ async fn main() -> anyhow::Result<()> {
let ig = harness::integrations::Integrations::load(&repl::proj_dir()); let ig = harness::integrations::Integrations::load(&repl::proj_dir());
post_integrations(&ig, &path, &out, jira, false, None).await; post_integrations(&ig, &path, &out, jira, false, None).await;
} }
Cmd::Greybox { repo, url, models, creds, focus, max_agents, vote_n, chain_depth, recon, quick, offline, subscription, mcp, only, verbose } => { Cmd::Greybox { repo, url, models, creds, focus, max_agents, vote_n, chain_depth, recon, quick, research, offline, subscription, mcp, only, verbose } => {
let repo = resolve_source(&base, &repo)?; // local path OR github URL/owner/repo let repo = resolve_source(&base, &repo)?; // local path OR github URL/owner/repo
let url = if url.starts_with("http") { url } else { format!("https://{url}") }; let url = if url.starts_with("http") { url } else { format!("https://{url}") };
let mut cfg = RunConfig::new(&url); let mut cfg = RunConfig::new(&url);
@@ -914,6 +921,7 @@ async fn main() -> anyhow::Result<()> {
cfg.verbose = verbose; cfg.verbose = verbose;
cfg.instructions = focus; cfg.instructions = focus;
cfg.pinned = parse_only(&only); cfg.pinned = parse_only(&only);
cfg.research = research;
if quick { apply_quick(&mut cfg); } if quick { apply_quick(&mut cfg); }
if !models.is_empty() { if !models.is_empty() {
cfg.models = models; cfg.models = models;
+14 -2
View File
@@ -152,7 +152,7 @@ pub(crate) const ACCEPTED: &[&str] = &[
"/history", "/idle", "/inscope", "/instructions", "/integration", "/integrations", "/key", "/log", "/history", "/idle", "/inscope", "/instructions", "/integration", "/integrations", "/key", "/log",
"/logs", "/mcp", "/memory", "/model", "/models", "/objective", "/objectives", "/observe", "/logs", "/mcp", "/memory", "/model", "/models", "/objective", "/objectives", "/observe",
"/observe-only", "/offline", "/observe-only", "/offline",
"/onboard", "/only", "/oos", "/outofscope", "/policy", "/providers", "/proxy", "/quick", "/economy", "/eco", "/q", "/quit", "/recon", "/onboard", "/only", "/oos", "/outofscope", "/policy", "/providers", "/proxy", "/research", "/quick", "/economy", "/eco", "/q", "/quit", "/recon",
"/pause", "/repo", "/report", "/results", "/resume", "/retest", "/revalidate", "/run", "/runs", "/pause", "/repo", "/report", "/results", "/resume", "/retest", "/revalidate", "/run", "/runs",
"/scope", "/scope-out", "/show", "/status", "/stop", "/sub", "/subscription", "/target", "/scope", "/scope-out", "/show", "/status", "/stop", "/sub", "/subscription", "/target",
"/temp-email", "/tempmail", "/theme", "/timeout", "/ua", "/url", "/useragent", "/validate", "/temp-email", "/tempmail", "/theme", "/timeout", "/ua", "/url", "/useragent", "/validate",
@@ -163,7 +163,7 @@ pub(crate) const ACCEPTED: &[&str] = &[
const COMMANDS: &[&str] = &[ const COMMANDS: &[&str] = &[
"/help", "/onboard", "/show", "/config", "/providers", "/model", "/key", "/sub", "/target", "/help", "/onboard", "/show", "/config", "/providers", "/model", "/key", "/sub", "/target",
"/repo", "/auth", "/creds", "/focus", "/objective", "/scope-out", "/attach", "/context", "/mcp", "/offline", "/repo", "/auth", "/creds", "/focus", "/objective", "/scope-out", "/attach", "/context", "/mcp", "/offline",
"/quick", "/economy", "/eco", "/votes", "/chain", "/recon", "/tempmail", "/timeout", "/proxy", "/burp", "/ua", "/agents", "/only", "/theme", "/clear", "/run", "/stop", "/pause", "/continue", "/runs", "/results", "/report", "/research", "/quick", "/economy", "/eco", "/votes", "/chain", "/recon", "/tempmail", "/timeout", "/proxy", "/burp", "/ua", "/agents", "/only", "/theme", "/clear", "/run", "/stop", "/pause", "/continue", "/runs", "/results", "/report",
"/status", "/logs", "/diff", "/retest", "/validate", "/finding", "/expand", "/integrations", "/status", "/logs", "/diff", "/retest", "/validate", "/finding", "/expand", "/integrations",
"/memory", "/forget", "/graph", "/inscope", "/observe", "/guardrail", "/policy", "/memory", "/forget", "/graph", "/inscope", "/observe", "/guardrail", "/policy",
"/capability", "/audit", "/quit", "/capability", "/audit", "/quit",
@@ -268,6 +268,7 @@ struct Session {
max_agents: usize, max_agents: usize,
chain_depth: usize, chain_depth: usize,
recon_intensity: usize, recon_intensity: usize,
research: bool,
/// Opt-in disposable email (mail.tm) for register flows needing a confirmation code. /// Opt-in disposable email (mail.tm) for register flows needing a confirmation code.
temp_email: bool, temp_email: bool,
/// Idle guardrail: stop a run if no NEW finding lands in this many seconds /// Idle guardrail: stop a run if no NEW finding lands in this many seconds
@@ -320,6 +321,7 @@ impl Default for Session {
max_agents: 0, max_agents: 0,
chain_depth: 2, chain_depth: 2,
recon_intensity: 3, recon_intensity: 3,
research: false,
temp_email: false, temp_email: false,
idle_secs: 300, // 5-minute idle guardrail by default idle_secs: 300, // 5-minute idle guardrail by default
proxy: None, proxy: None,
@@ -846,6 +848,13 @@ pub async fn repl(base: &Path, auth: SessionAuth) -> anyhow::Result<()> {
if arg.is_empty() { println!(" recon intensity: {} ({}) — set with /recon <1-4> [1 quick · 2 standard · 3 deep · 4 exhaustive]", s.recon_intensity, lvl(s.recon_intensity)); } if arg.is_empty() { println!(" recon intensity: {} ({}) — set with /recon <1-4> [1 quick · 2 standard · 3 deep · 4 exhaustive]", s.recon_intensity, lvl(s.recon_intensity)); }
else { s.recon_intensity = arg.parse::<usize>().unwrap_or(s.recon_intensity).clamp(1, 4); println!(" recon intensity: {} ({}) — more rounds, more enumeration, auto-installs tools", s.recon_intensity, lvl(s.recon_intensity)); } else { s.recon_intensity = arg.parse::<usize>().unwrap_or(s.recon_intensity).clamp(1, 4); println!(" recon intensity: {} ({}) — more rounds, more enumeration, auto-installs tools", s.recon_intensity, lvl(s.recon_intensity)); }
} }
"/research" => {
match arg.trim() {
"on" | "true" | "1" => { s.research = true; println!(" \x1b[1;36m🔬 research mode ON\x1b[0m — whitebox/greybox will hunt a NOVEL, CVE-reportable bug (known-CVE dedup + patch-diff variant analysis)"); }
"off" | "false" | "0" => { s.research = false; println!(" research mode off"); }
_ => println!(" research mode: {} — /research on|off (for whitebox/greybox: find a new CVE, not a known one)", if s.research { "\x1b[36mon\x1b[0m" } else { "\x1b[2moff\x1b[0m" }),
}
}
"/quick" | "/economy" | "/eco" => { "/quick" | "/economy" | "/eco" => {
// Economy preset for a short, low-cost test — the single switch // Economy preset for a short, low-cost test — the single switch
// for "fast and cheap" instead of tuning each knob. The big // for "fast and cheap" instead of tuning each knob. The big
@@ -1471,6 +1480,7 @@ async fn run(base: &Path, s: &Session, history: &mut Vec<RunRecord>) {
cfg.vote_n = s.vote_n; cfg.vote_n = s.vote_n;
cfg.chain_depth = s.chain_depth; cfg.chain_depth = s.chain_depth;
cfg.recon_intensity = s.recon_intensity; cfg.recon_intensity = s.recon_intensity;
cfg.research = s.research;
cfg.temp_email = s.temp_email; cfg.temp_email = s.temp_email;
cfg.proxy = s.proxy.clone(); cfg.proxy = s.proxy.clone();
cfg.user_agent = s.user_agent.clone(); cfg.user_agent = s.user_agent.clone();
@@ -1560,6 +1570,7 @@ async fn start_background(base: &Path, s: &Session, reader: &mut Reader,
cfg.vote_n = s.vote_n; cfg.vote_n = s.vote_n;
cfg.chain_depth = s.chain_depth; cfg.chain_depth = s.chain_depth;
cfg.recon_intensity = s.recon_intensity; cfg.recon_intensity = s.recon_intensity;
cfg.research = s.research;
cfg.temp_email = s.temp_email; cfg.temp_email = s.temp_email;
cfg.proxy = s.proxy.clone(); cfg.proxy = s.proxy.clone();
cfg.user_agent = s.user_agent.clone(); cfg.user_agent = s.user_agent.clone();
@@ -2236,6 +2247,7 @@ fn help() {
h("/votes <n>", "number of validator votes per finding"); h("/votes <n>", "number of validator votes per finding");
h("/chain <n>", "attack-chain depth (post-exploitation pivots; 0 = off)"); h("/chain <n>", "attack-chain depth (post-exploitation pivots; 0 = off)");
h("/recon <1-4>", "recon intensity: 1 quick · 2 standard · 3 deep · 4 exhaustive (installs tools)"); h("/recon <1-4>", "recon intensity: 1 quick · 2 standard · 3 deep · 4 exhaustive (installs tools)");
h("/research", "whitebox/greybox: hunt a NOVEL, CVE-reportable bug (known-CVE dedup + patch-diff variant analysis)");
h("/quick", "economy preset: short, low-cost run (1 voter · 1 chain round · light recon · ≤6 agents)"); h("/quick", "economy preset: short, low-cost run (1 voter · 1 chain round · light recon · ≤6 agents)");
h("/tempmail on|off", "opt-in disposable inbox (mail.tm) to read a register confirmation code"); h("/tempmail on|off", "opt-in disposable inbox (mail.tm) to read a register confirmation code");
h("/timeout <min>", "idle guardrail: stop if no new finding in <min> (0 = off)"); h("/timeout <min>", "idle guardrail: stop if no new finding in <min> (0 = off)");
+46 -2
View File
@@ -623,6 +623,34 @@ const WHITEBOX_DOCTRINE: &str = "MODE: WHITE-BOX STATIC SOURCE REVIEW. You are r
- Repro PoC (optional but valued): when a finding warrants it, WRITE a proof/repro script to $NEUROSPLOIT_POCS — e.g. the exact malicious input + the request/CLI call that would trigger the sink, or a unit-style harness exercising the vulnerable function — with a header comment (file:line it proves, how to run). Cite the PoC path in the evidence. Mark clearly that it demonstrates the code path (static-derived), not a live hit.\n\ - Repro PoC (optional but valued): when a finding warrants it, WRITE a proof/repro script to $NEUROSPLOIT_POCS — e.g. the exact malicious input + the request/CLI call that would trigger the sink, or a unit-style harness exercising the vulnerable function — with a header comment (file:line it proves, how to run). Cite the PoC path in the evidence. Mark clearly that it demonstrates the code path (static-derived), not a live hit.\n\
- Calibrate: High/Critical only when the sink is reachable and exploitable from untrusted input; guarded/unreachable code is Low or a lead.\n\n"; - Calibrate: High/Critical only when the sink is reachable and exploitable from untrusted input; guarded/unreachable code is Low or a lead.\n\n";
/// Vulnerability-RESEARCH doctrine: turn a source review into a hunt for a
/// NOVEL, CVE-reportable bug. Prepended (after the white-box doctrine) when the
/// operator runs research mode (`--research`, `/research`, or a focus/objective
/// that asks for a new CVE / 0-day / patch bypass). The whole point is NOVELTY:
/// do not re-report a known CVE — use the known ones as a map to find what the
/// fix missed, a variant, or a reintroduction.
const WHITEBOX_RESEARCH_DOCTRINE: &str = "MISSION: VULNERABILITY RESEARCH — find a NOVEL, CVE-reportable issue in this codebase, not a known one.\n\
- PIN THE VERSION FIRST: read the version (package.json/VERSION/__init__/composer.json/go.mod/tag) and the exact commit. Every claim is against THIS version/commit; note it in evidence.\n\
- RESEARCH KNOWN CVEs (de-duplicate): before reporting anything, build the set of ALREADY-KNOWN issues for this project+version — read SECURITY.md, CHANGELOG/release notes, the security advisories (GHSA), CVE/NVD, the issue tracker and recent security commits (`git log --oneline`, grep messages for CVE/security/fix/vuln/XSS/RCE/injection). A finding that matches a known CVE for this version is NOT novel — drop it or recast it ONLY as a patch-bypass/variant (below). State which known CVEs you checked against.\n\
- PATCH-DIFF / N-DAY -> 0-DAY (the highest-yield path): take a recent SECURITY fix (its commit) and study the diff. Ask: did the patch fix the ROOT CAUSE or just one path? Look for (1) incomplete fixes — another reachable sink the patch did not cover, a bypass of the new check (different encoding, type juggling, alternate parser, case/Unicode, second-order input); (2) the SAME bug pattern elsewhere in the tree (variant analysis — grep the fixed sink's shape across the repo); (3) reintroduction in a later commit. A proven bypass of an existing patch IS novel and reportable.\n\
- SOURCE->SINK with reachability: trace untrusted input (request params, headers, body, deserialized objects, file names, env, IPC, config) to a dangerous sink (SQL/exec/eval/template/path/deserialize/SSRF/XXE/prototype/unsafe-reflection). Only call it a vuln when the path is REACHABLE from an untrusted entrypoint without an effective sanitizer; record the full path `entry -> … -> sink`.\n\
- NOVELTY GATE (strict): report a finding ONLY if (a) it does not match a known CVE for this version, OR (b) it is a concrete bypass/variant of a patched issue. For each, state explicitly: 'novel: <why>' and 'checked-against: <CVEs/advisories/commits>'. No speculation — high-confidence, evidence-backed only.\n\
- WRITE A PoC: produce a minimal, SAFE proof (a failing unit test, a crafted input + the exact call reaching the sink, or a request) to $NEUROSPLOIT_POCS and cite it. Where a running instance is available (greybox), CONFIRM the source-derived bug dynamically — but never run destructive payloads.\n\
- REPORT for disclosure: each finding carries file:line, root-cause analysis, the version/commit, the novelty justification, CVSS vector, PoC path, suggested fix, and the project's disclosure channel (SECURITY.md / security@). Prefer a few solid, novel, reportable bugs over a long list of known or speculative ones.\n\n";
/// True when the engagement is asking for vulnerability research / a new CVE,
/// either via the explicit flag or from natural-language focus/objective.
fn is_research_intent(cfg: &RunConfig) -> bool {
if cfg.research { return true; }
let hay = format!("{} {}",
cfg.instructions.clone().unwrap_or_default(),
cfg.objective.clone().unwrap_or_default()).to_lowercase();
["new cve", "novel", "0-day", "0day", "zero-day", "zero day", "patch bypass",
"patch-bypass", "variant analysis", "n-day", "nday", "reportable", "cve research",
"vulnerability research", "find a cve", "nova cve", "pesquisa de vuln"]
.iter().any(|k| hay.contains(k))
}
/// Methodology directions for a modern JS SPA backed by a REST/GraphQL API /// Methodology directions for a modern JS SPA backed by a REST/GraphQL API
/// (Angular/React/Vue front + Node/Express-style API — the shape of OWASP Juice /// (Angular/React/Vue front + Node/Express-style API — the shape of OWASP Juice
/// Shop and many real apps). These are DIRECTIONS on HOW to hunt each vuln class, /// Shop and many real apps). These are DIRECTIONS on HOW to hunt each vuln class,
@@ -1167,6 +1195,10 @@ pub async fn run_whitebox(cfg: RunConfig, lib: &Library, pool: &ModelPool, tx: S
let context = collect_repo_context(Path::new(&cfg.target), 200, 120_000); let context = collect_repo_context(Path::new(&cfg.target), 200, 120_000);
let bytes = context.len(); let bytes = context.len();
let research = is_research_intent(&cfg);
if research {
let _ = tx.send("notify: 🔬 research mode — hunting a NOVEL, CVE-reportable issue (known-CVE dedup + patch-diff variant analysis)".into()).await;
}
let _ = tx.send(format!("collected {} bytes of source context", bytes)).await; let _ = tx.send(format!("collected {} bytes of source context", bytes)).await;
if bytes == 0 { if bytes == 0 {
let _ = tx.send("no readable source found at the given path".into()).await; let _ = tx.send("no readable source found at the given path".into()).await;
@@ -1214,7 +1246,12 @@ pub async fn run_whitebox(cfg: RunConfig, lib: &Library, pool: &ModelPool, tx: S
); );
// Prepend the white-box doctrine so code agents stay in static // Prepend the white-box doctrine so code agents stay in static
// source-review mode and never hallucinate live/black-box actions. // source-review mode and never hallucinate live/black-box actions.
let sys = format!("{}{}", WHITEBOX_DOCTRINE, ag.system); // In research mode, add the novelty/patch-diff hunting doctrine.
let sys = if research {
format!("{}{}{}", WHITEBOX_DOCTRINE, WHITEBOX_RESEARCH_DOCTRINE, ag.system)
} else {
format!("{}{}", WHITEBOX_DOCTRINE, ag.system)
};
match pool.complete_routed(Task::Exploit, &ag.name, &sys, &user).await { match pool.complete_routed(Task::Exploit, &ag.name, &sys, &user).await {
Ok((m, text)) => { Ok((m, text)) => {
let f = extract_findings(&text, &ag.name); let f = extract_findings(&text, &ag.name);
@@ -1267,6 +1304,10 @@ pub async fn run_greybox(cfg: RunConfig, lib: &Library, pool: &ModelPool, tx: Se
// ---- 2. Review the source for leads ------------------------------- // ---- 2. Review the source for leads -------------------------------
let context = collect_repo_context(Path::new(&repo), 200, 90_000); let context = collect_repo_context(Path::new(&repo), 200, 90_000);
let _ = tx.send(format!("collected {} bytes of source for code review", context.len())).await; let _ = tx.send(format!("collected {} bytes of source for code review", context.len())).await;
let gb_research = is_research_intent(&cfg);
if gb_research {
let _ = tx.send("notify: 🔬 research mode — code review hunts a novel, CVE-reportable bug, then confirms it live".into()).await;
}
let mut rl = cfg.rl_path.as_ref().map(|p| RlState::load(Path::new(p))).unwrap_or_default(); let mut rl = cfg.rl_path.as_ref().map(|p| RlState::load(Path::new(p))).unwrap_or_default();
let mut code_leads = String::new(); let mut code_leads = String::new();
@@ -1284,7 +1325,10 @@ pub async fn run_greybox(cfg: RunConfig, lib: &Library, pool: &ModelPool, tx: Se
where endpoint is file:line.", where endpoint is file:line.",
ag.user.replace("{target}", "the repository").replace("{recon_json}", "{}"), ctx ag.user.replace("{target}", "the repository").replace("{recon_json}", "{}"), ctx
); );
match pool.complete_routed(Task::Select, &ag.name, &ag.system, &user).await { // Research mode steers the code-review half of greybox too:
// find a novel, reportable bug, then confirm it live.
let sys = if gb_research { format!("{}{}", WHITEBOX_RESEARCH_DOCTRINE, ag.system) } else { ag.system.clone() };
match pool.complete_routed(Task::Select, &ag.name, &sys, &user).await {
Ok((_, text)) => { let f = extract_findings(&text, &ag.name); Ok((_, text)) => { let f = extract_findings(&text, &ag.name);
let _ = txc.send(format!("review {} → {} lead(s)", ag.name, f.len())).await; f } let _ = txc.send(format!("review {} → {} lead(s)", ag.name, f.len())).await; f }
Err(_) => vec![], Err(_) => vec![],
@@ -213,6 +213,12 @@ pub struct RunConfig {
/// more recon rounds, more active enumeration, and auto-installing tools. /// more recon rounds, more active enumeration, and auto-installing tools.
#[serde(default = "default_recon")] #[serde(default = "default_recon")]
pub recon_intensity: usize, pub recon_intensity: usize,
/// Vulnerability-research mode: hunt for NOVEL, CVE-reportable issues in a
/// source repo — research known CVEs/advisories to de-duplicate, do
/// patch-diff variant analysis (incomplete-fix bypasses, sibling sinks),
/// and gate strictly on novelty. Steers whitebox/greybox.
#[serde(default)]
pub research: bool,
/// Opt-in: when the app requires email confirmation to register, allow the /// Opt-in: when the app requires email confirmation to register, allow the
/// agent to use a free disposable-inbox API (mail.tm) to read the code/link. /// agent to use a free disposable-inbox API (mail.tm) to read the code/link.
/// Off by default. Account creation is still capped by the safety guardrail. /// Off by default. Account creation is still capped by the safety guardrail.
@@ -315,6 +321,7 @@ impl RunConfig {
repo: None, repo: None,
pinned: Vec::new(), pinned: Vec::new(),
chain_depth: 2, chain_depth: 2,
research: false,
proxy: None, proxy: None,
user_agent: None, user_agent: None,
recon_intensity: 3, recon_intensity: 3,
+1
View File
@@ -533,6 +533,7 @@ async function startExploitation() {
chainDepth: Number($('#fieldChain').value), chainDepth: Number($('#fieldChain').value),
recon: Number($('#fieldRecon').value), recon: Number($('#fieldRecon').value),
quick: $('#fieldQuick') ? $('#fieldQuick').checked : false, quick: $('#fieldQuick') ? $('#fieldQuick').checked : false,
research: $('#fieldResearch') ? $('#fieldResearch').checked : false,
subscription: state.authMode === 'subscription', subscription: state.authMode === 'subscription',
mcp: $('#fieldMcp').checked, mcp: $('#fieldMcp').checked,
agents: [...state.selected], agents: [...state.selected],
+3
View File
@@ -221,6 +221,9 @@
<div class="section-desc">Optional. Left on <em>unlimited</em>, the run behaves exactly as it always has — full depth, no cap.</div> <div class="section-desc">Optional. Left on <em>unlimited</em>, the run behaves exactly as it always has — full depth, no cap.</div>
</div> </div>
<div class="field-row"> <div class="field-row">
<div class="check-row"><input type="checkbox" id="fieldResearch" /> <label for="fieldResearch"><b>🔬 Research mode</b> — hunt a novel, CVE-reportable bug (whitebox/greybox)</label>
<div class="field-help">Known-CVE de-dup + patch-diff variant analysis. Best with a source repo; reports only genuinely new issues (or a concrete patch bypass).</div>
</div>
<div class="check-row check-quick"><input type="checkbox" id="fieldQuick" /> <label for="fieldQuick"><b>⚡ Quick mode</b> — short, low-cost test</label> <div class="check-row check-quick"><input type="checkbox" id="fieldQuick" /> <label for="fieldQuick"><b>⚡ Quick mode</b> — short, low-cost test</label>
<div class="field-help">Economy preset: 1 voter, 1 chain round, light recon, ≤6 agents, eco budget. The big token saver. Wins over the settings below.</div> <div class="field-help">Economy preset: 1 voter, 1 chain round, light recon, ≤6 agents, eco budget. The big token saver. Wins over the settings below.</div>
</div> </div>
+2
View File
@@ -882,6 +882,7 @@ function sanitizeLaunch(body) {
sandbox: !!body.sandbox, sandbox: !!body.sandbox,
typesafe: body.typesafe, typesafe: body.typesafe,
quick: !!body.quick, quick: !!body.quick,
research: !!body.research,
}; };
} }
@@ -902,6 +903,7 @@ function buildReplScript(body) {
if (body.objective) lines.push(`/objective ${body.objective}`); if (body.objective) lines.push(`/objective ${body.objective}`);
if (body.outOfScope) lines.push(`/scope-out ${body.outOfScope}`); if (body.outOfScope) lines.push(`/scope-out ${body.outOfScope}`);
if (body.creds) lines.push(`/creds ${body.creds}`); if (body.creds) lines.push(`/creds ${body.creds}`);
if (body.research) lines.push(`/research on`);
lines.push((body.agents || []).length ? `/only ${body.agents.join(',')}` : '/only clear'); lines.push((body.agents || []).length ? `/only ${body.agents.join(',')}` : '/only clear');
// Economy preset last, so it wins over the per-knob settings above. // Economy preset last, so it wins over the per-knob settings above.
if (body.quick) lines.push('/quick'); if (body.quick) lines.push('/quick');