feat(container,coverage): Strix-1.6.2-inspired capabilities

- Container image scanning: new `container` mode + 4 skills (vuln, secret,
  misconfig, SBOM) driving trivy/grype/syft headless, read-only. Scans an OCI
  ref / tar / Dockerfile for vulnerable packages (CVE/fixed-in/KEV), exposed
  secrets in any layer, Dockerfile+runtime misconfig, and writes an SBOM in
  both SPDX and CycloneDX to the run's sbom/. Also exposed as an MCP tool
  (neurosploit_container).
- Coverage report: every run writes coverage.md — which agents ran (tested
  surface), findings per agent, and the high-value classes NOT covered — so the
  reader sees the engagement's reach. Added to the assurance bundle.
- Login-verification evidence: doctrine now requires capturing the login
  request/response + a Playwright screenshot and recording success/failure
  before authenticated testing.
- HTTP traffic export: `neurosploit traffic <run>` turns the intercepted
  flows.jsonl into a traffic.http archive for external tools.

Not ported: Asset Discovery (enterprise-only, skipped per request).

383 tests.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
CyberSecurityUP
2026-09-23 01:27:26 -03:00
co-authored by Claude Opus 5
parent 1adc882f6d
commit e49595b8bf
12 changed files with 318 additions and 7 deletions
+31 -1
View File
@@ -13,7 +13,7 @@
<img src="https://img.shields.io/badge/License-MIT-green?style=flat-square"> <img src="https://img.shields.io/badge/License-MIT-green?style=flat-square">
<img src="https://img.shields.io/badge/MD%20Agents-458-red?style=flat-square"> <img src="https://img.shields.io/badge/MD%20Agents-458-red?style=flat-square">
<img src="https://img.shields.io/badge/Models-18%20providers-success?style=flat-square"> <img src="https://img.shields.io/badge/Models-18%20providers-success?style=flat-square">
<img src="https://img.shields.io/badge/Modes-Black%20%7C%20White%20%7C%20Grey%20%7C%20Host%20%7C%20AI%20%7C%20Mobile-9cf?style=flat-square"> <img src="https://img.shields.io/badge/Modes-Black%20%7C%20White%20%7C%20Grey%20%7C%20Host%20%7C%20AI%20%7C%20Mobile%20%7C%20Container-9cf?style=flat-square">
<img src="https://img.shields.io/badge/Auth-API%20key%20%7C%20Subscription-orange?style=flat-square"> <img src="https://img.shields.io/badge/Auth-API%20key%20%7C%20Subscription-orange?style=flat-square">
</p> </p>
@@ -46,6 +46,7 @@ Control TUI**.
| **AI / LLM red-team** | `neurosploit aitest <ai-url>` | jailbreaks & prompt injection + OWASP LLM Top 10 / MCP against a live AI agent | | **AI / LLM red-team** | `neurosploit aitest <ai-url>` | jailbreaks & prompt injection + OWASP LLM Top 10 / MCP against a live AI agent |
| **AI Skills / n8n** | `neurosploit skills <file\|folder>` | white-box audit of Skill/plugin & n8n workflow definitions | | **AI Skills / n8n** | `neurosploit skills <file\|folder>` | white-box audit of Skill/plugin & n8n workflow definitions |
| **Mobile / Binary** | `neurosploit mobile <app.apk\|app.ipa\|binary>` | reverse-engineer a local artifact: RASP, root/JB, pinning, anti-debug, obfuscation, secrets (Ghidra headless / MobSF / Frida) | | **Mobile / Binary** | `neurosploit mobile <app.apk\|app.ipa\|binary>` | reverse-engineer a local artifact: RASP, root/JB, pinning, anti-debug, obfuscation, secrets (Ghidra headless / MobSF / Frida) |
| **Container** | `neurosploit container <image:tag>` | scan an OCI image for vulnerable packages, exposed secrets, misconfig + emit an SBOM (SPDX/CycloneDX) via trivy/grype/syft |
| **Mission Control** | `neurosploit tui <url>` | live TUI panels + composer during the run | | **Mission Control** | `neurosploit tui <url>` | live TUI panels + composer during the run |
| **Interactive** | `neurosploit` | persistent REPL session (resumes per project) | | **Interactive** | `neurosploit` | persistent REPL session (resumes per project) |
@@ -713,6 +714,35 @@ non-destructively.
--- ---
## 📦 Container image scanning
```bash
neurosploit container myorg/app:1.4 --subscription --model anthropic:claude-opus-4-8 -v
neurosploit container ./image.tar
```
Scans an OCI image (registry ref, local tar, or Dockerfile) with trivy / grype /
syft headless: **vulnerable OS + language packages** (CVE, fixed-in, KEV),
**exposed secrets** in any layer, **Dockerfile/runtime misconfig** (root user,
unpinned base, curl-pipe-sh, secrets in ENV), and an **SBOM in both SPDX and
CycloneDX** written to the run's `sbom/` folder. Read-only — never pushes,
deletes or modifies a registry.
## 🧾 Coverage & traffic
Every run writes `coverage.md` — which agents ran (the tested surface), how many
findings each produced, and which high-value classes were **not** covered — so a
reader sees the engagement's reach, not just its findings. Login flows capture
verification evidence (the request/response + a Playwright screenshot) before
authenticated testing. With `--intercept own`, archived HTTP traffic exports to
a `.http` file:
```bash
neurosploit traffic <run> # flows.jsonl -> traffic.http
```
---
## 🔌 Run it as an MCP server ## 🔌 Run it as an MCP server
Drive NeuroSploit from Claude Code, Codex or Cursor as tools: Drive NeuroSploit from Claude Code, Codex or Cursor as tools:
@@ -0,0 +1,18 @@
# Container Image & Dockerfile Misconfiguration
## User Prompt
You are scanning the container image **{target}** (an OCI image reference, a local tar, or a Dockerfile) for: Container Image & Dockerfile Misconfiguration. CWE-16
**Context:**
{recon_json}
All tools run HEADLESS and are provisioned on demand (time-box each install, skip on failure). Only scan images you are authorized to scan.
### Method
1. Scan config with `trivy image --scanners misconfig` and `trivy config <Dockerfile|dir>`; optionally `hadolint <Dockerfile>`.
2. Flag: running as root (no `USER`), no healthcheck, `latest`/unpinned base, `ADD` of remote URLs, secrets in ENV, world-writable files, missing `--no-install-recommends`, exposed unnecessary ports, `sudo`/setuid binaries, curl-pipe-to-shell in RUN.
3. Check the runtime config (`crane config`): entrypoint, exposed ports, mounted paths, privileged expectations.
4. Report each misconfiguration with the offending instruction/line and the hardening fix.
Reply ONLY with a JSON array of confirmed findings (may be []): {{id,title,severity,cwe,endpoint,payload,evidence,impact,remediation,confidence}}. `endpoint` = the image ref + layer/path the finding lives in. Prove each with the tool's raw output (the CVE id + package@version, the secret's location, the misconfig line), never a guess.
## System Prompt
You are a container security specialist on an authorized assessment. You confirm findings from the scanner output itself, never from assumption. Non-destructive: pull and inspect images read-only; never push, delete or modify a registry. Redact any secret you find to a masked sample in the report.
@@ -0,0 +1,18 @@
# Container SBOM Generation
## User Prompt
You are scanning the container image **{target}** (an OCI image reference, a local tar, or a Dockerfile) for: Container SBOM Generation. CWE-1357
**Context:**
{recon_json}
All tools run HEADLESS and are provisioned on demand (time-box each install, skip on failure). Only scan images you are authorized to scan.
### Method
1. Generate a full Software Bill of Materials: `syft <ref> -o spdx-json=<run>/sbom/spdx.json -o cyclonedx-json=<run>/sbom/cyclonedx.json` (or `trivy image --format cyclonedx`).
2. Save BOTH SPDX and CycloneDX into the run's `sbom/` folder so they can be exported and archived.
3. Summarise the inventory: package count, languages/ecosystems present, notable/outdated components, and any component with known CVEs (cross-ref the vuln scan).
4. Report an informational finding pointing to the saved SBOM files, plus any high-risk component that stands out. Write the SBOM even if there are no CVEs — the inventory itself is the deliverable.
Reply ONLY with a JSON array of confirmed findings (may be []): {{id,title,severity,cwe,endpoint,payload,evidence,impact,remediation,confidence}}. `endpoint` = the image ref + layer/path the finding lives in. Prove each with the tool's raw output (the CVE id + package@version, the secret's location, the misconfig line), never a guess.
## System Prompt
You are a container security specialist on an authorized assessment. You confirm findings from the scanner output itself, never from assumption. Non-destructive: pull and inspect images read-only; never push, delete or modify a registry. Redact any secret you find to a masked sample in the report.
+18
View File
@@ -0,0 +1,18 @@
# Container Image Secret Scan
## User Prompt
You are scanning the container image **{target}** (an OCI image reference, a local tar, or a Dockerfile) for: Container Image Secret Scan. CWE-798
**Context:**
{recon_json}
All tools run HEADLESS and are provisioned on demand (time-box each install, skip on failure). Only scan images you are authorized to scan.
### Method
1. Scan every layer for exposed secrets: `trivy image --scanners secret -f json <ref>`, and/or extract layers (`docker save` / `crane export`) and run `trufflehog filesystem <dir>` / `gitleaks`.
2. Classify hits: API keys, cloud credentials, private keys/certs, tokens, DB passwords, .env/.npmrc/.dockercfg/kubeconfig baked into a layer, build-time ARG/ENV secrets left in history.
3. Check image history for secrets passed as build args: `docker history --no-trunc <ref>` / `crane config <ref>`.
4. Report each secret with its layer/path and a masked sample; note whether it is live/rotatable.
Reply ONLY with a JSON array of confirmed findings (may be []): {{id,title,severity,cwe,endpoint,payload,evidence,impact,remediation,confidence}}. `endpoint` = the image ref + layer/path the finding lives in. Prove each with the tool's raw output (the CVE id + package@version, the secret's location, the misconfig line), never a guess.
## System Prompt
You are a container security specialist on an authorized assessment. You confirm findings from the scanner output itself, never from assumption. Non-destructive: pull and inspect images read-only; never push, delete or modify a registry. Redact any secret you find to a masked sample in the report.
+18
View File
@@ -0,0 +1,18 @@
# Container Image Vulnerability Scan
## User Prompt
You are scanning the container image **{target}** (an OCI image reference, a local tar, or a Dockerfile) for: Container Image Vulnerability Scan. CWE-1104
**Context:**
{recon_json}
All tools run HEADLESS and are provisioned on demand (time-box each install, skip on failure). Only scan images you are authorized to scan.
### Method
1. Pull/inspect the image read-only. Scan for vulnerable OS + language packages with `trivy image <ref>` (or `grype <ref>`): `trivy image --scanners vuln --severity CRITICAL,HIGH,MEDIUM -f json <ref>`.
2. For each CVE: record the package@version, the fixed version, the CVE id and severity, and whether it is actually reachable (installed + in an executable layer). Prioritise KEV/known-exploited and those with a fix available.
3. Cross-check the base image age/EOL: `trivy image --scanners vuln` plus the base image tag; flag an outdated or EOL base (e.g. an old Debian/Alpine/Ubuntu release).
4. Report each meaningful CVE as a finding (package, version, fixed-in, CVE, severity) and the outdated-base issue separately.
Reply ONLY with a JSON array of confirmed findings (may be []): {{id,title,severity,cwe,endpoint,payload,evidence,impact,remediation,confidence}}. `endpoint` = the image ref + layer/path the finding lives in. Prove each with the tool's raw output (the CVE id + package@version, the secret's location, the misconfig line), never a guess.
## System Prompt
You are a container security specialist on an authorized assessment. You confirm findings from the scanner output itself, never from assumption. Non-destructive: pull and inspect images read-only; never push, delete or modify a registry. Redact any secret you find to a masked sample in the report.
+64 -2
View File
@@ -197,6 +197,12 @@ enum Cmd {
/// Run id (`ns-…`) or a path to the run directory. /// Run id (`ns-…`) or a path to the run directory.
run: String, run: String,
}, },
/// Export a run's archived HTTP traffic (from flows.jsonl) as a .http file
/// for external inspection tools.
Traffic {
/// Run id or path.
run: String,
},
/// Verify a finished run's audit trail — the hash chain and, with --anchor, /// Verify a finished run's audit trail — the hash chain and, with --anchor,
/// the signed anchors that catch truncation and silent rebuilds. /// the signed anchors that catch truncation and silent rebuilds.
Audit { Audit {
@@ -394,6 +400,27 @@ enum Cmd {
#[arg(short, long)] #[arg(short, long)]
verbose: bool, verbose: bool,
}, },
/// Container: scan an OCI image (repo:tag / tar / Dockerfile) for vulnerable
/// packages, exposed secrets and misconfigurations, and emit an SBOM
/// (SPDX + CycloneDX). Uses trivy / grype / syft headless.
Container {
/// Image reference, local tar, or Dockerfile path.
image: String,
#[arg(long = "model")]
models: Vec<String>,
#[arg(long, default_value_t = 0)]
max_agents: usize,
#[arg(long, default_value_t = 1)]
vote_n: usize,
#[arg(long)]
offline: bool,
#[arg(long)]
subscription: bool,
#[arg(long)]
focus: Option<String>,
#[arg(short, long)]
verbose: bool,
},
Host { Host {
/// Target host or IP. /// Target host or IP.
target: String, target: String,
@@ -780,6 +807,7 @@ async fn main() -> anyhow::Result<()> {
} }
} }
Cmd::Audit { run, anchor } => handle_audit(&base, &run, anchor)?, Cmd::Audit { run, anchor } => handle_audit(&base, &run, anchor)?,
Cmd::Traffic { run } => handle_traffic(&base, &run)?,
Cmd::Assurance { run, verify } => handle_assurance(&base, &run, verify)?, Cmd::Assurance { run, verify } => handle_assurance(&base, &run, verify)?,
Cmd::Compliance { run, framework, include_leads } => handle_compliance(&base, &run, &framework, include_leads)?, Cmd::Compliance { run, framework, include_leads } => handle_compliance(&base, &run, &framework, include_leads)?,
Cmd::Poc { run, repeats, apply } => handle_poc(&base, &run, repeats, apply).await?, Cmd::Poc { run, repeats, apply } => handle_poc(&base, &run, repeats, apply).await?,
@@ -900,6 +928,18 @@ async fn main() -> anyhow::Result<()> {
let out = run_mode(&base, cfg, false, Mode::Mobile).await?; let out = run_mode(&base, cfg, false, Mode::Mobile).await?;
print_findings(&out); print_findings(&out);
} }
Cmd::Container { image, models, max_agents, vote_n, offline, subscription, focus, verbose } => {
let mut cfg = RunConfig::new(&image);
cfg.max_agents = max_agents;
cfg.vote_n = vote_n;
cfg.offline = offline;
cfg.subscription = subscription;
cfg.verbose = verbose;
cfg.instructions = focus;
if !models.is_empty() { cfg.models = models; }
let out = run_mode(&base, cfg, false, Mode::Container).await?;
print_findings(&out);
}
Cmd::Host { target, models, creds, focus, max_agents, vote_n, chain_depth, recon, offline, subscription, verbose } => { Cmd::Host { target, models, creds, focus, max_agents, vote_n, chain_depth, recon, offline, subscription, verbose } => {
let mut cfg = RunConfig::new(&target); let mut cfg = RunConfig::new(&target);
cfg.max_agents = max_agents; cfg.max_agents = max_agents;
@@ -1104,7 +1144,7 @@ pub(crate) async fn apply_creds(cfg: &mut RunConfig, path: Option<&str>) {
} }
#[derive(Clone, Copy, PartialEq)] #[derive(Clone, Copy, PartialEq)]
pub(crate) enum Mode { Black, White, Grey, Host, Ai, Skills, Mobile } pub(crate) enum Mode { Black, White, Grey, Host, Ai, Skills, Mobile, Container }
pub(crate) async fn run_greybox_engagement(base: &Path, cfg: RunConfig, mcp: bool) -> anyhow::Result<RunOutput> { pub(crate) async fn run_greybox_engagement(base: &Path, cfg: RunConfig, mcp: bool) -> anyhow::Result<RunOutput> {
run_mode(base, cfg, mcp, Mode::Grey).await run_mode(base, cfg, mcp, Mode::Grey).await
@@ -1204,7 +1244,7 @@ pub(crate) fn spawn_engagement(base: &Path, mut cfg: RunConfig, mcp: bool, mode:
println!(" │ repo : {}", cfg.repo.clone().unwrap_or_default()); println!(" │ repo : {}", cfg.repo.clone().unwrap_or_default());
} }
println!(" └─ mode : {}{}{}", println!(" └─ mode : {}{}{}",
match mode { Mode::White => "white-box", Mode::Grey => "greybox", Mode::Host => "host/infra", Mode::Ai => "ai/llm", Mode::Skills => "skills/n8n audit", Mode::Mobile => "mobile/binary", Mode::Black => "black-box" }, match mode { Mode::White => "white-box", Mode::Grey => "greybox", Mode::Host => "host/infra", Mode::Ai => "ai/llm", Mode::Skills => "skills/n8n audit", Mode::Mobile => "mobile/binary", Mode::Container => "container-scan", Mode::Black => "black-box" },
if cfg.subscription { " · subscription" } else { " · api" }, if cfg.subscription { " · subscription" } else { " · api" },
if mcp { " · mcp" } else { "" }); if mcp { " · mcp" } else { "" });
@@ -1244,6 +1284,7 @@ pub(crate) fn spawn_engagement(base: &Path, mut cfg: RunConfig, mcp: bool, mode:
Mode::Ai => harness::pipeline::run_ai(cfg, &lib, &pool, tx).await, Mode::Ai => harness::pipeline::run_ai(cfg, &lib, &pool, tx).await,
Mode::Skills => harness::pipeline::run_skills_audit(cfg, &lib, &pool, tx).await, Mode::Skills => harness::pipeline::run_skills_audit(cfg, &lib, &pool, tx).await,
Mode::Mobile => harness::run_mobile(cfg, &lib, &pool, tx).await, Mode::Mobile => harness::run_mobile(cfg, &lib, &pool, tx).await,
Mode::Container => harness::run_container(cfg, &lib, &pool, tx).await,
Mode::Black => harness::run(cfg, &lib, &pool, tx).await, Mode::Black => harness::run(cfg, &lib, &pool, tx).await,
} }
}); });
@@ -1595,6 +1636,27 @@ fn handle_assurance(base: &std::path::Path, run: &str, verify: bool) -> anyhow::
Ok(()) Ok(())
} }
fn handle_traffic(base: &std::path::Path, run: &str) -> anyhow::Result<()> {
let dir = resolve_run(base, run)?;
let flows = std::fs::read_to_string(dir.join("flows.jsonl"))
.map_err(|e| anyhow::anyhow!("no flows.jsonl in {} (was the run started with --intercept own?): {e}", dir.display()))?;
let mut out = String::from("# NeuroSploit HTTP traffic archive\n# One exchange per block; bodies are not captured for tunnelled HTTPS.\n\n");
let mut n = 0usize;
for line in flows.lines().filter(|l| !l.trim().is_empty()) {
let v: serde_json::Value = match serde_json::from_str(line) { Ok(x) => x, Err(_) => continue };
let method = v.get("method").and_then(|x| x.as_str()).unwrap_or("GET");
let url = v.get("url").and_then(|x| x.as_str()).unwrap_or("");
let status = v.get("status").and_then(|x| x.as_u64()).unwrap_or(0);
let ct = v.get("content_type").and_then(|x| x.as_str()).unwrap_or("");
out.push_str(&format!("### {method} {url}\n=> HTTP {status} {ct}\n\n"));
n += 1;
}
let dest = dir.join("traffic.http");
std::fs::write(&dest, out)?;
println!(" exported {n} exchange(s) -> {}", dest.display());
Ok(())
}
fn handle_audit(base: &std::path::Path, run: &str, anchor: bool) -> anyhow::Result<()> { fn handle_audit(base: &std::path::Path, run: &str, anchor: bool) -> anyhow::Result<()> {
let dir = resolve_run(base, run)?; let dir = resolve_run(base, run)?;
let log = harness::audit::AuditLog::open(dir.join("audit.jsonl")); let log = harness::audit::AuditLog::open(dir.join("audit.jsonl"));
+9 -1
View File
@@ -100,7 +100,8 @@ fn tool_list() -> Value {
{ "name": "neurosploit_report", "description": "Read a finished run's Markdown report.", "inputSchema": { "type": "object", "properties": { "run": { "type": "string" } }, "required": ["run"] } }, { "name": "neurosploit_report", "description": "Read a finished run's Markdown report.", "inputSchema": { "type": "object", "properties": { "run": { "type": "string" } }, "required": ["run"] } },
{ "name": "neurosploit_rebuild", "description": "Rebuild a run's report artifacts from its findings (no model calls).", "inputSchema": { "type": "object", "properties": { "run": { "type": "string" } }, "required": ["run"] } }, { "name": "neurosploit_rebuild", "description": "Rebuild a run's report artifacts from its findings (no model calls).", "inputSchema": { "type": "object", "properties": { "run": { "type": "string" } }, "required": ["run"] } },
{ "name": "neurosploit_internal", "description": "Internal-network / Active Directory attack-graph analysis: paths to crown jewels and the choke point to fix first.", "inputSchema": { "type": "object", "properties": { "graph": { "type": "string", "description": "Path to a graph JSON" }, "scaffold": { "type": "string", "description": "Domain to scaffold, e.g. corp.local" }, "from": { "type": "string", "description": "Foothold node id" } } } }, { "name": "neurosploit_internal", "description": "Internal-network / Active Directory attack-graph analysis: paths to crown jewels and the choke point to fix first.", "inputSchema": { "type": "object", "properties": { "graph": { "type": "string", "description": "Path to a graph JSON" }, "scaffold": { "type": "string", "description": "Domain to scaffold, e.g. corp.local" }, "from": { "type": "string", "description": "Foothold node id" } } } },
{ "name": "neurosploit_compliance", "description": "Map a finished run's findings onto PCI-DSS, HIPAA or SOC 2 controls.", "inputSchema": { "type": "object", "properties": { "run": { "type": "string" }, "framework": { "type": "string", "enum": ["pci-dss","hipaa","soc2"] } }, "required": ["run"] } } { "name": "neurosploit_compliance", "description": "Map a finished run's findings onto PCI-DSS, HIPAA or SOC 2 controls.", "inputSchema": { "type": "object", "properties": { "run": { "type": "string" }, "framework": { "type": "string", "enum": ["pci-dss","hipaa","soc2"] } }, "required": ["run"] } },
{ "name": "neurosploit_container", "description": "Scan an OCI container image (repo:tag / tar / Dockerfile) for vulnerable packages, secrets, misconfig and emit an SBOM.", "inputSchema": { "type": "object", "properties": { "image": { "type": "string" }, "model": { "type": "string" }, "subscription": { "type": "boolean" } }, "required": ["image"] } }
]) ])
} }
@@ -143,6 +144,13 @@ fn handle_call(id: Option<Value>, req: &Value, exe: &std::path::Path) -> Value {
if let Some(sc) = s("scaffold") { argv.push("--scaffold".into()); argv.push(sc); } if let Some(sc) = s("scaffold") { argv.push("--scaffold".into()); argv.push(sc); }
if let Some(fr) = s("from") { argv.push("--from".into()); argv.push(fr); } if let Some(fr) = s("from") { argv.push("--from".into()); argv.push(fr); }
} }
"neurosploit_container" => {
let Some(image) = s("image") else { return tool_err(id, "image is required") };
argv.push("container".into()); argv.push(image);
if let Some(m) = s("model") { argv.push("--model".into()); argv.push(m); }
if b("subscription") { argv.push("--subscription".into()); }
argv.push("-v".into());
}
"neurosploit_compliance" => { "neurosploit_compliance" => {
let Some(run) = s("run") else { return tool_err(id, "run is required") }; let Some(run) = s("run") else { return tool_err(id, "run is required") };
argv.push("compliance".into()); argv.push(run); argv.push("compliance".into()); argv.push(run);
+2 -1
View File
@@ -148,7 +148,7 @@ pub async fn run(base: &Path, mut cfg: RunConfig, mcp: bool, mode: Mode) -> anyh
let (tx, mut rx) = tokio::sync::mpsc::channel::<String>(512); let (tx, mut rx) = tokio::sync::mpsc::channel::<String>(512);
let models = cfg.models.join(", "); let models = cfg.models.join(", ");
let mode_s = match mode { Mode::White => "white-box", Mode::Grey => "greybox", Mode::Host => "host/infra", Mode::Ai => "ai/llm", Mode::Skills => "skills/n8n", Mode::Mobile => "mobile/binary", Mode::Black => "black-box" }; let mode_s = match mode { Mode::White => "white-box", Mode::Grey => "greybox", Mode::Host => "host/infra", Mode::Ai => "ai/llm", Mode::Skills => "skills/n8n", Mode::Mobile => "mobile/binary", Mode::Container => "container-scan", Mode::Black => "black-box" };
let target_s = cfg.target.clone(); let target_s = cfg.target.clone();
// ---- terminal setup FIRST: on a non-TTY this errors before we spawn any // ---- terminal setup FIRST: on a non-TTY this errors before we spawn any
@@ -166,6 +166,7 @@ pub async fn run(base: &Path, mut cfg: RunConfig, mcp: bool, mode: Mode) -> anyh
Mode::Ai => harness::pipeline::run_ai(cfg, &lib, &pool, tx).await, Mode::Ai => harness::pipeline::run_ai(cfg, &lib, &pool, tx).await,
Mode::Skills => harness::pipeline::run_skills_audit(cfg, &lib, &pool, tx).await, Mode::Skills => harness::pipeline::run_skills_audit(cfg, &lib, &pool, tx).await,
Mode::Mobile => harness::run_mobile(cfg, &lib, &pool, tx).await, Mode::Mobile => harness::run_mobile(cfg, &lib, &pool, tx).await,
Mode::Container => harness::run_container(cfg, &lib, &pool, tx).await,
Mode::Black => harness::run(cfg, &lib, &pool, tx).await, Mode::Black => harness::run(cfg, &lib, &pool, tx).await,
} }
}); });
+3 -1
View File
@@ -28,12 +28,13 @@ pub struct Library {
/// AI/LLM/agent/MCP/skills security agents (OWASP LLM Top 10, MCP risks…). /// AI/LLM/agent/MCP/skills security agents (OWASP LLM Top 10, MCP risks…).
pub ai: Vec<Agent>, pub ai: Vec<Agent>,
pub mobile: Vec<Agent>, pub mobile: Vec<Agent>,
pub container: Vec<Agent>,
} }
impl Library { impl Library {
pub fn total(&self) -> usize { pub fn total(&self) -> usize {
self.vulns.len() + self.meta.len() + self.recon.len() + self.code.len() self.vulns.len() + self.meta.len() + self.recon.len() + self.code.len()
+ self.infra.len() + self.chains.len() + self.ai.len() + self.mobile.len() + self.infra.len() + self.chains.len() + self.ai.len() + self.mobile.len() + self.container.len()
} }
} }
@@ -49,6 +50,7 @@ pub fn load(base: &Path) -> Library {
chains: load_dir(&root.join("chains"), "chain"), chains: load_dir(&root.join("chains"), "chain"),
ai: load_dir(&root.join("ai"), "ai"), ai: load_dir(&root.join("ai"), "ai"),
mobile: load_dir(&root.join("mobile"), "mobile"), mobile: load_dir(&root.join("mobile"), "mobile"),
container: load_dir(&root.join("container"), "container"),
} }
} }
@@ -85,6 +85,7 @@ const KNOWN: &[(&str, &str)] = &[
("out-of-scope-findings.json", "findings quarantined for being outside scope (P2)"), ("out-of-scope-findings.json", "findings quarantined for being outside scope (P2)"),
("flows.jsonl", "intercepted request/response flows"), ("flows.jsonl", "intercepted request/response flows"),
("meta.json", "target metadata"), ("meta.json", "target metadata"),
("coverage.md", "what was tested and what was not"),
]; ];
fn hash_file(path: &Path) -> Option<(String, u64)> { fn hash_file(path: &Path) -> Option<(String, u64)> {
+1 -1
View File
@@ -58,7 +58,7 @@ pub use models::{
cli_binary_for, ensure_playwright_mcp, installed_cli_backends, mcp_supported, provider_for, cli_binary_for, ensure_playwright_mcp, installed_cli_backends, mcp_supported, provider_for,
providers, write_mcp_config, ChatClient, ModelRef, Provider, providers, write_mcp_config, ChatClient, ModelRef, Provider,
}; };
pub use pipeline::{run_greybox, run_host, run_mobile, run_whitebox, RunOutput}; pub use pipeline::{run_container, run_greybox, run_host, run_mobile, run_whitebox, RunOutput};
pub use pipeline::run; pub use pipeline::run;
pub use knowledge_graph::{EdgeKind, KnowledgeGraph, NodeKind}; pub use knowledge_graph::{EdgeKind, KnowledgeGraph, NodeKind};
pub use memory::{Memory, Query as MemoryQuery, Tier as MemoryTier}; pub use memory::{Memory, Query as MemoryQuery, Tier as MemoryTier};
@@ -461,6 +461,8 @@ fn engagement_ops(cfg: &RunConfig) -> String {
or \"unauthenticated\" (proven with no session), and `account` to which user/role you used. In grey-box, be \ or \"unauthenticated\" (proven with no session), and `account` to which user/role you used. In grey-box, be \
explicit about which findings needed a login. In black-box, record in `how`/evidence exactly what you did \ explicit about which findings needed a login. In black-box, record in `how`/evidence exactly what you did \
to create the user.\n\ to create the user.\n\
- LOGIN VERIFICATION EVIDENCE: when you authenticate (register or use given creds), CAPTURE proof the login actually worked BEFORE deep testing — save the login request/response pair to the evidence, a Playwright screenshot of the post-login page to $NEUROSPLOIT_POCS/../evidence/login-<role>.png, and record in the finding/vault whether login SUCCEEDED or FAILED and why. A pentest run should be able to show it was logged in (or explain why it could not) before claiming authenticated findings.
- HTTP TRAFFIC: the harness archives every request/response the replay layer makes; when you prove a finding, keep the exact request AND response in its `evidence` so the report can show the raw exchange behind it.
- {temp}\n\ - {temp}\n\
{oob}{sms}{waf}\n" {oob}{sms}{waf}\n"
) )
@@ -2556,6 +2558,9 @@ async fn finish(cfg: RunConfig, _lib: &Library, pool: &ModelPool, recon: String,
let _ = tx.send(n).await; let _ = tx.send(n).await;
} }
// Coverage report (Strix-style): what was tested, how, and what was NOT —
// so a reader can see the engagement's reach, not just its findings.
write_coverage(&cfg, &selected, &findings);
let artifacts = persist(&cfg, &recon, &transcript, &findings); let artifacts = persist(&cfg, &recon, &transcript, &findings);
if !artifacts.is_empty() { if !artifacts.is_empty() {
let _ = tx.send(format!("notify: evidence saved → {}", cfg.workdir.clone().unwrap_or_default())).await; let _ = tx.send(format!("notify: evidence saved → {}", cfg.workdir.clone().unwrap_or_default())).await;
@@ -2692,6 +2697,53 @@ fn provenance_key() -> Option<Vec<u8>> {
} }
/// Write recon/exploit/findings/report as json+md for downstream reuse. /// Write recon/exploit/findings/report as json+md for downstream reuse.
/// Write `coverage.md`: which agents ran (the tested surface), how many
/// findings each produced, and which high-value classes were NOT covered by the
/// selected agents. This is the "what did the pentest actually test" view.
fn write_coverage(cfg: &RunConfig, selected: &[Agent], findings: &[Finding]) {
let Some(dir) = cfg.workdir.as_deref() else { return };
let mut per_agent: std::collections::BTreeMap<&str, usize> = std::collections::BTreeMap::new();
for f in findings { *per_agent.entry(f.agent.as_str()).or_insert(0) += 1; }
let mut s = format!("# Coverage — {}\n\n", cfg.target);
s.push_str(&format!("{} agent(s) ran against the target. This is what was tested and what was not.\n\n", selected.len()));
s.push_str("## Tested\n\n| Agent | Class | Findings |\n|---|---|---|\n");
for a in selected {
let n = per_agent.get(a.name.as_str()).copied().unwrap_or(0);
s.push_str(&format!("| `{}` | {} | {} |\n", a.name, if a.cwe.is_empty() { a.title.as_str() } else { a.cwe.as_str() }, n));
}
// High-value classes a black-box web engagement should reach; flag any the
// selected agent names did not obviously cover, as an honest "not tested".
const EXPECTED: &[(&str, &[&str])] = &[
("SQL injection", &["sqli"]),
("Cross-site scripting", &["xss"]),
("Access control / IDOR / BOLA", &["idor", "bola", "bfla", "access"]),
("Authentication / session", &["auth", "jwt", "session", "login"]),
("SSRF", &["ssrf"]),
("Open redirect", &["redirect"]),
("File upload / path traversal", &["upload", "lfi", "traversal", "path"]),
("CSRF", &["csrf"]),
("Injection (cmd/template/XXE)", &["command", "ssti", "template", "xxe", "injection"]),
("Rate limiting", &["rate", "brute"]),
("Security misconfig / headers", &["header", "cors", "misconfig", "clickjack"]),
("Secrets / disclosure", &["secret", "disclosure", "exposure", "info"]),
];
let names: String = selected.iter().map(|a| a.name.to_lowercase()).collect::<Vec<_>>().join(" ");
let mut not_tested = Vec::new();
for (label, kws) in EXPECTED {
if !kws.iter().any(|k| names.contains(k)) { not_tested.push(*label); }
}
s.push_str("\n## Not tested (no agent selected for these classes)\n\n");
if not_tested.is_empty() {
s.push_str("All high-value web classes had at least one agent selected.\n");
} else {
for c in &not_tested { s.push_str(&format!("- {c}\n")); }
s.push_str("\n> These were out of the selected agent set for this run (recon-driven or `--only`). Re-run without a narrow focus, or add the agents explicitly, to cover them.\n");
}
let _ = std::fs::write(std::path::Path::new(dir).join("coverage.md"), s);
}
fn persist(cfg: &RunConfig, recon: &str, transcript: &str, findings: &[Finding]) -> Vec<String> { fn persist(cfg: &RunConfig, recon: &str, transcript: &str, findings: &[Finding]) -> Vec<String> {
let Some(dir) = &cfg.workdir else { return vec![] }; let Some(dir) = &cfg.workdir else { return vec![] };
let dir = PathBuf::from(dir); let dir = PathBuf::from(dir);
@@ -3289,6 +3341,89 @@ fn collect_repo_context(root: &Path, max_files: usize, max_bytes: usize) -> Stri
out out
} }
const CONTAINER_RECON_SYS: &str = "You are a container security recon specialist on an AUTHORIZED assessment of an OCI container image (a registry ref like repo/name:tag, a local tar, or a Dockerfile). Identify the image, its base image and OS, the layer count, the package ecosystems present, the entrypoint/exposed ports, and whether it runs as root. Pull/inspect read-only (trivy/syft/crane/docker). Do not push, delete or modify anything. Reply with a compact JSON object (image, base, os, layers, ecosystems, runs_as_root, ports). No prose.";
const CONTAINER_TOOLING: &str = "TOOLING (all read-only, provision on demand, time-box installs): `trivy image` (vuln/secret/misconfig scanners, JSON output), `grype` (vulns), `syft` (SBOM: SPDX + CycloneDX), `crane`/`skopeo` (inspect/export without a daemon), `docker save`/`docker history`, `trufflehog`/`gitleaks` (secrets in extracted layers), `hadolint` (Dockerfile). Write SBOMs into $NEUROSPLOIT_POCS/../sbom or the run's sbom/ folder. Never push/delete/modify a registry; redact secrets to a masked sample.\n\n";
/// Container image engagement: scan an OCI image (or Dockerfile) for vulnerable
/// packages, exposed secrets, misconfigurations, and produce an SBOM. Mirrors
/// the host pipeline; the target is an image ref and the agent set is `container`.
pub async fn run_container(cfg: RunConfig, lib: &Library, pool: &ModelPool, tx: Sender<String>) -> RunOutput {
pool.set_progress(tx.clone());
// The SBOM lands next to the run's pocs.
if let Some(w) = cfg.workdir.as_deref() { let _ = std::fs::create_dir_all(std::path::Path::new(w).join("sbom")); }
let _ = tx.send(format!("CONTAINER - image: {} - {} container agents - models: {}", cfg.target, lib.container.len(),
pool.candidates.iter().map(|m| m.label()).collect::<Vec<_>>().join(", "))).await;
let recon = if cfg.offline {
"{}".to_string()
} else {
let user = format!("{}{}Image: {}", operator_directives(&cfg), CONTAINER_TOOLING, cfg.target);
match pool.complete_routed(Task::Recon, "recon", CONTAINER_RECON_SYS, &user).await {
Ok((m, t)) => { let _ = tx.send(format!("recon complete via {}", m.label())).await; t }
Err(e) => { let _ = tx.send(format!("recon failed ({e})")).await; "{}".to_string() }
}
};
let mut rl = cfg.rl_path.as_ref().map(|p| RlState::load(Path::new(p))).unwrap_or_default();
let mut ranked: Vec<Agent> = lib.container.clone();
ranked.sort_by(|a, b| rl.weight(&b.name).partial_cmp(&rl.weight(&a.name)).unwrap_or(std::cmp::Ordering::Equal));
let cap = if cfg.max_agents > 0 { cfg.max_agents.min(ranked.len()) } else { ranked.len() };
let focus = cfg.instructions.clone().unwrap_or_default();
if cfg.offline {
let selected: Vec<Agent> = ranked.into_iter().take(cap).collect();
let _ = tx.send(format!("offline: selected {} container agent(s); no live scan", selected.len())).await;
let artifacts = persist(&cfg, &recon, "", &[]);
return RunOutput { target: cfg.target.clone(), workdir: cfg.workdir.clone().unwrap_or_default(), findings: vec![],
agents_ran: selected.iter().map(|a| a.name.clone()).collect(), candidates: 0, recon, artifacts, denied: None };
}
// Container agents are complementary (vuln/secret/misconfig/sbom) — run them
// all rather than a recon-based subset.
let selected: Vec<Agent> = ranked.into_iter().take(cap).collect();
let _ = tx.send(format!("running {} container agent(s): {}", selected.len(),
selected.iter().map(|a| a.name.clone()).collect::<Vec<_>>().join(", "))).await;
let target = cfg.target.clone();
let verbose = cfg.verbose;
let directives = operator_directives(&cfg);
let recon_ctx: String = recon.chars().take(3000).collect();
let raw: Vec<(String, String, Vec<Finding>)> = stream::iter(selected.iter().cloned())
.map(|ag| {
let target = target.clone();
let recon = recon_ctx.clone();
let directives = directives.clone();
let txc = tx.clone();
async move {
if pool.stop_exploiting() { return (ag.name.clone(), String::new(), vec![]); }
if verbose { let _ = txc.send(format!(" launching agent: {} ({})", ag.name, ag.title.replace(" Agent", ""))).await; }
let user = format!(
"AUTHORIZED container scan of {target}. Proceed and PROVE each issue with the scanner's raw output.\n\n{directives}{tooling}{react}{safety}{body}\n\nReply ONLY a JSON array of confirmed findings (may be []): {{id,title,severity,cwe,endpoint,payload,evidence,impact,remediation,confidence}}.",
target = target, directives = directives, tooling = CONTAINER_TOOLING, react = REACT_DOCTRINE, safety = SAFETY_DOCTRINE,
body = ag.user.replace("{target}", &target).replace("{recon_json}", &recon),
);
match pool.complete_routed(Task::Exploit, &ag.name, &ag.system, &user).await {
Ok((m, text)) => {
let f = extract_findings(&text, &ag.name);
let _ = txc.send(format!("scan {} via {} -> {} finding(s)", ag.name, m.label(), f.len())).await;
for c in &f { if let Ok(j) = serde_json::to_string(c) { let _ = txc.send(format!("finding_json: {j}")).await; } }
(ag.name.clone(), text, f)
}
Err(e) => { let _ = txc.send(format!("scan {} failed: {e}", ag.name)).await; (ag.name.clone(), format!("ERROR: {e}"), vec![]) }
}
}
})
.buffer_unordered(cfg.concurrency)
.collect::<Vec<_>>().await;
let transcript = transcript_of(&raw);
let candidates = dedup_findings(raw.iter().flat_map(|(_, _, f)| f.clone()).collect());
let _ = tx.send(format!("{} finding(s) (deduped) - validating", candidates.len())).await;
let findings = validate(candidates, pool, VOTE_SYS, effective_vote_n(&cfg), &tx).await;
finish(cfg, lib, pool, recon, transcript, findings, selected, &mut rl, crate::grounding::GroundMode::Empirical, String::new(), tx).await
}
const MOBILE_RECON_SYS: &str = "You are a mobile/binary reverse-engineering recon specialist on an AUTHORIZED assessment of a LOCAL artifact (a binary, APK or IPA on disk). Identify format/arch, package metadata, entry points, protection layers (RASP/anti-tamper, root/JB and anti-debug detection, TLS pinning, obfuscation/packing), the attack surface (exported components, URL schemes, entitlements, linked frameworks) and hardcoded secrets/endpoints. Run everything HEADLESS (MobSF REST/Docker, Ghidra analyzeHeadless, apktool, jadx, otool/nm, r2). Do not ask permission; proceed. Reply with a compact JSON object (format, arch, package, protections, surface, secrets). No prose."; const MOBILE_RECON_SYS: &str = "You are a mobile/binary reverse-engineering recon specialist on an AUTHORIZED assessment of a LOCAL artifact (a binary, APK or IPA on disk). Identify format/arch, package metadata, entry points, protection layers (RASP/anti-tamper, root/JB and anti-debug detection, TLS pinning, obfuscation/packing), the attack surface (exported components, URL schemes, entitlements, linked frameworks) and hardcoded secrets/endpoints. Run everything HEADLESS (MobSF REST/Docker, Ghidra analyzeHeadless, apktool, jadx, otool/nm, r2). Do not ask permission; proceed. Reply with a compact JSON object (format, arch, package, protections, surface, secrets). No prose.";
const MOBILE_TOOLING: &str = "TOOLING (all HEADLESS; provision on demand, time-box installs): APK/IPA static -> MobSF via its REST API (Docker image), `apktool`, `jadx`, `apkleaks`; binaries -> Ghidra `analyzeHeadless`, `radare2`/`rizin`, `binwalk`, `checksec`, `nm`/`otool`/`objdump`, `class-dump`; dynamic -> `frida`/`objection` for detection/pinning/anti-debug bypass; secrets -> `trufflehog`/`gitleaks`. Never require a GUI or an X display. Analyse and instrument non-destructively; never exfiltrate real user data.\n\ const MOBILE_TOOLING: &str = "TOOLING (all HEADLESS; provision on demand, time-box installs): APK/IPA static -> MobSF via its REST API (Docker image), `apktool`, `jadx`, `apkleaks`; binaries -> Ghidra `analyzeHeadless`, `radare2`/`rizin`, `binwalk`, `checksec`, `nm`/`otool`/`objdump`, `class-dump`; dynamic -> `frida`/`objection` for detection/pinning/anti-debug bypass; secrets -> `trufflehog`/`gitleaks`. Never require a GUI or an X display. Analyse and instrument non-destructively; never exfiltrate real user data.\n\