v4.2.2: validation keeps reproducible findings, /pocs command, version bump

Field feedback: real, reproducible findings (missing HSTS, insecure cookie flags,
internal IP leaked in a header) were being down-rated/dropped by the adversarial
opinion-vote. The user's rule: never discard something real; another person must
be able to reproduce the same finding.

Validation:
- has_http_receipt(): a finding whose proof is a captured HTTP response (status
  line / security headers / Set-Cookie / a header the class is proven by) or a
  file:line citation is REPRODUCIBLE by definition.
- validate(): a finding unanimously rejected by the vote but carrying such a
  receipt is NO LONGER dropped — it is kept as needs-review (capped to what the
  receipt alone proves), because "the response lacks HSTS" is a fact, not a
  story. grounded_receipt() now also recognizes an in-prose HTTP receipt.
- reproducibility: a kept finding at a URL with no explicit repro steps gets a
  minimal pasteable `curl -i` so anyone can reproduce the exact finding.

Also:
- /pocs (aliases /poc /evidence /artifacts): list a run's synthesized/written
  PoC + evidence files with their paths (was: "unknown command /pocs").
- version bumped to 4.2.2 across CLI/clap/web; stale v4.2.1/v4.1.0 labels fixed.

423 tests passing.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
This commit is contained in:
CyberSecurityUPandClaude Opus 4.8 committed 2026-10-04 00:40:12 -03:00
1 parent 45ad7bf359
commit b55778d6f6
9 files changed
+141 -34

No files matched your search

+18 -3
View File
@@ -1,4 +1,4 @@
<h1 align="center">🧠 NeuroSploit v4.2.1</h1>
<h1 align="center">🧠 NeuroSploit v4.2.2</h1>
<p align="center">
<a href="https://github.com/JoasASantos/NeuroSploit/stargazers"><img src="https://img.shields.io/github/stars/JoasASantos/NeuroSploit?style=for-the-badge&logo=github&color=8b5cf6" alt="Stars"></a>
@@ -8,7 +8,7 @@
</p>
<p align="center">
<img src="https://img.shields.io/badge/Version-4.2.1-blue?style=flat-square">
<img src="https://img.shields.io/badge/Version-4.2.2-blue?style=flat-square">
<img src="https://img.shields.io/badge/Harness-Rust%20%7C%20tokio-e6b673?style=flat-square">
<img src="https://img.shields.io/badge/License-MIT-green?style=flat-square">
<img src="https://img.shields.io/badge/MD%20Agents-479-red?style=flat-square">
@@ -52,7 +52,22 @@ Control TUI**.
### Highlights
> **New in v4.2.1** — **SARIF 2.1.0 export**: every run now writes `report.sarif`
> **New in v4.2.2** — **free, LLM-directed exploration**: an exploit agent's named
> class is a starting point, not a cage — it maps what the app actually does and
> reports any class it can prove, with **authentication / identity** (login, signup,
> password reset, MFA, OAuth/OIDC/SAML, JWT, session) as a first-class target and
> business-logic / multi-step flows pursued on its own judgment; agent **selection**
> now covers the surface instead of collapsing into one family. **WAF-aware User-Agent**
> (`/ua browser`) uses a realistic browser UA for accuracy behind a CDN while keeping
> attribution in the `X-NeuroSploit-Scan` header. **Importable engagement configs**:
> `/authorize <hosts…>` sets the whole scope in one line (no bug-bounty program needed),
> `/scope-file <yaml>` imports scope **and** target/models/focus/classes from one file;
> `/class idor,sqli,xss,ssrf` focuses a run on vuln classes. Plus **PoC/evidence files
> synthesized from recorded evidence** even on the API-key path (empty `pocs/`·`evidence/`
> fixed), **model-refusal detection** (a declined technique is reported as such, not a
> parse error), and version strings read from the build so they never go stale.
>
> **Also in v4.2.x** — **SARIF 2.1.0 export**: every run now writes `report.sarif`
> next to the Markdown/JSON/HTML/PDF, and `neurosploit sarif <run>` (re)emits it
> on demand, so findings drop straight into GitHub / Azure DevOps code-scanning
> as severity-coloured, CWE-linked alerts (also exposed over MCP). Plus stronger
+2 -2
View File
@@ -1,4 +1,4 @@
# NeuroSploit — Tutorial & User Guide (v4.2.1)
# NeuroSploit — Tutorial & User Guide (v4.2.2)
A complete, hands-on guide to installing, configuring and running NeuroSploit —
the autonomous, multi-model penetration-testing harness.
@@ -100,7 +100,7 @@ Agents **degrade gracefully**: if `rustscan` is absent they use `nmap`; if neith
### Verify
```bash
neurosploit --version # neurosploit 4.2.1
neurosploit --version # neurosploit 4.2.2
neurosploit agents # {"vulns":255,...,"ai":30,...,"total":480}
neurosploit models # all providers & models
```
+2 -2
View File
@@ -940,7 +940,7 @@ dependencies = [
[[package]]
name = "neurosploit"
version = "4.2.1"
version = "4.2.2"
dependencies = [
"anyhow",
"clap",
@@ -957,7 +957,7 @@ dependencies = [
[[package]]
name = "neurosploit-harness"
version = "4.2.1"
version = "4.2.2"
dependencies = [
"anyhow",
"base64",
+1 -1
View File
@@ -3,7 +3,7 @@ members = ["crates/harness", "app"]
resolver = "2"
[workspace.package]
version = "4.2.1"
version = "4.2.2"
edition = "2021"
license = "MIT"
repository = "https://github.com/JoasASantos/NeuroSploit"
+2 -2
View File
@@ -13,8 +13,8 @@ use std::path::{Path, PathBuf};
#[command(
name = "neurosploit",
version,
about = "NeuroSploit v4.2.1 — multi-model autonomous pentest harness",
long_about = "NeuroSploit v4.2.1 — a Rust multi-model harness that drives a pool of LLMs \
about = "NeuroSploit v4.2.2 — multi-model autonomous pentest harness",
long_about = "NeuroSploit v4.2.2 — a Rust multi-model harness that drives a pool of LLMs \
(API key or local subscription: Claude/Codex/Gemini/Grok/OpenCode/Hermes) to autonomously test a target. \
After recon it INTELLIGENTLY selects only the agents matching the discovered surface, runs \
them in parallel, then validates every finding by cross-model voting before reporting.\n\n\
+27 -2
View File
@@ -148,7 +148,7 @@ struct LiveCheckpoint {
pub(crate) const ACCEPTED: &[&str] = &[
"/?", "/agents", "/attach", "/audit", "/auth", "/burp", "/cap", "/capability", "/chain", "/changed", "/clear", "/config",
"/context", "/continue", "/creds", "/diff", "/exclude", "/exit", "/expand", "/feed",
"/finding", "/findings", "/focus", "/forget", "/full", "/go", "/goal", "/graph", "/guardrail", "/guardrails", "/help",
"/pocs", "/poc", "/evidence", "/artifacts", "/finding", "/findings", "/focus", "/forget", "/full", "/go", "/goal", "/graph", "/guardrail", "/guardrails", "/help",
"/history", "/idle", "/inscope", "/instructions", "/integration", "/integrations", "/key", "/log",
"/authorize", "/grant", "/inscope-set", "/scope-file", "/scopefile", "/import-scope", "/authorization", "/authz", "/program", "/logs", "/mcp", "/memory", "/model", "/models", "/objective", "/objectives", "/observe",
"/observe-only", "/offline",
@@ -164,7 +164,7 @@ const COMMANDS: &[&str] = &[
"/help", "/onboard", "/show", "/config", "/providers", "/model", "/key", "/sub", "/target",
"/scope-file", "/authorization", "/class", "/repo", "/auth", "/creds", "/focus", "/objective", "/scope-out", "/attach", "/context", "/mcp", "/offline",
"/class", "/research", "/quick", "/economy", "/eco", "/votes", "/chain", "/recon", "/tempmail", "/timeout", "/proxy", "/burp", "/ua", "/agents", "/only", "/theme", "/clear", "/run", "/stop", "/pause", "/continue", "/runs", "/results", "/report",
"/status", "/logs", "/diff", "/retest", "/validate", "/finding", "/expand", "/integrations",
"/status", "/logs", "/diff", "/retest", "/validate", "/finding", "/pocs", "/expand", "/integrations",
"/memory", "/forget", "/graph", "/inscope", "/observe", "/guardrail", "/policy",
"/capability", "/audit", "/quit",
];
@@ -1277,6 +1277,31 @@ pub async fn repl(base: &Path, auth: SessionAuth) -> anyhow::Result<()> {
if live_now { println!(" \x1b[2m(run still streaming in background — /logs for what happened while browsing)\x1b[0m"); }
}
}
"/pocs" | "/poc" | "/evidence" | "/artifacts" => {
// List the PoC and evidence artifacts a run produced (synthesized
// or agent-written), with their paths so they can be retrieved.
let h = history.lock().unwrap();
let rec = if arg.trim().is_empty() { h.last() } else { pick(&h, arg) };
match rec {
None => println!(" no run yet — /runs to list, or /pocs <id> after a run"),
Some(r) if r.workdir.is_empty() => println!(" run #{} has no workdir on record", r.id),
Some(r) => {
let dir = std::path::Path::new(&r.workdir);
let mut any = false;
for (sub, label) in [("pocs", "PoC scripts"), ("evidence", "evidence")] {
let p = dir.join(sub);
let mut files: Vec<String> = std::fs::read_dir(&p).map(|rd| rd.filter_map(|e| e.ok()).map(|e| e.file_name().to_string_lossy().to_string()).collect()).unwrap_or_default();
files.sort();
if !files.is_empty() {
any = true;
println!(" \x1b[1m{label}\x1b[0m ({}):", p.display());
for f in files { println!(" {}", p.join(&f).display()); }
}
}
if !any { println!(" no pocs/ or evidence/ files for run #{} ({})", r.id, dir.display()); }
}
}
}
"/finding" | "/findings" => {
// Build the finding pool: live run if active, else a past run.
let pool: Vec<Finding> = match &active {
+84 -17
View File
@@ -1935,26 +1935,35 @@ async fn validate(candidates: Vec<Finding>, pool: &ModelPool, sys: &str, vote_n:
f.review_status = "needs-review".into();
f.review_reason = format!("voter rejected the narrative, mechanic retained: {}", f.review_reason.trim());
f.confidence = f.confidence.min(0.5);
} else if grounded_receipt(&f) {
// Unanimously rejected, but the MECHANISM was demonstrated —
// a real engagement rejected "no rate limiting on the reset
// flow" because the agent claimed email flooding and only
// proved that 25 requests went through unthrottled. The
// claim was inflated; the measurement was real, and
// discarding it hid a genuine gap from the report.
//
// So an over-claimed finding is capped and flagged rather
// than deleted: the reader gets the fact, not the story
// that was built on it.
} else if grounded_receipt(&f) || has_http_receipt(&f) {
// Unanimously rejected by the opinion-vote, but the finding
// carries a concrete, REPRODUCIBLE receipt (a captured HTTP
// response, a header/cookie the class is proven by, or a
// file:line citation). Another person can reproduce this
// exact finding with one request — so it is NOT dropped.
// The adversarial validator is tuned to reject low-impact and
// "theoretical" issues, but "the response lacks HSTS" or "the
// cookie has no Secure flag" is a FACT, not a story. We cap
// the severity to what the receipt alone proves and flag it
// for human review, never delete it.
let cap = "Low";
let was = f.severity.clone();
f.severity = cap.to_string();
if sev_rank(&f.severity) < sev_rank(cap) { f.severity = cap.to_string(); }
f.review_status = "needs-review".into();
f.review_reason = format!(
"impact not demonstrated — capped from {was} to {cap}. Validator: {}",
f.review_reason.trim()
);
f.confidence = f.confidence.min(0.5);
f.review_reason = if was == f.severity {
format!("kept — reproducible receipt present; impact not independently demonstrated. Validator: {}", f.review_reason.trim())
} else {
format!("kept & capped from {was} to {} — reproducible receipt present, impact not demonstrated. Validator: {}", f.severity, f.review_reason.trim())
};
f.confidence = f.confidence.max(0.3).min(0.6);
}
// Reproducibility: a kept finding that reaches a URL endpoint but
// carries no explicit repro steps gets a minimal, pasteable one,
// so another person can reproduce the exact finding.
if (f.validated || f.review_status == "needs-review")
&& f.repro_steps.is_empty()
&& (f.endpoint.starts_with("http://") || f.endpoint.starts_with("https://")) {
f.repro_steps = vec![format!("curl -i -s '{}' # inspect the response (status + headers + body) that proves this finding", f.endpoint)];
}
let label = if f.validated { "CONFIRMED" } else if f.review_status == "needs-review" { "needs-review" } else { "rejected" };
let _ = txc.send(format!("vote {} → {} ({})", f.title, label, f.votes)).await;
@@ -2154,9 +2163,39 @@ fn grounded_receipt(f: &Finding) -> bool {
if f.review_reason.contains("receipt_missing") || f.votes.contains("receipt_missing") {
return false;
}
if has_http_receipt(f) {
return true;
}
crate::grounding::ground(f, "", crate::grounding::GroundMode::Either).ok
}
/// True when the finding's evidence carries a concrete, reproducible HTTP
/// receipt — a captured response (status line / headers / Set-Cookie) or a
/// `file:line` code citation. This is the test for "another person can
/// reproduce this exact finding": if the proof is the response itself (a missing
/// security header, an insecure cookie flag, an internal IP leaked in a header,
/// a status code), it is reproducible with a single request and must never be
/// discarded by an opinion-based vote — demoted to needs-review at worst.
fn has_http_receipt(f: &Finding) -> bool {
if f.evidence_data.is_some() { return true; }
let hay = format!("{}\n{}", f.evidence, f.payload).to_lowercase();
// A captured HTTP response or the headers/fields these deterministic classes
// are proven by.
const SIGNS: &[&str] = &[
"http/1.1", "http/2", "http/1.0", "status: ", "status code", "status=",
"set-cookie", "strict-transport-security", "x-frame-options",
"content-security-policy", "access-control-allow-origin", "location:",
"server:", "www-authenticate", "< http", "=> http", "response:", "200 ok",
"301 ", "302 ", "401 ", "403 ", "404 ", "500 ", "curl ",
];
if SIGNS.iter().any(|s| hay.contains(s)) { return true; }
// A white-box file:line citation is also a reproducible receipt.
f.endpoint.contains(':') && f.endpoint.chars().any(|c| c.is_ascii_digit())
&& (f.endpoint.contains(".rs") || f.endpoint.contains(".py") || f.endpoint.contains(".js")
|| f.endpoint.contains(".php") || f.endpoint.contains(".java") || f.endpoint.contains(".go")
|| f.endpoint.contains(".ts") || f.endpoint.contains(".rb"))
}
/// Adversarial refutation pass: every confirmed **High/Critical** finding is
/// re-examined by a skeptical panel that tries to prove it's a false positive.
/// A finding that fails to withstand a majority of skeptics is dropped. Lower
@@ -4440,6 +4479,34 @@ mod extraction_tests {
assert!(!looks_like_refusal("No vulnerabilities were found in the tested endpoints."));
}
#[test]
fn a_response_backed_finding_is_a_reproducible_receipt_and_never_dropped() {
// Missing HSTS: the proof is the response headers — reproducible with one
// request, must never be discarded by an opinion-vote.
let hsts = Finding {
title: "No Strict-Transport-Security header".into(),
severity: "Low".into(), cwe: "CWE-319".into(),
endpoint: "https://scapi.rockstargames.com/".into(),
evidence: "HTTP/2 200\nserver: cloudflare\n(no strict-transport-security header present)".into(),
..Default::default()
};
assert!(has_http_receipt(&hsts), "a captured response IS a receipt");
// Insecure cookie flags — Set-Cookie in evidence.
let cookie = Finding {
title: "Cookie without Secure/HttpOnly".into(), severity: "Low".into(), cwe: "CWE-614".into(),
endpoint: "https://scapi.rockstargames.com/".into(),
evidence: "Set-Cookie: bal=1; path=/ (no Secure, no HttpOnly, no SameSite)".into(),
..Default::default()
};
assert!(has_http_receipt(&cookie));
// Pure prose with no receipt is NOT a reproducible receipt.
let vague = Finding { title: "Maybe vulnerable".into(), evidence: "the app seems insecure".into(), ..Default::default() };
assert!(!has_http_receipt(&vague));
// A white-box file:line citation IS a receipt.
let wb = Finding { title: "SQLi".into(), endpoint: "src/db.py:42".into(), evidence: "query = f\"...{id}\"".into(), ..Default::default() };
assert!(has_http_receipt(&wb));
}
/// The bare-array happy path for agent selection.
#[test]
fn a_string_array_parses_plain() {
+2 -2
View File
@@ -3,7 +3,7 @@
<head>
<meta charset="utf-8" />
<meta name="viewport" content="width=device-width, initial-scale=1" />
<title>NeuroSploit v4.2.1 — Console</title>
<title>NeuroSploit v4.2.2 — Console</title>
<link rel="icon" href="data:image/svg+xml,<svg xmlns=%22http://www.w3.org/2000/svg%22 viewBox=%220 0 100 100%22><text y=%22.9em%22 font-size=%2290%22>🧠</text></svg>">
<link rel="stylesheet" href="/vendor/xterm.css" />
<link rel="stylesheet" href="/style.css" />
@@ -33,7 +33,7 @@
<div class="sb-groups" id="sbGroups"><!-- populated by app.js --></div>
<div class="sb-bottom">
<span class="sb-version" id="sbVersion">v4.2.1</span>
<span class="sb-version" id="sbVersion">v4.2.2</span>
<div class="sb-bottom-actions">
<button class="icon-btn" id="btnOpenAuth" title="Auth &amp; API keys">🔑</button>
<button class="icon-btn" id="btnOpenRepl" title="Open terminal (Ctrl+`)">❭_</button>
+3 -3
View File
@@ -1,7 +1,7 @@
#!/usr/bin/env node
'use strict';
/**
* NeuroSploit v4.2.1 — web console backend.
* NeuroSploit v4.2.2 — web console backend.
*
* Zero-dependency Node HTTP server that:
* - serves the static SPA in ./public
@@ -1487,7 +1487,7 @@ const server = http.createServer(async (req, res) => {
}
if (req.method === 'GET' && p === '/api/meta') {
return sendJson(res, 200, { version: '4.2.1', binary: BIN, root: ROOT });
return sendJson(res, 200, { version: "4.2.2", binary: BIN, root: ROOT });
}
// ---- providers / API keys (in-memory only, never persisted) ----
@@ -1521,7 +1521,7 @@ const server = http.createServer(async (req, res) => {
loadPersistedJobs();
server.listen(PORT, () => {
console.log(`NeuroSploit v4.2.1 web console → http://localhost:${PORT}`);
console.log(`NeuroSploit v4.2.2 web console → http://localhost:${PORT}`);
console.log(` binary : ${BIN || '(not found — build neurosploit-rs first)'}`);
console.log(` agents : ${AGENTS_DIR}`);
console.log(` runs : ${RUNS_DIR}`);