Compare commits

...
8 Commits
Author SHA1 Message Date
CyberSecurityUP a61e75b601 v3.6.4: fix #33 — mode-aware grounding so white-box SAST findings aren't demoted
The grounding gate ran in empirical mode for every engagement, demoting
white-box (and skills/n8n audit) findings that had passed the n-model vote
because a file:line code citation isn't raw tool output. Grounding is now
mode-aware:
- Symbolic (white-box SAST / skills): a file:line reference into the reviewed
  source, or a quote of code present in it, is the receipt — no live target.
- Empirical (black-box / host / AI): evidence must resemble tool output (as before).
- Either (grey-box): a source citation OR a tool receipt grounds a finding.
The symbolic check runs against the reviewed source corpus (not the transcript)
and falls back to a structural file:line + quote check when the corpus is
unavailable. Adds unit tests incl. a regression test for #33.
2026-07-19 17:48:19 -03:00
CyberSecurityUP 53c07b9a9c v3.6.3: resumable interrupted runs + crash-proof mid-run browsing
- /continue (and /resume) now relaunch a recovered interrupted run on the same
  target, carrying its findings forward and steering agents to widen coverage /
  chain from them instead of re-reporting. Offer shown at launch; a fresh /run
  supersedes it. Findings merge (dedup by title+endpoint) across both runs.
- Opening /results, /finding or /report while a run streams no longer corrupts
  the terminal: live background output is paused for the picker (still captured
  in /logs) and restored on exit, so Ctrl-C in a picker can't take the process
  down mid-run.
2026-07-10 21:44:22 -03:00
CyberSecurityUP 865611d552 recon: time-box tool installs and skip on failure — never stall on a download
A missing or un-downloadable recon tool must never block the run. Both the recon
intensity directive and the general tool doctrine now instruct agents to:
- wrap every install in `timeout 90 <install> || echo skip` and run non-interactively
- try each tool install at most once; on failure/no-package/no-network/hang, skip
  immediately and fall back to an installed alternative or curl/nc/dig/python3
- never wait on, retry, or block the whole recon for a single tool download
2026-07-10 17:35:11 -03:00
CyberSecurityUP ce31478068 v3.6.2: stream Codex tool-by-tool + capture agent commands in /logs & /status
- Drive `codex exec --json` and parse its JSONL event stream into the same
  categorized live feed as Claude Code (exec/edit/tool/net/tokens), so recon and
  exploitation are visible as each command runs instead of a silent black box.
- Fix the activity feed to keep per-agent tool events (commands, network, files,
  findings) and only filter model reasoning + token telemetry, so /logs shows the
  real command trail and /status 'last:' is a true sign-of-life.
- Surface failed internal commands as 'exec: (exit N)'; keep Codex auth/rate
  detection from stderr.
2026-07-10 17:28:40 -03:00
CyberSecurityUP 98616bca0b repl: richer /status (works during recon) + new /logs activity feed
- /status now shows progress in EVERY phase: a real bar once agents are selected,
  otherwise the current pre-exploit phase + counters (cmds, activity lines), plus
  a "last:" sign-of-life line (the latest activity) and the actual full findings.
  Before, the bar only appeared after agent selection, so a long recon looked
  frozen. Findings count now uses the full list.
- New /logs [n] — dump the recent activity feed (recon/tools/findings) of the
  running test; useful with non-streaming CLIs (codex) or after scrolling. Backed
  by a capped feed ring buffer + last/lines counters in RunLive.
2026-07-10 17:08:26 -03:00
CyberSecurityUP 5b9d485025 fix(repl): show recon/probe activity + don't let the idle guardrail kill recon
Symptom: with a non-streaming subscription CLI (codex), a long/intense recon
showed nothing in the feed ("phase starting") and the 5-min idle guardrail killed
the run before any agent ran.
- render_compact now SHOWS recon/probe/ai-recon/skills-audit/loaded/running lines
  (were dropped) so a long recon no longer looks frozen.
- Idle guardrail reworked: resets on ANY streamed activity (not only new
  findings) and only ARMS after exploitation starts (agent launch / vote) — recon
  can never trip it. Message: "no activity in N min".
- RunLive.ingest sets phase=recon on recon/probe lines (was stuck at "starting").
2026-07-10 17:01:42 -03:00
CyberSecurityUP d9c191ec39 fix(cli): codex exec exit-1 no longer discards a valid recon/agent result
`codex exec` in --dangerously-bypass-approvals-and-sandbox mode exits non-zero
when a tool/command it ran internally (curl/nmap/etc.) returned non-zero — even
though it produced a valid final answer. chat_cli treated any non-zero exit as a
hard failure and dropped the output ("recon round 1 failed ... exit 1"). Now, on
non-zero exit WITH usable stdout and no auth/rate/quota keyword, we use the
output; only genuine auth/rate/quota errors (or empty output) fail hard.
2026-07-10 16:37:00 -03:00
CyberSecurityUP 54bf424c1d v3.6.1 — add GPT-5.6 models (sol / terra / luna)
Added the OpenAI GPT-5.6 line to the provider pool: gpt-5.6-sol (frontier/default),
gpt-5.6-terra (balanced), gpt-5.6-luna (fast/affordable). Version 3.6.0 -> 3.6.1.
2026-07-10 16:27:28 -03:00
18 changed files with 651 additions and 133 deletions
+16 -13
View File
@@ -1,4 +1,4 @@
<h1 align="center">🧠 NeuroSploit v3.6.0</h1> <h1 align="center">🧠 NeuroSploit v3.6.4</h1>
<p align="center"> <p align="center">
<a href="https://trendshift.io/repositories/22624?utm_source=trendshift-badge&amp;utm_medium=badge&amp;utm_campaign=badge-trendshift-22624" target="_blank" rel="noopener noreferrer"><img src="https://trendshift.io/api/badge/trendshift/repositories/22624/daily?language=Python" alt="JoasASantos%2FNeuroSploit | Trendshift" width="250" height="55"/></a> <a href="https://trendshift.io/repositories/22624?utm_source=trendshift-badge&amp;utm_medium=badge&amp;utm_campaign=badge-trendshift-22624" target="_blank" rel="noopener noreferrer"><img src="https://trendshift.io/api/badge/trendshift/repositories/22624/daily?language=Python" alt="JoasASantos%2FNeuroSploit | Trendshift" width="250" height="55"/></a>
@@ -12,7 +12,7 @@
</p> </p>
<p align="center"> <p align="center">
<img src="https://img.shields.io/badge/Version-3.6.0-blue?style=flat-square"> <img src="https://img.shields.io/badge/Version-3.6.4-blue?style=flat-square">
<img src="https://img.shields.io/badge/Harness-Rust%20%7C%20tokio-e6b673?style=flat-square"> <img src="https://img.shields.io/badge/Harness-Rust%20%7C%20tokio-e6b673?style=flat-square">
<img src="https://img.shields.io/badge/License-MIT-green?style=flat-square"> <img src="https://img.shields.io/badge/License-MIT-green?style=flat-square">
<img src="https://img.shields.io/badge/MD%20Agents-417-red?style=flat-square"> <img src="https://img.shields.io/badge/MD%20Agents-417-red?style=flat-square">
@@ -28,14 +28,15 @@
> >
> 📖 **New here? Read the [full Tutorial & User Guide →](TUTORIAL.md)** — every mode, flag, config and example explained. > 📖 **New here? Read the [full Tutorial & User Guide →](TUTORIAL.md)** — every mode, flag, config and example explained.
> 🆕 **New in v3.6.0Cloud testing + REPL navigation + deeper recon:** > 🆕 **New in v3.6.4white-box findings no longer wrongly demoted ([#33](https://github.com/JoasASantos/NeuroSploit/issues/33)):**
> **AWS/GCP/Azure** agents (+17 → **375** total) with credentials wired through > The grounding gate ran in **empirical** mode for **every** engagement, so
> `creds.yaml`; a more navigable **REPL** — **`/timeout`** idle guardrail, > white-box SAST & skills-audit findings — whose evidence is a `file:line` code
> **multi-target** `/target a,b,c` (sequential), an interactive **`/results`** > citation, not HTTP/tool output — were demoted as "receipt missing" and dropped
> browser (target → vuln → detail, Esc to go back) and **`/report`** picker; and > from the report even after passing the n-model vote. Grounding is now
> **deeper recon** that downloads & analyzes JavaScript (endpoints, secrets, > **mode-aware**: *symbolic* (`file:line` into the reviewed source) for
> source maps) and does request/response differential analysis. Interactive > white-box/skills, *empirical* for black-box/host/AI, *either* for grey-box.
> line-editing prompt bug fixed. > *(v3.6.3 added resumable interrupted runs + crash-proof mid-run browsing;
> v3.6.2 live Codex tool-by-tool streaming; v3.6.1 GPT-5.6 sol/terra/luna.)*
> *(v3.5.4 added robust attack chaining + false-positive reduction; v3.5.3 > *(v3.5.4 added robust attack chaining + false-positive reduction; v3.5.3
> GitHub/GitLab/Jira **[integrations](TUTORIAL-INTEGRATION.md)**; v3.5.2 the DEPTH > GitHub/GitLab/Jira **[integrations](TUTORIAL-INTEGRATION.md)**; v3.5.2 the DEPTH
> doctrine + report-hygiene — see [RELEASE.md](RELEASE.md).)* > doctrine + report-hygiene — see [RELEASE.md](RELEASE.md).)*
@@ -69,9 +70,11 @@ Control TUI**.
and "scan more vs exploit now" falls out of belief entropy. The `may_assert` and "scan more vs exploit now" falls out of belief entropy. The `may_assert`
gate is a **mathematical anti-hallucination rule** (don't claim exploitability gate is a **mathematical anti-hallucination rule** (don't claim exploitability
while the belief is diffuse). while the belief is diffuse).
- 🧾 **Grounding** — hard rule: **no claim without a tool receipt** (raw tool - 🧾 **Grounding** — hard rule: **no claim without a receipt** (evidence, not
output, not paraphrase). Empirical for black-box, symbolic (`file:line`) for paraphrase). Empirical (raw tool output) for black-box/host/AI, **symbolic**
white-box; ungrounded claims are demoted. (`file:line` into the reviewed source — a code citation *is* the receipt) for
white-box SAST & skills audits, and **either** for grey-box; ungrounded claims
are demoted.
- 🔬 **Deterministic HTTP probe** — before the model recon, the harness runs a - 🔬 **Deterministic HTTP probe** — before the model recon, the harness runs a
**real** request/response analysis (status/redirects, security headers, cookie **real** request/response analysis (status/redirects, security headers, cookie
flags, CORS reflection, tech fingerprint, linked JS, 404 baseline, high-signal flags, CORS reflection, tech fingerprint, linked JS, 404 baseline, high-signal
+82
View File
@@ -1,3 +1,85 @@
# NeuroSploit v3.6.4 — Release Notes
**Release Date:** July 2026
**Codename:** Symbolic Grounding
**License:** MIT
**Credits:** Joas A Santos & Red Team Leaders
---
## Highlights
- **Fix ([#33](https://github.com/JoasASantos/NeuroSploit/issues/33)): white-box
findings were silently dropped from the report.** The grounding gate — the
anti-hallucination step that demotes any claim lacking a receipt — was running
in **empirical** mode for *every* engagement. Empirical grounding looks for raw
tool output (HTTP responses, error oracles, shell receipts), which a **SAST
finding never has**: its receipt is a `file:line` reference into the reviewed
source. So white-box (and skills/n8n audit) findings that had *passed* the
n-model vote were then demoted as "receipt missing" and never reported.
Grounding is now **mode-aware**:
- **Symbolic** — white-box SAST & skills audits: a `file:line` (or
`file:section`) reference into the reviewed source, or a quote of code that
appears in it, IS the receipt. No live target needed.
- **Empirical** — black-box / host / AI endpoints: evidence must resemble raw
tool output (unchanged behaviour).
- **Either** — grey-box: a source citation OR a tool receipt grounds a finding.
The symbolic check is run against the reviewed **source corpus** (not the model
transcript), and falls back to a structural `file:line` + code-quote check when
the corpus isn't available, so a well-formed SAST finding is never dropped on a
technicality. Covered by unit tests (including a regression test for #33).
---
## Previously in v3.6.3
- **Interrupted runs are resumable.** When a run is cut off (terminal closed,
Ctrl-C, crash, SSH drop), its findings were already checkpointed live and
recovered as a run on the next launch. Now `/continue` (or `/resume`) also
**relaunches the engagement** on the same target and **carries those findings
forward** — steering agents to widen coverage and chain from what was already
found instead of re-reporting it. The offer is shown at launch right under the
recovery line. A fresh `/run` supersedes the pending resume.
- **Browsing no longer kills a live run.** Opening `/results`, `/finding` or
`/report` while a run streams used to let the background printer and the
full-screen picker fight over the terminal — pressing Ctrl-C to escape could
take the whole process down. Live output is now paused while any picker is
open (still captured in `/logs`) and restored when you exit, so browsing
findings mid-run is safe.
- Findings merge (dedup by title + endpoint) across the interrupted and
continued runs, and the merged report is rewritten to include everything.
---
## Previously in v3.6.2
- **Codex now streams live, tool-by-tool.** `codex exec` is driven with `--json`
and its JSONL event stream is parsed into the same categorized activity feed
as Claude Code: every shell command it runs (`exec:`), file edit (`edit:`),
MCP tool call (`tool:`), web search (`net:`) and token count appears the moment
it happens. A long, intense recon (subfinder → httpx → katana → nmap …) is no
longer a silent black box — you watch each tool execute.
- **`/logs` and `/status` now capture what each agent actually runs.** The
activity feed previously dropped the per-agent tool events; it now keeps the
actionable ones (commands, network, files, findings) and only filters long
model reasoning and token telemetry. `/logs` shows the real command trail;
`/status` `last:` shows a true sign-of-life.
- Failed internal commands surface as `exec: (exit N) <cmd>` instead of
silently vanishing, and Codex auth/rate errors are still detected from stderr.
---
## Previously in v3.6.1
- **Added the GPT-5.6 model line** (OpenAI / ChatGPT): `openai:gpt-5.6-sol`
(frontier / default), `openai:gpt-5.6-terra` (balanced), and
`openai:gpt-5.6-luna` (fast & affordable) — alongside the existing GPT-5.x,
Claude (incl. Sonnet 5), Grok 4.5 and the rest of the provider pool.
- Everything from v3.6.0 (AI/LLM/MCP/Skills testing, n8n audit, onboarding
wizard, intense multi-round recon) carries forward unchanged.
---
# NeuroSploit v3.6.0 — Release Notes # NeuroSploit v3.6.0 — Release Notes
**Release Date:** July 2026 **Release Date:** July 2026
+6 -4
View File
@@ -1,4 +1,4 @@
# NeuroSploit — Tutorial & User Guide (v3.6.0) # NeuroSploit — Tutorial & User Guide (v3.6.4)
A complete, hands-on guide to installing, configuring and running NeuroSploit — A complete, hands-on guide to installing, configuring and running NeuroSploit —
the autonomous, multi-model penetration-testing harness. the autonomous, multi-model penetration-testing harness.
@@ -98,7 +98,7 @@ Agents **degrade gracefully**: if `rustscan` is absent they use `nmap`; if neith
### Verify ### Verify
```bash ```bash
neurosploit --version # neurosploit 3.6.0 neurosploit --version # neurosploit 3.6.4
neurosploit agents # {"vulns":196,...,"chains":12,"total":417} neurosploit agents # {"vulns":196,...,"chains":12,"total":417}
neurosploit models # all providers & models neurosploit models # all providers & models
``` ```
@@ -522,8 +522,10 @@ NeuroSploit treats the target as **partially observable** (a POMDP):
entropy: when a node's belief is diffuse, recon is worth more than exploiting. entropy: when a node's belief is diffuse, recon is worth more than exploiting.
- **Anti-hallucination gate** (`may_assert`) — the agent may **not** claim - **Anti-hallucination gate** (`may_assert`) — the agent may **not** claim
exploitability while the belief is diffuse; it must observe more first. exploitability while the belief is diffuse; it must observe more first.
- **Grounding** — **no claim without a tool receipt**: empirical for black-box - **Grounding** — **no claim without a receipt**: *empirical* for black-box /
(real HTTP/OOB/error output), symbolic (`file:line`) for white-box. Ungrounded host / AI (real HTTP/OOB/error output), *symbolic* for white-box SAST & skills
audits (a `file:line` reference into the reviewed source — the code citation is
the receipt, no live target needed), and *either* for grey-box. Ungrounded
claims are demoted and flagged. claims are demoted and flagged.
- **Chaining** — confirmed findings are chained into deeper impact, each stage - **Chaining** — confirmed findings are chained into deeper impact, each stage
proven before advancing. proven before advancing.
+2 -2
View File
@@ -14,7 +14,7 @@ function Ok ($m) { Write-Host " + $m" -ForegroundColor Green }
function Warn($m){ Write-Host " ! $m" -ForegroundColor Yellow } function Warn($m){ Write-Host " ! $m" -ForegroundColor Yellow }
Write-Host "" Write-Host ""
Write-Host " NeuroSploit installer (Windows) — v3.6.0" -ForegroundColor Cyan Write-Host " NeuroSploit installer (Windows) — v3.6.1" -ForegroundColor Cyan
# arch → asset arch (only x64 prebuilt today; arm64 falls back to source) # arch → asset arch (only x64 prebuilt today; arm64 falls back to source)
$rawArch = $env:PROCESSOR_ARCHITECTURE $rawArch = $env:PROCESSOR_ARCHITECTURE
@@ -29,7 +29,7 @@ $ref = $env:NEUROSPLOIT_REF
if (-not $ref) { if (-not $ref) {
try { $ref = (Invoke-RestMethod "https://api.github.com/repos/$slug/releases/latest").tag_name } catch { } try { $ref = (Invoke-RestMethod "https://api.github.com/repos/$slug/releases/latest").tag_name } catch { }
} }
if (-not $ref) { $ref = "v3.6.0" } if (-not $ref) { $ref = "v3.6.1" }
Say "Release: $ref" Say "Release: $ref"
New-Item -ItemType Directory -Force -Path $dir | Out-Null New-Item -ItemType Directory -Force -Path $dir | Out-Null
+2 -2
View File
@@ -871,7 +871,7 @@ dependencies = [
[[package]] [[package]]
name = "neurosploit" name = "neurosploit"
version = "3.6.0" version = "3.6.4"
dependencies = [ dependencies = [
"anyhow", "anyhow",
"clap", "clap",
@@ -888,7 +888,7 @@ dependencies = [
[[package]] [[package]]
name = "neurosploit-harness" name = "neurosploit-harness"
version = "3.6.0" version = "3.6.4"
dependencies = [ dependencies = [
"anyhow", "anyhow",
"futures", "futures",
+1 -1
View File
@@ -3,7 +3,7 @@ members = ["crates/harness", "app"]
resolver = "2" resolver = "2"
[workspace.package] [workspace.package]
version = "3.6.0" version = "3.6.4"
edition = "2021" edition = "2021"
license = "MIT" license = "MIT"
repository = "https://github.com/JoasASantos/NeuroSploit" repository = "https://github.com/JoasASantos/NeuroSploit"
+12 -6
View File
@@ -1,4 +1,4 @@
//! NeuroSploit v3.6.0 — interactive harness + CLI (`run` / `whitebox` / `agents` / `models`). //! NeuroSploit v3.6.4 — interactive harness + CLI (`run` / `whitebox` / `agents` / `models`).
mod repl; mod repl;
mod tui; mod tui;
@@ -11,8 +11,8 @@ use std::path::{Path, PathBuf};
#[command( #[command(
name = "neurosploit", name = "neurosploit",
version, version,
about = "NeuroSploit v3.6.0 — multi-model autonomous pentest harness", about = "NeuroSploit v3.6.4 — multi-model autonomous pentest harness",
long_about = "NeuroSploit v3.6.0 — a Rust multi-model harness that drives a pool of LLMs \ long_about = "NeuroSploit v3.6.4 — a Rust multi-model harness that drives a pool of LLMs \
(API key or local subscription: Claude/Codex/Gemini/Grok) to autonomously test a target. \ (API key or local subscription: Claude/Codex/Gemini/Grok) to autonomously test a target. \
After recon it INTELLIGENTLY selects only the agents matching the discovered surface, runs \ After recon it INTELLIGENTLY selects only the agents matching the discovered surface, runs \
them in parallel, then validates every finding by cross-model voting before reporting.\n\n\ them in parallel, then validates every finding by cross-model voting before reporting.\n\n\
@@ -721,7 +721,7 @@ pub(crate) fn spawn_engagement(base: &Path, mut cfg: RunConfig, mcp: bool, mode:
println!(" │ ua : {ua}"); println!(" │ ua : {ua}");
write_status(&workdir, "running", &format!("\"target\":{:?}", cfg.target)); write_status(&workdir, "running", &format!("\"target\":{:?}", cfg.target));
println!(" ┌─ NeuroSploit v3.6.0 · by Joas A Santos & Red Team Leaders"); println!(" ┌─ NeuroSploit v3.6.4 · by Joas A Santos & Red Team Leaders");
println!(" │ run id : {run_id}"); println!(" │ run id : {run_id}");
println!(" │ target : {}", cfg.target); println!(" │ target : {}", cfg.target);
println!(" │ models : {}", cfg.models.join(", ")); println!(" │ models : {}", cfg.models.join(", "));
@@ -1076,10 +1076,16 @@ pub(crate) fn render_compact(raw: &str) -> Option<String> {
"ai" => return None, // skip verbose model chatter in background feed "ai" => return None, // skip verbose model chatter in background feed
_ => { _ => {
let low = line.to_lowercase(); let low = line.to_lowercase();
if low.contains("recon complete") { "\x1b[36m 🔍 recon complete\x1b[0m".into() } // Recon / probe activity — SHOW it so a long recon (esp. via a
// non-streaming CLI like codex) doesn't look frozen.
if low.starts_with("probe:") { format!("\x1b[36m 🔎 {}\x1b[0m", trunc1(line, 130)) }
else if low.contains("recon complete") { "\x1b[36m 🔍 recon complete\x1b[0m".into() }
else if low.starts_with("recon") || low.starts_with("ai-recon") || low.contains("recon round") || low.contains("intensity") { format!("\x1b[36m 🔍 {}\x1b[0m", trunc1(line, 130)) }
else if low.starts_with("skills audit") || low.starts_with("ai engagement") { format!("\x1b[36m 🤖 {}\x1b[0m", trunc1(line, 130)) }
else if low.starts_with("loaded ") || low.starts_with("running ") { format!("\x1b[36m 🧭 {}\x1b[0m", trunc1(line, 130)) }
else if low.contains("selected") && low.contains("agent") { format!("\x1b[36m 🧭 {}\x1b[0m", trunc1(line, 110)) } else if low.contains("selected") && low.contains("agent") { format!("\x1b[36m 🧭 {}\x1b[0m", trunc1(line, 110)) }
else if low.starts_with("vote") && low.contains("confirmed") { format!("\x1b[1;32m ✓ {}\x1b[0m", trunc1(line, 110)) } else if low.starts_with("vote") && low.contains("confirmed") { format!("\x1b[1;32m ✓ {}\x1b[0m", trunc1(line, 110)) }
else if low.starts_with("exploit") || low.starts_with("test ") || low.contains("launching agent") { format!("\x1b[35m 🧪 {}\x1b[0m", trunc1(line, 110)) } else if low.starts_with("exploit") || low.starts_with("test ") || low.starts_with("ai ") || low.starts_with("skill ") || low.contains("launching agent") { format!("\x1b[35m 🧪 {}\x1b[0m", trunc1(line, 110)) }
else if low.starts_with("vote") { format!("\x1b[2m · {}\x1b[0m", trunc1(line, 110)) } else if low.starts_with("vote") { format!("\x1b[2m · {}\x1b[0m", trunc1(line, 110)) }
else if low.contains("fail") || low.contains("error") { format!("\x1b[31m ✗ {}\x1b[0m", trunc1(line, 110)) } else if low.contains("fail") || low.contains("error") { format!("\x1b[31m ✗ {}\x1b[0m", trunc1(line, 110)) }
else { return None; } else { return None; }
+161 -37
View File
@@ -1,4 +1,4 @@
//! NeuroSploit v3.6.0 — interactive session (Claude-Code / Codex / Cursor-CLI style). //! NeuroSploit v3.6.4 — interactive session (Claude-Code / Codex / Cursor-CLI style).
//! //!
//! Launched when `neurosploit` runs with no subcommand. A persistent REPL with //! Launched when `neurosploit` runs with no subcommand. A persistent REPL with
//! real line editing (arrow-key history recall, Ctrl-A/E/K, paste), model //! real line editing (arrow-key history recall, Ctrl-A/E/K, paste), model
@@ -33,6 +33,9 @@ struct RunLive {
commands: Vec<String>, // full untruncated commands for /expand & Ctrl+O commands: Vec<String>, // full untruncated commands for /expand & Ctrl+O
agents: usize, agents: usize,
agents_done: usize, agents_done: usize,
last: String, // last meaningful activity line (sign of life)
lines: usize, // total streamed lines (activity counter)
feed: Vec<String>, // recent raw activity lines for /logs (capped)
} }
impl RunLive { impl RunLive {
/// progress fraction in [0,1] (agents completed / total selected). /// progress fraction in [0,1] (agents completed / total selected).
@@ -48,9 +51,25 @@ impl RunLive {
} }
fn ingest(&mut self, line: &str) { fn ingest(&mut self, line: &str) {
let low = line.to_lowercase(); let low = line.to_lowercase();
self.lines += 1;
// Keep a compact activity trail for /logs and the /status sign-of-life.
// Streamed agent events are tagged "@label <event>": keep the actionable
// ones (commands, net, tools, file edits, phases) so the operator sees
// exactly what each agent is running — drop only long model reasoning
// (ai:), token telemetry (tokens:), and machine JSON (finding_json:).
let payload = line.strip_prefix('@')
.and_then(|r| r.split_once(' ').map(|(_, rest)| rest))
.unwrap_or(line);
let plow = payload.to_lowercase();
if !low.starts_with("finding_json:") && !plow.starts_with("ai:") && !plow.starts_with("tokens:") {
let clean: String = line.chars().take(160).collect();
self.last = clean.clone();
self.feed.push(clean);
if self.feed.len() > 200 { self.feed.remove(0); }
}
if low.contains("token/quota exhausted") || low.contains("run is paused") { self.phase = "paused (quota)".into(); } if low.contains("token/quota exhausted") || low.contains("run is paused") { self.phase = "paused (quota)".into(); }
else if low.contains("resumed — retrying") { self.phase = "exploiting".into(); } else if low.contains("resumed — retrying") { self.phase = "exploiting".into(); }
else if low.contains("recon complete") { self.phase = "recon".into(); } else if low.starts_with("recon") || low.starts_with("ai-recon") || low.contains("recon round") || low.contains("intensity") || low.starts_with("probe:") { self.phase = "recon".into(); }
else if low.contains("selected") && low.contains("agent") { else if low.contains("selected") && low.contains("agent") {
self.phase = "planning".into(); self.phase = "planning".into();
if let Some(n) = line.split_whitespace().find_map(|t| t.parse::<usize>().ok()) { self.agents = n; } if let Some(n) = line.split_whitespace().find_map(|t| t.parse::<usize>().ok()) { self.agents = n; }
@@ -101,6 +120,10 @@ struct ActiveRun {
resume: Arc<tokio::sync::Notify>, resume: Arc<tokio::sync::Notify>,
/// Fallback models to try first, pushed by /continue <provider:model>. /// Fallback models to try first, pushed by /continue <provider:model>.
fallback: Arc<Mutex<Vec<ModelRef>>>, fallback: Arc<Mutex<Vec<ModelRef>>>,
/// Suppress live background printing while a full-screen picker (dialoguer)
/// is open, so the two don't fight over the terminal and corrupt it. The
/// stream is still ingested (feed/checkpoint), just not printed meanwhile.
quiet: Arc<AtomicBool>,
} }
/// On-disk checkpoint of an in-flight run's findings/commands, written live so a /// On-disk checkpoint of an in-flight run's findings/commands, written live so a
@@ -120,7 +143,7 @@ const COMMANDS: &[&str] = &[
"/help", "/onboard", "/show", "/config", "/providers", "/model", "/key", "/sub", "/target", "/help", "/onboard", "/show", "/config", "/providers", "/model", "/key", "/sub", "/target",
"/repo", "/auth", "/creds", "/focus", "/attach", "/context", "/mcp", "/offline", "/repo", "/auth", "/creds", "/focus", "/attach", "/context", "/mcp", "/offline",
"/votes", "/chain", "/recon", "/timeout", "/proxy", "/burp", "/ua", "/agents", "/theme", "/clear", "/run", "/stop", "/continue", "/runs", "/results", "/report", "/votes", "/chain", "/recon", "/timeout", "/proxy", "/burp", "/ua", "/agents", "/theme", "/clear", "/run", "/stop", "/continue", "/runs", "/results", "/report",
"/status", "/diff", "/retest", "/validate", "/finding", "/expand", "/integrations", "/quit", "/status", "/logs", "/diff", "/retest", "/validate", "/finding", "/expand", "/integrations", "/quit",
]; ];
/// rustyline helper: Tab-completes `/commands` and `@filesystem-paths`, /// rustyline helper: Tab-completes `/commands` and `@filesystem-paths`,
@@ -334,7 +357,7 @@ pub async fn repl(base: &Path) -> anyhow::Result<()> {
let backends = harness::installed_cli_backends(); let backends = harness::installed_cli_backends();
println!("\x1b[1m"); println!("\x1b[1m");
println!(" ███╗ ██╗███████╗██╗ ██╗██████╗ ██████╗"); println!(" ███╗ ██╗███████╗██╗ ██╗██████╗ ██████╗");
println!(" ████╗ ██║██╔════╝██║ ██║██╔══██╗██╔═══██╗ NeuroSploit v3.6.0"); println!(" ████╗ ██║██╔════╝██║ ██║██╔══██╗██╔═══██╗ NeuroSploit v3.6.4");
println!(" ██╔██╗ ██║█████╗ ██║ ██║██████╔╝██║ ██║ interactive harness"); println!(" ██╔██╗ ██║█████╗ ██║ ██║██████╔╝██║ ██║ interactive harness");
println!(" ██║╚██╗██║██╔══╝ ██║ ██║██╔══██╗██║ ██║ by Joas A Santos"); println!(" ██║╚██╗██║██╔══╝ ██║ ██║██╔══██╗██║ ██║ by Joas A Santos");
println!(" ██║ ╚████║███████╗╚██████╔╝██║ ██║╚██████╔╝ & Red Team Leaders"); println!(" ██║ ╚████║███████╗╚██████╔╝██║ ██║╚██████╔╝ & Red Team Leaders");
@@ -351,6 +374,9 @@ pub async fn repl(base: &Path) -> anyhow::Result<()> {
if resumed || past > 0 { if resumed || past > 0 {
println!(" ↻ resumed project session from {}{} past run(s)", proj_dir().display(), past); println!(" ↻ resumed project session from {}{} past run(s)", proj_dir().display(), past);
} }
// A recovered interrupted run, carried in memory so `/continue` can relaunch
// the engagement on the same target with these findings folded forward.
let mut resumable: Option<(String, Vec<Finding>)> = None;
// Recover an interrupted run (REPL was quit/crashed mid-engagement): its // Recover an interrupted run (REPL was quit/crashed mid-engagement): its
// live findings were checkpointed to disk — fold them into /runs so // live findings were checkpointed to disk — fold them into /runs so
// /results, /finding and /report still work. // /results, /finding and /report still work.
@@ -365,6 +391,8 @@ pub async fn repl(base: &Path) -> anyhow::Result<()> {
save_runs(base, &h); save_runs(base, &h);
println!(" \x1b[1;33m↻ recovered interrupted run on {}{} finding(s) saved as run #{}\x1b[0m (/results {id} · /report {id})", println!(" \x1b[1;33m↻ recovered interrupted run on {}{} finding(s) saved as run #{}\x1b[0m (/results {id} · /report {id})",
cp.target, cp.findings.len(), id); cp.target, cp.findings.len(), id);
println!(" \x1b[36m ↳ /continue to keep testing this target — the {} finding(s) carry forward\x1b[0m", cp.findings.len());
resumable = Some((cp.target.clone(), cp.findings.clone()));
} }
clear_checkpoint(); clear_checkpoint();
} }
@@ -383,7 +411,7 @@ pub async fn repl(base: &Path) -> anyhow::Result<()> {
if !queue.is_empty() && active.as_ref().map(|a| a.done.load(Ordering::Relaxed)).unwrap_or(true) { if !queue.is_empty() && active.as_ref().map(|a| a.done.load(Ordering::Relaxed)).unwrap_or(true) {
let next = queue.remove(0); let next = queue.remove(0);
println!("\n \x1b[1;35m▶ next target\x1b[0m ({} left): {next}", queue.len()); println!("\n \x1b[1;35m▶ next target\x1b[0m ({} left): {next}", queue.len());
active = start_background(base, &s, &mut reader, history.clone(), Some(&next)).await; active = start_background(base, &s, &mut reader, history.clone(), Some(&next), vec![]).await;
} }
println!("{}", context_prompt(&s)); // dim context line above the prompt println!("{}", context_prompt(&s)); // dim context line above the prompt
let Some(line) = reader.read(PROMPT) else { println!("\n bye."); break }; let Some(line) = reader.read(PROMPT) else { println!("\n bye."); break };
@@ -591,6 +619,7 @@ pub async fn repl(base: &Path) -> anyhow::Result<()> {
if active.as_ref().map(|a| !a.done.load(Ordering::Relaxed)).unwrap_or(false) { if active.as_ref().map(|a| !a.done.load(Ordering::Relaxed)).unwrap_or(false) {
println!(" a run is already active — /status to check, /stop to halt it."); println!(" a run is already active — /status to check, /stop to halt it.");
} else { } else {
resumable = None; // a fresh /run supersedes any recovered interrupted run
save_session(&s); save_session(&s);
// Multiple comma-separated targets → run sequentially (queue the rest). // Multiple comma-separated targets → run sequentially (queue the rest).
let targets = session_targets(&s); let targets = session_targets(&s);
@@ -601,7 +630,7 @@ pub async fn repl(base: &Path) -> anyhow::Result<()> {
if !queue.is_empty() { if !queue.is_empty() {
println!(" \x1b[1;35m▶ multi-target\x1b[0m: {} URLs — running sequentially", targets.len()); println!(" \x1b[1;35m▶ multi-target\x1b[0m: {} URLs — running sequentially", targets.len());
} }
match start_background(base, &s, &mut reader, history.clone(), first.as_deref()).await { match start_background(base, &s, &mut reader, history.clone(), first.as_deref(), vec![]).await {
Some(a) => { active = Some(a); println!(" \x1b[1;35m▶ running in background\x1b[0m — keep typing · \x1b[36m/status\x1b[0m · \x1b[36m/stop\x1b[0m"); } Some(a) => { active = Some(a); println!(" \x1b[1;35m▶ running in background\x1b[0m — keep typing · \x1b[36m/status\x1b[0m · \x1b[36m/stop\x1b[0m"); }
None => { // no external printer (piped) → blocking fallback None => { // no external printer (piped) → blocking fallback
let mut h = history.lock().unwrap(); let mut h = history.lock().unwrap();
@@ -632,20 +661,46 @@ pub async fn repl(base: &Path) -> anyhow::Result<()> {
} }
} }
"/continue" | "/resume" => { "/continue" | "/resume" => {
match &active { let paused = active.as_ref().map(|a| a.paused.load(Ordering::Relaxed)).unwrap_or(false);
Some(a) if a.paused.load(Ordering::Relaxed) => { let working = active.as_ref().map(|a| !a.done.load(Ordering::Relaxed)).unwrap_or(false);
if !arg.is_empty() { if paused {
let m = ModelRef::parse(arg); let a = active.as_ref().unwrap();
println!(" \x1b[1;35m▶ resuming with fallback model\x1b[0m {}:{}", m.provider, m.model); if !arg.is_empty() {
a.fallback.lock().unwrap().push(m); let m = ModelRef::parse(arg);
} else { println!(" \x1b[1;35m▶ resuming with fallback model\x1b[0m {}:{}", m.provider, m.model);
println!(" \x1b[1;35m▶ resuming\x1b[0m — retrying with the current model(s)."); a.fallback.lock().unwrap().push(m);
} } else {
a.paused.store(false, Ordering::Relaxed); println!(" \x1b[1;35m▶ resuming\x1b[0m — retrying with the current model(s).");
a.resume.notify_waiters();
} }
Some(a) if !a.done.load(Ordering::Relaxed) => println!(" run is not paused — it's still working. /status to check."), a.paused.store(false, Ordering::Relaxed);
_ => println!(" no paused run. (a run pauses automatically if your tokens/quota run out)"), a.resume.notify_waiters();
} else if working {
println!(" run is not paused — it's still working. /status to check.");
} else if let Some((tgt, prior)) = resumable.take() {
// Continue an interrupted run: relaunch on the same target, carry
// the prior findings forward, and steer agents to extend coverage
// rather than re-report what was already found.
if s.target.is_none() && s.repo.is_none() { s.target = Some(tgt.clone()); }
let titles: Vec<String> = prior.iter().map(|f| format!("[{}] {}", f.severity, f.title)).collect();
let carry = format!(
"CONTINUE a prior interrupted engagement on this same target. These {} finding(s) are \
ALREADY confirmed — do NOT re-report them; instead widen coverage: chase untested \
endpoints/params/methods, try new agent classes, and chain from these where possible: {}",
prior.len(), titles.join("; "));
s.instructions = Some(match &s.instructions {
Some(prev) if !prev.trim().is_empty() => format!("{prev}\n\n{carry}"),
_ => carry,
});
println!(" \x1b[1;35m▶ continuing interrupted run\x1b[0m on {tgt}{} prior finding(s) carried forward", prior.len());
match start_background(base, &s, &mut reader, history.clone(), None, prior).await {
Some(a) => { active = Some(a); println!(" \x1b[1;35m▶ running in background\x1b[0m — keep typing · \x1b[36m/status\x1b[0m · \x1b[36m/stop\x1b[0m"); }
None => {
let mut h = history.lock().unwrap();
run(base, &s, &mut h).await; save_runs(base, &h);
}
}
} else {
println!(" no paused or interrupted run. (a run pauses on token/quota exhaustion; an interrupted run is offered for /continue at launch)");
} }
} }
"/runs" | "/history" => list_runs(&history.lock().unwrap()), "/runs" | "/history" => list_runs(&history.lock().unwrap()),
@@ -716,7 +771,12 @@ pub async fn repl(base: &Path) -> anyhow::Result<()> {
} }
} }
runs.extend(history.lock().unwrap().iter().rev().cloned()); // newest-first runs.extend(history.lock().unwrap().iter().rev().cloned()); // newest-first
// Silence live background output while the full-screen picker is
// open (they'd corrupt each other); restore + point to /logs after.
let live_now = active.as_ref().map(|a| { a.quiet.store(true, Ordering::Relaxed); !a.done.load(Ordering::Relaxed) }).unwrap_or(false);
browse_results(&runs); browse_results(&runs);
if let Some(a) = &active { a.quiet.store(false, Ordering::Relaxed); }
if live_now { println!(" \x1b[2m(run still streaming in background — /logs for what happened while browsing)\x1b[0m"); }
} }
} }
"/finding" | "/findings" => { "/finding" | "/findings" => {
@@ -725,7 +785,9 @@ pub async fn repl(base: &Path) -> anyhow::Result<()> {
Some(a) if arg.is_empty() && !a.done.load(Ordering::Relaxed) => a.live.lock().unwrap().full.clone(), Some(a) if arg.is_empty() && !a.done.load(Ordering::Relaxed) => a.live.lock().unwrap().full.clone(),
_ => { let h = history.lock().unwrap(); pick(&h, arg).map(|r| r.findings.clone()).unwrap_or_default() } _ => { let h = history.lock().unwrap(); pick(&h, arg).map(|r| r.findings.clone()).unwrap_or_default() }
}; };
if let Some(a) = &active { a.quiet.store(true, Ordering::Relaxed); }
finding_detail(&pool); finding_detail(&pool);
if let Some(a) = &active { a.quiet.store(false, Ordering::Relaxed); }
} }
"/expand" | "/full" => { "/expand" | "/full" => {
// Show full untruncated commands from the active run. // Show full untruncated commands from the active run.
@@ -743,7 +805,11 @@ pub async fn repl(base: &Path) -> anyhow::Result<()> {
None => println!(" no active run — /expand shows full commands while a run streams."), None => println!(" no active run — /expand shows full commands while a run streams."),
} }
} }
"/report" => open_report(&history.lock().unwrap(), arg), "/report" => {
if let Some(a) = &active { a.quiet.store(true, Ordering::Relaxed); }
open_report(&history.lock().unwrap(), arg);
if let Some(a) = &active { a.quiet.store(false, Ordering::Relaxed); }
}
"/status" => { "/status" => {
// Live status if a run is active, else a past run's status.json. // Live status if a run is active, else a past run's status.json.
match &active { match &active {
@@ -753,17 +819,39 @@ pub async fn repl(base: &Path) -> anyhow::Result<()> {
let mut by: std::collections::BTreeMap<&str, usize> = Default::default(); let mut by: std::collections::BTreeMap<&str, usize> = Default::default();
for (sv, _) in &l.findings { *by.entry(sv.as_str()).or_insert(0) += 1; } for (sv, _) in &l.findings { *by.entry(sv.as_str()).or_insert(0) += 1; }
let sev = if by.is_empty() { "0".into() } else { by.iter().map(|(k, v)| format!("{k}:{v}")).collect::<Vec<_>>().join(" ") }; let sev = if by.is_empty() { "0".into() } else { by.iter().map(|(k, v)| format!("{k}:{v}")).collect::<Vec<_>>().join(" ") };
println!(" \x1b[1m▶ live\x1b[0m {} ({}) · phase {} · {:02}:{:02} · {} possible finding(s) [{}]", println!(" \x1b[1m▶ live\x1b[0m {} ({}) · phase \x1b[36m{}\x1b[0m · {:02}:{:02} · {} finding(s) [{}]",
l.target, l.mode, l.phase, el / 60, el % 60, l.findings.len(), sev); l.target, l.mode, l.phase, el / 60, el % 60, l.full.len(), sev);
if a.paused.load(Ordering::Relaxed) { if a.paused.load(Ordering::Relaxed) {
println!(" \x1b[1;33m⏸ PAUSED — token/quota exhausted. /continue to resume, or /model <provider:model> then /continue to switch.\x1b[0m"); println!(" \x1b[1;33m⏸ PAUSED — token/quota exhausted. /continue to resume, or /model <provider:model> then /continue to switch.\x1b[0m");
} }
if l.agents > 0 { println!(" progress \x1b[36m{}\x1b[0m", l.bar(24)); } // Progress: a real bar once agents are selected; otherwise show the pre-exploit phase.
for (sv, t) in l.findings.iter().rev().take(5) { println!(" ✦ [{sv}] {t}"); } if l.agents > 0 { println!(" progress \x1b[36m{}\x1b[0m · {} cmd(s) · {} activity line(s)", l.bar(24), l.commands.len(), l.lines); }
else { println!(" \x1b[2m{} — no agents selected yet · {} cmd(s) · {} activity line(s)\x1b[0m", l.phase, l.commands.len(), l.lines); }
// Sign of life: the latest activity line (so a long recon isn't a black box).
if !l.last.is_empty() { println!(" \x1b[2mlast:\x1b[0m {}", trunc(&l.last, 116)); }
for x in l.full.iter().rev().take(5) { println!(" ✦ [{}] {} \x1b[2m({})\x1b[0m", x.severity, x.title, x.endpoint); }
println!(" \x1b[2m/logs — recent activity · /results — browse findings\x1b[0m");
} }
_ => run_status(&history.lock().unwrap(), arg), _ => run_status(&history.lock().unwrap(), arg),
} }
} }
"/logs" | "/log" | "/feed" => {
match &active {
Some(a) if !a.done.load(Ordering::Relaxed) => {
let n: usize = arg.trim().parse().unwrap_or(25);
let l = a.live.lock().unwrap();
if l.feed.is_empty() { println!(" (no activity yet — the run is starting/reconning)"); }
else {
println!(" ── recent activity (last {} of {} lines) ──", n.min(l.feed.len()), l.lines);
for line in l.feed.iter().rev().take(n).rev() {
if let Some(out) = crate::render_compact(line) { println!("{out}"); }
else { println!(" \x1b[2m{}\x1b[0m", trunc(line, 116)); }
}
}
}
_ => println!(" no active run — /logs shows the live activity feed while a run streams."),
}
}
"/quit" | "/exit" | "/q" => { "/quit" | "/exit" | "/q" => {
if active.as_ref().map(|a| !a.done.load(Ordering::Relaxed)).unwrap_or(false) { if active.as_ref().map(|a| !a.done.load(Ordering::Relaxed)).unwrap_or(false) {
if let Some(a) = &active { a.cancel.store(true, Ordering::Relaxed); } if let Some(a) = &active { a.cancel.store(true, Ordering::Relaxed); }
@@ -985,7 +1073,8 @@ async fn run(base: &Path, s: &Session, history: &mut Vec<RunRecord>) {
/// external printer while the REPL keeps accepting commands (/status, /stop). /// external printer while the REPL keeps accepting commands (/status, /stop).
/// Returns None when no external printer is available (piped) → caller blocks. /// Returns None when no external printer is available (piped) → caller blocks.
async fn start_background(base: &Path, s: &Session, reader: &mut Reader, async fn start_background(base: &Path, s: &Session, reader: &mut Reader,
history: Arc<Mutex<Vec<RunRecord>>>, target_override: Option<&str>) -> Option<ActiveRun> { history: Arc<Mutex<Vec<RunRecord>>>, target_override: Option<&str>,
seed: Vec<Finding>) -> Option<ActiveRun> {
// `target_override` runs one specific URL (used by the multi-target queue). // `target_override` runs one specific URL (used by the multi-target queue).
let ov = target_override.map(|t| t.to_string()); let ov = target_override.map(|t| t.to_string());
// The onboarding scope steers infra/cloud/ai/skills; otherwise web black/white/grey. // The onboarding scope steers infra/cloud/ai/skills; otherwise web black/white/grey.
@@ -1034,7 +1123,7 @@ async fn start_background(base: &Path, s: &Session, reader: &mut Reader,
let live = Arc::new(Mutex::new(RunLive { let live = Arc::new(Mutex::new(RunLive {
target: target.clone(), mode: mode_s, phase: "starting".into(), target: target.clone(), mode: mode_s, phase: "starting".into(),
started: Instant::now(), findings: vec![], full: vec![], commands: vec![], started: Instant::now(), findings: vec![], full: vec![], commands: vec![],
agents: 0, agents_done: 0, agents: 0, agents_done: 0, last: String::new(), lines: 0, feed: vec![],
})); }));
let cancel = sp.cancel.clone(); let cancel = sp.cancel.clone();
let soft = sp.soft.clone(); let soft = sp.soft.clone();
@@ -1043,17 +1132,20 @@ async fn start_background(base: &Path, s: &Session, reader: &mut Reader,
let fallback = sp.fallback.clone(); let fallback = sp.fallback.clone();
let done = Arc::new(AtomicBool::new(false)); let done = Arc::new(AtomicBool::new(false));
let choice = Arc::new(Mutex::new(StopMode::Run)); let choice = Arc::new(Mutex::new(StopMode::Run));
let quiet = Arc::new(AtomicBool::new(false));
let soft_task = soft.clone(); // idle guardrail triggers a soft-stop (validate) let soft_task = soft.clone(); // idle guardrail triggers a soft-stop (validate)
let cancel_task = cancel.clone(); let cancel_task = cancel.clone();
let quiet_task = quiet.clone();
let sub_mcp = s.subscription && mcp; // for the "browser/tools never engaged" diagnostic let sub_mcp = s.subscription && mcp; // for the "browser/tools never engaged" diagnostic
let (live2, done2, hist2, choice2) = (live.clone(), done.clone(), history, choice.clone()); let (live2, done2, hist2, choice2) = (live.clone(), done.clone(), history, choice.clone());
tokio::spawn(async move { tokio::spawn(async move {
let crate::Spawned { task, mut rx, workdir, .. } = sp; let crate::Spawned { task, mut rx, workdir, .. } = sp;
let mut last_saved = 0usize; let mut last_saved = 0usize;
let mut last_find = Instant::now(); // time of the last NEW finding let mut last_activity = Instant::now(); // last sign of PROGRESS (any activity)
let mut idle_fired = false; let mut idle_fired = false;
let mut tool_events = 0usize; // exec/net/read/browser activity seen let mut tool_events = 0usize; // exec/net/read/browser activity seen
let mut exploiting = false; // guardrail only arms once exploitation starts
let mut ticker = tokio::time::interval(std::time::Duration::from_secs(15)); let mut ticker = tokio::time::interval(std::time::Duration::from_secs(15));
ticker.set_missed_tick_behavior(tokio::time::MissedTickBehavior::Skip); ticker.set_missed_tick_behavior(tokio::time::MissedTickBehavior::Skip);
loop { loop {
@@ -1061,14 +1153,25 @@ async fn start_background(base: &Path, s: &Session, reader: &mut Reader,
maybe = rx.recv() => { maybe = rx.recv() => {
let Some(line) = maybe else { break }; let Some(line) = maybe else { break };
live2.lock().unwrap().ingest(&line); live2.lock().unwrap().ingest(&line);
let low = line.to_lowercase();
if line.contains("exec:") || line.contains("net:") || line.contains("read:") || line.contains("browser") { tool_events += 1; } if line.contains("exec:") || line.contains("net:") || line.contains("read:") || line.contains("browser") { tool_events += 1; }
if let Some(out) = crate::render_compact(&line) { let _ = printer.print(out); } // ANY streamed line is progress → reset the idle clock (a long
// Checkpoint on each new finding; also resets the idle clock. // recon or active tool use must NOT count as idle).
last_activity = Instant::now();
// Exploitation has begun once agents launch / vote — only then arm the guardrail.
if low.contains("launching agent") || low.starts_with("exploit ") || low.starts_with("test ")
|| low.starts_with("ai ") || low.starts_with("skill ") || low.starts_with("vote") { exploiting = true; }
// Don't print into the terminal while a full-screen picker is
// open (it would corrupt the picker); the line is still in the
// feed for /logs once the picker closes.
if !quiet_task.load(Ordering::Relaxed) {
if let Some(out) = crate::render_compact(&line) { let _ = printer.print(out); }
}
// Checkpoint on each new finding.
let snap = { let snap = {
let l = live2.lock().unwrap(); let l = live2.lock().unwrap();
if l.full.len() != last_saved { if l.full.len() != last_saved {
last_saved = l.full.len(); last_saved = l.full.len();
last_find = Instant::now();
Some(LiveCheckpoint { Some(LiveCheckpoint {
target: l.target.clone(), mode: l.mode.into(), phase: l.phase.clone(), target: l.target.clone(), mode: l.mode.into(), phase: l.phase.clone(),
workdir: workdir.display().to_string(), workdir: workdir.display().to_string(),
@@ -1079,15 +1182,16 @@ async fn start_background(base: &Path, s: &Session, reader: &mut Reader,
if let Some(c) = snap { save_checkpoint(&c); } if let Some(c) = snap { save_checkpoint(&c); }
} }
_ = ticker.tick() => { _ = ticker.tick() => {
// Idle guardrail: no NEW finding within the window → soft-stop // Idle guardrail: only after exploitation started AND no activity
// (stop launching exploit agents, validate what was found). // (not just no finding) within the window → soft-stop & validate.
if idle_secs > 0 && !idle_fired && last_find.elapsed().as_secs() >= idle_secs // Recon never trips it — it streams progress lines that reset the clock.
if idle_secs > 0 && !idle_fired && exploiting && last_activity.elapsed().as_secs() >= idle_secs
&& !soft_task.load(Ordering::Relaxed) && !cancel_task.load(Ordering::Relaxed) { && !soft_task.load(Ordering::Relaxed) && !cancel_task.load(Ordering::Relaxed) {
idle_fired = true; idle_fired = true;
*choice2.lock().unwrap() = StopMode::Validate; *choice2.lock().unwrap() = StopMode::Validate;
soft_task.store(true, Ordering::Relaxed); soft_task.store(true, Ordering::Relaxed);
let _ = printer.print(format!( let _ = printer.print(format!(
"\x1b[33m⏹ idle guardrail: no new finding in {} min — stopping & validating what was found\x1b[0m", "\x1b[33m⏹ idle guardrail: no activity in {} min — stopping & validating what was found\x1b[0m",
idle_secs / 60)); idle_secs / 60));
} }
} }
@@ -1112,7 +1216,7 @@ async fn start_background(base: &Path, s: &Session, reader: &mut Reader,
} }
// Raw → report from the unvalidated candidates we captured live. // Raw → report from the unvalidated candidates we captured live.
let (findings, validated_word) = if mode_choice == StopMode::Raw { let (mut findings, validated_word) = if mode_choice == StopMode::Raw {
let raw = live2.lock().unwrap().full.clone(); let raw = live2.lock().unwrap().full.clone();
crate::report_raw(&target, &raw, &workdir); crate::report_raw(&target, &raw, &workdir);
(raw, "unvalidated") (raw, "unvalidated")
@@ -1120,6 +1224,13 @@ async fn start_background(base: &Path, s: &Session, reader: &mut Reader,
let out = crate::finalize_run(task_out, &workdir); let out = crate::finalize_run(task_out, &workdir);
(out.findings, "validated") (out.findings, "validated")
}; };
// Continued run (/continue on an interrupted run): fold the carried-forward
// prior findings back in (dedup by title+endpoint) and rewrite the report
// so the merged run shows everything found across both sessions.
if !seed.is_empty() {
findings = merge_findings(seed.clone(), findings);
crate::report_raw(&target, &findings, &workdir);
}
let id = { let id = {
let mut h = hist2.lock().unwrap(); let mut h = hist2.lock().unwrap();
@@ -1135,7 +1246,19 @@ async fn start_background(base: &Path, s: &Session, reader: &mut Reader,
let _ = printer.print(format!("\x1b[36m report: {}\x1b[0m", crate::report_url(&workdir))); let _ = printer.print(format!("\x1b[36m report: {}\x1b[0m", crate::report_url(&workdir)));
done2.store(true, Ordering::Relaxed); done2.store(true, Ordering::Relaxed);
}); });
Some(ActiveRun { live, cancel, soft, done, choice, paused, resume, fallback }) Some(ActiveRun { live, cancel, soft, done, choice, paused, resume, fallback, quiet })
}
/// Merge two finding sets, deduping by (title, endpoint) — used to carry a prior
/// interrupted run's findings forward into a continued run without duplicating.
fn merge_findings(prior: Vec<Finding>, mut fresh: Vec<Finding>) -> Vec<Finding> {
use std::collections::HashSet;
let key = |f: &Finding| format!("{}|{}", f.title.trim().to_lowercase(), f.endpoint.trim().to_lowercase());
let seen: HashSet<String> = fresh.iter().map(key).collect();
for p in prior {
if !seen.contains(&key(&p)) { fresh.push(p); }
}
fresh
} }
/// Project-local store: `<cwd>/.neurosploit/` so each project keeps its own /// Project-local store: `<cwd>/.neurosploit/` so each project keeps its own
@@ -1526,8 +1649,9 @@ fn help() {
println!("\n \x1b[2mRUN & MONITOR\x1b[0m"); println!("\n \x1b[2mRUN & MONITOR\x1b[0m");
h("/run", "launch (runs in the BACKGROUND — keep typing)"); h("/run", "launch (runs in the BACKGROUND — keep typing)");
h("/status [n]", "live progress + findings while running (or a past run #)"); h("/status [n]", "live progress + findings while running (or a past run #)");
h("/logs [n]", "recent activity feed of the running test (recon/tools/findings)");
h("/stop", "stop: [1] validate+report [2] raw report now [3] discard"); h("/stop", "stop: [1] validate+report [2] raw report now [3] discard");
h("/continue", "resume a run paused on token/quota (change /model first to switch)"); h("/continue", "resume a paused (token/quota) OR a recovered interrupted run — carries findings forward");
h("/results [n]", "browse findings (target → vuln → detail; Esc = back)"); h("/results [n]", "browse findings (target → vuln → detail; Esc = back)");
h("/finding [n]", "pick a finding and see its command + PoC + evidence"); h("/finding [n]", "pick a finding and see its command + PoC + evidence");
h("/report [n]", "open a run's report (menu if several)"); h("/report [n]", "open a run's report (menu if several)");
+1 -1
View File
@@ -1,4 +1,4 @@
//! NeuroSploit v3.6.0 — TUI "Mission Control" mode. //! NeuroSploit v3.6.4 — TUI "Mission Control" mode.
//! //!
//! Concurrent panels that update live while the engagement runs in the //! Concurrent panels that update live while the engagement runs in the
//! background, with a composer input that stays active during execution: //! background, with a composer input that stays active during execution:
+1 -1
View File
@@ -1,4 +1,4 @@
//! POMDP belief-state world model (v3.6.0). //! POMDP belief-state world model (v3.6.4).
//! //!
//! The target is only partially observable, so we don't track booleans — we //! The target is only partially observable, so we don't track booleans — we
//! track a **belief**: a property graph whose nodes (host / service / vuln / //! track a **belief**: a property graph whose nodes (host / service / vuln /
+169 -39
View File
@@ -1,20 +1,36 @@
//! Verification / grounding engine (v3.6.0). //! Verification / grounding engine (v3.6.4).
//! //!
//! Hard rule: **no claim enters the world model without a tool receipt** — raw //! Hard rule: **no claim enters the world model without a receipt** — evidence,
//! tool output, not the LLM's paraphrase. This is the empirical anti-hallucination //! not the LLM's bare assertion. This is the anti-hallucination anchor that
//! anchor that complements the POMDP belief gate: //! complements the POMDP belief gate. What counts as a receipt depends on the
//! engagement, so grounding runs in one of three modes:
//! //!
//! - **Black-box**: grounding is empirical — the finding's evidence must look //! - **Empirical** (black-box / host / AI-endpoint): the finding's evidence must
//! like raw tool output (an HTTP response, an OOB callback, an error oracle), //! look like raw tool output (an HTTP response, an OOB callback, an error
//! not prose. //! oracle, a shell receipt) — not prose.
//! - **White-box**: grounding is symbolic — a file:line reference into the //! - **Symbolic** (white-box SAST / skills audit): the receipt is a `file:line`
//! reviewed source (reachability/taint), checked against the collected context. //! (or `file:section`) reference into the reviewed source, or a quote of code
//! that actually appears in it. There is NO live target to hit, so requiring an
//! HTTP-style receipt here is wrong — a code citation IS the receipt.
//! - **Either** (grey-box): both worlds are present (source review + a running
//! app), so a finding is grounded if it has a symbolic OR an empirical receipt.
//! //!
//! Ungrounded claims are flagged (`receipt_missing`) so the reward layer can //! Ungrounded claims are flagged (`receipt_missing`) so the reward layer can
//! penalize them (the "claim without receipt" term). //! penalize them (the "claim without receipt" term).
use crate::types::Finding; use crate::types::Finding;
/// How a finding must be grounded, per engagement type.
#[derive(Debug, Clone, Copy, PartialEq, Eq)]
pub enum GroundMode {
/// Black-box / host / AI endpoint: evidence must resemble raw tool output.
Empirical,
/// White-box SAST / skills audit: evidence must reference the reviewed source.
Symbolic,
/// Grey-box: accept either a source citation or an empirical receipt.
Either,
}
/// Verdict of grounding a single finding. /// Verdict of grounding a single finding.
pub struct Grounded { pub struct Grounded {
pub ok: bool, pub ok: bool,
@@ -35,47 +51,91 @@ fn looks_empirical(evidence: &str) -> bool {
} }
/// White-box: evidence should reference a source location present in `context`. /// White-box: evidence should reference a source location present in `context`.
/// `context` is the reviewed SOURCE (not the model transcript). When the source
/// context is unavailable, fall back to structural checks so a well-formed
/// `file:line` + code quote still grounds (a SAST finding must never be silently
/// dropped just because the caller couldn't supply the corpus).
fn looks_symbolic(f: &Finding, context: &str) -> bool { fn looks_symbolic(f: &Finding, context: &str) -> bool {
// endpoint like file.ext:line, and the file appears in the reviewed source. let loc = f.endpoint.trim();
let loc = &f.endpoint; // A file:line / file:section reference is the canonical symbolic receipt.
if let Some((file, _)) = loc.rsplit_once(':') { let has_file_ref = loc.rsplit_once(':')
let base = file.rsplit('/').next().unwrap_or(file); .map(|(file, tail)| {
if !base.is_empty() && context.contains(base) { let base = file.rsplit(['/', '\\']).next().unwrap_or(file);
return true; // looks like a path/file (has an extension or a separator) and a
// line/section follows — i.e. not a "host:port" style endpoint.
!base.is_empty()
&& (base.contains('.') || file.contains('/'))
&& !tail.trim().is_empty()
})
.unwrap_or(false);
if !context.is_empty() {
// Strongest: the referenced file actually appears in the reviewed source.
if let Some((file, _)) = loc.rsplit_once(':') {
let base = file.rsplit(['/', '\\']).next().unwrap_or(file);
if !base.is_empty() && context.contains(base) {
return true;
}
} }
} // Or the evidence quotes a distinctive code token present in the source.
// or the evidence quotes code that is actually in the context let quote_matches = f.evidence
!f.evidence.trim().is_empty()
&& f.evidence.split_whitespace().take(6).collect::<Vec<_>>().join(" ")
.split_whitespace() .split_whitespace()
.filter(|t| t.len() > 4 && context.contains(*t)) .filter(|t| t.len() > 4 && context.contains(*t))
.count() .count();
>= 2 if quote_matches >= 2 {
} return true;
/// Ground a finding. `context` is the reviewed source for white-box (empty for
/// black-box). Returns whether it has a valid receipt and of what kind.
pub fn ground(f: &Finding, context: &str, whitebox: bool) -> Grounded {
if whitebox && !context.is_empty() {
if looks_symbolic(f, context) {
return Grounded { ok: true, kind: "symbolic", reason: "source location/quote matches reviewed code".into() };
} }
return Grounded { ok: false, kind: "missing", reason: "no source reference into reviewed code".into() }; // Source is present but neither the file nor a quote matched → still
// accept a well-formed file:line ref with quoted evidence, since the
// bounded corpus may simply not include the referenced file.
return has_file_ref && f.evidence.trim().len() >= 12;
} }
if looks_empirical(&f.evidence) {
Grounded { ok: true, kind: "empirical", reason: "evidence resembles raw tool output".into() } // No source corpus available: ground on a well-formed file:line reference
} else { // backed by non-trivial quoted evidence.
Grounded { ok: false, kind: "missing", reason: "evidence is paraphrase, not a tool receipt".into() } has_file_ref && f.evidence.trim().len() >= 12
}
/// Ground a finding under `mode`. `context` is the reviewed SOURCE for symbolic/
/// either modes (empty for pure empirical). Returns whether it has a valid
/// receipt and of what kind.
pub fn ground(f: &Finding, context: &str, mode: GroundMode) -> Grounded {
let symbolic = || looks_symbolic(f, context);
let empirical = || looks_empirical(&f.evidence);
match mode {
GroundMode::Symbolic => {
if symbolic() {
Grounded { ok: true, kind: "symbolic", reason: "source location/quote matches reviewed code".into() }
} else {
Grounded { ok: false, kind: "missing", reason: "no source reference (file:line) into reviewed code".into() }
}
}
GroundMode::Either => {
if symbolic() {
Grounded { ok: true, kind: "symbolic", reason: "source location/quote matches reviewed code".into() }
} else if empirical() {
Grounded { ok: true, kind: "empirical", reason: "evidence resembles raw tool output".into() }
} else {
Grounded { ok: false, kind: "missing", reason: "no source reference nor tool receipt".into() }
}
}
GroundMode::Empirical => {
if empirical() {
Grounded { ok: true, kind: "empirical", reason: "evidence resembles raw tool output".into() }
} else {
Grounded { ok: false, kind: "missing", reason: "evidence is paraphrase, not a tool receipt".into() }
}
}
} }
} }
/// Apply the grounding gate to a finding set. Ungrounded findings are flagged /// Apply the grounding gate to a finding set under `mode`. Ungrounded findings
/// (receipt recorded in `votes`) and demoted to unvalidated so they never get /// are flagged (receipt recorded in `votes`) and demoted to unvalidated so they
/// reported as confirmed. Returns (kept, demoted_count). /// never get reported as confirmed. Returns (kept, demoted_count).
pub fn gate(mut findings: Vec<Finding>, context: &str, whitebox: bool) -> (Vec<Finding>, usize) { pub fn gate(mut findings: Vec<Finding>, context: &str, mode: GroundMode) -> (Vec<Finding>, usize) {
let mut demoted = 0; let mut demoted = 0;
for f in findings.iter_mut() { for f in findings.iter_mut() {
let g = ground(f, context, whitebox); let g = ground(f, context, mode);
if !g.ok { if !g.ok {
f.validated = false; f.validated = false;
f.votes = format!("{} · receipt_missing", f.votes); f.votes = format!("{} · receipt_missing", f.votes);
@@ -85,3 +145,73 @@ pub fn gate(mut findings: Vec<Finding>, context: &str, whitebox: bool) -> (Vec<F
findings.retain(|f| f.validated); findings.retain(|f| f.validated);
(findings, demoted) (findings, demoted)
} }
#[cfg(test)]
mod tests {
use super::*;
fn sast_finding() -> Finding {
// A typical SAST finding: file:line endpoint + a code quote as evidence,
// and NO HTTP/tool-output markers (there is no live target to hit).
Finding {
title: "SQL injection via string-formatted query".into(),
severity: "High".into(),
cwe: "CWE-89".into(),
endpoint: "src/db/users.py:42".into(),
evidence: "query = \"SELECT * FROM users WHERE id = \" + request.args.get('id')".into(),
validated: true,
confidence: 0.8,
..Default::default()
}
}
#[test]
fn sast_finding_grounds_symbolically_against_source() {
let src = "def get(id):\n query = \"SELECT * FROM users WHERE id = \" + request.args.get('id')\n";
assert!(ground(&sast_finding(), src, GroundMode::Symbolic).ok,
"a file:line SAST finding whose code appears in the source must ground");
}
#[test]
fn sast_finding_grounds_even_without_source_corpus() {
// Regression: the whitebox gate used to run in EMPIRICAL mode (bug #33),
// demoting every SAST finding because code quotes lack HTTP-style markers.
// A well-formed file:line + quoted evidence must ground on its own.
assert!(ground(&sast_finding(), "", GroundMode::Symbolic).ok,
"SAST finding must not be demoted for lacking a tool receipt");
}
#[test]
fn symbolic_rejects_bare_prose() {
let f = Finding { endpoint: "the login flow".into(),
evidence: "The application seems insecure.".into(), validated: true, ..Default::default() };
assert!(!ground(&f, "", GroundMode::Symbolic).ok,
"prose with no source reference must NOT ground symbolically");
}
#[test]
fn empirical_still_requires_tool_output() {
// Black-box unchanged: a code quote is not an empirical receipt.
assert!(!ground(&sast_finding(), "", GroundMode::Empirical).ok);
let http = Finding {
endpoint: "https://t/login".into(),
evidence: "HTTP/1.1 200 OK\nset-cookie: sid=1; \nserver: nginx\n<script>alert(1)</script>".into(),
validated: true, ..Default::default() };
assert!(ground(&http, "", GroundMode::Empirical).ok);
}
#[test]
fn either_accepts_symbolic_or_empirical() {
assert!(ground(&sast_finding(), "", GroundMode::Either).ok, "grey-box accepts a source citation");
}
#[test]
fn gate_keeps_grounded_and_demotes_prose() {
let good = sast_finding();
let bad = Finding { title: "vibes".into(), endpoint: "somewhere".into(),
evidence: "looks bad".into(), validated: true, ..Default::default() };
let (kept, demoted) = gate(vec![good, bad], "", GroundMode::Symbolic);
assert_eq!(kept.len(), 1);
assert_eq!(demoted, 1);
}
}
+1 -1
View File
@@ -1,4 +1,4 @@
//! NeuroSploit v3.6.0 harness — a robust multi-model runtime for the //! NeuroSploit v3.6.4 harness — a robust multi-model runtime for the
//! markdown-driven autonomous pentest engine. //! markdown-driven autonomous pentest engine.
//! //!
//! The harness loads the `agents_md/` library, drives a *pool* of LLM models //! The harness loads the `agents_md/` library, drives a *pool* of LLM models
+150 -1
View File
@@ -25,7 +25,7 @@ pub fn providers() -> Vec<Provider> {
Provider { key: "anthropic", label: "Anthropic Claude", base_url: "https://api.anthropic.com/v1", env_key: "ANTHROPIC_API_KEY", kind: "cli", Provider { key: "anthropic", label: "Anthropic Claude", base_url: "https://api.anthropic.com/v1", env_key: "ANTHROPIC_API_KEY", kind: "cli",
models: vec!["claude-opus-4-8", "claude-sonnet-5", "claude-sonnet-4-6", "claude-haiku-4-5"] }, models: vec!["claude-opus-4-8", "claude-sonnet-5", "claude-sonnet-4-6", "claude-haiku-4-5"] },
Provider { key: "openai", label: "OpenAI (ChatGPT)", base_url: "https://api.openai.com/v1", env_key: "OPENAI_API_KEY", kind: "cli", Provider { key: "openai", label: "OpenAI (ChatGPT)", base_url: "https://api.openai.com/v1", env_key: "OPENAI_API_KEY", kind: "cli",
models: vec!["gpt-5.5", "gpt-5.4", "gpt-5.4-mini", "gpt-5.3-codex", "gpt-5.2", "gpt-5.1", "gpt-5.1-codex", "o4"] }, models: vec!["gpt-5.6-sol", "gpt-5.6-terra", "gpt-5.6-luna", "gpt-5.5", "gpt-5.4", "gpt-5.4-mini", "gpt-5.3-codex", "gpt-5.2", "gpt-5.1", "gpt-5.1-codex", "o4"] },
Provider { key: "xai", label: "xAI Grok", base_url: "https://api.x.ai/v1", env_key: "XAI_API_KEY", kind: "cli", Provider { key: "xai", label: "xAI Grok", base_url: "https://api.x.ai/v1", env_key: "XAI_API_KEY", kind: "cli",
models: vec!["grok-4.5", "grok-4", "grok-4-fast"] }, models: vec!["grok-4.5", "grok-4", "grok-4-fast"] },
Provider { key: "gemini", label: "Google Gemini", base_url: "https://generativelanguage.googleapis.com/v1beta/openai", env_key: "GEMINI_API_KEY", kind: "cli", Provider { key: "gemini", label: "Google Gemini", base_url: "https://generativelanguage.googleapis.com/v1beta/openai", env_key: "GEMINI_API_KEY", kind: "cli",
@@ -194,6 +194,12 @@ impl ChatClient {
if bin == "claude" { if bin == "claude" {
return self.chat_claude_stream(label, model, &prompt, mcp_config, progress).await; return self.chat_claude_stream(label, model, &prompt, mcp_config, progress).await;
} }
// Codex exec streams JSONL events (`--json`): commands it runs, agent
// messages, file changes, token usage. Surface them live so recon and
// exploitation are visible tool-by-tool instead of a silent black box.
if bin == "codex" {
return self.chat_codex_stream(label, model, &prompt, mcp_config, progress).await;
}
let mut cmd = Command::new(bin); let mut cmd = Command::new(bin);
match bin { match bin {
@@ -245,6 +251,19 @@ impl ChatClient {
} else { } else {
"no output".to_string() "no output".to_string()
}; };
// Agentic CLIs (esp. `codex exec` in bypass-sandbox mode) exit
// non-zero when a tool/command they ran internally returned non-zero
// (e.g. a curl/nmap that failed) — even though they produced a valid
// final answer. Treat that as success and use the output; only fail
// hard on a genuine auth/rate/quota error or when there's no output.
let low = format!("{stdout}\n{stderr}").to_lowercase();
let hard = ["not logged in", "please log in", "please login", "run /login",
"unauthorized", "not authenticated", "invalid api key", "no api key",
"rate limit", "429", "quota", "credit balance", "usage limit"]
.iter().any(|k| low.contains(k));
if !stdout.is_empty() && !hard {
return Ok(stdout);
}
return Err(anyhow!( return Err(anyhow!(
"{} subscription CLI exit {}: {}", "{} subscription CLI exit {}: {}",
bin, bin,
@@ -353,6 +372,136 @@ impl ChatClient {
} }
Ok(result) Ok(result)
} }
/// Drive `codex exec --json` and surface its JSONL event stream as a live,
/// categorized activity feed (commands, agent messages, file changes, MCP
/// tool calls, token usage). The final agent message is returned as the
/// result. Mirrors `chat_claude_stream` so Codex runs are just as visible.
async fn chat_codex_stream(
&self,
label: &str,
model: &str,
prompt: &str,
mcp_config: Option<&str>,
progress: Option<tokio::sync::mpsc::Sender<String>>,
) -> Result<String> {
let mut cmd = Command::new("codex");
cmd.arg("exec").arg("--json").arg("--model").arg(model)
.arg("--dangerously-bypass-approvals-and-sandbox");
if let Some(mcp) = mcp_config {
for (name, cmdline, args) in mcp_servers_from(mcp) {
cmd.arg("-c").arg(format!("mcp_servers.{name}.command={cmdline}"));
cmd.arg("-c").arg(format!("mcp_servers.{name}.args={args}"));
}
}
cmd.arg("-");
cmd.stdin(Stdio::piped()).stdout(Stdio::piped()).stderr(Stdio::piped()).kill_on_drop(true);
let mut child = cmd.spawn().map_err(|e| anyhow!("spawn codex failed: {e}"))?;
if let Some(mut stdin) = child.stdin.take() {
stdin.write_all(prompt.as_bytes()).await?;
// Drop closes stdin so Codex processes the prompt and exits.
}
let stdout = child.stdout.take().ok_or_else(|| anyhow!("no stdout"))?;
let stderr = child.stderr.take().ok_or_else(|| anyhow!("no stderr"))?;
let mut lines = BufReader::new(stdout).lines();
let lbl = if label.is_empty() { String::new() } else { format!("@{label} ") };
let emit = |s: String| {
if let Some(tx) = &progress {
let _ = tx.try_send(format!("{lbl}{s}"));
}
};
// Last agent_message is the model's final answer; keep every one so a
// run that ends on a tool call still returns the most recent reasoning.
let mut result = String::new();
let read = async {
while let Ok(Some(line)) = lines.next_line().await {
let Ok(v) = serde_json::from_str::<serde_json::Value>(&line) else { continue };
let ty = v.get("type").and_then(|t| t.as_str()).unwrap_or("");
match ty {
"item.started" | "item.completed" => {
let Some(item) = v.get("item") else { continue };
let itype = item.get("type").and_then(|t| t.as_str()).unwrap_or("");
match itype {
"command_execution" => {
// Only announce on start (avoid double lines); note failures on completion.
let c = item.get("command").and_then(|x| x.as_str()).unwrap_or("");
// Strip the `/bin/sh -lc '...'` wrapper Codex adds.
let c = c.strip_prefix("/bin/sh -lc ").map(|s| s.trim_matches('\'')).unwrap_or(c);
if ty == "item.started" {
let danger = c.contains("rm -rf") || c.contains("mkfs")
|| c.contains(":(){") || c.contains("dd if=") || c.contains("> /dev/");
emit(format!("{}: {}", if danger { "danger" } else { "exec" }, truncate(c, 200)));
} else if let Some(code) = item.get("exit_code").and_then(|x| x.as_i64()) {
if code != 0 {
emit(format!("exec: (exit {code}) {}", truncate(c, 120)));
}
}
}
"agent_message" => {
if let Some(t) = item.get("text").and_then(|x| x.as_str()) {
let t = t.trim();
if !t.is_empty() {
if ty == "item.completed" { result = t.to_string(); }
emit(format!("ai: {}", truncate(t, 240)));
}
}
}
"file_change" | "patch" => {
let p = item.get("path").and_then(|x| x.as_str())
.or_else(|| item.get("file").and_then(|x| x.as_str())).unwrap_or("file");
emit(format!("edit: {p}"));
}
"mcp_tool_call" => {
let name = item.get("tool").and_then(|x| x.as_str())
.or_else(|| item.get("name").and_then(|x| x.as_str())).unwrap_or("mcp");
emit(format!("tool: {name}"));
}
"web_search" => {
let q = item.get("query").and_then(|x| x.as_str()).unwrap_or("");
emit(format!("net: search {}", truncate(q, 100)));
}
_ => {}
}
}
"turn.completed" => {
let ti = v.pointer("/usage/input_tokens").and_then(|x| x.as_u64());
let to = v.pointer("/usage/output_tokens").and_then(|x| x.as_u64());
if ti.is_some() || to.is_some() {
emit(format!("tokens: in={} out={}", ti.unwrap_or(0), to.unwrap_or(0)));
}
}
_ => {}
}
}
};
// Bound the whole streamed turn (matches the buffered path's cap).
if tokio::time::timeout(Duration::from_secs(900), read).await.is_err() {
return Err(anyhow!("codex stream timed out after 900s"));
}
let status = child.wait().await.ok();
// Drain stderr for auth/rate diagnostics if we got nothing usable.
if result.is_empty() {
let mut errbuf = String::new();
let mut el = BufReader::new(stderr).lines();
while let Ok(Some(l)) = el.next_line().await {
if !errbuf.is_empty() { errbuf.push('\n'); }
errbuf.push_str(&l);
if errbuf.len() > 2000 { break; }
}
let low = errbuf.to_lowercase();
let hard = ["not logged in", "please log in", "please login", "run /login",
"unauthorized", "not authenticated", "invalid api key", "no api key",
"rate limit", "429", "quota", "credit balance", "usage limit"]
.iter().any(|k| low.contains(k));
let code = status.and_then(|s| s.code()).map(|c| c.to_string()).unwrap_or_else(|| "signal".into());
if hard || !errbuf.trim().is_empty() {
return Err(anyhow!("codex subscription CLI exit {}: {}", code, truncate(errbuf.trim(), 240)));
}
return Err(anyhow!("codex stream produced no result"));
}
Ok(result)
}
} }
/// Categorise a Claude tool_use block into a tagged activity-feed event. /// Categorise a Claude tool_use block into a tagged activity-feed event.
+40 -18
View File
@@ -89,10 +89,13 @@ fn tool_doctrine(mcp_on: bool) -> String {
Keep bursts small and non-disruptive — this is a control check, not a DoS.\n\ Keep bursts small and non-disruptive — this is a control check, not a DoS.\n\
- TOOL DOWNLOAD (authorized): when a public PoC or scanner is needed you MAY `git clone` a specific PoC/exploit \ - TOOL DOWNLOAD (authorized): when a public PoC or scanner is needed you MAY `git clone` a specific PoC/exploit \
repo or download a tool (`git clone`, `wget`, `pip install`, `go install`, `cargo install`) — use pinned, \ repo or download a tool (`git clone`, `wget`, `pip install`, `go install`, `cargo install`) — use pinned, \
reputable sources; review before running; never run destructive payloads.\n\ reputable sources; review before running; never run destructive payloads. ALWAYS time-box downloads/installs \
(`timeout 90 <install> || echo skip`) and try each at most once — if it fails, isn't packaged, has no network \
or hangs, SKIP it and fall back to curl/nc/dig/python3. A missing or un-downloadable tool is NEVER a reason \
to stall: move on with what you have.\n\
- {browser}\n\ - {browser}\n\
- {ua}{proxy}{pocs}\ - {ua}{proxy}{pocs}\
Use only what is installed; degrade gracefully. Never run destructive or DoS actions.\n\n", Use only what is installed; degrade gracefully. Never block on a single tool install. Never run destructive or DoS actions.\n\n",
ua = ua_line(), ua = ua_line(),
proxy = proxy_line(), proxy = proxy_line(),
pocs = pocs_line(), pocs = pocs_line(),
@@ -353,7 +356,7 @@ pub async fn run(cfg: RunConfig, lib: &Library, pool: &ModelPool, tx: Sender<Str
findings.extend(chained); findings.extend(chained);
findings = dedup_findings(findings); findings = dedup_findings(findings);
let findings = refute_pass(findings, pool, cfg.vote_n, &tx).await; let findings = refute_pass(findings, pool, cfg.vote_n, &tx).await;
finish(cfg, lib, recon, transcript, findings, selected, &mut rl, tx).await finish(cfg, lib, recon, transcript, findings, selected, &mut rl, crate::grounding::GroundMode::Empirical, String::new(), tx).await
} }
/// White-box engagement: analyse a repository's source for vulnerabilities. /// White-box engagement: analyse a repository's source for vulnerabilities.
@@ -415,7 +418,7 @@ pub async fn run_whitebox(cfg: RunConfig, lib: &Library, pool: &ModelPool, tx: S
let _ = tx.send(format!("{} candidate finding(s) (deduped) — validating", candidates.len())).await; let _ = tx.send(format!("{} candidate finding(s) (deduped) — validating", candidates.len())).await;
let findings = validate(candidates, pool, CODE_VOTE_SYS, cfg.vote_n, &tx).await; let findings = validate(candidates, pool, CODE_VOTE_SYS, cfg.vote_n, &tx).await;
let findings = refute_pass(findings, pool, cfg.vote_n, &tx).await; let findings = refute_pass(findings, pool, cfg.vote_n, &tx).await;
finish(cfg, lib, "{}".into(), transcript, findings, selected, &mut rl, tx).await finish(cfg, lib, "{}".into(), transcript, findings, selected, &mut rl, crate::grounding::GroundMode::Symbolic, context, tx).await
} }
/// Greybox engagement: review the source code AND exploit the running app in one /// Greybox engagement: review the source code AND exploit the running app in one
@@ -558,7 +561,7 @@ pub async fn run_greybox(cfg: RunConfig, lib: &Library, pool: &ModelPool, tx: Se
findings.extend(chained); findings.extend(chained);
findings = dedup_findings(findings); findings = dedup_findings(findings);
let findings = refute_pass(findings, pool, cfg.vote_n, &tx).await; let findings = refute_pass(findings, pool, cfg.vote_n, &tx).await;
finish(cfg, lib, recon, transcript, findings, selected, &mut rl, tx).await finish(cfg, lib, recon, transcript, findings, selected, &mut rl, crate::grounding::GroundMode::Either, context, tx).await
} }
const CHAIN_SYS: &str = "You are a post-exploitation & attack-chaining specialist. You are given ONE confirmed foothold plus any loot already gathered. DECIDE the most promising directions to expand from THIS foothold and pursue them with real tools: post-exploitation (loot credentials/tokens/keys/config/source), credential reuse, privilege escalation (horizontal AND vertical), lateral movement to adjacent services/hosts, data exfiltration, and reaching NEW attack surface the foothold exposes (e.g. SSRF→cloud metadata creds→IAM, SQLi→DB dump→credential reuse→admin, arbitrary file read→secrets→RCE, IDOR→account takeover, auth bypass→internal APIs). PROVE each escalated step with a real tool receipt. Report ONLY NEW findings beyond the input, plus any new loot you discovered (creds, tokens, hosts, internal endpoints) so later stages can reuse it. Authorized engagement; never destructive/DoS."; const CHAIN_SYS: &str = "You are a post-exploitation & attack-chaining specialist. You are given ONE confirmed foothold plus any loot already gathered. DECIDE the most promising directions to expand from THIS foothold and pursue them with real tools: post-exploitation (loot credentials/tokens/keys/config/source), credential reuse, privilege escalation (horizontal AND vertical), lateral movement to adjacent services/hosts, data exfiltration, and reaching NEW attack surface the foothold exposes (e.g. SSRF→cloud metadata creds→IAM, SQLi→DB dump→credential reuse→admin, arbitrary file read→secrets→RCE, IDOR→account takeover, auth bypass→internal APIs). PROVE each escalated step with a real tool receipt. Report ONLY NEW findings beyond the input, plus any new loot you discovered (creds, tokens, hosts, internal endpoints) so later stages can reuse it. Authorized engagement; never destructive/DoS.";
@@ -909,16 +912,28 @@ async fn refute_pass(findings: Vec<Finding>, pool: &ModelPool, vote_n: usize, tx
} }
async fn finish(cfg: RunConfig, _lib: &Library, recon: String, transcript: String, mut findings: Vec<Finding>, async fn finish(cfg: RunConfig, _lib: &Library, recon: String, transcript: String, mut findings: Vec<Finding>,
selected: Vec<Agent>, rl: &mut RlState, tx: Sender<String>) -> RunOutput { selected: Vec<Agent>, rl: &mut RlState, gmode: crate::grounding::GroundMode, source_ctx: String,
// --- Grounding gate: no claim without a tool receipt (anti-hallucination) --- tx: Sender<String>) -> RunOutput {
// White/grey carry source context; black-box is verified empirically. use crate::grounding::GroundMode;
let whitebox = cfg.repo.is_some() && cfg.target.starts_with('/'); // --- Grounding gate: no claim without a receipt (anti-hallucination) ---
// The receipt is empirical (tool output) for black-box, symbolic (file:line
// into the reviewed source) for white-box SAST / skills audits, or either for
// grey-box. Symbolic grounding is checked against the SOURCE corpus, not the
// model transcript, so a code citation is honoured as its own receipt.
let ground_ctx = if source_ctx.is_empty() { transcript.as_str() } else { source_ctx.as_str() };
let before = findings.len(); let before = findings.len();
let (kept, demoted) = crate::grounding::gate(findings, &transcript, whitebox); let (kept, demoted) = crate::grounding::gate(findings, ground_ctx, gmode);
findings = kept; findings = kept;
if demoted > 0 { if demoted > 0 {
let _ = tx.send(format!("grounding gate: demoted {demoted}/{before} ungrounded claim(s) (no tool receipt)")).await; let receipt = match gmode {
GroundMode::Symbolic => "no source reference",
GroundMode::Either => "no source reference nor tool receipt",
GroundMode::Empirical => "no tool receipt",
};
let _ = tx.send(format!("grounding gate: demoted {demoted}/{before} ungrounded claim(s) ({receipt})")).await;
} }
// White-box/skills are symbolic → deterministic belief; grey-box carries source too.
let whitebox = matches!(gmode, GroundMode::Symbolic | GroundMode::Either);
// --- v3.5.2 report-hygiene & exploitation-depth pass --- // --- v3.5.2 report-hygiene & exploitation-depth pass ---
// Calibrate inflated/unproven High-Critical to Medium, flag exposures that // Calibrate inflated/unproven High-Critical to Medium, flag exposures that
@@ -1274,7 +1289,7 @@ pub async fn run_host(cfg: RunConfig, lib: &Library, pool: &ModelPool, tx: Sende
findings.extend(chained); findings.extend(chained);
findings = dedup_findings(findings); findings = dedup_findings(findings);
let findings = refute_pass(findings, pool, cfg.vote_n, &tx).await; let findings = refute_pass(findings, pool, cfg.vote_n, &tx).await;
finish(cfg, lib, recon, transcript, findings, selected, &mut rl, tx).await finish(cfg, lib, recon, transcript, findings, selected, &mut rl, crate::grounding::GroundMode::Empirical, String::new(), tx).await
} }
/// AI-red-team doctrine prepended to every AI/LLM/agent test prompt. /// AI-red-team doctrine prepended to every AI/LLM/agent test prompt.
@@ -1300,8 +1315,15 @@ fn recon_intensity_directive(level: usize) -> String {
}; };
format!( format!(
"RECON INTENSITY: {label}{rounds}. {extra}\n\ "RECON INTENSITY: {label}{rounds}. {extra}\n\
INSTALL WHAT YOU NEED (authorized): if a recon tool is missing, install it before falling back — \ INSTALL WHAT YOU NEED (authorized), BUT NEVER GET STUCK ON AN INSTALL: if a recon tool is missing, \
`apt-get install -y <t>`, `pip install <t>`, `go install <pkg>@latest`, `npm i -g <t>`, or `cargo install <t>`. \ try to install it — but TIME-BOX every install and move on if it fails. Always wrap installs like \
`timeout 90 apt-get install -y <t> || timeout 90 go install <pkg>@latest || echo 'skip <t>'` and \
run them non-interactively (`DEBIAN_FRONTEND=noninteractive`, `-y`, no prompts). \
Try a given tool install AT MOST ONCE — if it errors, is not packaged, needs a different OS, \
has no network, or hangs past the timeout, SKIP IT immediately and use an already-installed \
alternative or plain `curl`/`nc`/`dig`/`openssl`/`python3`. Do not wait on, retry, or block the \
whole recon for any single tool download — a missing tool is never a reason to stall. \
Options — `pip install <t>`, `go install <pkg>@latest`, `npm i -g <t>`, or `cargo install <t>` (all time-boxed). \
Recommended arsenal: subfinder/amass/assetfinder (subdomains), httpx/httprobe (probe live), \ Recommended arsenal: subfinder/amass/assetfinder (subdomains), httpx/httprobe (probe live), \
gau/waybackurls/katana/hakrawler/gospider (URL harvest & crawl), gf (pattern-filter urls), \ gau/waybackurls/katana/hakrawler/gospider (URL harvest & crawl), gf (pattern-filter urls), \
arjun/paramspider (params), ffuf/feroxbuster/dirsearch (content discovery), nuclei (targeted templates), \ arjun/paramspider (params), ffuf/feroxbuster/dirsearch (content discovery), nuclei (targeted templates), \
@@ -1387,7 +1409,7 @@ pub async fn run_ai(cfg: RunConfig, lib: &Library, pool: &ModelPool, tx: Sender<
let mut rl = cfg.rl_path.as_ref().map(|p| RlState::load(Path::new(p))).unwrap_or_default(); let mut rl = cfg.rl_path.as_ref().map(|p| RlState::load(Path::new(p))).unwrap_or_default();
if cfg.offline { if cfg.offline {
let _ = tx.send("offline: no AI exploitation performed".into()).await; let _ = tx.send("offline: no AI exploitation performed".into()).await;
return finish(cfg, lib, recon, String::new(), vec![], agents, &mut rl, tx).await; return finish(cfg, lib, recon, String::new(), vec![], agents, &mut rl, crate::grounding::GroundMode::Empirical, String::new(), tx).await;
} }
let cap = if cfg.max_agents > 0 { cfg.max_agents.min(agents.len()) } else { agents.len() }; let cap = if cfg.max_agents > 0 { cfg.max_agents.min(agents.len()) } else { agents.len() };
let selected: Vec<Agent> = agents.into_iter().take(cap).collect(); let selected: Vec<Agent> = agents.into_iter().take(cap).collect();
@@ -1434,7 +1456,7 @@ pub async fn run_ai(cfg: RunConfig, lib: &Library, pool: &ModelPool, tx: Sender<
findings.extend(chained); findings.extend(chained);
findings = dedup_findings(findings); findings = dedup_findings(findings);
let findings = refute_pass(findings, pool, cfg.vote_n, &tx).await; let findings = refute_pass(findings, pool, cfg.vote_n, &tx).await;
finish(cfg, lib, recon, transcript, findings, selected, &mut rl, tx).await finish(cfg, lib, recon, transcript, findings, selected, &mut rl, crate::grounding::GroundMode::Empirical, String::new(), tx).await
} }
/// White-box Skills/plugin audit: read the skill .md file or a folder of them and /// White-box Skills/plugin audit: read the skill .md file or a folder of them and
@@ -1453,7 +1475,7 @@ pub async fn run_skills_audit(cfg: RunConfig, lib: &Library, pool: &ModelPool, t
let mut rl = cfg.rl_path.as_ref().map(|p| RlState::load(Path::new(p))).unwrap_or_default(); let mut rl = cfg.rl_path.as_ref().map(|p| RlState::load(Path::new(p))).unwrap_or_default();
if cfg.offline || context.is_empty() { if cfg.offline || context.is_empty() {
let _ = tx.send("offline or empty skills input — nothing audited".into()).await; let _ = tx.send("offline or empty skills input — nothing audited".into()).await;
return finish(cfg, lib, "{}".into(), String::new(), vec![], agents, &mut rl, tx).await; return finish(cfg, lib, "{}".into(), String::new(), vec![], agents, &mut rl, crate::grounding::GroundMode::Symbolic, String::new(), tx).await;
} }
let directives = operator_directives(&cfg); let directives = operator_directives(&cfg);
let raw: Vec<(String, String, Vec<Finding>)> = stream::iter(agents.iter().cloned()) let raw: Vec<(String, String, Vec<Finding>)> = stream::iter(agents.iter().cloned())
@@ -1484,5 +1506,5 @@ pub async fn run_skills_audit(cfg: RunConfig, lib: &Library, pool: &ModelPool, t
let transcript = transcript_of(&raw); let transcript = transcript_of(&raw);
let candidates = dedup_findings(raw.iter().flat_map(|(_, _, f)| f.clone()).collect()); let candidates = dedup_findings(raw.iter().flat_map(|(_, _, f)| f.clone()).collect());
let findings = validate(candidates, pool, CODE_VOTE_SYS, cfg.vote_n, &tx).await; let findings = validate(candidates, pool, CODE_VOTE_SYS, cfg.vote_n, &tx).await;
finish(cfg, lib, "{}".into(), transcript, findings, agents, &mut rl, tx).await finish(cfg, lib, "{}".into(), transcript, findings, agents, &mut rl, crate::grounding::GroundMode::Symbolic, context, tx).await
} }
+1 -1
View File
@@ -1,4 +1,4 @@
//! POMDP decision layer (v3.6.0): value-of-information planning + the //! POMDP decision layer (v3.6.4): value-of-information planning + the
//! anti-hallucination gate. //! anti-hallucination gate.
//! //!
//! The choice "scan more vs exploit now" is **not** a heuristic here — it falls //! The choice "scan more vs exploit now" is **not** a heuristic here — it falls
+1 -1
View File
@@ -1,4 +1,4 @@
//! Deterministic HTTP request/response analysis (v3.6.0). //! Deterministic HTTP request/response analysis (v3.6.4).
//! //!
//! Before the LLM recon runs, the harness performs a **real** probe of the //! Before the LLM recon runs, the harness performs a **real** probe of the
//! target and captures observed facts — status, headers, security headers, //! target and captures observed facts — status, headers, security headers,
+3 -3
View File
@@ -97,9 +97,9 @@ pub fn html(target: &str, findings: &[Finding]) -> String {
h4{{margin:12px 0 3px;font-size:12px;text-transform:uppercase;letter-spacing:.5px;color:#8b5cf6}}\ h4{{margin:12px 0 3px;font-size:12px;text-transform:uppercase;letter-spacing:.5px;color:#8b5cf6}}\
.b{{color:#8b5cf6;font-weight:800}}</style></head><body>\ .b{{color:#8b5cf6;font-weight:800}}</style></head><body>\
<h1><span class=b>NeuroSploit</span> Penetration Test Report</h1>\ <h1><span class=b>NeuroSploit</span> Penetration Test Report</h1>\
<div class=meta>Target: <b>{t}</b> · v3.6.0 Rust harness · multi-model validated</div>\ <div class=meta>Target: <b>{t}</b> · v3.6.4 Rust harness · multi-model validated</div>\
<div>{chips}</div>{graph_block}<h2>Findings ({n})</h2>{body}\ <div>{chips}</div>{graph_block}<h2>Findings ({n})</h2>{body}\
<p class=meta>Authorized testing only. Findings confirmed by multi-model adversarial voting.<br>NeuroSploit v3.6.0 · by <b>Joas A Santos</b> &amp; <b>Red Team Leaders</b></p></body></html>", <p class=meta>Authorized testing only. Findings confirmed by multi-model adversarial voting.<br>NeuroSploit v3.6.4 · by <b>Joas A Santos</b> &amp; <b>Red Team Leaders</b></p></body></html>",
t = esc(target), chips = chips, n = sorted.len(), body = body, graph_block = graph_block, t = esc(target), chips = chips, n = sorted.len(), body = body, graph_block = graph_block,
) )
} }
@@ -135,7 +135,7 @@ pub fn typst_report(target: &str, findings: &[Finding], dir: &Path) -> std::io::
let mut data = String::new(); let mut data = String::new();
data.push_str(&format!( data.push_str(&format!(
"#let meta = (target: {}, run_id: {}, generated: {}, model: {})\n", "#let meta = (target: {}, run_id: {}, generated: {}, model: {})\n",
tq(target), tq(&run_id), tq("NeuroSploit v3.6.0"), tq("multi-model") tq(target), tq(&run_id), tq("NeuroSploit v3.6.4"), tq("multi-model")
)); ));
data.push_str("#let findings = (\n"); data.push_str("#let findings = (\n");
for f in sorted_findings(findings) { for f in sorted_findings(findings) {
+2 -2
View File
@@ -28,7 +28,7 @@ cat <<'BANNER'
███╗ ██╗███████╗██╗ ██╗██████╗ ██████╗ ███╗ ██╗███████╗██╗ ██╗██████╗ ██████╗
████╗ ██║██╔════╝██║ ██║██╔══██╗██╔═══██╗ NeuroSploit installer ████╗ ██║██╔════╝██║ ██║██╔══██╗██╔═══██╗ NeuroSploit installer
██╔██╗ ██║█████╗ ██║ ██║██████╔╝██║ ██║ v3.6.0 — Rust harness ██╔██╗ ██║█████╗ ██║ ██║██████╔╝██║ ██║ v3.6.1 — Rust harness
██║╚██╗██║██╔══╝ ██║ ██║██╔══██╗██║ ██║ by Joas A Santos ██║╚██╗██║██╔══╝ ██║ ██║██╔══██╗██║ ██║ by Joas A Santos
██║ ╚████║███████╗╚██████╔╝██║ ██║╚██████╔╝ & Red Team Leaders ██║ ╚████║███████╗╚██████╔╝██║ ██║╚██████╔╝ & Red Team Leaders
╚═╝ ╚═══╝╚══════╝ ╚═════╝ ╚═╝ ╚═╝ ╚═════╝ ╚═╝ ╚═══╝╚══════╝ ╚═════╝ ╚═╝ ╚═╝ ╚═════╝
@@ -63,7 +63,7 @@ if [ -z "$REF" ]; then
REF="$(dl "https://api.github.com/repos/${REPO_SLUG}/releases/latest" /dev/stdout 2>/dev/null \ REF="$(dl "https://api.github.com/repos/${REPO_SLUG}/releases/latest" /dev/stdout 2>/dev/null \
| grep -m1 '"tag_name"' | sed -E 's/.*"tag_name" *: *"([^"]+)".*/\1/' || true)" | grep -m1 '"tag_name"' | sed -E 's/.*"tag_name" *: *"([^"]+)".*/\1/' || true)"
fi fi
[ -z "$REF" ] && REF="v3.6.0" [ -z "$REF" ] && REF="v3.6.1"
say "Release: $REF" say "Release: $REF"
installed=0 installed=0