Compare commits

..
6 Commits
Author SHA1 Message Date
CyberSecurityUPandClaude Opus 5 69c5e3ddb9 feat(3.6.9): OpenCode Zen + Nous Research (Hermes) providers
Add two new model providers, both usable via API key or --subscription
(local CLI login, no key):

- opencode: OpenCode Zen gateway (OPENCODE_API_KEY, opencode.ai/zen/v1).
  Subscription mode drives the `opencode` CLI (`opencode run --auto`).
  Supports the Playwright MCP (--mcp): our .mcp.json is converted to
  OpenCode's own config schema and injected via OPENCODE_CONFIG.

- nous: Nous Research / Hermes models (NOUS_API_KEY,
  inference-api.nousresearch.com/v1). Subscription mode drives the
  `hermes` CLI (NousResearch/hermes-agent) on the user's Nous Portal
  OAuth login (`hermes setup --portal`), via `hermes chat -q`. No
  CLI-level MCP hook — falls back to Hermes's own built-in toolsets
  (web/terminal/computer-use).

Both wired into cli_binary_for, installed_cli_backends, cli_login_status
(prompt passed as argv, not stdin — neither CLI reads stdin for this).

Bump version 3.6.8 -> 3.6.9 across Cargo.toml, README, TUTORIAL, setup.sh,
install.ps1, and in-binary version strings. README/.env.example updated
with the new provider rows and subscription-login table.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HHFAVCHMvRkTy9Wgw7SayG
2026-08-11 23:47:12 -03:00
CyberSecurityUPandClaude Opus 4.6 1f8ccb6f9e fix(3.6.8): recon time budget — 5min cap prevents recon from eating entire run
- Add RECON_TOTAL_BUDGET_SECS (300s) total wall-clock cap across all rounds
- Per-round budget directive in prompt: 30-50 commands max, stop early if enough intel
- Elapsed time check between rounds: skip remaining if budget exhausted
- Remaining time communicated to follow-up rounds for self-pacing
- RELEASE.md updated with recon budget section

Previously: subscription CLI recon ran 150+ commands over 15 min, exploitation never started.
Now: recon caps at 5 min total, then proceeds to agent exploitation.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-08-08 10:10:43 -03:00
CyberSecurityUPandClaude Opus 4.6 3c49a83578 fix(3.6.8): auth resilience — circuit breaker pauses run on token revocation, preserves findings
- Add is_auth_failure() detector (401, OAuth revoked, session expired, invalid key)
- Circuit breaker: 3 consecutive auth failures auto-pause instead of burning 66 agents
- Auth-aware park_exhausted(): clear message + fallback provider switch via /continue
- No retry burn on auth errors (immediate return like exhaustion)
- Recon preserves HTTP probe facts when model auth fails
- REPL phase tracking: paused (auth) distinct from paused (quota)
- RELEASE.md updated with auth resilience section

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-08-08 09:14:21 -03:00
CyberSecurityUPandClaude Opus 4.6 e956b482b9 fix(3.6.8): JSON parse resilience + diagnostics for local model failures
- extract_findings: log when model output has no JSON (was silent drop)
- extract_findings: auto-fix trailing-comma JSON (common LLM mistake)
- pipeline: emit response tail when agent returns 0 parseable findings
- Helps diagnose why small/local models produce 0 findings on valid targets

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-08-07 09:51:50 -03:00
CyberSecurityUPandClaude Opus 4.6 105c62af61 docs: bump version references to 3.6.8 across README, TUTORIAL, setup.sh, install.ps1
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011739wMqPJJPttTLLX6YoQH
2026-08-06 15:27:22 -03:00
CyberSecurityUPandClaude Opus 4.6 a0a477a2bf fix(3.6.8): better Ollama error messages, empty-evidence findings go to needs-review, single-model vote warning
- models.rs: detect connection-refused and timeout on local providers
  (ollama/litellm/llamacpp), show actionable error instead of raw reqwest
- pipeline.rs: findings with empty evidence skip adversarial vote (which
  always rejects per 'default to rejected' prompt) and go straight to
  needs-review for human triage
- pipeline.rs: warn when single-model panel + vote_n=1 (same model
  validates its own findings = weaker validation)
- Bump version 3.6.7 → 3.6.8

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011739wMqPJJPttTLLX6YoQH
2026-08-06 15:21:37 -03:00
15 changed files with 412 additions and 61 deletions
+10
View File
@@ -51,6 +51,16 @@ TOGETHER_API_KEY=
# openrouter: https://openrouter.ai/keys # openrouter: https://openrouter.ai/keys
OPENROUTER_API_KEY= OPENROUTER_API_KEY=
# opencode: https://opencode.ai/auth (OpenCode Zen gateway)
# Or skip the key entirely and use --subscription with the
# `opencode` CLI logged into your own Zen/plan account.
OPENCODE_API_KEY=
# nous: Nous Portal (https://portal.nousresearch.com) — Hermes models.
# Or skip the key entirely and use --subscription with the
# `hermes` CLI (`hermes setup --portal` for OAuth login).
NOUS_API_KEY=
# ollama: local, no key needed. Override the endpoint if not default: # ollama: local, no key needed. Override the endpoint if not default:
#OLLAMA_BASE_URL=http://localhost:11434/v1 #OLLAMA_BASE_URL=http://localhost:11434/v1
+1
View File
@@ -108,3 +108,4 @@ data/repl_history.txt
# Cloned source repos (whitebox/greybox from a git URL) # Cloned source repos (whitebox/greybox from a git URL)
repos/ repos/
neurosploit-rs/repos/ neurosploit-rs/repos/
target/
+12 -2
View File
@@ -1,4 +1,4 @@
<h1 align="center">🧠 NeuroSploit v3.6.7</h1> <h1 align="center">🧠 NeuroSploit v3.6.9</h1>
<p align="center"> <p align="center">
<a href="https://trendshift.io/repositories/22624?utm_source=trendshift-badge&amp;utm_medium=badge&amp;utm_campaign=badge-trendshift-22624" target="_blank" rel="noopener noreferrer"><img src="https://trendshift.io/api/badge/trendshift/repositories/22624/daily?language=Python" alt="JoasASantos%2FNeuroSploit | Trendshift" width="250" height="55"/></a> <a href="https://trendshift.io/repositories/22624?utm_source=trendshift-badge&amp;utm_medium=badge&amp;utm_campaign=badge-trendshift-22624" target="_blank" rel="noopener noreferrer"><img src="https://trendshift.io/api/badge/trendshift/repositories/22624/daily?language=Python" alt="JoasASantos%2FNeuroSploit | Trendshift" width="250" height="55"/></a>
@@ -12,7 +12,7 @@
</p> </p>
<p align="center"> <p align="center">
<img src="https://img.shields.io/badge/Version-3.6.7-blue?style=flat-square"> <img src="https://img.shields.io/badge/Version-3.6.9-blue?style=flat-square">
<img src="https://img.shields.io/badge/Harness-Rust%20%7C%20tokio-e6b673?style=flat-square"> <img src="https://img.shields.io/badge/Harness-Rust%20%7C%20tokio-e6b673?style=flat-square">
<img src="https://img.shields.io/badge/License-MIT-green?style=flat-square"> <img src="https://img.shields.io/badge/License-MIT-green?style=flat-square">
<img src="https://img.shields.io/badge/MD%20Agents-435-red?style=flat-square"> <img src="https://img.shields.io/badge/MD%20Agents-435-red?style=flat-square">
@@ -433,6 +433,8 @@ export GROQ_API_KEY=... # groq:*
export TOGETHER_API_KEY=... # together:* export TOGETHER_API_KEY=... # together:*
export MOONSHOT_API_KEY=... # moonshot:* (Kimi K3/K2) export MOONSHOT_API_KEY=... # moonshot:* (Kimi K3/K2)
export OPENROUTER_API_KEY=... # openrouter:* export OPENROUTER_API_KEY=... # openrouter:*
export OPENCODE_API_KEY=... # opencode:* (OpenCode Zen gateway)
export NOUS_API_KEY=... # nous:* (Nous Portal — Hermes)
# ollama / llamacpp need no key (local) # ollama / llamacpp need no key (local)
# then run via API (note: NO --subscription) # then run via API (note: NO --subscription)
@@ -462,6 +464,8 @@ Or put the keys in a `.env` and source it (`cp .env.example .env`; edit; `set -a
| `together:` | `TOGETHER_API_KEY` | api.together.xyz | | `together:` | `TOGETHER_API_KEY` | api.together.xyz |
| `moonshot:` | `MOONSHOT_API_KEY` | api.moonshot.ai | | `moonshot:` | `MOONSHOT_API_KEY` | api.moonshot.ai |
| `openrouter:` | `OPENROUTER_API_KEY` | openrouter.ai | | `openrouter:` | `OPENROUTER_API_KEY` | openrouter.ai |
| `opencode:` | `OPENCODE_API_KEY` | opencode.ai/zen (OpenCode Zen gateway) |
| `nous:` | `NOUS_API_KEY` | inference-api.nousresearch.com (Hermes 4) |
| `ollama:` | _(none)_ | localhost:11434 | | `ollama:` | _(none)_ | localhost:11434 |
| `llamacpp:` | _(none)_ | localhost:8080 | | `llamacpp:` | _(none)_ | localhost:8080 |
@@ -484,6 +488,12 @@ install and log into one of the CLIs first:
| `openai:` | `codex` | `codex` login | | `openai:` | `codex` | `codex` login |
| `gemini:` | `gemini` | `gemini` login | | `gemini:` | `gemini` | `gemini` login |
| `xai:` | `grok` | `grok` login | | `xai:` | `grok` | `grok` login |
| `opencode:` | `opencode` | `opencode auth login` (or `/connect` in the TUI) — Zen/plan account |
| `nous:` | `hermes` | `hermes setup --portal` — Nous Portal OAuth |
`opencode:` also gets the Playwright MCP (`--mcp`) like anthropic/openai do.
`nous:` relies on Hermes's own built-in toolsets (web/terminal/computer-use)
instead — it has no CLI-level MCP hook.
```bash ```bash
./target/release/neurosploit run http://testphp.vulnweb.com/ \ ./target/release/neurosploit run http://testphp.vulnweb.com/ \
+63 -2
View File
@@ -1,11 +1,72 @@
# NeuroSploit v3.6.7 — Release Notes # NeuroSploit v3.6.8 — Release Notes
**Release Date:** August 2026 **Release Date:** August 2026
**Codename:** Chain & Exploit **Codename:** Chain & Exploit
**License:** MIT **License:** MIT
**Credits:** Joas A Santos & Red Team Leaders **Credits:** Joas A Santos & Red Team Leaders
## Highlights ## v3.6.8 — Auth resilience, circuit breaker, recon budget, Ollama error handling, empty-evidence validation
### Recon Time Budget (NEW)
- **5-minute total recon budget.** Recon phase is now time-boxed to 300 seconds
across ALL rounds. Previously, a single recon round could run 150+ commands
over 15 minutes via subscription CLI, leaving no time for exploitation.
- **Per-round budget directive.** Each recon round receives a prompt instruction
with its share of the time budget (e.g. "~100 seconds for this round") and a
command count guideline (30-50 commands max). The model is instructed to
prioritise high-signal actions and stop early when enough intel is gathered.
- **Elapsed time check between rounds.** Before starting each follow-up round,
the pipeline checks elapsed time. If the budget is exhausted, recon stops
immediately and proceeds to exploitation with the intelligence gathered so far.
### Auth Resilience & Circuit Breaker (NEW)
- **`is_auth_failure()` detector.** New function recognises OAuth token revocation
(401), session expiry, invalid/revoked API keys, and "not logged in" errors from
subscription CLIs. Distinct from `is_exhaustion()` (quota/rate-limit) — auth
failures are non-recoverable without re-login or provider switch.
- **Circuit breaker (3 consecutive auth failures → auto-pause).** A shared atomic
counter tracks consecutive auth failures across ALL agents. After 3 failures the
pool pauses the run BEFORE burning through the remaining agents on a dead token.
Previously, a revoked OAuth token caused all 66+ agents to silently return 0
findings with no pause or warning.
- **Auth-aware park: findings preserved, fallback offered.** When auth fails the
run parks with a clear message:
`⏸ authentication failed (...). Run is PAUSED — all findings so far are SAFE.`
The user can `/continue openai:gpt-5.1` (or any provider) to switch and resume.
All `LiveCheckpoint` findings on disk are preserved across the pause.
- **No retry burn on auth failure.** `one()` returns immediately on auth errors
instead of retrying 3 times against a dead token (same as quota exhaustion).
- **Recon preserves probe facts on auth failure.** When model recon fails with an
auth error, the HTTP probe data is still returned and the pipeline continues
with probe-only intelligence instead of silently dropping everything.
- **Phase tracking for auth pauses.** The REPL status line shows `paused (auth)`
(distinct from `paused (quota)`) so the operator knows the root cause at a glance.
### Bugfixes
- **Better Ollama/local provider error messages.** Connection-refused and timeout
errors now name the provider, URL, and suggest checking if the server is running.
Previously showed raw reqwest errors.
- **Empty-evidence findings skip the vote and go straight to `needs-review`.**
Findings with no evidence are unverifiable by the adversarial validator (which
always rejects "no evidence" per its system prompt). Now they bypass the vote
and are flagged for human review instead of being silently dropped.
- **Single-model + vote_n=1 warning.** When only one model is configured and
vote_n is 1, the pipeline emits a warning that validation is weaker (same model
validates its own findings).
- **JSON parse resilience for local models.** `extract_findings` now logs when a
model returns text but no parseable JSON (previously silent drop — 0 findings
with no diagnostic). Also auto-fixes trailing-comma JSON (`[...,]`) which small
models commonly produce.
- **Visible diagnostics when agents return 0 findings.** Pipeline emits the
response tail so the operator can see what the model actually returned (helps
debug model quality issues with local/small models).
---
## v3.6.7 Highlights
- **CVE exploitation pipeline — 4 new agents.** `cve_version_fingerprint` (pin - **CVE exploitation pipeline — 4 new agents.** `cve_version_fingerprint` (pin
exact versions) → `cve_research_analyst` (map to NVD/GHSA, judge reachability) → exact versions) → `cve_research_analyst` (map to NVD/GHSA, judge reachability) →
+2 -2
View File
@@ -1,4 +1,4 @@
# NeuroSploit — Tutorial & User Guide (v3.6.5) # NeuroSploit — Tutorial & User Guide (v3.6.9)
A complete, hands-on guide to installing, configuring and running NeuroSploit — A complete, hands-on guide to installing, configuring and running NeuroSploit —
the autonomous, multi-model penetration-testing harness. the autonomous, multi-model penetration-testing harness.
@@ -98,7 +98,7 @@ Agents **degrade gracefully**: if `rustscan` is absent they use `nmap`; if neith
### Verify ### Verify
```bash ```bash
neurosploit --version # neurosploit 3.6.5 neurosploit --version # neurosploit 3.6.9
neurosploit agents # {"vulns":241,...,"ai":30,...,"total":430} neurosploit agents # {"vulns":241,...,"ai":30,...,"total":430}
neurosploit models # all providers & models neurosploit models # all providers & models
``` ```
+2 -2
View File
@@ -14,7 +14,7 @@ function Ok ($m) { Write-Host " + $m" -ForegroundColor Green }
function Warn($m){ Write-Host " ! $m" -ForegroundColor Yellow } function Warn($m){ Write-Host " ! $m" -ForegroundColor Yellow }
Write-Host "" Write-Host ""
Write-Host " NeuroSploit installer (Windows) — v3.6.1" -ForegroundColor Cyan Write-Host " NeuroSploit installer (Windows) — v3.6.9" -ForegroundColor Cyan
# arch → asset arch (only x64 prebuilt today; arm64 falls back to source) # arch → asset arch (only x64 prebuilt today; arm64 falls back to source)
$rawArch = $env:PROCESSOR_ARCHITECTURE $rawArch = $env:PROCESSOR_ARCHITECTURE
@@ -29,7 +29,7 @@ $ref = $env:NEUROSPLOIT_REF
if (-not $ref) { if (-not $ref) {
try { $ref = (Invoke-RestMethod "https://api.github.com/repos/$slug/releases/latest").tag_name } catch { } try { $ref = (Invoke-RestMethod "https://api.github.com/repos/$slug/releases/latest").tag_name } catch { }
} }
if (-not $ref) { $ref = "v3.6.1" } if (-not $ref) { $ref = "v3.6.9" }
Say "Release: $ref" Say "Release: $ref"
New-Item -ItemType Directory -Force -Path $dir | Out-Null New-Item -ItemType Directory -Force -Path $dir | Out-Null
+2 -2
View File
@@ -871,7 +871,7 @@ dependencies = [
[[package]] [[package]]
name = "neurosploit" name = "neurosploit"
version = "3.6.7" version = "3.6.9"
dependencies = [ dependencies = [
"anyhow", "anyhow",
"clap", "clap",
@@ -888,7 +888,7 @@ dependencies = [
[[package]] [[package]]
name = "neurosploit-harness" name = "neurosploit-harness"
version = "3.6.7" version = "3.6.9"
dependencies = [ dependencies = [
"anyhow", "anyhow",
"futures", "futures",
+1 -1
View File
@@ -3,7 +3,7 @@ members = ["crates/harness", "app"]
resolver = "2" resolver = "2"
[workspace.package] [workspace.package]
version = "3.6.7" version = "3.6.9"
edition = "2021" edition = "2021"
license = "MIT" license = "MIT"
repository = "https://github.com/JoasASantos/NeuroSploit" repository = "https://github.com/JoasASantos/NeuroSploit"
+6 -6
View File
@@ -1,4 +1,4 @@
//! NeuroSploit v3.6.7 — interactive harness + CLI (`run` / `whitebox` / `agents` / `models`). //! NeuroSploit v3.6.9 — interactive harness + CLI (`run` / `whitebox` / `agents` / `models`).
mod repl; mod repl;
mod tui; mod tui;
@@ -11,9 +11,9 @@ use std::path::{Path, PathBuf};
#[command( #[command(
name = "neurosploit", name = "neurosploit",
version, version,
about = "NeuroSploit v3.6.7 — multi-model autonomous pentest harness", about = "NeuroSploit v3.6.9 — multi-model autonomous pentest harness",
long_about = "NeuroSploit v3.6.7 — a Rust multi-model harness that drives a pool of LLMs \ long_about = "NeuroSploit v3.6.9 — a Rust multi-model harness that drives a pool of LLMs \
(API key or local subscription: Claude/Codex/Gemini/Grok) to autonomously test a target. \ (API key or local subscription: Claude/Codex/Gemini/Grok/OpenCode/Hermes) to autonomously test a target. \
After recon it INTELLIGENTLY selects only the agents matching the discovered surface, runs \ After recon it INTELLIGENTLY selects only the agents matching the discovered surface, runs \
them in parallel, then validates every finding by cross-model voting before reporting.\n\n\ them in parallel, then validates every finding by cross-model voting before reporting.\n\n\
Run with NO arguments for an interactive wizard.\n\n\ Run with NO arguments for an interactive wizard.\n\n\
@@ -54,7 +54,7 @@ enum Cmd {
recon: usize, recon: usize,
#[arg(long)] #[arg(long)]
offline: bool, offline: bool,
/// Use local agentic CLI subscription (Claude/Codex/Gemini/Grok login). /// Use local agentic CLI subscription (Claude/Codex/Gemini/Grok/OpenCode/Hermes login).
#[arg(long)] #[arg(long)]
subscription: bool, subscription: bool,
/// Enable Playwright MCP (auto-installed if missing; backends that don't /// Enable Playwright MCP (auto-installed if missing; backends that don't
@@ -765,7 +765,7 @@ pub(crate) fn spawn_engagement(base: &Path, mut cfg: RunConfig, mcp: bool, mode:
println!(" │ ua : {ua}"); println!(" │ ua : {ua}");
write_status(&workdir, "running", &format!("\"target\":{:?}", cfg.target)); write_status(&workdir, "running", &format!("\"target\":{:?}", cfg.target));
println!(" ┌─ NeuroSploit v3.6.7 · by Joas A Santos & Red Team Leaders"); println!(" ┌─ NeuroSploit v3.6.9 · by Joas A Santos & Red Team Leaders");
println!(" │ run id : {run_id}"); println!(" │ run id : {run_id}");
println!(" │ target : {}", cfg.target); println!(" │ target : {}", cfg.target);
println!(" │ models : {}", cfg.models.join(", ")); println!(" │ models : {}", cfg.models.join(", "));
+3 -2
View File
@@ -1,4 +1,4 @@
//! NeuroSploit v3.6.7 — interactive session (Claude-Code / Codex / Cursor-CLI style). //! NeuroSploit v3.6.9 — interactive session (Claude-Code / Codex / Cursor-CLI style).
//! //!
//! Launched when `neurosploit` runs with no subcommand. A persistent REPL with //! Launched when `neurosploit` runs with no subcommand. A persistent REPL with
//! real line editing (arrow-key history recall, Ctrl-A/E/K, paste), model //! real line editing (arrow-key history recall, Ctrl-A/E/K, paste), model
@@ -68,6 +68,7 @@ impl RunLive {
if self.feed.len() > 200 { self.feed.remove(0); } if self.feed.len() > 200 { self.feed.remove(0); }
} }
if low.contains("token/quota exhausted") || low.contains("run is paused") { self.phase = "paused (quota)".into(); } if low.contains("token/quota exhausted") || low.contains("run is paused") { self.phase = "paused (quota)".into(); }
else if low.contains("authentication failed") || low.contains("auth failed") || low.contains("circuit breaker") { self.phase = "paused (auth)".into(); }
else if low.contains("resumed — retrying") { self.phase = "exploiting".into(); } else if low.contains("resumed — retrying") { self.phase = "exploiting".into(); }
else if low.starts_with("recon") || low.starts_with("ai-recon") || low.contains("recon round") || low.contains("intensity") || low.starts_with("probe:") { self.phase = "recon".into(); } else if low.starts_with("recon") || low.starts_with("ai-recon") || low.contains("recon round") || low.contains("intensity") || low.starts_with("probe:") { self.phase = "recon".into(); }
else if low.contains("selected") && low.contains("agent") { else if low.contains("selected") && low.contains("agent") {
@@ -370,7 +371,7 @@ pub async fn repl(base: &Path) -> anyhow::Result<()> {
let backends = harness::installed_cli_backends(); let backends = harness::installed_cli_backends();
println!("\x1b[1m"); println!("\x1b[1m");
println!(" ███╗ ██╗███████╗██╗ ██╗██████╗ ██████╗"); println!(" ███╗ ██╗███████╗██╗ ██╗██████╗ ██████╗");
println!(" ████╗ ██║██╔════╝██║ ██║██╔══██╗██╔═══██╗ NeuroSploit v3.6.7"); println!(" ████╗ ██║██╔════╝██║ ██║██╔══██╗██╔═══██╗ NeuroSploit v3.6.9");
println!(" ██╔██╗ ██║█████╗ ██║ ██║██████╔╝██║ ██║ interactive harness"); println!(" ██╔██╗ ██║█████╗ ██║ ██║██████╔╝██║ ██║ interactive harness");
println!(" ██║╚██╗██║██╔══╝ ██║ ██║██╔══██╗██║ ██║ by Joas A Santos"); println!(" ██║╚██╗██║██╔══╝ ██║ ██║██╔══██╗██║ ██║ by Joas A Santos");
println!(" ██║ ╚████║███████╗╚██████╔╝██║ ██║╚██████╔╝ & Red Team Leaders"); println!(" ██║ ╚████║███████╗╚██████╔╝██║ ██║╚██████╔╝ & Red Team Leaders");
+1 -1
View File
@@ -1,4 +1,4 @@
//! NeuroSploit v3.6.7 — TUI "Mission Control" mode. //! NeuroSploit v3.6.9 — TUI "Mission Control" mode.
//! //!
//! Concurrent panels that update live while the engagement runs in the //! Concurrent panels that update live while the engagement runs in the
//! background, with a composer input that stays active during execution: //! background, with a composer input that stays active during execution:
+126 -10
View File
@@ -52,6 +52,21 @@ pub fn providers() -> Vec<Provider> {
models: vec!["gpt-4o", "claude-3-7-sonnet", "gemini/gemini-2.5-pro"] }, models: vec!["gpt-4o", "claude-3-7-sonnet", "gemini/gemini-2.5-pro"] },
Provider { key: "openrouter", label: "OpenRouter", base_url: "https://openrouter.ai/api/v1", env_key: "OPENROUTER_API_KEY", kind: "api", Provider { key: "openrouter", label: "OpenRouter", base_url: "https://openrouter.ai/api/v1", env_key: "OPENROUTER_API_KEY", kind: "api",
models: vec!["anthropic/claude-opus-4-8", "qwen/qwen-2.5-coder-32b-instruct", "deepseek/deepseek-r1", "meta-llama/llama-3.3-70b-instruct"] }, models: vec!["anthropic/claude-opus-4-8", "qwen/qwen-2.5-coder-32b-instruct", "deepseek/deepseek-r1", "meta-llama/llama-3.3-70b-instruct"] },
// OpenCode Zen — the curated OpenAI-compatible gateway behind the
// `opencode` CLI (https://opencode.ai/zen). Works two ways, like
// anthropic/openai/xai/gemini above: as a plain API-key provider here,
// or (with --subscription) driven through the locally-installed
// `opencode` agentic CLI on the user's own Zen/plan login — no key
// needed in that mode. `kind: "cli"` reflects the latter.
Provider { key: "opencode", label: "OpenCode Zen", base_url: "https://opencode.ai/zen/v1", env_key: "OPENCODE_API_KEY", kind: "cli",
models: vec!["claude-opus-5", "claude-sonnet-5", "gpt-5.6-sol", "gpt-5.5", "gemini-3-pro", "grok-4.5", "deepseek-v4-pro", "qwen3.7-max", "kimi-k3"] },
// Nous Research — Hermes models via the Nous Portal. As an API-key
// provider here (OpenAI-compatible `inference-api.nousresearch.com`),
// or (with --subscription) driven through the `hermes` CLI
// (NousResearch/hermes-agent) on the user's OAuth Portal login
// (`hermes setup --portal`) — 300+ routed frontier models, no key.
Provider { key: "nous", label: "Nous Research (Hermes)", base_url: "https://inference-api.nousresearch.com/v1", env_key: "NOUS_API_KEY", kind: "cli",
models: vec!["Hermes-4-405B", "Hermes-4-70B", "DeepHermes-3-Mistral-24B-Preview"] },
// Azure OpenAI (OpenAI-compatible). Set AZURE_OPENAI_ENDPOINT (e.g. // Azure OpenAI (OpenAI-compatible). Set AZURE_OPENAI_ENDPOINT (e.g.
// https://<resource>.openai.azure.com), optionally AZURE_OPENAI_API_VERSION // https://<resource>.openai.azure.com), optionally AZURE_OPENAI_API_VERSION
// (default 2024-10-21), and use `azure:<your-deployment-name>` as the model. // (default 2024-10-21), and use `azure:<your-deployment-name>` as the model.
@@ -163,7 +178,20 @@ impl ChatClient {
if !key.is_empty() { if !key.is_empty() {
if azure { req = req.header("api-key", &key); } else { req = req.bearer_auth(&key); } if azure { req = req.header("api-key", &key); } else { req = req.bearer_auth(&key); }
} }
let resp = req.send().await?; let resp = req.send().await.map_err(|e| {
if e.is_connect() {
let local = matches!(p.key, "ollama" | "litellm" | "llamacpp");
if local {
anyhow!("{} connection refused at {} — is the server running? ({})", p.key, url, e)
} else {
anyhow!("{} connection error: {}", p.key, e)
}
} else if e.is_timeout() {
anyhow!("{} request timed out (120s) for model '{}' — model may be too large for available memory", p.key, m.model)
} else {
anyhow!("{} request error: {}", p.key, e)
}
})?;
let status = resp.status(); let status = resp.status();
let text = resp.text().await.unwrap_or_default(); let text = resp.text().await.unwrap_or_default();
if !status.is_success() { if !status.is_success() {
@@ -213,6 +241,11 @@ impl ChatClient {
} }
let mut cmd = Command::new(bin); let mut cmd = Command::new(bin);
// Most agentic CLIs here take the prompt on stdin; opencode and hermes
// take it as a trailing positional argument instead — track which so we
// don't also pipe it into stdin below (that would just hang the child
// waiting on a request it already got as an argv value).
let mut prompt_via_stdin = true;
match bin { match bin {
// Codex non-interactive exec (uses the ChatGPT/Codex login), prompt on stdin. // Codex non-interactive exec (uses the ChatGPT/Codex login), prompt on stdin.
"codex" => { "codex" => {
@@ -237,13 +270,46 @@ impl ChatClient {
"grok" => { "grok" => {
cmd.arg("--model").arg(model); cmd.arg("--model").arg(model);
} }
// OpenCode CLI (`opencode run`) — non-interactive one-shot, prompt
// as a positional arg, not stdin. `--auto` auto-approves anything
// not explicitly denied (our equivalent of --dangerously-skip-permissions).
// MCP (Playwright) is injected via a generated opencode.json pointed
// at through OPENCODE_CONFIG rather than a CLI flag (opencode has none).
"opencode" => {
prompt_via_stdin = false;
cmd.arg("run").arg("--model").arg(model).arg("--auto");
if let Some(mcp) = mcp_config {
match write_opencode_mcp_config(mcp) {
Ok(cfg) => { cmd.env("OPENCODE_CONFIG", cfg); }
Err(e) => eprintln!(" [!] opencode MCP config failed: {e}"),
}
}
cmd.arg(&prompt);
}
// Hermes Agent CLI (NousResearch/hermes-agent) — single-query mode.
// `-q` is the prompt-supplying flag (not a stdin read); `--provider
// nous` pins the Nous Portal OAuth login; `-Q` quiets banner/spinner
// for programmatic use; `--yolo` bypasses dangerous-command prompts.
// No CLI-level MCP hook — Hermes falls back to its own built-in
// toolsets (web/terminal/computer-use) rather than our Playwright MCP.
"hermes" => {
prompt_via_stdin = false;
cmd.arg("chat").arg("-m").arg(model).arg("--provider").arg("nous")
.arg("-Q").arg("--yolo").arg("-q").arg(&prompt);
}
_ => {} _ => {}
} }
cmd.stdin(Stdio::piped()).stdout(Stdio::piped()).stderr(Stdio::piped()).kill_on_drop(true); cmd.stdin(Stdio::piped()).stdout(Stdio::piped()).stderr(Stdio::piped()).kill_on_drop(true);
let mut child = cmd.spawn().map_err(|e| anyhow!("spawn {} failed: {}", bin, e))?; let mut child = cmd.spawn().map_err(|e| anyhow!("spawn {} failed: {}", bin, e))?;
if let Some(mut stdin) = child.stdin.take() { if prompt_via_stdin {
stdin.write_all(prompt.as_bytes()).await?; if let Some(mut stdin) = child.stdin.take() {
// Drop closes stdin so the CLI processes the prompt and exits. stdin.write_all(prompt.as_bytes()).await?;
// Drop closes stdin so the CLI processes the prompt and exits.
}
} else {
// Prompt went in as an argv value; close stdin immediately so
// nothing lingers waiting on it (opencode/hermes never read it).
drop(child.stdin.take());
} }
// Cap a single agentic CLI turn so a stuck tool-loop can't hang the run. // Cap a single agentic CLI turn so a stuck tool-loop can't hang the run.
let out = match tokio::time::timeout(Duration::from_secs(600), child.wait_with_output()).await { let out = match tokio::time::timeout(Duration::from_secs(600), child.wait_with_output()).await {
@@ -564,6 +630,8 @@ pub fn cli_binary_for(provider: &str) -> Option<&'static str> {
"openai" => Some("codex"), "openai" => Some("codex"),
"xai" => Some("grok"), "xai" => Some("grok"),
"gemini" => Some("gemini"), "gemini" => Some("gemini"),
"opencode" => Some("opencode"),
"nous" => Some("hermes"),
_ => None, _ => None,
} }
} }
@@ -577,7 +645,7 @@ pub fn binary_in_path(name: &str) -> bool {
/// Which subscription CLI backends are installed locally. /// Which subscription CLI backends are installed locally.
pub fn installed_cli_backends() -> Vec<&'static str> { pub fn installed_cli_backends() -> Vec<&'static str> {
["claude", "codex", "grok", "gemini"].into_iter().filter(|b| binary_in_path(b)).collect() ["claude", "codex", "grok", "gemini", "opencode", "hermes"].into_iter().filter(|b| binary_in_path(b)).collect()
} }
/// Login state of a subscription CLI backend. /// Login state of a subscription CLI backend.
@@ -597,15 +665,23 @@ pub async fn cli_login_status(provider: &str) -> LoginStatus {
let Some(bin) = cli_binary_for(provider) else { return LoginStatus::NotInstalled }; let Some(bin) = cli_binary_for(provider) else { return LoginStatus::NotInstalled };
if !binary_in_path(bin) { return LoginStatus::NotInstalled; } if !binary_in_path(bin) { return LoginStatus::NotInstalled; }
let mut cmd = Command::new(bin); let mut cmd = Command::new(bin);
// opencode/hermes take the probe prompt as an argv value, not stdin.
let prompt_via_stdin = !matches!(bin, "opencode" | "hermes");
match bin { match bin {
"claude" => { cmd.arg("-p").arg("--output-format").arg("text").arg("--dangerously-skip-permissions"); } "claude" => { cmd.arg("-p").arg("--output-format").arg("text").arg("--dangerously-skip-permissions"); }
"codex" => { cmd.arg("exec").arg("--dangerously-bypass-approvals-and-sandbox").arg("-"); } "codex" => { cmd.arg("exec").arg("--dangerously-bypass-approvals-and-sandbox").arg("-"); }
"opencode" => { cmd.arg("run").arg("--auto").arg("Reply with exactly: OK"); }
"hermes" => { cmd.arg("chat").arg("--provider").arg("nous").arg("-Q").arg("--yolo").arg("-q").arg("Reply with exactly: OK"); }
_ => { cmd.arg("-p"); } // grok / gemini: prompt on stdin _ => { cmd.arg("-p"); } // grok / gemini: prompt on stdin
} }
cmd.stdin(Stdio::piped()).stdout(Stdio::piped()).stderr(Stdio::piped()).kill_on_drop(true); cmd.stdin(Stdio::piped()).stdout(Stdio::piped()).stderr(Stdio::piped()).kill_on_drop(true);
let mut child = match cmd.spawn() { Ok(c) => c, Err(_) => return LoginStatus::Unknown }; let mut child = match cmd.spawn() { Ok(c) => c, Err(_) => return LoginStatus::Unknown };
if let Some(mut stdin) = child.stdin.take() { if prompt_via_stdin {
let _ = stdin.write_all(b"Reply with exactly: OK").await; if let Some(mut stdin) = child.stdin.take() {
let _ = stdin.write_all(b"Reply with exactly: OK").await;
}
} else {
drop(child.stdin.take());
} }
let out = match tokio::time::timeout(Duration::from_secs(45), child.wait_with_output()).await { let out = match tokio::time::timeout(Duration::from_secs(45), child.wait_with_output()).await {
Ok(Ok(o)) => o, Ok(Ok(o)) => o,
@@ -629,10 +705,50 @@ pub async fn cli_login_status(provider: &str) -> LoginStatus {
} }
/// Does this provider's agentic CLI accept a Playwright MCP config? /// Does this provider's agentic CLI accept a Playwright MCP config?
/// Claude Code and Codex do; Gemini/Grok CLIs don't take an MCP-config flag, so /// Claude Code, Codex, and OpenCode do (OpenCode via a generated
/// they fall back to their own built-in tools. /// `opencode.json` + `OPENCODE_CONFIG`, see `write_opencode_mcp_config`).
/// Gemini/Grok/Hermes have no CLI-level MCP hook, so they fall back to their
/// own built-in tools (Hermes ships web/terminal/computer-use natively).
pub fn mcp_supported(provider: &str) -> bool { pub fn mcp_supported(provider: &str) -> bool {
matches!(provider, "anthropic" | "openai") matches!(provider, "anthropic" | "openai" | "opencode")
}
/// Convert our `.mcp.json` (`{"mcpServers": {name: {command, args}}}`) into
/// OpenCode's own config schema (`{"mcp": {name: {"type":"local","command":
/// [command, ...args], "enabled": true}}}`) and write it next to the source
/// file. OpenCode has no `--mcp-config` flag; it's pointed at a config file
/// via the `OPENCODE_CONFIG` env var instead (set by the `opencode` arm of
/// `chat_cli`), so this doesn't touch the user's own `opencode.json`.
fn write_opencode_mcp_config(mcp_json_path: &str) -> Result<std::path::PathBuf> {
let txt = std::fs::read_to_string(mcp_json_path)
.map_err(|e| anyhow!("read {mcp_json_path}: {e}"))?;
let v: serde_json::Value = serde_json::from_str(&txt)
.map_err(|e| anyhow!("parse {mcp_json_path}: {e}"))?;
let servers = v.get("mcpServers").cloned().unwrap_or(v);
let mut mcp = serde_json::Map::new();
if let Some(obj) = servers.as_object() {
for (name, s) in obj {
let command = s.get("command").and_then(|c| c.as_str()).unwrap_or("").to_string();
if command.is_empty() { continue; }
let mut argv = vec![serde_json::Value::String(command)];
if let Some(args) = s.get("args").and_then(|a| a.as_array()) {
argv.extend(args.iter().cloned());
}
mcp.insert(name.clone(), serde_json::json!({
"type": "local",
"command": argv,
"enabled": true
}));
}
}
let cfg = serde_json::json!({
"$schema": "https://opencode.ai/config.json",
"mcp": mcp
});
let path = std::path::Path::new(mcp_json_path).with_file_name("opencode.json");
std::fs::write(&path, serde_json::to_string_pretty(&cfg).unwrap_or_default())
.map_err(|e| anyhow!("write {}: {e}", path.display()))?;
Ok(path)
} }
/// Best-effort ensure the Playwright MCP server is available locally. Requires /// Best-effort ensure the Playwright MCP server is available locally. Requires
+96 -10
View File
@@ -598,6 +598,10 @@ pub async fn run(cfg: RunConfig, lib: &Library, pool: &ModelPool, tx: Sender<Str
Ok((m, text)) => { Ok((m, text)) => {
let f = extract_findings(&text, &ag.name); let f = extract_findings(&text, &ag.name);
let _ = txc.send(format!("exploit {} via {}{} candidate(s)", ag.name, m.label(), f.len())).await; let _ = txc.send(format!("exploit {} via {}{} candidate(s)", ag.name, m.label(), f.len())).await;
if f.is_empty() && !text.trim().is_empty() && text.trim() != "[]" {
let tail: String = text.chars().rev().take(120).collect::<String>().chars().rev().collect();
let _ = txc.send(format!("⚠ agent {} returned text but 0 parseable findings (model may have produced malformed JSON). Tail: {:?}", ag.name, tail)).await;
}
// Live findings feed: surface each candidate the moment it appears. // Live findings feed: surface each candidate the moment it appears.
for c in &f { for c in &f {
let _ = txc.send(format!("finding: [{}] {} @ {}", c.severity, c.title, c.endpoint)).await; let _ = txc.send(format!("finding: [{}] {} @ {}", c.severity, c.title, c.endpoint)).await;
@@ -606,7 +610,12 @@ pub async fn run(cfg: RunConfig, lib: &Library, pool: &ModelPool, tx: Sender<Str
(ag.name.clone(), text, f) (ag.name.clone(), text, f)
} }
Err(e) => { Err(e) => {
let _ = txc.send(format!("exploit {} failed: {e}", ag.name)).await; let is_auth = crate::pool::is_auth_failure(&e);
if is_auth {
let _ = txc.send(format!("⚠ exploit {} auth failed — findings so far are SAFE, run is pausing: {e}", ag.name)).await;
} else {
let _ = txc.send(format!("exploit {} failed: {e}", ag.name)).await;
}
(ag.name.clone(), format!("ERROR: {e}"), vec![]) (ag.name.clone(), format!("ERROR: {e}"), vec![])
} }
} }
@@ -619,6 +628,9 @@ pub async fn run(cfg: RunConfig, lib: &Library, pool: &ModelPool, tx: Sender<Str
let transcript = transcript_of(&raw); let transcript = transcript_of(&raw);
let candidates = dedup_findings(raw.iter().flat_map(|(_, _, f)| f.clone()).collect()); let candidates = dedup_findings(raw.iter().flat_map(|(_, _, f)| f.clone()).collect());
let _ = tx.send(format!("{} candidate finding(s) (deduped) — validating by {}-model vote", candidates.len(), cfg.vote_n)).await; let _ = tx.send(format!("{} candidate finding(s) (deduped) — validating by {}-model vote", candidates.len(), cfg.vote_n)).await;
if pool.candidates.len() == 1 && cfg.vote_n <= 1 {
let _ = tx.send("⚠ single-model panel with vote_n=1 — validation is weaker (same model validates its own findings). Consider --vote-n 2 or adding a second model for cross-validation.".into()).await;
}
// ---- 4. Validate by N-model voting --------------------------------- // ---- 4. Validate by N-model voting ---------------------------------
let mut findings = validate(candidates, pool, VOTE_SYS, cfg.vote_n, &tx).await; let mut findings = validate(candidates, pool, VOTE_SYS, cfg.vote_n, &tx).await;
@@ -1148,9 +1160,26 @@ fn heuristic_select(ranked: &[Agent], recon: &str, focus: &str, cap: usize) -> V
} }
async fn validate(candidates: Vec<Finding>, pool: &ModelPool, sys: &str, vote_n: usize, tx: &Sender<String>) -> Vec<Finding> { async fn validate(candidates: Vec<Finding>, pool: &ModelPool, sys: &str, vote_n: usize, tx: &Sender<String>) -> Vec<Finding> {
// Fast-track: findings with no evidence are unverifiable — skip the vote
// and flag for human review instead of wasting a validator call that will
// always reject ("default to rejected when uncertain" + empty evidence).
let (have_evidence, no_evidence): (Vec<_>, Vec<_>) = candidates.into_iter().partition(|f| {
let e = f.evidence.trim();
!e.is_empty() && e != "N/A" && e != "n/a" && e != "none" && e != "-"
});
let mut flagged: Vec<Finding> = no_evidence.into_iter().map(|mut f| {
f.validated = false;
f.review_status = "needs-review".into();
f.review_reason = "no concrete evidence provided by agent — manual verification required".into();
f.votes = "0/0".into();
f
}).collect();
for f in &flagged {
let _ = tx.send(format!("vote {} → needs-review (no evidence)", f.title)).await;
}
// Prefer a model other than the primary (likely finder) to adjudicate. // Prefer a model other than the primary (likely finder) to adjudicate.
let finder = pool.candidates.first().map(|m| m.label()); let finder = pool.candidates.first().map(|m| m.label());
let validated: Vec<Finding> = stream::iter(candidates) let validated: Vec<Finding> = stream::iter(have_evidence)
.map(|mut f| { .map(|mut f| {
let txc = tx.clone(); let txc = tx.clone();
let finder = finder.clone(); let finder = finder.clone();
@@ -1184,7 +1213,9 @@ async fn validate(candidates: Vec<Finding>, pool: &ModelPool, sys: &str, vote_n:
.collect() .collect()
.await; .await;
// Keep confirmed AND needs-review (human decides); drop only zero-support noise. // Keep confirmed AND needs-review (human decides); drop only zero-support noise.
validated.into_iter().filter(|f| f.validated || f.review_status == "needs-review").collect() // Include no-evidence flagged findings so the human loop sees them.
flagged.extend(validated.into_iter().filter(|f| f.validated || f.review_status == "needs-review"));
flagged
} }
/// Adversarial refutation pass: every confirmed **High/Critical** finding is /// Adversarial refutation pass: every confirmed **High/Critical** finding is
@@ -1436,12 +1467,27 @@ fn extract_findings(text: &str, agent: &str) -> Vec<Finding> {
(Some(a), Some(b)) if b > a => &text[a..=b], (Some(a), Some(b)) if b > a => &text[a..=b],
_ => match (text.find('{'), text.rfind('}')) { _ => match (text.find('{'), text.rfind('}')) {
(Some(a), Some(b)) if b > a => &text[a..=b], (Some(a), Some(b)) if b > a => &text[a..=b],
_ => return vec![], _ => {
if !text.trim().is_empty() && text.trim() != "[]" {
eprintln!("[extract_findings] agent {agent}: model returned text but no JSON array/object found (len={}); raw tail: {:?}",
text.len(), &text[text.len().saturating_sub(200)..]);
}
return vec![];
}
}, },
}; };
let val: serde_json::Value = match serde_json::from_str(slice) { let val: serde_json::Value = match serde_json::from_str(slice) {
Ok(v) => v, Ok(v) => v,
Err(_) => return vec![], Err(e) => {
eprintln!("[extract_findings] agent {agent}: JSON parse failed: {e}; slice head: {:?}",
&slice[..slice.len().min(300)]);
// Attempt to salvage: strip trailing comma before ] (common LLM mistake)
let fixed = slice.replace(",]", "]").replace(",}", "}");
match serde_json::from_str(&fixed) {
Ok(v) => v,
Err(_) => return vec![],
}
}
}; };
let items: Vec<serde_json::Value> = match val { let items: Vec<serde_json::Value> = match val {
serde_json::Value::Array(a) => a, serde_json::Value::Array(a) => a,
@@ -1842,6 +1888,11 @@ fn recon_intensity_directive(level: usize) -> String {
(8) TLS/headers/cookies. Report counts (how many subdomains/urls/params/endpoints you actually found).\n\n") (8) TLS/headers/cookies. Report counts (how many subdomains/urls/params/endpoints you actually found).\n\n")
} }
/// Max wall-clock seconds for the ENTIRE recon phase (all rounds combined).
/// This prevents recon from eating the whole run — exploitation must start.
/// Per-round cap = total / (rounds + 1) so later rounds get equal time.
const RECON_TOTAL_BUDGET_SECS: u64 = 300; // 5 minutes total
/// Intense, multi-round recon: an initial deep pass, then follow-up rounds that /// Intense, multi-round recon: an initial deep pass, then follow-up rounds that
/// EXPAND the surface (chase discovered subdomains/endpoints/params, install /// EXPAND the surface (chase discovered subdomains/endpoints/params, install
/// tools, dig where the previous round found signal). Returns the merged recon /// tools, dig where the previous round found signal). Returns the merged recon
@@ -1853,22 +1904,53 @@ async fn deep_recon(cfg: &RunConfig, pool: &ModelPool, probe_facts: &str, tx: &S
let intensity_dir = recon_intensity_directive(intensity); let intensity_dir = recon_intensity_directive(intensity);
let dir = operator_directives(cfg); let dir = operator_directives(cfg);
let mut accum = format!("OBSERVED HTTP PROBE:\n{probe_facts}"); let mut accum = format!("OBSERVED HTTP PROBE:\n{probe_facts}");
let recon_start = std::time::Instant::now();
let total_rounds = 1 + extra_rounds;
// Time-budget directive: subscription CLIs run commands autonomously, so they
// need an explicit cap to avoid running 150+ commands in a single round.
let budget_dir = format!(
"TIME BUDGET: you have ~{budget_secs} seconds for THIS recon round. Be EFFICIENT: \
prioritise high-signal actions (JS analysis, API mapping, SQLi/auth probes) over \
exhaustive crawling. AIM for 30-50 commands max per round enough to map the surface, \
not so many that exploitation never starts. STOP EARLY if you have enough intel to \
select agents. When done, EMIT YOUR RESULTS IMMEDIATELY do not start another pass.\n\n",
budget_secs = RECON_TOTAL_BUDGET_SECS / total_rounds as u64,
);
// Initial deep pass. // Initial deep pass.
let user = format!("{dir}{intensity_dir}{doctrine}OBSERVED HTTP PROBE (build on these, verify, go deeper):\n{probe_facts}\n\nTarget: {}", cfg.target); let user = format!("{dir}{budget_dir}{intensity_dir}{doctrine}OBSERVED HTTP PROBE (build on these, verify, go deeper):\n{probe_facts}\n\nTarget: {}", cfg.target);
let _ = tx.send(format!("recon: intensity {} — actively enumerating (installing tools as needed)…", intensity)).await; let _ = tx.send(format!("recon: intensity {} — actively enumerating (budget {}s total, {} round(s))…", intensity, RECON_TOTAL_BUDGET_SECS, total_rounds)).await;
match pool.complete_routed(Task::Recon, "recon", RECON_SYS, &user).await { match pool.complete_routed(Task::Recon, "recon", RECON_SYS, &user).await {
Ok((m, t)) => { let _ = tx.send(format!("recon round 1 complete via {}", m.label())).await; accum.push_str(&format!("\n\nMODEL RECON (round 1):\n{t}")); } Ok((m, t)) => { let _ = tx.send(format!("recon round 1 complete via {}", m.label())).await; accum.push_str(&format!("\n\nMODEL RECON (round 1):\n{t}")); }
Err(e) => { let _ = tx.send(format!("recon round 1 failed ({e}) — probe facts only")).await; return accum; } Err(e) => {
let is_auth = crate::pool::is_auth_failure(&e);
if is_auth {
let _ = tx.send(format!("recon round 1 auth failed ({e}) — continuing with probe facts; run will pause before exploit phase")).await;
accum.push_str(&format!("\n\nMODEL RECON (round 1):\n{e}"));
} else {
let _ = tx.send(format!("recon round 1 failed ({e}) — probe facts only")).await;
}
return accum;
}
} }
// Follow-up expansion rounds — each digs further using what's known so far. // Follow-up expansion rounds — each digs further using what's known so far.
for r in 0..extra_rounds { for r in 0..extra_rounds {
if pool.stop_exploiting() { break; } if pool.stop_exploiting() { break; }
// Time budget check: if recon has already consumed the total budget, stop
// and proceed to exploitation with whatever intelligence we gathered.
let elapsed = recon_start.elapsed().as_secs();
if elapsed >= RECON_TOTAL_BUDGET_SECS {
let _ = tx.send(format!("recon: time budget exhausted ({elapsed}s/{RECON_TOTAL_BUDGET_SECS}s) — proceeding to exploitation with current intel")).await;
break;
}
let remaining = RECON_TOTAL_BUDGET_SECS - elapsed;
let round = r + 2; let round = r + 2;
let known: String = accum.chars().rev().take(3000).collect::<String>().chars().rev().collect(); let known: String = accum.chars().rev().take(3000).collect::<String>().chars().rev().collect();
let follow = format!( let follow = format!(
"{dir}{intensity_dir}{doctrine}CONTINUE the recon — this is round {round}. Here is what recon has found so far:\n{known}\n\n\ "{dir}TIME BUDGET: you have ~{remaining} seconds remaining for recon. Be CONCISE — focus only on the highest-value leads.\n\n\
{intensity_dir}{doctrine}CONTINUE the recon this is round {round}. Here is what recon has found so far:\n{known}\n\n\
Now EXPAND: pick the most promising leads and go deeper resolve & probe any NEW subdomains/hosts, crawl \ Now EXPAND: pick the most promising leads and go deeper resolve & probe any NEW subdomains/hosts, crawl \
and harvest URLs for endpoints not yet mapped, run content/parameter discovery where you saw interesting \ and harvest URLs for endpoints not yet mapped, run content/parameter discovery where you saw interesting \
paths, fingerprint exact versions of anything unclear, and enumerate the API/GraphQL further. Install any \ paths, fingerprint exact versions of anything unclear, and enumerate the API/GraphQL further. Install any \
@@ -1880,7 +1962,11 @@ async fn deep_recon(cfg: &RunConfig, pool: &ModelPool, probe_facts: &str, tx: &S
if novel.len() > 20 { let _ = tx.send(format!("recon round {round} via {} — expanded surface", m.label())).await; accum.push_str(&format!("\n\nMODEL RECON (round {round}):\n{novel}")); } if novel.len() > 20 { let _ = tx.send(format!("recon round {round} via {} — expanded surface", m.label())).await; accum.push_str(&format!("\n\nMODEL RECON (round {round}):\n{novel}")); }
else { let _ = tx.send(format!("recon round {round}: no new surface — recon converged")).await; break; } else { let _ = tx.send(format!("recon round {round}: no new surface — recon converged")).await; break; }
} }
Err(e) => { let _ = tx.send(format!("recon round {round} failed ({e})")).await; break; } Err(e) => {
let is_auth = crate::pool::is_auth_failure(&e);
let _ = tx.send(format!("recon round {round} {} ({e})", if is_auth { "auth failed" } else { "failed" })).await;
break;
}
} }
} }
accum accum
+85 -19
View File
@@ -20,6 +20,24 @@ pub fn is_exhaustion(e: &anyhow::Error) -> bool {
.any(|k| s.contains(k)) .any(|k| s.contains(k))
} }
/// Does this error look like an **authentication / authorization failure**
/// (revoked OAuth, expired session, invalid API key) — distinct from transient
/// quota/rate issues? Auth failures are non-recoverable without re-login, so
/// the run should pause immediately and offer fallback providers.
pub fn is_auth_failure(e: &anyhow::Error) -> bool {
let s = format!("{e:#}").to_lowercase();
[
"401", "403", "unauthorized", "token has been revoked",
"token revoked", "access token", "oauth", "session expired",
"not authenticated", "not logged in", "please log in",
"please login", "invalid api key", "invalid_api_key",
"api key expired", "authentication failed", "failed to authenticate",
"run /login",
]
.iter()
.any(|k| s.contains(k))
}
/// Task type used by the model router to pick the best model for the step. /// Task type used by the model router to pick the best model for the step.
#[derive(Clone, Copy, Debug)] #[derive(Clone, Copy, Debug)]
pub enum Task { pub enum Task {
@@ -64,6 +82,10 @@ pub struct ModelPool {
/// Fallback models the user added via `/continue <provider:model>` while /// Fallback models the user added via `/continue <provider:model>` while
/// paused — tried first on the next attempt. /// paused — tried first on the next attempt.
fallback: Arc<Mutex<Vec<ModelRef>>>, fallback: Arc<Mutex<Vec<ModelRef>>>,
/// Circuit breaker: consecutive auth/exhaustion failures across agents.
/// When this exceeds `AUTH_FAIL_THRESHOLD`, the pool auto-pauses instead of
/// burning through the remaining agents on a dead token.
consecutive_auth_fails: Arc<std::sync::atomic::AtomicUsize>,
} }
impl ModelPool { impl ModelPool {
@@ -96,9 +118,20 @@ impl ModelPool {
paused: Arc::new(AtomicBool::new(false)), paused: Arc::new(AtomicBool::new(false)),
resume: Arc::new(Notify::new()), resume: Arc::new(Notify::new()),
fallback: Arc::new(Mutex::new(Vec::new())), fallback: Arc::new(Mutex::new(Vec::new())),
consecutive_auth_fails: Arc::new(std::sync::atomic::AtomicUsize::new(0)),
} }
} }
/// Reset the consecutive auth-failure counter (called on any successful completion).
fn reset_auth_fails(&self) {
self.consecutive_auth_fails.store(0, std::sync::atomic::Ordering::Relaxed);
}
/// Increment the consecutive auth-failure counter and return the new count.
fn inc_auth_fails(&self) -> usize {
self.consecutive_auth_fails.fetch_add(1, std::sync::atomic::Ordering::Relaxed) + 1
}
/// Attach a progress channel so the subscription CLI streams structured /// Attach a progress channel so the subscription CLI streams structured
/// activity (commands run, files read, tools called) live. /// activity (commands run, files read, tools called) live.
pub fn set_progress(&self, tx: tokio::sync::mpsc::Sender<String>) { pub fn set_progress(&self, tx: tokio::sync::mpsc::Sender<String>) {
@@ -146,20 +179,32 @@ impl ModelPool {
self.paused.load(Ordering::Relaxed) self.paused.load(Ordering::Relaxed)
} }
/// Park the run on token/quota exhaustion: keep ALL state, emit a notice, /// Consecutive auth/quota failures needed to trip the circuit breaker and
/// and wait until the user runs `/continue` (or cancels). Returns when the /// auto-pause the run. Low threshold: 3 consecutive failures on the same
/// run should retry (pause cleared) or give up (cancelled). /// provider is enough signal that the token is dead.
async fn park_exhausted(&self, err: &anyhow::Error) { const AUTH_FAIL_THRESHOLD: usize = 3;
/// Park the run on token/quota exhaustion or auth failure: keep ALL state,
/// emit a notice, and wait until the user runs `/continue` (or cancels).
/// Returns when the run should retry (pause cleared) or give up (cancelled).
async fn park_exhausted(&self, err: &anyhow::Error, is_auth: bool) {
self.paused.store(true, Ordering::Relaxed); self.paused.store(true, Ordering::Relaxed);
if let Some(tx) = self.progress() { if let Some(tx) = self.progress() {
let msg = format!("{err:#}"); let msg = format!("{err:#}");
let short = msg.lines().next().unwrap_or(&msg); let short = msg.lines().next().unwrap_or(&msg);
let _ = tx let notice = if is_auth {
.send(format!( format!(
"notify: ⏸ authentication failed ({}). Run is PAUSED — all findings so far are SAFE. \
Fix: /continue <provider:model> to switch provider, or re-login and /continue.",
short.chars().take(120).collect::<String>()
)
} else {
format!(
"notify: ⏸ token/quota exhausted ({}). Run is PAUSED — type /continue when your quota renews, or switch with /model <provider:model> then /continue.", "notify: ⏸ token/quota exhausted ({}). Run is PAUSED — type /continue when your quota renews, or switch with /model <provider:model> then /continue.",
short.chars().take(120).collect::<String>() short.chars().take(120).collect::<String>()
)) )
.await; };
let _ = tx.send(notice).await;
} }
while self.paused.load(Ordering::Relaxed) && !self.is_cancelled() { while self.paused.load(Ordering::Relaxed) && !self.is_cancelled() {
let notified = self.resume.notified(); let notified = self.resume.notified();
@@ -169,8 +214,9 @@ impl ModelPool {
} }
} }
if !self.is_cancelled() { if !self.is_cancelled() {
self.reset_auth_fails(); // user resumed, reset counter
if let Some(tx) = self.progress() { if let Some(tx) = self.progress() {
let _ = tx.send("notify: ▶ resumed — retrying exhausted step.".to_string()).await; let _ = tx.send("notify: ▶ resumed — retrying with updated credentials/model.".to_string()).await;
} }
} }
} }
@@ -212,9 +258,9 @@ impl ModelPool {
}; };
match r { match r {
Ok(t) => return Ok(t), Ok(t) => return Ok(t),
// Don't burn retries on exhaustion — surface it so the caller // Don't burn retries on exhaustion or auth failure — surface
// can park and let the user /continue. // immediately so the caller can park and let the user /continue.
Err(e) if is_exhaustion(&e) => return Err(e), Err(e) if is_exhaustion(&e) || is_auth_failure(&e) => return Err(e),
Err(e) => last = e, Err(e) => last = e,
} }
} }
@@ -233,6 +279,20 @@ impl ModelPool {
if self.is_cancelled() { if self.is_cancelled() {
return Err(anyhow!("cancelled")); return Err(anyhow!("cancelled"));
} }
// Circuit breaker: if we've seen N consecutive auth failures across
// agents, pause immediately — don't burn another agent on a dead token.
let fail_count = self.consecutive_auth_fails.load(std::sync::atomic::Ordering::Relaxed);
if fail_count >= Self::AUTH_FAIL_THRESHOLD && !self.is_cancelled() {
self.park_exhausted(
&anyhow!("circuit breaker: {} consecutive auth failures — token/session likely dead", fail_count),
true,
).await;
if self.is_cancelled() {
return Err(anyhow!("cancelled"));
}
// After resume, retry with potentially new fallback models.
continue;
}
// User-supplied fallback models (via /continue) are tried first. // User-supplied fallback models (via /continue) are tried first.
let mut order = self.route(task); let mut order = self.route(task);
if let Ok(fb) = self.fallback.lock() { if let Ok(fb) = self.fallback.lock() {
@@ -244,25 +304,31 @@ impl ModelPool {
} }
let mut last = anyhow!("no candidate models"); let mut last = anyhow!("no candidate models");
let mut exhausted = false; let mut exhausted = false;
let mut auth_failed = false;
for m in &order { for m in &order {
if self.is_cancelled() { if self.is_cancelled() {
return Err(anyhow!("cancelled")); return Err(anyhow!("cancelled"));
} }
match self.one(label, m, system, user).await { match self.one(label, m, system, user).await {
Ok(text) => return Ok((m.clone(), text)), Ok(text) => {
self.reset_auth_fails(); // success resets circuit breaker
return Ok((m.clone(), text));
}
Err(e) => { Err(e) => {
if is_exhaustion(&e) { if is_auth_failure(&e) {
auth_failed = true;
self.inc_auth_fails();
} else if is_exhaustion(&e) {
exhausted = true; exhausted = true;
} }
last = e; last = e;
} }
} }
} }
// Every candidate failed. If it was token/quota exhaustion, park the // Every candidate failed. Park the run (keeping all state) so the user
// run until the user runs /continue, then retry the whole order (now // can fix auth or wait for quota renewal, then /continue.
// including any fallback model they added). Otherwise, give up. if (auth_failed || exhausted) && !self.is_cancelled() {
if exhausted && !self.is_cancelled() { self.park_exhausted(&last, auth_failed).await;
self.park_exhausted(&last).await;
continue; continue;
} }
return Err(last); return Err(last);
+2 -2
View File
@@ -28,7 +28,7 @@ cat <<'BANNER'
███╗ ██╗███████╗██╗ ██╗██████╗ ██████╗ ███╗ ██╗███████╗██╗ ██╗██████╗ ██████╗
████╗ ██║██╔════╝██║ ██║██╔══██╗██╔═══██╗ NeuroSploit installer ████╗ ██║██╔════╝██║ ██║██╔══██╗██╔═══██╗ NeuroSploit installer
██╔██╗ ██║█████╗ ██║ ██║██████╔╝██║ ██║ v3.6.1 — Rust harness ██╔██╗ ██║█████╗ ██║ ██║██████╔╝██║ ██║ v3.6.9 — Rust harness
██║╚██╗██║██╔══╝ ██║ ██║██╔══██╗██║ ██║ by Joas A Santos ██║╚██╗██║██╔══╝ ██║ ██║██╔══██╗██║ ██║ by Joas A Santos
██║ ╚████║███████╗╚██████╔╝██║ ██║╚██████╔╝ & Red Team Leaders ██║ ╚████║███████╗╚██████╔╝██║ ██║╚██████╔╝ & Red Team Leaders
╚═╝ ╚═══╝╚══════╝ ╚═════╝ ╚═╝ ╚═╝ ╚═════╝ ╚═╝ ╚═══╝╚══════╝ ╚═════╝ ╚═╝ ╚═╝ ╚═════╝
@@ -63,7 +63,7 @@ if [ -z "$REF" ]; then
REF="$(dl "https://api.github.com/repos/${REPO_SLUG}/releases/latest" /dev/stdout 2>/dev/null \ REF="$(dl "https://api.github.com/repos/${REPO_SLUG}/releases/latest" /dev/stdout 2>/dev/null \
| grep -m1 '"tag_name"' | sed -E 's/.*"tag_name" *: *"([^"]+)".*/\1/' || true)" | grep -m1 '"tag_name"' | sed -E 's/.*"tag_name" *: *"([^"]+)".*/\1/' || true)"
fi fi
[ -z "$REF" ] && REF="v3.6.1" [ -z "$REF" ] && REF="v3.6.9"
say "Release: $REF" say "Release: $REF"
installed=0 installed=0