mirror of
https://github.com/CyberSecurityUP/NeuroSploit.git
synced 2026-10-08 09:01:11 +02:00
35bf7ea1a6db64ca17b5cb95517faedad808e4ef
3
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
01eac8f0c3 |
fix: robust LLM response handling & JSON extraction (#46)
* fix(pipeline): robust LLM JSON extraction (json5 + truncation repair)
Model replies that did not exactly match the expected JSON syntax were
either dropped silently or surfaced as "[extract_findings] ... JSON parse
failed" / "... no JSON array/object found". Both came from the same two
weak stages in extract_findings: a greedy first-'['-to-last-']' span that
captured prose, and a salvage pass that only stripped trailing commas.
Add a shared, string/escape-aware extractor (crates/harness/src/json_extract.rs):
- locate balanced [..]/{..} regions, ignoring brackets inside prose/strings,
preferring fenced blocks (last wins);
- parse leniently: serde_json first, then json5 (trailing commas, comments,
single quotes, unquoted keys);
- repair token-limit truncation by closing the open structure, keeping the
complete findings instead of discarding the whole batch.
Route extract_findings, reported_nothing, extract_chain, parse_string_array
and prosecutor::parse_verdict through it. Make the diagnostic tail()
char-boundary-safe (the old slice could panic on UTF-8). Add regression
tests for single quotes/comments, capitalised ```JSON fences, truncated
arrays and pure prose.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix(models): robust LLM response handling + higher token/timeout limits
Harden the OpenAI-compatible chat client against the empty-content and
parse failures hit with reasoning models (GLM/DeepSeek via OpenRouter)
during whitebox runs:
- Accept message `content` as a string, an array of content parts, or a
`reasoning_content` fallback; surface `finish_reason` and empty-choices
errors instead of an opaque "no content in response".
- Stop masking mid-stream body-read failures as a bogus "EOF while
parsing"; report read timeouts and empty bodies explicitly, and
reassemble SSE-framed responses some gateways return unrequested.
- Raise reasoning-model max_tokens to 32768 and the HTTP timeout to 300s;
both overridable via NEUROSPLOIT_MAX_TOKENS / NEUROSPLOIT_HTTP_TIMEOUT.
Adds unit tests for content extraction, token sizing, and SSE reassembly.
Cargo.lock syncs the json5 entry from the prior extraction commit.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix(pipeline): unwrap findings/selection replies wrapped in an object
The json5 + truncation work made JSON *parsing* robust, but the *shape*
handling after it still dropped data when a model wrapped its answer in an
object instead of returning the bare array we asked for. Most visible on
black-box runs, where a long tool-use turn ends with the model narrating
into a report object.
extract_findings treated any object as ONE finding, so a real batch returned
as `{"findings":[…]}` (or `{"vulnerabilities":[…]}`, …) became a single
title-less "finding", was filtered out, and surfaced to the operator as
"returned text but 0 parseable findings" while the findings were lost. Add
findings_items() to normalise the shape: an array is the list; an object with
a title is one bare finding; otherwise an object wrapping a known findings key
unwraps to that array. reported_nothing() now recognises the same wrapper keys
so an empty `{"vulnerabilities":[]}` reads as an honest negative.
Two more consumers of the same class:
- parse_string_array (agent selection) accepted only a bare array of strings,
so a wrapped `{"agents":[…]}` or elements-as-objects `[{"name":"sqli"}]`
silently fell back to RL ranking. Now unwraps the wrapper and pulls the
string from object elements.
- extract_chain hard-coded the "findings" key for its object branch, dropping
the sibling `loot` under any other wrapper key. Now checks all wrapper keys.
Add regression tests for each shape.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
|
||
|
|
3456c32f4d |
feat(harness): risk model, engagement policies, capability tokens, audit trail
Four pieces that together answer "may this action happen, under whose authority, and can we prove afterwards what we did". policy.rs — effective_risk per action, exactly as specified: (action_risk + asset_criticality + protocol_risk + privilege_level + blast_radius) × environment_multiplier Every term is named and kept on the result, so the number can be explained rather than argued with. Three policies sit on it: SafetyPolicy (ceilings, approval thresholds, hard prohibitions), ReasoningPolicy (baseline before payload, bounded hypotheses, evidence before escalation, explicit stop conditions) and ProofOfImpactPolicy (what a severity must carry before it may be that severity). OT/ICS/SCADA is treated as its own regime, not web testing on odd ports. Industrial protocols authenticate nothing — a Modbus write is the protocol working as intended, addressed to a device that may be holding a valve — and scanners crash PLCs by sending unexpected data at line rate. So the OT profile blocks writes, disruptive actions, fuzzing and exploit payloads outright, caps the rate at ~1 req/s, and refuses the function codes that stop a CPU (Modbus 5/6/8/15/16/22/23/43, S7 start/stop, DNP3 restart/stop). Safety instrumented systems are off limits in every profile. A test caught a calibration error worth keeping: a plain READ of a critical PLC scores 3.6 on this formula, so the obvious tight ceiling would have refused exactly the observation OT findings come from. In an industrial environment it is the KIND of action that is forbidden, not the arithmetic — the ceiling catches extremes and the low approval threshold makes anything past trivial observation a human's decision. capability.rs — HMAC-signed grants: who authorized what, against which hosts, in which environment, until when. The harness verifies the signature before reading a single claim (a well-formed token from the wrong key must never get to influence what the harness believes), refuses expired and not-yet-valid tokens, and treats the grant as a CEILING: constrain() intersects it with local configuration, so config can narrow authorization and never widen it. Tokens carry no secrets — the payload is readable by anyone holding it. audit.rs — one structured record per action, in the specified shape (timestamp, agent, hypothesis, action, target, policy_decision, operator, tool, result, evidence_hash, capability_token). Two things make it worth having: it is hash-chained, so removing or editing an entry breaks every hash that follows and verify() says which one; and it records REFUSALS, because a trail containing only what happened cannot demonstrate restraint. Only the grant's id is recorded, never the token — the trail gets shared. Hard kill conditions end a run outright: target unresponsive after our traffic, sustained 5xx, out-of-scope request, forbidden industrial function code, safety system addressed, capability expired mid-run, repeated policy violations, budget exhausted, operator stop. Failures BEFORE the target ever answered do not count — nothing listening is not the same as knocked over. The OT switch trips far sooner: a PLC missing two requests already warrants stopping. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
56d3f0c723 |
NeuroSploit v3.4.0 — Rust multi-model harness + Axum dashboard
New cargo workspace `neurosploit-rs/` (single `neurosploit` binary): harness crate: - models.rs: 11 OpenAI-compatible providers / 31 models (Claude, GPT, Grok, NVIDIA NIM, DeepSeek, Mistral, Qwen, Groq, Together, OpenRouter, Ollama) - pool.rs: ModelPool with bounded concurrency, provider failover, and N-model validator voting (the panel doubles as the jury) - agents.rs: loads the existing agents_md/ library (213 agents) - pipeline.rs: recon → parallel exploit (semaphore-bounded) → N-model adversarial vote → score; streams live progress over a channel - report.rs: HTML report - tokio + reqwest(rustls); offline mode runs the pipeline without API keys app binary: - clap CLI: serve | run | agents | models (run supports --model x N, --vote-n, --max-agents, --offline) - axum web dashboard with multi-model panel, live console, findings, agent browser, embedded report; single binary serves the SPA (no npm/build) Verified: cargo build clean; agents/models/offline-run CLI; server endpoints (/api/info, /api/run lifecycle, /report); dashboard + live run in Playwright. Docs: README v3.4.0 callout + RELEASE.md notes. target/ gitignored. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |