mirror of
https://github.com/CyberSecurityUP/NeuroSploit.git
synced 2026-10-03 22:46:57 +02:00
* fix(pipeline): robust LLM JSON extraction (json5 + truncation repair)
Model replies that did not exactly match the expected JSON syntax were
either dropped silently or surfaced as "[extract_findings] ... JSON parse
failed" / "... no JSON array/object found". Both came from the same two
weak stages in extract_findings: a greedy first-'['-to-last-']' span that
captured prose, and a salvage pass that only stripped trailing commas.
Add a shared, string/escape-aware extractor (crates/harness/src/json_extract.rs):
- locate balanced [..]/{..} regions, ignoring brackets inside prose/strings,
preferring fenced blocks (last wins);
- parse leniently: serde_json first, then json5 (trailing commas, comments,
single quotes, unquoted keys);
- repair token-limit truncation by closing the open structure, keeping the
complete findings instead of discarding the whole batch.
Route extract_findings, reported_nothing, extract_chain, parse_string_array
and prosecutor::parse_verdict through it. Make the diagnostic tail()
char-boundary-safe (the old slice could panic on UTF-8). Add regression
tests for single quotes/comments, capitalised ```JSON fences, truncated
arrays and pure prose.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix(models): robust LLM response handling + higher token/timeout limits
Harden the OpenAI-compatible chat client against the empty-content and
parse failures hit with reasoning models (GLM/DeepSeek via OpenRouter)
during whitebox runs:
- Accept message `content` as a string, an array of content parts, or a
`reasoning_content` fallback; surface `finish_reason` and empty-choices
errors instead of an opaque "no content in response".
- Stop masking mid-stream body-read failures as a bogus "EOF while
parsing"; report read timeouts and empty bodies explicitly, and
reassemble SSE-framed responses some gateways return unrequested.
- Raise reasoning-model max_tokens to 32768 and the HTTP timeout to 300s;
both overridable via NEUROSPLOIT_MAX_TOKENS / NEUROSPLOIT_HTTP_TIMEOUT.
Adds unit tests for content extraction, token sizing, and SSE reassembly.
Cargo.lock syncs the json5 entry from the prior extraction commit.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix(pipeline): unwrap findings/selection replies wrapped in an object
The json5 + truncation work made JSON *parsing* robust, but the *shape*
handling after it still dropped data when a model wrapped its answer in an
object instead of returning the bare array we asked for. Most visible on
black-box runs, where a long tool-use turn ends with the model narrating
into a report object.
extract_findings treated any object as ONE finding, so a real batch returned
as `{"findings":[…]}` (or `{"vulnerabilities":[…]}`, …) became a single
title-less "finding", was filtered out, and surfaced to the operator as
"returned text but 0 parseable findings" while the findings were lost. Add
findings_items() to normalise the shape: an array is the list; an object with
a title is one bare finding; otherwise an object wrapping a known findings key
unwraps to that array. reported_nothing() now recognises the same wrapper keys
so an empty `{"vulnerabilities":[]}` reads as an honest negative.
Two more consumers of the same class:
- parse_string_array (agent selection) accepted only a bare array of strings,
so a wrapped `{"agents":[…]}` or elements-as-objects `[{"name":"sqli"}]`
silently fell back to RL ranking. Now unwraps the wrapper and pulls the
string from object elements.
- extract_chain hard-coded the "findings" key for its object branch, dropping
the sibling `loot` under any other wrapper key. Now checks all wrapper keys.
Add regression tests for each shape.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>