Files
NeuroSploit/neurosploit-rs/crates
CyberSecurityUPandClaude Opus 5 b4da49fae4 fix(harness): stop discarding findings when an agent narrates before answering
Both bugs surfaced during a live black-box engagement.

1. extract_findings took the span from the first '[' to the last ']'. One
   agent's reply opened with the prose line "[low] Antiforgery cookie missing
   Secure flag" and put its real findings in a fenced ```json block further
   down — so the span started inside prose, failed to parse, and every finding
   that agent had proven was thrown away. Fenced blocks are now tried first
   (last one wins: models narrate, then answer), with the span kept only as a
   fallback and the trailing-comma salvage preserved.

2. "no findings" was reported as malformed JSON. The guard compared the raw
   text to "[]", but models wrap the empty array in a fence, so every honest
   negative result was logged as a parse failure — which teaches an operator to
   ignore a warning that sometimes means a real one. reported_nothing() now
   recognises a bare [], a fenced [], and {"findings": []}.

The second bug made the first one harder to see: the log was already full of
"malformed JSON" warnings for agents that had simply found nothing, so the one
warning that meant a genuine loss looked like more of the same.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-13 19:46:20 -03:00
..