mirror of
https://github.com/CyberSecurityUP/NeuroSploit.git
synced 2026-09-29 20:41:51 +02:00
A live black-box engagement returned 21 "deduped" candidates for about ten actual issues. One cookie problem came back five times and one missing header four, because the key was `cwe|endpoint|title[..40]` and every agent writes those differently: `CWE-614` vs `CWE-614 (Sensitive Cookie in HTTPS Session Without Secure Attribute)`, `https://host/` vs `GET https://host/ (and /Account/Login)`. Worse than the inflated count, the severities disagreed — the same issue arrived Low from one agent and Medium from another, which is indefensible in front of a client. The key is now (CWE number, normalized endpoint) plus a title-similarity check, because grouping on the first two alone over-merges: missing `nosniff`, `Referrer-Policy` and `Permissions-Policy` are all CWE-693 on `/` and are three separate fixes. Titles merge at Jaccard >= 0.4 over meaningful words — calibrated on this run's real output, where two phrasings of the cookie issue score 0.44 and the two header findings score 0.33. The survivor keeps the HIGHEST severity with the fullest evidence, and inherits whatever the duplicates knew that it did not (remediation, repro steps, structured evidence). Agreement between independent agents is recorded as "corroborated by …" and nudges confidence up: several agents reaching the same conclusion separately is a reason to trust a finding, not a reason to print it five times. Tests use the actual titles, CWEs and endpoints from the engagement, including the case that must NOT merge (three rate-limit findings on three different endpoints — login spraying and reset-email flooding are different problems). Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>