Files
NeuroSploit/neurosploit-rs
CyberSecurityUPandClaude Opus 5 36c9e05ea7 feat(uncertainty): collect more evidence instead of handing a human a thin verdict
Everything upstream produced a verdict and stopped. When the evidence was thin
the answer was needs-review — which hands a reviewer the same thin evidence and
asks them to do the collecting. That is the wrong party: the harness still has
the target, the session and the tooling; the reviewer has a paragraph.

The uncertainty engine scores how undecided each candidate is by counting what
its class needs against what it has, which makes "how sure are we?" a
measurement rather than another model's opinion. A finding that is undecided
AND whose gap is obtainable gets one more collection round before judgement.

Two rules keep the loop from becoming a treadmill:
- Only obtainable gaps trigger a round. A missing baseline is one request away;
  a confirmed account behind an email gate is not, and retrying it forever
  burns budget while nothing changes. Unreachable gaps are recorded, never
  retried.
- Rounds are bounded per finding (two by default) and the remaining gap is
  written into the review reason, so a needs-review says exactly what was
  missing instead of sending a human looking from scratch.

Actions are ordered by cost: a baseline is one request, a browser run costs
seconds and a process. Spending the expensive step before the cheap one has had
a chance to decide it is how a budget goes without buying anything.

Merging is monotonic — a later run that did not see the marker does not unsee
an earlier one that did.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-14 00:54:01 -03:00
..