Files
NeuroSploit/neurosploit-rs
CyberSecurityUPandClaude Opus 5 22f2a3894d feat(claims): separate mechanic from impact so an overstatement stops deleting the observation
The Arena engagement rejected "no rate limiting on the password-reset flow"
outright. The agent had proved 25 requests accepted with no 429, no
Retry-After, no RateLimit-* — and then titled it "reset-email flooding". The
voter judged the claimed impact unproven and discarded everything, so a real
missing control never reached the report.

That was structural, not a bad call. The judge got a prose paragraph and one
accept/reject lever, while agents reliably walk: control absent -> abuse
possible -> impact plausible -> impact written as fact. The chain has to break
at step two, and that needs the finding to arrive as separable claims.

claims.rs adds:
- An evidence ledger (E01, E02, …) so a verdict is auditable: "supported by
  E01-E27" is checkable, "the evidence looks convincing" is not. A claim citing
  an id that was never recorded is REJECT_INVALID_EVIDENCE — worse than citing
  nothing, because it looks supported.
- Mechanic and impact as separate claims, each with its own citations. An
  asserted status never outruns its evidence: a model may downgrade itself and
  can never upgrade past what it cited.
- Six structured decisions instead of accept/reject. DOWNGRADE_UNPROVEN_IMPACT
  and DOWNGRADE_SCOPE_LIMITATION cannot discard — that is enforced by
  Decision::discards(), not by an instruction a model could reinterpret.
- Impact preconditions: "email flooding" needs account_exists +
  account_confirmed + email_delivery_observed. 25 accepted requests prove
  throttling was not observed; they do not prove mail was delivered. The
  difference is now computed, not argued.
- A rewriter, because lowering severity is not enough: a report headed
  "Password Reset Email Flooding" still asserts flooding whatever number sits
  beside it. The title is rebuilt from the mechanic, the impact prose becomes
  Observed / Not demonstrated / Potential impact, and the claimed consequence
  survives only as clearly labelled potential.
- still_security_relevant(): strip the unproven impact and ask whether anything
  remains. "The reset endpoint has no observable rate limiting" does; "the
  application responded" does not. That decides retain-vs-reject.

Wired into the pipeline ahead of the voters, and the voter can no longer delete
a finding that arrived with claims — it can only mark the narrative rejected
while the mechanic stands.

A test caught an inverted comparison in the severity cap that silently left a
High finding at High: a cap must lower and never raise.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-14 00:27:29 -03:00
..