feat(typesafe): System One calibrated adjudication (RLCD tier)

Integrates TypeSafe's System One model (Jev) as an optional, calibrated
adjudicator — the RLCD (Reinforcement Learning for Calibrated Decisions) tier:
typed judgments with probabilities where the harness needs a number, not prose.

- typesafe.rs: HTTP client for POST /v1/systemone (Bearer TYPESAFE_API_KEY,
  model jev-latest), with Choice/Noul/Score primitives, retry on 429/529, and
  parsed answers exposing the probability distribution + confidence.
  adjudicate() asks a Choice {confirmed/needs-review/rejected} plus an
  impact-demonstrated Noul over a finding; calibrated_confidence() folds
  demonstrated impact into the number, wants_review() gates a split distribution.
- pipeline: an optional pass (runs when TYPESAFE_API_KEY is set, off with
  NEUROSPLOIT_TYPESAFE=off) adjudicates each finding over its EVIDENCE — never
  its narrative — refining confidence and the needs-review boundary. Additive:
  a deterministic validator still rules; TypeSafe can only lower confidence or
  flag for review, never resurrect a rejected claim. Audited per finding.
- env.example + README document it; the web console inherits the key via env.

Where the model stack maps in NeuroSploit: LM/BERT ≈ the deterministic
validators (no model), RLHF chat ≈ the exploit/recon agents, RLVR reasoning ≈
the DeepReasoning budget tier, RLCD ≈ this calibrated adjudication.

373 tests (+5).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
CyberSecurityUP
2026-09-19 17:56:34 -03:00
co-authored by Claude Opus 5
parent 2e95556df5
commit c3de51d508
5 changed files with 415 additions and 0 deletions
+15
View File
@@ -525,6 +525,21 @@ neurosploit provenance scan report.pdf.txt # is this ours? which build?
neurosploit provenance verify runs/ns-… # manifest vs findings
```
### TypeSafe System One — calibrated adjudication (RLCD)
When `TYPESAFE_API_KEY` is set, each finding is adjudicated by TypeSafe's
System One model (Jev) — a **calibrated decision** over its *evidence*, not its
prose: a `Choice` of `{confirmed, needs-review, rejected}` with a probability
distribution, plus a `Noul` on whether real impact was demonstrated. The result
refines the finding's confidence and moves borderline cases to needs-review.
It is **additive**: a deterministic validator still rules (a rejected finding
stays rejected), and TypeSafe can only lower confidence or flag for review,
never resurrect a claim. Every adjudication is written to the audit trail.
Disable with `NEUROSPLOIT_TYPESAFE=off`. This is the RLCD (Reinforcement
Learning for Calibrated Decisions) tier of the model stack — typed judgments
where the harness needs a number, not a paragraph.
### Scope-evasion resistance, evidence integrity, untrusted output
Three hardening passes, all enforced in code: