mirror of
https://github.com/CyberSecurityUP/NeuroSploit.git
synced 2026-09-29 12:31:43 +02:00
feat(typesafe): System One calibrated adjudication (RLCD tier)
Integrates TypeSafe's System One model (Jev) as an optional, calibrated
adjudicator — the RLCD (Reinforcement Learning for Calibrated Decisions) tier:
typed judgments with probabilities where the harness needs a number, not prose.
- typesafe.rs: HTTP client for POST /v1/systemone (Bearer TYPESAFE_API_KEY,
model jev-latest), with Choice/Noul/Score primitives, retry on 429/529, and
parsed answers exposing the probability distribution + confidence.
adjudicate() asks a Choice {confirmed/needs-review/rejected} plus an
impact-demonstrated Noul over a finding; calibrated_confidence() folds
demonstrated impact into the number, wants_review() gates a split distribution.
- pipeline: an optional pass (runs when TYPESAFE_API_KEY is set, off with
NEUROSPLOIT_TYPESAFE=off) adjudicates each finding over its EVIDENCE — never
its narrative — refining confidence and the needs-review boundary. Additive:
a deterministic validator still rules; TypeSafe can only lower confidence or
flag for review, never resurrect a rejected claim. Audited per finding.
- env.example + README document it; the web console inherits the key via env.
Where the model stack maps in NeuroSploit: LM/BERT ≈ the deterministic
validators (no model), RLHF chat ≈ the exploit/recon agents, RLVR reasoning ≈
the DeepReasoning budget tier, RLCD ≈ this calibrated adjudication.
373 tests (+5).
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
co-authored by
Claude Opus 5
parent
2e95556df5
commit
c3de51d508
@@ -525,6 +525,21 @@ neurosploit provenance scan report.pdf.txt # is this ours? which build?
|
||||
neurosploit provenance verify runs/ns-… # manifest vs findings
|
||||
```
|
||||
|
||||
### TypeSafe System One — calibrated adjudication (RLCD)
|
||||
|
||||
When `TYPESAFE_API_KEY` is set, each finding is adjudicated by TypeSafe's
|
||||
System One model (Jev) — a **calibrated decision** over its *evidence*, not its
|
||||
prose: a `Choice` of `{confirmed, needs-review, rejected}` with a probability
|
||||
distribution, plus a `Noul` on whether real impact was demonstrated. The result
|
||||
refines the finding's confidence and moves borderline cases to needs-review.
|
||||
|
||||
It is **additive**: a deterministic validator still rules (a rejected finding
|
||||
stays rejected), and TypeSafe can only lower confidence or flag for review,
|
||||
never resurrect a claim. Every adjudication is written to the audit trail.
|
||||
Disable with `NEUROSPLOIT_TYPESAFE=off`. This is the RLCD (Reinforcement
|
||||
Learning for Calibrated Decisions) tier of the model stack — typed judgments
|
||||
where the harness needs a number, not a paragraph.
|
||||
|
||||
### Scope-evasion resistance, evidence integrity, untrusted output
|
||||
|
||||
Three hardening passes, all enforced in code:
|
||||
|
||||
Reference in New Issue
Block a user