Addresses the benchmark's honest edge (a genuine BOLA credential dump graded
Low because evidence_data was null). Two fixes so criticals like it are not
recalibrated away:
- attack_graph::backfill_evidence — when evidence_data is null but the agent
recorded a proof in prose, copy that text into the structured slot the grader
reads (no fabrication, just relocation). Called first in enrich().
- attack_graph::data_class — classifies the demonstrated data (none/data/
sensitive) by scanning every evidence slot for credential/key/PII/payment
signatures. cvss_graded now grants the confidentiality receipt when sensitive
data was shown, even on a thin structured receipt — the KIND of data is itself
the impact.
- TypeSafe adjudication adds a `data_sensitivity` Score (public → PII → secrets),
carried on Adjudication. The pipeline regrade only strips impact when the
model was unconvinced AND no sensitive data was shown AND data_sensitivity is
low; a demonstrated credential/PII exposure keeps its severity.
articles/ — LinkedIn article (PT, no em-dashes) in Markdown + DOCX: explains
TypeSafe/System One/Jev, NeuroSploit, how to configure TypeSafe, the step-by-step
benchmark, results, the refinements this forced, and offensive-security use cases.
383 tests.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>