The 13-target benchmark left 3 misses. Root-caused and fixed the two that were
coverage gaps (the third was single-run variance, already handled by the
session-limit fix):
- CRLF header injection (web_crlf_header_go): the agent confirmed the open
redirect on /go?url= and stopped; the CRLF payload was never generated. The
open_redirect skill now tests %0d%0a header injection on the SAME param, and
CHAIN_DOCTRINE says a param landing in a Location header must also be tested
for response splitting. chain.rs: CWE-113/93/644 now provide capabilities;
attack_graph maps their kill-chain stage.
- Second-order SQLi (web_sqli_second_order): the sink was behind /admin, which
the customer account could not reach. CHAIN_DOCTRINE now teaches the
precondition pattern (store the payload, trigger from every identity, escalate
first if the trigger page needs a role you lack, else report as a chained
lead). chain.rs: CWE-564 requires PrivilegedContext so it chains after privesc.
BENCHMARK.md: added the TypeSafe calibrated-adjudication row; dropped the
"genuinely ahead" prose (the table is the summary); condensed the rest
188 -> 89 lines; refreshed scale (27 validators, 47 modules, 383 tests).
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Addresses the benchmark's honest edge (a genuine BOLA credential dump graded
Low because evidence_data was null). Two fixes so criticals like it are not
recalibrated away:
- attack_graph::backfill_evidence — when evidence_data is null but the agent
recorded a proof in prose, copy that text into the structured slot the grader
reads (no fabrication, just relocation). Called first in enrich().
- attack_graph::data_class — classifies the demonstrated data (none/data/
sensitive) by scanning every evidence slot for credential/key/PII/payment
signatures. cvss_graded now grants the confidentiality receipt when sensitive
data was shown, even on a thin structured receipt — the KIND of data is itself
the impact.
- TypeSafe adjudication adds a `data_sensitivity` Score (public → PII → secrets),
carried on Adjudication. The pipeline regrade only strips impact when the
model was unconvinced AND no sensitive data was shown AND data_sensitivity is
low; a demonstrated credential/PII exposure keeps its severity.
articles/ — LinkedIn article (PT, no em-dashes) in Markdown + DOCX: explains
TypeSafe/System One/Jev, NeuroSploit, how to configure TypeSafe, the step-by-step
benchmark, results, the refinements this forced, and offensive-security use cases.
383 tests.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Honesty audit found four modules written but not on the runtime path. Two mattered
and are now wired; two are noted.
cvss.rs — was NOT called; finding.cvss came from the old attack_graph ladder.
Now attack_graph::cvss_graded() bridges the class shape + demonstrated rung into
crate::cvss::grade (the FIRST-verbatim v3.1 equation), and enrich() sets
finding.cvss from the demonstrated vector, recording the potential ceiling in
the impact text. The class ladder remains only as a fallback for findings with
no evidence to grade.
waf.rs — the deterministic classifier was NOT run on any real exchange (only
WAF_OPS prompt text reached the agent). Now poc.rs classifies each re-run: a PoC
answered by a WAF/CDN is Unverifiable, not "gone" — closing a false-demotion
where an edge block looked like a fix.
TypeSafe (System One) extended per the build-with docs:
- CVSS via System One: when impact_demonstrated < 0.5, the finding's CVSS is
re-graded with impact receipts stripped — the calibrated judgment, not just
the rung, decides the demonstrated number.
- Agent selection: typesafe_prune_agents() asks one batched request (the
fan-out pattern), a Noul per chosen agent, and drops only those it calibrates
as clearly irrelevant (p < 0.25), never prunes to empty. Additive over the
LLM selection; skipped without a key.
Still shelf-ware, flagged honestly (not wired): inbox.rs (mail.tm/SMS happens
via agent prompt instructions, the Rust client is unused) and pomdp.rs
(redundant — belief.rs is the one on the path).
374 tests.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The Arena engagement produced 24 findings, a graph with 83 edges, and exactly
one chain edge. Two defects, both a step nobody was doing rather than a model
reasoning badly.
chains_from came back empty on every finding. An agent works one vulnerability
and has no view of what the other twelve found, so asking it to link its result
to findings it never saw was asking for something it cannot know. Chaining now
happens after the whole set is visible, on rules about ENABLEMENT: what one
weakness yields that another needs. Account enumeration yields valid
identities; absent throttling turns them into unlimited guesses; a permissive
password policy makes the guessing land. None is severe alone, and that
sequence is how accounts get taken over — on the real data it now reads
CWE-307 <- CWE-204, CWE-208 and CWE-614 <- CWE-319.
The CWE->stage fallback sent 23 of 24 findings to initial-access, so the kill
chain had one populated column and drew a star. Enumeration and side channels
are discovery; missing throttling, password policy, cookie flags and session
fixation are credential-access; hardening headers are recon. The same run now
spreads across credential-access 13, discovery 6, initial-access 5.
Two bugs the tests and the real data caught:
- CWE-614 both yields session material and needs it, so a class chained to
itself: duplicates formed a circular "attack path" from a cookie flag to the
same cookie flag. A weakness class no longer enables itself.
- apply_links only fills an empty chains_from, which is right for asserted
chains and wrong for links written by an older version of these rules — a
report kept the circular link through two rebuilds because nothing was
allowed to touch it. repair() now drops links that cannot be true whoever
wrote them: self-references, same-class links, dangling ids, cross-host
links.
enrich() still only fills empty fields during a run (an agent's judgement
should survive), but a rebuild applies the current mappings via remap_stages —
otherwise a finished run is frozen with whatever taxonomy existed that day.
path_for() gives the per-vulnerability view: what precedes this finding, what
it enables, and the narrative to print beside it.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
"SQLi = Critical" was the shortcut. The same weakness is a different finding
depending on how far the evidence took it, and the report has to be able to
defend the difference.
A ladder is read off the recorded observations:
reached the component → the mechanic is proven, impact is not
read data → confidentiality impact is real
read SENSITIVE data → and it is high
wrote (read-back) → integrity impact is real
executed code → the system is compromised
crossed to a second system → scope changes
The class now sets the CEILING and the evidence sets the score, so an injection
that reached the interpreter and extracted nothing no longer scores like one
that returned credentials.
Only observations climb it. "Could lead to remote code execution" stays at the
bottom rung — a test asserts exactly that, because prose is where inflation
enters.
Temporal metrics come from facts the engagement owns: E from whether a runnable
PoC exists, RC from the validation verdict (needs-review is Reasonable, never
Confirmed). They only ever lower the score.
A bug the tests caught: the first rung kept the class's availability impact, so
"reached" scored ABOVE "read data" — the ladder inverted at its first step.
Two older tests encoded the behaviour this replaces ("a bare CWE-89 must be
critical"). They now assert the new contract instead: a class name alone earns
no critical, and command execution scores like command execution only when
execution was observed.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Driven by the Arena Hockey engagement, where all 24 findings shipped with an
empty CVSS field and the PDF ran payloads off the page edge.
CVSS
- Derived deterministically from what the harness knows: the weakness class
sets the impact shape, the PROVEN exploitability sets attack complexity, and
the auth context sets privileges required. The vector is emitted with the
score, because a score without its vector cannot be checked and an unchecked
score is just a bigger adjective.
- The base equation is the v3.1 specification verbatim, including round-up and
the scope-changed privileges table. Tests anchor it against known values
(9.8 unauthenticated RCE, 10.0 with scope change, 6.1 reflected XSS, 0.0 for
no impact).
- An unknown weakness stays conservative — guessing high impact from a class
nobody mapped is how reports get inflated. An agent-supplied score is never
overwritten.
PDF
- Steps are passed as an ARRAY and rendered as a real numbered list, one command
per box. The previous template flattened them into a single `raw` block, which
rendered five separate commands as one run-on paragraph.
- Finding blocks are breakable, so a long evidence dump flows to the next page
instead of off the bottom of this one.
- `wrappable()` inserts zero-width breaks so encoded payloads wrap. The first
attempt broke prose mid-word ("rota ted", "lockoutOnFailu re=false") by
breaking every N characters regardless of context; it now works per token and
leaves anything that fits on a line exactly as it was.
- rebuild() re-enriches before rendering, so a run that finished before a
mapping existed picks it up instead of reprinting the gap forever.
Over-claimed findings are capped, not deleted
- The engagement rejected "no rate limiting on the password-reset flow" because
the agent claimed inbox flooding and only proved 25 unthrottled requests. The
claim was inflated; the measurement was real, and dropping it hid a genuine
gap. A unanimously rejected finding that still carries a checkable receipt is
now capped to Low and flagged for review, with the validator's reason
attached — the reader gets the fact without the story built on it.
- The agent contract now says impact must be what was MEASURED, and warns that
inflating it costs the whole finding.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
- Finding enriched with owasp / mitre / kill-chain stage / exploitability /
business_impact / chains_from (attack-path edges).
- attack_graph module: derive OWASP Top 10 + MITRE ATT&CK technique + kill-chain
stage from CWE (heuristic, no extra model call); render a Mermaid attack-path
flowchart (findings grouped by stage, explicit + implicit edges) and an ASCII
kill chain for the REPL.
- enrich() runs in finish() for every engagement.
- HTML report gains an "Attack Path & Kill Chain" section (Mermaid via CDN, dark)
plus a stage/sev/OWASP/MITRE/exploitability table.
- REPL print_findings shows the ASCII kill-chain + severity summary after a run.
- models: add GPT-5.5, GPT-5.4, GPT-5.4-mini, GPT-5.3-codex, GPT-5.2.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>