docs: drop links to the internal benchmark folder

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
CyberSecurityUP
2026-09-20 19:11:17 -03:00
co-authored by Claude Opus 5
parent d5d136ef34
commit d752e252e6
2 changed files with 6 additions and 7 deletions
+3 -5
View File
@@ -64,9 +64,7 @@ Control TUI**.
> (`--compliance pci-dss,hipaa,soc2`); an **internal-network / AD attack graph**;
> a **reasoning-budget governor** (`--budget`); and **TypeSafe System One**
> (`--typesafe on|off|auto`) as a calibrated confirmation + adjudication layer.
> 27 deterministic per-CWE validators, 446 agents. See
> [benchmarks/typesafe-2026-09-20](benchmarks/typesafe-2026-09-20/) for a
> with/without measurement.
> 27 deterministic per-CWE validators, 446 agents.
- 🧠 **POMDP belief + anti-hallucination gate** — findings aren't booleans; a
property-graph belief carries probabilities, and `may_assert` refuses to claim
@@ -505,8 +503,8 @@ neurosploit run https://app --typesafe off # the identical pipeline, no TypeS
```
`--typesafe auto` (default) is on when the key is set. Each run's `meta.json`
records `"typesafe": true|false` — a clean with/without measurement, one of
which lives in [`benchmarks/typesafe-2026-09-20/`](benchmarks/typesafe-2026-09-20/).
records `"typesafe": true|false` — a clean with/without measurement you can run
against your own target.
### Scope-evasion resistance, evidence integrity, untrusted output
+3 -2
View File
@@ -800,8 +800,9 @@ finding with a calibrated `{confirmed/needs-review/rejected}` judgment over the
*evidence*, re-grades CVSS when impact isn't demonstrated, prunes irrelevant
agents, and runs a code-owned confirmation loop over enumerable classes. It is
**additive** — a deterministic validator still rules; TypeSafe can only lower
confidence or flag for review, never resurrect a rejected claim. A with/without
measurement lives in [`benchmarks/typesafe-2026-09-20/`](benchmarks/typesafe-2026-09-20/).
confidence or flag for review, never resurrect a rejected claim. Run it with
`--typesafe on` and `--typesafe off` against the same target to measure the
difference (`meta.json` records which mode ran).
### Internal network / AD & reasoning budget