release: v4.1.0 — assurance layer, TypeSafe, hardening + benchmark

Version bumped to 4.1.0 across the workspace, binaries, web console and Typst
template.

README: new "New in v4.1.0" summary; trimmed the verbose highlight bullets and
the TypeSafe section; removed the anti-plagiarism/provenance section (provenance
stays in the code, just not front-and-centre in the README); TypeSafe promoted
to its own top-level section; agent count 446.

TUTORIAL: new section 17 "Assurance & authorization" covering the target gate,
--scope-file, evidence-graded CVSS, audit anchoring + assurance bundle, sandbox,
intercept proxy, PoC re-validation, compliance mapping, TypeSafe, and the
internal/AD graph + budget governor.

benchmarks/typesafe-2026-09-20/: the with/without TypeSafe measurement —
report.html, scorer, both runs' findings/assurance/meta/logs, and a README.
No secrets committed (env-only during the runs, verified clean).

381 tests.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
CyberSecurityUP
2026-09-20 12:25:23 -03:00
co-authored by Claude Opus 5
parent d2ec0a112d
commit 088d133c80
28 changed files with 4292 additions and 126 deletions
+54 -99
View File
@@ -1,4 +1,4 @@
<h1 align="center">🧠 NeuroSploit v4.0.0</h1>
<h1 align="center">🧠 NeuroSploit v4.1.0</h1>
<p align="center">
<a href="https://github.com/JoasASantos/NeuroSploit/stargazers"><img src="https://img.shields.io/github/stars/JoasASantos/NeuroSploit?style=for-the-badge&logo=github&color=8b5cf6" alt="Stars"></a>
@@ -8,10 +8,10 @@
</p>
<p align="center">
<img src="https://img.shields.io/badge/Version-4.0.0-blue?style=flat-square">
<img src="https://img.shields.io/badge/Version-4.1.0-blue?style=flat-square">
<img src="https://img.shields.io/badge/Harness-Rust%20%7C%20tokio-e6b673?style=flat-square">
<img src="https://img.shields.io/badge/License-MIT-green?style=flat-square">
<img src="https://img.shields.io/badge/MD%20Agents-435-red?style=flat-square">
<img src="https://img.shields.io/badge/MD%20Agents-446-red?style=flat-square">
<img src="https://img.shields.io/badge/Models-18%20providers-success?style=flat-square">
<img src="https://img.shields.io/badge/Modes-Black%20%7C%20White%20%7C%20Grey%20%7C%20Host%20%7C%20AI-9cf?style=flat-square">
<img src="https://img.shields.io/badge/Auth-API%20key%20%7C%20Subscription-orange?style=flat-square">
@@ -50,50 +50,43 @@ Control TUI**.
### Highlights
- 🧠 **POMDP belief + value-of-information** — the target is partially observable,
so findings aren't booleans: a property-graph **belief** carries probabilities,
and "scan more vs exploit now" falls out of belief entropy. The `may_assert`
gate is a **mathematical anti-hallucination rule** (don't claim exploitability
while the belief is diffuse).
- 🧾 **Grounding** — hard rule: **no claim without a receipt** (evidence, not
paraphrase). Empirical (raw tool output) for black-box/host/AI, **symbolic**
(`file:line` into the reviewed source — a code citation *is* the receipt) for
white-box SAST & skills audits, and **either** for grey-box; ungrounded claims
are demoted.
- 🔬 **Deterministic HTTP probe** — before the model recon, the harness runs a
**real** request/response analysis (status/redirects, security headers, cookie
flags, CORS reflection, tech fingerprint, linked JS, 404 baseline, high-signal
paths) and feeds those observed facts into recon, so agent selection and
exploitation decisions are grounded in evidence — not the model's guess.
- 🔗 **Attack chaining — any primitive pivots.** 13 multi-stage chain agents
(SQLi→RCE→LPE, SSRF→cloud creds, upload→LFI→RCE→LPE, CVE→RCE→pivot, …) **plus a
chaining doctrine** that turns *any* confirmed foothold into the next step:
reduce it to a primitive (exec / read / write / request-forgery / identity /
secret) and pivot — file-upload→RCE, SSRF→metadata creds, IDOR→takeover — reusing
looted creds and reasoning about **business logic** (payment/tenancy/workflow
abuse). Each stage proven; strictly non-destructive (no data loss, no DB
overwrite, no DoS).
- ☁️ **Cloud testing** — AWS / GCP / Azure agents that drive the provider CLIs
(`aws`/`gcloud`/`az`). Connect via `creds.yaml`: AWS keys, a Google
service-account JSON, or an Azure service principal — see
[Cloud credentials](#cloud-credentials-awsgcpazure).
- 🤖 **LLM red-teaming** — 30 AI agents that jailbreak & prompt-inject a live AI
system across scenarios: **AdvPrefix**, **PAIR**, **TAP**, **Crescendo**,
many-shot, persona/DAN, encoding/obfuscation, refusal-suppression; plus
**indirect injection** (RAG/web/email/tool output), **goal hijacking**,
tool/function-call abuse, and system-prompt exfiltration. Each runs an
attacker→**LLM-judge** loop (baseline refusal → technique → verdict) and proves
the bypass with a **benign, redacted** receipt. Maps to OWASP LLM Top 10 (2025),
MCP threats & OWASP AI Exchange; Skill/plugin & **n8n** files audited white-box.
- 🧰 **Misconfig & CVE hunting → exploitation, safely** — a full CVE pipeline:
**version fingerprint** (pin exact versions) → **research analyst** (map to
NVD/GHSA CVEs, judge reachability) → **PoC finder** (locate/vet/adapt a public
PoC) → **exploit scripter** (write a custom exploit when none exists). Every PoC
is written to the run's **`pocs/` folder and referenced in the report** so
findings are reproducible. Plus absurd-misconfig agents (exposed `.git`/`.env`,
debug/actuator, default creds, dashboards, CORS) and rate-limit testing — all
under a strict **data-safety/PII guardrail** (no destructive/state-changing
actions; PII proven with a masked sample, never dumped).
> **New in v4.1.0** — evidence-graded CVSS computed from the FIRST v3.1 equation
> (not guessed by class); a **target-authorization gate** (default-deny, refuses a
> target outside the capability grant before any recon); **audit anchoring** that
> detects truncation & silent rebuilds; a signed **assurance bundle** (P1–P5 in one
> manifest per run); **scope-evasion resistance** (alt-IP-encoding normalization,
> redirect-to-private-IP block, DNS-rebinding guard); **evidence-integrity** checks
> (cross-target / reused-receipt / foreign-marker / orphan-claim rejection);
> **untrusted-output taint** (prompt-injection stripping + data fencing); a
> **`--scope-file` YAML loader** + web Scoping/Guardrails UI; a **Kali sandbox**
> (`--sandbox`), **intercept proxy** (`--intercept burp|caido|zap|mitmproxy|own`),
> **PoC re-validation** (`--revalidate-poc`), **compliance mapping**
> (`--compliance pci-dss,hipaa,soc2`); an **internal-network / AD attack graph**;
> a **reasoning-budget governor** (`--budget`); and **TypeSafe System One**
> (`--typesafe on|off|auto`) as a calibrated confirmation + adjudication layer.
> 27 deterministic per-CWE validators, 446 agents. See
> [benchmarks/typesafe-2026-09-20](benchmarks/typesafe-2026-09-20/) for a
> with/without measurement.
- 🧠 **POMDP belief + anti-hallucination gate** — findings aren't booleans; a
property-graph belief carries probabilities, and `may_assert` refuses to claim
exploitability while the belief is diffuse.
- 🧾 **Grounding — no claim without a receipt.** Empirical (raw tool output) or
symbolic (`file:line` into the reviewed source); ungrounded claims are demoted.
- 🔬 **Deterministic HTTP probe** feeds observed facts (headers, cookies, CORS,
fingerprint, JS, 404 baseline) into recon — decisions grounded in evidence,
not the model's guess.
- 🔗 **Attack chaining — any primitive pivots.** Reduce a foothold to a primitive
(exec/read/write/request-forgery/identity/secret) and pivot; each stage proven,
strictly non-destructive.
- ☁️ **Cloud testing** — AWS / GCP / Azure agents driving `aws`/`gcloud`/`az` via
`creds.yaml` ([details](#cloud-credentials-awsgcpazure)).
- 🤖 **LLM red-teaming** — jailbreak & prompt-inject a live AI system (AdvPrefix,
PAIR, TAP, Crescendo, indirect injection, goal hijacking) via an attacker→judge
loop; maps to OWASP LLM Top 10.
- 🧰 **Misconfig & CVE pipeline** — fingerprint → CVE research → PoC finder →
exploit scripter; every PoC written to `pocs/` and referenced in the report,
under a strict data-safety/PII guardrail.
- 🎯 **Re-test one vulnerability** — `--only <agent>` (repeatable /
comma-separated) runs exactly the agent(s) you name and skips recon-based
selection — re-test a single finding fast. Works on `run` / `whitebox` /
@@ -495,63 +488,25 @@ neurosploit run https://app.example --creds creds.yaml \
Each finding is proven with the **authorized vs unauthorized** request pair, under
the data-safety guardrail (read-only, PII masked).
## 🏷️ Identification & attribution (anti-plagiarism)
## 🧮 TypeSafe System One — calibrated adjudication
Every request is tagged with an identifying **User-Agent** (default
`NeuroSploit/<ver> …`, change with **`/ua`** or `NEUROSPLOIT_UA`) plus an
`X-NeuroSploit-Scan` header, and every finding is **stamped** "Identified and
validated by NeuroSploit" — so provenance travels in the traffic, the finding
text, `findings.json` and the report footer.
### Provenance — which build made this, and does it still match
Attribution that survives someone else's copy-paste:
- **`JOASNSCOPE`** leads every canary the harness mints, so a marker that
turns up later — in a response body, a customer's log, somebody else's
report — extracts whole and names the build that made it.
- **Per-build fingerprint** (`neurosploit provenance show`), plus an optional
per-customer build id via `NEUROSPLOIT_CUSTOMER_ID`.
- **`findings.json` is stamped** with `_engine`, and a **signed
`provenance.json`** ships beside it (`NEUROSPLOIT_PROVENANCE_KEY`).
- **Structural signature** over the finding set's shape — it survives
rewording and reformatting, but not a changed result.
- **Prompts are watermarked** at the single model-pool chokepoint
(`NEUROSPLOIT_WATERMARK=off` to disable).
Set `TYPESAFE_API_KEY` and NeuroSploit adjudicates each finding with TypeSafe's
System One model (Jev): a calibrated `{confirmed, needs-review, rejected}`
judgment over the *evidence* (not the prose), plus a check on whether real
impact was demonstrated. It refines confidence, re-grades CVSS when impact is
unproven, and runs a code-owned confirmation loop over enumerable classes
(XSS/SQLi/redirect/traversal/SSRF/IDOR). **Additive** — a deterministic
validator still rules; TypeSafe only lowers confidence or flags for review,
never resurrects a rejected claim.
```bash
neurosploit provenance show # this build's identity
neurosploit provenance scan report.pdf.txt # is this ours? which build?
neurosploit provenance verify runs/ns-… # manifest vs findings
neurosploit run https://app --typesafe on # calibrated adjudication + confirmation
neurosploit run https://app --typesafe off # the identical pipeline, no TypeSafe (A/B)
```
### TypeSafe System One — calibrated adjudication (RLCD)
When `TYPESAFE_API_KEY` is set, each finding is adjudicated by TypeSafe's
System One model (Jev) — a **calibrated decision** over its *evidence*, not its
prose: a `Choice` of `{confirmed, needs-review, rejected}` with a probability
distribution, plus a `Noul` on whether real impact was demonstrated. The result
refines the finding's confidence and moves borderline cases to needs-review.
It is **additive**: a deterministic validator still rules (a rejected finding
stays rejected), and TypeSafe can only lower confidence or flag for review,
never resurrect a claim. Every adjudication is written to the audit trail.
Disable with `NEUROSPLOIT_TYPESAFE=off`. This is the RLCD (Reinforcement
Learning for Calibrated Decisions) tier of the model stack — typed judgments
where the harness needs a number, not a paragraph.
**As an additional confirmation strategy** (`typesafe_agent`), a code-owned loop
where TypeSafe picks the next payload (`Choice`) and judges the real response
(`Noul`) over the replay engine — for enumerable classes (XSS, SQLi, open
redirect, path traversal, SSRF, IDOR). It runs only on findings the LLM path
left unconfirmed or in needs-review (the recall lever), can only raise a finding
to confirmed with a calibrated probability, never downgrades, and refuses edge
(WAF) responses. It is **not** a discovery agent — System One does not generate.
**Flag & A/B.** `--typesafe on|off|auto` (default auto = on when the key is set).
`off` runs the *identical* pipeline without it, and the run's `meta.json` records
`"typesafe": true|false` — so a with/without pair against the same target is a
clean measurement of what it adds.
`--typesafe auto` (default) is on when the key is set. Each run's `meta.json`
records `"typesafe": true|false` — a clean with/without measurement, one of
which lives in [`benchmarks/typesafe-2026-09-20/`](benchmarks/typesafe-2026-09-20/).
### Scope-evasion resistance, evidence integrity, untrusted output
+108 -4
View File
@@ -1,4 +1,4 @@
# NeuroSploit — Tutorial & User Guide (v4.0.0)
# NeuroSploit — Tutorial & User Guide (v4.1.0)
A complete, hands-on guide to installing, configuring and running NeuroSploit —
the autonomous, multi-model penetration-testing harness.
@@ -31,7 +31,8 @@ the autonomous, multi-model penetration-testing harness.
14. [The agent library](#14-the-agent-library)
15. [Playwright MCP & extra tools](#15-playwright-mcp--extra-tools)
16. [Tips, tuning & troubleshooting](#16-tips-tuning--troubleshooting)
17. [Command & flag reference](#17-command--flag-reference)
17. [Assurance & authorization (v4.1.0)](#17-assurance--authorization)
18. [Command & flag reference](#18-command--flag-reference)
---
@@ -99,7 +100,7 @@ Agents **degrade gracefully**: if `rustscan` is absent they use `nmap`; if neith
### Verify
```bash
neurosploit --version # neurosploit 4.0.0
neurosploit --version # neurosploit 4.1.0
neurosploit agents # {"vulns":241,...,"ai":30,...,"total":430}
neurosploit models # all providers & models
```
@@ -713,7 +714,110 @@ back to `curl`. You can add more MCP servers by placing a `mcp.servers.json`
---
## 17. Command & flag reference
## 17. Assurance & authorization
v4.1.0 adds a layer of controls that make a run **defensible**, not just
productive. All are enforced in code (not prompt text) and every decision lands
in the hash-chained audit trail.
### Target authorization gate (default-deny)
Before any recon, the target is checked against the capability grant — protocol,
host, port, URL prefix. A signed token that does not cover the target **refuses
the run** and exits non-zero:
```bash
# mint a grant for one host, then run against a different one → refused
neurosploit capability issue --scope app.example.com --issuer you --subject op --hours 8
neurosploit run https://other.example.com --capability-token <tok>
# ⛔ DENY_TARGET_OUTSIDE_GRANT — other.example.com is outside the authorized scope
# (non-zero exit; nothing was tested; the denial is audited)
```
Loopback (`localhost`/`127.0.0.1`) is exempt — it is unambiguous.
### Hard scope from a file
```bash
neurosploit run https://app.example.com --scope-file scope.yaml
```
```yaml
# scope.yaml — enforced in code; a capability token still caps it
hard: [ app.example.com, "*.staging.example.com", 10.20.30.0/24 ]
exclude: [ payments.example.com ]
soft:
observe_only: [ cdn.example.com ]
allow_destructive_methods: false
max_requests_per_minute: 240
forbidden_payloads: [ "drop table", "rm -rf /" ]
notes: [ "SOW-2026-0142; window 02:00-06:00 UTC" ]
```
Alt-IP encodings (`0x7f000001`, `2130706433`, `0177.0.0.1`,
`::ffff:127.0.0.1`) all normalize to dotted-quad, so an exclude can't be dodged
by re-spelling; redirects to a private/loopback address are refused; a
DNS-rebinding guard refuses a name that re-resolves to a new internal address.
### Evidence-graded CVSS
The score is computed from the FIRST v3.1 equation, and each impact metric is
graded against a receipt. SQLi that reached the interpreter but extracted
nothing scores **demonstrated 0 / potential 9.8** — never a manufactured
critical. The vector travels with the number in the report.
### Audit anchoring & the assurance bundle
```bash
neurosploit audit <run> --anchor # chain + signed anchors: catch truncation/rebuild/forgery
neurosploit assurance <run> # P1–P5 in one manifest (authorization/enforcement/evidence/integrity/provenance)
neurosploit assurance <run> --verify # re-hash every artifact + check the signature
```
Set `NEUROSPLOIT_ANCHOR_DIR` to also write anchors to external append-only
(ideally WORM) storage, and `NEUROSPLOIT_PROVENANCE_KEY` to sign them.
### Tooling: sandbox · proxy · PoC re-validation · compliance
```bash
--sandbox # run agent commands in a Kali container (docker/podman)
--intercept burp|caido|zap|mitmproxy # route through a tool …
--intercept own | own+burp # … or the harness's own recording interceptor
--revalidate-poc # re-run each PoC; demote what no longer reproduces
--compliance pci-dss,hipaa,soc2 # map findings onto control requirements in the report
```
### TypeSafe System One (calibrated confirmation)
```bash
export TYPESAFE_API_KEY=... # then:
neurosploit run https://app --typesafe on # calibrated adjudication + confirmation loop
neurosploit run https://app --typesafe off # the identical pipeline, no TypeSafe (for A/B)
```
`--typesafe auto` (default) turns it on when the key is set. It adjudicates each
finding with a calibrated `{confirmed/needs-review/rejected}` judgment over the
*evidence*, re-grades CVSS when impact isn't demonstrated, prunes irrelevant
agents, and runs a code-owned confirmation loop over enumerable classes. It is
**additive** — a deterministic validator still rules; TypeSafe can only lower
confidence or flag for review, never resurrect a rejected claim. A with/without
measurement lives in [`benchmarks/typesafe-2026-09-20/`](benchmarks/typesafe-2026-09-20/).
### Internal network / AD & reasoning budget
```bash
neurosploit internal --graph g.json --scaffold corp.local --from foothold --mermaid
neurosploit run https://app --budget eco|balanced|aggressive # ration reasoning; default unlimited
```
The internal graph models an engagement as `Asset → Exposure → Weakness →
Credential → Privilege → Movement → Crown Jewel` and answers the question a
CVSS-sorted list can't: **which single edge, removed, cuts the most paths to the
crown jewels** (`choke_points`).
---
## 18. Command & flag reference
```
neurosploit # interactive REPL (resumes per project)
+83
View File
@@ -0,0 +1,83 @@
# NeuroSploit × TypeSafe — benchmark (2026-09-20)
Two identical NeuroSploit engagements against the same vulnerable target — one
plain, one with **TypeSafe System One (Jev)** as a calibrated confirmation
layer. Same model, same focus, same 13 seeded vulnerabilities. Only the
`--typesafe` flag differs.
Open **`report.html`** for the full visual write-up.
## Setup
| | |
|---|---|
| Harness | NeuroSploit v4.0.0 |
| Model | `claude-opus-4-8` (subscription) |
| Target | NimbusCart / BenchMarkBurpAT · `http://localhost:3000` |
| Mode | black-box, `--recon 2`, `--vote-n 1`, `--max-agents 15` |
| Ground truth | 13 seeded scenarios (IDOR/BOLA, SQLi ×5, XSS ×4, open redirect, CRLF) |
| Solver | none — the LLM discovered and confirmed everything live |
Run commands (the only difference is `--typesafe`):
```bash
# A — no TypeSafe
NEUROSPLOIT_TYPESAFE=off neurosploit run http://localhost:3000 \
--subscription --model anthropic:claude-opus-4-8 \
--typesafe off --recon 2 --max-agents 15 --vote-n 1 --focus "<13 endpoints>" -v
# B — with TypeSafe (TYPESAFE_API_KEY set in env, never committed)
NEUROSPLOIT_TYPESAFE=on neurosploit run http://localhost:3000 \
--subscription --model anthropic:claude-opus-4-8 \
--typesafe on --recon 2 --max-agents 15 --vote-n 1 --focus "<13 endpoints>" -v
```
## Result
| Metric | A — no TypeSafe | B — TypeSafe |
|---|---|---|
| Targets hit | **10 / 13** | 9 / 13 |
| Findings | 16 | **18** |
| Wall-clock | 32m 12s | **26m 53s** |
| Criticals | 5 | 2 (recalibrated) |
| Belief-gate holds (POMDP) | 3 | — |
| Assurance P1–P5 | all present | all present |
| Model cost | $0 (subscription) | $0 + TypeSafe ≪ $5 |
Union coverage (both runs): **11 / 13**. Neither reached `web_sqli_second_order`
or `web_crlf_header_go`.
## Reading it honestly
- **Recall is a tie** — 10 vs 9 is within run-to-run variance at `vote-n 1`.
TypeSafe is a judgment layer, not a recall multiplier.
- **B surfaced 2 real net-new findings** the plain run missed (`config.json`
API-key exposure CWE-200, no-lockout brute force CWE-307) and caught
`web_idor_invoice`.
- **TypeSafe recalibrated severity** — 5 class-inflated Criticals → 2 evidence-
backed ones. On this target it *under-rated* one genuine critical (the BOLA
credential dump: A = Critical 9.1, B = Low). Calibration is a dial toward
defensibility, not a correctness oracle.
- **Harness gap found & fixed**: an earlier B collapsed to 0 findings when the
subscription hit a session limit mid-run — NeuroSploit treated the limit
message as a normal (exit-0) response and burned every agent. Now the
session-limit sentinel parks the run (`fix(models)`).
## Confounders
Single samples, not averages. `vote-n 1` = no cross-model agreement in either
arm. Recall scored by class + endpoint-keyword match (coverage, not graded
proof). One target. Treat as one honest data point, not a leaderboard.
## Files
```
report.html the visual write-up
score.py the scorer (class + endpoint keyword match vs the 13 targets)
scores.txt scorer output for both runs
run_a_no_typesafe/ findings.json · assurance.json · meta.json · report.html · run.log
run_b_typesafe/ findings.json · assurance.json · meta.json · report.html · run.log
```
The TypeSafe API key and any subscription tokens are **not** in these files
(env-only during the runs; verified clean before commit).
+282
View File
@@ -0,0 +1,282 @@
<title>NeuroSploit × TypeSafe Benchmark</title>
<meta name="description" content="Head-to-head of NeuroSploit against a vulnerable target, with and without TypeSafe System One as a confirmation layer.">
<link rel="stylesheet" href="https://fonts.googleapis.com/css2?family=IBM+Plex+Sans:wght@400;500;600&family=IBM+Plex+Mono:wght@400;500;600&family=Chivo:wght@600;700;800&display=swap">
<style>
:root{
--ground:#f4f2f7; --surface:#ffffff; --surface-2:#eceaf3; --line:#ddd8e8;
--ink:#1a1726; --muted:#6b6580; --faint:#938da6;
--accent:#6d4bd8; /* neuro violet */
--a:#c2701c; /* run A — amber (no typesafe) */
--b:#0e8f86; /* run B — teal (typesafe) */
--crit:#c8324a; --high:#d9743a; --med:#c2a01c; --low:#4a76c4; --info:#7b7590; --good:#1f9d68;
--shadow:0 1px 2px rgba(26,23,38,.06),0 6px 20px rgba(26,23,38,.06);
/* severity — vivid, identical in both themes (severity is not theme-relative) */
--sev-crit:#e5484d; --sev-high:#f76b15; --sev-med:#f5b301; --sev-low:#3e7bfa; --sev-info:#8b8698;
}
:root:not([data-theme="light"]){ @media (prefers-color-scheme:dark){
--ground:#0f0e17; --surface:#191627; --surface-2:#211d33; --line:#2e2942;
--ink:#eceaf5; --muted:#a49dbd; --faint:#736c8f;
--accent:#a78bfa; --a:#f0a95e; --b:#4fd6c6;
--crit:#f26d7d; --high:#f0965e; --med:#e6cd52; --low:#7aa6f0; --info:#9a93b4; --good:#5ee0a0;
--shadow:0 1px 2px rgba(0,0,0,.4),0 8px 30px rgba(0,0,0,.35);
}}
:root[data-theme="dark"]{
--ground:#0f0e17; --surface:#191627; --surface-2:#211d33; --line:#2e2942;
--ink:#eceaf5; --muted:#a49dbd; --faint:#736c8f;
--accent:#a78bfa; --a:#f0a95e; --b:#4fd6c6;
--crit:#f26d7d; --high:#f0965e; --med:#e6cd52; --low:#7aa6f0; --info:#9a93b4; --good:#5ee0a0;
--shadow:0 1px 2px rgba(0,0,0,.4),0 8px 30px rgba(0,0,0,.35);
}
*{box-sizing:border-box}
body{background:var(--ground);color:var(--ink);font-family:"IBM Plex Sans",system-ui,sans-serif;line-height:1.55;
-webkit-font-smoothing:antialiased;margin:0}
.wrap{max-width:1000px;margin:0 auto;padding:clamp(24px,5vw,64px) clamp(18px,4vw,40px)}
h1,h2,h3{font-family:"Chivo","IBM Plex Sans",sans-serif;text-wrap:balance;line-height:1.1;margin:0}
code,.mono,.num{font-family:"IBM Plex Mono",ui-monospace,monospace;font-variant-numeric:tabular-nums}
.eyebrow{font-family:"IBM Plex Mono",monospace;font-size:12px;letter-spacing:.18em;text-transform:uppercase;color:var(--accent);font-weight:600}
/* header */
header{border-bottom:1px solid var(--line);padding-bottom:28px;margin-bottom:36px}
h1{font-size:clamp(30px,5.5vw,50px);font-weight:800;margin:10px 0 8px;letter-spacing:-.02em}
.sub{color:var(--muted);font-size:16px;max-width:64ch}
.meta{display:flex;flex-wrap:wrap;gap:8px 18px;margin-top:18px;font-family:"IBM Plex Mono",monospace;font-size:12.5px;color:var(--faint)}
.meta b{color:var(--ink);font-weight:500}
/* thesis tiles */
.thesis{display:grid;grid-template-columns:repeat(auto-fit,minmax(150px,1fr));gap:14px;margin:34px 0}
.tile{background:var(--surface);border:1px solid var(--line);border-radius:12px;padding:18px 18px 16px;box-shadow:var(--shadow)}
.tile .k{font-family:"IBM Plex Mono",monospace;font-size:11px;letter-spacing:.1em;text-transform:uppercase;color:var(--faint)}
.tile .v{font-family:"Chivo",sans-serif;font-weight:800;font-size:30px;letter-spacing:-.02em;margin-top:6px;display:flex;align-items:baseline;gap:8px}
.tile .u{font-size:13px;font-weight:500;color:var(--muted);font-family:"IBM Plex Sans"}
.tile .note{font-size:12.5px;color:var(--muted);margin-top:4px}
.swatchA{color:var(--a)} .swatchB{color:var(--b)}
section{margin:44px 0}
h2{font-size:22px;font-weight:700;margin-bottom:4px}
.lead{color:var(--muted);font-size:15px;margin:6px 0 20px;max-width:70ch}
/* comparison table */
.cmp{width:100%;border-collapse:collapse;font-size:14.5px;background:var(--surface);border:1px solid var(--line);border-radius:12px;overflow:hidden;box-shadow:var(--shadow)}
.cmp th,.cmp td{padding:12px 16px;text-align:left;border-bottom:1px solid var(--line)}
.cmp thead th{font-family:"IBM Plex Mono",monospace;font-size:11.5px;letter-spacing:.08em;text-transform:uppercase;color:var(--faint);font-weight:600;background:var(--surface-2)}
.cmp tbody tr:last-child td{border-bottom:none}
.cmp td.metric{color:var(--muted)}
.cmp td .num{font-weight:600;font-size:15px}
.colA{color:var(--a)} .colB{color:var(--b)}
.win{position:relative}
.win::after{content:"▲";font-size:9px;margin-left:6px;vertical-align:middle;color:var(--good)}
/* per-scenario grid */
.scen{display:grid;grid-template-columns:1fr auto auto;gap:0;background:var(--surface);border:1px solid var(--line);border-radius:12px;overflow:hidden;box-shadow:var(--shadow)}
.scen .row{display:contents}
.scen .cell{padding:10px 16px;border-bottom:1px solid var(--line);display:flex;align-items:center;gap:10px}
.scen .row:last-child .cell{border-bottom:none}
.scen .head .cell{font-family:"IBM Plex Mono",monospace;font-size:11px;letter-spacing:.08em;text-transform:uppercase;color:var(--faint);background:var(--surface-2);font-weight:600}
.scen .idc{font-family:"IBM Plex Mono",monospace;font-size:13px}
.scen .cls{font-size:11px;color:var(--faint);font-family:"IBM Plex Mono";margin-left:auto;padding-left:10px}
.mk{width:60px;justify-content:center;font-family:"IBM Plex Mono";font-size:13px;font-weight:600}
.hit{color:var(--good)} .miss{color:var(--crit);opacity:.7}
.hdrA{color:var(--a)} .hdrB{color:var(--b)}
/* severity bars */
.sev-wrap{display:grid;grid-template-columns:1fr 1fr;gap:18px}
@media(max-width:640px){.sev-wrap{grid-template-columns:1fr}}
.sevcard{background:var(--surface);border:1px solid var(--line);border-radius:12px;padding:18px;box-shadow:var(--shadow)}
.sevcard h3{font-size:14px;font-family:"IBM Plex Mono";letter-spacing:.05em;margin-bottom:14px;display:flex;align-items:center;gap:8px}
.dot{width:9px;height:9px;border-radius:50%;display:inline-block}
.bar{display:flex;align-items:center;gap:12px;margin:9px 0;font-size:13px}
.bar .lab{width:70px;color:var(--muted);font-family:"IBM Plex Mono";font-size:11.5px;display:flex;align-items:center;gap:7px}
.bar .lab .sw{width:9px;height:9px;border-radius:2px;flex:none}
.bar .track{flex:1;height:22px;background:var(--surface-2);border-radius:5px;overflow:hidden;border:1px solid var(--line)}
.bar .fill{height:100%;border-radius:4px;min-width:6px;box-shadow:inset 0 0 0 1px rgba(255,255,255,.08)}
.bar .n{width:22px;text-align:right;font-family:"IBM Plex Mono";font-weight:700;font-size:14px}
.sev-legend{display:flex;flex-wrap:wrap;gap:14px;margin:0 0 18px;font-family:"IBM Plex Mono";font-size:11.5px;color:var(--muted)}
.sev-legend span{display:inline-flex;align-items:center;gap:6px}
.sev-legend i{width:11px;height:11px;border-radius:3px;display:inline-block}
/* callout */
.callout{background:var(--surface);border:1px solid var(--line);border-left:3px solid var(--accent);border-radius:10px;padding:20px 22px;box-shadow:var(--shadow)}
.callout h3{font-size:16px;margin-bottom:10px}
.callout p{margin:8px 0;font-size:14.5px;color:var(--ink)}
.callout .contrast{display:grid;grid-template-columns:1fr 1fr;gap:14px;margin-top:14px}
@media(max-width:560px){.callout .contrast{grid-template-columns:1fr}}
.cbox{background:var(--surface-2);border-radius:8px;padding:12px 14px}
.cbox .t{font-family:"IBM Plex Mono";font-size:11px;text-transform:uppercase;letter-spacing:.08em;margin-bottom:6px}
.cbox .r{font-size:13px;color:var(--muted)}
.cbox .g{font-size:20px;font-family:"Chivo";font-weight:800;margin-top:4px}
ul.take{list-style:none;padding:0;margin:0;display:flex;flex-direction:column;gap:12px}
ul.take li{background:var(--surface);border:1px solid var(--line);border-radius:10px;padding:14px 16px;font-size:14.5px;display:flex;gap:12px;box-shadow:var(--shadow)}
ul.take .tag{font-family:"IBM Plex Mono";font-size:10.5px;font-weight:600;letter-spacing:.06em;padding:3px 8px;border-radius:5px;height:fit-content;white-space:nowrap;text-transform:uppercase}
.tag.even{background:color-mix(in srgb,var(--info) 22%,transparent);color:var(--info)}
.tag.plus{background:color-mix(in srgb,var(--good) 20%,transparent);color:var(--good)}
.tag.minus{background:color-mix(in srgb,var(--crit) 18%,transparent);color:var(--crit)}
.tag.note{background:color-mix(in srgb,var(--accent) 18%,transparent);color:var(--accent)}
.disclaim{margin-top:44px;padding-top:22px;border-top:1px solid var(--line);color:var(--faint);font-size:12.5px;line-height:1.6}
.disclaim b{color:var(--muted)}
a{color:var(--accent)}
</style>
<div class="wrap">
<header>
<div class="eyebrow">NeuroSploit · assurance benchmark · 2026-09-20</div>
<h1>Does TypeSafe make the run better?</h1>
<p class="sub">Two identical NeuroSploit engagements against the same vulnerable target — one plain,
one with TypeSafe System One (Jev) as a calibrated confirmation layer. Same model, same focus,
same 13 seeded vulnerabilities. Only the <code>--typesafe</code> flag differs.</p>
<div class="meta">
<span>target <b>NimbusCart (BenchMarkBurpAT)</b> · localhost:3000</span>
<span>model <b>claude-opus-4-8</b> (subscription)</span>
<span>recon <b>2</b> · vote-n <b>1</b> · max-agents <b>15</b></span>
<span>ground truth <b>13 targets</b></span>
</div>
</header>
<div class="thesis">
<div class="tile">
<div class="k">Recall — no TypeSafe</div>
<div class="v swatchA">10<span class="u">/13</span></div>
<div class="note">16 findings · 32m12s</div>
</div>
<div class="tile">
<div class="k">Recall — with TypeSafe</div>
<div class="v swatchB">9<span class="u">/13</span></div>
<div class="note">18 findings · 26m53s</div>
</div>
<div class="tile">
<div class="k">Union coverage</div>
<div class="v">11<span class="u">/13</span></div>
<div class="note">the two runs together</div>
</div>
<div class="tile">
<div class="k">TypeSafe recalibrated</div>
<div class="v swatchB">9</div>
<div class="note">findings, calibrated confidence</div>
</div>
</div>
<section>
<h2>Head to head</h2>
<p class="lead">The recall is a tie inside the noise; the real difference is <em>shape</em>. TypeSafe was
faster, surfaced two real findings the plain run missed, and pulled inflated severities down toward what the
evidence actually demonstrated — at the cost of being conservative enough to drop two scenarios and under-rate
one genuine critical.</p>
<div style="overflow-x:auto">
<table class="cmp">
<thead><tr><th>Metric</th><th class="colA">A — no TypeSafe</th><th class="colB">B — TypeSafe</th></tr></thead>
<tbody>
<tr><td class="metric">Seeded targets hit</td><td class="colA win"><span class="num">10 / 13</span></td><td class="colB"><span class="num">9 / 13</span></td></tr>
<tr><td class="metric">Total findings reported</td><td class="colA"><span class="num">16</span></td><td class="colB win"><span class="num">18</span></td></tr>
<tr><td class="metric">Findings beyond the 13 targets</td><td class="colA"><span class="num">6</span></td><td class="colB win"><span class="num">9</span> <span style="color:var(--muted);font-size:12px">(2 real: config leak, no-lockout)</span></td></tr>
<tr><td class="metric">Wall-clock time</td><td class="colA"><span class="num">32m 12s</span></td><td class="colB win"><span class="num">26m 53s</span></td></tr>
<tr><td class="metric">Criticals reported</td><td class="colA"><span class="num">5</span></td><td class="colB"><span class="num">2</span> <span style="color:var(--muted);font-size:12px">(recalibrated)</span></td></tr>
<tr><td class="metric">Belief-gate holds (POMDP)</td><td class="colA"><span class="num">3</span></td><td class="colB"><span class="num">—</span></td></tr>
<tr><td class="metric">Assurance P1–P5</td><td class="colA"><span class="num">all present</span></td><td class="colB"><span class="num">all present</span></td></tr>
<tr><td class="metric">Model API cost</td><td class="colA"><span class="num">$0</span> subscription</td><td class="colB"><span class="num">$0</span> + TypeSafe ≪ $5</td></tr>
</tbody>
</table>
</div>
</section>
<section>
<h2>Per-scenario coverage</h2>
<p class="lead">Each seeded vulnerability, and whether each run confirmed it. Neither run reached the
second-order SQLi or the CRLF header injection — the two that need a multi-step chain the single-vote
config didn't pursue.</p>
<div class="scen">
<div class="row head">
<div class="cell">Scenario</div>
<div class="cell mk hdrA">A</div>
<div class="cell mk hdrB">B·TS</div>
</div>
<!-- rows -->
<div class="row"><div class="cell"><span class="idc">web_sqli_login_bypass</span><span class="cls">SQLi</span></div><div class="cell mk hit">✓</div><div class="cell mk hit">✓</div></div>
<div class="row"><div class="cell"><span class="idc">web_sqli_union_search</span><span class="cls">SQLi</span></div><div class="cell mk hit">✓</div><div class="cell mk miss">✕</div></div>
<div class="row"><div class="cell"><span class="idc">web_sqli_blind_boolean</span><span class="cls">SQLi</span></div><div class="cell mk hit">✓</div><div class="cell mk hit">✓</div></div>
<div class="row"><div class="cell"><span class="idc">web_sqli_blind_time</span><span class="cls">SQLi</span></div><div class="cell mk hit">✓</div><div class="cell mk miss">✕</div></div>
<div class="row"><div class="cell"><span class="idc">web_sqli_second_order</span><span class="cls">SQLi</span></div><div class="cell mk miss">✕</div><div class="cell mk miss">✕</div></div>
<div class="row"><div class="cell"><span class="idc">web_xss_reflected_search</span><span class="cls">XSS</span></div><div class="cell mk hit">✓</div><div class="cell mk hit">✓</div></div>
<div class="row"><div class="cell"><span class="idc">web_xss_stored_review</span><span class="cls">XSS</span></div><div class="cell mk hit">✓</div><div class="cell mk hit">✓</div></div>
<div class="row"><div class="cell"><span class="idc">web_xss_svg_upload</span><span class="cls">XSS</span></div><div class="cell mk hit">✓</div><div class="cell mk hit">✓</div></div>
<div class="row"><div class="cell"><span class="idc">web_xss_dom_redirect</span><span class="cls">XSS</span></div><div class="cell mk hit">✓</div><div class="cell mk hit">✓</div></div>
<div class="row"><div class="cell"><span class="idc">web_idor_invoice</span><span class="cls">IDOR</span></div><div class="cell mk miss">✕</div><div class="cell mk hit">✓</div></div>
<div class="row"><div class="cell"><span class="idc">api_bola_orders</span><span class="cls">BOLA</span></div><div class="cell mk hit">✓</div><div class="cell mk hit">✓</div></div>
<div class="row"><div class="cell"><span class="idc">web_open_redirect_login</span><span class="cls">Redirect</span></div><div class="cell mk hit">✓</div><div class="cell mk hit">✓</div></div>
<div class="row"><div class="cell"><span class="idc">web_crlf_header_go</span><span class="cls">CRLF</span></div><div class="cell mk miss">✕</div><div class="cell mk miss">✕</div></div>
</div>
</section>
<section>
<h2>Severity shape</h2>
<p class="lead">The clearest effect of TypeSafe: the severity distribution flattens. The plain run stacks
five Criticals; the calibrated run keeps two and pushes the rest down to where the demonstrated-impact
evidence puts them.</p>
<div class="sev-legend"><span><i style="background:#e5484d"></i>Critical</span><span><i style="background:#f76b15"></i>High</span><span><i style="background:#f5b301"></i>Medium</span><span><i style="background:#3e7bfa"></i>Low</span><span><i style="background:#8b8698"></i>Info</span></div>
<div class="sev-wrap">
<div class="sevcard">
<h3 class="swatchA"><span class="dot" style="background:var(--a)"></span> A — no TypeSafe · 16</h3>
<div class="bar"><span class="lab"><i class="sw" style="background:#e5484d"></i>Critical</span><span class="track"><span class="fill" style="width:100%;background:#e5484d"></span></span><span class="n" style="color:#e5484d">5</span></div>
<div class="bar"><span class="lab"><i class="sw" style="background:#f76b15"></i>High</span><span class="track"><span class="fill" style="width:80%;background:#f76b15"></span></span><span class="n" style="color:#f76b15">4</span></div>
<div class="bar"><span class="lab"><i class="sw" style="background:#f5b301"></i>Medium</span><span class="track"><span class="fill" style="width:20%;background:#f5b301"></span></span><span class="n" style="color:#f5b301">1</span></div>
<div class="bar"><span class="lab"><i class="sw" style="background:#3e7bfa"></i>Low</span><span class="track"><span class="fill" style="width:40%;background:#3e7bfa"></span></span><span class="n" style="color:#3e7bfa">2</span></div>
<div class="bar"><span class="lab"><i class="sw" style="background:#8b8698"></i>Info</span><span class="track"><span class="fill" style="width:80%;background:#8b8698"></span></span><span class="n" style="color:#8b8698">4</span></div>
</div>
<div class="sevcard">
<h3 class="swatchB"><span class="dot" style="background:var(--b)"></span> B — TypeSafe · 18</h3>
<div class="bar"><span class="lab"><i class="sw" style="background:#e5484d"></i>Critical</span><span class="track"><span class="fill" style="width:40%;background:#e5484d"></span></span><span class="n" style="color:#e5484d">2</span></div>
<div class="bar"><span class="lab"><i class="sw" style="background:#f76b15"></i>High</span><span class="track"><span class="fill" style="width:80%;background:#f76b15"></span></span><span class="n" style="color:#f76b15">4</span></div>
<div class="bar"><span class="lab"><i class="sw" style="background:#f5b301"></i>Medium</span><span class="track"><span class="fill" style="width:60%;background:#f5b301"></span></span><span class="n" style="color:#f5b301">3</span></div>
<div class="bar"><span class="lab"><i class="sw" style="background:#3e7bfa"></i>Low</span><span class="track"><span class="fill" style="width:80%;background:#3e7bfa"></span></span><span class="n" style="color:#3e7bfa">4</span></div>
<div class="bar"><span class="lab"><i class="sw" style="background:#8b8698"></i>Info</span><span class="track"><span class="fill" style="width:100%;background:#8b8698"></span></span><span class="n" style="color:#8b8698">5</span></div>
</div>
</div>
</section>
<section>
<h2>What calibration actually did</h2>
<div class="callout">
<h3>The same BOLA, two severities</h3>
<p>Both runs found the object-level auth flaw on <code>GET /api/v2/users/:id</code> — a customer token
reads any user's full record, including the admin's plaintext password. The plain run rated it
<b>Critical (9.1)</b> on the class. TypeSafe, grading against the demonstrated-impact receipts and its
calibrated judgment, rated it <b>Low</b>.</p>
<div class="contrast">
<div class="cbox"><div class="t swatchA">A — class-graded</div><div class="g swatchA">Critical 9.1</div><div class="r">BOLA + excessive data exposure</div></div>
<div class="cbox"><div class="t swatchB">B — evidence-graded</div><div class="g swatchB">Low</div><div class="r">same finding, impact receipts weighted</div></div>
</div>
<p style="margin-top:14px"><b>Why it fired:</b> the severity is graded from the <em>structured</em> evidence
slot (the recorded request/response exchange), not the agent's prose. This finding proved the dump in its
narrative and claims ledger but left <code>evidence_data</code> null — so the demonstrated-impact rung saw no
machine-readable C/I/A receipt, and the calibrated grader dropped the impact metrics to <code>None</code>,
collapsing 9.1 → Low. The proof existed; it just wasn't in the slot the grader reads.</p>
<p>This is the honest edge: calibration removes inflated Criticals (good — most scanners over-rate by class),
but a receipt in the wrong slot gets under-rated. It is a dial toward defensibility, not a correctness
oracle — the operator still owns the final severity, and the fix is to make agents populate
<code>evidence_data</code> for impact, not to loosen the grader.</p>
</div>
</section>
<section>
<h2>Takeaways</h2>
<ul class="take">
<li><span class="tag even">tie</span><div><b>Recall is a wash.</b> 10 vs 9 of 13 is within run-to-run variance at vote-n 1. TypeSafe is not a recall multiplier — it is a judgment layer.</div></li>
<li><span class="tag plus">gain</span><div><b>Two real net-new findings.</b> The TypeSafe run surfaced a <code>config.json</code> API-key exposure (CWE-200) and a no-lockout brute-force (CWE-307) the plain run never reported — and it caught <code>web_idor_invoice</code>, which the plain run missed.</div></li>
<li><span class="tag plus">gain</span><div><b>Faster and calibrated.</b> 5m19s quicker, and it recalibrated 9 findings' confidence — collapsing five class-inflated Criticals to two evidence-backed ones.</div></li>
<li><span class="tag minus">cost</span><div><b>Conservatism has a price.</b> It dropped <code>union_search</code> and <code>blind_time</code>, and under-rated the credential-dump BOLA. A confirmation layer that demands receipts will sometimes discard a real thing it couldn't re-prove in-budget.</div></li>
<li><span class="tag note">cheap</span><div><b>Negligible cost.</b> TypeSafe adds no LLM tokens of its own — one probe call billed 319 in / 21 out. The whole run stayed far under the $5 budget.</div></li>
</ul>
</section>
<div class="disclaim">
<b>Method &amp; honesty.</b> Both runs: NeuroSploit v4.0.0, <code>claude-opus-4-8</code> via subscription,
black-box, recon intensity 2, single-model vote (<code>vote-n 1</code>), same natural-language focus naming
the 13 endpoints, no pre-baked solver — the LLM discovered and confirmed everything live. Recall is scored by
class + endpoint keyword match against the target's ground-truth list, so a match is coverage, not a graded
proof. <b>Confounders:</b> the two runs are single samples, not averages; an earlier TypeSafe run collapsed to
zero when the subscription hit a session limit mid-run (a real harness gap, since fixed — session-limit stdout
now parks the run instead of burning agents); vote-n 1 means no cross-model agreement in either arm. Treat this
as one honest data point on one target, not a leaderboard. <b>Not measured here:</b> multi-sample variance,
higher vote-n, and TypeSafe's agent-pruning effect on a broader agent set.
</div>
</div>
@@ -0,0 +1,122 @@
{
"engine": "neurosploit",
"version": "4.0.0",
"build": "49d3d3ceb1df",
"run": "ns-1789853137-localhost_3000",
"target": "http://localhost:3000",
"generated": 1789855069,
"findings": 16,
"artifacts": [
{
"name": "findings.json",
"present": true,
"sha256": "61ed87d0036ae5b2dd61ffd2c9ead34e11076f1cfbb8570f7554c072f3a76cb9",
"bytes": 91029,
"role": "the findings, each stamped with the engine build (P5)"
},
{
"name": "report.html",
"present": false,
"bytes": 0,
"role": "the human report"
},
{
"name": "recon.json",
"present": true,
"sha256": "8f5110c6d65cac10c4c04a8deacaf4cacbc8c8d320d18ed6cee236a7e61fe104",
"bytes": 15141,
"role": "reconnaissance facts"
},
{
"name": "audit.jsonl",
"present": true,
"sha256": "c6f63d9c2e70f59b05120a732ce157e23606ff388f232d31299545521818135b",
"bytes": 19254,
"role": "hash-chained decision log — every ALLOW/DENY (P1/P2/P4)"
},
{
"name": "audit.jsonl.anchors",
"present": true,
"sha256": "3b014612736ca9110c6342e61892286620ba2db605f82d278bec4de732e4bd1f",
"bytes": 213,
"role": "external anchors of the audit chain (P4)"
},
{
"name": "provenance.json",
"present": true,
"sha256": "c33f47d22ff52d82db3077f0eecceb105f73fb0f8cfe6f057c77a4696ddd1e4a",
"bytes": 297,
"role": "signed provenance manifest — build + structural signature (P5)"
},
{
"name": "out-of-scope-findings.json",
"present": true,
"sha256": "425df6ff6ac515da2b36f1bf4582d9acd8599e8c4e586e8dadc99b5601a06752",
"bytes": 17615,
"role": "findings quarantined for being outside scope (P2)"
},
{
"name": "flows.jsonl",
"present": false,
"bytes": 0,
"role": "intercepted request/response flows"
},
{
"name": "meta.json",
"present": true,
"sha256": "1e47c73f41061aef5e1943d3c8321f41349cf8e3588cfb1286a5627a226773cc",
"bytes": 198,
"role": "target metadata"
}
],
"properties": [
{
"id": "P1",
"name": "Signed authorization",
"status": "present",
"evidenced_by": [
"audit.jsonl"
],
"note": "capability recorded and decisions logged"
},
{
"id": "P2",
"name": "Scope enforcement",
"status": "present",
"evidenced_by": [
"audit.jsonl",
"out-of-scope-findings.json"
],
"note": "scope decisions recorded, including denials/quarantine"
},
{
"id": "P3",
"name": "Evidence & CVSS",
"status": "present",
"evidenced_by": [
"findings.json"
],
"note": "0/16 findings carry structured evidence · 14 with CVSS · 15 voted · 20 PoC(s) · 0 screenshot(s) · 13 evidence file(s)"
},
{
"id": "P4",
"name": "Audit integrity",
"status": "present",
"evidenced_by": [
"audit.jsonl",
"audit.jsonl.anchors"
],
"note": "hash chain plus signed anchors (truncation/rebuild detectable)"
},
{
"id": "P5",
"name": "Provenance",
"status": "present",
"evidenced_by": [
"provenance.json"
],
"note": "signed provenance manifest with structural signature"
}
],
"bundle_hash": "579449f887726db317b0169641be27dc5a11db46da698b87255d91dfaae9697f"
}
File diff suppressed because it is too large Load Diff
@@ -0,0 +1,10 @@
{
"asset": "NimbusCart Inc",
"brand": "NimbusCart Inc",
"server": "",
"status": 200,
"target": "http://localhost:3000",
"tech": [],
"title": "Home · NimbusCart",
"typesafe": false
}
File diff suppressed because one or more lines are too long
@@ -0,0 +1,120 @@
{
"engine": "neurosploit",
"version": "4.0.0",
"build": "49d3d3ceb1df",
"run": "ns-1789870577-localhost_3000",
"target": "http://localhost:3000",
"generated": 1789872190,
"findings": 18,
"artifacts": [
{
"name": "findings.json",
"present": true,
"sha256": "9827d2c67a851679ce8462fc1885d2c4ddfeb0a70180db293029f33d362d108f",
"bytes": 112054,
"role": "the findings, each stamped with the engine build (P5)"
},
{
"name": "report.html",
"present": false,
"bytes": 0,
"role": "the human report"
},
{
"name": "recon.json",
"present": true,
"sha256": "18a9d9456b5d239905a8c5a2d0647b272f8b9e5b5ff7f20c6ed26e7bf5164258",
"bytes": 11611,
"role": "reconnaissance facts"
},
{
"name": "audit.jsonl",
"present": true,
"sha256": "11e92303952192781686111381c17e37326ecedc1969b2c0d17b8d810fffebdd",
"bytes": 32535,
"role": "hash-chained decision log — every ALLOW/DENY (P1/P2/P4)"
},
{
"name": "audit.jsonl.anchors",
"present": true,
"sha256": "a0dd67ad35b530842fc0221ead9536b3ce19d45be251e8568b80909f4859d7cb",
"bytes": 213,
"role": "external anchors of the audit chain (P4)"
},
{
"name": "provenance.json",
"present": true,
"sha256": "ac45f0856813ca943713ff782a784e34ccc08038794fd5dfa00d6f060eac803c",
"bytes": 297,
"role": "signed provenance manifest — build + structural signature (P5)"
},
{
"name": "out-of-scope-findings.json",
"present": false,
"bytes": 0,
"role": "findings quarantined for being outside scope (P2)"
},
{
"name": "flows.jsonl",
"present": false,
"bytes": 0,
"role": "intercepted request/response flows"
},
{
"name": "meta.json",
"present": true,
"sha256": "c879fc77b942399b73b8050f258d00b4e671b38bdaa00ef7eebfac356484864c",
"bytes": 197,
"role": "target metadata"
}
],
"properties": [
{
"id": "P1",
"name": "Signed authorization",
"status": "present",
"evidenced_by": [
"audit.jsonl"
],
"note": "capability recorded and decisions logged"
},
{
"id": "P2",
"name": "Scope enforcement",
"status": "present",
"evidenced_by": [
"audit.jsonl"
],
"note": "scope decisions recorded, including denials/quarantine"
},
{
"id": "P3",
"name": "Evidence & CVSS",
"status": "present",
"evidenced_by": [
"findings.json"
],
"note": "0/18 findings carry structured evidence · 16 with CVSS · 17 voted · 16 PoC(s) · 0 screenshot(s) · 14 evidence file(s)"
},
{
"id": "P4",
"name": "Audit integrity",
"status": "present",
"evidenced_by": [
"audit.jsonl",
"audit.jsonl.anchors"
],
"note": "hash chain plus signed anchors (truncation/rebuild detectable)"
},
{
"id": "P5",
"name": "Provenance",
"status": "present",
"evidenced_by": [
"provenance.json"
],
"note": "signed provenance manifest with structural signature"
}
],
"bundle_hash": "85b0ef6f4789f08cb6bde0669b45aefb9f9f6bcb5f853a811c20f725a58e1eb1"
}
File diff suppressed because it is too large Load Diff
@@ -0,0 +1,10 @@
{
"asset": "NimbusCart Inc",
"brand": "NimbusCart Inc",
"server": "",
"status": 200,
"target": "http://localhost:3000",
"tech": [],
"title": "Home · NimbusCart",
"typesafe": true
}
File diff suppressed because one or more lines are too long
+85
View File
@@ -0,0 +1,85 @@
#!/usr/bin/env python3
# Score a NeuroSploit run against the user's 13 target scenarios (ground truth
# from BenchMarkBurpAT/SOLUTIONS.md). Matching is by class + endpoint keyword.
# No solver is involved — this only reads what the run reported.
import json, sys, re
# The 13 targets the user asked to attack: id -> (cwe-ish class, endpoint keywords)
TARGETS = {
"web_idor_invoice": ("idor", ["/account/invoice", "invoice"]),
"api_bola_orders": ("idor", ["/orders", "/v2/orders", "bola"]),
"web_sqli_login_bypass": ("sqli", ["/login"]),
"web_sqli_union_search": ("sqli", ["/shop/search", "search"]),
"web_sqli_blind_boolean": ("sqli", ["/shop/filter", "maxprice", "filter"]),
"web_sqli_blind_time": ("sqli", ["/support/feedback", "feedback", "comment"]),
"web_sqli_second_order": ("sqli", ["/account/profile", "/admin/search-users", "bio", "second"]),
"web_xss_reflected_search":("xss", ["/shop/search", "search"]),
"web_xss_stored_review": ("xss", ["/review", "/shop/product"]),
"web_xss_svg_upload": ("xss", ["/support/ticket", "/uploads", "svg"]),
"web_xss_dom_redirect": ("xss", ["/go", "dom", "?url", "name="]),
"web_open_redirect_login": ("redirect", ["/login", "next", "/go", "url="]),
"web_crlf_header_go": ("crlf", ["/go", "crlf", "header inject"]),
}
CLASS_CWE = {
"sqli": {"89","943","564"},
"xss": {"79","80","83","87"},
"idor": {"639","862","863","284","285","566","425","200"},
"redirect": {"601"},
"crlf": {"113","93"},
}
def classify(f):
cwe = "".join(ch for ch in f.get("cwe","") if ch.isdigit())
t = (f.get("title","")+" "+f.get("cwe","")).lower()
for cls, cwes in CLASS_CWE.items():
if cwe in cwes: return cls
for cls, kw in {"sqli":["sql inj","sqli"],"xss":["xss","cross-site scripting"],
"idor":["idor","bola","broken access","broken object"],
"redirect":["open redirect"],"crlf":["crlf","response splitting","header inject"]}.items():
if any(k in t for k in kw): return cls
return "other"
def endpoint_blob(f):
return " ".join(str(f.get(k,"")) for k in ("endpoint","title","payload","evidence")).lower()
def score(findings_path):
findings = json.load(open(findings_path))
hits = {} # target_id -> matched finding index
used = set()
for tid,(cls,kws) in TARGETS.items():
for i,f in enumerate(findings):
if i in used: continue
if classify(f)!=cls: continue
blob = endpoint_blob(f)
if any(kw.lower() in blob for kw in kws):
hits[tid]=i; used.add(i); break
tp = len(hits)
fn = [t for t in TARGETS if t not in hits]
# extra findings not matched to a target = out-of-scope-but-real OR noise;
# count as "extra" (not penalised as FP unless clearly bogus).
extra = [i for i in range(len(findings)) if i not in used]
return {
"total_findings": len(findings),
"targets_hit": tp,
"targets_total": len(TARGETS),
"recall": round(tp/len(TARGETS),3),
"hit_ids": sorted(hits.keys()),
"missed_ids": sorted(fn),
"extra_findings": len(extra),
}
if __name__=="__main__":
import os
for path in sys.argv[1:]:
fp = path if path.endswith(".json") else os.path.join(path,"findings.json")
try:
r = score(fp)
except Exception as e:
print(f"{path}: ERROR {e}"); continue
print(f"\n== {path} ==")
print(f" findings reported : {r['total_findings']}")
print(f" targets hit : {r['targets_hit']}/{r['targets_total']} (recall {r['recall']})")
print(f" hit : {', '.join(r['hit_ids']) or '—'}")
print(f" missed : {', '.join(r['missed_ids']) or '—'}")
print(f" extra findings : {r['extra_findings']}")
+14
View File
@@ -0,0 +1,14 @@
== /opt/neurosploit-rs/runs/ns-1789853137-localhost_3000 ==
findings reported : 16
targets hit : 10/13 (recall 0.769)
hit : api_bola_orders, web_open_redirect_login, web_sqli_blind_boolean, web_sqli_blind_time, web_sqli_login_bypass, web_sqli_union_search, web_xss_dom_redirect, web_xss_reflected_search, web_xss_stored_review, web_xss_svg_upload
missed : web_crlf_header_go, web_idor_invoice, web_sqli_second_order
extra findings : 6
== /opt/neurosploit-rs/runs/ns-1789855082-localhost_3000 ==
findings reported : 0
targets hit : 0/13 (recall 0.0)
hit : —
missed : api_bola_orders, web_crlf_header_go, web_idor_invoice, web_open_redirect_login, web_sqli_blind_boolean, web_sqli_blind_time, web_sqli_login_bypass, web_sqli_second_order, web_sqli_union_search, web_xss_dom_redirect, web_xss_reflected_search, web_xss_stored_review, web_xss_svg_upload
extra findings : 0
+2 -2
View File
@@ -929,7 +929,7 @@ dependencies = [
[[package]]
name = "neurosploit"
version = "4.0.0"
version = "4.1.0"
dependencies = [
"anyhow",
"clap",
@@ -946,7 +946,7 @@ dependencies = [
[[package]]
name = "neurosploit-harness"
version = "4.0.0"
version = "4.1.0"
dependencies = [
"anyhow",
"base64",
+1 -1
View File
@@ -3,7 +3,7 @@ members = ["crates/harness", "app"]
resolver = "2"
[workspace.package]
version = "4.0.0"
version = "4.1.0"
edition = "2021"
license = "MIT"
repository = "https://github.com/JoasASantos/NeuroSploit"
+4 -4
View File
@@ -1,4 +1,4 @@
//! NeuroSploit v4.0.0 — interactive harness + CLI (`run` / `whitebox` / `agents` / `models`).
//! NeuroSploit v4.1.0 — interactive harness + CLI (`run` / `whitebox` / `agents` / `models`).
mod rectify;
mod repl;
@@ -12,8 +12,8 @@ use std::path::{Path, PathBuf};
#[command(
name = "neurosploit",
version,
about = "NeuroSploit v4.0.0 — multi-model autonomous pentest harness",
long_about = "NeuroSploit v4.0.0 — a Rust multi-model harness that drives a pool of LLMs \
about = "NeuroSploit v4.1.0 — multi-model autonomous pentest harness",
long_about = "NeuroSploit v4.1.0 — a Rust multi-model harness that drives a pool of LLMs \
(API key or local subscription: Claude/Codex/Gemini/Grok/OpenCode/Hermes) to autonomously test a target. \
After recon it INTELLIGENTLY selects only the agents matching the discovered surface, runs \
them in parallel, then validates every finding by cross-model voting before reporting.\n\n\
@@ -1011,7 +1011,7 @@ pub(crate) fn spawn_engagement(base: &Path, mut cfg: RunConfig, mcp: bool, mode:
println!(" │ ua : {ua}");
write_status(&workdir, "running", &format!("\"target\":{:?}", cfg.target));
println!(" ┌─ NeuroSploit v4.0.0 · by Joas A Santos & Red Team Leaders");
println!(" ┌─ NeuroSploit v4.1.0 · by Joas A Santos & Red Team Leaders");
println!(" │ run id : {run_id}");
println!(" │ target : {}", cfg.target);
println!(" │ models : {}", cfg.models.join(", "));
+2 -2
View File
@@ -1,4 +1,4 @@
//! NeuroSploit v4.0.0 — interactive session (Claude-Code / Codex / Cursor-CLI style).
//! NeuroSploit v4.1.0 — interactive session (Claude-Code / Codex / Cursor-CLI style).
//!
//! Launched when `neurosploit` runs with no subcommand. A persistent REPL with
//! real line editing (arrow-key history recall, Ctrl-A/E/K, paste), model
@@ -440,7 +440,7 @@ pub async fn repl(base: &Path, auth: SessionAuth) -> anyhow::Result<()> {
let backends = harness::installed_cli_backends();
println!("\x1b[1m");
println!(" ███╗ ██╗███████╗██╗ ██╗██████╗ ██████╗");
println!(" ████╗ ██║██╔════╝██║ ██║██╔══██╗██╔═══██╗ NeuroSploit v4.0.0");
println!(" ████╗ ██║██╔════╝██║ ██║██╔══██╗██╔═══██╗ NeuroSploit v4.1.0");
println!(" ██╔██╗ ██║█████╗ ██║ ██║██████╔╝██║ ██║ interactive harness");
println!(" ██║╚██╗██║██╔══╝ ██║ ██║██╔══██╗██║ ██║ by Joas A Santos");
println!(" ██║ ╚████║███████╗╚██████╔╝██║ ██║╚██████╔╝ & Red Team Leaders");
+1 -1
View File
@@ -1,4 +1,4 @@
//! NeuroSploit v4.0.0 — TUI "Mission Control" mode.
//! NeuroSploit v4.1.0 — TUI "Mission Control" mode.
//!
//! Concurrent panels that update live while the engagement runs in the
//! background, with a composer input that stays active during execution:
+1 -1
View File
@@ -242,7 +242,7 @@ pub fn html_with_pocs(target: &str, findings: &[Finding], meta: &EngagementMeta,
<h2>Executive Summary</h2><div class=summary-grid>{summary_grid}</div>\
{vuln_summary}\
<h2>Findings ({n})</h2>{body}\
<p class=footer>Authorized testing only. Confirmed findings passed multi-model voting, receipt grounding and adversarial refute; \"needs-review\" are flagged for a human.<br>NeuroSploit v4.0.0 · by <b>Joas A Santos</b> &amp; <b>Red Team Leaders</b><br><span style=\"font-family:ui-monospace,monospace\">{provenance}</span></p></body></html>",
<p class=footer>Authorized testing only. Confirmed findings passed multi-model voting, receipt grounding and adversarial refute; \"needs-review\" are flagged for a human.<br>NeuroSploit v4.1.0 · by <b>Joas A Santos</b> &amp; <b>Red Team Leaders</b><br><span style=\"font-family:ui-monospace,monospace\">{provenance}</span></p></body></html>",
t = esc(target), n = sorted.len(), body = body, summary_grid = summary_grid, vuln_summary = vuln_summary,
// Which build produced this document. A report that circulates without
// it is a report nobody can trace back to the run that made it.
+2 -2
View File
@@ -1,4 +1,4 @@
// NeuroSploit v3.5.1 — Typst report template (blank, structured).
// NeuroSploit v4.1.0 — Typst report template (blank, structured).
//
// The harness generates `report.typ` per run by prepending a `findings` array
// and a `meta` dict, then including this template's rendering logic. This file
@@ -53,7 +53,7 @@
#set page(margin: 2cm, numbering: "1", footer: context [
#set text(size: 8pt, fill: gray)
NeuroSploit v4.0.0 · #meta.target · confidential
NeuroSploit v4.1.0 · #meta.target · confidential
#h(1fr)
// Build+run identity, so a page that circulates on its own still says which
// engagement produced it.
+1 -1
View File
@@ -22,7 +22,7 @@ run this only on a trusted machine/network, same trust model as the CLI itself.
Server/version info.
```json
{ "version": "4.0.0", "binary": "/opt/neurosploit-rs/neurosploit-rs/target/release/neurosploit", "root": "/opt/neurosploit-rs" }
{ "version": "4.1.0", "binary": "/opt/neurosploit-rs/neurosploit-rs/target/release/neurosploit", "root": "/opt/neurosploit-rs" }
```
---
+1 -1
View File
@@ -1,4 +1,4 @@
# NeuroSploit v4.0.0 — web console
# NeuroSploit v4.1.0 — web console
A browser UI for the `neurosploit` CLI harness: a 5-step engagement wizard (Asset → Scope & Auth
→ Leads → Model & Run → Review), a live structured findings view with a generative attack-path
+2 -2
View File
@@ -1,8 +1,8 @@
{
"name": "neurosploit-web",
"version": "4.0.0",
"version": "4.1.0",
"private": true,
"description": "NeuroSploit v4.0.0 web console — lead board + REPL, backed by the neurosploit CLI harness.",
"description": "NeuroSploit v4.1.0 web console — lead board + REPL, backed by the neurosploit CLI harness.",
"main": "server.js",
"scripts": {
"start": "node server.js"
+1 -1
View File
@@ -1,5 +1,5 @@
'use strict';
/* NeuroSploit v4.0.0 — web console frontend. Vanilla JS, no build step. */
/* NeuroSploit v4.1.0 — web console frontend. Vanilla JS, no build step. */
const $ = (sel, root = document) => root.querySelector(sel);
const $$ = (sel, root = document) => Array.from(root.querySelectorAll(sel));
+2 -2
View File
@@ -3,7 +3,7 @@
<head>
<meta charset="utf-8" />
<meta name="viewport" content="width=device-width, initial-scale=1" />
<title>NeuroSploit v4.0.0 — Console</title>
<title>NeuroSploit v4.1.0 — Console</title>
<link rel="icon" href="data:image/svg+xml,<svg xmlns=%22http://www.w3.org/2000/svg%22 viewBox=%220 0 100 100%22><text y=%22.9em%22 font-size=%2290%22>🧠</text></svg>">
<link rel="stylesheet" href="/vendor/xterm.css" />
<link rel="stylesheet" href="/style.css" />
@@ -33,7 +33,7 @@
<div class="sb-groups" id="sbGroups"><!-- populated by app.js --></div>
<div class="sb-bottom">
<span class="sb-version" id="sbVersion">v4.0.0</span>
<span class="sb-version" id="sbVersion">v4.1.0</span>
<div class="sb-bottom-actions">
<button class="icon-btn" id="btnOpenAuth" title="Auth &amp; API keys">🔑</button>
<button class="icon-btn" id="btnOpenRepl" title="Open terminal (Ctrl+`)">❭_</button>
+1 -1
View File
@@ -1,4 +1,4 @@
/* NeuroSploit v4.0.0 — web console.
/* NeuroSploit v4.1.0 — web console.
Visual direction: dense security-operations console (not a marketing SaaS
page). Borders over shadows, typography over color, two radii, one accent.
*/
+2 -2
View File
@@ -1,7 +1,7 @@
#!/usr/bin/env node
'use strict';
/**
* NeuroSploit v4.0.0 — web console backend.
* NeuroSploit v4.1.0 — web console backend.
*
* Zero-dependency Node HTTP server that:
* - serves the static SPA in ./public
@@ -1223,7 +1223,7 @@ const server = http.createServer(async (req, res) => {
});
server.listen(PORT, () => {
console.log(`NeuroSploit v4.0.0 web console → http://localhost:${PORT}`);
console.log(`NeuroSploit v4.1.0 web console → http://localhost:${PORT}`);
console.log(` binary : ${BIN || '(not found — build neurosploit-rs first)'}`);
console.log(` agents : ${AGENTS_DIR}`);
console.log(` runs : ${RUNS_DIR}`);