mirror of
https://github.com/CyberSecurityUP/NeuroSploit.git
synced 2026-09-30 04:51:50 +02:00
release: v4.1.0 — assurance layer, TypeSafe, hardening + benchmark
Version bumped to 4.1.0 across the workspace, binaries, web console and Typst template. README: new "New in v4.1.0" summary; trimmed the verbose highlight bullets and the TypeSafe section; removed the anti-plagiarism/provenance section (provenance stays in the code, just not front-and-centre in the README); TypeSafe promoted to its own top-level section; agent count 446. TUTORIAL: new section 17 "Assurance & authorization" covering the target gate, --scope-file, evidence-graded CVSS, audit anchoring + assurance bundle, sandbox, intercept proxy, PoC re-validation, compliance mapping, TypeSafe, and the internal/AD graph + budget governor. benchmarks/typesafe-2026-09-20/: the with/without TypeSafe measurement — report.html, scorer, both runs' findings/assurance/meta/logs, and a README. No secrets committed (env-only during the runs, verified clean). 381 tests. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
co-authored by
Claude Opus 5
parent
d2ec0a112d
commit
088d133c80
@@ -1,4 +1,4 @@
|
||||
<h1 align="center">🧠 NeuroSploit v4.0.0</h1>
|
||||
<h1 align="center">🧠 NeuroSploit v4.1.0</h1>
|
||||
|
||||
<p align="center">
|
||||
<a href="https://github.com/JoasASantos/NeuroSploit/stargazers"><img src="https://img.shields.io/github/stars/JoasASantos/NeuroSploit?style=for-the-badge&logo=github&color=8b5cf6" alt="Stars"></a>
|
||||
@@ -8,10 +8,10 @@
|
||||
</p>
|
||||
|
||||
<p align="center">
|
||||
<img src="https://img.shields.io/badge/Version-4.0.0-blue?style=flat-square">
|
||||
<img src="https://img.shields.io/badge/Version-4.1.0-blue?style=flat-square">
|
||||
<img src="https://img.shields.io/badge/Harness-Rust%20%7C%20tokio-e6b673?style=flat-square">
|
||||
<img src="https://img.shields.io/badge/License-MIT-green?style=flat-square">
|
||||
<img src="https://img.shields.io/badge/MD%20Agents-435-red?style=flat-square">
|
||||
<img src="https://img.shields.io/badge/MD%20Agents-446-red?style=flat-square">
|
||||
<img src="https://img.shields.io/badge/Models-18%20providers-success?style=flat-square">
|
||||
<img src="https://img.shields.io/badge/Modes-Black%20%7C%20White%20%7C%20Grey%20%7C%20Host%20%7C%20AI-9cf?style=flat-square">
|
||||
<img src="https://img.shields.io/badge/Auth-API%20key%20%7C%20Subscription-orange?style=flat-square">
|
||||
@@ -50,50 +50,43 @@ Control TUI**.
|
||||
|
||||
### Highlights
|
||||
|
||||
- 🧠 **POMDP belief + value-of-information** — the target is partially observable,
|
||||
so findings aren't booleans: a property-graph **belief** carries probabilities,
|
||||
and "scan more vs exploit now" falls out of belief entropy. The `may_assert`
|
||||
gate is a **mathematical anti-hallucination rule** (don't claim exploitability
|
||||
while the belief is diffuse).
|
||||
- 🧾 **Grounding** — hard rule: **no claim without a receipt** (evidence, not
|
||||
paraphrase). Empirical (raw tool output) for black-box/host/AI, **symbolic**
|
||||
(`file:line` into the reviewed source — a code citation *is* the receipt) for
|
||||
white-box SAST & skills audits, and **either** for grey-box; ungrounded claims
|
||||
are demoted.
|
||||
- 🔬 **Deterministic HTTP probe** — before the model recon, the harness runs a
|
||||
**real** request/response analysis (status/redirects, security headers, cookie
|
||||
flags, CORS reflection, tech fingerprint, linked JS, 404 baseline, high-signal
|
||||
paths) and feeds those observed facts into recon, so agent selection and
|
||||
exploitation decisions are grounded in evidence — not the model's guess.
|
||||
- 🔗 **Attack chaining — any primitive pivots.** 13 multi-stage chain agents
|
||||
(SQLi→RCE→LPE, SSRF→cloud creds, upload→LFI→RCE→LPE, CVE→RCE→pivot, …) **plus a
|
||||
chaining doctrine** that turns *any* confirmed foothold into the next step:
|
||||
reduce it to a primitive (exec / read / write / request-forgery / identity /
|
||||
secret) and pivot — file-upload→RCE, SSRF→metadata creds, IDOR→takeover — reusing
|
||||
looted creds and reasoning about **business logic** (payment/tenancy/workflow
|
||||
abuse). Each stage proven; strictly non-destructive (no data loss, no DB
|
||||
overwrite, no DoS).
|
||||
- ☁️ **Cloud testing** — AWS / GCP / Azure agents that drive the provider CLIs
|
||||
(`aws`/`gcloud`/`az`). Connect via `creds.yaml`: AWS keys, a Google
|
||||
service-account JSON, or an Azure service principal — see
|
||||
[Cloud credentials](#cloud-credentials-awsgcpazure).
|
||||
- 🤖 **LLM red-teaming** — 30 AI agents that jailbreak & prompt-inject a live AI
|
||||
system across scenarios: **AdvPrefix**, **PAIR**, **TAP**, **Crescendo**,
|
||||
many-shot, persona/DAN, encoding/obfuscation, refusal-suppression; plus
|
||||
**indirect injection** (RAG/web/email/tool output), **goal hijacking**,
|
||||
tool/function-call abuse, and system-prompt exfiltration. Each runs an
|
||||
attacker→**LLM-judge** loop (baseline refusal → technique → verdict) and proves
|
||||
the bypass with a **benign, redacted** receipt. Maps to OWASP LLM Top 10 (2025),
|
||||
MCP threats & OWASP AI Exchange; Skill/plugin & **n8n** files audited white-box.
|
||||
- 🧰 **Misconfig & CVE hunting → exploitation, safely** — a full CVE pipeline:
|
||||
**version fingerprint** (pin exact versions) → **research analyst** (map to
|
||||
NVD/GHSA CVEs, judge reachability) → **PoC finder** (locate/vet/adapt a public
|
||||
PoC) → **exploit scripter** (write a custom exploit when none exists). Every PoC
|
||||
is written to the run's **`pocs/` folder and referenced in the report** so
|
||||
findings are reproducible. Plus absurd-misconfig agents (exposed `.git`/`.env`,
|
||||
debug/actuator, default creds, dashboards, CORS) and rate-limit testing — all
|
||||
under a strict **data-safety/PII guardrail** (no destructive/state-changing
|
||||
actions; PII proven with a masked sample, never dumped).
|
||||
> **New in v4.1.0** — evidence-graded CVSS computed from the FIRST v3.1 equation
|
||||
> (not guessed by class); a **target-authorization gate** (default-deny, refuses a
|
||||
> target outside the capability grant before any recon); **audit anchoring** that
|
||||
> detects truncation & silent rebuilds; a signed **assurance bundle** (P1–P5 in one
|
||||
> manifest per run); **scope-evasion resistance** (alt-IP-encoding normalization,
|
||||
> redirect-to-private-IP block, DNS-rebinding guard); **evidence-integrity** checks
|
||||
> (cross-target / reused-receipt / foreign-marker / orphan-claim rejection);
|
||||
> **untrusted-output taint** (prompt-injection stripping + data fencing); a
|
||||
> **`--scope-file` YAML loader** + web Scoping/Guardrails UI; a **Kali sandbox**
|
||||
> (`--sandbox`), **intercept proxy** (`--intercept burp|caido|zap|mitmproxy|own`),
|
||||
> **PoC re-validation** (`--revalidate-poc`), **compliance mapping**
|
||||
> (`--compliance pci-dss,hipaa,soc2`); an **internal-network / AD attack graph**;
|
||||
> a **reasoning-budget governor** (`--budget`); and **TypeSafe System One**
|
||||
> (`--typesafe on|off|auto`) as a calibrated confirmation + adjudication layer.
|
||||
> 27 deterministic per-CWE validators, 446 agents. See
|
||||
> [benchmarks/typesafe-2026-09-20](benchmarks/typesafe-2026-09-20/) for a
|
||||
> with/without measurement.
|
||||
|
||||
- 🧠 **POMDP belief + anti-hallucination gate** — findings aren't booleans; a
|
||||
property-graph belief carries probabilities, and `may_assert` refuses to claim
|
||||
exploitability while the belief is diffuse.
|
||||
- 🧾 **Grounding — no claim without a receipt.** Empirical (raw tool output) or
|
||||
symbolic (`file:line` into the reviewed source); ungrounded claims are demoted.
|
||||
- 🔬 **Deterministic HTTP probe** feeds observed facts (headers, cookies, CORS,
|
||||
fingerprint, JS, 404 baseline) into recon — decisions grounded in evidence,
|
||||
not the model's guess.
|
||||
- 🔗 **Attack chaining — any primitive pivots.** Reduce a foothold to a primitive
|
||||
(exec/read/write/request-forgery/identity/secret) and pivot; each stage proven,
|
||||
strictly non-destructive.
|
||||
- ☁️ **Cloud testing** — AWS / GCP / Azure agents driving `aws`/`gcloud`/`az` via
|
||||
`creds.yaml` ([details](#cloud-credentials-awsgcpazure)).
|
||||
- 🤖 **LLM red-teaming** — jailbreak & prompt-inject a live AI system (AdvPrefix,
|
||||
PAIR, TAP, Crescendo, indirect injection, goal hijacking) via an attacker→judge
|
||||
loop; maps to OWASP LLM Top 10.
|
||||
- 🧰 **Misconfig & CVE pipeline** — fingerprint → CVE research → PoC finder →
|
||||
exploit scripter; every PoC written to `pocs/` and referenced in the report,
|
||||
under a strict data-safety/PII guardrail.
|
||||
- 🎯 **Re-test one vulnerability** — `--only <agent>` (repeatable /
|
||||
comma-separated) runs exactly the agent(s) you name and skips recon-based
|
||||
selection — re-test a single finding fast. Works on `run` / `whitebox` /
|
||||
@@ -495,63 +488,25 @@ neurosploit run https://app.example --creds creds.yaml \
|
||||
Each finding is proven with the **authorized vs unauthorized** request pair, under
|
||||
the data-safety guardrail (read-only, PII masked).
|
||||
|
||||
## 🏷️ Identification & attribution (anti-plagiarism)
|
||||
## 🧮 TypeSafe System One — calibrated adjudication
|
||||
|
||||
Every request is tagged with an identifying **User-Agent** (default
|
||||
`NeuroSploit/<ver> …`, change with **`/ua`** or `NEUROSPLOIT_UA`) plus an
|
||||
`X-NeuroSploit-Scan` header, and every finding is **stamped** "Identified and
|
||||
validated by NeuroSploit" — so provenance travels in the traffic, the finding
|
||||
text, `findings.json` and the report footer.
|
||||
|
||||
### Provenance — which build made this, and does it still match
|
||||
|
||||
Attribution that survives someone else's copy-paste:
|
||||
|
||||
- **`JOASNSCOPE`** leads every canary the harness mints, so a marker that
|
||||
turns up later — in a response body, a customer's log, somebody else's
|
||||
report — extracts whole and names the build that made it.
|
||||
- **Per-build fingerprint** (`neurosploit provenance show`), plus an optional
|
||||
per-customer build id via `NEUROSPLOIT_CUSTOMER_ID`.
|
||||
- **`findings.json` is stamped** with `_engine`, and a **signed
|
||||
`provenance.json`** ships beside it (`NEUROSPLOIT_PROVENANCE_KEY`).
|
||||
- **Structural signature** over the finding set's shape — it survives
|
||||
rewording and reformatting, but not a changed result.
|
||||
- **Prompts are watermarked** at the single model-pool chokepoint
|
||||
(`NEUROSPLOIT_WATERMARK=off` to disable).
|
||||
Set `TYPESAFE_API_KEY` and NeuroSploit adjudicates each finding with TypeSafe's
|
||||
System One model (Jev): a calibrated `{confirmed, needs-review, rejected}`
|
||||
judgment over the *evidence* (not the prose), plus a check on whether real
|
||||
impact was demonstrated. It refines confidence, re-grades CVSS when impact is
|
||||
unproven, and runs a code-owned confirmation loop over enumerable classes
|
||||
(XSS/SQLi/redirect/traversal/SSRF/IDOR). **Additive** — a deterministic
|
||||
validator still rules; TypeSafe only lowers confidence or flags for review,
|
||||
never resurrects a rejected claim.
|
||||
|
||||
```bash
|
||||
neurosploit provenance show # this build's identity
|
||||
neurosploit provenance scan report.pdf.txt # is this ours? which build?
|
||||
neurosploit provenance verify runs/ns-… # manifest vs findings
|
||||
neurosploit run https://app --typesafe on # calibrated adjudication + confirmation
|
||||
neurosploit run https://app --typesafe off # the identical pipeline, no TypeSafe (A/B)
|
||||
```
|
||||
|
||||
### TypeSafe System One — calibrated adjudication (RLCD)
|
||||
|
||||
When `TYPESAFE_API_KEY` is set, each finding is adjudicated by TypeSafe's
|
||||
System One model (Jev) — a **calibrated decision** over its *evidence*, not its
|
||||
prose: a `Choice` of `{confirmed, needs-review, rejected}` with a probability
|
||||
distribution, plus a `Noul` on whether real impact was demonstrated. The result
|
||||
refines the finding's confidence and moves borderline cases to needs-review.
|
||||
|
||||
It is **additive**: a deterministic validator still rules (a rejected finding
|
||||
stays rejected), and TypeSafe can only lower confidence or flag for review,
|
||||
never resurrect a claim. Every adjudication is written to the audit trail.
|
||||
Disable with `NEUROSPLOIT_TYPESAFE=off`. This is the RLCD (Reinforcement
|
||||
Learning for Calibrated Decisions) tier of the model stack — typed judgments
|
||||
where the harness needs a number, not a paragraph.
|
||||
|
||||
**As an additional confirmation strategy** (`typesafe_agent`), a code-owned loop
|
||||
where TypeSafe picks the next payload (`Choice`) and judges the real response
|
||||
(`Noul`) over the replay engine — for enumerable classes (XSS, SQLi, open
|
||||
redirect, path traversal, SSRF, IDOR). It runs only on findings the LLM path
|
||||
left unconfirmed or in needs-review (the recall lever), can only raise a finding
|
||||
to confirmed with a calibrated probability, never downgrades, and refuses edge
|
||||
(WAF) responses. It is **not** a discovery agent — System One does not generate.
|
||||
|
||||
**Flag & A/B.** `--typesafe on|off|auto` (default auto = on when the key is set).
|
||||
`off` runs the *identical* pipeline without it, and the run's `meta.json` records
|
||||
`"typesafe": true|false` — so a with/without pair against the same target is a
|
||||
clean measurement of what it adds.
|
||||
`--typesafe auto` (default) is on when the key is set. Each run's `meta.json`
|
||||
records `"typesafe": true|false` — a clean with/without measurement, one of
|
||||
which lives in [`benchmarks/typesafe-2026-09-20/`](benchmarks/typesafe-2026-09-20/).
|
||||
|
||||
### Scope-evasion resistance, evidence integrity, untrusted output
|
||||
|
||||
|
||||
+108
-4
@@ -1,4 +1,4 @@
|
||||
# NeuroSploit — Tutorial & User Guide (v4.0.0)
|
||||
# NeuroSploit — Tutorial & User Guide (v4.1.0)
|
||||
|
||||
A complete, hands-on guide to installing, configuring and running NeuroSploit —
|
||||
the autonomous, multi-model penetration-testing harness.
|
||||
@@ -31,7 +31,8 @@ the autonomous, multi-model penetration-testing harness.
|
||||
14. [The agent library](#14-the-agent-library)
|
||||
15. [Playwright MCP & extra tools](#15-playwright-mcp--extra-tools)
|
||||
16. [Tips, tuning & troubleshooting](#16-tips-tuning--troubleshooting)
|
||||
17. [Command & flag reference](#17-command--flag-reference)
|
||||
17. [Assurance & authorization (v4.1.0)](#17-assurance--authorization)
|
||||
18. [Command & flag reference](#18-command--flag-reference)
|
||||
|
||||
---
|
||||
|
||||
@@ -99,7 +100,7 @@ Agents **degrade gracefully**: if `rustscan` is absent they use `nmap`; if neith
|
||||
### Verify
|
||||
|
||||
```bash
|
||||
neurosploit --version # neurosploit 4.0.0
|
||||
neurosploit --version # neurosploit 4.1.0
|
||||
neurosploit agents # {"vulns":241,...,"ai":30,...,"total":430}
|
||||
neurosploit models # all providers & models
|
||||
```
|
||||
@@ -713,7 +714,110 @@ back to `curl`. You can add more MCP servers by placing a `mcp.servers.json`
|
||||
|
||||
---
|
||||
|
||||
## 17. Command & flag reference
|
||||
## 17. Assurance & authorization
|
||||
|
||||
v4.1.0 adds a layer of controls that make a run **defensible**, not just
|
||||
productive. All are enforced in code (not prompt text) and every decision lands
|
||||
in the hash-chained audit trail.
|
||||
|
||||
### Target authorization gate (default-deny)
|
||||
|
||||
Before any recon, the target is checked against the capability grant — protocol,
|
||||
host, port, URL prefix. A signed token that does not cover the target **refuses
|
||||
the run** and exits non-zero:
|
||||
|
||||
```bash
|
||||
# mint a grant for one host, then run against a different one → refused
|
||||
neurosploit capability issue --scope app.example.com --issuer you --subject op --hours 8
|
||||
neurosploit run https://other.example.com --capability-token <tok>
|
||||
# ⛔ DENY_TARGET_OUTSIDE_GRANT — other.example.com is outside the authorized scope
|
||||
# (non-zero exit; nothing was tested; the denial is audited)
|
||||
```
|
||||
|
||||
Loopback (`localhost`/`127.0.0.1`) is exempt — it is unambiguous.
|
||||
|
||||
### Hard scope from a file
|
||||
|
||||
```bash
|
||||
neurosploit run https://app.example.com --scope-file scope.yaml
|
||||
```
|
||||
|
||||
```yaml
|
||||
# scope.yaml — enforced in code; a capability token still caps it
|
||||
hard: [ app.example.com, "*.staging.example.com", 10.20.30.0/24 ]
|
||||
exclude: [ payments.example.com ]
|
||||
soft:
|
||||
observe_only: [ cdn.example.com ]
|
||||
allow_destructive_methods: false
|
||||
max_requests_per_minute: 240
|
||||
forbidden_payloads: [ "drop table", "rm -rf /" ]
|
||||
notes: [ "SOW-2026-0142; window 02:00-06:00 UTC" ]
|
||||
```
|
||||
|
||||
Alt-IP encodings (`0x7f000001`, `2130706433`, `0177.0.0.1`,
|
||||
`::ffff:127.0.0.1`) all normalize to dotted-quad, so an exclude can't be dodged
|
||||
by re-spelling; redirects to a private/loopback address are refused; a
|
||||
DNS-rebinding guard refuses a name that re-resolves to a new internal address.
|
||||
|
||||
### Evidence-graded CVSS
|
||||
|
||||
The score is computed from the FIRST v3.1 equation, and each impact metric is
|
||||
graded against a receipt. SQLi that reached the interpreter but extracted
|
||||
nothing scores **demonstrated 0 / potential 9.8** — never a manufactured
|
||||
critical. The vector travels with the number in the report.
|
||||
|
||||
### Audit anchoring & the assurance bundle
|
||||
|
||||
```bash
|
||||
neurosploit audit <run> --anchor # chain + signed anchors: catch truncation/rebuild/forgery
|
||||
neurosploit assurance <run> # P1–P5 in one manifest (authorization/enforcement/evidence/integrity/provenance)
|
||||
neurosploit assurance <run> --verify # re-hash every artifact + check the signature
|
||||
```
|
||||
|
||||
Set `NEUROSPLOIT_ANCHOR_DIR` to also write anchors to external append-only
|
||||
(ideally WORM) storage, and `NEUROSPLOIT_PROVENANCE_KEY` to sign them.
|
||||
|
||||
### Tooling: sandbox · proxy · PoC re-validation · compliance
|
||||
|
||||
```bash
|
||||
--sandbox # run agent commands in a Kali container (docker/podman)
|
||||
--intercept burp|caido|zap|mitmproxy # route through a tool …
|
||||
--intercept own | own+burp # … or the harness's own recording interceptor
|
||||
--revalidate-poc # re-run each PoC; demote what no longer reproduces
|
||||
--compliance pci-dss,hipaa,soc2 # map findings onto control requirements in the report
|
||||
```
|
||||
|
||||
### TypeSafe System One (calibrated confirmation)
|
||||
|
||||
```bash
|
||||
export TYPESAFE_API_KEY=... # then:
|
||||
neurosploit run https://app --typesafe on # calibrated adjudication + confirmation loop
|
||||
neurosploit run https://app --typesafe off # the identical pipeline, no TypeSafe (for A/B)
|
||||
```
|
||||
|
||||
`--typesafe auto` (default) turns it on when the key is set. It adjudicates each
|
||||
finding with a calibrated `{confirmed/needs-review/rejected}` judgment over the
|
||||
*evidence*, re-grades CVSS when impact isn't demonstrated, prunes irrelevant
|
||||
agents, and runs a code-owned confirmation loop over enumerable classes. It is
|
||||
**additive** — a deterministic validator still rules; TypeSafe can only lower
|
||||
confidence or flag for review, never resurrect a rejected claim. A with/without
|
||||
measurement lives in [`benchmarks/typesafe-2026-09-20/`](benchmarks/typesafe-2026-09-20/).
|
||||
|
||||
### Internal network / AD & reasoning budget
|
||||
|
||||
```bash
|
||||
neurosploit internal --graph g.json --scaffold corp.local --from foothold --mermaid
|
||||
neurosploit run https://app --budget eco|balanced|aggressive # ration reasoning; default unlimited
|
||||
```
|
||||
|
||||
The internal graph models an engagement as `Asset → Exposure → Weakness →
|
||||
Credential → Privilege → Movement → Crown Jewel` and answers the question a
|
||||
CVSS-sorted list can't: **which single edge, removed, cuts the most paths to the
|
||||
crown jewels** (`choke_points`).
|
||||
|
||||
---
|
||||
|
||||
## 18. Command & flag reference
|
||||
|
||||
```
|
||||
neurosploit # interactive REPL (resumes per project)
|
||||
|
||||
@@ -0,0 +1,83 @@
|
||||
# NeuroSploit × TypeSafe — benchmark (2026-09-20)
|
||||
|
||||
Two identical NeuroSploit engagements against the same vulnerable target — one
|
||||
plain, one with **TypeSafe System One (Jev)** as a calibrated confirmation
|
||||
layer. Same model, same focus, same 13 seeded vulnerabilities. Only the
|
||||
`--typesafe` flag differs.
|
||||
|
||||
Open **`report.html`** for the full visual write-up.
|
||||
|
||||
## Setup
|
||||
|
||||
| | |
|
||||
|---|---|
|
||||
| Harness | NeuroSploit v4.0.0 |
|
||||
| Model | `claude-opus-4-8` (subscription) |
|
||||
| Target | NimbusCart / BenchMarkBurpAT · `http://localhost:3000` |
|
||||
| Mode | black-box, `--recon 2`, `--vote-n 1`, `--max-agents 15` |
|
||||
| Ground truth | 13 seeded scenarios (IDOR/BOLA, SQLi ×5, XSS ×4, open redirect, CRLF) |
|
||||
| Solver | none — the LLM discovered and confirmed everything live |
|
||||
|
||||
Run commands (the only difference is `--typesafe`):
|
||||
|
||||
```bash
|
||||
# A — no TypeSafe
|
||||
NEUROSPLOIT_TYPESAFE=off neurosploit run http://localhost:3000 \
|
||||
--subscription --model anthropic:claude-opus-4-8 \
|
||||
--typesafe off --recon 2 --max-agents 15 --vote-n 1 --focus "<13 endpoints>" -v
|
||||
|
||||
# B — with TypeSafe (TYPESAFE_API_KEY set in env, never committed)
|
||||
NEUROSPLOIT_TYPESAFE=on neurosploit run http://localhost:3000 \
|
||||
--subscription --model anthropic:claude-opus-4-8 \
|
||||
--typesafe on --recon 2 --max-agents 15 --vote-n 1 --focus "<13 endpoints>" -v
|
||||
```
|
||||
|
||||
## Result
|
||||
|
||||
| Metric | A — no TypeSafe | B — TypeSafe |
|
||||
|---|---|---|
|
||||
| Targets hit | **10 / 13** | 9 / 13 |
|
||||
| Findings | 16 | **18** |
|
||||
| Wall-clock | 32m 12s | **26m 53s** |
|
||||
| Criticals | 5 | 2 (recalibrated) |
|
||||
| Belief-gate holds (POMDP) | 3 | — |
|
||||
| Assurance P1–P5 | all present | all present |
|
||||
| Model cost | $0 (subscription) | $0 + TypeSafe ≪ $5 |
|
||||
|
||||
Union coverage (both runs): **11 / 13**. Neither reached `web_sqli_second_order`
|
||||
or `web_crlf_header_go`.
|
||||
|
||||
## Reading it honestly
|
||||
|
||||
- **Recall is a tie** — 10 vs 9 is within run-to-run variance at `vote-n 1`.
|
||||
TypeSafe is a judgment layer, not a recall multiplier.
|
||||
- **B surfaced 2 real net-new findings** the plain run missed (`config.json`
|
||||
API-key exposure CWE-200, no-lockout brute force CWE-307) and caught
|
||||
`web_idor_invoice`.
|
||||
- **TypeSafe recalibrated severity** — 5 class-inflated Criticals → 2 evidence-
|
||||
backed ones. On this target it *under-rated* one genuine critical (the BOLA
|
||||
credential dump: A = Critical 9.1, B = Low). Calibration is a dial toward
|
||||
defensibility, not a correctness oracle.
|
||||
- **Harness gap found & fixed**: an earlier B collapsed to 0 findings when the
|
||||
subscription hit a session limit mid-run — NeuroSploit treated the limit
|
||||
message as a normal (exit-0) response and burned every agent. Now the
|
||||
session-limit sentinel parks the run (`fix(models)`).
|
||||
|
||||
## Confounders
|
||||
|
||||
Single samples, not averages. `vote-n 1` = no cross-model agreement in either
|
||||
arm. Recall scored by class + endpoint-keyword match (coverage, not graded
|
||||
proof). One target. Treat as one honest data point, not a leaderboard.
|
||||
|
||||
## Files
|
||||
|
||||
```
|
||||
report.html the visual write-up
|
||||
score.py the scorer (class + endpoint keyword match vs the 13 targets)
|
||||
scores.txt scorer output for both runs
|
||||
run_a_no_typesafe/ findings.json · assurance.json · meta.json · report.html · run.log
|
||||
run_b_typesafe/ findings.json · assurance.json · meta.json · report.html · run.log
|
||||
```
|
||||
|
||||
The TypeSafe API key and any subscription tokens are **not** in these files
|
||||
(env-only during the runs; verified clean before commit).
|
||||
@@ -0,0 +1,282 @@
|
||||
<title>NeuroSploit × TypeSafe Benchmark</title>
|
||||
<meta name="description" content="Head-to-head of NeuroSploit against a vulnerable target, with and without TypeSafe System One as a confirmation layer.">
|
||||
<link rel="stylesheet" href="https://fonts.googleapis.com/css2?family=IBM+Plex+Sans:wght@400;500;600&family=IBM+Plex+Mono:wght@400;500;600&family=Chivo:wght@600;700;800&display=swap">
|
||||
<style>
|
||||
:root{
|
||||
--ground:#f4f2f7; --surface:#ffffff; --surface-2:#eceaf3; --line:#ddd8e8;
|
||||
--ink:#1a1726; --muted:#6b6580; --faint:#938da6;
|
||||
--accent:#6d4bd8; /* neuro violet */
|
||||
--a:#c2701c; /* run A — amber (no typesafe) */
|
||||
--b:#0e8f86; /* run B — teal (typesafe) */
|
||||
--crit:#c8324a; --high:#d9743a; --med:#c2a01c; --low:#4a76c4; --info:#7b7590; --good:#1f9d68;
|
||||
--shadow:0 1px 2px rgba(26,23,38,.06),0 6px 20px rgba(26,23,38,.06);
|
||||
/* severity — vivid, identical in both themes (severity is not theme-relative) */
|
||||
--sev-crit:#e5484d; --sev-high:#f76b15; --sev-med:#f5b301; --sev-low:#3e7bfa; --sev-info:#8b8698;
|
||||
}
|
||||
:root:not([data-theme="light"]){ @media (prefers-color-scheme:dark){
|
||||
--ground:#0f0e17; --surface:#191627; --surface-2:#211d33; --line:#2e2942;
|
||||
--ink:#eceaf5; --muted:#a49dbd; --faint:#736c8f;
|
||||
--accent:#a78bfa; --a:#f0a95e; --b:#4fd6c6;
|
||||
--crit:#f26d7d; --high:#f0965e; --med:#e6cd52; --low:#7aa6f0; --info:#9a93b4; --good:#5ee0a0;
|
||||
--shadow:0 1px 2px rgba(0,0,0,.4),0 8px 30px rgba(0,0,0,.35);
|
||||
}}
|
||||
:root[data-theme="dark"]{
|
||||
--ground:#0f0e17; --surface:#191627; --surface-2:#211d33; --line:#2e2942;
|
||||
--ink:#eceaf5; --muted:#a49dbd; --faint:#736c8f;
|
||||
--accent:#a78bfa; --a:#f0a95e; --b:#4fd6c6;
|
||||
--crit:#f26d7d; --high:#f0965e; --med:#e6cd52; --low:#7aa6f0; --info:#9a93b4; --good:#5ee0a0;
|
||||
--shadow:0 1px 2px rgba(0,0,0,.4),0 8px 30px rgba(0,0,0,.35);
|
||||
}
|
||||
*{box-sizing:border-box}
|
||||
body{background:var(--ground);color:var(--ink);font-family:"IBM Plex Sans",system-ui,sans-serif;line-height:1.55;
|
||||
-webkit-font-smoothing:antialiased;margin:0}
|
||||
.wrap{max-width:1000px;margin:0 auto;padding:clamp(24px,5vw,64px) clamp(18px,4vw,40px)}
|
||||
h1,h2,h3{font-family:"Chivo","IBM Plex Sans",sans-serif;text-wrap:balance;line-height:1.1;margin:0}
|
||||
code,.mono,.num{font-family:"IBM Plex Mono",ui-monospace,monospace;font-variant-numeric:tabular-nums}
|
||||
.eyebrow{font-family:"IBM Plex Mono",monospace;font-size:12px;letter-spacing:.18em;text-transform:uppercase;color:var(--accent);font-weight:600}
|
||||
|
||||
/* header */
|
||||
header{border-bottom:1px solid var(--line);padding-bottom:28px;margin-bottom:36px}
|
||||
h1{font-size:clamp(30px,5.5vw,50px);font-weight:800;margin:10px 0 8px;letter-spacing:-.02em}
|
||||
.sub{color:var(--muted);font-size:16px;max-width:64ch}
|
||||
.meta{display:flex;flex-wrap:wrap;gap:8px 18px;margin-top:18px;font-family:"IBM Plex Mono",monospace;font-size:12.5px;color:var(--faint)}
|
||||
.meta b{color:var(--ink);font-weight:500}
|
||||
|
||||
/* thesis tiles */
|
||||
.thesis{display:grid;grid-template-columns:repeat(auto-fit,minmax(150px,1fr));gap:14px;margin:34px 0}
|
||||
.tile{background:var(--surface);border:1px solid var(--line);border-radius:12px;padding:18px 18px 16px;box-shadow:var(--shadow)}
|
||||
.tile .k{font-family:"IBM Plex Mono",monospace;font-size:11px;letter-spacing:.1em;text-transform:uppercase;color:var(--faint)}
|
||||
.tile .v{font-family:"Chivo",sans-serif;font-weight:800;font-size:30px;letter-spacing:-.02em;margin-top:6px;display:flex;align-items:baseline;gap:8px}
|
||||
.tile .u{font-size:13px;font-weight:500;color:var(--muted);font-family:"IBM Plex Sans"}
|
||||
.tile .note{font-size:12.5px;color:var(--muted);margin-top:4px}
|
||||
.swatchA{color:var(--a)} .swatchB{color:var(--b)}
|
||||
|
||||
section{margin:44px 0}
|
||||
h2{font-size:22px;font-weight:700;margin-bottom:4px}
|
||||
.lead{color:var(--muted);font-size:15px;margin:6px 0 20px;max-width:70ch}
|
||||
|
||||
/* comparison table */
|
||||
.cmp{width:100%;border-collapse:collapse;font-size:14.5px;background:var(--surface);border:1px solid var(--line);border-radius:12px;overflow:hidden;box-shadow:var(--shadow)}
|
||||
.cmp th,.cmp td{padding:12px 16px;text-align:left;border-bottom:1px solid var(--line)}
|
||||
.cmp thead th{font-family:"IBM Plex Mono",monospace;font-size:11.5px;letter-spacing:.08em;text-transform:uppercase;color:var(--faint);font-weight:600;background:var(--surface-2)}
|
||||
.cmp tbody tr:last-child td{border-bottom:none}
|
||||
.cmp td.metric{color:var(--muted)}
|
||||
.cmp td .num{font-weight:600;font-size:15px}
|
||||
.colA{color:var(--a)} .colB{color:var(--b)}
|
||||
.win{position:relative}
|
||||
.win::after{content:"▲";font-size:9px;margin-left:6px;vertical-align:middle;color:var(--good)}
|
||||
|
||||
/* per-scenario grid */
|
||||
.scen{display:grid;grid-template-columns:1fr auto auto;gap:0;background:var(--surface);border:1px solid var(--line);border-radius:12px;overflow:hidden;box-shadow:var(--shadow)}
|
||||
.scen .row{display:contents}
|
||||
.scen .cell{padding:10px 16px;border-bottom:1px solid var(--line);display:flex;align-items:center;gap:10px}
|
||||
.scen .row:last-child .cell{border-bottom:none}
|
||||
.scen .head .cell{font-family:"IBM Plex Mono",monospace;font-size:11px;letter-spacing:.08em;text-transform:uppercase;color:var(--faint);background:var(--surface-2);font-weight:600}
|
||||
.scen .idc{font-family:"IBM Plex Mono",monospace;font-size:13px}
|
||||
.scen .cls{font-size:11px;color:var(--faint);font-family:"IBM Plex Mono";margin-left:auto;padding-left:10px}
|
||||
.mk{width:60px;justify-content:center;font-family:"IBM Plex Mono";font-size:13px;font-weight:600}
|
||||
.hit{color:var(--good)} .miss{color:var(--crit);opacity:.7}
|
||||
.hdrA{color:var(--a)} .hdrB{color:var(--b)}
|
||||
|
||||
/* severity bars */
|
||||
.sev-wrap{display:grid;grid-template-columns:1fr 1fr;gap:18px}
|
||||
@media(max-width:640px){.sev-wrap{grid-template-columns:1fr}}
|
||||
.sevcard{background:var(--surface);border:1px solid var(--line);border-radius:12px;padding:18px;box-shadow:var(--shadow)}
|
||||
.sevcard h3{font-size:14px;font-family:"IBM Plex Mono";letter-spacing:.05em;margin-bottom:14px;display:flex;align-items:center;gap:8px}
|
||||
.dot{width:9px;height:9px;border-radius:50%;display:inline-block}
|
||||
.bar{display:flex;align-items:center;gap:12px;margin:9px 0;font-size:13px}
|
||||
.bar .lab{width:70px;color:var(--muted);font-family:"IBM Plex Mono";font-size:11.5px;display:flex;align-items:center;gap:7px}
|
||||
.bar .lab .sw{width:9px;height:9px;border-radius:2px;flex:none}
|
||||
.bar .track{flex:1;height:22px;background:var(--surface-2);border-radius:5px;overflow:hidden;border:1px solid var(--line)}
|
||||
.bar .fill{height:100%;border-radius:4px;min-width:6px;box-shadow:inset 0 0 0 1px rgba(255,255,255,.08)}
|
||||
.bar .n{width:22px;text-align:right;font-family:"IBM Plex Mono";font-weight:700;font-size:14px}
|
||||
.sev-legend{display:flex;flex-wrap:wrap;gap:14px;margin:0 0 18px;font-family:"IBM Plex Mono";font-size:11.5px;color:var(--muted)}
|
||||
.sev-legend span{display:inline-flex;align-items:center;gap:6px}
|
||||
.sev-legend i{width:11px;height:11px;border-radius:3px;display:inline-block}
|
||||
|
||||
/* callout */
|
||||
.callout{background:var(--surface);border:1px solid var(--line);border-left:3px solid var(--accent);border-radius:10px;padding:20px 22px;box-shadow:var(--shadow)}
|
||||
.callout h3{font-size:16px;margin-bottom:10px}
|
||||
.callout p{margin:8px 0;font-size:14.5px;color:var(--ink)}
|
||||
.callout .contrast{display:grid;grid-template-columns:1fr 1fr;gap:14px;margin-top:14px}
|
||||
@media(max-width:560px){.callout .contrast{grid-template-columns:1fr}}
|
||||
.cbox{background:var(--surface-2);border-radius:8px;padding:12px 14px}
|
||||
.cbox .t{font-family:"IBM Plex Mono";font-size:11px;text-transform:uppercase;letter-spacing:.08em;margin-bottom:6px}
|
||||
.cbox .r{font-size:13px;color:var(--muted)}
|
||||
.cbox .g{font-size:20px;font-family:"Chivo";font-weight:800;margin-top:4px}
|
||||
|
||||
ul.take{list-style:none;padding:0;margin:0;display:flex;flex-direction:column;gap:12px}
|
||||
ul.take li{background:var(--surface);border:1px solid var(--line);border-radius:10px;padding:14px 16px;font-size:14.5px;display:flex;gap:12px;box-shadow:var(--shadow)}
|
||||
ul.take .tag{font-family:"IBM Plex Mono";font-size:10.5px;font-weight:600;letter-spacing:.06em;padding:3px 8px;border-radius:5px;height:fit-content;white-space:nowrap;text-transform:uppercase}
|
||||
.tag.even{background:color-mix(in srgb,var(--info) 22%,transparent);color:var(--info)}
|
||||
.tag.plus{background:color-mix(in srgb,var(--good) 20%,transparent);color:var(--good)}
|
||||
.tag.minus{background:color-mix(in srgb,var(--crit) 18%,transparent);color:var(--crit)}
|
||||
.tag.note{background:color-mix(in srgb,var(--accent) 18%,transparent);color:var(--accent)}
|
||||
|
||||
.disclaim{margin-top:44px;padding-top:22px;border-top:1px solid var(--line);color:var(--faint);font-size:12.5px;line-height:1.6}
|
||||
.disclaim b{color:var(--muted)}
|
||||
a{color:var(--accent)}
|
||||
</style>
|
||||
|
||||
<div class="wrap">
|
||||
<header>
|
||||
<div class="eyebrow">NeuroSploit · assurance benchmark · 2026-09-20</div>
|
||||
<h1>Does TypeSafe make the run better?</h1>
|
||||
<p class="sub">Two identical NeuroSploit engagements against the same vulnerable target — one plain,
|
||||
one with TypeSafe System One (Jev) as a calibrated confirmation layer. Same model, same focus,
|
||||
same 13 seeded vulnerabilities. Only the <code>--typesafe</code> flag differs.</p>
|
||||
<div class="meta">
|
||||
<span>target <b>NimbusCart (BenchMarkBurpAT)</b> · localhost:3000</span>
|
||||
<span>model <b>claude-opus-4-8</b> (subscription)</span>
|
||||
<span>recon <b>2</b> · vote-n <b>1</b> · max-agents <b>15</b></span>
|
||||
<span>ground truth <b>13 targets</b></span>
|
||||
</div>
|
||||
</header>
|
||||
|
||||
<div class="thesis">
|
||||
<div class="tile">
|
||||
<div class="k">Recall — no TypeSafe</div>
|
||||
<div class="v swatchA">10<span class="u">/13</span></div>
|
||||
<div class="note">16 findings · 32m12s</div>
|
||||
</div>
|
||||
<div class="tile">
|
||||
<div class="k">Recall — with TypeSafe</div>
|
||||
<div class="v swatchB">9<span class="u">/13</span></div>
|
||||
<div class="note">18 findings · 26m53s</div>
|
||||
</div>
|
||||
<div class="tile">
|
||||
<div class="k">Union coverage</div>
|
||||
<div class="v">11<span class="u">/13</span></div>
|
||||
<div class="note">the two runs together</div>
|
||||
</div>
|
||||
<div class="tile">
|
||||
<div class="k">TypeSafe recalibrated</div>
|
||||
<div class="v swatchB">9</div>
|
||||
<div class="note">findings, calibrated confidence</div>
|
||||
</div>
|
||||
</div>
|
||||
|
||||
<section>
|
||||
<h2>Head to head</h2>
|
||||
<p class="lead">The recall is a tie inside the noise; the real difference is <em>shape</em>. TypeSafe was
|
||||
faster, surfaced two real findings the plain run missed, and pulled inflated severities down toward what the
|
||||
evidence actually demonstrated — at the cost of being conservative enough to drop two scenarios and under-rate
|
||||
one genuine critical.</p>
|
||||
<div style="overflow-x:auto">
|
||||
<table class="cmp">
|
||||
<thead><tr><th>Metric</th><th class="colA">A — no TypeSafe</th><th class="colB">B — TypeSafe</th></tr></thead>
|
||||
<tbody>
|
||||
<tr><td class="metric">Seeded targets hit</td><td class="colA win"><span class="num">10 / 13</span></td><td class="colB"><span class="num">9 / 13</span></td></tr>
|
||||
<tr><td class="metric">Total findings reported</td><td class="colA"><span class="num">16</span></td><td class="colB win"><span class="num">18</span></td></tr>
|
||||
<tr><td class="metric">Findings beyond the 13 targets</td><td class="colA"><span class="num">6</span></td><td class="colB win"><span class="num">9</span> <span style="color:var(--muted);font-size:12px">(2 real: config leak, no-lockout)</span></td></tr>
|
||||
<tr><td class="metric">Wall-clock time</td><td class="colA"><span class="num">32m 12s</span></td><td class="colB win"><span class="num">26m 53s</span></td></tr>
|
||||
<tr><td class="metric">Criticals reported</td><td class="colA"><span class="num">5</span></td><td class="colB"><span class="num">2</span> <span style="color:var(--muted);font-size:12px">(recalibrated)</span></td></tr>
|
||||
<tr><td class="metric">Belief-gate holds (POMDP)</td><td class="colA"><span class="num">3</span></td><td class="colB"><span class="num">—</span></td></tr>
|
||||
<tr><td class="metric">Assurance P1–P5</td><td class="colA"><span class="num">all present</span></td><td class="colB"><span class="num">all present</span></td></tr>
|
||||
<tr><td class="metric">Model API cost</td><td class="colA"><span class="num">$0</span> subscription</td><td class="colB"><span class="num">$0</span> + TypeSafe ≪ $5</td></tr>
|
||||
</tbody>
|
||||
</table>
|
||||
</div>
|
||||
</section>
|
||||
|
||||
<section>
|
||||
<h2>Per-scenario coverage</h2>
|
||||
<p class="lead">Each seeded vulnerability, and whether each run confirmed it. Neither run reached the
|
||||
second-order SQLi or the CRLF header injection — the two that need a multi-step chain the single-vote
|
||||
config didn't pursue.</p>
|
||||
<div class="scen">
|
||||
<div class="row head">
|
||||
<div class="cell">Scenario</div>
|
||||
<div class="cell mk hdrA">A</div>
|
||||
<div class="cell mk hdrB">B·TS</div>
|
||||
</div>
|
||||
<!-- rows -->
|
||||
<div class="row"><div class="cell"><span class="idc">web_sqli_login_bypass</span><span class="cls">SQLi</span></div><div class="cell mk hit">✓</div><div class="cell mk hit">✓</div></div>
|
||||
<div class="row"><div class="cell"><span class="idc">web_sqli_union_search</span><span class="cls">SQLi</span></div><div class="cell mk hit">✓</div><div class="cell mk miss">✕</div></div>
|
||||
<div class="row"><div class="cell"><span class="idc">web_sqli_blind_boolean</span><span class="cls">SQLi</span></div><div class="cell mk hit">✓</div><div class="cell mk hit">✓</div></div>
|
||||
<div class="row"><div class="cell"><span class="idc">web_sqli_blind_time</span><span class="cls">SQLi</span></div><div class="cell mk hit">✓</div><div class="cell mk miss">✕</div></div>
|
||||
<div class="row"><div class="cell"><span class="idc">web_sqli_second_order</span><span class="cls">SQLi</span></div><div class="cell mk miss">✕</div><div class="cell mk miss">✕</div></div>
|
||||
<div class="row"><div class="cell"><span class="idc">web_xss_reflected_search</span><span class="cls">XSS</span></div><div class="cell mk hit">✓</div><div class="cell mk hit">✓</div></div>
|
||||
<div class="row"><div class="cell"><span class="idc">web_xss_stored_review</span><span class="cls">XSS</span></div><div class="cell mk hit">✓</div><div class="cell mk hit">✓</div></div>
|
||||
<div class="row"><div class="cell"><span class="idc">web_xss_svg_upload</span><span class="cls">XSS</span></div><div class="cell mk hit">✓</div><div class="cell mk hit">✓</div></div>
|
||||
<div class="row"><div class="cell"><span class="idc">web_xss_dom_redirect</span><span class="cls">XSS</span></div><div class="cell mk hit">✓</div><div class="cell mk hit">✓</div></div>
|
||||
<div class="row"><div class="cell"><span class="idc">web_idor_invoice</span><span class="cls">IDOR</span></div><div class="cell mk miss">✕</div><div class="cell mk hit">✓</div></div>
|
||||
<div class="row"><div class="cell"><span class="idc">api_bola_orders</span><span class="cls">BOLA</span></div><div class="cell mk hit">✓</div><div class="cell mk hit">✓</div></div>
|
||||
<div class="row"><div class="cell"><span class="idc">web_open_redirect_login</span><span class="cls">Redirect</span></div><div class="cell mk hit">✓</div><div class="cell mk hit">✓</div></div>
|
||||
<div class="row"><div class="cell"><span class="idc">web_crlf_header_go</span><span class="cls">CRLF</span></div><div class="cell mk miss">✕</div><div class="cell mk miss">✕</div></div>
|
||||
</div>
|
||||
</section>
|
||||
|
||||
<section>
|
||||
<h2>Severity shape</h2>
|
||||
<p class="lead">The clearest effect of TypeSafe: the severity distribution flattens. The plain run stacks
|
||||
five Criticals; the calibrated run keeps two and pushes the rest down to where the demonstrated-impact
|
||||
evidence puts them.</p>
|
||||
<div class="sev-legend"><span><i style="background:#e5484d"></i>Critical</span><span><i style="background:#f76b15"></i>High</span><span><i style="background:#f5b301"></i>Medium</span><span><i style="background:#3e7bfa"></i>Low</span><span><i style="background:#8b8698"></i>Info</span></div>
|
||||
<div class="sev-wrap">
|
||||
<div class="sevcard">
|
||||
<h3 class="swatchA"><span class="dot" style="background:var(--a)"></span> A — no TypeSafe · 16</h3>
|
||||
<div class="bar"><span class="lab"><i class="sw" style="background:#e5484d"></i>Critical</span><span class="track"><span class="fill" style="width:100%;background:#e5484d"></span></span><span class="n" style="color:#e5484d">5</span></div>
|
||||
<div class="bar"><span class="lab"><i class="sw" style="background:#f76b15"></i>High</span><span class="track"><span class="fill" style="width:80%;background:#f76b15"></span></span><span class="n" style="color:#f76b15">4</span></div>
|
||||
<div class="bar"><span class="lab"><i class="sw" style="background:#f5b301"></i>Medium</span><span class="track"><span class="fill" style="width:20%;background:#f5b301"></span></span><span class="n" style="color:#f5b301">1</span></div>
|
||||
<div class="bar"><span class="lab"><i class="sw" style="background:#3e7bfa"></i>Low</span><span class="track"><span class="fill" style="width:40%;background:#3e7bfa"></span></span><span class="n" style="color:#3e7bfa">2</span></div>
|
||||
<div class="bar"><span class="lab"><i class="sw" style="background:#8b8698"></i>Info</span><span class="track"><span class="fill" style="width:80%;background:#8b8698"></span></span><span class="n" style="color:#8b8698">4</span></div>
|
||||
</div>
|
||||
<div class="sevcard">
|
||||
<h3 class="swatchB"><span class="dot" style="background:var(--b)"></span> B — TypeSafe · 18</h3>
|
||||
<div class="bar"><span class="lab"><i class="sw" style="background:#e5484d"></i>Critical</span><span class="track"><span class="fill" style="width:40%;background:#e5484d"></span></span><span class="n" style="color:#e5484d">2</span></div>
|
||||
<div class="bar"><span class="lab"><i class="sw" style="background:#f76b15"></i>High</span><span class="track"><span class="fill" style="width:80%;background:#f76b15"></span></span><span class="n" style="color:#f76b15">4</span></div>
|
||||
<div class="bar"><span class="lab"><i class="sw" style="background:#f5b301"></i>Medium</span><span class="track"><span class="fill" style="width:60%;background:#f5b301"></span></span><span class="n" style="color:#f5b301">3</span></div>
|
||||
<div class="bar"><span class="lab"><i class="sw" style="background:#3e7bfa"></i>Low</span><span class="track"><span class="fill" style="width:80%;background:#3e7bfa"></span></span><span class="n" style="color:#3e7bfa">4</span></div>
|
||||
<div class="bar"><span class="lab"><i class="sw" style="background:#8b8698"></i>Info</span><span class="track"><span class="fill" style="width:100%;background:#8b8698"></span></span><span class="n" style="color:#8b8698">5</span></div>
|
||||
</div>
|
||||
</div>
|
||||
</section>
|
||||
|
||||
<section>
|
||||
<h2>What calibration actually did</h2>
|
||||
<div class="callout">
|
||||
<h3>The same BOLA, two severities</h3>
|
||||
<p>Both runs found the object-level auth flaw on <code>GET /api/v2/users/:id</code> — a customer token
|
||||
reads any user's full record, including the admin's plaintext password. The plain run rated it
|
||||
<b>Critical (9.1)</b> on the class. TypeSafe, grading against the demonstrated-impact receipts and its
|
||||
calibrated judgment, rated it <b>Low</b>.</p>
|
||||
<div class="contrast">
|
||||
<div class="cbox"><div class="t swatchA">A — class-graded</div><div class="g swatchA">Critical 9.1</div><div class="r">BOLA + excessive data exposure</div></div>
|
||||
<div class="cbox"><div class="t swatchB">B — evidence-graded</div><div class="g swatchB">Low</div><div class="r">same finding, impact receipts weighted</div></div>
|
||||
</div>
|
||||
<p style="margin-top:14px"><b>Why it fired:</b> the severity is graded from the <em>structured</em> evidence
|
||||
slot (the recorded request/response exchange), not the agent's prose. This finding proved the dump in its
|
||||
narrative and claims ledger but left <code>evidence_data</code> null — so the demonstrated-impact rung saw no
|
||||
machine-readable C/I/A receipt, and the calibrated grader dropped the impact metrics to <code>None</code>,
|
||||
collapsing 9.1 → Low. The proof existed; it just wasn't in the slot the grader reads.</p>
|
||||
<p>This is the honest edge: calibration removes inflated Criticals (good — most scanners over-rate by class),
|
||||
but a receipt in the wrong slot gets under-rated. It is a dial toward defensibility, not a correctness
|
||||
oracle — the operator still owns the final severity, and the fix is to make agents populate
|
||||
<code>evidence_data</code> for impact, not to loosen the grader.</p>
|
||||
</div>
|
||||
</section>
|
||||
|
||||
<section>
|
||||
<h2>Takeaways</h2>
|
||||
<ul class="take">
|
||||
<li><span class="tag even">tie</span><div><b>Recall is a wash.</b> 10 vs 9 of 13 is within run-to-run variance at vote-n 1. TypeSafe is not a recall multiplier — it is a judgment layer.</div></li>
|
||||
<li><span class="tag plus">gain</span><div><b>Two real net-new findings.</b> The TypeSafe run surfaced a <code>config.json</code> API-key exposure (CWE-200) and a no-lockout brute-force (CWE-307) the plain run never reported — and it caught <code>web_idor_invoice</code>, which the plain run missed.</div></li>
|
||||
<li><span class="tag plus">gain</span><div><b>Faster and calibrated.</b> 5m19s quicker, and it recalibrated 9 findings' confidence — collapsing five class-inflated Criticals to two evidence-backed ones.</div></li>
|
||||
<li><span class="tag minus">cost</span><div><b>Conservatism has a price.</b> It dropped <code>union_search</code> and <code>blind_time</code>, and under-rated the credential-dump BOLA. A confirmation layer that demands receipts will sometimes discard a real thing it couldn't re-prove in-budget.</div></li>
|
||||
<li><span class="tag note">cheap</span><div><b>Negligible cost.</b> TypeSafe adds no LLM tokens of its own — one probe call billed 319 in / 21 out. The whole run stayed far under the $5 budget.</div></li>
|
||||
</ul>
|
||||
</section>
|
||||
|
||||
<div class="disclaim">
|
||||
<b>Method & honesty.</b> Both runs: NeuroSploit v4.0.0, <code>claude-opus-4-8</code> via subscription,
|
||||
black-box, recon intensity 2, single-model vote (<code>vote-n 1</code>), same natural-language focus naming
|
||||
the 13 endpoints, no pre-baked solver — the LLM discovered and confirmed everything live. Recall is scored by
|
||||
class + endpoint keyword match against the target's ground-truth list, so a match is coverage, not a graded
|
||||
proof. <b>Confounders:</b> the two runs are single samples, not averages; an earlier TypeSafe run collapsed to
|
||||
zero when the subscription hit a session limit mid-run (a real harness gap, since fixed — session-limit stdout
|
||||
now parks the run instead of burning agents); vote-n 1 means no cross-model agreement in either arm. Treat this
|
||||
as one honest data point on one target, not a leaderboard. <b>Not measured here:</b> multi-sample variance,
|
||||
higher vote-n, and TypeSafe's agent-pruning effect on a broader agent set.
|
||||
</div>
|
||||
</div>
|
||||
@@ -0,0 +1,122 @@
|
||||
{
|
||||
"engine": "neurosploit",
|
||||
"version": "4.0.0",
|
||||
"build": "49d3d3ceb1df",
|
||||
"run": "ns-1789853137-localhost_3000",
|
||||
"target": "http://localhost:3000",
|
||||
"generated": 1789855069,
|
||||
"findings": 16,
|
||||
"artifacts": [
|
||||
{
|
||||
"name": "findings.json",
|
||||
"present": true,
|
||||
"sha256": "61ed87d0036ae5b2dd61ffd2c9ead34e11076f1cfbb8570f7554c072f3a76cb9",
|
||||
"bytes": 91029,
|
||||
"role": "the findings, each stamped with the engine build (P5)"
|
||||
},
|
||||
{
|
||||
"name": "report.html",
|
||||
"present": false,
|
||||
"bytes": 0,
|
||||
"role": "the human report"
|
||||
},
|
||||
{
|
||||
"name": "recon.json",
|
||||
"present": true,
|
||||
"sha256": "8f5110c6d65cac10c4c04a8deacaf4cacbc8c8d320d18ed6cee236a7e61fe104",
|
||||
"bytes": 15141,
|
||||
"role": "reconnaissance facts"
|
||||
},
|
||||
{
|
||||
"name": "audit.jsonl",
|
||||
"present": true,
|
||||
"sha256": "c6f63d9c2e70f59b05120a732ce157e23606ff388f232d31299545521818135b",
|
||||
"bytes": 19254,
|
||||
"role": "hash-chained decision log — every ALLOW/DENY (P1/P2/P4)"
|
||||
},
|
||||
{
|
||||
"name": "audit.jsonl.anchors",
|
||||
"present": true,
|
||||
"sha256": "3b014612736ca9110c6342e61892286620ba2db605f82d278bec4de732e4bd1f",
|
||||
"bytes": 213,
|
||||
"role": "external anchors of the audit chain (P4)"
|
||||
},
|
||||
{
|
||||
"name": "provenance.json",
|
||||
"present": true,
|
||||
"sha256": "c33f47d22ff52d82db3077f0eecceb105f73fb0f8cfe6f057c77a4696ddd1e4a",
|
||||
"bytes": 297,
|
||||
"role": "signed provenance manifest — build + structural signature (P5)"
|
||||
},
|
||||
{
|
||||
"name": "out-of-scope-findings.json",
|
||||
"present": true,
|
||||
"sha256": "425df6ff6ac515da2b36f1bf4582d9acd8599e8c4e586e8dadc99b5601a06752",
|
||||
"bytes": 17615,
|
||||
"role": "findings quarantined for being outside scope (P2)"
|
||||
},
|
||||
{
|
||||
"name": "flows.jsonl",
|
||||
"present": false,
|
||||
"bytes": 0,
|
||||
"role": "intercepted request/response flows"
|
||||
},
|
||||
{
|
||||
"name": "meta.json",
|
||||
"present": true,
|
||||
"sha256": "1e47c73f41061aef5e1943d3c8321f41349cf8e3588cfb1286a5627a226773cc",
|
||||
"bytes": 198,
|
||||
"role": "target metadata"
|
||||
}
|
||||
],
|
||||
"properties": [
|
||||
{
|
||||
"id": "P1",
|
||||
"name": "Signed authorization",
|
||||
"status": "present",
|
||||
"evidenced_by": [
|
||||
"audit.jsonl"
|
||||
],
|
||||
"note": "capability recorded and decisions logged"
|
||||
},
|
||||
{
|
||||
"id": "P2",
|
||||
"name": "Scope enforcement",
|
||||
"status": "present",
|
||||
"evidenced_by": [
|
||||
"audit.jsonl",
|
||||
"out-of-scope-findings.json"
|
||||
],
|
||||
"note": "scope decisions recorded, including denials/quarantine"
|
||||
},
|
||||
{
|
||||
"id": "P3",
|
||||
"name": "Evidence & CVSS",
|
||||
"status": "present",
|
||||
"evidenced_by": [
|
||||
"findings.json"
|
||||
],
|
||||
"note": "0/16 findings carry structured evidence · 14 with CVSS · 15 voted · 20 PoC(s) · 0 screenshot(s) · 13 evidence file(s)"
|
||||
},
|
||||
{
|
||||
"id": "P4",
|
||||
"name": "Audit integrity",
|
||||
"status": "present",
|
||||
"evidenced_by": [
|
||||
"audit.jsonl",
|
||||
"audit.jsonl.anchors"
|
||||
],
|
||||
"note": "hash chain plus signed anchors (truncation/rebuild detectable)"
|
||||
},
|
||||
{
|
||||
"id": "P5",
|
||||
"name": "Provenance",
|
||||
"status": "present",
|
||||
"evidenced_by": [
|
||||
"provenance.json"
|
||||
],
|
||||
"note": "signed provenance manifest with structural signature"
|
||||
}
|
||||
],
|
||||
"bundle_hash": "579449f887726db317b0169641be27dc5a11db46da698b87255d91dfaae9697f"
|
||||
}
|
||||
File diff suppressed because it is too large
Load Diff
@@ -0,0 +1,10 @@
|
||||
{
|
||||
"asset": "NimbusCart Inc",
|
||||
"brand": "NimbusCart Inc",
|
||||
"server": "",
|
||||
"status": 200,
|
||||
"target": "http://localhost:3000",
|
||||
"tech": [],
|
||||
"title": "Home · NimbusCart",
|
||||
"typesafe": false
|
||||
}
|
||||
File diff suppressed because one or more lines are too long
@@ -0,0 +1,120 @@
|
||||
{
|
||||
"engine": "neurosploit",
|
||||
"version": "4.0.0",
|
||||
"build": "49d3d3ceb1df",
|
||||
"run": "ns-1789870577-localhost_3000",
|
||||
"target": "http://localhost:3000",
|
||||
"generated": 1789872190,
|
||||
"findings": 18,
|
||||
"artifacts": [
|
||||
{
|
||||
"name": "findings.json",
|
||||
"present": true,
|
||||
"sha256": "9827d2c67a851679ce8462fc1885d2c4ddfeb0a70180db293029f33d362d108f",
|
||||
"bytes": 112054,
|
||||
"role": "the findings, each stamped with the engine build (P5)"
|
||||
},
|
||||
{
|
||||
"name": "report.html",
|
||||
"present": false,
|
||||
"bytes": 0,
|
||||
"role": "the human report"
|
||||
},
|
||||
{
|
||||
"name": "recon.json",
|
||||
"present": true,
|
||||
"sha256": "18a9d9456b5d239905a8c5a2d0647b272f8b9e5b5ff7f20c6ed26e7bf5164258",
|
||||
"bytes": 11611,
|
||||
"role": "reconnaissance facts"
|
||||
},
|
||||
{
|
||||
"name": "audit.jsonl",
|
||||
"present": true,
|
||||
"sha256": "11e92303952192781686111381c17e37326ecedc1969b2c0d17b8d810fffebdd",
|
||||
"bytes": 32535,
|
||||
"role": "hash-chained decision log — every ALLOW/DENY (P1/P2/P4)"
|
||||
},
|
||||
{
|
||||
"name": "audit.jsonl.anchors",
|
||||
"present": true,
|
||||
"sha256": "a0dd67ad35b530842fc0221ead9536b3ce19d45be251e8568b80909f4859d7cb",
|
||||
"bytes": 213,
|
||||
"role": "external anchors of the audit chain (P4)"
|
||||
},
|
||||
{
|
||||
"name": "provenance.json",
|
||||
"present": true,
|
||||
"sha256": "ac45f0856813ca943713ff782a784e34ccc08038794fd5dfa00d6f060eac803c",
|
||||
"bytes": 297,
|
||||
"role": "signed provenance manifest — build + structural signature (P5)"
|
||||
},
|
||||
{
|
||||
"name": "out-of-scope-findings.json",
|
||||
"present": false,
|
||||
"bytes": 0,
|
||||
"role": "findings quarantined for being outside scope (P2)"
|
||||
},
|
||||
{
|
||||
"name": "flows.jsonl",
|
||||
"present": false,
|
||||
"bytes": 0,
|
||||
"role": "intercepted request/response flows"
|
||||
},
|
||||
{
|
||||
"name": "meta.json",
|
||||
"present": true,
|
||||
"sha256": "c879fc77b942399b73b8050f258d00b4e671b38bdaa00ef7eebfac356484864c",
|
||||
"bytes": 197,
|
||||
"role": "target metadata"
|
||||
}
|
||||
],
|
||||
"properties": [
|
||||
{
|
||||
"id": "P1",
|
||||
"name": "Signed authorization",
|
||||
"status": "present",
|
||||
"evidenced_by": [
|
||||
"audit.jsonl"
|
||||
],
|
||||
"note": "capability recorded and decisions logged"
|
||||
},
|
||||
{
|
||||
"id": "P2",
|
||||
"name": "Scope enforcement",
|
||||
"status": "present",
|
||||
"evidenced_by": [
|
||||
"audit.jsonl"
|
||||
],
|
||||
"note": "scope decisions recorded, including denials/quarantine"
|
||||
},
|
||||
{
|
||||
"id": "P3",
|
||||
"name": "Evidence & CVSS",
|
||||
"status": "present",
|
||||
"evidenced_by": [
|
||||
"findings.json"
|
||||
],
|
||||
"note": "0/18 findings carry structured evidence · 16 with CVSS · 17 voted · 16 PoC(s) · 0 screenshot(s) · 14 evidence file(s)"
|
||||
},
|
||||
{
|
||||
"id": "P4",
|
||||
"name": "Audit integrity",
|
||||
"status": "present",
|
||||
"evidenced_by": [
|
||||
"audit.jsonl",
|
||||
"audit.jsonl.anchors"
|
||||
],
|
||||
"note": "hash chain plus signed anchors (truncation/rebuild detectable)"
|
||||
},
|
||||
{
|
||||
"id": "P5",
|
||||
"name": "Provenance",
|
||||
"status": "present",
|
||||
"evidenced_by": [
|
||||
"provenance.json"
|
||||
],
|
||||
"note": "signed provenance manifest with structural signature"
|
||||
}
|
||||
],
|
||||
"bundle_hash": "85b0ef6f4789f08cb6bde0669b45aefb9f9f6bcb5f853a811c20f725a58e1eb1"
|
||||
}
|
||||
File diff suppressed because it is too large
Load Diff
@@ -0,0 +1,10 @@
|
||||
{
|
||||
"asset": "NimbusCart Inc",
|
||||
"brand": "NimbusCart Inc",
|
||||
"server": "",
|
||||
"status": 200,
|
||||
"target": "http://localhost:3000",
|
||||
"tech": [],
|
||||
"title": "Home · NimbusCart",
|
||||
"typesafe": true
|
||||
}
|
||||
File diff suppressed because one or more lines are too long
@@ -0,0 +1,85 @@
|
||||
#!/usr/bin/env python3
|
||||
# Score a NeuroSploit run against the user's 13 target scenarios (ground truth
|
||||
# from BenchMarkBurpAT/SOLUTIONS.md). Matching is by class + endpoint keyword.
|
||||
# No solver is involved — this only reads what the run reported.
|
||||
import json, sys, re
|
||||
|
||||
# The 13 targets the user asked to attack: id -> (cwe-ish class, endpoint keywords)
|
||||
TARGETS = {
|
||||
"web_idor_invoice": ("idor", ["/account/invoice", "invoice"]),
|
||||
"api_bola_orders": ("idor", ["/orders", "/v2/orders", "bola"]),
|
||||
"web_sqli_login_bypass": ("sqli", ["/login"]),
|
||||
"web_sqli_union_search": ("sqli", ["/shop/search", "search"]),
|
||||
"web_sqli_blind_boolean": ("sqli", ["/shop/filter", "maxprice", "filter"]),
|
||||
"web_sqli_blind_time": ("sqli", ["/support/feedback", "feedback", "comment"]),
|
||||
"web_sqli_second_order": ("sqli", ["/account/profile", "/admin/search-users", "bio", "second"]),
|
||||
"web_xss_reflected_search":("xss", ["/shop/search", "search"]),
|
||||
"web_xss_stored_review": ("xss", ["/review", "/shop/product"]),
|
||||
"web_xss_svg_upload": ("xss", ["/support/ticket", "/uploads", "svg"]),
|
||||
"web_xss_dom_redirect": ("xss", ["/go", "dom", "?url", "name="]),
|
||||
"web_open_redirect_login": ("redirect", ["/login", "next", "/go", "url="]),
|
||||
"web_crlf_header_go": ("crlf", ["/go", "crlf", "header inject"]),
|
||||
}
|
||||
|
||||
CLASS_CWE = {
|
||||
"sqli": {"89","943","564"},
|
||||
"xss": {"79","80","83","87"},
|
||||
"idor": {"639","862","863","284","285","566","425","200"},
|
||||
"redirect": {"601"},
|
||||
"crlf": {"113","93"},
|
||||
}
|
||||
|
||||
def classify(f):
|
||||
cwe = "".join(ch for ch in f.get("cwe","") if ch.isdigit())
|
||||
t = (f.get("title","")+" "+f.get("cwe","")).lower()
|
||||
for cls, cwes in CLASS_CWE.items():
|
||||
if cwe in cwes: return cls
|
||||
for cls, kw in {"sqli":["sql inj","sqli"],"xss":["xss","cross-site scripting"],
|
||||
"idor":["idor","bola","broken access","broken object"],
|
||||
"redirect":["open redirect"],"crlf":["crlf","response splitting","header inject"]}.items():
|
||||
if any(k in t for k in kw): return cls
|
||||
return "other"
|
||||
|
||||
def endpoint_blob(f):
|
||||
return " ".join(str(f.get(k,"")) for k in ("endpoint","title","payload","evidence")).lower()
|
||||
|
||||
def score(findings_path):
|
||||
findings = json.load(open(findings_path))
|
||||
hits = {} # target_id -> matched finding index
|
||||
used = set()
|
||||
for tid,(cls,kws) in TARGETS.items():
|
||||
for i,f in enumerate(findings):
|
||||
if i in used: continue
|
||||
if classify(f)!=cls: continue
|
||||
blob = endpoint_blob(f)
|
||||
if any(kw.lower() in blob for kw in kws):
|
||||
hits[tid]=i; used.add(i); break
|
||||
tp = len(hits)
|
||||
fn = [t for t in TARGETS if t not in hits]
|
||||
# extra findings not matched to a target = out-of-scope-but-real OR noise;
|
||||
# count as "extra" (not penalised as FP unless clearly bogus).
|
||||
extra = [i for i in range(len(findings)) if i not in used]
|
||||
return {
|
||||
"total_findings": len(findings),
|
||||
"targets_hit": tp,
|
||||
"targets_total": len(TARGETS),
|
||||
"recall": round(tp/len(TARGETS),3),
|
||||
"hit_ids": sorted(hits.keys()),
|
||||
"missed_ids": sorted(fn),
|
||||
"extra_findings": len(extra),
|
||||
}
|
||||
|
||||
if __name__=="__main__":
|
||||
import os
|
||||
for path in sys.argv[1:]:
|
||||
fp = path if path.endswith(".json") else os.path.join(path,"findings.json")
|
||||
try:
|
||||
r = score(fp)
|
||||
except Exception as e:
|
||||
print(f"{path}: ERROR {e}"); continue
|
||||
print(f"\n== {path} ==")
|
||||
print(f" findings reported : {r['total_findings']}")
|
||||
print(f" targets hit : {r['targets_hit']}/{r['targets_total']} (recall {r['recall']})")
|
||||
print(f" hit : {', '.join(r['hit_ids']) or '—'}")
|
||||
print(f" missed : {', '.join(r['missed_ids']) or '—'}")
|
||||
print(f" extra findings : {r['extra_findings']}")
|
||||
@@ -0,0 +1,14 @@
|
||||
|
||||
== /opt/neurosploit-rs/runs/ns-1789853137-localhost_3000 ==
|
||||
findings reported : 16
|
||||
targets hit : 10/13 (recall 0.769)
|
||||
hit : api_bola_orders, web_open_redirect_login, web_sqli_blind_boolean, web_sqli_blind_time, web_sqli_login_bypass, web_sqli_union_search, web_xss_dom_redirect, web_xss_reflected_search, web_xss_stored_review, web_xss_svg_upload
|
||||
missed : web_crlf_header_go, web_idor_invoice, web_sqli_second_order
|
||||
extra findings : 6
|
||||
|
||||
== /opt/neurosploit-rs/runs/ns-1789855082-localhost_3000 ==
|
||||
findings reported : 0
|
||||
targets hit : 0/13 (recall 0.0)
|
||||
hit : —
|
||||
missed : api_bola_orders, web_crlf_header_go, web_idor_invoice, web_open_redirect_login, web_sqli_blind_boolean, web_sqli_blind_time, web_sqli_login_bypass, web_sqli_second_order, web_sqli_union_search, web_xss_dom_redirect, web_xss_reflected_search, web_xss_stored_review, web_xss_svg_upload
|
||||
extra findings : 0
|
||||
Generated
+2
-2
@@ -929,7 +929,7 @@ dependencies = [
|
||||
|
||||
[[package]]
|
||||
name = "neurosploit"
|
||||
version = "4.0.0"
|
||||
version = "4.1.0"
|
||||
dependencies = [
|
||||
"anyhow",
|
||||
"clap",
|
||||
@@ -946,7 +946,7 @@ dependencies = [
|
||||
|
||||
[[package]]
|
||||
name = "neurosploit-harness"
|
||||
version = "4.0.0"
|
||||
version = "4.1.0"
|
||||
dependencies = [
|
||||
"anyhow",
|
||||
"base64",
|
||||
|
||||
@@ -3,7 +3,7 @@ members = ["crates/harness", "app"]
|
||||
resolver = "2"
|
||||
|
||||
[workspace.package]
|
||||
version = "4.0.0"
|
||||
version = "4.1.0"
|
||||
edition = "2021"
|
||||
license = "MIT"
|
||||
repository = "https://github.com/JoasASantos/NeuroSploit"
|
||||
|
||||
@@ -1,4 +1,4 @@
|
||||
//! NeuroSploit v4.0.0 — interactive harness + CLI (`run` / `whitebox` / `agents` / `models`).
|
||||
//! NeuroSploit v4.1.0 — interactive harness + CLI (`run` / `whitebox` / `agents` / `models`).
|
||||
|
||||
mod rectify;
|
||||
mod repl;
|
||||
@@ -12,8 +12,8 @@ use std::path::{Path, PathBuf};
|
||||
#[command(
|
||||
name = "neurosploit",
|
||||
version,
|
||||
about = "NeuroSploit v4.0.0 — multi-model autonomous pentest harness",
|
||||
long_about = "NeuroSploit v4.0.0 — a Rust multi-model harness that drives a pool of LLMs \
|
||||
about = "NeuroSploit v4.1.0 — multi-model autonomous pentest harness",
|
||||
long_about = "NeuroSploit v4.1.0 — a Rust multi-model harness that drives a pool of LLMs \
|
||||
(API key or local subscription: Claude/Codex/Gemini/Grok/OpenCode/Hermes) to autonomously test a target. \
|
||||
After recon it INTELLIGENTLY selects only the agents matching the discovered surface, runs \
|
||||
them in parallel, then validates every finding by cross-model voting before reporting.\n\n\
|
||||
@@ -1011,7 +1011,7 @@ pub(crate) fn spawn_engagement(base: &Path, mut cfg: RunConfig, mcp: bool, mode:
|
||||
println!(" │ ua : {ua}");
|
||||
write_status(&workdir, "running", &format!("\"target\":{:?}", cfg.target));
|
||||
|
||||
println!(" ┌─ NeuroSploit v4.0.0 · by Joas A Santos & Red Team Leaders");
|
||||
println!(" ┌─ NeuroSploit v4.1.0 · by Joas A Santos & Red Team Leaders");
|
||||
println!(" │ run id : {run_id}");
|
||||
println!(" │ target : {}", cfg.target);
|
||||
println!(" │ models : {}", cfg.models.join(", "));
|
||||
|
||||
@@ -1,4 +1,4 @@
|
||||
//! NeuroSploit v4.0.0 — interactive session (Claude-Code / Codex / Cursor-CLI style).
|
||||
//! NeuroSploit v4.1.0 — interactive session (Claude-Code / Codex / Cursor-CLI style).
|
||||
//!
|
||||
//! Launched when `neurosploit` runs with no subcommand. A persistent REPL with
|
||||
//! real line editing (arrow-key history recall, Ctrl-A/E/K, paste), model
|
||||
@@ -440,7 +440,7 @@ pub async fn repl(base: &Path, auth: SessionAuth) -> anyhow::Result<()> {
|
||||
let backends = harness::installed_cli_backends();
|
||||
println!("\x1b[1m");
|
||||
println!(" ███╗ ██╗███████╗██╗ ██╗██████╗ ██████╗");
|
||||
println!(" ████╗ ██║██╔════╝██║ ██║██╔══██╗██╔═══██╗ NeuroSploit v4.0.0");
|
||||
println!(" ████╗ ██║██╔════╝██║ ██║██╔══██╗██╔═══██╗ NeuroSploit v4.1.0");
|
||||
println!(" ██╔██╗ ██║█████╗ ██║ ██║██████╔╝██║ ██║ interactive harness");
|
||||
println!(" ██║╚██╗██║██╔══╝ ██║ ██║██╔══██╗██║ ██║ by Joas A Santos");
|
||||
println!(" ██║ ╚████║███████╗╚██████╔╝██║ ██║╚██████╔╝ & Red Team Leaders");
|
||||
|
||||
@@ -1,4 +1,4 @@
|
||||
//! NeuroSploit v4.0.0 — TUI "Mission Control" mode.
|
||||
//! NeuroSploit v4.1.0 — TUI "Mission Control" mode.
|
||||
//!
|
||||
//! Concurrent panels that update live while the engagement runs in the
|
||||
//! background, with a composer input that stays active during execution:
|
||||
|
||||
@@ -242,7 +242,7 @@ pub fn html_with_pocs(target: &str, findings: &[Finding], meta: &EngagementMeta,
|
||||
<h2>Executive Summary</h2><div class=summary-grid>{summary_grid}</div>\
|
||||
{vuln_summary}\
|
||||
<h2>Findings ({n})</h2>{body}\
|
||||
<p class=footer>Authorized testing only. Confirmed findings passed multi-model voting, receipt grounding and adversarial refute; \"needs-review\" are flagged for a human.<br>NeuroSploit v4.0.0 · by <b>Joas A Santos</b> & <b>Red Team Leaders</b><br><span style=\"font-family:ui-monospace,monospace\">{provenance}</span></p></body></html>",
|
||||
<p class=footer>Authorized testing only. Confirmed findings passed multi-model voting, receipt grounding and adversarial refute; \"needs-review\" are flagged for a human.<br>NeuroSploit v4.1.0 · by <b>Joas A Santos</b> & <b>Red Team Leaders</b><br><span style=\"font-family:ui-monospace,monospace\">{provenance}</span></p></body></html>",
|
||||
t = esc(target), n = sorted.len(), body = body, summary_grid = summary_grid, vuln_summary = vuln_summary,
|
||||
// Which build produced this document. A report that circulates without
|
||||
// it is a report nobody can trace back to the run that made it.
|
||||
|
||||
@@ -1,4 +1,4 @@
|
||||
// NeuroSploit v3.5.1 — Typst report template (blank, structured).
|
||||
// NeuroSploit v4.1.0 — Typst report template (blank, structured).
|
||||
//
|
||||
// The harness generates `report.typ` per run by prepending a `findings` array
|
||||
// and a `meta` dict, then including this template's rendering logic. This file
|
||||
@@ -53,7 +53,7 @@
|
||||
|
||||
#set page(margin: 2cm, numbering: "1", footer: context [
|
||||
#set text(size: 8pt, fill: gray)
|
||||
NeuroSploit v4.0.0 · #meta.target · confidential
|
||||
NeuroSploit v4.1.0 · #meta.target · confidential
|
||||
#h(1fr)
|
||||
// Build+run identity, so a page that circulates on its own still says which
|
||||
// engagement produced it.
|
||||
|
||||
+1
-1
@@ -22,7 +22,7 @@ run this only on a trusted machine/network, same trust model as the CLI itself.
|
||||
Server/version info.
|
||||
|
||||
```json
|
||||
{ "version": "4.0.0", "binary": "/opt/neurosploit-rs/neurosploit-rs/target/release/neurosploit", "root": "/opt/neurosploit-rs" }
|
||||
{ "version": "4.1.0", "binary": "/opt/neurosploit-rs/neurosploit-rs/target/release/neurosploit", "root": "/opt/neurosploit-rs" }
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
+1
-1
@@ -1,4 +1,4 @@
|
||||
# NeuroSploit v4.0.0 — web console
|
||||
# NeuroSploit v4.1.0 — web console
|
||||
|
||||
A browser UI for the `neurosploit` CLI harness: a 5-step engagement wizard (Asset → Scope & Auth
|
||||
→ Leads → Model & Run → Review), a live structured findings view with a generative attack-path
|
||||
|
||||
+2
-2
@@ -1,8 +1,8 @@
|
||||
{
|
||||
"name": "neurosploit-web",
|
||||
"version": "4.0.0",
|
||||
"version": "4.1.0",
|
||||
"private": true,
|
||||
"description": "NeuroSploit v4.0.0 web console — lead board + REPL, backed by the neurosploit CLI harness.",
|
||||
"description": "NeuroSploit v4.1.0 web console — lead board + REPL, backed by the neurosploit CLI harness.",
|
||||
"main": "server.js",
|
||||
"scripts": {
|
||||
"start": "node server.js"
|
||||
|
||||
+1
-1
@@ -1,5 +1,5 @@
|
||||
'use strict';
|
||||
/* NeuroSploit v4.0.0 — web console frontend. Vanilla JS, no build step. */
|
||||
/* NeuroSploit v4.1.0 — web console frontend. Vanilla JS, no build step. */
|
||||
|
||||
const $ = (sel, root = document) => root.querySelector(sel);
|
||||
const $$ = (sel, root = document) => Array.from(root.querySelectorAll(sel));
|
||||
|
||||
@@ -3,7 +3,7 @@
|
||||
<head>
|
||||
<meta charset="utf-8" />
|
||||
<meta name="viewport" content="width=device-width, initial-scale=1" />
|
||||
<title>NeuroSploit v4.0.0 — Console</title>
|
||||
<title>NeuroSploit v4.1.0 — Console</title>
|
||||
<link rel="icon" href="data:image/svg+xml,<svg xmlns=%22http://www.w3.org/2000/svg%22 viewBox=%220 0 100 100%22><text y=%22.9em%22 font-size=%2290%22>🧠</text></svg>">
|
||||
<link rel="stylesheet" href="/vendor/xterm.css" />
|
||||
<link rel="stylesheet" href="/style.css" />
|
||||
@@ -33,7 +33,7 @@
|
||||
<div class="sb-groups" id="sbGroups"><!-- populated by app.js --></div>
|
||||
|
||||
<div class="sb-bottom">
|
||||
<span class="sb-version" id="sbVersion">v4.0.0</span>
|
||||
<span class="sb-version" id="sbVersion">v4.1.0</span>
|
||||
<div class="sb-bottom-actions">
|
||||
<button class="icon-btn" id="btnOpenAuth" title="Auth & API keys">🔑</button>
|
||||
<button class="icon-btn" id="btnOpenRepl" title="Open terminal (Ctrl+`)">❭_</button>
|
||||
|
||||
@@ -1,4 +1,4 @@
|
||||
/* NeuroSploit v4.0.0 — web console.
|
||||
/* NeuroSploit v4.1.0 — web console.
|
||||
Visual direction: dense security-operations console (not a marketing SaaS
|
||||
page). Borders over shadows, typography over color, two radii, one accent.
|
||||
*/
|
||||
|
||||
+2
-2
@@ -1,7 +1,7 @@
|
||||
#!/usr/bin/env node
|
||||
'use strict';
|
||||
/**
|
||||
* NeuroSploit v4.0.0 — web console backend.
|
||||
* NeuroSploit v4.1.0 — web console backend.
|
||||
*
|
||||
* Zero-dependency Node HTTP server that:
|
||||
* - serves the static SPA in ./public
|
||||
@@ -1223,7 +1223,7 @@ const server = http.createServer(async (req, res) => {
|
||||
});
|
||||
|
||||
server.listen(PORT, () => {
|
||||
console.log(`NeuroSploit v4.0.0 web console → http://localhost:${PORT}`);
|
||||
console.log(`NeuroSploit v4.1.0 web console → http://localhost:${PORT}`);
|
||||
console.log(` binary : ${BIN || '(not found — build neurosploit-rs first)'}`);
|
||||
console.log(` agents : ${AGENTS_DIR}`);
|
||||
console.log(` runs : ${RUNS_DIR}`);
|
||||
|
||||
Reference in New Issue
Block a user