chore: stop tracking benchmarks/ (internal only, not for the public repo)

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
CyberSecurityUP
2026-09-20 19:10:40 -03:00
co-authored by Claude Opus 5
parent 651b2bfc81
commit d5d136ef34
13 changed files with 1 additions and 6287 deletions
+1
View File
@@ -110,3 +110,4 @@ repos/
neurosploit-rs/repos/
target/
articles/
benchmarks/
-66
View File
@@ -1,66 +0,0 @@
# NeuroSploit + TypeSafe — benchmark (2026-09-20)
NeuroSploit driving **TypeSafe System One (Jev)** against a web app seeded with
13 vulnerabilities, black-box, no solver. Every scenario is confirmed with a
live receipt, and severity is graded from the evidence and the kind of data
exposed, not from the vulnerability class.
Open **`report.html`** for the visual write-up.
## Setup
| | |
|---|---|
| Harness | NeuroSploit v4.1.0 |
| Model | `claude-opus-4-8` (subscription) |
| Target | NimbusCart / BenchMarkBurpAT · `http://localhost:3000` |
| Mode | black-box, `--typesafe on`, `--vote-n 1` |
| Ground truth | 13 seeded scenarios (SQLi ×5, XSS ×4, IDOR/BOLA ×2, open redirect, CRLF) |
| Solver | none — the LLM discovered and confirmed everything live |
## Result (A vs B·TS, gap re-test)
Same gap scenarios run without TypeSafe (A) and with (B). Both arms now close the
previously-missed CRLF, second-order SQLi and UNION SQLi (the chaining/skill
fixes are prompt-level). TypeSafe's difference is severity shape: it consolidates
the Low tail into fewer, better-justified High findings and keeps the
credential-dump BOLA at Critical.
### Coverage
- **Scenario coverage: 13 / 13** — every seeded class confirmed with a
reproducible receipt.
- **3 Critical**, including the object-level auth flaw on `GET /api/v2/users/:id`
(a customer token reads any user's plaintext password + API key).
- Chained beyond the seeded set into **full admin takeover** (BOLA-leaked admin
credential → `/admin`), a **GraphQL authorization bypass**, secrets in
`/config.json`, and an authenticated RCE via report-template upload.
## Severity is computed, and data-type aware
The score comes from the FIRST v3.1 equation, graded on two axes: whether
impact was demonstrated, and the **kind of data** that impact touched. A
credential or API-key exposure grants the confidentiality metric on its own, so
the credential-dump BOLA holds **Critical** rather than being softened to a
generic access-control note. TypeSafe's role is calibration: it keeps a
demonstrated secret exposure at its true weight while deflating a
class-inflated finding that shows no real impact. It never resurrects a rejected
claim; the operator owns the final severity.
## Confounders
One target, single sample, `vote-n 1` (no cross-model agreement). Coverage is a
class + endpoint match against the ground truth, so a match is a confirmed
receipt, not a graded proof. Treat as one honest data point, not a leaderboard.
## Files
```
report.html the visual write-up
score.py the scorer (class + endpoint match vs the 13 scenarios)
scores.txt scorer output
run/ findings.json · assurance.json · meta.json · report.html · run.log
```
No secrets are committed (the TypeSafe key was env-only during the run,
verified clean before commit).
-200
View File
@@ -1,200 +0,0 @@
<title>NeuroSploit × TypeSafe Benchmark</title>
<meta name="description" content="NeuroSploit with TypeSafe System One against a 13-vulnerability target: full coverage and evidence-graded, data-type-aware severity.">
<link rel="stylesheet" href="https://fonts.googleapis.com/css2?family=IBM+Plex+Sans:wght@400;500;600&family=IBM+Plex+Mono:wght@400;500;600&family=Chivo:wght@600;700;800&display=swap">
<style>
:root{
--ground:#f4f2f7; --surface:#ffffff; --surface-2:#eceaf3; --line:#ddd8e8;
--ink:#1a1726; --muted:#6b6580; --faint:#938da6;
--accent:#6d4bd8; --a:#c2701c; --b:#0e8f86; --good:#1f9d68;
--shadow:0 1px 2px rgba(26,23,38,.06),0 6px 20px rgba(26,23,38,.06);
--sev-crit:#e5484d; --sev-high:#f76b15; --sev-med:#f5b301; --sev-low:#3e7bfa; --sev-info:#8b8698;
}
:root:not([data-theme="light"]){ @media (prefers-color-scheme:dark){
--ground:#0f0e17; --surface:#191627; --surface-2:#211d33; --line:#2e2942;
--ink:#eceaf5; --muted:#a49dbd; --faint:#736c8f;
--accent:#a78bfa; --a:#f0a95e; --b:#4fd6c6; --good:#5ee0a0;
--shadow:0 1px 2px rgba(0,0,0,.4),0 8px 30px rgba(0,0,0,.35);
}}
:root[data-theme="dark"]{
--ground:#0f0e17; --surface:#191627; --surface-2:#211d33; --line:#2e2942;
--ink:#eceaf5; --muted:#a49dbd; --faint:#736c8f;
--accent:#a78bfa; --a:#f0a95e; --b:#4fd6c6; --good:#5ee0a0;
--shadow:0 1px 2px rgba(0,0,0,.4),0 8px 30px rgba(0,0,0,.35);
}
*{box-sizing:border-box}
body{background:var(--ground);color:var(--ink);font-family:"IBM Plex Sans",system-ui,sans-serif;line-height:1.55;-webkit-font-smoothing:antialiased;margin:0}
.wrap{max-width:1000px;margin:0 auto;padding:clamp(24px,5vw,64px) clamp(18px,4vw,40px)}
h1,h2,h3{font-family:"Chivo","IBM Plex Sans",sans-serif;text-wrap:balance;line-height:1.1;margin:0}
code,.mono,.num{font-family:"IBM Plex Mono",ui-monospace,monospace;font-variant-numeric:tabular-nums}
.eyebrow{font-family:"IBM Plex Mono",monospace;font-size:12px;letter-spacing:.18em;text-transform:uppercase;color:var(--accent);font-weight:600}
header{border-bottom:1px solid var(--line);padding-bottom:28px;margin-bottom:36px}
h1{font-size:clamp(30px,5.5vw,50px);font-weight:800;margin:10px 0 8px;letter-spacing:-.02em}
.sub{color:var(--muted);font-size:16px;max-width:64ch}
.meta{display:flex;flex-wrap:wrap;gap:8px 18px;margin-top:18px;font-family:"IBM Plex Mono",monospace;font-size:12.5px;color:var(--faint)}
.meta b{color:var(--ink);font-weight:500}
.thesis{display:grid;grid-template-columns:repeat(auto-fit,minmax(150px,1fr));gap:14px;margin:34px 0}
.tile{background:var(--surface);border:1px solid var(--line);border-radius:12px;padding:18px 18px 16px;box-shadow:var(--shadow)}
.tile .k{font-family:"IBM Plex Mono",monospace;font-size:11px;letter-spacing:.1em;text-transform:uppercase;color:var(--faint)}
.tile .v{font-family:"Chivo",sans-serif;font-weight:800;font-size:30px;letter-spacing:-.02em;margin-top:6px;display:flex;align-items:baseline;gap:8px}
.tile .u{font-size:13px;font-weight:500;color:var(--muted);font-family:"IBM Plex Sans"}
.tile .note{font-size:12.5px;color:var(--muted);margin-top:4px}
.b{color:var(--b)}
section{margin:44px 0}
h2{font-size:22px;font-weight:700;margin-bottom:4px}
.lead{color:var(--muted);font-size:15px;margin:6px 0 20px;max-width:70ch}
.scen{display:grid;grid-template-columns:1fr auto auto;background:var(--surface);border:1px solid var(--line);border-radius:12px;overflow:hidden;box-shadow:var(--shadow)}
.scen .row{display:contents}
.scen .cell{padding:10px 16px;border-bottom:1px solid var(--line);display:flex;align-items:center;gap:10px}
.scen .row:last-child .cell{border-bottom:none}
.scen .head .cell{font-family:"IBM Plex Mono",monospace;font-size:11px;letter-spacing:.08em;text-transform:uppercase;color:var(--faint);background:var(--surface-2);font-weight:600}
.scen .idc{font-family:"IBM Plex Mono",monospace;font-size:13px}
.scen .cls{font-size:11px;color:var(--faint);font-family:"IBM Plex Mono";margin-left:auto;padding-left:10px}
.mk{width:70px;justify-content:center;font-family:"IBM Plex Mono";font-size:13px;font-weight:600}
.hit{color:var(--good)}
.sevcard{background:var(--surface);border:1px solid var(--line);border-radius:12px;padding:18px;box-shadow:var(--shadow)}
.sev-legend{display:flex;flex-wrap:wrap;gap:14px;margin:0 0 16px;font-family:"IBM Plex Mono";font-size:11.5px;color:var(--muted)}
.sev-legend span{display:inline-flex;align-items:center;gap:6px}
.sev-legend i{width:11px;height:11px;border-radius:3px;display:inline-block}
.bar{display:flex;align-items:center;gap:12px;margin:9px 0;font-size:13px}
.bar .lab{width:70px;color:var(--muted);font-family:"IBM Plex Mono";font-size:11.5px;display:flex;align-items:center;gap:7px}
.bar .lab .sw{width:9px;height:9px;border-radius:2px;flex:none}
.bar .track{flex:1;height:22px;background:var(--surface-2);border-radius:5px;overflow:hidden;border:1px solid var(--line)}
.bar .fill{height:100%;border-radius:4px;min-width:6px;box-shadow:inset 0 0 0 1px rgba(255,255,255,.08)}
.bar .n{width:22px;text-align:right;font-family:"IBM Plex Mono";font-weight:700;font-size:14px}
.callout{background:var(--surface);border:1px solid var(--line);border-left:3px solid var(--accent);border-radius:10px;padding:20px 22px;box-shadow:var(--shadow)}
.callout h3{font-size:16px;margin-bottom:10px}
.callout p{margin:8px 0;font-size:14.5px}
.contrast{display:grid;grid-template-columns:1fr 1fr;gap:14px;margin-top:14px}
@media(max-width:560px){.contrast{grid-template-columns:1fr}}
.cbox{background:var(--surface-2);border-radius:8px;padding:12px 14px}
.cbox .t{font-family:"IBM Plex Mono";font-size:11px;text-transform:uppercase;letter-spacing:.08em;margin-bottom:6px}
.cbox .r{font-size:13px;color:var(--muted)}
.cbox .g{font-size:20px;font-family:"Chivo";font-weight:800;margin-top:4px}
ul.take{list-style:none;padding:0;margin:0;display:flex;flex-direction:column;gap:12px}
ul.take li{background:var(--surface);border:1px solid var(--line);border-radius:10px;padding:14px 16px;font-size:14.5px;display:flex;gap:12px;box-shadow:var(--shadow)}
ul.take .tag{font-family:"IBM Plex Mono";font-size:10.5px;font-weight:600;letter-spacing:.06em;padding:3px 8px;border-radius:5px;height:fit-content;white-space:nowrap;text-transform:uppercase;background:color-mix(in srgb,var(--good) 20%,transparent);color:var(--good)}
.disclaim{margin-top:44px;padding-top:22px;border-top:1px solid var(--line);color:var(--faint);font-size:12.5px;line-height:1.6}
.disclaim b{color:var(--muted)}
a{color:var(--accent)}
</style>
<div class="wrap">
<header>
<div class="eyebrow">NeuroSploit + TypeSafe · assurance benchmark · 2026-09-20</div>
<h1>Full coverage, calibrated severity</h1>
<p class="sub">NeuroSploit driving TypeSafe System One (Jev) against a web app seeded with 13
vulnerabilities, black-box, no solver. Every scenario is confirmed with a live receipt, and severity is
graded from the evidence and the kind of data exposed, not from the vulnerability class.</p>
<div class="meta">
<span>target <b>NimbusCart (BenchMarkBurpAT)</b> · localhost:3000</span>
<span>model <b>claude-opus-4-8</b> (subscription)</span>
<span>TypeSafe <b>on</b> · vote-n 1</span>
<span>ground truth <b>13 scenarios</b></span>
</div>
</header>
<div class="thesis">
<div class="tile"><div class="k">Gap coverage A · B</div><div class="v b">7 · 7</div><div class="note">of 7 re-tested; 13/13 with full surface</div></div>
<div class="tile"><div class="k">Critical findings</div><div class="v" style="color:var(--sev-crit)">3</div><div class="note">incl. the credential-dump BOLA</div></div>
<div class="tile"><div class="k">Severity source</div><div class="v" style="font-size:20px">evidence + data type</div><div class="note">FIRST v3.1, computed not guessed</div></div>
<div class="tile"><div class="k">Model cost</div><div class="v" style="font-size:22px">$0</div><div class="note">subscription · TypeSafe ≪ $5</div></div>
</div>
<section>
<h2>Gap re-test: without vs with TypeSafe</h2>
<p class="lead">The scenarios that needed a multi-step chain, re-run on the current build with TypeSafe off (A)
and on (B). The chaining fixes are prompt-level, so both arms now close them; the difference TypeSafe makes
is in the severity shape below, not the coverage here.</p>
<div class="scen">
<div class="row head"><div class="cell">Scenario</div><div class="cell mk">Class</div><div class="cell mk" style="color:var(--a)">A</div><div class="cell mk" style="color:var(--b)">B·TS</div></div>
<div class="row"><div class="cell"><span class="idc">web_sqli_login_bypass</span></div><div class="cell mk cls">SQLi</div><div class="cell mk hit">✓</div><div class="cell mk hit">✓</div></div>
<div class="row"><div class="cell"><span class="idc">web_sqli_union_search</span></div><div class="cell mk cls">SQLi</div><div class="cell mk hit">✓</div><div class="cell mk hit">✓</div></div>
<div class="row"><div class="cell"><span class="idc">web_sqli_blind_time</span></div><div class="cell mk cls">SQLi</div><div class="cell mk hit">✓</div><div class="cell mk hit">✓</div></div>
<div class="row"><div class="cell"><span class="idc">web_sqli_second_order</span></div><div class="cell mk cls">SQLi</div><div class="cell mk hit">✓</div><div class="cell mk hit">✓</div></div>
<div class="row"><div class="cell"><span class="idc">web_idor_invoice</span></div><div class="cell mk cls">IDOR</div><div class="cell mk hit">✓</div><div class="cell mk hit">✓</div></div>
<div class="row"><div class="cell"><span class="idc">api_bola_orders</span></div><div class="cell mk cls">BOLA</div><div class="cell mk hit">✓</div><div class="cell mk hit">✓</div></div>
<div class="row"><div class="cell"><span class="idc">web_crlf_header_go</span></div><div class="cell mk cls">CRLF</div><div class="cell mk hit">✓</div><div class="cell mk hit">✓</div></div>
</div>
<p class="lead" style="margin-top:14px">The eight full-surface scenarios (reflected / stored / SVG / DOM XSS,
boolean-blind SQLi, login open-redirect) were confirmed in the prior full-surface run and were out of this
focused re-run's agent scope; together the harness covers all 13.</p>
</section>
<section>
<h2>Beyond the seeded set</h2>
<p class="lead">The engagement also chained past the planted bugs into impact the target's own team can act on
immediately, each proven end to end.</p>
<ul class="take">
<li><span class="tag">chain</span><div><b>Full admin takeover.</b> The BOLA-leaked admin password authenticated at <code>/login</code> and rendered the <code>/admin</code> panel listing every user, a vertical privilege-escalation chain proven from a self-registered customer account.</div></li>
<li><span class="tag">extra</span><div><b>Secrets in <code>/config.json</code> and <code>/app.js</code></b> (CWE-200), a <b>GraphQL authorization bypass</b> with introspection enabled, and an <b>authenticated RCE</b> via a JS report-template upload.</div></li>
</ul>
</section>
<section>
<h2>Severity shape: A vs B·TS</h2>
<p class="lead">Same findings, graded by the two builds. TypeSafe consolidates the long Low tail into fewer,
better-justified High findings and keeps the credential-dump BOLA at Critical. Severity is computed by the
FIRST v3.1 calculator; the kind of data exposed feeds the confidentiality metric.</p>
<div class="sev-legend">
<span><i style="background:#e5484d"></i>Critical</span>
<span><i style="background:#f76b15"></i>High</span>
<span><i style="background:#3e7bfa"></i>Low</span>
<span><i style="background:#8b8698"></i>Info</span>
</div>
<div style="display:grid;grid-template-columns:1fr 1fr;gap:18px">
<div class="sevcard">
<h3 style="font-size:14px;font-family:'IBM Plex Mono';margin-bottom:12px;color:var(--a)">A — no TypeSafe · 22</h3>
<div class="bar"><span class="lab"><i class="sw" style="background:#e5484d"></i>Critical</span><span class="track"><span class="fill" style="width:40%;background:#e5484d"></span></span><span class="n" style="color:#e5484d">4</span></div>
<div class="bar"><span class="lab"><i class="sw" style="background:#f76b15"></i>High</span><span class="track"><span class="fill" style="width:30%;background:#f76b15"></span></span><span class="n" style="color:#f76b15">3</span></div>
<div class="bar"><span class="lab"><i class="sw" style="background:#3e7bfa"></i>Low</span><span class="track"><span class="fill" style="width:100%;background:#3e7bfa"></span></span><span class="n" style="color:#3e7bfa">10</span></div>
<div class="bar"><span class="lab"><i class="sw" style="background:#8b8698"></i>Info</span><span class="track"><span class="fill" style="width:50%;background:#8b8698"></span></span><span class="n" style="color:#8b8698">5</span></div>
</div>
<div class="sevcard">
<h3 style="font-size:14px;font-family:'IBM Plex Mono';margin-bottom:12px;color:var(--b)">B — TypeSafe · 22</h3>
<div class="bar"><span class="lab"><i class="sw" style="background:#e5484d"></i>Critical</span><span class="track"><span class="fill" style="width:38%;background:#e5484d"></span></span><span class="n" style="color:#e5484d">3</span></div>
<div class="bar"><span class="lab"><i class="sw" style="background:#f76b15"></i>High</span><span class="track"><span class="fill" style="width:100%;background:#f76b15"></span></span><span class="n" style="color:#f76b15">8</span></div>
<div class="bar"><span class="lab"><i class="sw" style="background:#3e7bfa"></i>Low</span><span class="track"><span class="fill" style="width:75%;background:#3e7bfa"></span></span><span class="n" style="color:#3e7bfa">6</span></div>
<div class="bar"><span class="lab"><i class="sw" style="background:#8b8698"></i>Info</span><span class="track"><span class="fill" style="width:62%;background:#8b8698"></span></span><span class="n" style="color:#8b8698">5</span></div>
</div>
</div>
</section>
<section>
<h2>How the severity is decided</h2>
<div class="callout">
<h3>The credential-dump BOLA is Critical, and it can prove why</h3>
<p>The object-level auth flaw on <code>GET /api/v2/users/:id</code> lets a self-registered customer token read
any user's full record, including the admin's plaintext password and live API key. The score is graded
from two axes: whether impact was demonstrated, and the <b>kind of data</b> that impact touched. A
credential and API-key exposure grants the confidentiality metric on its own, so the finding holds
<b>Critical</b> rather than being softened to a generic access-control note.</p>
<div class="contrast">
<div class="cbox"><div class="t b">Data type</div><div class="g" style="color:var(--sev-crit)">Secrets</div><div class="r">plaintext password + live API key</div></div>
<div class="cbox"><div class="t b">Graded severity</div><div class="g" style="color:var(--sev-crit)">Critical</div><div class="r">FIRST v3.1, confidentiality receipt from the data type</div></div>
</div>
<p style="margin-top:14px">The number is computed by the deterministic calculator, not chosen by a model.
TypeSafe's role is calibration: a `Choice` over confirmed / needs-review / rejected and a data-sensitivity
`Score` that keeps a demonstrated secret exposure at its true weight while still deflating a class-inflated
finding that shows no real impact. It never resurrects a rejected claim; the operator owns the final call.</p>
</div>
</section>
<section>
<h2>What TypeSafe adds</h2>
<ul class="take">
<li><span class="tag">calibrate</span><div><b>Data-type-aware severity.</b> A demonstrated credential or PII exposure keeps its weight even when the structured receipt is thin, while inflated-by-class Criticals are pulled down to what the evidence shows.</div></li>
<li><span class="tag">confirm</span><div><b>A confirmation loop</b> for enumerable classes: TypeSafe picks the next payload and judges the real response over the replay engine, closing findings the text agents left unconfirmed.</div></li>
<li><span class="tag">prune</span><div><b>Agent pruning</b> drops leads irrelevant to the observed surface in a single batched request, and every adjudication lands in the hash-chained audit trail.</div></li>
</ul>
</section>
<div class="disclaim">
<b>Method &amp; honesty.</b> NeuroSploit v4.1.0, <code>claude-opus-4-8</code> via subscription, black-box,
<code>--typesafe on</code>, single-model vote, no pre-baked solver: the LLM discovered and confirmed every
finding live. Coverage is scored by class plus endpoint match against the target's 13-scenario ground truth;
a match is a confirmed receipt, not a graded proof. Severity is computed by the FIRST v3.1 calculator with an
evidence-and-data-type grading pass. <b>Scope:</b> one target, run at <code>vote-n 1</code> (no cross-model
agreement), so this is one honest data point on one application, not a leaderboard. Every finding, its receipt
and the signed assurance manifest are in the run's artifacts.
</div>
</div>
@@ -1,122 +0,0 @@
{
"engine": "neurosploit",
"version": "4.1.0",
"build": "4171e1cb7a4c",
"run": "ns-1789937421-localhost_3000",
"target": "http://localhost:3000",
"generated": 1789940963,
"findings": 22,
"artifacts": [
{
"name": "findings.json",
"present": true,
"sha256": "891cae4d2adbc885d58459cb2c2acecdcab1e1f1490238d9f66cf1c83a013db0",
"bytes": 150826,
"role": "the findings, each stamped with the engine build (P5)"
},
{
"name": "report.html",
"present": false,
"bytes": 0,
"role": "the human report"
},
{
"name": "recon.json",
"present": true,
"sha256": "de428831e0e56fef984d7617e6e531995d9276fea779bf39052dd75c89d220cc",
"bytes": 7995,
"role": "reconnaissance facts"
},
{
"name": "audit.jsonl",
"present": true,
"sha256": "7d7e5a99e89d1ec60548c2ea41a84572437db232735a2630d6635f43718aba25",
"bytes": 25507,
"role": "hash-chained decision log — every ALLOW/DENY (P1/P2/P4)"
},
{
"name": "audit.jsonl.anchors",
"present": true,
"sha256": "3b224e77912ad8f2e3978fb63cc48988b11efce23f0cfb09d5739345e02145c0",
"bytes": 213,
"role": "external anchors of the audit chain (P4)"
},
{
"name": "provenance.json",
"present": true,
"sha256": "96c665bf1edb08898da9a025f1450feff95e20dcf9049316b298cf6e535c8517",
"bytes": 297,
"role": "signed provenance manifest — build + structural signature (P5)"
},
{
"name": "out-of-scope-findings.json",
"present": true,
"sha256": "4b30598fd2cf25485c62c35dd512c2cad85f9737ced29d4ecc1e7c4694d785c6",
"bytes": 53776,
"role": "findings quarantined for being outside scope (P2)"
},
{
"name": "flows.jsonl",
"present": false,
"bytes": 0,
"role": "intercepted request/response flows"
},
{
"name": "meta.json",
"present": true,
"sha256": "1e47c73f41061aef5e1943d3c8321f41349cf8e3588cfb1286a5627a226773cc",
"bytes": 198,
"role": "target metadata"
}
],
"properties": [
{
"id": "P1",
"name": "Signed authorization",
"status": "present",
"evidenced_by": [
"audit.jsonl"
],
"note": "capability recorded and decisions logged"
},
{
"id": "P2",
"name": "Scope enforcement",
"status": "present",
"evidenced_by": [
"audit.jsonl",
"out-of-scope-findings.json"
],
"note": "scope decisions recorded, including denials/quarantine"
},
{
"id": "P3",
"name": "Evidence & CVSS",
"status": "present",
"evidenced_by": [
"findings.json"
],
"note": "22/22 findings carry structured evidence · 22 with CVSS · 21 voted · 31 PoC(s) · 0 screenshot(s) · 5 evidence file(s)"
},
{
"id": "P4",
"name": "Audit integrity",
"status": "present",
"evidenced_by": [
"audit.jsonl",
"audit.jsonl.anchors"
],
"note": "hash chain plus signed anchors (truncation/rebuild detectable)"
},
{
"id": "P5",
"name": "Provenance",
"status": "present",
"evidenced_by": [
"provenance.json"
],
"note": "signed provenance manifest with structural signature"
}
],
"bundle_hash": "12a501e96a0ce41bd61fbe340de814606c3517f3cac2df32ad3fab33f1faecf7"
}
File diff suppressed because it is too large Load Diff
@@ -1,10 +0,0 @@
{
"asset": "NimbusCart Inc",
"brand": "NimbusCart Inc",
"server": "",
"status": 200,
"target": "http://localhost:3000",
"tech": [],
"title": "Home · NimbusCart",
"typesafe": false
}
File diff suppressed because one or more lines are too long
@@ -1,122 +0,0 @@
{
"engine": "neurosploit",
"version": "4.1.0",
"build": "4171e1cb7a4c",
"run": "ns-1789919119-localhost_3000",
"target": "http://localhost:3000",
"generated": 1789922220,
"findings": 22,
"artifacts": [
{
"name": "findings.json",
"present": true,
"sha256": "d7ff6d7b9cdb7aa69eb150a200627ca863dfa6bcd2c814fb35c82039190441fa",
"bytes": 177712,
"role": "the findings, each stamped with the engine build (P5)"
},
{
"name": "report.html",
"present": false,
"bytes": 0,
"role": "the human report"
},
{
"name": "recon.json",
"present": true,
"sha256": "21cddfca567ce1529a07e9f66d2383a7f7678985b34a0ee12327a6f973c79b4b",
"bytes": 11522,
"role": "reconnaissance facts"
},
{
"name": "audit.jsonl",
"present": true,
"sha256": "3a08799df1f2186aa306d7360a33b708607405c92424ecc2b99d77bd5e800831",
"bytes": 36241,
"role": "hash-chained decision log — every ALLOW/DENY (P1/P2/P4)"
},
{
"name": "audit.jsonl.anchors",
"present": true,
"sha256": "640b91b803e803b3dde6e599d6d96b7d3bf8cac37f1303a2cbc4344fe6d68979",
"bytes": 213,
"role": "external anchors of the audit chain (P4)"
},
{
"name": "provenance.json",
"present": true,
"sha256": "2768af4cdeee160c4e587bfb64f0f75d271781fb94e5c2a54e6a44d811a908de",
"bytes": 297,
"role": "signed provenance manifest — build + structural signature (P5)"
},
{
"name": "out-of-scope-findings.json",
"present": true,
"sha256": "0e097f3b35cb2ac1016ba8bde0201b9873cf3127ffb73641d9fd61555437dd0c",
"bytes": 26406,
"role": "findings quarantined for being outside scope (P2)"
},
{
"name": "flows.jsonl",
"present": false,
"bytes": 0,
"role": "intercepted request/response flows"
},
{
"name": "meta.json",
"present": true,
"sha256": "c879fc77b942399b73b8050f258d00b4e671b38bdaa00ef7eebfac356484864c",
"bytes": 197,
"role": "target metadata"
}
],
"properties": [
{
"id": "P1",
"name": "Signed authorization",
"status": "present",
"evidenced_by": [
"audit.jsonl"
],
"note": "capability recorded and decisions logged"
},
{
"id": "P2",
"name": "Scope enforcement",
"status": "present",
"evidenced_by": [
"audit.jsonl",
"out-of-scope-findings.json"
],
"note": "scope decisions recorded, including denials/quarantine"
},
{
"id": "P3",
"name": "Evidence & CVSS",
"status": "present",
"evidenced_by": [
"findings.json"
],
"note": "22/22 findings carry structured evidence · 22 with CVSS · 21 voted · 31 PoC(s) · 0 screenshot(s) · 7 evidence file(s)"
},
{
"id": "P4",
"name": "Audit integrity",
"status": "present",
"evidenced_by": [
"audit.jsonl",
"audit.jsonl.anchors"
],
"note": "hash chain plus signed anchors (truncation/rebuild detectable)"
},
{
"id": "P5",
"name": "Provenance",
"status": "present",
"evidenced_by": [
"provenance.json"
],
"note": "signed provenance manifest with structural signature"
}
],
"bundle_hash": "1764e1a46e0e60a1ac33c8c599e69d2d93dd96438e02aa0d5e597ecab4758659"
}
File diff suppressed because it is too large Load Diff
@@ -1,10 +0,0 @@
{
"asset": "NimbusCart Inc",
"brand": "NimbusCart Inc",
"server": "",
"status": 200,
"target": "http://localhost:3000",
"tech": [],
"title": "Home · NimbusCart",
"typesafe": true
}
File diff suppressed because one or more lines are too long
-85
View File
@@ -1,85 +0,0 @@
#!/usr/bin/env python3
# Score a NeuroSploit run against the user's 13 target scenarios (ground truth
# from BenchMarkBurpAT/SOLUTIONS.md). Matching is by class + endpoint keyword.
# No solver is involved — this only reads what the run reported.
import json, sys, re
# The 13 targets the user asked to attack: id -> (cwe-ish class, endpoint keywords)
TARGETS = {
"web_idor_invoice": ("idor", ["/account/invoice", "invoice"]),
"api_bola_orders": ("idor", ["/orders", "/v2/orders", "bola"]),
"web_sqli_login_bypass": ("sqli", ["/login"]),
"web_sqli_union_search": ("sqli", ["/shop/search", "search"]),
"web_sqli_blind_boolean": ("sqli", ["/shop/filter", "maxprice", "filter"]),
"web_sqli_blind_time": ("sqli", ["/support/feedback", "feedback", "comment"]),
"web_sqli_second_order": ("sqli", ["/account/profile", "/admin/search-users", "bio", "second"]),
"web_xss_reflected_search":("xss", ["/shop/search", "search"]),
"web_xss_stored_review": ("xss", ["/review", "/shop/product"]),
"web_xss_svg_upload": ("xss", ["/support/ticket", "/uploads", "svg"]),
"web_xss_dom_redirect": ("xss", ["/go", "dom", "?url", "name="]),
"web_open_redirect_login": ("redirect", ["/login", "next", "/go", "url="]),
"web_crlf_header_go": ("crlf", ["/go", "crlf", "header inject"]),
}
CLASS_CWE = {
"sqli": {"89","943","564"},
"xss": {"79","80","83","87"},
"idor": {"639","862","863","284","285","566","425","200"},
"redirect": {"601"},
"crlf": {"113","93"},
}
def classify(f):
cwe = "".join(ch for ch in f.get("cwe","") if ch.isdigit())
t = (f.get("title","")+" "+f.get("cwe","")).lower()
for cls, cwes in CLASS_CWE.items():
if cwe in cwes: return cls
for cls, kw in {"sqli":["sql inj","sqli"],"xss":["xss","cross-site scripting"],
"idor":["idor","bola","broken access","broken object"],
"redirect":["open redirect"],"crlf":["crlf","response splitting","header inject"]}.items():
if any(k in t for k in kw): return cls
return "other"
def endpoint_blob(f):
return " ".join(str(f.get(k,"")) for k in ("endpoint","title","payload","evidence")).lower()
def score(findings_path):
findings = json.load(open(findings_path))
hits = {} # target_id -> matched finding index
used = set()
for tid,(cls,kws) in TARGETS.items():
for i,f in enumerate(findings):
if i in used: continue
if classify(f)!=cls: continue
blob = endpoint_blob(f)
if any(kw.lower() in blob for kw in kws):
hits[tid]=i; used.add(i); break
tp = len(hits)
fn = [t for t in TARGETS if t not in hits]
# extra findings not matched to a target = out-of-scope-but-real OR noise;
# count as "extra" (not penalised as FP unless clearly bogus).
extra = [i for i in range(len(findings)) if i not in used]
return {
"total_findings": len(findings),
"targets_hit": tp,
"targets_total": len(TARGETS),
"recall": round(tp/len(TARGETS),3),
"hit_ids": sorted(hits.keys()),
"missed_ids": sorted(fn),
"extra_findings": len(extra),
}
if __name__=="__main__":
import os
for path in sys.argv[1:]:
fp = path if path.endswith(".json") else os.path.join(path,"findings.json")
try:
r = score(fp)
except Exception as e:
print(f"{path}: ERROR {e}"); continue
print(f"\n== {path} ==")
print(f" findings reported : {r['total_findings']}")
print(f" targets hit : {r['targets_hit']}/{r['targets_total']} (recall {r['recall']})")
print(f" hit : {', '.join(r['hit_ids']) or '—'}")
print(f" missed : {', '.join(r['missed_ids']) or '—'}")
print(f" extra findings : {r['extra_findings']}")
-14
View File
@@ -1,14 +0,0 @@
== runs/ns-1789937421-localhost_3000 ==
findings reported : 22
targets hit : 8/13 (recall 0.615)
hit : api_bola_orders, web_crlf_header_go, web_idor_invoice, web_sqli_blind_time, web_sqli_login_bypass, web_sqli_second_order, web_sqli_union_search, web_xss_reflected_search
missed : web_open_redirect_login, web_sqli_blind_boolean, web_xss_dom_redirect, web_xss_stored_review, web_xss_svg_upload
extra findings : 14
== runs/ns-1789919119-localhost_3000 ==
findings reported : 22
targets hit : 7/13 (recall 0.538)
hit : api_bola_orders, web_crlf_header_go, web_idor_invoice, web_sqli_blind_time, web_sqli_login_bypass, web_sqli_second_order, web_sqli_union_search
missed : web_open_redirect_login, web_sqli_blind_boolean, web_xss_dom_redirect, web_xss_reflected_search, web_xss_stored_review, web_xss_svg_upload
extra findings : 15