mirror of
https://github.com/CyberSecurityUP/NeuroSploit.git
synced 2026-09-28 20:11:49 +02:00
chore: stop tracking benchmarks/ (internal only, not for the public repo)
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
co-authored by
Claude Opus 5
parent
651b2bfc81
commit
d5d136ef34
@@ -110,3 +110,4 @@ repos/
|
||||
neurosploit-rs/repos/
|
||||
target/
|
||||
articles/
|
||||
benchmarks/
|
||||
|
||||
@@ -1,66 +0,0 @@
|
||||
# NeuroSploit + TypeSafe — benchmark (2026-09-20)
|
||||
|
||||
NeuroSploit driving **TypeSafe System One (Jev)** against a web app seeded with
|
||||
13 vulnerabilities, black-box, no solver. Every scenario is confirmed with a
|
||||
live receipt, and severity is graded from the evidence and the kind of data
|
||||
exposed, not from the vulnerability class.
|
||||
|
||||
Open **`report.html`** for the visual write-up.
|
||||
|
||||
## Setup
|
||||
|
||||
| | |
|
||||
|---|---|
|
||||
| Harness | NeuroSploit v4.1.0 |
|
||||
| Model | `claude-opus-4-8` (subscription) |
|
||||
| Target | NimbusCart / BenchMarkBurpAT · `http://localhost:3000` |
|
||||
| Mode | black-box, `--typesafe on`, `--vote-n 1` |
|
||||
| Ground truth | 13 seeded scenarios (SQLi ×5, XSS ×4, IDOR/BOLA ×2, open redirect, CRLF) |
|
||||
| Solver | none — the LLM discovered and confirmed everything live |
|
||||
|
||||
## Result (A vs B·TS, gap re-test)
|
||||
|
||||
Same gap scenarios run without TypeSafe (A) and with (B). Both arms now close the
|
||||
previously-missed CRLF, second-order SQLi and UNION SQLi (the chaining/skill
|
||||
fixes are prompt-level). TypeSafe's difference is severity shape: it consolidates
|
||||
the Low tail into fewer, better-justified High findings and keeps the
|
||||
credential-dump BOLA at Critical.
|
||||
|
||||
### Coverage
|
||||
|
||||
- **Scenario coverage: 13 / 13** — every seeded class confirmed with a
|
||||
reproducible receipt.
|
||||
- **3 Critical**, including the object-level auth flaw on `GET /api/v2/users/:id`
|
||||
(a customer token reads any user's plaintext password + API key).
|
||||
- Chained beyond the seeded set into **full admin takeover** (BOLA-leaked admin
|
||||
credential → `/admin`), a **GraphQL authorization bypass**, secrets in
|
||||
`/config.json`, and an authenticated RCE via report-template upload.
|
||||
|
||||
## Severity is computed, and data-type aware
|
||||
|
||||
The score comes from the FIRST v3.1 equation, graded on two axes: whether
|
||||
impact was demonstrated, and the **kind of data** that impact touched. A
|
||||
credential or API-key exposure grants the confidentiality metric on its own, so
|
||||
the credential-dump BOLA holds **Critical** rather than being softened to a
|
||||
generic access-control note. TypeSafe's role is calibration: it keeps a
|
||||
demonstrated secret exposure at its true weight while deflating a
|
||||
class-inflated finding that shows no real impact. It never resurrects a rejected
|
||||
claim; the operator owns the final severity.
|
||||
|
||||
## Confounders
|
||||
|
||||
One target, single sample, `vote-n 1` (no cross-model agreement). Coverage is a
|
||||
class + endpoint match against the ground truth, so a match is a confirmed
|
||||
receipt, not a graded proof. Treat as one honest data point, not a leaderboard.
|
||||
|
||||
## Files
|
||||
|
||||
```
|
||||
report.html the visual write-up
|
||||
score.py the scorer (class + endpoint match vs the 13 scenarios)
|
||||
scores.txt scorer output
|
||||
run/ findings.json · assurance.json · meta.json · report.html · run.log
|
||||
```
|
||||
|
||||
No secrets are committed (the TypeSafe key was env-only during the run,
|
||||
verified clean before commit).
|
||||
@@ -1,200 +0,0 @@
|
||||
<title>NeuroSploit × TypeSafe Benchmark</title>
|
||||
<meta name="description" content="NeuroSploit with TypeSafe System One against a 13-vulnerability target: full coverage and evidence-graded, data-type-aware severity.">
|
||||
<link rel="stylesheet" href="https://fonts.googleapis.com/css2?family=IBM+Plex+Sans:wght@400;500;600&family=IBM+Plex+Mono:wght@400;500;600&family=Chivo:wght@600;700;800&display=swap">
|
||||
<style>
|
||||
:root{
|
||||
--ground:#f4f2f7; --surface:#ffffff; --surface-2:#eceaf3; --line:#ddd8e8;
|
||||
--ink:#1a1726; --muted:#6b6580; --faint:#938da6;
|
||||
--accent:#6d4bd8; --a:#c2701c; --b:#0e8f86; --good:#1f9d68;
|
||||
--shadow:0 1px 2px rgba(26,23,38,.06),0 6px 20px rgba(26,23,38,.06);
|
||||
--sev-crit:#e5484d; --sev-high:#f76b15; --sev-med:#f5b301; --sev-low:#3e7bfa; --sev-info:#8b8698;
|
||||
}
|
||||
:root:not([data-theme="light"]){ @media (prefers-color-scheme:dark){
|
||||
--ground:#0f0e17; --surface:#191627; --surface-2:#211d33; --line:#2e2942;
|
||||
--ink:#eceaf5; --muted:#a49dbd; --faint:#736c8f;
|
||||
--accent:#a78bfa; --a:#f0a95e; --b:#4fd6c6; --good:#5ee0a0;
|
||||
--shadow:0 1px 2px rgba(0,0,0,.4),0 8px 30px rgba(0,0,0,.35);
|
||||
}}
|
||||
:root[data-theme="dark"]{
|
||||
--ground:#0f0e17; --surface:#191627; --surface-2:#211d33; --line:#2e2942;
|
||||
--ink:#eceaf5; --muted:#a49dbd; --faint:#736c8f;
|
||||
--accent:#a78bfa; --a:#f0a95e; --b:#4fd6c6; --good:#5ee0a0;
|
||||
--shadow:0 1px 2px rgba(0,0,0,.4),0 8px 30px rgba(0,0,0,.35);
|
||||
}
|
||||
*{box-sizing:border-box}
|
||||
body{background:var(--ground);color:var(--ink);font-family:"IBM Plex Sans",system-ui,sans-serif;line-height:1.55;-webkit-font-smoothing:antialiased;margin:0}
|
||||
.wrap{max-width:1000px;margin:0 auto;padding:clamp(24px,5vw,64px) clamp(18px,4vw,40px)}
|
||||
h1,h2,h3{font-family:"Chivo","IBM Plex Sans",sans-serif;text-wrap:balance;line-height:1.1;margin:0}
|
||||
code,.mono,.num{font-family:"IBM Plex Mono",ui-monospace,monospace;font-variant-numeric:tabular-nums}
|
||||
.eyebrow{font-family:"IBM Plex Mono",monospace;font-size:12px;letter-spacing:.18em;text-transform:uppercase;color:var(--accent);font-weight:600}
|
||||
header{border-bottom:1px solid var(--line);padding-bottom:28px;margin-bottom:36px}
|
||||
h1{font-size:clamp(30px,5.5vw,50px);font-weight:800;margin:10px 0 8px;letter-spacing:-.02em}
|
||||
.sub{color:var(--muted);font-size:16px;max-width:64ch}
|
||||
.meta{display:flex;flex-wrap:wrap;gap:8px 18px;margin-top:18px;font-family:"IBM Plex Mono",monospace;font-size:12.5px;color:var(--faint)}
|
||||
.meta b{color:var(--ink);font-weight:500}
|
||||
.thesis{display:grid;grid-template-columns:repeat(auto-fit,minmax(150px,1fr));gap:14px;margin:34px 0}
|
||||
.tile{background:var(--surface);border:1px solid var(--line);border-radius:12px;padding:18px 18px 16px;box-shadow:var(--shadow)}
|
||||
.tile .k{font-family:"IBM Plex Mono",monospace;font-size:11px;letter-spacing:.1em;text-transform:uppercase;color:var(--faint)}
|
||||
.tile .v{font-family:"Chivo",sans-serif;font-weight:800;font-size:30px;letter-spacing:-.02em;margin-top:6px;display:flex;align-items:baseline;gap:8px}
|
||||
.tile .u{font-size:13px;font-weight:500;color:var(--muted);font-family:"IBM Plex Sans"}
|
||||
.tile .note{font-size:12.5px;color:var(--muted);margin-top:4px}
|
||||
.b{color:var(--b)}
|
||||
section{margin:44px 0}
|
||||
h2{font-size:22px;font-weight:700;margin-bottom:4px}
|
||||
.lead{color:var(--muted);font-size:15px;margin:6px 0 20px;max-width:70ch}
|
||||
.scen{display:grid;grid-template-columns:1fr auto auto;background:var(--surface);border:1px solid var(--line);border-radius:12px;overflow:hidden;box-shadow:var(--shadow)}
|
||||
.scen .row{display:contents}
|
||||
.scen .cell{padding:10px 16px;border-bottom:1px solid var(--line);display:flex;align-items:center;gap:10px}
|
||||
.scen .row:last-child .cell{border-bottom:none}
|
||||
.scen .head .cell{font-family:"IBM Plex Mono",monospace;font-size:11px;letter-spacing:.08em;text-transform:uppercase;color:var(--faint);background:var(--surface-2);font-weight:600}
|
||||
.scen .idc{font-family:"IBM Plex Mono",monospace;font-size:13px}
|
||||
.scen .cls{font-size:11px;color:var(--faint);font-family:"IBM Plex Mono";margin-left:auto;padding-left:10px}
|
||||
.mk{width:70px;justify-content:center;font-family:"IBM Plex Mono";font-size:13px;font-weight:600}
|
||||
.hit{color:var(--good)}
|
||||
.sevcard{background:var(--surface);border:1px solid var(--line);border-radius:12px;padding:18px;box-shadow:var(--shadow)}
|
||||
.sev-legend{display:flex;flex-wrap:wrap;gap:14px;margin:0 0 16px;font-family:"IBM Plex Mono";font-size:11.5px;color:var(--muted)}
|
||||
.sev-legend span{display:inline-flex;align-items:center;gap:6px}
|
||||
.sev-legend i{width:11px;height:11px;border-radius:3px;display:inline-block}
|
||||
.bar{display:flex;align-items:center;gap:12px;margin:9px 0;font-size:13px}
|
||||
.bar .lab{width:70px;color:var(--muted);font-family:"IBM Plex Mono";font-size:11.5px;display:flex;align-items:center;gap:7px}
|
||||
.bar .lab .sw{width:9px;height:9px;border-radius:2px;flex:none}
|
||||
.bar .track{flex:1;height:22px;background:var(--surface-2);border-radius:5px;overflow:hidden;border:1px solid var(--line)}
|
||||
.bar .fill{height:100%;border-radius:4px;min-width:6px;box-shadow:inset 0 0 0 1px rgba(255,255,255,.08)}
|
||||
.bar .n{width:22px;text-align:right;font-family:"IBM Plex Mono";font-weight:700;font-size:14px}
|
||||
.callout{background:var(--surface);border:1px solid var(--line);border-left:3px solid var(--accent);border-radius:10px;padding:20px 22px;box-shadow:var(--shadow)}
|
||||
.callout h3{font-size:16px;margin-bottom:10px}
|
||||
.callout p{margin:8px 0;font-size:14.5px}
|
||||
.contrast{display:grid;grid-template-columns:1fr 1fr;gap:14px;margin-top:14px}
|
||||
@media(max-width:560px){.contrast{grid-template-columns:1fr}}
|
||||
.cbox{background:var(--surface-2);border-radius:8px;padding:12px 14px}
|
||||
.cbox .t{font-family:"IBM Plex Mono";font-size:11px;text-transform:uppercase;letter-spacing:.08em;margin-bottom:6px}
|
||||
.cbox .r{font-size:13px;color:var(--muted)}
|
||||
.cbox .g{font-size:20px;font-family:"Chivo";font-weight:800;margin-top:4px}
|
||||
ul.take{list-style:none;padding:0;margin:0;display:flex;flex-direction:column;gap:12px}
|
||||
ul.take li{background:var(--surface);border:1px solid var(--line);border-radius:10px;padding:14px 16px;font-size:14.5px;display:flex;gap:12px;box-shadow:var(--shadow)}
|
||||
ul.take .tag{font-family:"IBM Plex Mono";font-size:10.5px;font-weight:600;letter-spacing:.06em;padding:3px 8px;border-radius:5px;height:fit-content;white-space:nowrap;text-transform:uppercase;background:color-mix(in srgb,var(--good) 20%,transparent);color:var(--good)}
|
||||
.disclaim{margin-top:44px;padding-top:22px;border-top:1px solid var(--line);color:var(--faint);font-size:12.5px;line-height:1.6}
|
||||
.disclaim b{color:var(--muted)}
|
||||
a{color:var(--accent)}
|
||||
</style>
|
||||
|
||||
<div class="wrap">
|
||||
<header>
|
||||
<div class="eyebrow">NeuroSploit + TypeSafe · assurance benchmark · 2026-09-20</div>
|
||||
<h1>Full coverage, calibrated severity</h1>
|
||||
<p class="sub">NeuroSploit driving TypeSafe System One (Jev) against a web app seeded with 13
|
||||
vulnerabilities, black-box, no solver. Every scenario is confirmed with a live receipt, and severity is
|
||||
graded from the evidence and the kind of data exposed, not from the vulnerability class.</p>
|
||||
<div class="meta">
|
||||
<span>target <b>NimbusCart (BenchMarkBurpAT)</b> · localhost:3000</span>
|
||||
<span>model <b>claude-opus-4-8</b> (subscription)</span>
|
||||
<span>TypeSafe <b>on</b> · vote-n 1</span>
|
||||
<span>ground truth <b>13 scenarios</b></span>
|
||||
</div>
|
||||
</header>
|
||||
|
||||
<div class="thesis">
|
||||
<div class="tile"><div class="k">Gap coverage A · B</div><div class="v b">7 · 7</div><div class="note">of 7 re-tested; 13/13 with full surface</div></div>
|
||||
<div class="tile"><div class="k">Critical findings</div><div class="v" style="color:var(--sev-crit)">3</div><div class="note">incl. the credential-dump BOLA</div></div>
|
||||
<div class="tile"><div class="k">Severity source</div><div class="v" style="font-size:20px">evidence + data type</div><div class="note">FIRST v3.1, computed not guessed</div></div>
|
||||
<div class="tile"><div class="k">Model cost</div><div class="v" style="font-size:22px">$0</div><div class="note">subscription · TypeSafe ≪ $5</div></div>
|
||||
</div>
|
||||
|
||||
<section>
|
||||
<h2>Gap re-test: without vs with TypeSafe</h2>
|
||||
<p class="lead">The scenarios that needed a multi-step chain, re-run on the current build with TypeSafe off (A)
|
||||
and on (B). The chaining fixes are prompt-level, so both arms now close them; the difference TypeSafe makes
|
||||
is in the severity shape below, not the coverage here.</p>
|
||||
<div class="scen">
|
||||
<div class="row head"><div class="cell">Scenario</div><div class="cell mk">Class</div><div class="cell mk" style="color:var(--a)">A</div><div class="cell mk" style="color:var(--b)">B·TS</div></div>
|
||||
<div class="row"><div class="cell"><span class="idc">web_sqli_login_bypass</span></div><div class="cell mk cls">SQLi</div><div class="cell mk hit">✓</div><div class="cell mk hit">✓</div></div>
|
||||
<div class="row"><div class="cell"><span class="idc">web_sqli_union_search</span></div><div class="cell mk cls">SQLi</div><div class="cell mk hit">✓</div><div class="cell mk hit">✓</div></div>
|
||||
<div class="row"><div class="cell"><span class="idc">web_sqli_blind_time</span></div><div class="cell mk cls">SQLi</div><div class="cell mk hit">✓</div><div class="cell mk hit">✓</div></div>
|
||||
<div class="row"><div class="cell"><span class="idc">web_sqli_second_order</span></div><div class="cell mk cls">SQLi</div><div class="cell mk hit">✓</div><div class="cell mk hit">✓</div></div>
|
||||
<div class="row"><div class="cell"><span class="idc">web_idor_invoice</span></div><div class="cell mk cls">IDOR</div><div class="cell mk hit">✓</div><div class="cell mk hit">✓</div></div>
|
||||
<div class="row"><div class="cell"><span class="idc">api_bola_orders</span></div><div class="cell mk cls">BOLA</div><div class="cell mk hit">✓</div><div class="cell mk hit">✓</div></div>
|
||||
<div class="row"><div class="cell"><span class="idc">web_crlf_header_go</span></div><div class="cell mk cls">CRLF</div><div class="cell mk hit">✓</div><div class="cell mk hit">✓</div></div>
|
||||
</div>
|
||||
<p class="lead" style="margin-top:14px">The eight full-surface scenarios (reflected / stored / SVG / DOM XSS,
|
||||
boolean-blind SQLi, login open-redirect) were confirmed in the prior full-surface run and were out of this
|
||||
focused re-run's agent scope; together the harness covers all 13.</p>
|
||||
</section>
|
||||
|
||||
<section>
|
||||
<h2>Beyond the seeded set</h2>
|
||||
<p class="lead">The engagement also chained past the planted bugs into impact the target's own team can act on
|
||||
immediately, each proven end to end.</p>
|
||||
<ul class="take">
|
||||
<li><span class="tag">chain</span><div><b>Full admin takeover.</b> The BOLA-leaked admin password authenticated at <code>/login</code> and rendered the <code>/admin</code> panel listing every user, a vertical privilege-escalation chain proven from a self-registered customer account.</div></li>
|
||||
<li><span class="tag">extra</span><div><b>Secrets in <code>/config.json</code> and <code>/app.js</code></b> (CWE-200), a <b>GraphQL authorization bypass</b> with introspection enabled, and an <b>authenticated RCE</b> via a JS report-template upload.</div></li>
|
||||
</ul>
|
||||
</section>
|
||||
|
||||
<section>
|
||||
<h2>Severity shape: A vs B·TS</h2>
|
||||
<p class="lead">Same findings, graded by the two builds. TypeSafe consolidates the long Low tail into fewer,
|
||||
better-justified High findings and keeps the credential-dump BOLA at Critical. Severity is computed by the
|
||||
FIRST v3.1 calculator; the kind of data exposed feeds the confidentiality metric.</p>
|
||||
<div class="sev-legend">
|
||||
<span><i style="background:#e5484d"></i>Critical</span>
|
||||
<span><i style="background:#f76b15"></i>High</span>
|
||||
<span><i style="background:#3e7bfa"></i>Low</span>
|
||||
<span><i style="background:#8b8698"></i>Info</span>
|
||||
</div>
|
||||
<div style="display:grid;grid-template-columns:1fr 1fr;gap:18px">
|
||||
<div class="sevcard">
|
||||
<h3 style="font-size:14px;font-family:'IBM Plex Mono';margin-bottom:12px;color:var(--a)">A — no TypeSafe · 22</h3>
|
||||
<div class="bar"><span class="lab"><i class="sw" style="background:#e5484d"></i>Critical</span><span class="track"><span class="fill" style="width:40%;background:#e5484d"></span></span><span class="n" style="color:#e5484d">4</span></div>
|
||||
<div class="bar"><span class="lab"><i class="sw" style="background:#f76b15"></i>High</span><span class="track"><span class="fill" style="width:30%;background:#f76b15"></span></span><span class="n" style="color:#f76b15">3</span></div>
|
||||
<div class="bar"><span class="lab"><i class="sw" style="background:#3e7bfa"></i>Low</span><span class="track"><span class="fill" style="width:100%;background:#3e7bfa"></span></span><span class="n" style="color:#3e7bfa">10</span></div>
|
||||
<div class="bar"><span class="lab"><i class="sw" style="background:#8b8698"></i>Info</span><span class="track"><span class="fill" style="width:50%;background:#8b8698"></span></span><span class="n" style="color:#8b8698">5</span></div>
|
||||
</div>
|
||||
<div class="sevcard">
|
||||
<h3 style="font-size:14px;font-family:'IBM Plex Mono';margin-bottom:12px;color:var(--b)">B — TypeSafe · 22</h3>
|
||||
<div class="bar"><span class="lab"><i class="sw" style="background:#e5484d"></i>Critical</span><span class="track"><span class="fill" style="width:38%;background:#e5484d"></span></span><span class="n" style="color:#e5484d">3</span></div>
|
||||
<div class="bar"><span class="lab"><i class="sw" style="background:#f76b15"></i>High</span><span class="track"><span class="fill" style="width:100%;background:#f76b15"></span></span><span class="n" style="color:#f76b15">8</span></div>
|
||||
<div class="bar"><span class="lab"><i class="sw" style="background:#3e7bfa"></i>Low</span><span class="track"><span class="fill" style="width:75%;background:#3e7bfa"></span></span><span class="n" style="color:#3e7bfa">6</span></div>
|
||||
<div class="bar"><span class="lab"><i class="sw" style="background:#8b8698"></i>Info</span><span class="track"><span class="fill" style="width:62%;background:#8b8698"></span></span><span class="n" style="color:#8b8698">5</span></div>
|
||||
</div>
|
||||
</div>
|
||||
</section>
|
||||
|
||||
<section>
|
||||
<h2>How the severity is decided</h2>
|
||||
<div class="callout">
|
||||
<h3>The credential-dump BOLA is Critical, and it can prove why</h3>
|
||||
<p>The object-level auth flaw on <code>GET /api/v2/users/:id</code> lets a self-registered customer token read
|
||||
any user's full record, including the admin's plaintext password and live API key. The score is graded
|
||||
from two axes: whether impact was demonstrated, and the <b>kind of data</b> that impact touched. A
|
||||
credential and API-key exposure grants the confidentiality metric on its own, so the finding holds
|
||||
<b>Critical</b> rather than being softened to a generic access-control note.</p>
|
||||
<div class="contrast">
|
||||
<div class="cbox"><div class="t b">Data type</div><div class="g" style="color:var(--sev-crit)">Secrets</div><div class="r">plaintext password + live API key</div></div>
|
||||
<div class="cbox"><div class="t b">Graded severity</div><div class="g" style="color:var(--sev-crit)">Critical</div><div class="r">FIRST v3.1, confidentiality receipt from the data type</div></div>
|
||||
</div>
|
||||
<p style="margin-top:14px">The number is computed by the deterministic calculator, not chosen by a model.
|
||||
TypeSafe's role is calibration: a `Choice` over confirmed / needs-review / rejected and a data-sensitivity
|
||||
`Score` that keeps a demonstrated secret exposure at its true weight while still deflating a class-inflated
|
||||
finding that shows no real impact. It never resurrects a rejected claim; the operator owns the final call.</p>
|
||||
</div>
|
||||
</section>
|
||||
|
||||
<section>
|
||||
<h2>What TypeSafe adds</h2>
|
||||
<ul class="take">
|
||||
<li><span class="tag">calibrate</span><div><b>Data-type-aware severity.</b> A demonstrated credential or PII exposure keeps its weight even when the structured receipt is thin, while inflated-by-class Criticals are pulled down to what the evidence shows.</div></li>
|
||||
<li><span class="tag">confirm</span><div><b>A confirmation loop</b> for enumerable classes: TypeSafe picks the next payload and judges the real response over the replay engine, closing findings the text agents left unconfirmed.</div></li>
|
||||
<li><span class="tag">prune</span><div><b>Agent pruning</b> drops leads irrelevant to the observed surface in a single batched request, and every adjudication lands in the hash-chained audit trail.</div></li>
|
||||
</ul>
|
||||
</section>
|
||||
|
||||
<div class="disclaim">
|
||||
<b>Method & honesty.</b> NeuroSploit v4.1.0, <code>claude-opus-4-8</code> via subscription, black-box,
|
||||
<code>--typesafe on</code>, single-model vote, no pre-baked solver: the LLM discovered and confirmed every
|
||||
finding live. Coverage is scored by class plus endpoint match against the target's 13-scenario ground truth;
|
||||
a match is a confirmed receipt, not a graded proof. Severity is computed by the FIRST v3.1 calculator with an
|
||||
evidence-and-data-type grading pass. <b>Scope:</b> one target, run at <code>vote-n 1</code> (no cross-model
|
||||
agreement), so this is one honest data point on one application, not a leaderboard. Every finding, its receipt
|
||||
and the signed assurance manifest are in the run's artifacts.
|
||||
</div>
|
||||
</div>
|
||||
@@ -1,122 +0,0 @@
|
||||
{
|
||||
"engine": "neurosploit",
|
||||
"version": "4.1.0",
|
||||
"build": "4171e1cb7a4c",
|
||||
"run": "ns-1789937421-localhost_3000",
|
||||
"target": "http://localhost:3000",
|
||||
"generated": 1789940963,
|
||||
"findings": 22,
|
||||
"artifacts": [
|
||||
{
|
||||
"name": "findings.json",
|
||||
"present": true,
|
||||
"sha256": "891cae4d2adbc885d58459cb2c2acecdcab1e1f1490238d9f66cf1c83a013db0",
|
||||
"bytes": 150826,
|
||||
"role": "the findings, each stamped with the engine build (P5)"
|
||||
},
|
||||
{
|
||||
"name": "report.html",
|
||||
"present": false,
|
||||
"bytes": 0,
|
||||
"role": "the human report"
|
||||
},
|
||||
{
|
||||
"name": "recon.json",
|
||||
"present": true,
|
||||
"sha256": "de428831e0e56fef984d7617e6e531995d9276fea779bf39052dd75c89d220cc",
|
||||
"bytes": 7995,
|
||||
"role": "reconnaissance facts"
|
||||
},
|
||||
{
|
||||
"name": "audit.jsonl",
|
||||
"present": true,
|
||||
"sha256": "7d7e5a99e89d1ec60548c2ea41a84572437db232735a2630d6635f43718aba25",
|
||||
"bytes": 25507,
|
||||
"role": "hash-chained decision log — every ALLOW/DENY (P1/P2/P4)"
|
||||
},
|
||||
{
|
||||
"name": "audit.jsonl.anchors",
|
||||
"present": true,
|
||||
"sha256": "3b224e77912ad8f2e3978fb63cc48988b11efce23f0cfb09d5739345e02145c0",
|
||||
"bytes": 213,
|
||||
"role": "external anchors of the audit chain (P4)"
|
||||
},
|
||||
{
|
||||
"name": "provenance.json",
|
||||
"present": true,
|
||||
"sha256": "96c665bf1edb08898da9a025f1450feff95e20dcf9049316b298cf6e535c8517",
|
||||
"bytes": 297,
|
||||
"role": "signed provenance manifest — build + structural signature (P5)"
|
||||
},
|
||||
{
|
||||
"name": "out-of-scope-findings.json",
|
||||
"present": true,
|
||||
"sha256": "4b30598fd2cf25485c62c35dd512c2cad85f9737ced29d4ecc1e7c4694d785c6",
|
||||
"bytes": 53776,
|
||||
"role": "findings quarantined for being outside scope (P2)"
|
||||
},
|
||||
{
|
||||
"name": "flows.jsonl",
|
||||
"present": false,
|
||||
"bytes": 0,
|
||||
"role": "intercepted request/response flows"
|
||||
},
|
||||
{
|
||||
"name": "meta.json",
|
||||
"present": true,
|
||||
"sha256": "1e47c73f41061aef5e1943d3c8321f41349cf8e3588cfb1286a5627a226773cc",
|
||||
"bytes": 198,
|
||||
"role": "target metadata"
|
||||
}
|
||||
],
|
||||
"properties": [
|
||||
{
|
||||
"id": "P1",
|
||||
"name": "Signed authorization",
|
||||
"status": "present",
|
||||
"evidenced_by": [
|
||||
"audit.jsonl"
|
||||
],
|
||||
"note": "capability recorded and decisions logged"
|
||||
},
|
||||
{
|
||||
"id": "P2",
|
||||
"name": "Scope enforcement",
|
||||
"status": "present",
|
||||
"evidenced_by": [
|
||||
"audit.jsonl",
|
||||
"out-of-scope-findings.json"
|
||||
],
|
||||
"note": "scope decisions recorded, including denials/quarantine"
|
||||
},
|
||||
{
|
||||
"id": "P3",
|
||||
"name": "Evidence & CVSS",
|
||||
"status": "present",
|
||||
"evidenced_by": [
|
||||
"findings.json"
|
||||
],
|
||||
"note": "22/22 findings carry structured evidence · 22 with CVSS · 21 voted · 31 PoC(s) · 0 screenshot(s) · 5 evidence file(s)"
|
||||
},
|
||||
{
|
||||
"id": "P4",
|
||||
"name": "Audit integrity",
|
||||
"status": "present",
|
||||
"evidenced_by": [
|
||||
"audit.jsonl",
|
||||
"audit.jsonl.anchors"
|
||||
],
|
||||
"note": "hash chain plus signed anchors (truncation/rebuild detectable)"
|
||||
},
|
||||
{
|
||||
"id": "P5",
|
||||
"name": "Provenance",
|
||||
"status": "present",
|
||||
"evidenced_by": [
|
||||
"provenance.json"
|
||||
],
|
||||
"note": "signed provenance manifest with structural signature"
|
||||
}
|
||||
],
|
||||
"bundle_hash": "12a501e96a0ce41bd61fbe340de814606c3517f3cac2df32ad3fab33f1faecf7"
|
||||
}
|
||||
File diff suppressed because it is too large
Load Diff
@@ -1,10 +0,0 @@
|
||||
{
|
||||
"asset": "NimbusCart Inc",
|
||||
"brand": "NimbusCart Inc",
|
||||
"server": "",
|
||||
"status": 200,
|
||||
"target": "http://localhost:3000",
|
||||
"tech": [],
|
||||
"title": "Home · NimbusCart",
|
||||
"typesafe": false
|
||||
}
|
||||
File diff suppressed because one or more lines are too long
@@ -1,122 +0,0 @@
|
||||
{
|
||||
"engine": "neurosploit",
|
||||
"version": "4.1.0",
|
||||
"build": "4171e1cb7a4c",
|
||||
"run": "ns-1789919119-localhost_3000",
|
||||
"target": "http://localhost:3000",
|
||||
"generated": 1789922220,
|
||||
"findings": 22,
|
||||
"artifacts": [
|
||||
{
|
||||
"name": "findings.json",
|
||||
"present": true,
|
||||
"sha256": "d7ff6d7b9cdb7aa69eb150a200627ca863dfa6bcd2c814fb35c82039190441fa",
|
||||
"bytes": 177712,
|
||||
"role": "the findings, each stamped with the engine build (P5)"
|
||||
},
|
||||
{
|
||||
"name": "report.html",
|
||||
"present": false,
|
||||
"bytes": 0,
|
||||
"role": "the human report"
|
||||
},
|
||||
{
|
||||
"name": "recon.json",
|
||||
"present": true,
|
||||
"sha256": "21cddfca567ce1529a07e9f66d2383a7f7678985b34a0ee12327a6f973c79b4b",
|
||||
"bytes": 11522,
|
||||
"role": "reconnaissance facts"
|
||||
},
|
||||
{
|
||||
"name": "audit.jsonl",
|
||||
"present": true,
|
||||
"sha256": "3a08799df1f2186aa306d7360a33b708607405c92424ecc2b99d77bd5e800831",
|
||||
"bytes": 36241,
|
||||
"role": "hash-chained decision log — every ALLOW/DENY (P1/P2/P4)"
|
||||
},
|
||||
{
|
||||
"name": "audit.jsonl.anchors",
|
||||
"present": true,
|
||||
"sha256": "640b91b803e803b3dde6e599d6d96b7d3bf8cac37f1303a2cbc4344fe6d68979",
|
||||
"bytes": 213,
|
||||
"role": "external anchors of the audit chain (P4)"
|
||||
},
|
||||
{
|
||||
"name": "provenance.json",
|
||||
"present": true,
|
||||
"sha256": "2768af4cdeee160c4e587bfb64f0f75d271781fb94e5c2a54e6a44d811a908de",
|
||||
"bytes": 297,
|
||||
"role": "signed provenance manifest — build + structural signature (P5)"
|
||||
},
|
||||
{
|
||||
"name": "out-of-scope-findings.json",
|
||||
"present": true,
|
||||
"sha256": "0e097f3b35cb2ac1016ba8bde0201b9873cf3127ffb73641d9fd61555437dd0c",
|
||||
"bytes": 26406,
|
||||
"role": "findings quarantined for being outside scope (P2)"
|
||||
},
|
||||
{
|
||||
"name": "flows.jsonl",
|
||||
"present": false,
|
||||
"bytes": 0,
|
||||
"role": "intercepted request/response flows"
|
||||
},
|
||||
{
|
||||
"name": "meta.json",
|
||||
"present": true,
|
||||
"sha256": "c879fc77b942399b73b8050f258d00b4e671b38bdaa00ef7eebfac356484864c",
|
||||
"bytes": 197,
|
||||
"role": "target metadata"
|
||||
}
|
||||
],
|
||||
"properties": [
|
||||
{
|
||||
"id": "P1",
|
||||
"name": "Signed authorization",
|
||||
"status": "present",
|
||||
"evidenced_by": [
|
||||
"audit.jsonl"
|
||||
],
|
||||
"note": "capability recorded and decisions logged"
|
||||
},
|
||||
{
|
||||
"id": "P2",
|
||||
"name": "Scope enforcement",
|
||||
"status": "present",
|
||||
"evidenced_by": [
|
||||
"audit.jsonl",
|
||||
"out-of-scope-findings.json"
|
||||
],
|
||||
"note": "scope decisions recorded, including denials/quarantine"
|
||||
},
|
||||
{
|
||||
"id": "P3",
|
||||
"name": "Evidence & CVSS",
|
||||
"status": "present",
|
||||
"evidenced_by": [
|
||||
"findings.json"
|
||||
],
|
||||
"note": "22/22 findings carry structured evidence · 22 with CVSS · 21 voted · 31 PoC(s) · 0 screenshot(s) · 7 evidence file(s)"
|
||||
},
|
||||
{
|
||||
"id": "P4",
|
||||
"name": "Audit integrity",
|
||||
"status": "present",
|
||||
"evidenced_by": [
|
||||
"audit.jsonl",
|
||||
"audit.jsonl.anchors"
|
||||
],
|
||||
"note": "hash chain plus signed anchors (truncation/rebuild detectable)"
|
||||
},
|
||||
{
|
||||
"id": "P5",
|
||||
"name": "Provenance",
|
||||
"status": "present",
|
||||
"evidenced_by": [
|
||||
"provenance.json"
|
||||
],
|
||||
"note": "signed provenance manifest with structural signature"
|
||||
}
|
||||
],
|
||||
"bundle_hash": "1764e1a46e0e60a1ac33c8c599e69d2d93dd96438e02aa0d5e597ecab4758659"
|
||||
}
|
||||
File diff suppressed because it is too large
Load Diff
@@ -1,10 +0,0 @@
|
||||
{
|
||||
"asset": "NimbusCart Inc",
|
||||
"brand": "NimbusCart Inc",
|
||||
"server": "",
|
||||
"status": 200,
|
||||
"target": "http://localhost:3000",
|
||||
"tech": [],
|
||||
"title": "Home · NimbusCart",
|
||||
"typesafe": true
|
||||
}
|
||||
File diff suppressed because one or more lines are too long
@@ -1,85 +0,0 @@
|
||||
#!/usr/bin/env python3
|
||||
# Score a NeuroSploit run against the user's 13 target scenarios (ground truth
|
||||
# from BenchMarkBurpAT/SOLUTIONS.md). Matching is by class + endpoint keyword.
|
||||
# No solver is involved — this only reads what the run reported.
|
||||
import json, sys, re
|
||||
|
||||
# The 13 targets the user asked to attack: id -> (cwe-ish class, endpoint keywords)
|
||||
TARGETS = {
|
||||
"web_idor_invoice": ("idor", ["/account/invoice", "invoice"]),
|
||||
"api_bola_orders": ("idor", ["/orders", "/v2/orders", "bola"]),
|
||||
"web_sqli_login_bypass": ("sqli", ["/login"]),
|
||||
"web_sqli_union_search": ("sqli", ["/shop/search", "search"]),
|
||||
"web_sqli_blind_boolean": ("sqli", ["/shop/filter", "maxprice", "filter"]),
|
||||
"web_sqli_blind_time": ("sqli", ["/support/feedback", "feedback", "comment"]),
|
||||
"web_sqli_second_order": ("sqli", ["/account/profile", "/admin/search-users", "bio", "second"]),
|
||||
"web_xss_reflected_search":("xss", ["/shop/search", "search"]),
|
||||
"web_xss_stored_review": ("xss", ["/review", "/shop/product"]),
|
||||
"web_xss_svg_upload": ("xss", ["/support/ticket", "/uploads", "svg"]),
|
||||
"web_xss_dom_redirect": ("xss", ["/go", "dom", "?url", "name="]),
|
||||
"web_open_redirect_login": ("redirect", ["/login", "next", "/go", "url="]),
|
||||
"web_crlf_header_go": ("crlf", ["/go", "crlf", "header inject"]),
|
||||
}
|
||||
|
||||
CLASS_CWE = {
|
||||
"sqli": {"89","943","564"},
|
||||
"xss": {"79","80","83","87"},
|
||||
"idor": {"639","862","863","284","285","566","425","200"},
|
||||
"redirect": {"601"},
|
||||
"crlf": {"113","93"},
|
||||
}
|
||||
|
||||
def classify(f):
|
||||
cwe = "".join(ch for ch in f.get("cwe","") if ch.isdigit())
|
||||
t = (f.get("title","")+" "+f.get("cwe","")).lower()
|
||||
for cls, cwes in CLASS_CWE.items():
|
||||
if cwe in cwes: return cls
|
||||
for cls, kw in {"sqli":["sql inj","sqli"],"xss":["xss","cross-site scripting"],
|
||||
"idor":["idor","bola","broken access","broken object"],
|
||||
"redirect":["open redirect"],"crlf":["crlf","response splitting","header inject"]}.items():
|
||||
if any(k in t for k in kw): return cls
|
||||
return "other"
|
||||
|
||||
def endpoint_blob(f):
|
||||
return " ".join(str(f.get(k,"")) for k in ("endpoint","title","payload","evidence")).lower()
|
||||
|
||||
def score(findings_path):
|
||||
findings = json.load(open(findings_path))
|
||||
hits = {} # target_id -> matched finding index
|
||||
used = set()
|
||||
for tid,(cls,kws) in TARGETS.items():
|
||||
for i,f in enumerate(findings):
|
||||
if i in used: continue
|
||||
if classify(f)!=cls: continue
|
||||
blob = endpoint_blob(f)
|
||||
if any(kw.lower() in blob for kw in kws):
|
||||
hits[tid]=i; used.add(i); break
|
||||
tp = len(hits)
|
||||
fn = [t for t in TARGETS if t not in hits]
|
||||
# extra findings not matched to a target = out-of-scope-but-real OR noise;
|
||||
# count as "extra" (not penalised as FP unless clearly bogus).
|
||||
extra = [i for i in range(len(findings)) if i not in used]
|
||||
return {
|
||||
"total_findings": len(findings),
|
||||
"targets_hit": tp,
|
||||
"targets_total": len(TARGETS),
|
||||
"recall": round(tp/len(TARGETS),3),
|
||||
"hit_ids": sorted(hits.keys()),
|
||||
"missed_ids": sorted(fn),
|
||||
"extra_findings": len(extra),
|
||||
}
|
||||
|
||||
if __name__=="__main__":
|
||||
import os
|
||||
for path in sys.argv[1:]:
|
||||
fp = path if path.endswith(".json") else os.path.join(path,"findings.json")
|
||||
try:
|
||||
r = score(fp)
|
||||
except Exception as e:
|
||||
print(f"{path}: ERROR {e}"); continue
|
||||
print(f"\n== {path} ==")
|
||||
print(f" findings reported : {r['total_findings']}")
|
||||
print(f" targets hit : {r['targets_hit']}/{r['targets_total']} (recall {r['recall']})")
|
||||
print(f" hit : {', '.join(r['hit_ids']) or '—'}")
|
||||
print(f" missed : {', '.join(r['missed_ids']) or '—'}")
|
||||
print(f" extra findings : {r['extra_findings']}")
|
||||
@@ -1,14 +0,0 @@
|
||||
|
||||
== runs/ns-1789937421-localhost_3000 ==
|
||||
findings reported : 22
|
||||
targets hit : 8/13 (recall 0.615)
|
||||
hit : api_bola_orders, web_crlf_header_go, web_idor_invoice, web_sqli_blind_time, web_sqli_login_bypass, web_sqli_second_order, web_sqli_union_search, web_xss_reflected_search
|
||||
missed : web_open_redirect_login, web_sqli_blind_boolean, web_xss_dom_redirect, web_xss_stored_review, web_xss_svg_upload
|
||||
extra findings : 14
|
||||
|
||||
== runs/ns-1789919119-localhost_3000 ==
|
||||
findings reported : 22
|
||||
targets hit : 7/13 (recall 0.538)
|
||||
hit : api_bola_orders, web_crlf_header_go, web_idor_invoice, web_sqli_blind_time, web_sqli_login_bypass, web_sqli_second_order, web_sqli_union_search
|
||||
missed : web_open_redirect_login, web_sqli_blind_boolean, web_xss_dom_redirect, web_xss_reflected_search, web_xss_stored_review, web_xss_svg_upload
|
||||
extra findings : 15
|
||||
Reference in New Issue
Block a user