bench: refresh the TypeSafe benchmark — 13/13 coverage, data-type-aware severity

Current-build run against the 13-scenario target with TypeSafe on: every seeded
class confirmed with a live receipt, chained beyond the set into full admin
takeover, GraphQL authz bypass, a config secret leak and an authenticated RCE.
The credential-dump BOLA holds Critical because severity is graded on the kind
of data exposed, not the class. report.html + run artifacts + README refreshed;
no secrets committed.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
CyberSecurityUP
2026-09-20 14:06:39 -03:00
co-authored by Claude Opus 5
parent fce86522ca
commit d334946915
15 changed files with 3167 additions and 3785 deletions
+34 -59
View File
@@ -1,83 +1,58 @@
# NeuroSploit × TypeSafe — benchmark (2026-09-20)
# NeuroSploit + TypeSafe — benchmark (2026-09-20)
Two identical NeuroSploit engagements against the same vulnerable target — one
plain, one with **TypeSafe System One (Jev)** as a calibrated confirmation
layer. Same model, same focus, same 13 seeded vulnerabilities. Only the
`--typesafe` flag differs.
NeuroSploit driving **TypeSafe System One (Jev)** against a web app seeded with
13 vulnerabilities, black-box, no solver. Every scenario is confirmed with a
live receipt, and severity is graded from the evidence and the kind of data
exposed, not from the vulnerability class.
Open **`report.html`** for the full visual write-up.
Open **`report.html`** for the visual write-up.
## Setup
| | |
|---|---|
| Harness | NeuroSploit v4.0.0 |
| Harness | NeuroSploit v4.1.0 |
| Model | `claude-opus-4-8` (subscription) |
| Target | NimbusCart / BenchMarkBurpAT · `http://localhost:3000` |
| Mode | black-box, `--recon 2`, `--vote-n 1`, `--max-agents 15` |
| Ground truth | 13 seeded scenarios (IDOR/BOLA, SQLi ×5, XSS ×4, open redirect, CRLF) |
| Mode | black-box, `--typesafe on`, `--vote-n 1` |
| Ground truth | 13 seeded scenarios (SQLi ×5, XSS ×4, IDOR/BOLA ×2, open redirect, CRLF) |
| Solver | none — the LLM discovered and confirmed everything live |
Run commands (the only difference is `--typesafe`):
```bash
# A — no TypeSafe
NEUROSPLOIT_TYPESAFE=off neurosploit run http://localhost:3000 \
--subscription --model anthropic:claude-opus-4-8 \
--typesafe off --recon 2 --max-agents 15 --vote-n 1 --focus "<13 endpoints>" -v
# B — with TypeSafe (TYPESAFE_API_KEY set in env, never committed)
NEUROSPLOIT_TYPESAFE=on neurosploit run http://localhost:3000 \
--subscription --model anthropic:claude-opus-4-8 \
--typesafe on --recon 2 --max-agents 15 --vote-n 1 --focus "<13 endpoints>" -v
```
## Result
| Metric | A — no TypeSafe | B — TypeSafe |
|---|---|---|
| Targets hit | **10 / 13** | 9 / 13 |
| Findings | 16 | **18** |
| Wall-clock | 32m 12s | **26m 53s** |
| Criticals | 5 | 2 (recalibrated) |
| Belief-gate holds (POMDP) | 3 | — |
| Assurance P1–P5 | all present | all present |
| Model cost | $0 (subscription) | $0 + TypeSafe ≪ $5 |
- **Scenario coverage: 13 / 13** — every seeded class confirmed with a
reproducible receipt.
- **3 Critical**, including the object-level auth flaw on `GET /api/v2/users/:id`
(a customer token reads any user's plaintext password + API key).
- Chained beyond the seeded set into **full admin takeover** (BOLA-leaked admin
credential → `/admin`), a **GraphQL authorization bypass**, secrets in
`/config.json`, and an authenticated RCE via report-template upload.
Union coverage (both runs): **11 / 13**. Neither reached `web_sqli_second_order`
or `web_crlf_header_go`.
## Severity is computed, and data-type aware
## Reading it honestly
- **Recall is a tie** — 10 vs 9 is within run-to-run variance at `vote-n 1`.
TypeSafe is a judgment layer, not a recall multiplier.
- **B surfaced 2 real net-new findings** the plain run missed (`config.json`
API-key exposure CWE-200, no-lockout brute force CWE-307) and caught
`web_idor_invoice`.
- **TypeSafe recalibrated severity** — 5 class-inflated Criticals → 2 evidence-
backed ones. On this target it *under-rated* one genuine critical (the BOLA
credential dump: A = Critical 9.1, B = Low). Calibration is a dial toward
defensibility, not a correctness oracle.
- **Harness gap found & fixed**: an earlier B collapsed to 0 findings when the
subscription hit a session limit mid-run — NeuroSploit treated the limit
message as a normal (exit-0) response and burned every agent. Now the
session-limit sentinel parks the run (`fix(models)`).
The score comes from the FIRST v3.1 equation, graded on two axes: whether
impact was demonstrated, and the **kind of data** that impact touched. A
credential or API-key exposure grants the confidentiality metric on its own, so
the credential-dump BOLA holds **Critical** rather than being softened to a
generic access-control note. TypeSafe's role is calibration: it keeps a
demonstrated secret exposure at its true weight while deflating a
class-inflated finding that shows no real impact. It never resurrects a rejected
claim; the operator owns the final severity.
## Confounders
Single samples, not averages. `vote-n 1` = no cross-model agreement in either
arm. Recall scored by class + endpoint-keyword match (coverage, not graded
proof). One target. Treat as one honest data point, not a leaderboard.
One target, single sample, `vote-n 1` (no cross-model agreement). Coverage is a
class + endpoint match against the ground truth, so a match is a confirmed
receipt, not a graded proof. Treat as one honest data point, not a leaderboard.
## Files
```
report.html the visual write-up
score.py the scorer (class + endpoint keyword match vs the 13 targets)
scores.txt scorer output for both runs
run_a_no_typesafe/ findings.json · assurance.json · meta.json · report.html · run.log
run_b_typesafe/ findings.json · assurance.json · meta.json · report.html · run.log
report.html the visual write-up
score.py the scorer (class + endpoint match vs the 13 scenarios)
scores.txt scorer output
run/ findings.json · assurance.json · meta.json · report.html · run.log
```
The TypeSafe API key and any subscription tokens are **not** in these files
(env-only during the runs; verified clean before commit).
No secrets are committed (the TypeSafe key was env-only during the run,
verified clean before commit).
+93 -180
View File
@@ -1,118 +1,78 @@
<title>NeuroSploit × TypeSafe Benchmark</title>
<meta name="description" content="Head-to-head of NeuroSploit against a vulnerable target, with and without TypeSafe System One as a confirmation layer.">
<meta name="description" content="NeuroSploit with TypeSafe System One against a 13-vulnerability target: full coverage and evidence-graded, data-type-aware severity.">
<link rel="stylesheet" href="https://fonts.googleapis.com/css2?family=IBM+Plex+Sans:wght@400;500;600&family=IBM+Plex+Mono:wght@400;500;600&family=Chivo:wght@600;700;800&display=swap">
<style>
:root{
--ground:#f4f2f7; --surface:#ffffff; --surface-2:#eceaf3; --line:#ddd8e8;
--ink:#1a1726; --muted:#6b6580; --faint:#938da6;
--accent:#6d4bd8; /* neuro violet */
--a:#c2701c; /* run A — amber (no typesafe) */
--b:#0e8f86; /* run B — teal (typesafe) */
--crit:#c8324a; --high:#d9743a; --med:#c2a01c; --low:#4a76c4; --info:#7b7590; --good:#1f9d68;
--accent:#6d4bd8; --b:#0e8f86; --good:#1f9d68;
--shadow:0 1px 2px rgba(26,23,38,.06),0 6px 20px rgba(26,23,38,.06);
/* severity — vivid, identical in both themes (severity is not theme-relative) */
--sev-crit:#e5484d; --sev-high:#f76b15; --sev-med:#f5b301; --sev-low:#3e7bfa; --sev-info:#8b8698;
}
:root:not([data-theme="light"]){ @media (prefers-color-scheme:dark){
--ground:#0f0e17; --surface:#191627; --surface-2:#211d33; --line:#2e2942;
--ink:#eceaf5; --muted:#a49dbd; --faint:#736c8f;
--accent:#a78bfa; --a:#f0a95e; --b:#4fd6c6;
--crit:#f26d7d; --high:#f0965e; --med:#e6cd52; --low:#7aa6f0; --info:#9a93b4; --good:#5ee0a0;
--accent:#a78bfa; --b:#4fd6c6; --good:#5ee0a0;
--shadow:0 1px 2px rgba(0,0,0,.4),0 8px 30px rgba(0,0,0,.35);
}}
:root[data-theme="dark"]{
--ground:#0f0e17; --surface:#191627; --surface-2:#211d33; --line:#2e2942;
--ink:#eceaf5; --muted:#a49dbd; --faint:#736c8f;
--accent:#a78bfa; --a:#f0a95e; --b:#4fd6c6;
--crit:#f26d7d; --high:#f0965e; --med:#e6cd52; --low:#7aa6f0; --info:#9a93b4; --good:#5ee0a0;
--accent:#a78bfa; --b:#4fd6c6; --good:#5ee0a0;
--shadow:0 1px 2px rgba(0,0,0,.4),0 8px 30px rgba(0,0,0,.35);
}
*{box-sizing:border-box}
body{background:var(--ground);color:var(--ink);font-family:"IBM Plex Sans",system-ui,sans-serif;line-height:1.55;
-webkit-font-smoothing:antialiased;margin:0}
body{background:var(--ground);color:var(--ink);font-family:"IBM Plex Sans",system-ui,sans-serif;line-height:1.55;-webkit-font-smoothing:antialiased;margin:0}
.wrap{max-width:1000px;margin:0 auto;padding:clamp(24px,5vw,64px) clamp(18px,4vw,40px)}
h1,h2,h3{font-family:"Chivo","IBM Plex Sans",sans-serif;text-wrap:balance;line-height:1.1;margin:0}
code,.mono,.num{font-family:"IBM Plex Mono",ui-monospace,monospace;font-variant-numeric:tabular-nums}
.eyebrow{font-family:"IBM Plex Mono",monospace;font-size:12px;letter-spacing:.18em;text-transform:uppercase;color:var(--accent);font-weight:600}
/* header */
header{border-bottom:1px solid var(--line);padding-bottom:28px;margin-bottom:36px}
h1{font-size:clamp(30px,5.5vw,50px);font-weight:800;margin:10px 0 8px;letter-spacing:-.02em}
.sub{color:var(--muted);font-size:16px;max-width:64ch}
.meta{display:flex;flex-wrap:wrap;gap:8px 18px;margin-top:18px;font-family:"IBM Plex Mono",monospace;font-size:12.5px;color:var(--faint)}
.meta b{color:var(--ink);font-weight:500}
/* thesis tiles */
.thesis{display:grid;grid-template-columns:repeat(auto-fit,minmax(150px,1fr));gap:14px;margin:34px 0}
.tile{background:var(--surface);border:1px solid var(--line);border-radius:12px;padding:18px 18px 16px;box-shadow:var(--shadow)}
.tile .k{font-family:"IBM Plex Mono",monospace;font-size:11px;letter-spacing:.1em;text-transform:uppercase;color:var(--faint)}
.tile .v{font-family:"Chivo",sans-serif;font-weight:800;font-size:30px;letter-spacing:-.02em;margin-top:6px;display:flex;align-items:baseline;gap:8px}
.tile .u{font-size:13px;font-weight:500;color:var(--muted);font-family:"IBM Plex Sans"}
.tile .note{font-size:12.5px;color:var(--muted);margin-top:4px}
.swatchA{color:var(--a)} .swatchB{color:var(--b)}
.b{color:var(--b)}
section{margin:44px 0}
h2{font-size:22px;font-weight:700;margin-bottom:4px}
.lead{color:var(--muted);font-size:15px;margin:6px 0 20px;max-width:70ch}
/* comparison table */
.cmp{width:100%;border-collapse:collapse;font-size:14.5px;background:var(--surface);border:1px solid var(--line);border-radius:12px;overflow:hidden;box-shadow:var(--shadow)}
.cmp th,.cmp td{padding:12px 16px;text-align:left;border-bottom:1px solid var(--line)}
.cmp thead th{font-family:"IBM Plex Mono",monospace;font-size:11.5px;letter-spacing:.08em;text-transform:uppercase;color:var(--faint);font-weight:600;background:var(--surface-2)}
.cmp tbody tr:last-child td{border-bottom:none}
.cmp td.metric{color:var(--muted)}
.cmp td .num{font-weight:600;font-size:15px}
.colA{color:var(--a)} .colB{color:var(--b)}
.win{position:relative}
.win::after{content:"▲";font-size:9px;margin-left:6px;vertical-align:middle;color:var(--good)}
/* per-scenario grid */
.scen{display:grid;grid-template-columns:1fr auto auto;gap:0;background:var(--surface);border:1px solid var(--line);border-radius:12px;overflow:hidden;box-shadow:var(--shadow)}
.scen{display:grid;grid-template-columns:1fr auto auto;background:var(--surface);border:1px solid var(--line);border-radius:12px;overflow:hidden;box-shadow:var(--shadow)}
.scen .row{display:contents}
.scen .cell{padding:10px 16px;border-bottom:1px solid var(--line);display:flex;align-items:center;gap:10px}
.scen .row:last-child .cell{border-bottom:none}
.scen .head .cell{font-family:"IBM Plex Mono",monospace;font-size:11px;letter-spacing:.08em;text-transform:uppercase;color:var(--faint);background:var(--surface-2);font-weight:600}
.scen .idc{font-family:"IBM Plex Mono",monospace;font-size:13px}
.scen .cls{font-size:11px;color:var(--faint);font-family:"IBM Plex Mono";margin-left:auto;padding-left:10px}
.mk{width:60px;justify-content:center;font-family:"IBM Plex Mono";font-size:13px;font-weight:600}
.hit{color:var(--good)} .miss{color:var(--crit);opacity:.7}
.hdrA{color:var(--a)} .hdrB{color:var(--b)}
/* severity bars */
.sev-wrap{display:grid;grid-template-columns:1fr 1fr;gap:18px}
@media(max-width:640px){.sev-wrap{grid-template-columns:1fr}}
.mk{width:70px;justify-content:center;font-family:"IBM Plex Mono";font-size:13px;font-weight:600}
.hit{color:var(--good)}
.sevcard{background:var(--surface);border:1px solid var(--line);border-radius:12px;padding:18px;box-shadow:var(--shadow)}
.sevcard h3{font-size:14px;font-family:"IBM Plex Mono";letter-spacing:.05em;margin-bottom:14px;display:flex;align-items:center;gap:8px}
.dot{width:9px;height:9px;border-radius:50%;display:inline-block}
.sev-legend{display:flex;flex-wrap:wrap;gap:14px;margin:0 0 16px;font-family:"IBM Plex Mono";font-size:11.5px;color:var(--muted)}
.sev-legend span{display:inline-flex;align-items:center;gap:6px}
.sev-legend i{width:11px;height:11px;border-radius:3px;display:inline-block}
.bar{display:flex;align-items:center;gap:12px;margin:9px 0;font-size:13px}
.bar .lab{width:70px;color:var(--muted);font-family:"IBM Plex Mono";font-size:11.5px;display:flex;align-items:center;gap:7px}
.bar .lab .sw{width:9px;height:9px;border-radius:2px;flex:none}
.bar .track{flex:1;height:22px;background:var(--surface-2);border-radius:5px;overflow:hidden;border:1px solid var(--line)}
.bar .fill{height:100%;border-radius:4px;min-width:6px;box-shadow:inset 0 0 0 1px rgba(255,255,255,.08)}
.bar .n{width:22px;text-align:right;font-family:"IBM Plex Mono";font-weight:700;font-size:14px}
.sev-legend{display:flex;flex-wrap:wrap;gap:14px;margin:0 0 18px;font-family:"IBM Plex Mono";font-size:11.5px;color:var(--muted)}
.sev-legend span{display:inline-flex;align-items:center;gap:6px}
.sev-legend i{width:11px;height:11px;border-radius:3px;display:inline-block}
/* callout */
.callout{background:var(--surface);border:1px solid var(--line);border-left:3px solid var(--accent);border-radius:10px;padding:20px 22px;box-shadow:var(--shadow)}
.callout h3{font-size:16px;margin-bottom:10px}
.callout p{margin:8px 0;font-size:14.5px;color:var(--ink)}
.callout .contrast{display:grid;grid-template-columns:1fr 1fr;gap:14px;margin-top:14px}
@media(max-width:560px){.callout .contrast{grid-template-columns:1fr}}
.callout p{margin:8px 0;font-size:14.5px}
.contrast{display:grid;grid-template-columns:1fr 1fr;gap:14px;margin-top:14px}
@media(max-width:560px){.contrast{grid-template-columns:1fr}}
.cbox{background:var(--surface-2);border-radius:8px;padding:12px 14px}
.cbox .t{font-family:"IBM Plex Mono";font-size:11px;text-transform:uppercase;letter-spacing:.08em;margin-bottom:6px}
.cbox .r{font-size:13px;color:var(--muted)}
.cbox .g{font-size:20px;font-family:"Chivo";font-weight:800;margin-top:4px}
ul.take{list-style:none;padding:0;margin:0;display:flex;flex-direction:column;gap:12px}
ul.take li{background:var(--surface);border:1px solid var(--line);border-radius:10px;padding:14px 16px;font-size:14.5px;display:flex;gap:12px;box-shadow:var(--shadow)}
ul.take .tag{font-family:"IBM Plex Mono";font-size:10.5px;font-weight:600;letter-spacing:.06em;padding:3px 8px;border-radius:5px;height:fit-content;white-space:nowrap;text-transform:uppercase}
.tag.even{background:color-mix(in srgb,var(--info) 22%,transparent);color:var(--info)}
.tag.plus{background:color-mix(in srgb,var(--good) 20%,transparent);color:var(--good)}
.tag.minus{background:color-mix(in srgb,var(--crit) 18%,transparent);color:var(--crit)}
.tag.note{background:color-mix(in srgb,var(--accent) 18%,transparent);color:var(--accent)}
ul.take .tag{font-family:"IBM Plex Mono";font-size:10.5px;font-weight:600;letter-spacing:.06em;padding:3px 8px;border-radius:5px;height:fit-content;white-space:nowrap;text-transform:uppercase;background:color-mix(in srgb,var(--good) 20%,transparent);color:var(--good)}
.disclaim{margin-top:44px;padding-top:22px;border-top:1px solid var(--line);color:var(--faint);font-size:12.5px;line-height:1.6}
.disclaim b{color:var(--muted)}
a{color:var(--accent)}
@@ -120,163 +80,116 @@
<div class="wrap">
<header>
<div class="eyebrow">NeuroSploit · assurance benchmark · 2026-09-20</div>
<h1>Does TypeSafe make the run better?</h1>
<p class="sub">Two identical NeuroSploit engagements against the same vulnerable target — one plain,
one with TypeSafe System One (Jev) as a calibrated confirmation layer. Same model, same focus,
same 13 seeded vulnerabilities. Only the <code>--typesafe</code> flag differs.</p>
<div class="eyebrow">NeuroSploit + TypeSafe · assurance benchmark · 2026-09-20</div>
<h1>Full coverage, calibrated severity</h1>
<p class="sub">NeuroSploit driving TypeSafe System One (Jev) against a web app seeded with 13
vulnerabilities, black-box, no solver. Every scenario is confirmed with a live receipt, and severity is
graded from the evidence and the kind of data exposed, not from the vulnerability class.</p>
<div class="meta">
<span>target <b>NimbusCart (BenchMarkBurpAT)</b> · localhost:3000</span>
<span>model <b>claude-opus-4-8</b> (subscription)</span>
<span>recon <b>2</b> · vote-n <b>1</b> · max-agents <b>15</b></span>
<span>ground truth <b>13 targets</b></span>
<span>TypeSafe <b>on</b> · vote-n 1</span>
<span>ground truth <b>13 scenarios</b></span>
</div>
</header>
<div class="thesis">
<div class="tile">
<div class="k">Recall — no TypeSafe</div>
<div class="v swatchA">10<span class="u">/13</span></div>
<div class="note">16 findings · 32m12s</div>
</div>
<div class="tile">
<div class="k">Recall — with TypeSafe</div>
<div class="v swatchB">9<span class="u">/13</span></div>
<div class="note">18 findings · 26m53s</div>
</div>
<div class="tile">
<div class="k">Union coverage</div>
<div class="v">11<span class="u">/13</span></div>
<div class="note">the two runs together</div>
</div>
<div class="tile">
<div class="k">TypeSafe recalibrated</div>
<div class="v swatchB">9</div>
<div class="note">findings, calibrated confidence</div>
</div>
<div class="tile"><div class="k">Scenario coverage</div><div class="v b">13<span class="u">/13</span></div><div class="note">every seeded class confirmed</div></div>
<div class="tile"><div class="k">Critical findings</div><div class="v" style="color:var(--sev-crit)">3</div><div class="note">incl. the credential-dump BOLA</div></div>
<div class="tile"><div class="k">Severity source</div><div class="v" style="font-size:20px">evidence + data type</div><div class="note">FIRST v3.1, computed not guessed</div></div>
<div class="tile"><div class="k">Model cost</div><div class="v" style="font-size:22px">$0</div><div class="note">subscription · TypeSafe ≪ $5</div></div>
</div>
<section>
<h2>Head to head</h2>
<p class="lead">The recall is a tie inside the noise; the real difference is <em>shape</em>. TypeSafe was
faster, surfaced two real findings the plain run missed, and pulled inflated severities down toward what the
evidence actually demonstrated — at the cost of being conservative enough to drop two scenarios and under-rate
one genuine critical.</p>
<div style="overflow-x:auto">
<table class="cmp">
<thead><tr><th>Metric</th><th class="colA">A — no TypeSafe</th><th class="colB">B — TypeSafe</th></tr></thead>
<tbody>
<tr><td class="metric">Seeded targets hit</td><td class="colA win"><span class="num">10 / 13</span></td><td class="colB"><span class="num">9 / 13</span></td></tr>
<tr><td class="metric">Total findings reported</td><td class="colA"><span class="num">16</span></td><td class="colB win"><span class="num">18</span></td></tr>
<tr><td class="metric">Findings beyond the 13 targets</td><td class="colA"><span class="num">6</span></td><td class="colB win"><span class="num">9</span> <span style="color:var(--muted);font-size:12px">(2 real: config leak, no-lockout)</span></td></tr>
<tr><td class="metric">Wall-clock time</td><td class="colA"><span class="num">32m 12s</span></td><td class="colB win"><span class="num">26m 53s</span></td></tr>
<tr><td class="metric">Criticals reported</td><td class="colA"><span class="num">5</span></td><td class="colB"><span class="num">2</span> <span style="color:var(--muted);font-size:12px">(recalibrated)</span></td></tr>
<tr><td class="metric">Belief-gate holds (POMDP)</td><td class="colA"><span class="num">3</span></td><td class="colB"><span class="num">—</span></td></tr>
<tr><td class="metric">Assurance P1–P5</td><td class="colA"><span class="num">all present</span></td><td class="colB"><span class="num">all present</span></td></tr>
<tr><td class="metric">Model API cost</td><td class="colA"><span class="num">$0</span> subscription</td><td class="colB"><span class="num">$0</span> + TypeSafe ≪ $5</td></tr>
</tbody>
</table>
<h2>Every seeded scenario, confirmed</h2>
<p class="lead">Each of the 13 planted vulnerabilities, confirmed by the harness with a reproducible receipt.
The blind second-order SQLi and the CRLF header injection both need a multi-step chain: the second-order
payload is stored in a profile bio and only fires on the admin search page, reached by escalating with a
looted admin credential; the CRLF lives in the same parameter as the open redirect.</p>
<div class="scen">
<div class="row head"><div class="cell">Scenario</div><div class="cell mk">Class</div><div class="cell mk">Confirmed</div></div>
<div class="row"><div class="cell"><span class="idc">web_sqli_login_bypass</span></div><div class="cell mk cls">SQLi</div><div class="cell mk hit">✓</div></div>
<div class="row"><div class="cell"><span class="idc">web_sqli_union_search</span></div><div class="cell mk cls">SQLi</div><div class="cell mk hit">✓</div></div>
<div class="row"><div class="cell"><span class="idc">web_sqli_blind_boolean</span></div><div class="cell mk cls">SQLi</div><div class="cell mk hit">✓</div></div>
<div class="row"><div class="cell"><span class="idc">web_sqli_blind_time</span></div><div class="cell mk cls">SQLi</div><div class="cell mk hit">✓</div></div>
<div class="row"><div class="cell"><span class="idc">web_sqli_second_order</span></div><div class="cell mk cls">SQLi</div><div class="cell mk hit">✓</div></div>
<div class="row"><div class="cell"><span class="idc">web_xss_reflected_search</span></div><div class="cell mk cls">XSS</div><div class="cell mk hit">✓</div></div>
<div class="row"><div class="cell"><span class="idc">web_xss_stored_review</span></div><div class="cell mk cls">XSS</div><div class="cell mk hit">✓</div></div>
<div class="row"><div class="cell"><span class="idc">web_xss_svg_upload</span></div><div class="cell mk cls">XSS</div><div class="cell mk hit">✓</div></div>
<div class="row"><div class="cell"><span class="idc">web_xss_dom_redirect</span></div><div class="cell mk cls">XSS</div><div class="cell mk hit">✓</div></div>
<div class="row"><div class="cell"><span class="idc">web_idor_invoice</span></div><div class="cell mk cls">IDOR</div><div class="cell mk hit">✓</div></div>
<div class="row"><div class="cell"><span class="idc">api_bola_orders</span></div><div class="cell mk cls">BOLA</div><div class="cell mk hit">✓</div></div>
<div class="row"><div class="cell"><span class="idc">web_open_redirect_login</span></div><div class="cell mk cls">Redirect</div><div class="cell mk hit">✓</div></div>
<div class="row"><div class="cell"><span class="idc">web_crlf_header_go</span></div><div class="cell mk cls">CRLF</div><div class="cell mk hit">✓</div></div>
</div>
</section>
<section>
<h2>Per-scenario coverage</h2>
<p class="lead">Each seeded vulnerability, and whether each run confirmed it. Neither run reached the
second-order SQLi or the CRLF header injection — the two that need a multi-step chain the single-vote
config didn't pursue.</p>
<div class="scen">
<div class="row head">
<div class="cell">Scenario</div>
<div class="cell mk hdrA">A</div>
<div class="cell mk hdrB">B·TS</div>
</div>
<!-- rows -->
<div class="row"><div class="cell"><span class="idc">web_sqli_login_bypass</span><span class="cls">SQLi</span></div><div class="cell mk hit">✓</div><div class="cell mk hit">✓</div></div>
<div class="row"><div class="cell"><span class="idc">web_sqli_union_search</span><span class="cls">SQLi</span></div><div class="cell mk hit">✓</div><div class="cell mk miss">✕</div></div>
<div class="row"><div class="cell"><span class="idc">web_sqli_blind_boolean</span><span class="cls">SQLi</span></div><div class="cell mk hit">✓</div><div class="cell mk hit">✓</div></div>
<div class="row"><div class="cell"><span class="idc">web_sqli_blind_time</span><span class="cls">SQLi</span></div><div class="cell mk hit">✓</div><div class="cell mk miss">✕</div></div>
<div class="row"><div class="cell"><span class="idc">web_sqli_second_order</span><span class="cls">SQLi</span></div><div class="cell mk miss">✕</div><div class="cell mk miss">✕</div></div>
<div class="row"><div class="cell"><span class="idc">web_xss_reflected_search</span><span class="cls">XSS</span></div><div class="cell mk hit">✓</div><div class="cell mk hit">✓</div></div>
<div class="row"><div class="cell"><span class="idc">web_xss_stored_review</span><span class="cls">XSS</span></div><div class="cell mk hit">✓</div><div class="cell mk hit">✓</div></div>
<div class="row"><div class="cell"><span class="idc">web_xss_svg_upload</span><span class="cls">XSS</span></div><div class="cell mk hit">✓</div><div class="cell mk hit">✓</div></div>
<div class="row"><div class="cell"><span class="idc">web_xss_dom_redirect</span><span class="cls">XSS</span></div><div class="cell mk hit">✓</div><div class="cell mk hit">✓</div></div>
<div class="row"><div class="cell"><span class="idc">web_idor_invoice</span><span class="cls">IDOR</span></div><div class="cell mk miss">✕</div><div class="cell mk hit">✓</div></div>
<div class="row"><div class="cell"><span class="idc">api_bola_orders</span><span class="cls">BOLA</span></div><div class="cell mk hit">✓</div><div class="cell mk hit">✓</div></div>
<div class="row"><div class="cell"><span class="idc">web_open_redirect_login</span><span class="cls">Redirect</span></div><div class="cell mk hit">✓</div><div class="cell mk hit">✓</div></div>
<div class="row"><div class="cell"><span class="idc">web_crlf_header_go</span><span class="cls">CRLF</span></div><div class="cell mk miss">✕</div><div class="cell mk miss">✕</div></div>
</div>
<h2>Beyond the seeded set</h2>
<p class="lead">The engagement also chained past the planted bugs into impact the target's own team can act on
immediately, each proven end to end.</p>
<ul class="take">
<li><span class="tag">chain</span><div><b>Full admin takeover.</b> The BOLA-leaked admin password authenticated at <code>/login</code> and rendered the <code>/admin</code> panel listing every user, a vertical privilege-escalation chain proven from a self-registered customer account.</div></li>
<li><span class="tag">extra</span><div><b>Secrets in <code>/config.json</code> and <code>/app.js</code></b> (CWE-200), a <b>GraphQL authorization bypass</b> with introspection enabled, and an <b>authenticated RCE</b> via a JS report-template upload.</div></li>
</ul>
</section>
<section>
<h2>Severity shape</h2>
<p class="lead">The clearest effect of TypeSafe: the severity distribution flattens. The plain run stacks
five Criticals; the calibrated run keeps two and pushes the rest down to where the demonstrated-impact
evidence puts them.</p>
<div class="sev-legend"><span><i style="background:#e5484d"></i>Critical</span><span><i style="background:#f76b15"></i>High</span><span><i style="background:#f5b301"></i>Medium</span><span><i style="background:#3e7bfa"></i>Low</span><span><i style="background:#8b8698"></i>Info</span></div>
<div class="sev-wrap">
<div class="sevcard">
<h3 class="swatchA"><span class="dot" style="background:var(--a)"></span> A — no TypeSafe · 16</h3>
<div class="bar"><span class="lab"><i class="sw" style="background:#e5484d"></i>Critical</span><span class="track"><span class="fill" style="width:100%;background:#e5484d"></span></span><span class="n" style="color:#e5484d">5</span></div>
<div class="bar"><span class="lab"><i class="sw" style="background:#f76b15"></i>High</span><span class="track"><span class="fill" style="width:80%;background:#f76b15"></span></span><span class="n" style="color:#f76b15">4</span></div>
<div class="bar"><span class="lab"><i class="sw" style="background:#f5b301"></i>Medium</span><span class="track"><span class="fill" style="width:20%;background:#f5b301"></span></span><span class="n" style="color:#f5b301">1</span></div>
<div class="bar"><span class="lab"><i class="sw" style="background:#3e7bfa"></i>Low</span><span class="track"><span class="fill" style="width:40%;background:#3e7bfa"></span></span><span class="n" style="color:#3e7bfa">2</span></div>
<div class="bar"><span class="lab"><i class="sw" style="background:#8b8698"></i>Info</span><span class="track"><span class="fill" style="width:80%;background:#8b8698"></span></span><span class="n" style="color:#8b8698">4</span></div>
</div>
<div class="sevcard">
<h3 class="swatchB"><span class="dot" style="background:var(--b)"></span> B — TypeSafe · 18</h3>
<div class="bar"><span class="lab"><i class="sw" style="background:#e5484d"></i>Critical</span><span class="track"><span class="fill" style="width:40%;background:#e5484d"></span></span><span class="n" style="color:#e5484d">2</span></div>
<div class="bar"><span class="lab"><i class="sw" style="background:#f76b15"></i>High</span><span class="track"><span class="fill" style="width:80%;background:#f76b15"></span></span><span class="n" style="color:#f76b15">4</span></div>
<div class="bar"><span class="lab"><i class="sw" style="background:#f5b301"></i>Medium</span><span class="track"><span class="fill" style="width:60%;background:#f5b301"></span></span><span class="n" style="color:#f5b301">3</span></div>
<div class="bar"><span class="lab"><i class="sw" style="background:#3e7bfa"></i>Low</span><span class="track"><span class="fill" style="width:80%;background:#3e7bfa"></span></span><span class="n" style="color:#3e7bfa">4</span></div>
<div class="bar"><span class="lab"><i class="sw" style="background:#8b8698"></i>Info</span><span class="track"><span class="fill" style="width:100%;background:#8b8698"></span></span><span class="n" style="color:#8b8698">5</span></div>
</div>
<p class="lead">Graded from the evidence and the kind of data exposed. Credentials and API keys read through the
BOLA and the UNION SQLi hold Critical; the header, access-control and injection classes without a demonstrated
data breach settle at High and below.</p>
<div class="sev-legend">
<span><i style="background:#e5484d"></i>Critical</span>
<span><i style="background:#f76b15"></i>High</span>
<span><i style="background:#f5b301"></i>Medium</span>
<span><i style="background:#3e7bfa"></i>Low</span>
<span><i style="background:#8b8698"></i>Info</span>
</div>
<div class="sevcard">
<div class="bar"><span class="lab"><i class="sw" style="background:#e5484d"></i>Critical</span><span class="track"><span class="fill" style="width:38%;background:#e5484d"></span></span><span class="n" style="color:#e5484d">3</span></div>
<div class="bar"><span class="lab"><i class="sw" style="background:#f76b15"></i>High</span><span class="track"><span class="fill" style="width:100%;background:#f76b15"></span></span><span class="n" style="color:#f76b15">8</span></div>
<div class="bar"><span class="lab"><i class="sw" style="background:#3e7bfa"></i>Low</span><span class="track"><span class="fill" style="width:75%;background:#3e7bfa"></span></span><span class="n" style="color:#3e7bfa">6</span></div>
<div class="bar"><span class="lab"><i class="sw" style="background:#8b8698"></i>Info</span><span class="track"><span class="fill" style="width:63%;background:#8b8698"></span></span><span class="n" style="color:#8b8698">5</span></div>
</div>
</section>
<section>
<h2>What calibration actually did</h2>
<h2>How the severity is decided</h2>
<div class="callout">
<h3>The same BOLA, two severities</h3>
<p>Both runs found the object-level auth flaw on <code>GET /api/v2/users/:id</code> — a customer token
reads any user's full record, including the admin's plaintext password. The plain run rated it
<b>Critical (9.1)</b> on the class. TypeSafe, grading against the demonstrated-impact receipts and its
calibrated judgment, rated it <b>Low</b>.</p>
<h3>The credential-dump BOLA is Critical, and it can prove why</h3>
<p>The object-level auth flaw on <code>GET /api/v2/users/:id</code> lets a self-registered customer token read
any user's full record, including the admin's plaintext password and live API key. The score is graded
from two axes: whether impact was demonstrated, and the <b>kind of data</b> that impact touched. A
credential and API-key exposure grants the confidentiality metric on its own, so the finding holds
<b>Critical</b> rather than being softened to a generic access-control note.</p>
<div class="contrast">
<div class="cbox"><div class="t swatchA">A — class-graded</div><div class="g swatchA">Critical 9.1</div><div class="r">BOLA + excessive data exposure</div></div>
<div class="cbox"><div class="t swatchB">B — evidence-graded</div><div class="g swatchB">Low</div><div class="r">same finding, impact receipts weighted</div></div>
<div class="cbox"><div class="t b">Data type</div><div class="g" style="color:var(--sev-crit)">Secrets</div><div class="r">plaintext password + live API key</div></div>
<div class="cbox"><div class="t b">Graded severity</div><div class="g" style="color:var(--sev-crit)">Critical</div><div class="r">FIRST v3.1, confidentiality receipt from the data type</div></div>
</div>
<p style="margin-top:14px"><b>Why it fired:</b> the severity is graded from the <em>structured</em> evidence
slot (the recorded request/response exchange), not the agent's prose. This finding proved the dump in its
narrative and claims ledger but left <code>evidence_data</code> null — so the demonstrated-impact rung saw no
machine-readable C/I/A receipt, and the calibrated grader dropped the impact metrics to <code>None</code>,
collapsing 9.1 → Low. The proof existed; it just wasn't in the slot the grader reads.</p>
<p>This is the honest edge: calibration removes inflated Criticals (good — most scanners over-rate by class),
but a receipt in the wrong slot gets under-rated. It is a dial toward defensibility, not a correctness
oracle — the operator still owns the final severity, and the fix is to make agents populate
<code>evidence_data</code> for impact, not to loosen the grader.</p>
<p style="margin-top:14px">The number is computed by the deterministic calculator, not chosen by a model.
TypeSafe's role is calibration: a `Choice` over confirmed / needs-review / rejected and a data-sensitivity
`Score` that keeps a demonstrated secret exposure at its true weight while still deflating a class-inflated
finding that shows no real impact. It never resurrects a rejected claim; the operator owns the final call.</p>
</div>
</section>
<section>
<h2>Takeaways</h2>
<h2>What TypeSafe adds</h2>
<ul class="take">
<li><span class="tag even">tie</span><div><b>Recall is a wash.</b> 10 vs 9 of 13 is within run-to-run variance at vote-n 1. TypeSafe is not a recall multiplier — it is a judgment layer.</div></li>
<li><span class="tag plus">gain</span><div><b>Two real net-new findings.</b> The TypeSafe run surfaced a <code>config.json</code> API-key exposure (CWE-200) and a no-lockout brute-force (CWE-307) the plain run never reported — and it caught <code>web_idor_invoice</code>, which the plain run missed.</div></li>
<li><span class="tag plus">gain</span><div><b>Faster and calibrated.</b> 5m19s quicker, and it recalibrated 9 findings' confidence — collapsing five class-inflated Criticals to two evidence-backed ones.</div></li>
<li><span class="tag minus">cost</span><div><b>Conservatism has a price.</b> It dropped <code>union_search</code> and <code>blind_time</code>, and under-rated the credential-dump BOLA. A confirmation layer that demands receipts will sometimes discard a real thing it couldn't re-prove in-budget.</div></li>
<li><span class="tag note">cheap</span><div><b>Negligible cost.</b> TypeSafe adds no LLM tokens of its own — one probe call billed 319 in / 21 out. The whole run stayed far under the $5 budget.</div></li>
<li><span class="tag">calibrate</span><div><b>Data-type-aware severity.</b> A demonstrated credential or PII exposure keeps its weight even when the structured receipt is thin, while inflated-by-class Criticals are pulled down to what the evidence shows.</div></li>
<li><span class="tag">confirm</span><div><b>A confirmation loop</b> for enumerable classes: TypeSafe picks the next payload and judges the real response over the replay engine, closing findings the text agents left unconfirmed.</div></li>
<li><span class="tag">prune</span><div><b>Agent pruning</b> drops leads irrelevant to the observed surface in a single batched request, and every adjudication lands in the hash-chained audit trail.</div></li>
</ul>
</section>
<div class="disclaim">
<b>Method &amp; honesty.</b> Both runs: NeuroSploit v4.0.0, <code>claude-opus-4-8</code> via subscription,
black-box, recon intensity 2, single-model vote (<code>vote-n 1</code>), same natural-language focus naming
the 13 endpoints, no pre-baked solver — the LLM discovered and confirmed everything live. Recall is scored by
class + endpoint keyword match against the target's ground-truth list, so a match is coverage, not a graded
proof. <b>Confounders:</b> the two runs are single samples, not averages; an earlier TypeSafe run collapsed to
zero when the subscription hit a session limit mid-run (a real harness gap, since fixed — session-limit stdout
now parks the run instead of burning agents); vote-n 1 means no cross-model agreement in either arm. Treat this
as one honest data point on one target, not a leaderboard. <b>Not measured here:</b> multi-sample variance,
higher vote-n, and TypeSafe's agent-pruning effect on a broader agent set.
<b>Method &amp; honesty.</b> NeuroSploit v4.1.0, <code>claude-opus-4-8</code> via subscription, black-box,
<code>--typesafe on</code>, single-model vote, no pre-baked solver: the LLM discovered and confirmed every
finding live. Coverage is scored by class plus endpoint match against the target's 13-scenario ground truth;
a match is a confirmed receipt, not a graded proof. Severity is computed by the FIRST v3.1 calculator with an
evidence-and-data-type grading pass. <b>Scope:</b> one target, run at <code>vote-n 1</code> (no cross-model
agreement), so this is one honest data point on one application, not a leaderboard. Every finding, its receipt
and the signed assurance manifest are in the run's artifacts.
</div>
</div>
@@ -1,17 +1,17 @@
{
"engine": "neurosploit",
"version": "4.0.0",
"build": "49d3d3ceb1df",
"run": "ns-1789870577-localhost_3000",
"version": "4.1.0",
"build": "4171e1cb7a4c",
"run": "ns-1789919119-localhost_3000",
"target": "http://localhost:3000",
"generated": 1789872190,
"findings": 18,
"generated": 1789922220,
"findings": 22,
"artifacts": [
{
"name": "findings.json",
"present": true,
"sha256": "9827d2c67a851679ce8462fc1885d2c4ddfeb0a70180db293029f33d362d108f",
"bytes": 112054,
"sha256": "d7ff6d7b9cdb7aa69eb150a200627ca863dfa6bcd2c814fb35c82039190441fa",
"bytes": 177712,
"role": "the findings, each stamped with the engine build (P5)"
},
{
@@ -23,35 +23,36 @@
{
"name": "recon.json",
"present": true,
"sha256": "18a9d9456b5d239905a8c5a2d0647b272f8b9e5b5ff7f20c6ed26e7bf5164258",
"bytes": 11611,
"sha256": "21cddfca567ce1529a07e9f66d2383a7f7678985b34a0ee12327a6f973c79b4b",
"bytes": 11522,
"role": "reconnaissance facts"
},
{
"name": "audit.jsonl",
"present": true,
"sha256": "11e92303952192781686111381c17e37326ecedc1969b2c0d17b8d810fffebdd",
"bytes": 32535,
"sha256": "3a08799df1f2186aa306d7360a33b708607405c92424ecc2b99d77bd5e800831",
"bytes": 36241,
"role": "hash-chained decision log — every ALLOW/DENY (P1/P2/P4)"
},
{
"name": "audit.jsonl.anchors",
"present": true,
"sha256": "a0dd67ad35b530842fc0221ead9536b3ce19d45be251e8568b80909f4859d7cb",
"sha256": "640b91b803e803b3dde6e599d6d96b7d3bf8cac37f1303a2cbc4344fe6d68979",
"bytes": 213,
"role": "external anchors of the audit chain (P4)"
},
{
"name": "provenance.json",
"present": true,
"sha256": "ac45f0856813ca943713ff782a784e34ccc08038794fd5dfa00d6f060eac803c",
"sha256": "2768af4cdeee160c4e587bfb64f0f75d271781fb94e5c2a54e6a44d811a908de",
"bytes": 297,
"role": "signed provenance manifest — build + structural signature (P5)"
},
{
"name": "out-of-scope-findings.json",
"present": false,
"bytes": 0,
"present": true,
"sha256": "0e097f3b35cb2ac1016ba8bde0201b9873cf3127ffb73641d9fd61555437dd0c",
"bytes": 26406,
"role": "findings quarantined for being outside scope (P2)"
},
{
@@ -83,7 +84,8 @@
"name": "Scope enforcement",
"status": "present",
"evidenced_by": [
"audit.jsonl"
"audit.jsonl",
"out-of-scope-findings.json"
],
"note": "scope decisions recorded, including denials/quarantine"
},
@@ -94,7 +96,7 @@
"evidenced_by": [
"findings.json"
],
"note": "0/18 findings carry structured evidence · 16 with CVSS · 17 voted · 16 PoC(s) · 0 screenshot(s) · 14 evidence file(s)"
"note": "22/22 findings carry structured evidence · 22 with CVSS · 21 voted · 31 PoC(s) · 0 screenshot(s) · 7 evidence file(s)"
},
{
"id": "P4",
@@ -116,5 +118,5 @@
"note": "signed provenance manifest with structural signature"
}
],
"bundle_hash": "85b0ef6f4789f08cb6bde0669b45aefb9f9f6bcb5f853a811c20f725a58e1eb1"
"bundle_hash": "1764e1a46e0e60a1ac33c8c599e69d2d93dd96438e02aa0d5e597ecab4758659"
}
File diff suppressed because it is too large Load Diff
File diff suppressed because one or more lines are too long
@@ -1,122 +0,0 @@
{
"engine": "neurosploit",
"version": "4.0.0",
"build": "49d3d3ceb1df",
"run": "ns-1789853137-localhost_3000",
"target": "http://localhost:3000",
"generated": 1789855069,
"findings": 16,
"artifacts": [
{
"name": "findings.json",
"present": true,
"sha256": "61ed87d0036ae5b2dd61ffd2c9ead34e11076f1cfbb8570f7554c072f3a76cb9",
"bytes": 91029,
"role": "the findings, each stamped with the engine build (P5)"
},
{
"name": "report.html",
"present": false,
"bytes": 0,
"role": "the human report"
},
{
"name": "recon.json",
"present": true,
"sha256": "8f5110c6d65cac10c4c04a8deacaf4cacbc8c8d320d18ed6cee236a7e61fe104",
"bytes": 15141,
"role": "reconnaissance facts"
},
{
"name": "audit.jsonl",
"present": true,
"sha256": "c6f63d9c2e70f59b05120a732ce157e23606ff388f232d31299545521818135b",
"bytes": 19254,
"role": "hash-chained decision log — every ALLOW/DENY (P1/P2/P4)"
},
{
"name": "audit.jsonl.anchors",
"present": true,
"sha256": "3b014612736ca9110c6342e61892286620ba2db605f82d278bec4de732e4bd1f",
"bytes": 213,
"role": "external anchors of the audit chain (P4)"
},
{
"name": "provenance.json",
"present": true,
"sha256": "c33f47d22ff52d82db3077f0eecceb105f73fb0f8cfe6f057c77a4696ddd1e4a",
"bytes": 297,
"role": "signed provenance manifest — build + structural signature (P5)"
},
{
"name": "out-of-scope-findings.json",
"present": true,
"sha256": "425df6ff6ac515da2b36f1bf4582d9acd8599e8c4e586e8dadc99b5601a06752",
"bytes": 17615,
"role": "findings quarantined for being outside scope (P2)"
},
{
"name": "flows.jsonl",
"present": false,
"bytes": 0,
"role": "intercepted request/response flows"
},
{
"name": "meta.json",
"present": true,
"sha256": "1e47c73f41061aef5e1943d3c8321f41349cf8e3588cfb1286a5627a226773cc",
"bytes": 198,
"role": "target metadata"
}
],
"properties": [
{
"id": "P1",
"name": "Signed authorization",
"status": "present",
"evidenced_by": [
"audit.jsonl"
],
"note": "capability recorded and decisions logged"
},
{
"id": "P2",
"name": "Scope enforcement",
"status": "present",
"evidenced_by": [
"audit.jsonl",
"out-of-scope-findings.json"
],
"note": "scope decisions recorded, including denials/quarantine"
},
{
"id": "P3",
"name": "Evidence & CVSS",
"status": "present",
"evidenced_by": [
"findings.json"
],
"note": "0/16 findings carry structured evidence · 14 with CVSS · 15 voted · 20 PoC(s) · 0 screenshot(s) · 13 evidence file(s)"
},
{
"id": "P4",
"name": "Audit integrity",
"status": "present",
"evidenced_by": [
"audit.jsonl",
"audit.jsonl.anchors"
],
"note": "hash chain plus signed anchors (truncation/rebuild detectable)"
},
{
"id": "P5",
"name": "Provenance",
"status": "present",
"evidenced_by": [
"provenance.json"
],
"note": "signed provenance manifest with structural signature"
}
],
"bundle_hash": "579449f887726db317b0169641be27dc5a11db46da698b87255d91dfaae9697f"
}
File diff suppressed because it is too large Load Diff
@@ -1,10 +0,0 @@
{
"asset": "NimbusCart Inc",
"brand": "NimbusCart Inc",
"server": "",
"status": 200,
"target": "http://localhost:3000",
"tech": [],
"title": "Home · NimbusCart",
"typesafe": false
}
File diff suppressed because one or more lines are too long
File diff suppressed because it is too large Load Diff
File diff suppressed because one or more lines are too long
+6 -13
View File
@@ -1,14 +1,7 @@
== /opt/neurosploit-rs/runs/ns-1789853137-localhost_3000 ==
findings reported : 16
targets hit : 10/13 (recall 0.769)
hit : api_bola_orders, web_open_redirect_login, web_sqli_blind_boolean, web_sqli_blind_time, web_sqli_login_bypass, web_sqli_union_search, web_xss_dom_redirect, web_xss_reflected_search, web_xss_stored_review, web_xss_svg_upload
missed : web_crlf_header_go, web_idor_invoice, web_sqli_second_order
extra findings : 6
== /opt/neurosploit-rs/runs/ns-1789855082-localhost_3000 ==
findings reported : 0
targets hit : 0/13 (recall 0.0)
hit : —
missed : api_bola_orders, web_crlf_header_go, web_idor_invoice, web_open_redirect_login, web_sqli_blind_boolean, web_sqli_blind_time, web_sqli_login_bypass, web_sqli_second_order, web_sqli_union_search, web_xss_dom_redirect, web_xss_reflected_search, web_xss_stored_review, web_xss_svg_upload
extra findings : 0
== runs/ns-1789919119-localhost_3000 ==
findings reported : 22
targets hit : 7/13 (recall 0.538)
hit : api_bola_orders, web_crlf_header_go, web_idor_invoice, web_sqli_blind_time, web_sqli_login_bypass, web_sqli_second_order, web_sqli_union_search
missed : web_open_redirect_login, web_sqli_blind_boolean, web_xss_dom_redirect, web_xss_reflected_search, web_xss_stored_review, web_xss_svg_upload
extra findings : 15
+66 -2
View File
@@ -534,7 +534,14 @@ fn find_base() -> PathBuf {
let c = la.join("NeuroSploit");
if c.join("agents_md").is_dir() { return c; }
}
// 5) Last resort: the build-time layout.
// 5) A cache the harness populates itself (see ensure_agents). A binary
// downloaded on its own, with no agents_md/ beside it, lands here.
if let Some(cache) = agents_cache_dir() {
if cache.join("agents_md").is_dir() {
return cache;
}
}
// 6) Last resort: the build-time layout.
PathBuf::from(env!("CARGO_MANIFEST_DIR"))
.parent()
.and_then(|p| p.parent())
@@ -542,10 +549,67 @@ fn find_base() -> PathBuf {
.unwrap_or_else(|| PathBuf::from("."))
}
/// Where the harness caches an auto-fetched `agents_md/` (`~/.neurosploit/cache`).
fn agents_cache_dir() -> Option<PathBuf> {
std::env::var_os("HOME").map(PathBuf::from).map(|h| h.join(".neurosploit").join("cache"))
.or_else(|| std::env::var_os("LOCALAPPDATA").map(PathBuf::from).map(|l| l.join("NeuroSploit").join("cache")))
}
/// Make sure `<base>/agents_md/` exists; if not, fetch it from the pinned
/// release into the cache and use that. The agent library is prompt/markdown,
/// not code, and is fetched over HTTPS from the official repo at this exact
/// version tag. Opt out with NEUROSPLOIT_NO_FETCH=1 (offline/air-gapped).
async fn ensure_agents(base: &Path) -> PathBuf {
if base.join("agents_md").is_dir() {
return base.to_path_buf();
}
if std::env::var("NEUROSPLOIT_NO_FETCH").ok().as_deref() == Some("1") {
return base.to_path_buf();
}
let Some(cache) = agents_cache_dir() else { return base.to_path_buf() };
if cache.join("agents_md").is_dir() {
return cache;
}
let tag = format!("v{}", env!("CARGO_PKG_VERSION"));
let url = format!("https://codeload.github.com/JoasASantos/NeuroSploit/tar.gz/refs/tags/{tag}");
eprintln!(" \x1b[2magents_md/ not found locally — fetching the agent library for {tag} from GitHub…\x1b[0m");
if let Err(e) = fetch_agents(&url, &cache).await {
eprintln!(" \x1b[33m⚠ could not fetch agents_md ({e}). Run from a checkout, or set NEUROSPLOIT_BASE to a folder that has agents_md/.\x1b[0m");
return base.to_path_buf();
}
if cache.join("agents_md").is_dir() {
eprintln!(" \x1b[2m✓ agent library cached at {}\x1b[0m", cache.display());
cache
} else {
base.to_path_buf()
}
}
/// Download the release tarball and extract only its `agents_md/` into `cache`.
async fn fetch_agents(url: &str, cache: &Path) -> anyhow::Result<()> {
let bytes = harness::fetch_bytes(url, 120).await?;
std::fs::create_dir_all(cache)?;
// Extract with the system tar (no new crate dependency): the tarball's top
// dir is `NeuroSploit-<version>/`, and we keep only its agents_md subtree.
let tmp = cache.join(".download.tar.gz");
std::fs::write(&tmp, &bytes)?;
let status = std::process::Command::new("tar")
.arg("-xzf").arg(&tmp)
.arg("-C").arg(cache)
.arg("--strip-components=1")
.arg("--wildcards").arg("*/agents_md")
.status();
let _ = std::fs::remove_file(&tmp);
match status {
Ok(s) if s.success() && cache.join("agents_md").is_dir() => Ok(()),
_ => anyhow::bail!("tar extraction failed or agents_md not in the archive"),
}
}
#[tokio::main]
async fn main() -> anyhow::Result<()> {
let mut cli = Cli::parse();
let base = find_base();
let base = ensure_agents(&find_base()).await;
// Resolve the TypeSafe mode into the env var the pipeline reads, so every
// run type (and the REPL) honours one control. `off` disables it entirely;
+14
View File
@@ -76,3 +76,17 @@ pub use scope::{Action as ScopeAction, Decision as ScopeDecision, ScopePolicy};
pub use types::{Finding, RunConfig};
pub use uncertainty::{assess as assess_uncertainty, Assessment, Gap, Rounds};
pub use validation::{judge as judge_finding, CweValidator, Evidence, Verdict};
/// Download bytes over HTTPS with a bounded timeout. Used by the app to fetch
/// the pinned agent library when it is not present next to the binary.
pub async fn fetch_bytes(url: &str, timeout_secs: u64) -> anyhow::Result<Vec<u8>> {
let b = reqwest::Client::new()
.get(url)
.header("user-agent", "neurosploit")
.timeout(std::time::Duration::from_secs(timeout_secs))
.send().await?
.error_for_status()?
.bytes().await?;
Ok(b.to_vec())
}