mirror of
https://github.com/CyberSecurityUP/NeuroSploit.git
synced 2026-10-10 18:04:04 +02:00
07f97076de8ba2cec803e88545b9048b9ad9e219
4
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
f1fb6b8bc7 |
feat: PoC validator, Kali sandbox, intercept proxy, compliance, +8 validators
Closes the three benchmark gaps and adds the two the user asked for. poc.rs — re-runs each finding's recorded proof and sorts it into reproduced / changed / gone / unverifiable. The last two are kept apart deliberately: a PoC that could not be tested (out of scope now, state-changing, nothing recorded) is never reported as one that failed. Never re-runs a mutating request to "confirm" it. Can only lower a finding's standing, never raise it. Wired as a run pass (--revalidate-poc) and a subcommand (neurosploit poc <run> --apply). proxy.rs — own recording forward proxy (HTTP in full; HTTPS tunnelled with honest metadata, no fake CA) that chains upstream to Burp / Caido / ZAP / mitmproxy. A bare tool routes straight through it; own+tool records here and forwards for full TLS interception. Flows -> flows.jsonl, distinct hosts become passive-discovery leads. Harness and agent child commands share one route. sandbox.rs — Kali docker/podman container: no host network, no mounted socket, no-new-privileges, workdir mounted, proxy/transport env inherited. A missing runtime is an explicit error, never a silent fallback to host execution — the whole point being to keep attack payloads off the operator's host. Subcommands sandbox up|exec|install|down. compliance.rs — maps confirmed findings onto PCI-DSS v4.0, HIPAA Security Rule and SOC 2 controls. Phrased as "bears on control X", never "compliant/non- compliant"; the disclaimer is rendered on top and absence of a finding is never presented as compliance. Report section + `neurosploit compliance <run>`. validation.rs — 8 new deterministic validators (19 -> 27 classes): verbose errors/stack traces (CWE-209), cleartext/HSTS (319), CRLF response splitting (113), dangerous HTTP methods (650), GraphQL introspection, exposed backup files (530), Host header injection (644), cacheable private responses (525). Each names exactly what it saw and rejects the classic false positives (a block page echoing a payload, the SPA served under a bogus path, a copyright year mistaken for a code). All wired through RunConfig, the CLI (global --intercept/--sandbox; run-level --revalidate-poc/--compliance) and the web console's Tooling & assurance block. 328 tests. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
408350539f |
feat(budget,provenance): reasoning budget modes and JOASNSCOPE provenance
Budget (opt-in, unlimited by default so an un-budgeted run is unchanged):
- crates/harness/src/budget.rs — modes, phase shares, Token Governor
- CLI: --budget/--token-limit/--deep-test-limit/--coverage-first/
--depth-first/--sample-per-route; same controls in the web wizard
- pipeline honours it: vote_n narrows, evidence rounds are capped
Run control parity in the web console:
- /pause in the REPL, backed by a pause gate in the model pool: in-flight
agents finish, then the run holds with every finding kept
- POST /api/exploit/:id/{pause,continue,report} + GET .../log
Provenance (crates/harness/src/provenance.rs):
- JOASNSCOPE sigil leads every canary, so a marker found in a response,
a log or someone else's report extracts whole and names its build
- per-build fingerprint, per-run id, optional per-customer build id
- findings.json stamped with _engine; signed provenance.json manifest
- structural signature survives rewording but not a changed result set
- prompts watermarked at the single pool chokepoint
- `neurosploit provenance show|scan|verify`
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
6475dba752 |
feat(validation): 13 more CWE validators, each with a rejection rule
Six classes had deterministic rules; the rest of a run still rested on models
voting. These thirteen cover the classes that produce the most false positives
in AI-driven testing, and each one is written around what *disproves* the
claim, because that is the part a language model skips:
SSTI an expression evaluated server-side whose result was never
sent — the payload echoing its own "result" is rejected
XXE entity content or an OOB callback; a parser error mentioning
entities shows the DTD was read, not that anything resolved
Open redirect 3xx WITH a Location off-site; a rendered link is not a redirect
CORS reflected Origin PLUS credentials; ACAO:* without credentials
exposes only what an anonymous client could already read, and
ACAO:* WITH credentials is refused by browsers anyway
Cookie flags fully decidable from Set-Cookie + scheme
Clickjacking neither X-Frame-Options nor CSP frame-ancestors
Auth bypass protected content with NO credentials sent — a "bypass" whose
request still carried a cookie is rejected, as is a redirect
to login
JWT forged token accepted AND privileged content returned
Rate limiting >= 20 attempts with no 429/Retry-After; five attempts prove
nothing about a limit that was never reached
Session fix. the session id surviving login unchanged
Mass assign. a read-back proving the field persisted — a 200 on the write
means nothing, APIs accept and ignore extra fields routinely
CSRF a cross-origin state change read back; a SameSite session
cookie means a browser would never attach it cross-site
Exposure a real secret/listing signature the baseline lacked; a
soft-404 mirroring the baseline page is rejected
Exchange gains response and request headers, because several of these classes
are decided by a header (Location, Set-Cookie, Access-Control-Allow-*) and the
body alone is not evidence for them.
A test asserts no two validators claim the same CWE — ambiguous ownership would
make routing depend on registration order, which is how a class silently gets
the wrong rule.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
093c87fbc6 |
feat(harness): enforced scope guard + deterministic Evidence & Validation Engine
Two gaps this closes, both found by reading what the code actually did.
Scope was never enforced
------------------------
`out_of_scope` was rendered into the prompt as "HARD CONSTRAINT — do NOT test…"
and nothing checked it. That is a request to a model, not a control: an agent
that decided a discovered subdomain was interesting, or that followed a
redirect off-target, was free to act and the operator found out by reading the
report.
scope.rs adds a guard in code:
- Hard scope: allowlist of hosts, *.wildcards, IPv4 CIDRs, URL prefixes, with
exclusions that always win. Defaults to the engagement's own target, so
discovery cannot widen authorization — finding a host is not permission to
attack it. An unconfigured policy is closed, not open.
- Soft scope: observe-only zones, destructive verbs (off by default), an
account-creation cap, a rate guard that warns rather than silently dropping
requests (a dropped request reads as "target unreachable"), and payload
classes refused even in scope because they damage the target instead of
demonstrating a bug.
- Enforced at the harness's own chokepoint (probe) and as a post-run audit:
findings proven against an unauthorized host are withheld from the report and
written to out-of-scope-findings.json as an incident to disclose, because
shipping one would launder the mistake.
- REPL: /inscope, /observe, /guardrail, /policy; /scope-out now promotes
host-shaped entries into enforced rules immediately, and says plainly when an
entry is prose the guard cannot enforce.
Validation was models checking models
-------------------------------------
N-model voting plus an adversarial refute pass share the failure mode of the
thing they check — agreement is not evidence, and a confident hallucination
survives a vote by being confident. grounding.rs helps but matches keywords
("http/", "status", "alert(") and cannot tell a real response from a plausible
transcript of one.
validation.rs asks a different question — does the recorded evidence
demonstrate THIS class? — with per-CWE rules and no model in the loop:
SQLi baseline/attack difference that reproduces >= 2x
XSS a browser executed a harness-chosen marker; reflection is not proof
IDOR identity B reads A's resource AND the body matches (a 200 returning a
login page is rejected, which is the classic false positive)
SSRF controlled callback or canary retrieval
LFI controlled marker or a file signature the baseline lacked
RCE a unique nonce in output/callback; reflected input is rejected
Absent evidence is never a pass, and a class with no rule is never
auto-confirmed. NEUROSPLOIT_VALIDATION=advisory (default) rejects
contradictions without demoting voted findings for missing artifacts;
enforcing makes the verdict the status. The evidence contract is injected into
exploit prompts so agents collect the artifacts while they still hold the
target.
Finding gains evidence_data so agents can emit structured artifacts alongside
the finding JSON.
Two bugs the tests caught while writing this: the scope guard treated a SAST
`src/auth.rs:42` endpoint as a host and quarantined valid source findings, and
two canaries minted in the same clock tick came out identical — a marker that
repeats would let a stale token vouch for a new finding.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
|