feat(harness): enforced scope guard + deterministic Evidence & Validation Engine

Two gaps this closes, both found by reading what the code actually did.

Scope was never enforced
------------------------
`out_of_scope` was rendered into the prompt as "HARD CONSTRAINT — do NOT test…"
and nothing checked it. That is a request to a model, not a control: an agent
that decided a discovered subdomain was interesting, or that followed a
redirect off-target, was free to act and the operator found out by reading the
report.

scope.rs adds a guard in code:
- Hard scope: allowlist of hosts, *.wildcards, IPv4 CIDRs, URL prefixes, with
  exclusions that always win. Defaults to the engagement's own target, so
  discovery cannot widen authorization — finding a host is not permission to
  attack it. An unconfigured policy is closed, not open.
- Soft scope: observe-only zones, destructive verbs (off by default), an
  account-creation cap, a rate guard that warns rather than silently dropping
  requests (a dropped request reads as "target unreachable"), and payload
  classes refused even in scope because they damage the target instead of
  demonstrating a bug.
- Enforced at the harness's own chokepoint (probe) and as a post-run audit:
  findings proven against an unauthorized host are withheld from the report and
  written to out-of-scope-findings.json as an incident to disclose, because
  shipping one would launder the mistake.
- REPL: /inscope, /observe, /guardrail, /policy; /scope-out now promotes
  host-shaped entries into enforced rules immediately, and says plainly when an
  entry is prose the guard cannot enforce.

Validation was models checking models
-------------------------------------
N-model voting plus an adversarial refute pass share the failure mode of the
thing they check — agreement is not evidence, and a confident hallucination
survives a vote by being confident. grounding.rs helps but matches keywords
("http/", "status", "alert(") and cannot tell a real response from a plausible
transcript of one.

validation.rs asks a different question — does the recorded evidence
demonstrate THIS class? — with per-CWE rules and no model in the loop:
  SQLi   baseline/attack difference that reproduces >= 2x
  XSS    a browser executed a harness-chosen marker; reflection is not proof
  IDOR   identity B reads A's resource AND the body matches (a 200 returning a
         login page is rejected, which is the classic false positive)
  SSRF   controlled callback or canary retrieval
  LFI    controlled marker or a file signature the baseline lacked
  RCE    a unique nonce in output/callback; reflected input is rejected
Absent evidence is never a pass, and a class with no rule is never
auto-confirmed. NEUROSPLOIT_VALIDATION=advisory (default) rejects
contradictions without demoting voted findings for missing artifacts;
enforcing makes the verdict the status. The evidence contract is injected into
exploit prompts so agents collect the artifacts while they still hold the
target.

Finding gains evidence_data so agents can emit structured artifacts alongside
the finding JSON.

Two bugs the tests caught while writing this: the scope guard treated a SAST
`src/auth.rs:42` endpoint as a host and quarantined valid source findings, and
two canaries minted in the same clock tick came out identical — a marker that
repeats would let a stale token vouch for a new finding.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
CyberSecurityUP
2026-09-13 15:03:21 -03:00
co-authored by Claude Opus 5
parent 6e1b73e036
commit 093c87fbc6
8 changed files with 1850 additions and 12 deletions
+53
View File
@@ -286,6 +286,59 @@ time. Two stores fix that, both under `.neurosploit/` in the project directory:
harness derived itself are marked `inferred` and drawn dashed in the web console. Secrets
never enter the graph; they stay in the vault.
### Scope: enforced, not requested
`out_of_scope` used to be a sentence in the prompt and nothing checked it — a
*request* to the model, not a control. Scope is now a guard in code
(`crates/harness/src/scope.rs`):
- **Hard scope** — an allowlist of hosts, `*.wildcards`, IPv4 CIDRs and URL
prefixes, plus exclusions that always win. It defaults to **the engagement's
target and nothing else**, so discovery can never widen the engagement:
finding a subdomain in a JS bundle is not authorization to test it.
- **Soft scope** — guardrails inside authorized territory: observe-only zones,
destructive HTTP verbs (off by default), account-creation cap, request-rate
guard, and payload classes that are never acceptable (data destruction, DoS)
— refused even against an in-scope host.
- Findings proven against a host outside the boundary are **withheld from the
report** and written to `out-of-scope-findings.json` as an incident to
disclose.
```
/inscope *.example.com 10.0.0.0/24 # authorize more
/scope-out payments.example.com # host-shaped entries become ENFORCED denials
/observe legacy.example.com # discovery allowed, interaction blocked
/guardrail destructive on · accounts 5 · rate 60
/policy # what is actually enforced
```
### Evidence & Validation Engine
Voting is models checking models, and a confident hallucination passes a vote by
being confident. `crates/harness/src/validation.rs` adds a deterministic layer
that never consults a model:
```
HYPOTHESIS → CANDIDATE → [ VALIDATION ENGINE ] → CONFIRMED | NEEDS_REVIEW | REJECTED
```
Per-CWE rules, because "is this real?" has a different answer per class:
| class | what confirms it |
|-------|------------------|
| SQLi (89/943) | baseline vs attack difference **that reproduces ≥2×** |
| XSS (79/80) | a real browser executed a **harness-chosen marker** — reflection alone is not proof |
| IDOR/BOLA (639/862/863) | identity B reads identity A's resource **and the body matches** (a 200 returning a login page is rejected) |
| SSRF (918) | controlled callback, or retrieval of a canary resource |
| LFI (22/23/98) | controlled file marker, or a file signature the baseline lacked |
| RCE (77/78/94) | a unique nonce in command output or a callback — reflected input is rejected |
Two rules keep it honest: absent evidence is **never** a pass (it becomes
`needs-review`), and a class with no rule is never auto-confirmed.
`NEUROSPLOIT_VALIDATION=advisory|enforcing|off` — advisory (default) rejects
contradictions but won't demote a voted finding merely for missing artifacts;
enforcing makes the verdict the status.
### Keeping a run going
- **Command rectification** — a mistyped command is corrected (`/staus` → `/status`), completed