Commit Graph
2 Commits
Author SHA1 Message Date
CyberSecurityUPandClaude Opus 5 b4903575c0 feat(assurance): target gate (default-deny), evidence-graded CVSS, audit anchoring, P1–P5 bundle
The five immediate priorities from the assurance review — the harness-core
ones, not the commercial/research items (Ed25519, enterprise mode, ablation,
multi-target benchmark are deferred, noted as such).

P1 — target authorization gate. scope.rs::validate_target checks protocol,
host, port and URL prefix before ANY recon. A capability token that does not
cover the CLI target now refuses the run with DENY_TARGET_OUTSIDE_GRANT,
audits it, and exits non-zero — closing the auto-trust-the-target bypass.
RunOutput carries a `denied` code the CLI turns into a non-zero exit.

P6 — cvss.rs: the FIRST v3.1 base equation verbatim (roundup, scope
coefficients), validated against first.org reference vectors (9.8, 6.1, 10.0,
7.8, 7.5, 5.3, 3.1). grade() drops any impact metric that raises severity
without a receipt to a *demonstrated* vector, keeping the *potential* one for
context — SQLi with no extraction scores 0 demonstrated / 9.8 potential, never
a manufactured critical.

P4 — audit.rs anchoring: signed checkpoints of the chain head, local and (with
NEUROSPLOIT_ANCHOR_DIR) external append-only. verify_anchored() catches
truncation (chain shorter than an anchor) and silent rebuilds (head hash no
longer matches), and forged anchors via signature. `neurosploit audit --anchor`.

P5 — assurance.rs bundle: one assurance.json per run — every artifact with its
SHA-256, which of P1–P5 it evidenced (present/partial/absent, never flattered),
a bundle hash and a signature. `neurosploit assurance <run> [--verify]`.

Also +8 deterministic validators earlier this session (19→27). Deferred and
documented: Ed25519 tokens (#3), enterprise mode (#25), benchmark/ablation
(#20/#21), model pinning + reproducibility (#22/#23), README claims taxonomy
(#24), per-agent seccomp (#14).

347 tests.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-18 21:25:21 -03:00
CyberSecurityUPandClaude Opus 5 3456c32f4d feat(harness): risk model, engagement policies, capability tokens, audit trail
Four pieces that together answer "may this action happen, under whose
authority, and can we prove afterwards what we did".

policy.rs — effective_risk per action, exactly as specified:
  (action_risk + asset_criticality + protocol_risk + privilege_level
   + blast_radius) × environment_multiplier
Every term is named and kept on the result, so the number can be explained
rather than argued with. Three policies sit on it: SafetyPolicy (ceilings,
approval thresholds, hard prohibitions), ReasoningPolicy (baseline before
payload, bounded hypotheses, evidence before escalation, explicit stop
conditions) and ProofOfImpactPolicy (what a severity must carry before it may
be that severity).

OT/ICS/SCADA is treated as its own regime, not web testing on odd ports.
Industrial protocols authenticate nothing — a Modbus write is the protocol
working as intended, addressed to a device that may be holding a valve — and
scanners crash PLCs by sending unexpected data at line rate. So the OT profile
blocks writes, disruptive actions, fuzzing and exploit payloads outright, caps
the rate at ~1 req/s, and refuses the function codes that stop a CPU (Modbus
5/6/8/15/16/22/23/43, S7 start/stop, DNP3 restart/stop). Safety instrumented
systems are off limits in every profile.

A test caught a calibration error worth keeping: a plain READ of a critical PLC
scores 3.6 on this formula, so the obvious tight ceiling would have refused
exactly the observation OT findings come from. In an industrial environment it
is the KIND of action that is forbidden, not the arithmetic — the ceiling
catches extremes and the low approval threshold makes anything past trivial
observation a human's decision.

capability.rs — HMAC-signed grants: who authorized what, against which hosts,
in which environment, until when. The harness verifies the signature before
reading a single claim (a well-formed token from the wrong key must never get
to influence what the harness believes), refuses expired and not-yet-valid
tokens, and treats the grant as a CEILING: constrain() intersects it with local
configuration, so config can narrow authorization and never widen it. Tokens
carry no secrets — the payload is readable by anyone holding it.

audit.rs — one structured record per action, in the specified shape (timestamp,
agent, hypothesis, action, target, policy_decision, operator, tool, result,
evidence_hash, capability_token). Two things make it worth having: it is
hash-chained, so removing or editing an entry breaks every hash that follows
and verify() says which one; and it records REFUSALS, because a trail
containing only what happened cannot demonstrate restraint. Only the grant's
id is recorded, never the token — the trail gets shared.

Hard kill conditions end a run outright: target unresponsive after our traffic,
sustained 5xx, out-of-scope request, forbidden industrial function code, safety
system addressed, capability expired mid-run, repeated policy violations,
budget exhausted, operator stop. Failures BEFORE the target ever answered do
not count — nothing listening is not the same as knocked over. The OT switch
trips far sooner: a PLC missing two requests already warrants stopping.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-13 17:56:31 -03:00