mirror of
https://github.com/CyberSecurityUP/NeuroSploit.git
synced 2026-09-29 20:41:51 +02:00
v4.2.1
19
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
7c25958827 |
feat(web): make FAIR top-contributor rows clickable to their finding
Each dashboard "Top contributors" row is now a button: clicking it opens the run it belongs to and pops that finding's detail modal (evidence/impact/PoC). Matches the finding by title, falling back to CWE; if it was recalibrated or merged, opens the run and says so. Fetches the run detail directly to avoid racing loadDetail's async fill. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
f82e3fe265 |
feat: deepen 268 exploitation skills; web session delete; CSS design system; JEV progress checkpoint
agents_md (skills):
- enrich all 255 vulns/ + 13 chains/ agents from thin one-liner stages to
concrete playbooks: exact tools/commands, per-stack decision points, benign
proof markers (unique OOB nonces, single reads, URLDNS-before-exec), explicit
proof criteria, false-positive/pitfall sections, and chaining hooks. Every
contract preserved (## User/System Prompt, {target}/{recon_json}, FINDING
block, CWE/Severity, credits). avg 37->53 lines; loader parses all 449.
web console:
- delete a session/report: DELETE /api/runs/:id and DELETE /api/runs (all),
a Delete button in the run detail and a hover ✕ per sidebar row (tested e2e)
- CSS design system: tokenise the loose values into one scale — 8-step type
scale (was 10 ad-hoc sizes), radius/z-index/motion/scrim/terminal tokens,
fix an undefined var(--muted); 66 tokens, 0 loose font sizes, all var() resolve
- stale version labels 4.0.0/4.2.0 -> 4.2.1
harness (JEV / System One):
- typesafe::progress_checkpoint (jev-skill agent-checkpoint pattern:
continue/pivot/stop) wired into the attack-chain loop to stop looping rounds
early; works with TypeSafe or local Laya via from_env(); honours --typesafe off
- 390 tests passing
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
|
||
|
|
4b71ac63a0 |
feat(mobile): binary/APK/IPA testing mode + 12 RE skills — v4.2.0
New `mobile` engagement mode: `neurosploit mobile <app.apk|app.ipa|binary>` reverse-engineers a local artifact with a dedicated `mobile` agent set, all headless and provisioned on demand (Ghidra analyzeHeadless, MobSF REST/Docker, Frida, apktool/jadx, radare2). Twelve original, generic skills (agents_md/mobile/, English): static binary triage, APK static analysis, IPA static analysis, RASP & anti-tamper mapping, root/jailbreak detection + bypass, TLS pinning detection + bypass, anti-debug detection + bypass, obfuscation analysis & deobfuscation, code-integrity / tamper-check bypass, hardcoded-secrets extraction, insecure local storage, and mobile network traffic analysis. Findings are proven from the artifact (decompilation or Frida trace), non-destructively. - agents.rs: new `mobile` Library category (loaded, counted). - pipeline.rs: run_mobile() mirroring the host pipeline with a mobile recon and headless tooling doctrine; exported from the crate. - CLI: `Cmd::Mobile` + `Mode::Mobile`, wired in main and the TUI. - README + TUTORIAL document the new test type; engagement-modes badge + table updated; "New in v4.2.0" note. Version bumped to 4.2.0 across the workspace. 383 tests. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
088d133c80 |
release: v4.1.0 — assurance layer, TypeSafe, hardening + benchmark
Version bumped to 4.1.0 across the workspace, binaries, web console and Typst template. README: new "New in v4.1.0" summary; trimmed the verbose highlight bullets and the TypeSafe section; removed the anti-plagiarism/provenance section (provenance stays in the code, just not front-and-centre in the README); TypeSafe promoted to its own top-level section; agent count 446. TUTORIAL: new section 17 "Assurance & authorization" covering the target gate, --scope-file, evidence-graded CVSS, audit anchoring + assurance bundle, sandbox, intercept proxy, PoC re-validation, compliance mapping, TypeSafe, and the internal/AD graph + budget governor. benchmarks/typesafe-2026-09-20/: the with/without TypeSafe measurement — report.html, scorer, both runs' findings/assurance/meta/logs, and a README. No secrets committed (env-only during the runs, verified clean). 381 tests. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
8894649ccb |
feat(scope): --scope-file YAML loader + web Scoping/Guardrails UI
Hard scoping was already enforced in code (every request passes
ScopePolicy::check_request; exclude beats allowlist; capability token caps
it; out-of-scope findings withheld + audited). What was missing was a way to
author that boundary from a file or the web form instead of only CLI flags.
- scope.rs: ScopePolicy::from_yaml / from_file — a dependency-free parser for
the friendly string format (app.example.com, *.wildcard, CIDR, url-prefix),
the same strings Pattern::parse already takes, NOT the raw serde {kind,value}
shape. Strict in one direction: an unreadable file errors, an empty hard list
authorizes nothing (a safe failure, but the operator's choice, not a typo).
- CLI: --scope-file <yaml>. Loaded before authorization so --in-scope adds to
it and the capability grant still caps it.
- Web: a full Scoping & Guardrails section in the Authorization tab — hard
scope, exclusions, observe-only, destructive-method + account-creation
toggles, max accounts, rate limit, forbidden payloads, notes. The server
materializes a scope YAML and passes --scope-file; notes stay labelled
"guidance, NOT enforced" so prose is never mistaken for a control.
- examples/scope.example.yaml documents the format.
End-to-end verified: web form -> YAML -> Rust loader -> enforced boundary.
332 tests (+4).
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
408350539f |
feat(budget,provenance): reasoning budget modes and JOASNSCOPE provenance
Budget (opt-in, unlimited by default so an un-budgeted run is unchanged):
- crates/harness/src/budget.rs — modes, phase shares, Token Governor
- CLI: --budget/--token-limit/--deep-test-limit/--coverage-first/
--depth-first/--sample-per-route; same controls in the web wizard
- pipeline honours it: vote_n narrows, evidence rounds are capped
Run control parity in the web console:
- /pause in the REPL, backed by a pause gate in the model pool: in-flight
agents finish, then the run holds with every finding kept
- POST /api/exploit/:id/{pause,continue,report} + GET .../log
Provenance (crates/harness/src/provenance.rs):
- JOASNSCOPE sigil leads every canary, so a marker found in a response,
a log or someone else's report extracts whole and names its build
- per-build fingerprint, per-run id, optional per-customer build id
- findings.json stamped with _engine; signed provenance.json manifest
- structural signature survives rewording but not a changed result set
- prompts watermarked at the single pool chokepoint
- `neurosploit provenance show|scan|verify`
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
481a4eb1b9 |
feat: wire capability tokens and the audit trail through CLI, REPL and web
The risk model, grants and hash-chained trail existed as modules nothing called. Now every engagement runs under them. Capability - `neurosploit capability issue|verify` mints and inspects grants. - `--capability-token` (global, so the REPL takes it too), `--in-scope`, `--environment`, `--policy` on `run`; verification happens at the command line, so an invalid grant fails with a readable message instead of halfway through an engagement. - The pipeline verifies before anything else and REFUSES to run on a token that does not verify — proceeding would mean acting on an authorization nobody can prove was issued. `effective_scope` then applies the grant as a ceiling. - Web: an Authorization tab carrying the token, extra hosts, environment and policy profile. The browser decodes the claims for display and says plainly that it is not verifying them — a "valid" badge from a party without the key would be the UI vouching for something it cannot check. A hole the smoke test found: `/inscope evil.test` inside a session under a grant WIDENED the scope past it — the one thing a capability token exists to prevent. The run itself would still have been constrained (the pipeline re-applies the grant), but `/policy` reported a boundary that was not real, and a tool that misreports its own limits is worse than one with none. Scope mutations now re-apply the ceiling and name what it refused. Session authorization also arrives from argv rather than a `/`-command, because a session that can widen its own grant is not constrained by one. Audit - One hash-chained record per action in `<run>/audit.jsonl`, in the specified shape, covering engagement start/end, validator rejections, findings that reach the report (with the hash of the evidence behind them) and findings withheld for being out of scope. - The run verifies its own chain at the end and says loudly if it is broken. - `/audit [n]` tails the trail and verifies it; the web offers it as a download next to the report, so "show me what the tool did" is a link. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
c3a6132153 |
feat(report): findings that answer where / why / how to fix, with a pasteable PoC
A finding is useful only if the reader can find the problem, see why it matters, fix it, and reproduce it without trusting us. The report answered the last one badly and the other three not at all: it printed a payload blob and an evidence blob, and "payload: ' OR 1=1--" tells a developer nothing about WHERE to look. A PoC script attached as a file is a black box unless you run it. Findings are now rendered in the order a reader works through them — where the problem is, what it means, how to fix it, then the proof — in the HTML report, the Markdown, the Typst/PDF and the web console's finding modal. The proof is numbered, pasteable steps: baseline request, attack request, how to read the result, with the real URL and the real payload. They come from the agent's `repro_steps` when it recorded them, and are derived from the endpoint/payload/identity pair otherwise, so every finding carries something runnable. The generated curl redacts Authorization/Cookie/API-key headers — a report gets shared, and a live session cookie inside one is a new bug. A PoC script is now offered as an extra artifact that automates the steps, never as the proof itself. Technical evidence is the measured difference, not a paraphrase: baseline vs attack status, size, timing and delta; how many repeats reproduced it; the controlled marker and whether a browser or a callback observed it; then each recorded exchange with the headers that decide a class (Location, Set-Cookie, Access-Control-*, X-Frame-Options, CSP, Retry-After) and a body excerpt. Finding gains `location` (the parameter/field/flow step, not just the URL) and `repro_steps`, and the agent contract now asks for them explicitly, along with impact tied to this app's data and remediation that names the control rather than saying "sanitise input". The web console offers the run's PDF when Typst produced one — and only then, since a dead download button is worse than none. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
4f277838c6 |
fix(web): responsive pass — real device sizes, scroll in the right containers
Audited at 390×844, 844×390 (phone landscape), 768×1024, 1024×768, 1440×900
and ≥1600px. What was actually broken:
- **The run header collapsed.** `justify-content: space-between` let the two
action buttons take the whole row, so at 390px the title wrapped one
character per line ("Test AspNe / t") and the facts line wrapped one word per
line. It now stacks below 680px, with the title block sized to its content
instead of stretching (the desktop `flex: 1` was what left a 180px void above
the buttons in the stacked layout).
- **The terminal header overflowed** its dock at 390px (557px of content in a
390px box) — it wraps now, and the status text drops out on narrow screens
where the coloured dot already carries it.
- **`100vh` is wrong on mobile.** It measures the viewport without the
collapsing address bar, so the wizard footer and its CTA sit underneath it.
Switched to `dvh` with the `vh` line kept as the fallback.
- **The dock took 82% of a phone in landscape** at its fixed 320px. It now
tracks the viewport (`clamp(180px, 42dvh, 340px)`, tighter still under
500px of height).
- **The off-canvas drawer had no way out but the button that opened it.**
Added a scrim that closes it, Esc, and auto-close when a run is picked —
and it closes itself if the window grows past the breakpoint, which
otherwise left a scrim over a sidebar that was no longer a drawer.
- **The stepper scrolls horizontally on a phone**, so advancing to an
off-screen step looked like nothing happened; the active step is scrolled
into view.
Device-type rules rather than width alone: `pointer: coarse` gets 38-44px hit
targets and 16px inputs (under 16px, iOS zooms the page on focus and breaks the
layout the user is typing into); `prefers-reduced-motion` drops the drawer
slide and the progress animation, which are decoration.
Scrolling stays where it belongs — one scroller per pane (`.wizard-body`,
`.run-body`, `.dash-body`, `.sb-groups`, `.modal-body`, `.term-host`), wide
tables scroll inside `.table-wrap`, and the page itself never scrolls
horizontally at any tested size.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
9d83cb6e30 |
feat: attack knowledge graph, layered memory, command rectification, FAIR dashboard
Backend ------- - knowledge_graph.rs — the durable structure under attack_graph's per-run view: typed entities (asset/endpoint/weakness/technique/finding/account/credential/ impact) joined by typed, weighted, provenance-carrying edges, accumulated across runs in .neurosploit/graph.json plus a per-run copy the report and web console can draw. Answers what a finding list can't: ranked attack paths, and the frontier of entities observed but never proven — where chaining should look next. Agents only sometimes fill chains_from, so progression is also inferred between adjacent kill-chain stages; those edges are marked inferred, weighted lower, and drawn dashed, because presenting a hypothesis as evidence is the graph lying about itself. Secrets stay in the vault, never the graph. - memory.rs — four tiers scoped by lifetime, not importance: working (one run), engagement (one target), technique (one agent/CWE), reusable (generalized). Promotion is evidence-gated and needs independent evidence at each step: a claim repeated within a run becomes engagement knowledge; one confirmed across runs becomes technique knowledge; one that held on two DIFFERENT targets is generalized into a reusable lesson with host-specific tokens stripped. Nothing is promoted on a single observation, which is exactly what a hallucination looks like. Recall is scored (overlap × past success × recency) and injected into recon/exploit prompts as leads to verify. Recalled memos are credited only when the run they informed actually found something. - rectify.rs — a mistyped command cost a full round trip through /help, at the worst possible moment during a live run. Accepted-as-typed wins over everything (so the /url alias is never "corrected" to /ua), then unique prefix, then Damerau-Levenshtein with a length-scaled budget, and a tie is reported rather than resolved. Arguments too: a bare host gets its scheme, an out-of-range count is clamped with a note instead of silently reverting, a near-miss model id is matched against the live catalog. - pool.rs — when every configured model is exhausted or its token is dead, try whatever else this machine can actually reach (an installed CLI subscription, or a provider whose key is in the environment) before parking. A run that stops on a box with three other usable backends stopped for no reason. - repl.rs — /memory, /forget, /graph; a recovered run resumes by itself where nobody is watching (piped stdin — the web console — or NEUROSPLOIT_AUTO_RESUME), since a `/continue` prompt there waits forever. Web --- - Attack path: the stage list was seven hardcoded values, so findings the harness staged outside it were silently dropped — 5 of 27 on a real run. Rewritten against the harness's own stage list with unknown stages kept, two-line labels (every node used to read "SQL Injection Authent…"), stage column headers, pan/zoom/fit, path highlighting, severity filter, and the run's graph.json used when present. - Dashboard: coverage, findings by severity, top weaknesses, and annualized loss exposure via FAIR — frequency from exploitability × validation confidence, magnitude from assumptions shown on screen and editable, reported as a range. The posture score saturates instead of subtracting, so it keeps discriminating past the first critical. - Run history groups into one folder per target with a filter, instead of one flat list that grows forever. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BvdGy9XtVWSdXDTa3FFLJv |
||
|
|
0ef0ce8d94 |
feat(web): xterm.js terminal dock + front-end QA pass
Replaces the floating REPL drawer with a docked terminal, and fixes the usability problems a screenshot audit of the console turned up. Terminal (the reason for the change): - The drawer rendered the harness into a <div>, so the server had to strip ANSI before sending it: colour, the box-drawn /status panel and the banner all arrived flattened, and long lines rewrapped mid-glyph. The stream is now sent verbatim and rendered by xterm.js (vendored, nothing fetched at runtime), decoded with a streaming UTF-8 decoder so a multi-byte character split across two reads survives. - The drawer floated bottom-right, directly over "Next →" and "Start Exploitation" — the wizard's primary buttons. The dock is a flex child of .main, so opening it shortens the view instead of covering it. Drag its top edge to resize; the height is remembered. - The child is spawned over a pipe, not a PTY, so it never echoes: line editing is local — echo, ←/→, Home/End, history, Tab completion over the slash commands, Ctrl+C/L/U/K/A/E. Ctrl-C is delivered as SIGINT by the server, since a raw 0x03 byte over a pipe interrupts nothing. - A target picker switches the terminal between a standalone REPL session and the engagement currently running, so mid-run instructions go to the same process doing the testing. QA fixes: - Findings tables sorted by severity (a LOW above a CRITICAL made a 27-row result unreadable), with sortable headers, a severity summary that doubles as a filter, a text filter, a sticky header, and horizontal scroll confined to the table instead of the whole page. - alert()/prompt() replaced by inline field errors, a custom-lead modal and toasts — a modal alert hid the very field it was complaining about. - Lead categories start collapsed (412 leads over ~30 categories); search auto-expands what it matches and shows per-category hit counts. - Sidebar rows truncate inside the rail (a long target URL used to spill past its border), and carry a worst-severity dot, finding count and age. - Past-run header shows when it ran, how many agents ran, the recon asset, PoC count and run id — two runs of one target were indistinguishable. - Off-canvas sidebar below 768px had no way to be opened; added the toggle. - Long evidence values (cookies, tokens) now wrap instead of running under the finding modal's edge. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BvdGy9XtVWSdXDTa3FFLJv |
||
|
|
4fbe608a7a |
feat(web): drive run/whitebox/greybox exploitation through a real REPL session
Root cause of "can't send prompts while a run streams": /api/exploit spawned a plain `neurosploit run ...` subprocess, and that CLI path (run_mode() in main.rs) never reads stdin - it only waits on the task or Ctrl-C. The ONLY thing in the harness that keeps accepting input while an engagement streams is the interactive REPL's background-run loop. So: - New startJobViaRepl(): for mode run/whitebox/greybox, spawns a bare `neurosploit` REPL session and scripts it via stdin (/target or /repo, /model, /sub, /mcp, /votes, /chain, /recon, /focus, /objective, /scope-out, /creds, /only <agents> or /only clear, then /run) instead of building CLI args. Same underlying pipeline, same tagged output lines, so all existing parsing (findings/phase/progress/runId) works unchanged. host/aitest/skills modes stay on the old one-shot startJob() - they need onboarding's scope picker, an interactive arrow-key menu that silently skips itself over a piped stdin, so they can't be scripted this way. - New POST /api/exploit/:id/input writes a line to the session's stdin - natural language, /status, /continue, anything the REPL accepts - and the live run view grows a "send prompt" box (in the Activity log tab) for it, shown only when the job reports interactive: true. - Stop, for an interactive job, now sends the REPL's own graceful '/stop\n1\n' (validate what's found, then report) instead of SIGINT - the REPL's own input loop has no signal handler, so SIGINT there would just kill the process outright and skip the report step. Non- interactive jobs still get SIGINT (run_mode() does catch that). - 'done' can no longer be process-exit only: an interactive session stays open after the engagement finishes (for /report, /continue, another /run), so ingestLine() now also flags done from the same "phase complete" content signal it already used for the phase field. Verified end-to-end: started an interactive job, confirmed `interactive: true` and a captured runId, sent /status and /agents mid- and post-run over the new /input endpoint (both accepted, session stayed alive and responsive after completion), and confirmed a non-interactive run is unaffected. Also: the missing "Activity log" tab a screenshot showed for a "running" engagement was the sidebar's detail-view fallback (2 tabs, no log) for a run whose Job object no longer exists in server memory - it happens when the Node process gets restarted while a spawned neurosploit child is still alive underneath it (an orphan from testing across many redeploys this session, not a code bug); the live view itself always had the tab. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0129WdYHccPsH27k5GGuwijd |
||
|
|
8047c66e8f |
fix(web): finding modal — dedupe repeated attribution text, prose vs code sections
impact and business_impact often carry identical text (both ending with the reporter's 'Identified and validated by NeuroSploit...' footer), and the modal concatenated them verbatim — the boilerplate line rendered twice, and whenever the two fields matched, so did the whole paragraph. - Strip the attribution sentence out of impact/business_impact/remediation/ evidence wherever it appears; surface it once, at the bottom of the modal, instead of embedded per field. - Skip business_impact entirely when it's identical to impact (the common case) instead of printing the same paragraph twice. - Split rendering into codeBlock() (endpoint/payload — monospace, looks like what it is: a request/curl) and proseBlock() (description/impact/ remediation — a readable paragraph, not a code box). Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0129WdYHccPsH27k5GGuwijd |
||
|
|
a083990ce4 |
fix(web): attack-graph canvas follows the app theme instead of forcing dark
Node/edge colors now use the same --sev-*-fg / --surface / --text / --border CSS custom properties as the rest of the console (set via inline style= attributes, since SVG presentation attributes don't resolve var()) — the graph reads correctly in light mode instead of always rendering as a dark canvas. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0129WdYHccPsH27k5GGuwijd |
||
|
|
7f365b4e88 |
feat(web): render the attack path as a real node graph, not flat cards
The 'Generative Attack Path Chaining' tab previously showed kill-chain stages as stacked cards in columns — with 1 finding (the common case early in a run) it looked like an empty list, nothing like an attack graph. Rewritten as an inline SVG node/edge graph on a fixed-dark canvas (matches attack-graph tools like NodeZero regardless of the app's own light/dark theme — bright severity colors read better against near-black): - Root node = the target, always present. - One node per confirmed finding, positioned in its kill-chain-stage column (falls back to a single flat column when no finding has a stage yet). - Edges: from the finding's chains_from parent when the harness set one, else fanned directly from root — never invents a specific relationship that doesn't exist in the data. - Per-node icon inferred from title/evidence/cwe/stage (key/shield/ person/host/db/impact), severity-colored border + corner tick. - Nodes are clickable — opens the same finding detail modal as the findings table (PoC included). Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0129WdYHccPsH27k5GGuwijd |
||
|
|
d42e9ff8e8 |
fix(web): global [hidden] bug, progress bar, F5 persistence, finding detail + PoC
Real front-end bugs found and fixed:
- [hidden] never worked on any element whose class also sets 'display'
(every .btn, .chip, ...): the browser's built-in '[hidden]{display:none}'
rule and an author rule of equal specificity tie, and the later one in the
cascade wins — so 'Next' stayed visible on the Review step alongside
'Start Exploitation', and 'Open report'/'Stop' rendered during 'starting'.
Fixed with a single global '[hidden]{display:none!important}' override.
- Progress bar was functionally correct but easy to miss (thin, 0%-width,
low-contrast track) and gave no feedback while the agent count is still
unknown (recon phase). Added a border for visibility and an indeterminate
sliding-segment state for the 'agents: ?' window.
- A live run watched in the browser was lost on F5 (jumped back to the
wizard) even though the job keeps running server-side. The active job id
now persists in localStorage; on load the app reconnects the SSE stream
(the server replays its full event buffer) instead of losing the view.
New:
- Findings are now clickable — a detail modal shows every Finding field
(CWE/CVSS/OWASP/MITRE/stage/exploitability/confidence/votes/review status/
auth context/account/agent), endpoint+payload, evidence, impact, business
impact, remediation, and chains_from — in both the live run and past-run
detail views.
- PoC surfacing: the finding modal looks up any script the run wrote to
pocs/ that's cited in the finding's evidence (per the harness's own
doctrine — see pipeline.rs change below), fetches and previews it inline,
with a link to open the raw file. Live runs poll for new PoC files every
5s once the run id is known.
- Pinned-leads confirmation: the live run header now states plainly how
many leads were pinned (and their names) or that selection is auto
(recon-driven) — this was previously buried in the scrolling activity log
behind the harness's unconditional 'Loaded 435 agents' library-size line,
which describes the full agent library, not what will actually run.
Harness doctrine (crates/harness/src/pipeline.rs, pocs_line()):
PoC-writing for black-box findings was previously conditioned on 'when an
issue needs a custom multi-step exploit/script' — vague enough that a
straightforward finding (single-request XSS/SQLi/IDOR) often got no PoC
file at all. Now required for every confirmed Medium+ finding, one
standalone .py/.sh script per finding, and explicit about citing the exact
file name in the finding's evidence field (which is what the web UI now
matches on to link a PoC to its finding).
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0129WdYHccPsH27k5GGuwijd
|
||
|
|
040170e54a |
fix(web): category master switch reads as partial, not off, when only some agents are deselected
Unchecking one agent in an 8-agent category (7/8 left on) rendered the category header switch fully unchecked — visually indistinguishable from 'category disabled', even though 7 of 8 agents were still on. The checkbox's checked state only had two positions; a partial selection had nowhere to render but off. Fix: set the master switch's .indeterminate property when 0 < selected < total, with its own CSS state (grey track, thumb parked halfway) instead of the on/off track+thumb. Clicking a checkbox out of indeterminate selects everything, per browser default — unchanged behavior, just an honest visual for the in-between state. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0129WdYHccPsH27k5GGuwijd |
||
|
|
bb659412fc |
feat(web): engagement wizard, model/auth picker, Auth & Keys menu, attack-path graph
Full frontend rewrite following a deliberate visual direction (dense security-operations console — borders over shadows, two radii, one accent, no gradients/glassmorphism) and fixing real bugs found in review: - EventSource on the exploit stream never called es.close() on 'done', so the browser silently reconnected and re-streamed the whole job (duplicate log lines/findings). Fixed. - Sidebar 'running' step indicator and openRun() matched ANY running run instead of the one belonging to the current job (by runId). Fixed. New: - 5-step engagement wizard (Asset -> Scope & Auth -> Leads -> Model & Run -> Review) replacing the single flat board — inspired by the Discovery/Plan/Exploit/Remediate stage model both a.security and terra.security use publicly. - Model is now a real dropdown sourced from /api/providers (mirrors harness::models::providers()), with an API-key vs. subscription toggle that disables subscription for API-only providers. - One Auth & Keys menu: target auth header + named roles (IDOR/BOLA/BFLA multi-identity testing) materialize into an ephemeral creds.yaml passed via --creds; per-provider API keys live in server memory only (never on disk) and are merged into every spawned child's env. - Generative Attack Path Chaining: findings rendered as kill-chain columns (recon -> initial-access -> ... -> impact) with chains_from resolved to parent titles, live in the run view and static in run detail. - Findings are now a proper table (severity/title/endpoint/CWE/agent/ confidence) instead of stacked cards. - Explicit light/dark theme toggle persisted in localStorage, defaulting to light (previously light only won when the OS wasn't in dark mode). - All UI strings in English. Backend additions: GET /api/providers, GET/POST/DELETE /api/keys, ephemeral creds.yaml generation for auth/roles, env override merged into every exploit-job and REPL child spawn. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0129WdYHccPsH27k5GGuwijd |
||
|
|
d1d1c71e24 |
feat(4.0.0): web console — lead board + live findings + real CLI REPL
New web/ app (zero npm deps, Node http built-ins only):
- server.js reads agents_md/ to build a categorized lead board (435 agents
auto-classified into Business Logic / Broken Access Control / Injection /
LLM Application / Auth & Session / SSRF / API / Cloud & Infra / etc.),
reads runs/ for history, and spawns the compiled neurosploit CLI binary
for every exploitation job — structured findings/phase/progress are parsed
from its stdout (finding_json:/phase lines), same signal the TUI uses.
- REPL drawer spawns `neurosploit` with no subcommand (real interactive
session, Reader::Plain over the piped stdin) and streams stdin/stdout —
every /command works exactly as in a terminal, nothing reimplemented.
- SSE endpoints for both job and REPL streams; run/finding/report assets
served under /api/runs/:id/asset/*.
- public/{index,app.js,style.css}: lead board with category toggles + custom
leads + Start Exploitation, live run view (progress/findings/log), run
detail view, REPL drawer — screenshot-inspired layout.
- web/API.md: full endpoint reference. web/README.md: quick start.
Bump version 3.6.9 -> 4.0.0 (Cargo.toml, CLI banners, README/TUTORIAL).
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0129WdYHccPsH27k5GGuwijd
|