Compare commits

..
Author SHA1 Message Date
george-keygraphandClaude Opus 5 cbf2019e7f docs: add a "Why 'Shannon'?" section to the README
Adds a short naming section under "What is Shannon?", directly after
"Why Shannon Exists", and syncs the same text into llms-full.txt.

Frames pentesting as an information problem, which is the actual link to
Claude Shannon, and closes on the "Hey Claude, run Shannon" line.

Actioned from the Aug 21 Shannon 3.0 launch planning session.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-31 15:48:36 -07:00
george-keygraphandClaude Opus 5 692440135a docs: mark native Windows as community-supported, WSL2 behaves like Linux
Per the Aug 27 CI/CD integration sync: WSL2 stays fully supported (it
behaves like Linux); native Windows moves from "not supported" to
community-supported, meaning contributions are welcome but Keygraph does
not actively develop or test against it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-31 10:19:30 -07:00
George Flores fabf56e574 Add files via upload 2026-08-31 10:19:30 -07:00
George Flores 85a29e2ea3 Delete assets/Shannon3GIF.gif 2026-08-31 10:19:30 -07:00
George Flores 82436dafe6 Rename Timeline 1.gif to Shannon3GIF.gif 2026-08-31 10:19:30 -07:00
George Flores 8a3fbc5286 Delete assets/Shannon3GIF.gif 2026-08-31 10:19:30 -07:00
George Flores d9748355e3 Add files via upload 2026-08-31 10:19:30 -07:00
George Flores 9a5f201f15 Rename Timeline 1.gif to Shannon3GIF.gif 2026-08-31 10:19:30 -07:00
George Flores 945fd2ce57 Delete assets/Shannon3GIF.gif 2026-08-31 10:19:30 -07:00
George Flores 4b26e677dd Add files via upload 2026-08-31 10:19:30 -07:00
George Flores c8ee22f2d1 Rename Shanon3GIF.gif to Shannon3GIF.gif 2026-08-31 10:19:30 -07:00
George Flores 529f6a7a95 Update README.md 2026-08-31 10:19:30 -07:00
George Flores 5a43c6127f Add files via upload 2026-08-31 10:19:30 -07:00
George Flores 9f7a349715 Update README.md 2026-08-31 10:19:30 -07:00
george-keygraphandClaude Opus 4.8 d017d74519 feat(cli): sunset wordmark and section hierarchy in terminal output
Add a section-chrome module (chrome.ts) and a log decorator
(log-render.ts) that give the CLI a coherent visual hierarchy beneath
the SHANNON splash, all drawn from the existing sunset ramp:

- splash.ts now sources the ramp from chrome.ts (one definition)
- start's launch info gets a solid-yellow rule under each section
  label, with aligned Label: values greyed
- streaming scan logs (logs, and start --follow) get one per-phase
  gutter bar that walks the ramp, dimmed timestamps, red on [ERROR],
  and the end-of-run summary framed in a corner panel

Named chrome.ts to avoid the existing ui.ts (spinner/step helpers).
Presentation only: the worker and workflow.log are untouched, the log
on disk stays plain text (tail/grep and the failure marker keep
working), and non-TTY / NO_COLOR output is byte-identical to before.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-08-31 10:19:30 -07:00
ezl-keygraph 6108de3cfc feat: bump pi harness to 0.84.2 to enable xAI subscription auth (#435) 2026-08-28 21:20:13 +05:30
ezl-keygraph 7e0464bf79 feat: support pentests with xAI (Grok) subscription auth (#434) 2026-08-28 21:11:24 +05:30
ezl-keygraph dc2a4fe4e8 feat: brand the npm page, CLI output, and reports (#432)
* docs: rebuild the npm package README on the main README's identity

* docs: point the README banner fallback at an asset that exists

* docs: declare the npm package author, homepage, and issue tracker

* docs(cli): retire "Framework" and settle on the canonical product line

* docs: describe the banner image in alt text instead of repeating the lockup

* feat(cli): print a plain-text banner when stdout is not a terminal

* feat(report): attribute the markdown report from a shared brand constant

* feat(cli): frame the plain-text banner with rules and split the version line

* docs: drop the URL from the npm author field
2026-08-27 18:53:56 +05:30
ezl-keygraph ed5659e2e2 fix(report): emit SARIF by default for exploit runs (#431)
* fix(report): emit SARIF by default for exploit runs, opt out with report.sarif: false

* docs: describe SARIF as on-by-default for exploit runs
2026-08-26 19:25:16 +05:30
ezl-keygraph f64a30040e ci: publish npm and beta via OIDC trusted publishing (#430) 2026-08-25 00:12:58 +05:30
ezl-keygraph b13788d8ef fix: terminate failed scans in Temporal and surface the reason when following (#429)
* fix(cli): skip splash screen off a TTY (e.g. CI)

* fix: terminate failed scans in Temporal and surface the reason when following

* fix(cli): indent embedded newlines within failure-error segments

* fix(worker): omit the Agent Breakdown section when no agents completed

* fix(cli): don't reprint the failure reason when the log already showed it

* fix(worker): indent embedded newlines within the workflow.log error block
2026-08-24 20:04:47 +05:30
George Flores 53118c6203 Merge pull request #427 from KeygraphHQ/docs/ci-sarif-common-questions
README update
2026-08-19 18:33:16 -07:00
George FloresandClaude Opus 5 af1ed2a563 README update
Documentation pass over the README and supporting docs, incorporating the
Aug 19 review with Parathan.

README:
- Dark/light banner and Discord/Keygraph buttons via <picture>
- Add a Common Questions section at the bottom of the page
- State one consistent position on model support and provider breadth
- Name the OpenAI Responses API alongside Chat Completions
- Frame local and self-hosted models as technically supported but not
  recommended, since capability varies once the harness opens every
  provider and model
- Describe SARIF as machine-readable output rather than a CI feature

Docs:
- ai-providers: drop the Claude-preference claim; explain that capability
  varies and the model should be evaluated against your own targets
- configuration: correct rating semantics stale since v2.2.0, since
  severity is now recorded in both exploitative and analysis-only runs
- safety: reframe the model-support caveat in the same terms
- worker: correct the stale rationale on the SARIF analysis-mode gate

CI/CD documentation is intentionally omitted until the GitHub Marketplace
action lands, so the README does not ship a hand-rolled npx wrapper that
is about to be replaced.

llms.txt and llms-full.txt regenerated from source, with one deliberate
exception: the "Is Shannon free?" and "Is Shannon free for startups and
nonprofits?" questions are kept in the llms-full.txt copy of the README
but not in the README itself. That section exists for agents, so a naive
regeneration of llms-full.txt would drop them; re-add them if you rebuild
the file from source.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-19 18:26:34 -07:00
ezl-keygraph 12d1c48a78 fix(cli): align usage command column in help output (#426) 2026-08-19 20:07:54 +05:30
ezl-keygraph dfb7c69d3b fix(cli): show splash screen on bare invocation and setup (#425) 2026-08-19 19:42:24 +05:30
ezl-keygraph d41ae9c20d feat(cli): overhaul commands and add live scan status (#424)
* refactor(cli): list workspaces natively instead of via the worker image

* feat(cli): preflight that Docker is installed and running

* feat(cli): stop scans by workspace or --all, terminating their Temporal workflows

* fix(worker): abort the running agent on cancellation so Temporal cancel takes effect

* refactor(cli): split destructive teardown out of stop into a reset command

* refactor(cli): centralise flag parsing and confirmation across commands

* fix(cli): pass provider credentials to docker by name to keep secrets out of argv

* feat(cli): add per-command help via <command> --help/-h and help <command>

* feat(cli): replace raw docker output with clack spinners for infra and scan teardown

* fix(cli): verify scan stop by re-querying container and workflow state instead of assuming success

* fix(cli): resolve running state before prompting on stop and report no-op stops honestly

* refactor(cli): show splash first and drive start with one spinner resolving to a clean line

* fix(cli): validate --url up front so a bad value fails cleanly instead of a late crash

* refactor(cli): centralize error reporting with fail() for expected errors and a crash handler that logs the stack and links the issue tracker

* feat(cli): add --json/--plain machine-readable output to workspaces and status

* refactor(cli): remove the workspaces command

* refactor(cli): remove the status command

* feat(cli): add 'progress <workspace>' — live scan progress from Temporal

* fix(cli): mark metric-less agents as skipped in progress, not done

* feat(cli): animate running agents in progress with a clack-style spinner

* feat(cli): rename progress->status, reveal agents as they run, show live per-agent elapsed

* fix(cli): mark passed-over phases as skipped live, not pending

* style(cli): rename status footer 'Wall-clock' to 'Time Taken', drop the parenthetical

* style(cli): drop '(sum of agents)' from status total cost line

* style(cli): green filled circle for completed, Shannon gold for running

* style(cli): use Shannon gold in place of green in status

* feat(cli): suggest closest command or flag on typo

* refactor(cli): single-source start help and drop ./repos bare-name shortcut

* feat(cli): name providers and fix in multi-provider credential error

* feat(cli): support --flag=value syntax and expand leading ~ in paths

* refactor(cli): centralize ANSI color codes in colors.ts

* feat(cli): add scans command listing completed scans with cost and duration

* fix(cli): keep stdout clean off-TTY for logs and start

* feat(cli): add repo link to top-level help

* feat(worker): record auth-validation metrics and register resume attempts early

* refactor(cli): share resume-aware workflow-id resolution and surface root-cause failures

* feat(cli): add status --json, auth phase, dashboard link, and stable live redraw

* refactor(cli): drop cost from status and scans output

* feat(worker): surface both PDF and markdown report at run root

* refactor(cli): normalize error/warning prefixing through fail and warn

* feat(cli): add version --json for machine-readable output

* refactor(cli): rename start --debug to --keep-container

* refactor(cli): point start's progress hint at status instead of the Temporal dashboard

* refactor(cli): centralize the mode-aware command prefix

* refactor(cli): trim start and logs output to durable facts off-TTY

* feat(cli): require typed confirmation for reset instead of --yes

reset permanently wipes all Temporal data and volumes — a severe,
irreversible action. Replace its default y/N confirm (bypassable with
--yes) with a typed-word confirmation that has no bypass, so the wipe
can only be triggered by a deliberate interactive answer.

* feat(cli): surface logs and status hints after start on a TTY

* feat(cli): exit 2 on usage errors, distinct from operational failures

* feat(cli): add start --follow to stream logs and exit on scan outcome

* refactor(cli): redesign splash with sunset-gradient wordmark and truecolor

* refactor(cli): remove the uninstall command

* docs: sync CLI docs with removed uninstall/workspaces, new scans and --follow

* docs: fix reset confirmation — typed confirm, not --yes/-y

* style(cli): restructure status footer with divider, aligned Logs/Temporal rows

* feat(cli): show splash in the status command

* fix(worker): validate auth-state shape, not entry count

* docs: correct reset confirmation and add markdown report to run-root docs
2026-08-18 15:46:25 +05:30
ezl-keygraph 1ae0a142f8 feat(worker): render PDF security reports via Typst (#421) 2026-08-12 15:06:43 +05:30
ezl-keygraph d4cc2ab974 feat: support pentests with Codex subscription auth (#419) 2026-08-10 15:27:19 +05:30
90 changed files with 4889 additions and 1055 deletions

No files matched your search

+10
View File
@@ -48,3 +48,13 @@ SHANNON_AI_MODEL=anthropic:claude-sonnet-4-6
# --- Misc --------------------------------------------------------------------
# Forward /etc/hosts entries into the worker container.
# SHANNON_FORWARD_HOSTS=false
# See the guide below to use an OpenAI subscription
# https://github.com/KeygraphHQ/shannon/blob/main/docs/ai-providers.md#openai-codex-chatgpt-pluspro-subscription
# SHANNON_USE_PI_AUTH=1
# SHANNON_AI_MODEL=openai-codex:gpt-5.5
# Or the guide below to use an xAI subscription
# https://github.com/KeygraphHQ/shannon/blob/main/docs/ai-providers.md#xai-grok-subscription
# SHANNON_USE_PI_AUTH=1
# SHANNON_AI_MODEL=xai:grok-4.6
-2
View File
@@ -189,8 +189,6 @@ jobs:
- name: Publish npm package
working-directory: apps/cli
env:
NODE_AUTH_TOKEN: ${{ secrets.NPM_TOKEN }}
run: |
if npm view "@keygraph/shannon@${{ needs.preflight.outputs.version }}" version 2>/dev/null; then
echo "Version already published, skipping"
-2
View File
@@ -201,8 +201,6 @@ jobs:
- name: Publish npm package
working-directory: apps/cli
env:
NODE_AUTH_TOKEN: ${{ secrets.NPM_TOKEN }}
run: |
if npm view "@keygraph/shannon@${{ needs.preflight.outputs.version }}" version 2>/dev/null; then
echo "Version already published, skipping"
+2
View File
@@ -5,3 +5,5 @@ credentials/
dist/
repos/
.turbo/
.DS_Store
+20 -18
View File
@@ -44,8 +44,8 @@ echo "ANTHROPIC_API_KEY=your-key" > .env
./shannon build
# Run
./shannon start -u <url> -r my-repo
./shannon start -u <url> -r my-repo -c ./apps/worker/configs/my-config.yaml
./shannon start -u <url> -r ./my-repo
./shannon start -u <url> -r ./my-repo -c ./apps/worker/configs/my-config.yaml
./shannon start -u <url> -r /any/path/to/repo
```
@@ -56,25 +56,24 @@ echo "ANTHROPIC_API_KEY=your-key" > .env
npx @keygraph/shannon setup
# Workspaces & Resume
./shannon start -u <url> -r my-repo -w my-audit # New named workspace
./shannon start -u <url> -r my-repo -w my-audit # Resume (same command)
./shannon workspaces # List all workspaces
./shannon start -u <url> -r ./my-repo -w my-audit # New named workspace
./shannon start -u <url> -r ./my-repo -w my-audit # Resume (same command)
# Monitor
./shannon logs <workspace> # Show a scan's live log
./shannon status # Show running scans
./shannon status <workspace> # Live phase/agent progress of one scan, read from Temporal (redraws, then exits)
# Dashboard: http://localhost:8233
# Stop
./shannon stop # Preserves scan data
./shannon stop --clean # Full cleanup including volumes (confirms first; --yes/-y to skip)
./shannon stop <workspace> # Stop one scan (confirms first; --yes/-y to skip)
./shannon stop --all # Stop all running scans (Temporal stays up; confirms first)
./shannon reset # Stop everything and wipe all Temporal data + volumes (type 'confirm' to proceed; cannot be skipped)
# Version
./shannon version # npx: package version; local: git SHA
# Image management
./shannon build [--no-cache] # Local mode: build worker image
npx @keygraph/shannon uninstall # npx mode: remove ~/.shannon/ (confirms first; --yes/-y to skip)
# Build TypeScript (development)
pnpm run build # Build all packages via Turborepo
@@ -85,7 +84,7 @@ pnpm biome:fix # Auto-fix lint, format, and import sorting
**Monorepo tooling:** pnpm workspaces, Turborepo for task orchestration, Biome for linting/formatting. TypeScript compiler options shared via `tsconfig.base.json` at the root. All packages extend it, overriding only `rootDir` and `outDir`. Shared devDependencies (`typescript`, `@types/node`, `turbo`, `@biomejs/biome`) are hoisted to the root workspace.
**Options:** `-c <file>` (YAML config), `-o <path>` (output directory), `-w <name>` (named workspace; auto-resumes if exists), `--pipeline-testing` (minimal prompts, 10s retries), `--debug` (preserve worker container after exit for log inspection), `--yes`/`-y` (skip the confirmation prompt on `stop --clean`/`uninstall`; required for non-interactive use)
**Options:** `-c <file>` (YAML config), `-o <path>` (output directory), `-w <name>` (named workspace; auto-resumes if exists), `--pipeline-testing` (minimal prompts, 10s retries), `--keep-container` (preserve worker container after exit for log inspection), `--yes`/`-y` (skip the confirmation prompt on `stop`; required for non-interactive use; `reset` requires a typed `confirm` and cannot be skipped)
## Architecture
@@ -97,9 +96,11 @@ apps/worker/ — @shannon/worker (private, Temporal worker + pipeline logic)
```
### CLI Package (`apps/cli/`)
Published as `@keygraph/shannon` on npm. Contains only Docker orchestration logic — no Temporal SDK, business logic, or prompts. Bundled with tsdown for single-file ESM output.
Published as `@keygraph/shannon` on npm. Contains Docker orchestration logic plus a read-only `@temporalio/client` reader (for `status`); no worker/pipeline business logic or prompts. Bundled with tsdown for single-file ESM output (deps stay external).
- `apps/cli/src/index.ts` — CLI dispatcher (`setup`, `start`, `stop`, `logs`, `workspaces`, `status`, `build`, `uninstall`, `version`)
- `apps/cli/src/index.ts` — CLI dispatcher (`setup`, `start`, `stop`, `reset`, `logs`, `status`, `build`, `version`)
- `apps/cli/src/temporal-client.ts` — `@temporalio/client` reader for `status`: connects to the frontend on `127.0.0.1:7233` (published by compose), `describeScan` (status + `pendingActivities` → running agents), `queryProgress` (live `getProgress` query → `PipelineState`), `getTerminalOutcome` (workflow `result()`). No worker of its own; scans are visible only within Temporal's ~24h retention (namespace default, unset in compose)
- `apps/cli/src/scan/` — `status` rendering: `pipeline.ts` (static phase/agent plan + `run*Agent` activity-type→agent map + mirrored `PipelineState`/`AgentMetrics` types; keep in sync with the worker), `render.ts` (one renderer for both the live query state and the terminal result)
- `apps/cli/src/mode.ts` — Auto-detection: local mode if `SHANNON_LOCAL=1` env var is set
- `apps/cli/src/docker.ts` — Compose lifecycle, image pull/build, ephemeral `docker run` worker spawning
- `apps/cli/src/home.ts` — State directory management (`~/.shannon/` for npx, `./` for local)
@@ -108,7 +109,7 @@ Published as `@keygraph/shannon` on npm. Contains only Docker orchestration logi
- `apps/cli/src/config/resolver.ts` — Cascading config (npx only): env vars → `~/.shannon/config.toml` (parsed with `smol-toml`)
- `apps/cli/src/config/writer.ts` — TOML serialization and secure file persistence (0o600)
- `apps/cli/src/commands/setup.ts` — Interactive TUI wizard (`@clack/prompts`) for provider credential setup (npx only)
- `apps/cli/src/paths.ts` — Repo/config path resolution (bare name → `./repos/<name>`, or any absolute/relative path)
- `apps/cli/src/paths.ts` — Repo/config path resolution (any absolute or relative path)
- `apps/cli/src/version.ts` — Version reporting (npx: `package.json` version; local: `git-<sha>`)
- `apps/cli/src/tty.ts` — Terminal capability detection: `requireInteractive` guard (fails fast off-TTY instead of hanging on a prompt), `supportsColor` color gating (`NO_COLOR`/`FORCE_COLOR`), and `stdoutIsTerminal` for spinner/cursor output
- `apps/cli/src/commands/` — Command handlers
@@ -151,12 +152,13 @@ Durable workflow orchestration with crash recovery, queryable progress, intellig
5. **Reporting** (`report`) — Executive-level security report
### Supporting Systems
- **Configuration** — YAML configs in `apps/worker/configs/` with JSON Schema validation (`config-schema.json`). Supports auth settings (MFA/TOTP), URL/code rule scoping (`rules.avoid`/`rules.focus`), run-scope steering (`vuln_classes`, `exploit`), free-form `rules_of_engagement`, and post-hoc `report` options (`min_severity`, `min_confidence`, `guidance`, and `sarif` to emit a SARIF 2.1.0 log via `apps/worker/src/services/sarif-renderer.ts`; exploit-only). `code_path` avoid rules are enforced via the `@gotgenes/pi-permission-system` extension: `apps/worker/src/temporal/activities.ts:syncCodePathDenyRules` writes a global `path` deny config once per workflow (`apps/worker/src/ai/pi/permission-system.ts:syncPermissionSystemConfig`), and the executor loads the extension when that config is present (`apps/worker/src/ai/pi/pi-executor.ts`), so denies fire across every tool and child `task` session. `vuln_classes`/`exploit` scope is locked into `session.json` on first run; resumes with a different scope fail fast (`persistOrValidateRunScope`). Credential resolution — local mode: env vars → `./.env`; npx mode: env vars → `~/.shannon/config.toml` (via `npx @keygraph/shannon setup`)
- **Configuration** — YAML configs in `apps/worker/configs/` with JSON Schema validation (`config-schema.json`). Supports auth settings (MFA/TOTP), URL/code rule scoping (`rules.avoid`/`rules.focus`), run-scope steering (`vuln_classes`, `exploit`), free-form `rules_of_engagement`, and post-hoc `report` options (`min_severity`, `min_confidence`, `guidance`, and `sarif` for a SARIF 2.1.0 log via `apps/worker/src/services/sarif-renderer.ts`, on by default for exploit runs and opt out with `report.sarif: false`). `code_path` avoid rules are enforced via the `@gotgenes/pi-permission-system` extension: `apps/worker/src/temporal/activities.ts:syncCodePathDenyRules` writes a global `path` deny config once per workflow (`apps/worker/src/ai/pi/permission-system.ts:syncPermissionSystemConfig`), and the executor loads the extension when that config is present (`apps/worker/src/ai/pi/pi-executor.ts`), so denies fire across every tool and child `task` session. `vuln_classes`/`exploit` scope is locked into `session.json` on first run; resumes with a different scope fail fast (`persistOrValidateRunScope`). Credential resolution — local mode: env vars → `./.env`; npx mode: env vars → `~/.shannon/config.toml` (via `npx @keygraph/shannon setup`)
- **Prompts** — Per-phase templates in `apps/worker/prompts/` with variable substitution (`{{TARGET_URL}}`, `{{CONFIG_CONTEXT}}`). Shared partials in `apps/worker/prompts/shared/` via `apps/worker/src/services/prompt-manager.ts`, including `_code-path-rules.txt` (focus/avoid `[FILE]`/`[GLOB]` routing) and `_rules-of-engagement.txt` (free-text engagement rules). When `exploit: false`, `apps/worker/src/services/findings-renderer.ts` deterministically converts each `*_exploitation_queue.json` into a `*_findings.md` for report assembly — no LLM in the loop
- **Agent Harness (pi)** — Uses the **pi harness** (`@earendil-works/pi-coding-agent`, requires Node ≥ 22.19) via `apps/worker/src/ai/pi/pi-executor.ts` (`runPiPrompt` → `createAgentSession`). Retry is split in `apps/worker/src/ai/pi/retry-settings.ts`: pi's agent-level loop is off so Temporal owns agent restarts, while `provider.maxRetries` stays on — pi reads the `provider` block independently of the `enabled` flag — so transport faults are absorbed in-session rather than costing a full agent re-run. `maxRetryDelayMs` is left at pi's 60s default. One model runs every phase, named by `SHANNON_AI_MODEL=<provider>:<model-id>` (default `anthropic:claude-sonnet-4-6`). `apps/worker/src/ai/models.ts` parses the spec — splitting on the **first** colon only, so Bedrock IDs keep theirs — and resolves it through pi's `ModelRuntime`. pi ships the `CredentialStore` interface but no in-memory implementation (its own reads `auth.json` from disk), so `RuntimeCredentialStore` in that file supplies one: credentials arrive as env vars in an ephemeral container and must never touch disk. `createModelRuntime(providerId, apiKey)` builds the runtime; `allowModelNetwork` stays at its default `false` so a scan never blocks on a catalog refresh. `resolveModelSelection()` is **async** because `ModelRuntime.create()` is. Any pi-ai provider id is accepted — `parseModelSpec` no longer rejects against a hardcoded list, so pi's registry is the authority (an unknown provider/model surfaces as a clear "not found in pi registry" error at preflight, which points to the browsable catalogue at `pi.dev/models` — `PI_CATALOG_URL` in `apps/worker/src/ai/models.ts`, appended to the not-found errors and shown in the setup wizard's "Other provider" hint). Four providers are **curated** (`CURATED_PROVIDERS`: `anthropic`, `openai`, `xai`, `amazon-bedrock`) with their own credential variables, config sections, and setup flows; each provider's API key env var is declared once in `PROVIDER_API_KEY_ENV` — Shannon uses each vendor's own variable name (`OPENAI_API_KEY`, `XAI_API_KEY`, …), never an invented one; Bedrock's entry is `AWS_BEARER_TOKEN_BEDROCK`, paired with `AWS_REGION`, which preflight requires separately as provider config rather than a credential. Any other provider uses the **generic** credential path: `SHANNON_AI_API_KEY` (`GENERIC_API_KEY_ENV`) supplies the key for any provider whose credential is a plain API key. Curated providers' own variables take precedence over it, and it also works as a fallback for them — Bedrock is the sole exception (it authenticates through its AWS_ variables, so the generic key never stands in for it). The CLI forwards `SHANNON_AI_API_KEY` in `COMMON_FORWARD_VARS` (it is provider-neutral, binding to whatever `SHANNON_AI_MODEL` names, so the "only one provider configured" guard counts only named credentials), and stores it under a generic `[provider]` config.toml section (`provider.api_key`). `npx @keygraph/shannon setup` exposes this as the "Other provider" option: free-text provider id + model id + key (a curated provider id is rejected there, since it has its own option). `SHANNON_AI_BASE_URL` overrides the endpoint for any provider (proxies/gateways); the credential is unchanged. `pointAtGateway` (`apps/worker/src/ai/models.ts`) applies the one dialect change: behind a base URL, `openai` follows `SHANNON_AI_OPENAI_FORMAT` (`chat-completions` default, or `responses`). On `chat-completions` it switches the API to `openai-completions` and drops the catalogue's Responses-shaped `compat` block so pi's `detectCompat` derives completions settings; on `responses` the descriptor is unchanged but for the endpoint. `resolveGatewayFormat` rejects the variable when the provider is not `openai` or no base URL is set, since it cannot take effect there. All other providers keep their API. The CLI mirrors the accepted values in `apps/cli/src/model-spec.ts`, forwards the variable in `COMMON_FORWARD_VARS`, and maps it to `openai.format` in config.toml. `buildEnvFlags` forwards only the selected provider's credential into the worker container. The CLI mirrors the parse rule and the provider/credential tables in `apps/cli/src/model-spec.ts` (it cannot import from the worker package); the two must stay in sync. pi ships no JSON-schema output or `Task`/`TodoWrite` built-ins, so structured queues are captured via a `submit_exploitation_queue` custom tool (`apps/worker/src/ai/queue-schemas.ts`), and `task` (child sessions scoped to `read`, `grep`, `find`, `ls`, `write`, and `bash` — no nested `task` or collector tools; `CHILD_TOOLS` in `apps/worker/src/ai/pi/task-tool.ts`) + `todo_write` (`apps/worker/src/ai/pi/session-tools.ts`) are provided as custom tools; the per-phase collectors are pi custom tools (TypeBox `defineTool` in `apps/worker/src/collectors/`). Shannon sets no thinking configuration at all — no `thinkingLevel` is passed to any `createAgentSession` call, so pi's own default applies. There Line truncated
- **Audit System** — Crash-safe append-only logging in `workspaces/{hostname}_{sessionId}/`. The run directory's top level holds only the human-facing report (`Security-Assessment-Report.md`, `FINAL_REPORT_FILENAME` in `apps/worker/src/paths.ts`); everything else — deliverables, per-agent logs, prompts, `session.json`, `workflow.log`, and browser artifacts — is nested under a hidden `.shannon/` internals dir (`INTERNAL_DIR`) so a customer sees only the report. Audit path helpers route through `generateInternalPath` (`apps/worker/src/audit/utils.ts`); the CLI nests the overlay backing dirs under the same `.shannon/` (`apps/cli/src/docker.ts`, `start.ts`). `session.json`/`workflow.log` reads use dual-read resolvers (`resolveSessionJsonPath`, `resolveRunFile`) that prefer `.shannon/` and fall back to the legacy run-root layout, so pre-restructure workspaces stay listable (`workspaces`/`logs`) without migration. Resuming a pre-restructure workspace upgrades it in place first: `migrateLegacyWorkspaceLayout` (`apps/cli/src/commands/start.ts`) renames the flat deliverables/logs/session entries into `.shannon/` (carrying the deliverables `.git` along) before the overlay dirs are mounted, so resume finds the old checkpoints instead of re-running every agent. The report is surfaced by copying the assembled `comprehensive_security_assessment_report.md` from the deliverables dir to the run root (`copyReportToRunRoot` in `apps/worker/src/services/reporting.ts`). WorkflowLogger (`apps/worker/src/audit/workflow-logger.ts`) provides unified human-readable per-workflow logs, backed by LogStream (`apps/worker/src/audit/log-stream.ts`) shared stream primitive
- **Pi Credential Reuse** — `SHANNON_USE_PI_AUTH=1` opts into reusing the host's Pi login, including an `openai-codex` ChatGPT Plus/Pro subscription (`SHANNON_AI_MODEL=openai-codex:<model-id>`) or an `xai` Grok subscription (`SHANNON_AI_MODEL=xai:<model-id>`); the mechanism is provider-agnostic and works for any Pi login. `apps/cli/src/env.ts` requires `~/.pi/agent/auth.json`; `start.ts` passes its path to `spawnWorker`, which mounts only that file read-write at `/tmp/.pi/agent/auth.json`. The flag itself is not forwarded: the worker detects the file with `piAuthPresent()` and passes its path to `ModelRuntime.create`. CLI and worker API-key presence checks are skipped on this path, but the normal preflight model probe still validates the credential. The image and UID-remapping entrypoint keep `/tmp/.pi/agent` owned by `pentest` so adjacent Pi/Shannon configuration remains writable. Refreshed OAuth state is persisted to the host for subsequent scans.
- **Audit System** — Crash-safe append-only logging in `workspaces/{hostname}_{sessionId}/`. The run directory's top level holds the human-facing report in both formats (`Security-Assessment-Report.pdf` and `Security-Assessment-Report.md`, `FINAL_REPORT_PDF_FILENAME`/`FINAL_REPORT_MD_FILENAME` in `apps/worker/src/paths.ts`); everything else — deliverables, per-agent logs, prompts, `session.json`, `workflow.log`, and browser artifacts — is nested under a hidden `.shannon/` internals dir (`INTERNAL_DIR`) so a customer sees only the report. Audit path helpers route through `generateInternalPath` (`apps/worker/src/audit/utils.ts`); the CLI nests the overlay backing dirs under the same `.shannon/` (`apps/cli/src/docker.ts`, `start.ts`). `session.json`/`workflow.log` reads use dual-read resolvers (`resolveSessionJsonPath`, `resolveRunFile`) that prefer `.shannon/` and fall back to the legacy run-root layout, so pre-restructure workspaces stay listable (`workspaces`/`logs`) without migration. Resuming a pre-restructure workspace upgrades it in place first: `migrateLegacyWorkspaceLayout` (`apps/cli/src/commands/start.ts`) renames the flat deliverables/logs/session entries into `.shannon/` (carrying the deliverables `.git` along) before the overlay dirs are mounted, so resume finds the old checkpoints instead of re-running every agent. The report agent writes structured findings to `report.json`, from which `report-renderer.ts` renders the assembled markdown and `report-json-adapter.ts` produces the Typst-shaped JSON that `pdf-renderer.ts` compiles into `comprehensive_security_assessment_report.pdf` using the bundled `apps/worker/templates/typst/report.typ` template (the `typst` binary is installed in the worker image). `copyReportToRunRoot` (`apps/worker/src/services/reporting.ts`) surfaces both the PDF and the markdown to the run root as `Security-Assessment-Report.pdf` and `Security-Assessment-Report.md`; the deliverables-dir copies remain as the git-checkpointed sources. PDF compilation is best-effort — a failure is logged and the run still completes. WorkflowLogger (`apps/worker/src/audit/workflow-logger.ts`) provides unified human-readable per-workflow logs, backed by LogStream (`apps/worker/src/audit/log-stream.ts`) shared stream primitive
- **Deliverables** — Saved to `.shannon/deliverables/` in the target repo via the `save-deliverable` CLI script (`apps/worker/src/scripts/save-deliverable.ts`)
- **Workspaces & Resume** — Named workspaces via `-w <name>` or auto-named from URL+timestamp. Resume detects completed agents via `session.json`. `loadResumeState()` in `apps/worker/src/temporal/activities.ts` validates deliverable existence, restores git checkpoints, and cleans up incomplete deliverables. Workspace listing via `apps/worker/src/temporal/workspaces.ts`
- **Workspaces & Resume** — Named workspaces via `-w <name>` or auto-named from URL+timestamp. Resume detects completed agents via `session.json`. `loadResumeState()` in `apps/worker/src/temporal/activities.ts` validates deliverable existence, restores git checkpoints, and cleans up incomplete deliverables
## Development Notes
@@ -246,9 +248,9 @@ Package managers are configured with a minimum release age (7 days). Requires pn
## Troubleshooting
- **"Repository not found"** — Pass a bare name (`-r my-repo`) for `./repos/my-repo`, or a path (`-r /path/to/repo`) for any directory
- **"Repository not found"** — Pass a path to the target repo (`-r /path/to/repo` or `-r ./my-repo`)
- **"Temporal not ready"** — Wait for health check or `docker compose logs temporal`
- **Worker not processing** — Check `docker ps --filter "name=shannon-worker-"`
- **Reset state** — `./shannon stop --clean`
- **Reset state** — `./shannon reset`
- **Local apps unreachable** — Use `host.docker.internal` instead of `localhost`
- **Container permissions** — On Linux, may need `sudo` for docker commands
+20 -2
View File
@@ -52,6 +52,8 @@ RUN apk update && apk add --no-cache \
curl \
ca-certificates \
shadow \
# Typst tarball decompression
xz \
# Language runtimes (minimal)
nodejs-22 \
npm \
@@ -73,6 +75,22 @@ RUN apk update && apk add --no-cache \
# Font rendering
fontconfig
# Install Typst (report PDF compilation)
ARG TYPST_VERSION=0.14.2
RUN case "$(uname -m)" in \
x86_64) TYPST_ARCH=x86_64-unknown-linux-musl ;; \
aarch64) TYPST_ARCH=aarch64-unknown-linux-musl ;; \
*) echo "unsupported arch $(uname -m)" && exit 1 ;; \
esac && \
mkdir -p /tmp/typst-dl /usr/local/bin && cd /tmp/typst-dl && \
curl -fsSL "https://github.com/typst/typst/releases/download/v${TYPST_VERSION}/typst-${TYPST_ARCH}.tar.xz" -o typst.tar.xz && \
xz -d typst.tar.xz && \
tar -xf typst.tar && \
mv "typst-${TYPST_ARCH}/typst" /usr/local/bin/typst && \
chmod +x /usr/local/bin/typst && \
cd / && rm -rf /tmp/typst-dl && \
typst --version
# Create non-root user
RUN addgroup -g 1001 pentest && \
adduser -u 1001 -G pentest -s /bin/bash -D pentest
@@ -107,12 +125,12 @@ RUN ln -s /app/apps/worker/dist/scripts/save-deliverable.js /usr/local/bin/save-
# Create directories for session data and ensure proper permissions
RUN mkdir -p /app/sessions /app/repos /app/workspaces && \
mkdir -p /tmp/.cache /tmp/.config /tmp/.npm && \
mkdir -p /tmp/.cache /tmp/.config /tmp/.npm /tmp/.pi/agent && \
chmod 777 /app && \
chmod 777 /tmp/.cache && \
chmod 777 /tmp/.config && \
chmod 777 /tmp/.npm && \
chown -R pentest:pentest /app /tmp/.claude
chown -R pentest:pentest /app /tmp/.claude /tmp/.pi
COPY entrypoint.sh /app/entrypoint.sh
RUN chmod +x /app/entrypoint.sh
+52 -10
View File
@@ -3,23 +3,26 @@
<div align="center">
<img src="./assets/github-banner.png" alt="Shannon - AI Pentester by Keygraph" width="100%">
# Shannon - AI Pentester by Keygraph
<picture>
<source media="(prefers-color-scheme: dark)" srcset="./assets/github-banner-dark.png">
<source media="(prefers-color-scheme: light)" srcset="./assets/github-banner-light.png">
<img src="./assets/github-banner-light.png" alt="Shannon, AI Pentester for Web Apps and APIs, by Keygraph" width="100%">
</picture>
<a href="https://trendshift.io/repositories/15604" target="_blank"><img src="https://trendshift.io/api/badge/repositories/15604" alt="KeygraphHQ%2Fshannon | Trendshift" style="width: 250px; height: 55px;" width="250" height="55"/></a>
Shannon is an autonomous, AI pentester for web applications and APIs. <br />
### Shannon is an autonomous, AI pentester for web applications and APIs.
It analyzes your source code, identifies attack paths, and executes real exploits to prove vulnerabilities before they reach production.
**This repository is Shannon Open Source: the full agent, run locally from your command line.**
---
<a href="https://discord.gg/9ZqQPuhJB7"><img src="./assets/discord.png" height="40" alt="Join Discord"></a>
<a href="https://keygraph.io/"><img src="./assets/Keygraph_Button.png" height="40" alt="Visit Keygraph.io"></a>
<a href="https://discord.gg/9ZqQPuhJB7"><picture><source media="(prefers-color-scheme: dark)" srcset="./assets/discord_button_dark.png"><source media="(prefers-color-scheme: light)" srcset="./assets/discord_button_light.png"><img src="./assets/discord_button_light.png" height="40" alt="Join Discord"></picture></a>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;<a href="https://keygraph.io/"><picture><source media="(prefers-color-scheme: dark)" srcset="./assets/keygraph_button_dark.png"><source media="(prefers-color-scheme: light)" srcset="./assets/keygraph_button_light.png"><img src="./assets/keygraph_button_light.png" height="40" alt="Visit Keygraph.io"></picture></a>
---
</div>
> [!TIP]
@@ -38,6 +41,7 @@ It analyzes your source code, identifies attack paths, and executes real exploit
- [License](#license)
- [About Keygraph](#about-keygraph)
- [Community and Support](#community-and-support)
- [Common Questions](#common-questions)
## What is Shannon?
@@ -53,10 +57,16 @@ Thanks to tools like Claude Code and Cursor, your team ships code non-stop. But
Shannon closes that gap by providing on-demand, automated penetration testing that can run against every build or release.
### Why "Shannon"?
It's named after Claude Shannon, the father of information theory. At its core, pentesting is an information problem: every probe reduces uncertainty about a system's state. The best tools maximize the signal gained from every request, turning those bits of knowledge into an exploit path.
Also, we wanted you to be able to say, "Hey Claude, run Shannon" to find all the security flaws in your vibe-coded app.
## Shannon in Action
<p align="center">
<img src="assets/shannon-action.gif" alt="Shannon running an autonomous pentest" width="100%">
<img src="assets/Shannon3GIF.gif" alt="Shannon running an autonomous pentest" width="100%">
</p>
Sample penetration test reports from intentionally vulnerable applications, produced by Shannon Open Source:
@@ -73,7 +83,7 @@ Sample penetration test reports from intentionally vulnerable applications, prod
- **Docker**: required for the worker container.
- **Node.js 18+**: required for the recommended `npx` workflow.
- **AI provider credentials**: Anthropic, OpenAI, xAI, or AWS Bedrock - or [any other provider](docs/ai-providers.md#any-other-provider). Claude models are recommended. For suggested model IDs per provider, plus gateways and custom base URLs, see [AI providers](docs/ai-providers.md#suggested-models).
- **AI provider credentials**: Shannon runs on Anthropic, OpenAI, xAI, AWS Bedrock, [any other provider](docs/ai-providers.md#any-other-provider) in the harness catalogue, and any endpoint that speaks the Anthropic Messages API or the OpenAI Chat Completions or Responses API through a [custom base URL](docs/ai-providers.md#custom-base-url). You bring your own key, and Keygraph never proxies your model traffic. Shannon is provider-agnostic. See [AI providers](docs/ai-providers.md#suggested-models) for suggested model IDs.
- **Cyber safeguards cleared with your provider**: Anthropic and OpenAI apply real-time safeguards to cyber-security workloads, which can interrupt a scan mid-run. Complete their guidance for legitimate security testers before your first run - see [AI providers](docs/ai-providers.md#cyber-safeguards-do-this-before-your-first-scan).
### Run Shannon
@@ -94,7 +104,11 @@ Shannon pulls the worker image from Docker Hub, starts the required local infras
For source builds, authenticated scans, provider-specific setup, and platform notes, see [Documentation](#documentation).
> [!TIP]
> **Prefer to run on your Claude Code subscription instead of API credits?** The [`shannon-v1`](https://github.com/KeygraphHQ/shannon/tree/shannon-v1) branch is the last release built on the Claude Agent SDK, so it accepts a Claude Code OAuth token. Generate one with `claude setup-token`, then run `npx @keygraph/shannon@1.9.0 setup` and pick **OAuth Token**. Pentests then cost nothing beyond your existing subscription.
> **Prefer to use a subscription instead of API credits?**
>
> - **OpenAI Codex:** The latest version of Shannon supports ChatGPT Plus and Pro subscriptions. Follow the [OpenAI Codex subscription setup guide](docs/ai-providers.md#openai-codex-chatgpt-pluspro-subscription) to get started.
> - **xAI (Grok):** The latest version of Shannon supports xAI subscriptions. Follow the [xAI subscription setup guide](docs/ai-providers.md#xai-grok-subscription) to get started.
> - **Claude Code:** The latest version of Shannon does not support Claude Code subscriptions. Follow the [Claude Code subscription setup guide](docs/ai-providers.md#claude-code-subscription) to use version `1.9.0`, which is the final release built on the Claude Agent SDK.
## Key Capabilities
@@ -104,6 +118,8 @@ For source builds, authenticated scans, provider-specific setup, and platform no
- **Authenticated testing**: configuration files can describe login flows, test credentials, TOTP, email-based login flows, focus areas, and rules of engagement.
- **OWASP-focused coverage**: Shannon targets exploitable Injection, XSS, SSRF, Broken Authentication, and Broken Authorization issues.
- **Resumable workspaces**: Shannon can resume interrupted runs without re-running completed agents.
- **Machine-readable output**: Shannon emits findings as structured JSON, and as SARIF 2.1.0 by default on exploit-mode scans (opt out with `report.sarif: "false"`). SARIF is the OASIS standard for static analysis results, so findings flow into any code scanning service, vulnerability management platform, security dashboard, or CI/CD pipeline that reads it.
- **Bring your own key, provider-agnostic**: Shannon runs on Anthropic, OpenAI, xAI, AWS Bedrock, and any endpoint speaking the Anthropic Messages API or the OpenAI Chat Completions or Responses API, including self-hosted models served through Ollama, vLLM, or LM Studio and gateways such as OpenRouter and LiteLLM. You supply the credentials, so source code and model traffic stay inside your infrastructure. Local and self-hosted models are technically supported but not recommended: they may not follow Shannon's instructions or tool-use constraints as reliably as frontier models, so take that path only if you know how your chosen model behaves.
## Editions
@@ -207,7 +223,7 @@ Important limitations:
- Shannon Open Source focuses on actively exploitable issues such as Injection, XSS, SSRF, Broken Authentication, and Broken Authorization. Broader static-analysis coverage, including vulnerable dependencies and insecure configurations, is delivered through the Keygraph platform.
- Findings still require human review. LLM-generated reports can contain weakly supported or incorrect details.
- Shannon is officially supported with Claude models. Smaller, alternative, or proxied non-Claude models may be incomplete or unstable.
- Anthropic, OpenAI, xAI, and AWS Bedrock are built-in providers, and any Anthropic Messages API or OpenAI Chat Completions or Responses API endpoint works through a custom base URL. Model capability varies, and a model that does not follow Shannon's instructions or tool-use constraints reliably will produce weaker results.
- A full run can take roughly 1 to 1.5 hours and may incur LLM API costs depending on model pricing and application complexity.
- Do not scan untrusted or adversarial codebases. AI-powered tools that read source code can be exposed to prompt injection.
@@ -246,6 +262,32 @@ Stay connected:
- [Twitter/X: @KeygraphHQ](https://twitter.com/KeygraphHQ)
- [LinkedIn: Keygraph](https://linkedin.com/company/keygraph)
## Common Questions
### Can I self-host Shannon?
Yes. Shannon Open Source runs entirely on your own infrastructure in an ephemeral Docker container. Your source code is mounted read-only and never leaves your environment.
### Does Shannon support bring your own key (BYOK)?
Yes, always. You provide the LLM credentials Shannon uses to run a pentest, in every deployment, open source and commercial. Keygraph never proxies your model traffic.
### Does Shannon output SARIF?
Yes. Shannon emits SARIF 2.1.0, the OASIS standard format for static analysis results, alongside structured JSON. Any SARIF consumer reads it: code scanning services, vulnerability management platforms, security dashboards, and CI/CD pipelines. It is written by default on exploit-mode scans; set `report.sarif` to `"false"` in your configuration file to opt out.
### Which AI providers does Shannon support?
Anthropic, OpenAI, xAI, and AWS Bedrock are built in and configured directly by provider ID. Beyond those, Shannon runs on any endpoint that implements the Anthropic Messages API or the OpenAI Chat Completions or Responses API, reached through a custom base URL. The rule is the API format, not the vendor. Shannon uses a single unified model setting throughout a pentest.
### Can I run Shannon on a local or self-hosted model?
Technically yes, but it is not recommended. Shannon works with local models served through Ollama, vLLM, or LM Studio, which expose an OpenAI-compatible endpoint, as well as routers such as OpenRouter and gateways such as LiteLLM. Point Shannon at the endpoint with a custom base URL. Capability varies, and a model that does not follow Shannon's instructions or tool-use constraints reliably will produce weaker pentests than a frontier model, so take this path only if you know how your chosen model behaves. See [AI providers](docs/ai-providers.md#custom-base-url).
### Does Shannon actually exploit vulnerabilities, or just scan?
Shannon executes real exploits. It reports a finding only when it has produced a working proof-of-concept, and discards hypotheses it cannot prove. It is a pentester, not a scanner.
<p align="center">
<b>Built by <a href="https://keygraph.io">Keygraph</a></b>
</p>
+49 -11
View File
@@ -1,22 +1,60 @@
<div align="center">
<img src="https://raw.githubusercontent.com/KeygraphHQ/shannon/main/assets/github-banner.png" alt="Shannon — AI Pentester for Web Applications and APIs" width="100%">
<img src="https://raw.githubusercontent.com/KeygraphHQ/shannon/main/assets/github-banner-light.png" alt="Shannon, AI Pentester for Web Apps and APIs, by Keygraph" width="100%">
# Shannon — AI Pentester by Keygraph
### Shannon is an autonomous, AI pentester for web applications and APIs.
Shannon is an autonomous, white-box AI pentester for web applications and APIs. <br />
It analyzes your source code, identifies attack vectors, and executes real exploits to prove vulnerabilities before they reach production.
It analyzes your source code, identifies attack paths, and executes real exploits to prove vulnerabilities before they reach production.
**This package is Shannon Open Source: the full agent, run locally from your command line.**
---
<a href="https://github.com/KeygraphHQ/shannon/discussions/categories/announcements"><img src="https://raw.githubusercontent.com/KeygraphHQ/shannon/main/assets/announcements.png" height="40" alt="Announcements"></a>
<a href="https://discord.gg/9ZqQPuhJB7"><img src="https://raw.githubusercontent.com/KeygraphHQ/shannon/main/assets/discord.png" height="40" alt="Join Discord"></a>
<a href="https://keygraph.io/"><img src="https://raw.githubusercontent.com/KeygraphHQ/shannon/main/assets/Keygraph_Button.png" height="40" alt="Visit Keygraph.io"></a>
<a href="https://www.linkedin.com/company/keygraph/"><img src="https://raw.githubusercontent.com/KeygraphHQ/shannon/main/assets/linkedin.png" height="40" alt="Follow Us on Linkedin"></a>
<a href="https://discord.gg/9ZqQPuhJB7"><img src="https://raw.githubusercontent.com/KeygraphHQ/shannon/main/assets/discord_button_light.png" height="40" alt="Join Discord"></a>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;<a href="https://keygraph.io/"><img src="https://raw.githubusercontent.com/KeygraphHQ/shannon/main/assets/keygraph_button_light.png" height="40" alt="Visit Keygraph.io"></a>
---
**Full README and usage guide**
[https://github.com/KeygraphHQ/shannon#readme](https://github.com/KeygraphHQ/shannon#readme)
</div>
## Quick Start
### Prerequisites
- **Docker**: required for the worker container.
- **Node.js 18+**: required for the recommended `npx` workflow.
- **AI provider credentials**: Shannon runs on Anthropic, OpenAI, xAI, AWS Bedrock, any other provider in the harness catalogue, and any endpoint that speaks the Anthropic Messages API or the OpenAI Chat Completions or Responses API through a custom base URL. You bring your own key, and Keygraph never proxies your model traffic. Shannon is provider-agnostic.
- **Cyber safeguards cleared with your provider**: Anthropic and OpenAI apply real-time safeguards to cyber-security workloads, which can interrupt a scan mid-run. Complete their guidance for legitimate security testers before your first run.
### Run Shannon
> **Warning:** Shannon actively executes exploits. Run it only against applications and environments you own or have explicit written authorization to test. Do not run Shannon against production systems.
```bash
# Configure credentials with the interactive wizard.
npx @keygraph/shannon setup
# Run a pentest against a source-available target.
npx @keygraph/shannon start -u https://your-app.com -r /path/to/your-repo
```
Shannon pulls the worker image from Docker Hub, starts the required local infrastructure, mounts the target repository read-only inside an ephemeral worker container, and writes results to a local workspace.
## Editions
Shannon ships in two ways. **Shannon Open Source** is this package: the standalone pentester you run yourself, on demand, and complete in that lane. The **Keygraph platform** is the commercial product that runs an enhanced build of Shannon continuously and closes the full AppSec lifecycle around it - code analysis, finding management, automated remediation, verification, and enterprise deployment.
## Documentation
**Full README, guides, and usage documentation:** [github.com/KeygraphHQ/shannon](https://github.com/KeygraphHQ/shannon#readme)
## License
Shannon Open Source is licensed under the [GNU Affero General Public License v3.0](https://github.com/KeygraphHQ/shannon/blob/main/LICENSE).
Commercial and enterprise licensing is available for organizations that need different license terms, commercial support, private redistribution, managed-service use, or broader deployment options, including the Keygraph platform.
For commercial licensing, contact [shannon@keygraph.io](mailto:shannon@keygraph.io).
<p align="center">
<b>Built by <a href="https://keygraph.io">Keygraph</a></b>
</p>
+7 -2
View File
@@ -1,7 +1,7 @@
{
"name": "@keygraph/shannon",
"version": "0.0.0",
"description": "Shannon - Autonomous white-box AI pentester for web applications and APIs by Keygraph",
"description": "Shannon is an autonomous white-box AI pentester for web applications and APIs, by Keygraph.",
"type": "module",
"main": "dist/index.mjs",
"bin": {
@@ -18,6 +18,7 @@
},
"dependencies": {
"@clack/prompts": "^1.1.0",
"@temporalio/client": "^1.11.0",
"chokidar": "^5.0.0",
"dotenv": "^17.3.1",
"smol-toml": "^1.6.1"
@@ -34,8 +35,12 @@
"appsec",
"keygraph"
],
"author": "",
"author": "Keygraph, Inc.",
"license": "AGPL-3.0-only",
"bugs": {
"url": "https://github.com/KeygraphHQ/shannon/issues"
},
"homepage": "https://github.com/KeygraphHQ/shannon#readme",
"repository": {
"type": "git",
"url": "git+https://github.com/KeygraphHQ/shannon.git",
+106
View File
@@ -0,0 +1,106 @@
/**
* Shared argument parsing for CLI commands.
*
* Every command declares which boolean flags, value options, and positionals it
* accepts; `parseArgs` resolves aliases, rejects anything unrecognized, and hands
* back a typed result. This centralizes the common flags (notably `--yes`/`-y`) so
* each command no longer re-hardcodes `args.includes('--yes')`, and it makes
* unknown flags and stray arguments fail loudly instead of being silently ignored.
*/
import { closestMatch } from './suggest.js';
/** Thrown when argv does not match a command's schema. The dispatcher formats it. */
export class ArgError extends Error {}
/** Tokens that set the "skip confirmation" flag, declared once for every command. */
export const YES_FLAGS = ['--yes', '-y'] as const;
export interface ArgSchema {
/** Boolean flags: result key -> accepted tokens (canonical plus any aliases). */
readonly booleans?: Record<string, readonly string[]>;
/** Value-taking options: result key -> accepted tokens. */
readonly values?: Record<string, readonly string[]>;
/** Maximum positional arguments allowed. Defaults to 0. */
readonly maxPositionals?: number;
/** Extra guidance appended to the error when too many positionals are given. */
readonly positionalHint?: string;
}
export interface ParsedArgs {
readonly flags: Record<string, boolean>;
readonly values: Record<string, string>;
readonly positionals: readonly string[];
}
/** Build a token -> result-key lookup from a schema section. */
function indexTokens(section: Record<string, readonly string[]>): Map<string, string> {
const byToken = new Map<string, string>();
for (const [key, tokens] of Object.entries(section)) {
for (const token of tokens) {
byToken.set(token, key);
}
}
return byToken;
}
export function parseArgs(argv: readonly string[], schema: ArgSchema): ParsedArgs {
const booleanByToken = indexTokens(schema.booleans ?? {});
const valueByToken = indexTokens(schema.values ?? {});
const maxPositionals = schema.maxPositionals ?? 0;
const flags: Record<string, boolean> = {};
const values: Record<string, string> = {};
const positionals: string[] = [];
for (let i = 0; i < argv.length; i++) {
const arg = argv[i];
if (arg === undefined) {
continue;
}
const equalsIndex = arg.startsWith('--') ? arg.indexOf('=') : -1;
const token = equalsIndex === -1 ? arg : arg.slice(0, equalsIndex);
const inlineValue = equalsIndex === -1 ? undefined : arg.slice(equalsIndex + 1);
const booleanKey = booleanByToken.get(token);
if (booleanKey !== undefined) {
if (inlineValue !== undefined) {
throw new ArgError(`Flag ${token} does not take a value`);
}
flags[booleanKey] = true;
continue;
}
const valueKey = valueByToken.get(token);
if (valueKey !== undefined) {
if (inlineValue !== undefined) {
values[valueKey] = inlineValue;
continue;
}
const next = argv[i + 1];
if (next === undefined || next.startsWith('-')) {
throw new ArgError(`Option ${token} requires a value`);
}
values[valueKey] = next;
i++;
continue;
}
if (arg.startsWith('-')) {
const suggestion = closestMatch(token, [...booleanByToken.keys(), ...valueByToken.keys()]);
const hint = suggestion ? `\nDid you mean '${suggestion}'?` : '';
throw new ArgError(`Unknown option: ${token}${hint}`);
}
positionals.push(arg);
}
if (positionals.length > maxPositionals) {
const extra = positionals[maxPositionals];
const hint = schema.positionalHint ? `\n${schema.positionalHint}` : '';
throw new ArgError(`Unexpected argument: ${extra}${hint}`);
}
return { flags, values, positionals };
}
+159
View File
@@ -0,0 +1,159 @@
/**
* Section chrome — the sunset-ramp hierarchy beneath the splash wordmark.
* (Distinct from ui.ts, which holds the spinner/step helpers.)
*
* Three treatments, one job each:
* rule() static blocks that print once — `start`, `status`, a log's header
* gutter() streaming output, where a section scrolls off the top of the screen
* panel() the single summary block at the end of a run
*
* Presentation only. None of this is ever written to workflow.log — the file on disk
* stays plain text so `tail`, `grep`, and the completion regex in commands/logs.ts keep
* working against it.
*/
import { supportsColor } from './tty.js';
/**
* Sunset ramp, yellow at the top row down to burnt orange at the base.
* The wordmark paints row i with stop i and edges it with stop i + 1; section chrome
* draws from the same seven stops so the hierarchy reads as one family.
* `xterm` is the 256-color approximation for terminals without 24-bit color.
*/
export const SUNSET: ReadonlyArray<{ rgb: readonly [number, number, number]; xterm: number }> = [
{ rgb: [247, 203, 45], xterm: 220 },
{ rgb: [246, 182, 38], xterm: 220 },
{ rgb: [245, 160, 32], xterm: 214 },
{ rgb: [242, 141, 28], xterm: 214 },
{ rgb: [238, 121, 24], xterm: 208 },
{ rgb: [231, 100, 21], xterm: 208 },
{ rgb: [222, 82, 19], xterm: 202 },
];
/** Half-block bar for streaming sections — the wordmark's █ at one eighth the weight. */
const BAR = '▌';
/** Columns reserved to the left of every section, matching the existing output grid. */
const INDENT = 2;
/** Rules stop here even in a wide terminal; a rule spanning 200 columns reads as a divider, not a header. */
const MAX_RULE = 64;
export interface Palette {
color: boolean;
RESET: string;
WHITE: string;
GRAY: string;
DIM: string;
RED: string;
/** The seven sunset stops, ready to emit. Empty strings when color is off. */
ramp: string[];
/** Ramp stop 0 — the solid yellow used for every static rule. */
YELLOW: string;
}
/**
* Build the escape set for the current terminal, degrading 24-bit → 256-color → bare text.
* Resolved per call rather than at import so NO_COLOR/FORCE_COLOR are honored whenever they land.
*/
export function palette(): Palette {
const color = supportsColor();
const truecolor = color && /truecolor|24bit/i.test(process.env.COLORTERM ?? '');
const ramp = SUNSET.map(({ rgb: [r, g, b], xterm }) => {
if (!color) return '';
return truecolor ? `\x1b[38;2;${r};${g};${b}m` : `\x1b[38;5;${xterm}m`;
});
return {
color,
RESET: color ? '\x1b[0m' : '',
WHITE: color ? '\x1b[1;97m' : '',
GRAY: color ? '\x1b[0;37m' : '',
DIM: color ? '\x1b[90m' : '',
RED: color ? '\x1b[0;31m' : '',
ramp,
YELLOW: ramp[0] ?? '',
};
}
/** Usable width, leaving the indent and a column of breathing room at the right edge. */
function columns(): number {
return process.stdout.columns && process.stdout.columns > 0 ? process.stdout.columns : 80;
}
/** Printed width of a string, ignoring any escapes already embedded in it. */
export function visibleWidth(text: string): number {
// biome-ignore lint/suspicious/noControlCharactersInRegex: matching SGR escapes is the point
return text.replace(/\x1b\[[0-9;]*m/g, '').length;
}
/**
* Option A — a static section header: the existing label, then a solid yellow rule.
* The label keeps whatever case and punctuation it already had; only the rule is added.
* Degrades to an undecorated label when the terminal is too narrow to carry one.
*/
export function rule(label: string, indent = INDENT): string {
const { WHITE, YELLOW, RESET } = palette();
const pad = ' '.repeat(indent);
const width = Math.min(MAX_RULE, columns() - indent - 1);
const dashes = width - visibleWidth(label) - 1;
if (dashes < 2) return `${pad}${WHITE}${label}${RESET}`;
return `${pad}${WHITE}${label}${RESET} ${YELLOW}${'─'.repeat(dashes)}${RESET}`;
}
/**
* Grey the label half of an aligned `Label: value` line, leaving the value at default
* weight. Purely additive: the string's own characters are never rewritten, so alignment
* that was already correct stays correct.
*/
export function field(line: string): string {
const { GRAY, RESET } = palette();
const match = /^(\s*)([A-Za-z][A-Za-z ]*:)(\s*)(.*)$/.exec(line);
if (!match) return line;
return `${match[1]}${GRAY}${match[2]}${RESET}${match[3]}${match[4]}`;
}
/**
* Option C — one line of a streaming section, carrying the section's bar in the gutter.
* `stop` indexes the sunset ramp and wraps, so consecutive sections stay distinguishable
* however many a run produces.
*/
export function gutter(text: string, stop: number, indent = INDENT): string {
const { ramp, RESET } = palette();
const color = ramp[((stop % ramp.length) + ramp.length) % ramp.length] ?? '';
const pad = ' '.repeat(indent);
// Trailing space is dropped on empty lines so sections don't emit trailing whitespace.
return text ? `${pad}${color}${BAR}${RESET} ${text}` : `${pad}${color}${BAR}${RESET}`;
}
/**
* Option E — the framed summary block, used once at the end of a run.
* Falls back to a rule plus indented lines when the terminal is too narrow to hold the
* frame, since a box that wraps is worse than no box at all.
*/
export function panel(title: string, body: string[], indent = INDENT): string[] {
const { WHITE, YELLOW, RESET } = palette();
const pad = ' '.repeat(indent);
const titleWidth = visibleWidth(title);
const widest = body.reduce((max, line) => Math.max(max, visibleWidth(line)), 0);
const available = columns() - indent - 6;
const inner = Math.max(titleWidth + 1, widest);
if (available < inner || available < titleWidth + 3) {
return [rule(title, indent), '', ...body.map((line) => `${pad} ${line}`)];
}
const frame = (s: string): string => `${YELLOW}${s}${RESET}`;
const top = `${pad}${frame('╭─')} ${WHITE}${title}${RESET} ${frame(`${'─'.repeat(inner + 1 - titleWidth)}╮`)}`;
const bottom = `${pad}${frame(`╰${'─'.repeat(inner + 4)}╯`)}`;
const rows = body.map((line) => {
const fill = ' '.repeat(inner - visibleWidth(line));
return `${pad}${frame('│')} ${line}${fill} ${frame('│')}`;
});
return [top, ...rows, bottom];
}
+34
View File
@@ -0,0 +1,34 @@
/**
* ANSI color and style escapes — the single source for the CLI's palette.
*
* Codes are plain constants; callers decide whether to emit them via `paint`
* (wrap-and-reset) or `gate` (prefix-or-empty), gating on `supportsColor()` from
* `tty.ts`. Cursor-control escapes live with their sole consumer, not here — this
* module is color only.
*/
export const RESET = '\x1b[0m';
/** Shannon brand gold — the running/completed accent, shared with the splash logo. */
export const GOLD = '\x1b[38;2;244;197;66m';
export const BOLD = '\x1b[1m';
export const RED = '\x1b[31m';
export const YELLOW = '\x1b[33m';
export const DIM = '\x1b[90m';
// The splash logo uses bolder variants of cyan/white/yellow than the progress tree.
export const CYAN = '\x1b[36;1m';
export const WHITE = '\x1b[1;37m';
export const GRAY = '\x1b[0;37m';
export const BOLD_YELLOW = '\x1b[1;33m';
/** Wrap `text` in `code` and reset, or return it unchanged when color is off. */
export function paint(text: string, code: string, enabled: boolean): string {
return enabled ? `${code}${text}${RESET}` : text;
}
/** A style code when color is on, or an empty string when off — for templates that interleave prefixes directly. */
export function gate(code: string, enabled: boolean): string {
return enabled ? code : '';
}
+8 -4
View File
@@ -3,13 +3,17 @@
* Requires a clone (Dockerfile in the working directory).
*/
import { buildImage, canBuildImage } from '../docker.js';
import { buildImage, canBuildImage, ensureDocker } from '../docker.js';
import { fail } from '../errors.js';
export function build(noCache: boolean, version: string): void {
ensureDocker();
if (!canBuildImage()) {
console.error('ERROR: Build is only available when running from the Shannon repository');
console.error(' (Dockerfile not found in current directory)');
process.exit(1);
fail(
'Build is only available when running from the Shannon repository',
' (Dockerfile not found in current directory)',
);
}
buildImage(noCache, version);
+132 -56
View File
@@ -1,18 +1,24 @@
/**
* `shannon logs` command — tail a scan's live log.
*
* Uses chokidar for reliable cross-platform file watching and
* bounded synchronous reads to prevent duplicate output.
* The log file is streamed for its content; completion is decided by Temporal (the
* workflow's status), so a worker that dies mid-run can't leave the tail hanging. Uses
* chokidar for reliable cross-platform file watching and bounded synchronous reads to
* prevent duplicate output.
*/
import fs from 'node:fs';
import path from 'node:path';
import { setTimeout as sleep } from 'node:timers/promises';
import { watch } from 'chokidar';
import { field } from '../chrome.js';
import { fail } from '../errors.js';
import { getWorkspacesDir } from '../home.js';
import { LogRenderer } from '../log-render.js';
import { resolveRunFile } from '../paths.js';
// Match the exact line the worker writes — anchored to prevent false positives from agent output
const COMPLETION_PATTERN = /^Scan (COMPLETED|FAILED)$/m;
import { resolveWorkflowId } from '../session.js';
import { waitForWorkflowClose } from '../temporal-client.js';
import { stdoutIsTerminal } from '../tty.js';
/** Read a byte range from a file and return it as a UTF-8 string. */
function readRange(filePath: string, start: number, end: number): string {
@@ -28,7 +34,7 @@ function readRange(filePath: string, start: number, end: number): string {
}
/** Resolve a workspace ID to its workflow.log path, or exit with an error. */
function resolveLogFile(workspaceId: string): string {
export function resolveLogFile(workspaceId: string): string {
const workspacesDir = getWorkspacesDir();
// 1. Direct match
@@ -49,59 +55,129 @@ function resolveLogFile(workspaceId: string): string {
if (fs.existsSync(namedPath)) return namedPath;
}
console.error(`ERROR: No scan found named: ${workspaceId}`);
console.error('');
console.error('Possible causes:');
console.error(" - The scan hasn't started yet");
console.error(' - The workspace name is incorrect');
console.error('');
console.error('Check the dashboard at http://localhost:8233 for scan details');
process.exit(1);
fail(
`No scan found named: ${workspaceId}`,
'',
'Possible causes:',
" - The scan hasn't started yet",
' - The workspace name is incorrect',
'',
'Check the dashboard at http://localhost:8233 for scan details',
);
}
export interface TailOptions {
/** Workflow whose Temporal status decides when the tail stops. Without it, only Ctrl-C ends the tail. */
readonly workflowId?: string;
/** Called if the tail ends because Temporal became unreachable, with the captured error. */
readonly onUnreachable?: (lastError: string) => void;
}
/** Outcome of a tail: whether the streamed log already contained the worker's `Scan FAILED` block. */
export interface TailResult {
readonly sawFailure: boolean;
}
// The worker writes this exact line at the head of its terminal failure summary.
const FAILURE_MARKER = /^Scan FAILED$/m;
/**
* Stream a scan's log to the terminal until the workflow closes (completion comes from Temporal,
* or Ctrl-C). A Temporal outage is warned about and, if sustained, ends the tail with a diagnostic.
* Never exits the process: plain `logs` exits; `start --follow` reads the workflow outcome first.
* Reports whether the log already showed the failure, so a caller need not print it a second time.
*/
export function tailUntilComplete(logFile: string, opts: TailOptions = {}): Promise<TailResult> {
return new Promise((resolve) => {
let position = 0;
let done = false;
let sawFailure = false;
// Decorates the streamed log for the terminal; a pass-through when colour is off, so
// piped/redirected output stays byte-identical and the failure check still sees raw text.
const renderer = new LogRenderer();
const controller = new AbortController();
let watcher: ReturnType<typeof watch> | undefined;
/** Output any new content appended since the last read. */
function flush(): void {
try {
const { size } = fs.statSync(logFile);
if (size <= position) return;
const data = readRange(logFile, position, size);
process.stdout.write(renderer.write(data));
position = size;
if (!sawFailure && FAILURE_MARKER.test(data)) {
sawFailure = true;
}
} catch {
// File not present yet or transiently unreadable — nothing to flush this round.
}
}
function finish(): void {
if (done) return;
done = true;
process.stdout.write(renderer.end());
controller.abort();
if (watcher) {
watcher.close().finally(() => resolve({ sawFailure }));
// Safety net — resolve anyway if watcher.close() stalls.
setTimeout(() => resolve({ sawFailure }), 1000).unref();
} else {
resolve({ sawFailure });
}
}
// 1. Output existing content, then stream anything appended.
flush();
watcher = watch(logFile, { persistent: true });
watcher.on('change', () => flush());
// 2. Ctrl-C stops watching.
process.on('SIGINT', finish);
// 3. Temporal decides completion. Without a workflow id, the tail relies on Ctrl-C alone.
if (opts.workflowId) {
waitForWorkflowClose(opts.workflowId, {
signal: controller.signal,
onConnectionTrouble: (lastError) => {
if (!done) console.error(`\n⚠ Lost contact with Temporal, retrying… (${lastError})`);
},
onReconnected: () => {
if (!done) console.error(' Reconnected to Temporal.');
},
})
.then(async (end) => {
if (done) return;
// Flush, let a just-written final summary land, then flush the tail once more.
flush();
await sleep(750).catch(() => {});
flush();
if (end.reason === 'unreachable') {
console.error('\nScan watch aborted: lost contact with Temporal.');
console.error(` Last error: ${end.lastError}`);
console.error(' Temporal may have crashed — check `docker compose logs temporal`.');
opts.onUnreachable?.(end.lastError);
}
finish();
})
.catch(() => {
// waitForWorkflowClose never rejects; guard only against an aborted race.
});
}
});
}
export function logs(workspaceId: string): void {
const logFile = resolveLogFile(workspaceId);
let position = 0;
const workflowId = resolveWorkflowId(workspaceId);
console.error(stdoutIsTerminal() ? field(`Tailing scan log: ${logFile}`) : 'Tailing scan log');
/**
* Output any new content appended since the last read.
* Returns true when the workflow completion marker is detected.
*/
function flush(): boolean {
try {
const { size } = fs.statSync(logFile);
if (size <= position) return false;
const data = readRange(logFile, position, size);
process.stdout.write(data);
position = size;
return COMPLETION_PATTERN.test(data);
} catch {
// File deleted or unreadable — treat as done
return true;
}
}
console.log(`Tailing scan log: ${logFile}`);
// 1. Output existing content
if (flush()) {
process.exit(0);
}
// 2. Watch for appended content via chokidar
const watcher = watch(logFile, { persistent: true });
const shutdown = (): void => {
watcher.close().finally(() => process.exit(0));
// Safety net — force exit if watcher.close() stalls
setTimeout(() => process.exit(0), 1000).unref();
};
watcher.on('change', () => {
if (flush()) shutdown();
});
process.on('SIGINT', shutdown);
let unreachable = false;
tailUntilComplete(logFile, {
...(workflowId ? { workflowId } : {}),
onUnreachable: () => {
unreachable = true;
},
}).finally(() => process.exit(unreachable ? 1 : 0));
}
+26
View File
@@ -0,0 +1,26 @@
/**
* `shannon reset` command — stop everything and wipe all Temporal data and volumes,
* returning the machine to a clean slate. The destructive counterpart to `stop`.
*/
import * as p from '@clack/prompts';
import { confirmByTyping } from '../confirm.js';
import { ensureDocker, runningContainers, stopContainers, stopInfra, WORKER_FILTER } from '../docker.js';
export async function reset(): Promise<void> {
ensureDocker();
console.log('This will stop all running scans and permanently remove all Temporal data and volumes.');
await confirmByTyping('reset', 'confirm');
const spinner = p.spinner();
spinner.start('Stopping scans');
const running = runningContainers(WORKER_FILTER);
await stopContainers(running);
spinner.stop(
running.length > 0 ? `Stopped ${running.length} scan${running.length === 1 ? '' : 's'}` : 'No scans running',
);
await stopInfra(true);
console.log('Reset complete.');
}
+197
View File
@@ -0,0 +1,197 @@
/**
* `shannon scans` command — list completed scans and where each report lives.
*
* A scan counts as completed when it produced a report. The report can live in any of a
* few locations depending on the version that ran it, so `findReport` probes them in order
* and the first hit is both the completion signal and the link target behind the workspace
* name. The date and wall-clock duration come from the run's session.json
* (createdAt/completedAt), with the report file's mtime as the date fallback for
* runs that lack a recorded time.
*
* Human-readable by default; `--json` emits the same rows as raw machine values on stdout.
*
* Filesystem-only (local ./workspaces/ or npx ~/.shannon/workspaces/ via getWorkspacesDir);
* no Temporal dependency.
*/
import fs from 'node:fs';
import path from 'node:path';
import { pathToFileURL } from 'node:url';
import { BOLD, GOLD, paint } from '../colors.js';
import { getWorkspacesDir } from '../home.js';
import { commandPrefix } from '../mode.js';
import { FINAL_REPORT_PDF_FILENAME, INTERNAL_DIR, resolveRunFile } from '../paths.js';
import { stdoutIsTerminal, supportsColor } from '../tty.js';
/** Assembled report in the deliverables dir. Must match ASSEMBLED_REPORT_FILENAME in the worker package. */
const ASSEMBLED_REPORT_FILENAME = 'comprehensive_security_assessment_report.md';
/** Run-root markdown surfaced by older versions, before the PDF. Kept so those runs still list. */
const FINAL_REPORT_MD_FILENAME = 'Security-Assessment-Report.md';
const DELIVERABLES_SUBDIR = 'deliverables';
/** One completed scan; raw values so the table and --json render from one source. */
interface ScanRow {
readonly workspace: string;
/** Completion time in ms — sort key and date source. */
readonly finishedMs: number;
/** Wall-clock duration (completedAt − createdAt) in ms, or null when unknown. */
readonly durationMs: number | null;
/** Absolute path to the report file — the link target behind the workspace name. */
readonly report: string;
}
/** The --json row shape: raw machine values, one per completed scan. */
interface JsonRow {
readonly workspace: string;
readonly finishedAt: string;
readonly durationMs: number | null;
readonly reportPath: string;
}
/** Compact wall-clock duration from milliseconds: "47s", "1m 32s", "1h 47m". */
function formatDuration(ms: number): string {
const totalSeconds = Math.round(ms / 1000);
if (totalSeconds < 60) {
return `${totalSeconds}s`;
}
const totalMinutes = Math.floor(totalSeconds / 60);
if (totalMinutes < 60) {
return `${totalMinutes}m ${totalSeconds % 60}s`;
}
return `${Math.floor(totalMinutes / 60)}h ${totalMinutes % 60}m`;
}
/**
* Wrap `text` in an OSC 8 hyperlink to `url` so a supporting terminal opens it on click,
* or return `text` unchanged. Terminals without OSC 8 simply show the text.
*/
function hyperlink(text: string, url: string): string {
return `\x1b]8;;${url}\x1b\\${text}\x1b]8;;\x1b\\`;
}
/** First existing report path for a run (newest-surfaced first), or null if it has none. */
function findReport(runDir: string): string | null {
const candidates = [
path.join(runDir, FINAL_REPORT_PDF_FILENAME),
path.join(runDir, FINAL_REPORT_MD_FILENAME),
path.join(runDir, INTERNAL_DIR, DELIVERABLES_SUBDIR, ASSEMBLED_REPORT_FILENAME),
path.join(runDir, DELIVERABLES_SUBDIR, ASSEMBLED_REPORT_FILENAME),
];
for (const candidate of candidates) {
if (fs.existsSync(candidate)) {
return candidate;
}
}
return null;
}
interface SessionData {
readonly session: { readonly createdAt?: string; readonly completedAt?: string };
}
/** Read a run's session.json (dual-read across layouts). Missing or unreadable → empty shape. */
function readSession(runDir: string): SessionData {
try {
const parsed = JSON.parse(fs.readFileSync(resolveRunFile(runDir, 'session.json'), 'utf8'));
return { session: parsed?.session ?? {} };
} catch {
return { session: {} };
}
}
/** Gather every workspace that has a report, one row each. */
function collectCompletedScans(workspacesDir: string): ScanRow[] {
let entries: fs.Dirent[];
try {
entries = fs.readdirSync(workspacesDir, { withFileTypes: true });
} catch {
// Workspaces directory does not exist yet — no scans have ever run.
return [];
}
const rows: ScanRow[] = [];
for (const entry of entries) {
if (!entry.isDirectory()) {
continue;
}
const runDir = path.join(workspacesDir, entry.name);
const reportPath = findReport(runDir);
if (!reportPath) {
continue;
}
const { session } = readSession(runDir);
const completedMs = Date.parse(session.completedAt ?? '');
const createdMs = Date.parse(session.createdAt ?? '');
const finishedMs = Number.isNaN(completedMs) ? fs.statSync(reportPath).mtimeMs : completedMs;
const durationMs = Number.isNaN(completedMs) || Number.isNaN(createdMs) ? null : completedMs - createdMs;
rows.push({ workspace: entry.name, finishedMs, durationMs, report: reportPath });
}
return rows;
}
function toJsonRow(row: ScanRow): JsonRow {
return {
workspace: row.workspace,
finishedAt: new Date(row.finishedMs).toISOString(),
durationMs: row.durationMs,
reportPath: row.report,
};
}
/** Print the completed scans as an aligned table with the workspace name linked to its report. */
function printTable(workspacesDir: string, rows: readonly ScanRow[]): void {
if (rows.length === 0) {
const prefix = commandPrefix();
console.log(`No completed scans yet. Run '${prefix} start -u <url> -r <path>' to begin.`);
return;
}
const color = supportsColor();
// On a terminal the workspace name is an OSC 8 hyperlink that opens its report; when
// piped there is nothing to click, so it prints as plain text.
const linkable = stdoutIsTerminal();
const table = rows.map((row) => ({
finished: new Date(row.finishedMs).toISOString().slice(0, 10),
duration: row.durationMs === null ? '—' : formatDuration(row.durationMs),
workspace: row.workspace,
report: row.report,
}));
const dateWidth = Math.max('FINISHED'.length, 'YYYY-MM-DD'.length);
const durationWidth = Math.max('DURATION'.length, ...table.map((row) => row.duration.length));
console.log(`\nCompleted scans in ${workspacesDir}:\n`);
const header = `${'FINISHED'.padEnd(dateWidth)} ${'DURATION'.padEnd(durationWidth)} WORKSPACE`;
console.log(paint(header, BOLD, color));
for (const row of table) {
const finished = row.finished.padEnd(dateWidth);
const duration = row.duration.padEnd(durationWidth);
const name = paint(row.workspace, GOLD, color);
const workspace = linkable ? hyperlink(name, pathToFileURL(row.report).href) : name;
console.log(`${finished} ${duration} ${workspace}`);
}
console.log('');
}
export function scans(opts: { readonly json: boolean }): void {
const workspacesDir = getWorkspacesDir();
const rows = collectCompletedScans(workspacesDir);
// Latest on top.
rows.sort((a, b) => b.finishedMs - a.finishedMs);
if (opts.json) {
console.log(JSON.stringify(rows.map(toJsonRow), null, 2));
return;
}
printTable(workspacesDir, rows);
}
+4 -1
View File
@@ -11,7 +11,9 @@ import path from 'node:path';
import * as p from '@clack/prompts';
import { type ShannonConfig, saveConfig } from '../config/writer.js';
import { CURATED_PROVIDERS, type CuratedProviderId, isCuratedProvider, type OpenAiFormat } from '../model-spec.js';
import { displaySplash } from '../splash.js';
import { requireInteractive } from '../tty.js';
import { getVersion } from '../version.js';
const SHANNON_HOME = path.join(os.homedir(), '.shannon');
@@ -63,7 +65,8 @@ function modelIdPlaceholder(provider: string): string | undefined {
export async function setup(): Promise<void> {
requireInteractive('setup', 'For non-interactive use, export credentials as env vars (e.g. ANTHROPIC_API_KEY).');
p.intro('Shannon Setup');
displaySplash(getVersion());
p.intro('Setup');
// 1. Select provider. "Custom Base URL" is a route, not a provider — it asks
// which API dialect the gateway speaks and configures that provider. "Other
+165 -96
View File
@@ -8,14 +8,29 @@
import { execFileSync } from 'node:child_process';
import fs from 'node:fs';
import path from 'node:path';
import { ensureImage, ensureInfra, randomSuffix, spawnWorker } from '../docker.js';
import { buildEnvFlags, loadEnv, validateCredentials } from '../env.js';
import { setTimeout as sleep } from 'node:timers/promises';
import * as p from '@clack/prompts';
import { field, rule } from '../chrome.js';
import { ensureDocker, ensureImage, ensureInfra, randomSuffix, spawnWorker } from '../docker.js';
import { buildEnvFlags, loadEnv, resolveHostPiAuthPath, shouldUsePiAuth, validateCredentials } from '../env.js';
import { fail } from '../errors.js';
import { getWorkspacesDir, initHome } from '../home.js';
import { isLocal } from '../mode.js';
import { commandPrefix, isLocal } from '../mode.js';
import { resolveModelSpec } from '../model-spec.js';
import { FINAL_REPORT_FILENAME, INTERNAL_DIR, resolveConfig, resolveRepo, resolveRunFile } from '../paths.js';
import { displaySplash } from '../splash.js';
import {
expandHome,
FINAL_REPORT_PDF_FILENAME,
INTERNAL_DIR,
resolveConfig,
resolveRepo,
resolveRunFile,
} from '../paths.js';
import { indentFailureSegments } from '../scan/failure.js';
import { resolveWorkflowId } from '../session.js';
import { displayPlainBanner, displaySplash } from '../splash.js';
import { getTerminalOutcome } from '../temporal-client.js';
import { stdoutIsTerminal } from '../tty.js';
import { tailUntilComplete } from './logs.js';
export interface StartArgs {
url: string;
@@ -24,7 +39,8 @@ export interface StartArgs {
workspace?: string;
output?: string;
pipelineTesting: boolean;
debug: boolean;
keepContainer: boolean;
follow: boolean;
version: string;
}
@@ -59,22 +75,34 @@ export async function start(args: StartArgs): Promise<void> {
// 2. Validate credentials
const creds = validateCredentials();
if (!creds.valid) {
console.error(`ERROR: ${creds.error}`);
process.exit(1);
fail(creds.error ?? 'Invalid credentials');
}
// 3. Resolve paths
const repo = resolveRepo(args.repo);
const config = args.config ? resolveConfig(args.config) : undefined;
// Inputs are valid — identify the run before the Docker/Temporal setup work.
const bannerVersion = isLocal() ? undefined : args.version;
if (stdoutIsTerminal()) {
displaySplash(bannerVersion);
} else {
displayPlainBanner(bannerVersion);
}
// 4. Ensure workspaces dir is writable by container user (UID 1001)
const workspacesDir = getWorkspacesDir();
fs.mkdirSync(workspacesDir, { recursive: true });
fs.chmodSync(workspacesDir, 0o777);
// 5. Ensure image (auto-build in dev, pull in npx) and start infra
// 5. Ensure Docker and the worker image are available (pull/build prints its own progress).
ensureDocker();
ensureImage(args.version);
await ensureInfra();
// One spinner spans the whole launch: bringing up Temporal and registering the worker.
const spinner = p.spinner();
spinner.start('Starting scan');
await ensureInfra(spinner);
// 6. Generate unique task queue and container name
const suffix = randomSuffix();
@@ -109,7 +137,7 @@ export async function start(args: StartArgs): Promise<void> {
fs.mkdirSync(path.join(repo.hostPath, '.playwright'), { recursive: true });
// 10. Resolve output directory
const outputDir = args.output ? path.resolve(args.output) : undefined;
const outputDir = args.output ? path.resolve(expandHome(args.output)) : undefined;
if (outputDir) {
fs.mkdirSync(outputDir, { recursive: true });
}
@@ -117,10 +145,7 @@ export async function start(args: StartArgs): Promise<void> {
// 11. Resolve prompts directory (local mode only)
const promptsDir = isLocal() ? path.resolve('apps/worker/prompts') : undefined;
// 12. Display splash screen
displaySplash(isLocal() ? undefined : args.version);
// 13. Spawn worker container
// 12. Spawn worker container
const proc = spawnWorker({
version: args.version,
url: args.url,
@@ -134,19 +159,18 @@ export async function start(args: StartArgs): Promise<void> {
...(outputDir && { outputDir }),
workspace,
...(args.pipelineTesting && { pipelineTesting: true }),
...(args.debug && { debug: true }),
...(args.keepContainer && { keepContainer: true }),
...(shouldUsePiAuth() && { piAuthHostPath: resolveHostPiAuthPath() }),
});
// 14. Bail if `docker run -d` itself fails (mount error, image missing, etc.)
// Bail if `docker run -d` itself fails (mount error, image missing, etc.)
const dockerExitCode = await new Promise<number>((resolve) => {
proc.once('exit', (code) => resolve(code ?? 1));
proc.once('error', (err) => {
console.error(`Failed to start the scan: ${err.message}`);
resolve(1);
});
proc.once('error', () => resolve(1));
});
if (dockerExitCode !== 0) {
spinner.error('Could not start the scan');
process.exit(1);
}
@@ -163,64 +187,23 @@ export async function start(args: StartArgs): Promise<void> {
}
}
// Poll for workflow to register in session.json. Off-TTY, skip the dots and
// clear-line escape so redirected logs stay clean.
const animate = stdoutIsTerminal();
process.stdout.write('Waiting for the scan to start...');
let workflowId = '';
let started = false;
let attempts = 0;
const pollInterval = setInterval(() => {
attempts++;
if (attempts > 60) {
clearInterval(pollInterval);
process.stdout.write('\n');
console.error('Timed out waiting for the scan to start');
process.exit(1);
}
try {
const session = JSON.parse(fs.readFileSync(sessionJson, 'utf-8'));
const resumeAttempts: { workflowId: string }[] = session.session?.resumeAttempts ?? [];
// Fresh: session.json appears with originalWorkflowId. Resume: new resumeAttempts entry.
const ready = isResume ? resumeAttempts.length > initialResumeCount : !!session.session?.originalWorkflowId;
if (ready) {
clearInterval(pollInterval);
started = true;
// Latest workflow ID: last resume attempt, or originalWorkflowId for fresh scans
workflowId = resumeAttempts.at(-1)?.workflowId ?? session.session?.originalWorkflowId ?? '';
// Clear the waiting line, or just break it off-TTY
process.stdout.write(animate ? '\r\x1b[K' : '\n');
printInfo(args, workspace, workflowId, repo.hostPath, workspacesDir);
return;
}
} catch {
// File doesn't exist yet
}
if (animate) process.stdout.write('.');
}, 2000);
// Stop the worker container only if it hasn't started yet
// Stop the worker only if the scan hasn't registered yet (e.g. Ctrl-C mid-startup).
let cleaned = false;
const cleanup = (): void => {
if (cleaned || started) return;
cleaned = true;
clearInterval(pollInterval);
console.log('\nStopping scan...');
spinner.stop('Stopping scan');
try {
execFileSync('docker', ['stop', containerName], { stdio: 'pipe' });
} catch {
// Container may have already exited
}
if (args.debug) {
printDebugHint(containerName);
if (args.keepContainer) {
printPreservedContainerHint(containerName);
}
};
process.on('SIGINT', () => {
cleanup();
process.exit(0);
@@ -230,53 +213,139 @@ export async function start(args: StartArgs): Promise<void> {
process.exit(0);
});
process.on('exit', cleanup);
// Poll for the workflow to register in session.json; the spinner resolves once it does.
spinner.message('Waiting for the scan to start');
for (let attempts = 0; attempts < 60; attempts++) {
try {
const session = JSON.parse(fs.readFileSync(sessionJson, 'utf-8'));
const resumeAttempts: { workflowId: string }[] = session.session?.resumeAttempts ?? [];
// Fresh: session.json appears with originalWorkflowId. Resume: new resumeAttempts entry.
const ready = isResume ? resumeAttempts.length > initialResumeCount : !!session.session?.originalWorkflowId;
if (ready) {
started = true;
spinner.stop(`Scan started — ${workspace}`);
printInfo(args, workspace, repo.hostPath, workspacesDir);
if (args.follow) {
await followScan(workspace, workspacesDir);
}
return;
}
} catch {
// File doesn't exist yet
}
await sleep(2000);
}
spinner.error('Timed out waiting for the scan to start');
process.exit(1);
}
function printDebugHint(containerName: string): void {
/**
* Follow a just-started scan (for `--follow`, aimed at CI): stream its log while Temporal drives
* completion, then exit on the workflow outcome — 0 if the assessment ran, 1 if the scan failed.
* That tracks whether the pipeline ran, not whether vulnerabilities were found. On failure the
* root-cause message is printed so a red CI build says why.
*/
async function followScan(workspace: string, workspacesDir: string): Promise<never> {
const logFile = resolveRunFile(path.join(workspacesDir, workspace), 'workflow.log');
const workflowId = resolveWorkflowId(workspace);
// The worker creates workflow.log as it starts; wait briefly so the first read doesn't
// mistake a not-yet-created file for an already-finished scan.
for (let attempts = 0; attempts < 30 && !fs.existsSync(logFile); attempts++) {
await sleep(1000);
}
if (stdoutIsTerminal()) {
console.error('\n Following scan log (Ctrl-C to stop watching):\n');
}
let temporalUnreachable = false;
const { sawFailure } = await tailUntilComplete(logFile, {
...(workflowId && { workflowId }),
onUnreachable: () => {
temporalUnreachable = true;
},
});
// The tail already printed the diagnostic; reading the outcome would only fail the same way.
if (temporalUnreachable) {
process.exit(1);
}
if (!workflowId) {
fail('Scan finished but its workflow id could not be resolved from session.json.');
}
try {
const outcome = await getTerminalOutcome(workflowId);
if (outcome.kind === 'failed') {
// Print the reason only when the streamed log didn't already show the worker's failure
// summary — otherwise the worker crashed before writing it, and this is the only report.
if (!sawFailure) {
console.error(`\nScan failed:\n${indentFailureSegments(outcome.message)}`);
}
process.exit(1);
}
process.exit(0);
} catch (err) {
const detail = err instanceof Error ? err.message : String(err);
fail('Could not read the scan outcome from Temporal at 127.0.0.1:7233.', ` ${detail}`);
}
}
function printPreservedContainerHint(containerName: string): void {
console.log('');
console.log(` Worker container preserved: ${containerName}`);
console.log(` Inspect logs: docker logs ${containerName}`);
console.log(` Remove: docker rm ${containerName}`);
console.log(field(` Worker container preserved: ${containerName}`));
console.log(field(` Inspect logs: docker logs ${containerName}`));
console.log(field(` Remove: docker rm ${containerName}`));
console.log('');
}
function printInfo(
args: StartArgs,
workspace: string,
workflowId: string,
repoPath: string,
workspacesDir: string,
): void {
const logsCmd = isLocal() ? `./shannon logs ${workspace}` : `npx @keygraph/shannon logs ${workspace}`;
const reportPath = path.join(workspacesDir, workspace, FINAL_REPORT_FILENAME);
function printInfo(args: StartArgs, workspace: string, repoPath: string, workspacesDir: string): void {
const interactive = stdoutIsTerminal();
console.log(' Scan started — it runs in the background, so you can close this terminal.');
console.log('');
console.log(` Target: ${args.url}`);
console.log(` Repository: ${repoPath}`);
console.log(` Workspace: ${workspace}`);
if (interactive && !args.follow) {
console.log(' It runs in the background — you can close this terminal.');
console.log('');
}
console.log(field(` Target: ${args.url}`));
console.log(field(` Repository: ${interactive ? repoPath : path.basename(repoPath)}`));
console.log(field(` Workspace: ${workspace}`));
if (args.config) {
console.log(` Config: ${path.resolve(args.config)}`);
console.log(field(` Config: ${interactive ? path.resolve(args.config) : path.basename(args.config)}`));
}
if (args.pipelineTesting) {
console.log(' Mode: Pipeline Testing');
console.log(field(' Mode: Pipeline Testing'));
}
const spec = resolveModelSpec();
if (typeof spec !== 'string') {
console.log(` Model: ${spec.providerId}:${spec.modelId}`);
console.log(field(` Model: ${spec.providerId}:${spec.modelId}`));
}
if (!interactive) {
return;
}
const reportPath = path.join(workspacesDir, workspace, FINAL_REPORT_PDF_FILENAME);
// When following, the scan log streams inline next, so the "run these to watch it" hints
// would only contradict that.
if (!args.follow) {
const prefix = commandPrefix();
console.log('');
console.log(rule('Watch scan progress:'));
console.log(field(` Live logs: ${prefix} logs ${workspace}`));
console.log(field(` Progress: ${prefix} status ${workspace}`));
}
console.log('');
console.log(' Watch scan progress:');
console.log(` Live logs: ${logsCmd}`);
if (workflowId) {
console.log(` Dashboard: http://localhost:8233/namespaces/default/workflows/${workflowId}`);
} else {
console.log(' Dashboard: http://localhost:8233');
}
console.log('');
console.log(' Report (when the scan finishes):');
console.log(rule('Report (when the scan finishes):'));
console.log(` ${reportPath}`);
console.log('');
}
+188 -16
View File
@@ -1,24 +1,196 @@
/**
* `shannon status` command — show running scans and Temporal health.
* `shannon status <workspace>` — one scan's live progress from Temporal.
*
* While the scan runs, polls Temporal and redraws the phase/agent tree on a
* terminal (a pipe or a finished scan gets a single frame). When the scan reaches
* a terminal state, prints the overall result and exits. Reads Temporal directly —
* no worker, no session files — so it needs Temporal up and shows scans within its
* ~24h retention window.
*/
import { isTemporalReady, listRunningWorkers } from '../docker.js';
import { setTimeout as sleep } from 'node:timers/promises';
import { fail } from '../errors.js';
import { isLocal } from '../mode.js';
import { type RenderInput, renderScan } from '../scan/render.js';
import { toStatusJson } from '../scan/status-json.js';
import { resolveWorkflowId } from '../session.js';
import { displaySplash } from '../splash.js';
import { describeScan, getTerminalOutcome, queryProgress, type ScanDescription } from '../temporal-client.js';
import { stdoutIsTerminal, supportsColor } from '../tty.js';
import { getVersion } from '../version.js';
export function status(): void {
// 1. Temporal health
const temporalUp = isTemporalReady();
console.log(`Temporal: ${temporalUp ? 'running' : 'not running'}`);
if (temporalUp) {
console.log(' Dashboard: http://localhost:8233');
const HIDE_CURSOR = '\x1b[?25l';
const SHOW_CURSOR = '\x1b[?25h';
/** Redraw cadence for the spinner animation; data is refreshed on the slower poll. */
const RENDER_MS = 120;
const POLL_MS = 1200;
/** Terminal = anything other than an open, running execution. */
function isTerminalStatus(status: string): boolean {
return status !== 'RUNNING' && status !== 'UNSPECIFIED';
}
// Match SGR color escapes (ESC[…m) so a line's on-screen width excludes them. Built from the ESC
// char code so the source carries no literal control character.
const ANSI_PATTERN = new RegExp(`${String.fromCharCode(27)}\\[[0-9;]*m`, 'g');
/**
* Physical terminal rows a frame occupies, so the live redraw moves the cursor up by the right
* amount. A line wider than the terminal wraps onto extra rows, so counting logical lines alone
* undercounts and the redraw drifts downward. Color escapes don't take screen columns, so strip them.
*/
function physicalRows(frame: string): number {
const columns = process.stdout.columns || 80;
return frame.split('\n').reduce((rows, line) => {
const width = line.replace(ANSI_PATTERN, '').length;
return rows + Math.max(1, Math.ceil(width / columns));
}, 0);
}
function exitCodeFor(input: RenderInput): number {
if (input.temporalStatus === 'FAILED' || input.temporalStatus === 'TIMED_OUT') return 1;
if (input.state?.status === 'failed') return 1;
return 0;
}
/** Live view of a running scan: its progress query plus the in-flight agents from describe. */
async function buildRunningInput(workspace: string, workflowId: string, desc: ScanDescription): Promise<RenderInput> {
const state = await queryProgress(workflowId);
return {
workspace,
workflowId,
temporalStatus: desc.status,
state,
running: desc.runningAgents,
...(desc.startedAt !== undefined && { startedAt: desc.startedAt }),
};
}
/** Final view of a closed scan: its result (or the failure) plus timing from describe. */
async function buildTerminalInput(workspace: string, workflowId: string, desc: ScanDescription): Promise<RenderInput> {
const outcome = await getTerminalOutcome(workflowId);
const timing = {
...(desc.startedAt !== undefined && { startedAt: desc.startedAt }),
...(desc.closedAt !== undefined && { endedAt: desc.closedAt }),
};
if (outcome.kind === 'success') {
return { workspace, workflowId, temporalStatus: desc.status, state: outcome.state, running: [], ...timing };
}
console.log('');
return {
workspace,
workflowId,
temporalStatus: desc.status,
state: null,
running: [],
failureMessage: outcome.message,
...timing,
};
}
// 2. Running scans
const workers = listRunningWorkers();
if (workers) {
console.log('Running scans:');
console.log(workers);
} else {
console.log('No scans running');
function printFrame(input: RenderInput): void {
const frame = renderScan(input, {
now: Date.now(),
color: supportsColor(),
unicode: stdoutIsTerminal(),
live: false,
frame: 0,
});
process.stdout.write(`${frame}\n`);
}
/**
* Poll Temporal and redraw until the scan reaches a terminal state, then print the
* final frame and exit. A fast ticker animates the running spinner off the cached
* snapshot; the network poll refreshes that snapshot on a slower cadence.
*/
async function watch(workspace: string, workflowId: string): Promise<never> {
let prevRows = 0;
let frame = 0;
let cached: RenderInput | null = null;
const draw = (input: RenderInput, live: boolean): void => {
const out = renderScan(input, { now: Date.now(), color: supportsColor(), unicode: true, live, frame });
if (prevRows > 0) process.stdout.write(`\x1b[${prevRows}A\x1b[0J`);
process.stdout.write(`${out}\n`);
prevRows = physicalRows(out);
};
process.on('exit', () => process.stdout.write(SHOW_CURSOR));
process.on('SIGINT', () => {
process.stdout.write('\n');
process.exit(0);
});
process.stdout.write(HIDE_CURSOR);
const ticker = setInterval(() => {
frame++;
if (cached) draw(cached, true);
}, RENDER_MS);
for (;;) {
const desc = await describeScan(workflowId);
if (!desc) {
clearInterval(ticker);
fail(`Scan "${workspace}" is no longer in Temporal.`);
}
if (isTerminalStatus(desc.status)) {
clearInterval(ticker);
const input = await buildTerminalInput(workspace, workflowId, desc);
draw(input, false);
process.exit(exitCodeFor(input));
}
cached = await buildRunningInput(workspace, workflowId, desc);
await sleep(POLL_MS);
}
}
/** Read one point-in-time snapshot from Temporal: the terminal result if closed, else live progress. */
async function snapshot(workspace: string, workflowId: string, desc: ScanDescription): Promise<RenderInput> {
return isTerminalStatus(desc.status)
? buildTerminalInput(workspace, workflowId, desc)
: buildRunningInput(workspace, workflowId, desc);
}
export async function status(workspace: string, opts: { readonly json: boolean }): Promise<void> {
// A resume spawns a new workflow id (recorded in session.json); resolve through there so status
// follows the current resume, not the superseded original. Fresh scans: the name is the id.
const workflowId = resolveWorkflowId(workspace) ?? workspace;
let desc: ScanDescription | null;
try {
desc = await describeScan(workflowId);
} catch {
fail('Could not reach Temporal at 127.0.0.1:7233.', 'Start Temporal (it comes up with a scan) and try again.');
}
if (!desc) {
fail(
`No scan found for "${workspace}".`,
'',
'Scans are visible while running and for ~24h after they finish (Temporal retention).',
);
}
// --json is always a single snapshot then exit, even on a TTY — it never enters the live watch loop.
if (opts.json) {
const input = await snapshot(workspace, workflowId, desc);
process.stdout.write(`${JSON.stringify(toStatusJson(input, Date.now()), null, 2)}\n`);
process.exit(exitCodeFor(input));
}
// Human-facing views open with the splash; skip it off a real terminal so piped output stays clean.
if (stdoutIsTerminal()) {
displaySplash(isLocal() ? undefined : getVersion());
}
// A finished scan, or output that isn't a live terminal, gets a single frame.
if (isTerminalStatus(desc.status) || !stdoutIsTerminal()) {
const input = await snapshot(workspace, workflowId, desc);
printFrame(input);
process.exit(exitCodeFor(input));
}
await watch(workspace, workflowId);
}
+118 -14
View File
@@ -1,23 +1,127 @@
/**
* `shannon stop` command — stop workers and infrastructure.
* `shannon stop` command — stop one scan by workspace, or every scan with --all.
* Never touches infra or data; to wipe Temporal state entirely, use `shannon reset`.
*/
import * as p from '@clack/prompts';
import { stopInfra, stopWorkers } from '../docker.js';
import { requireInteractive } from '../tty.js';
import { confirmOrExit } from '../confirm.js';
import {
anyRunningScanWorkflow,
ensureDocker,
isTemporalReady,
isWorkflowRunning,
runningContainers,
scanFilter,
stopContainers,
terminateAllWorkflows,
terminateWorkflow,
WORKER_FILTER,
} from '../docker.js';
import { fail, failUsage, warn } from '../errors.js';
import { commandPrefix } from '../mode.js';
import { resolveWorkflowId } from '../session.js';
export async function stop(clean: boolean, yes: boolean): Promise<void> {
if (clean && !yes) {
requireInteractive('stop --clean', 'Re-run with --yes to skip this confirmation.');
const confirmed = await p.confirm({
message: 'This will stop all running scans and remove the Temporal data. Continue?',
});
if (p.isCancel(confirmed) || !confirmed) {
p.cancel('Aborted.');
process.exit(0);
export interface StopOptions {
all: boolean;
yes: boolean;
workspace?: string;
}
/**
* Stop a single scan. Terminating the workflow both clears Temporal's record and
* brings the container down (the worker waits on the workflow result), so that runs
* first; `docker stop` is the fallback for the pre-registration window and an
* unreachable Temporal. The stop is then verified rather than assumed.
*/
async function stopSingleScan(workspace: string, yes: boolean): Promise<void> {
const workflowId = resolveWorkflowId(workspace);
const filter = scanFilter(workspace);
const temporalUp = isTemporalReady();
const initialContainers = runningContainers(filter);
const workflowRunning = Boolean(workflowId && temporalUp && isWorkflowRunning(workflowId));
// Resolve what is running before prompting, so we never confirm a no-op.
if (initialContainers.length === 0 && !workflowRunning) {
if (!workflowId) {
fail(`No scan found for workspace: ${workspace}`);
}
console.log(`Nothing was running for ${workspace}.`);
return;
}
stopWorkers();
stopInfra(clean);
await confirmOrExit('stop', `Stop the scan "${workspace}"?`, yes);
const spinner = p.spinner();
spinner.start(`Stopping scan ${workspace}`);
if (workflowId && workflowRunning) {
terminateWorkflow(workflowId, `Stopped via shannon stop ${workspace}`);
}
await stopContainers(runningContainers(filter));
const stillRunning = runningContainers(filter);
if (stillRunning.length > 0) {
spinner.error(`Scan ${workspace} may still be running`);
console.error(`${stillRunning.length} container(s) did not stop. Retry: ${commandPrefix()} stop ${workspace}`);
process.exit(1);
}
spinner.stop(`Stopped scan ${workspace}`);
if (workflowId && temporalUp && isWorkflowRunning(workflowId)) {
warn(`scan ${workspace} stopped, but its workflow is still Running in Temporal.`);
}
}
async function stopAllScans(yes: boolean): Promise<void> {
const temporalUp = isTemporalReady();
const initial = runningContainers(WORKER_FILTER);
// Resolve what is running before prompting, so we never confirm a no-op.
if (initial.length === 0) {
console.log('No running scans to stop.');
return;
}
await confirmOrExit('stop', 'This will stop all running scans. Continue?', yes);
const spinner = p.spinner();
spinner.start('Stopping all scans');
if (temporalUp) {
terminateAllWorkflows('Stopped via shannon stop --all');
}
await stopContainers(runningContainers(WORKER_FILTER));
const stillRunning = runningContainers(WORKER_FILTER);
if (stillRunning.length > 0) {
spinner.error(`Stopped ${initial.length - stillRunning.length} of ${initial.length} scans`);
console.error(`${stillRunning.length} container(s) did not stop. Retry: ${commandPrefix()} stop --all`);
process.exit(1);
}
spinner.stop(`Stopped ${initial.length} scan${initial.length === 1 ? '' : 's'}`);
if (temporalUp && anyRunningScanWorkflow()) {
warn('some scan workflows are still Running in Temporal — check http://localhost:8233');
}
}
export async function stop(opts: StopOptions): Promise<void> {
ensureDocker();
// Validate the target: exactly one of <workspace> or --all.
if (opts.all && opts.workspace) {
failUsage('Pass a workspace name or --all, not both.');
}
if (!opts.all && !opts.workspace) {
failUsage('Specify which scan to stop: `stop <workspace>`, or `stop --all` to stop every scan.');
}
if (opts.workspace) {
await stopSingleScan(opts.workspace, opts.yes);
} else {
await stopAllScans(opts.yes);
}
}
-55
View File
@@ -1,55 +0,0 @@
/**
* `npx @keygraph/shannon uninstall` command — remove ~/.shannon/ after confirmation (npx only).
*/
import fs from 'node:fs';
import os from 'node:os';
import path from 'node:path';
import * as p from '@clack/prompts';
import { stopInfra, stopWorkers } from '../docker.js';
import { requireInteractive } from '../tty.js';
const SHANNON_HOME = path.join(os.homedir(), '.shannon');
export async function uninstall(yes: boolean): Promise<void> {
const interactive = !yes;
if (interactive) p.intro('Shannon Uninstall');
if (!fs.existsSync(SHANNON_HOME)) {
const message = 'Nothing to remove. Shannon is not configured on this machine.';
if (interactive) {
p.log.info(message);
p.outro('Done.');
} else {
console.log(message);
}
return;
}
if (interactive) {
requireInteractive('uninstall', 'Re-run with --yes to skip this confirmation.');
const confirmed = await p.confirm({
message: 'This will permanently remove all past scan data, saved configurations, and API keys. Continue?',
});
if (p.isCancel(confirmed) || !confirmed) {
p.cancel('Aborted.');
process.exit(0);
}
}
// Stop any running containers first
stopWorkers();
stopInfra(false);
fs.rmSync(SHANNON_HOME, { recursive: true, force: true });
const done = 'All Shannon data has been removed.';
const hint = 'Shannon has been uninstalled. Run `npx @keygraph/shannon setup` to start fresh.';
if (interactive) {
p.log.success(done);
p.outro(hint);
} else {
console.log(done);
console.log(hint);
}
}
-35
View File
@@ -1,35 +0,0 @@
/**
* `shannon workspaces` command — list all workspaces.
*/
import { execFileSync } from 'node:child_process';
import os from 'node:os';
import { getWorkerImage } from '../docker.js';
import { getWorkspacesDir } from '../home.js';
export function workspaces(version: string): void {
const workspacesDir = getWorkspacesDir();
const image = getWorkerImage(version);
try {
execFileSync(
'docker',
[
'run',
'--rm',
'-v',
`${workspacesDir}:/app/workspaces`,
'-e',
'WORKSPACES_DIR=/app/workspaces',
image,
'node',
'apps/worker/dist/temporal/workspaces.js',
],
{ stdio: 'inherit', ...(os.platform() === 'win32' && { env: { ...process.env, MSYS_NO_PATHCONV: '1' } }) },
);
} catch {
console.error('ERROR: Failed to list workspaces. Is the Docker image available?');
console.error(` Run: docker pull ${image}`);
process.exit(1);
}
}
+9 -12
View File
@@ -7,6 +7,7 @@
import fs from 'node:fs';
import { parse as parseTOML } from 'smol-toml';
import { fail } from '../errors.js';
import { getConfigFile } from '../home.js';
import { getMode } from '../mode.js';
import {
@@ -100,10 +101,9 @@ function loadTOML(): TOMLConfig | null {
const mode = fs.statSync(configPath).mode;
if (mode & 0o077) {
const actual = (mode & 0o777).toString(8).padStart(3, '0');
console.error(
`\nYour config file is readable by other users on this machine (${actual}). Lock it down: chmod 600 ${configPath}\n`,
fail(
`Your config file is readable by other users on this machine (${actual}). Lock it down: chmod 600 ${configPath}`,
);
process.exit(1);
}
}
@@ -112,9 +112,7 @@ function loadTOML(): TOMLConfig | null {
return parseTOML(content) as TOMLConfig;
} catch (err) {
const message = err instanceof Error ? err.message : String(err);
console.error(`\nFailed to parse ${configPath}: ${message}`);
console.error(`\nRun 'npx @keygraph/shannon setup' to reconfigure.\n`);
process.exit(1);
fail(`Failed to parse ${configPath}: ${message}`, `Run 'npx @keygraph/shannon setup' to reconfigure.`);
}
}
@@ -256,12 +254,11 @@ export function resolveConfig(): void {
// Validate before injecting
const errors = validateConfig(toml);
if (errors.length > 0) {
console.error('\nInvalid configuration:');
for (const err of errors) {
console.error(` - ${err}`);
}
console.error(`\nRun 'npx @keygraph/shannon setup' to reconfigure.\n`);
process.exit(1);
fail(
'Invalid configuration:',
...errors.map((err) => ` - ${err}`),
`Run 'npx @keygraph/shannon setup' to reconfigure.`,
);
}
for (const mapping of CONFIG_MAP) {
+43
View File
@@ -0,0 +1,43 @@
/**
* Shared confirmation prompt for destructive or batch commands.
*
* `stop` and `reset` gate their action behind the same "confirm unless --yes"
* flow. Centralizing it here keeps the behavior identical across commands and
* impossible to change in only one place by accident.
*/
import * as p from '@clack/prompts';
import { requireInteractive } from './tty.js';
/**
* Ask the user to confirm an action, unless `yes` was passed. Off a TTY without
* `--yes`, fails fast rather than hanging on a prompt. Exits 0 if the user declines.
*/
export async function confirmOrExit(command: string, message: string, yes: boolean): Promise<void> {
if (yes) {
return;
}
requireInteractive(command, 'Re-run with --yes to skip this confirmation.');
const confirmed = await p.confirm({ message });
if (p.isCancel(confirmed) || !confirmed) {
p.cancel('Aborted.');
process.exit(0);
}
}
/**
* Severe-tier confirmation: the user must type `word` exactly to proceed. Unlike
* `confirmOrExit` there is no `--yes` bypass. Off a TTY it fails fast; exits 0 if declined.
*/
export async function confirmByTyping(command: string, word: string): Promise<void> {
requireInteractive(command, `'${command}' cannot be run non-interactively.`);
const typed = await p.text({
message: `Type ${word} to confirm — this cannot be undone:`,
validate: (value) => (value === word ? undefined : `Type ${word} to proceed, or press Ctrl-C to abort.`),
});
if (p.isCancel(typed) || typed !== word) {
p.cancel('Aborted.');
process.exit(0);
}
}
+146 -49
View File
@@ -12,14 +12,21 @@ import os from 'node:os';
import path from 'node:path';
import { setTimeout as sleep } from 'node:timers/promises';
import { fileURLToPath } from 'node:url';
import type { SpinnerResult } from '@clack/prompts';
import { envBool, PI_AUTH_CONTAINER_PATH } from './env.js';
import { fail } from './errors.js';
import { getMode, isDevMode } from './mode.js';
import { INTERNAL_DIR } from './paths.js';
import { runStep, spawnCaptured, surfaceOutput } from './ui.js';
const __dirname = path.dirname(fileURLToPath(import.meta.url));
const NPX_IMAGE_REPO = 'keygraph/shannon';
const DEV_IMAGE = 'shannon-worker';
/** Docker label stamped on each worker container, mapping it back to its workspace so a single scan can be stopped by name. */
const WORKSPACE_LABEL = 'shannon.workspace';
export function getWorkerImage(version: string): string {
return getMode() === 'local' ? DEV_IMAGE : `${NPX_IMAGE_REPO}:${version}`;
}
@@ -65,44 +72,77 @@ function runOutput(cmd: string, args: string[]): string {
}
}
/** Run a command asynchronously, resolving true on success. Never rejects. */
function spawnQuiet(cmd: string, args: string[]): Promise<boolean> {
return new Promise((resolve) => {
const child = spawn(cmd, args, { stdio: 'ignore' });
child.on('close', (code) => resolve(code === 0));
child.on('error', () => resolve(false));
});
}
const TEMPORAL_CONTAINER = 'shannon-temporal';
const TEMPORAL_ADDRESS = 'localhost:7233';
/** Query matching every running pentest scan workflow. */
const RUNNING_SCAN_QUERY = "ExecutionStatus = 'Running' AND WorkflowType = 'pentestPipelineWorkflow'";
/** Build `docker exec` args for a `temporal` CLI command run inside the Temporal container. */
function temporalCmd(...args: string[]): string[] {
return ['exec', TEMPORAL_CONTAINER, 'temporal', ...args, '--address', TEMPORAL_ADDRESS];
}
/**
* Verify Docker is installed and its daemon is running, exiting otherwise.
* `docker info` succeeds only when both are true. Call this before any command
* that shells out to Docker.
*/
export function ensureDocker(): void {
try {
execFileSync('docker', ['info'], { stdio: 'pipe' });
} catch {
fail(
'Docker must be installed and running. Start Docker and try again.',
'Install Docker: https://docs.docker.com/get-docker/',
);
}
}
/**
* Check if Temporal is running and healthy.
*/
export function isTemporalReady(): boolean {
const output = runOutput('docker', [
'exec',
'shannon-temporal',
'temporal',
'operator',
'cluster',
'health',
'--address',
'localhost:7233',
]);
const output = runOutput('docker', temporalCmd('operator', 'cluster', 'health'));
return output.includes('SERVING');
}
/**
* Ensure Temporal is running via compose.
*/
export async function ensureInfra(): Promise<void> {
export async function ensureInfra(spinner: SpinnerResult): Promise<void> {
if (isTemporalReady()) {
return;
}
// Drive the caller's spinner — the whole "start" flow is one spinner, not several.
spinner.message('Starting Temporal');
const composeFile = getComposeFile();
console.log('Starting Shannon infrastructure...');
execFileSync('docker', ['compose', '-f', composeFile, 'up', '-d'], { stdio: 'inherit' });
const result = await spawnCaptured('docker', ['compose', '-f', composeFile, 'up', '-d']);
if (!result.ok) {
spinner.error('Could not start Temporal');
surfaceOutput(result.output);
process.exit(1);
}
console.log('Waiting for Temporal to be ready...');
spinner.message('Waiting for Temporal to be ready');
for (let i = 0; i < 30; i++) {
if (isTemporalReady()) {
console.log('Temporal is ready!');
return;
}
await sleep(2000);
}
console.error('Timeout waiting for Temporal');
spinner.error('Temporal did not become ready in time');
process.exit(1);
}
@@ -137,10 +177,11 @@ export function ensureImage(version: string): void {
try {
execFileSync('docker', ['pull', image], { stdio: 'inherit' });
} catch {
console.error(`\nERROR: Failed to pull ${image}`);
console.error('The image may not be available for your platform yet.');
console.error('Check https://hub.docker.com/r/keygraph/shannon for available tags.');
process.exit(1);
fail(
`Failed to pull ${image}`,
'The image may not be available for your platform yet.',
'Check https://hub.docker.com/r/keygraph/shannon for available tags.',
);
}
pruneOldImages(version);
}
@@ -203,7 +244,7 @@ function shouldSkipHostsName(name: string, hostname: string): boolean {
* `host-gateway` so they target the host's loopback instead of the container's.
*/
function forwardEtcHostsFlags(): string[] {
if (process.env.SHANNON_FORWARD_HOSTS === 'false') return [];
if (!envBool('SHANNON_FORWARD_HOSTS', true)) return [];
if (os.platform() === 'win32') return [];
let content: string;
@@ -254,20 +295,24 @@ export interface WorkerOptions {
outputDir?: string;
workspace: string;
pipelineTesting?: boolean;
debug?: boolean;
keepContainer?: boolean;
piAuthHostPath?: string;
}
/**
* Spawn the worker container in detached mode and return the process.
* When `opts.debug` is true, omits `--rm` so the container persists for log inspection.
* When `opts.keepContainer` is true, omits `--rm` so the container persists for log inspection.
*/
export function spawnWorker(opts: WorkerOptions): ChildProcess {
const args = ['run', '-d'];
if (!opts.debug) {
if (!opts.keepContainer) {
args.push('--rm');
}
args.push('--name', opts.containerName, '--network', 'shannon-net');
// Tag with the workspace so `stop <workspace>` can target this scan's container
args.push('--label', `${WORKSPACE_LABEL}=${opts.workspace}`);
// Add host flag for Linux
args.push(...addHostFlag());
@@ -305,6 +350,11 @@ export function spawnWorker(opts: WorkerOptions): ChildProcess {
args.push('-v', `${opts.outputDir}:/app/output`);
}
// Reuse the host's pi credentials: mount only the auth file, allowing token refreshes to persist.
if (opts.piAuthHostPath) {
args.push('-v', `${opts.piAuthHostPath}:${PI_AUTH_CONTAINER_PATH}`);
}
// Environment
args.push(...opts.envFlags);
@@ -337,26 +387,86 @@ export function spawnWorker(opts: WorkerOptions): ChildProcess {
});
}
/**
* Stop all running shannon-worker-* containers.
*/
export function stopWorkers(): void {
const workers = runOutput('docker', ['ps', '-q', '--filter', 'name=shannon-worker-']);
if (!workers) return;
/** `docker ps --filter` args matching every running worker container. */
export const WORKER_FILTER: readonly string[] = ['--filter', 'name=shannon-worker-'];
const ids = workers.split('\n').filter(Boolean);
console.log('Stopping running scans...');
execFileSync('docker', ['stop', ...ids], { stdio: 'inherit' });
/** `docker ps --filter` args matching one scan's worker container(s), by workspace label. */
export function scanFilter(workspace: string): readonly string[] {
return ['--filter', `label=${WORKSPACE_LABEL}=${workspace}`];
}
/**
* Tear down the compose stack.
* IDs of running containers matching the filter. Re-querying this after a stop is
* the authoritative check for whether containers actually stopped — `docker stop`'s
* exit code can't distinguish "already gone" from "failed to stop".
*/
export function stopInfra(clean: boolean): void {
export function runningContainers(filter: readonly string[]): string[] {
const output = runOutput('docker', ['ps', '-q', ...filter]);
return output.split('\n').filter(Boolean);
}
/**
* Stop containers by ID, tolerating any that vanished between being listed and
* stopped (a `--rm` worker exiting is success, not an error). Async so a spinner
* can animate during docker's graceful-shutdown wait.
*/
export async function stopContainers(ids: string[]): Promise<void> {
await Promise.all(ids.map((id) => spawnQuiet('docker', ['stop', id])));
}
/**
* Terminate a Temporal workflow so a stopped scan doesn't linger as a running
* workflow with no worker. Best-effort: returns false if Temporal is unreachable
* or the workflow already closed. Requires Temporal to be up (guard with isTemporalReady).
*/
export function terminateWorkflow(workflowId: string, reason: string): boolean {
return runQuiet('docker', temporalCmd('workflow', 'terminate', '--workflow-id', workflowId, '--reason', reason));
}
/**
* Terminate every running pentest workflow in one batch, so `stop --all` doesn't
* leave workflows running with no worker. Best-effort: returns false if Temporal
* is unreachable. Requires Temporal to be up (guard with isTemporalReady).
*/
export function terminateAllWorkflows(reason: string): boolean {
return runQuiet(
'docker',
temporalCmd('workflow', 'terminate', '--query', RUNNING_SCAN_QUERY, '--reason', reason, '--yes'),
);
}
/**
* Whether a specific workflow is still in the Running state. Re-querying this after
* a terminate verifies it actually took effect, rather than trusting the terminate
* command's exit code. Requires Temporal to be up (guard with isTemporalReady).
*/
export function isWorkflowRunning(workflowId: string): boolean {
const query = `WorkflowId = '${workflowId}' AND ExecutionStatus = 'Running'`;
const output = runOutput('docker', temporalCmd('workflow', 'list', '--query', query));
return output.includes(workflowId);
}
/**
* Whether any pentest scan workflow is still Running — the `stop --all` counterpart
* to isWorkflowRunning. Requires Temporal to be up (guard with isTemporalReady).
*/
export function anyRunningScanWorkflow(): boolean {
const output = runOutput('docker', temporalCmd('workflow', 'list', '--query', RUNNING_SCAN_QUERY));
return output.includes('pentestPipelineWorkflow');
}
/**
* Tear down the compose stack. When `clean` is set, volumes are removed too.
*/
export async function stopInfra(clean: boolean): Promise<void> {
const composeFile = getComposeFile();
const args = ['compose', '-f', composeFile, 'down'];
if (clean) args.push('-v');
execFileSync('docker', args, { stdio: 'inherit' });
const label = clean ? 'Removing Temporal data and volumes' : 'Stopping Temporal';
const step = await runStep(label, 'docker', args);
if (!step.ok) {
fail(`${label} failed. See the output above.`);
}
}
/**
@@ -372,16 +482,3 @@ function pruneOldImages(currentVersion: string): void {
runQuiet('docker', ['rmi', `${NPX_IMAGE_REPO}:${tag}`]);
}
}
/**
* List running worker containers.
*/
export function listRunningWorkers(): string {
return runOutput('docker', [
'ps',
'--filter',
'name=shannon-worker-',
'--format',
'table {{.Names}}\t{{.Status}}\t{{.RunningFor}}',
]);
}
+70 -7
View File
@@ -5,6 +5,9 @@
* NPX mode: fills gaps from ~/.shannon/config.toml (no .env).
*/
import fs from 'node:fs';
import os from 'node:os';
import path from 'node:path';
import dotenv from 'dotenv';
import { resolveConfig } from './config/resolver.js';
import { getMode } from './mode.js';
@@ -41,6 +44,34 @@ function providerForwardVars(providerId: string): readonly string[] {
return [...PROVIDER_API_KEY_ENV[providerId], ...PROVIDER_EXTRA_ENV[providerId]];
}
/** Parse a user-facing boolean env var: `1`/`true` (any case) true, `0`/`false`/empty false, else the default. */
export function envBool(name: string, defaultValue: boolean): boolean {
const raw = process.env[name]?.trim().toLowerCase();
if (raw === undefined || raw === '') return defaultValue;
if (raw === '1' || raw === 'true') return true;
if (raw === '0' || raw === 'false') return false;
return defaultValue;
}
const USE_PI_AUTH_ENV = 'SHANNON_USE_PI_AUTH';
/** Where the host's auth.json is mounted: pi's standard location (worker HOME is /tmp), read natively. */
export const PI_AUTH_CONTAINER_PATH = '/tmp/.pi/agent/auth.json';
/** Host path to pi's credential file. */
export function resolveHostPiAuthPath(): string {
return path.join(os.homedir(), '.pi', 'agent', 'auth.json');
}
export function piAuthFlagEnabled(): boolean {
return envBool(USE_PI_AUTH_ENV, false);
}
/** Opted into pi auth via the flag, and the auth file exists to mount. */
export function shouldUsePiAuth(): boolean {
return piAuthFlagEnabled() && fs.existsSync(resolveHostPiAuthPath());
}
/**
* Load credentials into process.env.
* Local mode: loads ./.env via dotenv.
@@ -56,8 +87,9 @@ export function loadEnv(): void {
}
/**
* Build `-e KEY=VALUE` flags for docker run. Forwards the common vars plus only
* the selected provider's credentials.
* Build `-e` flags for docker run. Forwards the common vars plus only the
* selected provider's credentials, passed by name (`-e KEY`) so secret values
* stay out of the `docker run` argv; docker inherits them from this process's env.
*/
export function buildEnvFlags(): string[] {
const flags: string[] = ['-e', 'TEMPORAL_ADDRESS=shannon-temporal:7233'];
@@ -66,9 +98,8 @@ export function buildEnvFlags(): string[] {
const providerVars = typeof spec === 'string' ? [] : providerForwardVars(spec.providerId);
for (const key of [...COMMON_FORWARD_VARS, ...providerVars]) {
const value = process.env[key];
if (value) {
flags.push('-e', `${key}=${value}`);
if (process.env[key]) {
flags.push('-e', key);
}
}
@@ -110,6 +141,18 @@ export function validateCredentials(): CredentialValidation {
return { valid: false, error: spec };
}
// Pi-auth: skip the API-key checks, but the host auth file must exist to mount.
if (piAuthFlagEnabled()) {
const authPath = resolveHostPiAuthPath();
if (!fs.existsSync(authPath)) {
return {
valid: false,
error: `${USE_PI_AUTH_ENV} is set but no pi credentials were found at ${authPath}. Authenticate with pi first.`,
};
}
return { valid: true };
}
// 2. The selected provider must have a credential
if (!hasCredential(spec.providerId)) {
const requirement = isCuratedProvider(spec.providerId)
@@ -128,8 +171,28 @@ export function validateCredentials(): CredentialValidation {
// 3. Exactly one provider may be configured. Several complete credentials make
// the scan's provider depend on SHANNON_AI_MODEL alone, which is too easy to
// misread as "both are in play" and too easy to redirect by editing one line.
if (configuredProviders().length > 1) {
return { valid: false, error: 'Credentials for more than one provider are set.' };
const configured = configuredProviders();
if (configured.length > 1) {
const setKeys = (id: CuratedProviderId): string[] =>
PROVIDER_API_KEY_ENV[id].filter((name) => Boolean(process.env[name]));
const list = configured.map((id) => `${id} (${setKeys(id).join(', ')})`).join(' and ');
const others = configured.filter((id) => id !== spec.providerId);
const extraVars = others.flatMap(setKeys);
const dropHint =
getMode() === 'local'
? 'remove them from .env or unset them in your shell:'
: "unset them in your shell, or reconfigure with 'npx @keygraph/shannon setup':";
const lines = [`Credentials for more than one provider are set: ${list}.`];
if (extraVars.length > 0) {
lines.push(
`Shannon runs one provider per scan, selected by SHANNON_AI_MODEL ("${spec.providerId}:...").`,
`Keep ${spec.providerId} and drop the rest — ${dropHint}`,
` unset ${extraVars.join(' ')}`,
);
}
return { valid: false, error: lines.join('\n') };
}
return { valid: true };
+70
View File
@@ -0,0 +1,70 @@
/**
* Centralized error reporting.
*
* `fail` — an expected, user-fixable error (bad input, missing prerequisite):
* a clean message on stderr and a non-zero exit, never a stack trace.
* `failUsage` — a malformed invocation (unknown command, bad or missing
* arguments): the same clean message, but a distinct exit code so callers can
* tell a usage mistake from an operational failure.
* `crash` — an unexpected error (a bug): a brief message, the full stack written
* to a log file for a bug report, and a pointer to the issue tracker.
*/
import fs from 'node:fs';
import os from 'node:os';
import path from 'node:path';
const ISSUES_URL = 'https://github.com/KeygraphHQ/shannon/issues';
/** Report an expected, user-fixable error (with optional extra lines) and exit non-zero. */
export function fail(message: string, ...hints: string[]): never {
console.error(`ERROR: ${message}`);
for (const hint of hints) {
console.error(hint);
}
process.exit(1);
}
/** Report a usage/argument error (with optional extra lines) and exit 2. */
export function failUsage(message: string, ...hints: string[]): never {
console.error(`ERROR: ${message}`);
for (const hint of hints) {
console.error(hint);
}
process.exit(2);
}
/** Report a non-fatal warning on stderr (with optional extra lines) without exiting. */
export function warn(message: string, ...hints: string[]): void {
console.error(`WARNING: ${message}`);
for (const hint of hints) {
console.error(hint);
}
}
/** Report an unexpected error: brief message, full stack to a log file, plus the issue link. */
export function crash(error: unknown): never {
console.error(`ERROR: ${error instanceof Error ? error.message : String(error)}`);
if (process.env.DEBUG) {
console.error(error instanceof Error ? error.stack : String(error));
}
const logPath = writeCrashLog(error);
if (logPath) {
console.error(`Details written to ${logPath}`);
}
console.error(`If this looks like a bug, please report it: ${ISSUES_URL}`);
process.exit(1);
}
/** Write the full error and stack to a log file; return its path, or null if it can't be written. */
function writeCrashLog(error: unknown): string | null {
try {
const logPath = path.join(os.tmpdir(), 'shannon-error.log');
const detail = error instanceof Error && error.stack ? error.stack : String(error);
fs.writeFileSync(logPath, `${new Date().toISOString()}\n${detail}\n`);
return logPath;
} catch {
return null;
}
}
+145
View File
@@ -0,0 +1,145 @@
/**
* Per-command help text.
*
* `shannon <command> --help`, `shannon <command> -h`, and `shannon help <command>`
* all render the matching command's usage, so a user can discover a command's
* flags without scanning the global help. The global help lives in index.ts.
*/
import { commandPrefix, getMode } from './mode.js';
interface CommandHelp {
readonly usage: readonly string[];
readonly description: string;
readonly options?: readonly (readonly [string, string])[];
readonly examples?: readonly string[];
}
const YES_OPTION: readonly [string, string] = [
'-y, --yes',
'Skip the confirmation prompt (required for non-interactive use)',
];
const HELP_OPTION: readonly [string, string] = ['-h, --help', 'Show this help'];
/**
* `start`'s flags, the single source rendered by both the per-command help here
* and the global help in index.ts, so the two can never drift.
*/
export const START_OPTIONS: readonly (readonly [string, string])[] = [
['-u, --url <url>', 'Target URL (required)'],
['-r, --repo <path>', 'Repository path (required)'],
['-c, --config <path>', 'Configuration file (YAML)'],
['-o, --output <path>', 'Copy deliverables to this directory after the run'],
['-w, --workspace <name>', 'Named workspace (auto-resumes if it exists)'],
['-f, --follow', 'Stream the scan log until it finishes'],
['--pipeline-testing', 'Use minimal prompts for fast testing'],
['--keep-container', 'Preserve the worker container after exit for log inspection'],
];
const COMMAND_HELP: Readonly<Record<string, CommandHelp>> = {
start: {
usage: ['start -u <url> -r <path> [options]'],
description: 'Start a pentest scan.',
examples: [
'start -u https://example.com -r ./my-repo',
'start -u https://example.com -r /path/to/repo -c config.yaml -w q1-audit',
'start -u https://example.com -r ./my-repo --follow',
],
},
stop: {
usage: ['stop <workspace> [--yes]', 'stop --all [--yes]'],
description: 'Stop one scan by workspace, or every scan with --all (Temporal stays up).',
options: [['--all', 'Stop all running scans'], YES_OPTION],
examples: ['stop q1-audit', 'stop --all'],
},
reset: {
usage: ['reset'],
description: 'Stop everything and permanently remove all Temporal data and volumes.',
},
logs: {
usage: ['logs <workspace>'],
description: "Tail a scan's live log until it completes.",
examples: ['logs q1-audit'],
},
status: {
usage: ['status <workspace> [--json]'],
description:
"Show one scan's phase-by-phase progress, read live from Temporal. Watches and redraws until the scan finishes on a terminal; prints one frame when piped or already finished. With --json, prints a single machine-readable snapshot and exits.",
options: [['--json', 'Output a point-in-time snapshot as JSON, then exit']],
examples: ['status q1-audit', 'status q1-audit --json'],
},
scans: {
usage: ['scans [--json]'],
description: 'List completed scans and where each report lives.',
options: [['--json', 'Output the scan list as JSON']],
examples: ['scans', 'scans --json'],
},
build: {
usage: ['build [--no-cache]'],
description: 'Build the worker Docker image (local mode only).',
options: [['--no-cache', 'Build without using the Docker layer cache']],
},
setup: {
usage: ['setup'],
description: 'Configure provider credentials interactively (npx mode only).',
},
version: {
usage: ['version [--json]'],
description: 'Show the version. With --json, prints the version and mode as a machine-readable object.',
options: [['--json', 'Output the version and mode as JSON']],
examples: ['version', 'version --json'],
},
};
/** Commands that only exist in one mode; everything else is available in both. */
const MODE_ONLY: Readonly<Record<string, 'local' | 'npx'>> = {
build: 'local',
setup: 'npx',
};
/** Whether a command has its own help page (and so responds to `--help`/`-h`). */
export function isHelpableCommand(command: string): boolean {
return command in COMMAND_HELP;
}
/**
* User-facing command names available in the current mode, for "did you mean?"
* suggestions. Derived from the same table that backs per-command help, so the
* suggestion set can never drift from the commands that actually exist.
*/
export function availableCommands(): readonly string[] {
const mode = getMode();
const commands = Object.keys(COMMAND_HELP).filter((command) => (MODE_ONLY[command] ?? mode) === mode);
return [...commands, 'help'];
}
/** Print the help page for one command. No-op if the command has no page. */
export function printCommandHelp(command: string): void {
const help = COMMAND_HELP[command];
if (!help) return;
const prefix = commandPrefix();
const baseOptions = command === 'start' ? START_OPTIONS : (help.options ?? []);
const options = [...baseOptions, HELP_OPTION];
const flagWidth = Math.max(...options.map(([flag]) => flag.length));
const lines: string[] = ['', help.description, '', 'USAGE'];
for (const line of help.usage) {
lines.push(` ${prefix} ${line}`);
}
lines.push('', 'OPTIONS');
for (const [flag, desc] of options) {
lines.push(` ${flag.padEnd(flagWidth)} ${desc}`);
}
if (help.examples && help.examples.length > 0) {
lines.push('', 'EXAMPLES');
for (const example of help.examples) {
lines.push(` ${prefix} ${example}`);
}
}
lines.push('');
console.log(lines.join('\n'));
}
+206 -178
View File
@@ -1,5 +1,5 @@
/**
* Shannon CLI — AI Penetration Testing Framework
* Shannon CLI — AI Pentester for Web Apps and APIs
*
* Unified CLI supporting two modes:
* Local mode: Run from cloned repo — builds locally, mounts prompts, uses ./workspaces/
@@ -9,15 +9,21 @@
* in the current working directory.
*/
import { ArgError, parseArgs, YES_FLAGS } from './args.js';
import { build } from './commands/build.js';
import { logs } from './commands/logs.js';
import { reset } from './commands/reset.js';
import { scans } from './commands/scans.js';
import { setup } from './commands/setup.js';
import { start } from './commands/start.js';
import { status } from './commands/status.js';
import { stop } from './commands/stop.js';
import { uninstall } from './commands/uninstall.js';
import { workspaces } from './commands/workspaces.js';
import { getMode } from './mode.js';
import { crash, fail, failUsage } from './errors.js';
import { availableCommands, isHelpableCommand, printCommandHelp, START_OPTIONS } from './help.js';
import { commandPrefix, getMode, isLocal, type Mode } from './mode.js';
import { displaySplash } from './splash.js';
import { closestMatch } from './suggest.js';
import { stdoutIsTerminal } from './tty.js';
import { getVersion, getVersionLine } from './version.js';
function blockSudo(): void {
@@ -25,69 +31,73 @@ function blockSudo(): void {
const isRoot = process.geteuid?.() === 0;
if (!isSudo && !isRoot) return;
const linuxHints =
process.platform === 'linux'
? ['Configure Docker to run without sudo first:', 'https://docs.docker.com/engine/install/linux-postinstall']
: [];
if (isSudo) {
console.error('ERROR: Shannon must not be run with sudo.');
console.error('Re-run this command as your normal user.');
} else {
console.error('ERROR: Shannon must not be run as the root user.');
console.error('Switch to a regular user account and re-run this command.');
fail('Shannon must not be run with sudo.', 'Re-run this command as your normal user.', ...linuxHints);
}
if (process.platform === 'linux') {
console.error('Configure Docker to run without sudo first:');
console.error('https://docs.docker.com/engine/install/linux-postinstall');
}
process.exit(1);
fail(
'Shannon must not be run as the root user.',
'Switch to a regular user account and re-run this command.',
...linuxHints,
);
}
function showHelp(): void {
/** Render `start`'s flags for the global help, from the same source as `start --help`. */
function renderStartOptions(): string {
const flagWidth = Math.max(...START_OPTIONS.map(([flag]) => flag.length));
return START_OPTIONS.map(([flag, desc]) => ` ${flag.padEnd(flagWidth)} ${desc}`).join('\n');
}
/**
* Render the command list with the description column aligned. Padding is computed from the
* widest command, so it lines up regardless of the prefix (`npx @keygraph/shannon` vs `./shannon`).
*/
function renderUsage(prefix: string, mode: Mode): string {
const rows: ReadonlyArray<readonly [string, string]> = [
...(mode === 'local' ? [] : [[`${prefix} setup`, 'Configure credentials'] as const]),
[`${prefix} start --url <url> --repo <path> [options]`, 'Start a pentest scan'],
[`${prefix} stop <workspace> [--yes]`, 'Stop one scan'],
[`${prefix} stop --all [--yes]`, 'Stop all scans (Temporal stays up)'],
[`${prefix} reset`, 'Stop everything and wipe all Temporal data'],
[`${prefix} logs <workspace>`, "Show a scan's live log"],
[`${prefix} status <workspace> [--json]`, 'Live phase/agent progress of one scan'],
[`${prefix} scans [--json]`, 'List completed scans and their reports'],
...(mode === 'local' ? [[`${prefix} build [--no-cache]`, 'Build worker image'] as const] : []),
[`${prefix} version [--json]`, 'Show version'],
[`${prefix} help`, 'Show this help'],
];
const commandWidth = Math.max(...rows.map(([command]) => command.length));
return rows.map(([command, desc]) => ` ${command.padEnd(commandWidth)} ${desc}`).join('\n');
}
function showHelp(withSplash: boolean): void {
const mode = getMode();
const prefix = mode === 'local' ? './shannon' : 'npx @keygraph/shannon';
const prefix = commandPrefix();
console.log(`
Shannon - AI Penetration Testing Framework
const header = withSplash ? '' : '\nShannon — AI Pentester by Keygraph\n';
Usage:${
mode === 'local'
? ''
: `
${prefix} setup Configure credentials`
}
${prefix} start --url <url> --repo <path> [options] Start a pentest scan
${prefix} stop [--clean] [--yes] Stop all running scans
${prefix} workspaces List all workspaces
${prefix} logs <workspace> Show a scan's live log
${prefix} status Show running scans${
mode === 'local'
? `
${prefix} build [--no-cache] Build worker image`
: `
${prefix} uninstall [--yes] Remove ~/.shannon/ and all data`
}
${prefix} version Show version
${prefix} help Show this help
console.log(`${header}
Usage:
${renderUsage(prefix, mode)}
Options for 'start':
-u, --url <url> Target URL (required)
-r, --repo <path> Repository path${mode === 'local' ? ' or bare name' : ''} (required)
-c, --config <path> Configuration file (YAML)
-o, --output <path> Copy deliverables to this directory after run
-w, --workspace <name> Named workspace (auto-resumes if exists)
--pipeline-testing Use minimal prompts for fast testing
--debug Preserve worker container after exit for log inspection
${renderStartOptions()}
Examples:
${prefix} start -u https://example.com -r ${mode === 'local' ? 'my-repo' : './my-repo'}
${prefix} start -u https://example.com -r ./my-repo
${prefix} start -u https://example.com -r /path/to/repo -c config.yaml -w q1-audit
${prefix} logs q1-audit
${prefix} stop --clean
${
mode === 'local'
? `
State directory: ./workspaces/`
: `
State directory: ~/.shannon/`
}
Monitor scans at http://localhost:8233
${prefix} stop q1-audit
${prefix} reset
Run '${prefix} <command> --help' for help on a specific command.
Docs & source: https://github.com/KeygraphHQ/shannon
`);
}
@@ -98,150 +108,168 @@ interface ParsedStartArgs {
workspace?: string;
output?: string;
pipelineTesting: boolean;
debug: boolean;
keepContainer: boolean;
follow: boolean;
}
function parseStartArgs(argv: string[]): ParsedStartArgs {
let url = '';
let repo = '';
let config: string | undefined;
let workspace: string | undefined;
let output: string | undefined;
let pipelineTesting = false;
let debug = false;
const { flags, values } = parseArgs(argv, {
values: {
url: ['-u', '--url'],
repo: ['-r', '--repo'],
config: ['-c', '--config'],
output: ['-o', '--output'],
workspace: ['-w', '--workspace'],
},
booleans: {
pipelineTesting: ['--pipeline-testing'],
keepContainer: ['--keep-container'],
follow: ['-f', '--follow'],
},
});
for (let i = 0; i < argv.length; i++) {
const arg = argv[i];
const next = argv[i + 1];
switch (arg) {
case '-u':
case '--url':
if (next && !next.startsWith('-')) {
url = next;
i++;
}
break;
case '-r':
case '--repo':
if (next && !next.startsWith('-')) {
repo = next;
i++;
}
break;
case '-c':
case '--config':
if (next && !next.startsWith('-')) {
config = next;
i++;
}
break;
case '-w':
case '--workspace':
if (next && !next.startsWith('-')) {
workspace = next;
i++;
}
break;
case '-o':
case '--output':
if (next && !next.startsWith('-')) {
output = next;
i++;
}
break;
case '--pipeline-testing':
pipelineTesting = true;
break;
case '--debug':
debug = true;
break;
default:
console.error(`Unknown option: ${arg}`);
console.error(`Run "${getMode() === 'local' ? './shannon' : 'npx @keygraph/shannon'} help" for usage`);
process.exit(1);
}
const url = values.url ?? '';
const repo = values.repo ?? '';
if (!url || !repo) {
failUsage('--url and --repo are required', `Usage: ${commandPrefix()} start -u <url> -r <path>`);
}
if (!url || !repo) {
console.error('ERROR: --url and --repo are required');
console.error(`Usage: ${getMode() === 'local' ? './shannon' : 'npx @keygraph/shannon'} start -u <url> -r <path>`);
process.exit(1);
try {
new URL(url);
} catch {
failUsage(`invalid --url: ${url}`);
}
return {
url,
repo,
pipelineTesting,
debug,
...(config && { config }),
...(workspace && { workspace }),
...(output && { output }),
pipelineTesting: !!flags.pipelineTesting,
keepContainer: !!flags.keepContainer,
follow: !!flags.follow,
...(values.config && { config: values.config }),
...(values.workspace && { workspace: values.workspace }),
...(values.output && { output: values.output }),
};
}
// === Main Dispatch ===
blockSudo();
async function main(): Promise<void> {
// A reader that closes early (e.g. `shannon logs my-scan | head`) makes writes
// to stdout raise EPIPE. That's normal for a piped CLI, not a crash — exit quietly
// instead of letting Node dump an unhandled-error stack trace.
process.stdout.on('error', (err: NodeJS.ErrnoException) => {
if (err.code === 'EPIPE') process.exit(0);
throw err;
});
const args = process.argv.slice(2);
const command = args[0];
blockSudo();
switch (command) {
case 'start': {
const parsed = parseStartArgs(args.slice(1));
await start({ ...parsed, version: getVersion() });
break;
const args = process.argv.slice(2);
const command = args[0];
const rest = args.slice(1);
if (command === undefined || command === 'help' || command === '--help' || command === '-h') {
const topic = rest[0];
if (topic && isHelpableCommand(topic)) {
printCommandHelp(topic);
} else {
const bare = command === undefined;
if (bare && stdoutIsTerminal()) displaySplash(isLocal() ? undefined : getVersion());
showHelp(bare);
}
return;
}
case 'stop':
stop(args.includes('--clean'), args.includes('--yes') || args.includes('-y'));
break;
case 'logs': {
const workspaceId = args[1];
if (!workspaceId) {
console.error('ERROR: Workspace ID is required');
console.error(`Usage: ${getMode() === 'local' ? './shannon' : 'npx @keygraph/shannon'} logs <workspace>`);
process.exit(1);
}
logs(workspaceId);
break;
// Reachable from any invocation: `-h`/`--help` anywhere wins over the rest of the line.
if (isHelpableCommand(command) && (rest.includes('-h') || rest.includes('--help'))) {
printCommandHelp(command);
return;
}
case 'workspaces':
workspaces(getVersion());
break;
case 'status':
status();
break;
case 'setup':
if (getMode() === 'local') {
console.error('ERROR: setup is only available in npx mode. In local mode, use .env');
process.exit(1);
switch (command) {
case 'start': {
const parsed = parseStartArgs(rest);
await start({ ...parsed, version: getVersion() });
break;
}
setup();
break;
case 'build':
build(args.includes('--no-cache'), getVersion());
break;
case 'uninstall':
if (getMode() === 'local') {
console.error('ERROR: uninstall is only available in npx mode.');
process.exit(1);
case 'stop': {
const { flags, positionals } = parseArgs(rest, {
booleans: { all: ['--all'], yes: YES_FLAGS },
maxPositionals: 1,
});
await stop({ all: !!flags.all, yes: !!flags.yes, ...(positionals[0] && { workspace: positionals[0] }) });
break;
}
uninstall(args.includes('--yes') || args.includes('-y'));
break;
case 'version':
case '--version':
case '-v':
console.log(getVersionLine());
break;
case 'help':
case '--help':
case '-h':
case undefined:
showHelp();
break;
default:
console.error(`Unknown command: ${command}`);
showHelp();
process.exit(1);
case 'reset': {
// reset is all-or-nothing; a stray name likely means the user wanted `stop <name>`.
parseArgs(rest, {
positionalHint: 'reset takes no workspace argument. To stop one scan, use: stop <name>',
});
await reset();
break;
}
case 'logs': {
const { positionals } = parseArgs(rest, { maxPositionals: 1 });
const workspaceId = positionals[0];
if (!workspaceId) {
failUsage('Workspace ID is required', `Usage: ${commandPrefix()} logs <workspace>`);
}
logs(workspaceId);
break;
}
case 'status': {
const { flags, positionals } = parseArgs(rest, { booleans: { json: ['--json'] }, maxPositionals: 1 });
const workspaceId = positionals[0];
if (!workspaceId) {
failUsage('Workspace is required', `Usage: ${commandPrefix()} status <workspace> [--json]`);
}
await status(workspaceId, { json: !!flags.json });
break;
}
case 'scans': {
const { flags } = parseArgs(rest, { booleans: { json: ['--json'] } });
scans({ json: !!flags.json });
break;
}
case 'setup':
if (getMode() === 'local') {
fail('setup is only available in npx mode. In local mode, use .env');
}
parseArgs(rest, {});
await setup();
break;
case 'build': {
const { flags } = parseArgs(rest, { booleans: { noCache: ['--no-cache'] } });
build(!!flags.noCache, getVersion());
break;
}
case 'version':
case '--version':
case '-v': {
const { flags } = parseArgs(rest, { booleans: { json: ['--json'] } });
if (flags.json) {
console.log(JSON.stringify({ version: getVersion(), mode: getMode() }, null, 2));
} else {
console.log(getVersionLine());
}
break;
}
default: {
const prefix = commandPrefix();
const suggestion = closestMatch(command, availableCommands());
const hints = [
...(suggestion ? [`Did you mean '${suggestion}'?`] : []),
`Run '${prefix} help' to see available commands.`,
];
failUsage(`Unknown command: ${command}`, ...hints);
}
}
}
main().catch((err) => {
if (err instanceof ArgError) {
failUsage(err.message, `Run "${commandPrefix()} help" for usage`);
}
crash(err);
});
+167
View File
@@ -0,0 +1,167 @@
/**
* Decorates a tailed workflow.log for the terminal.
*
* The worker writes workflow.log as plain text and the CLI reads it back, so all of the
* chrome lives here on the read side. Nothing in this file changes what the log *says* —
* it adds the section treatments the plain file has no way to carry:
*
* the log header and RESUMED banner -> rule() (their ==== bars become the rule)
* everything between phases -> gutter() (one bar per phase, walking the ramp)
* the closing Scan COMPLETED block -> panel() (its ==== bars become the frame)
*
* When stdout is not a terminal the renderer is a pass-through and emits the file's bytes
* unchanged, so redirected logs, pipes, and CI keep grepping the same text they always did.
*/
import { field, gutter, palette, panel, rule } from './chrome.js';
/** The ==== bars that open and close a block; replaced by our own chrome. */
const BLOCK_BAR = /^={10,}\s*$/;
/** The ──── bar dividing a block's title from its body; replaced by the panel frame. */
const INNER_BAR = /^─{10,}\s*$/;
/** `[2026-08-26 17:04:11] ` — every streamed event line carries one. */
const TIMESTAMP = /^(\[\d{4}-\d{2}-\d{2} \d{2}:\d{2}:\d{2}\])( .*)$/;
/** A phase transition opens a new gutter section. */
const PHASE_START = /^\[[^\]]*\] \[PHASE\] Starting: /;
/** Titles of the two banner blocks, which take the static rule treatment. */
const BANNER_TITLE = /^(Shannon Pentest - Scan Log|RESUMED)$/;
/** Title of the final block, which takes the panel treatment. Mirrors logs.ts's completion regex. */
const COMPLETION_TITLE = /^Scan (COMPLETED|FAILED)$/;
/** Dim the timestamp so the message reads first; errors take the semantic red, not a ramp colour. */
function colorizeEvent(line: string): string {
const { DIM, RED, RESET } = palette();
const match = TIMESTAMP.exec(line);
if (!match) return line;
const rest = match[2] ?? '';
const body = rest.includes('[ERROR]') ? `${RED}${rest}${RESET}` : rest;
return `${DIM}${match[1]}${RESET}${body}`;
}
type Mode = 'stream' | 'banner' | 'summary';
export class LogRenderer {
/** Bytes past the last newline, held until the rest of the line arrives. */
private carry = '';
private mode: Mode = 'stream';
/** Suppresses a second consecutive blank line; starts true so the stream can't open on one. */
private lastBlank = true;
/** Current sunset stop for the gutter bar; advanced by each phase transition. */
private stop = 0;
private summaryTitle = '';
private summaryBody: string[] = [];
private readonly passthrough: boolean;
constructor() {
this.passthrough = !palette().color;
}
/** Decorate a chunk of newly appended log text. Incomplete trailing lines are held back. */
write(chunk: string): string {
if (this.passthrough) return chunk;
const text = this.carry + chunk;
const lines = text.split('\n');
// The final element is whatever followed the last newline — possibly a partial line.
this.carry = lines.pop() ?? '';
const out: string[] = [];
for (const line of lines) {
out.push(...this.renderLine(line));
}
return this.emit(out);
}
/** Join rendered lines, dropping blank runs left behind by the bars we removed. */
private emit(lines: string[]): string {
const kept: string[] = [];
for (const line of lines) {
const blank = line === '';
if (blank && this.lastBlank) continue;
this.lastBlank = blank;
kept.push(line);
}
return kept.length ? `${kept.join('\n')}\n` : '';
}
/** Flush a held partial line and close an unterminated summary block. */
end(): string {
if (this.passthrough) return '';
const out: string[] = [];
if (this.carry) {
out.push(...this.renderLine(this.carry));
this.carry = '';
}
if (this.mode === 'summary') {
out.push(...this.closeSummary());
}
return this.emit(out);
}
private renderLine(raw: string): string[] {
// Strip the \r from CRLF logs so it never lands in the middle of a decorated line.
const line = raw.endsWith('\r') ? raw.slice(0, -1) : raw;
// Only a ==== bar closes the summary; its ──── divider is chrome we replace, not a terminator.
if (BLOCK_BAR.test(line)) {
return this.mode === 'summary' ? this.closeSummary() : [];
}
if (INNER_BAR.test(line)) return [];
if (COMPLETION_TITLE.test(line)) {
this.mode = 'summary';
this.summaryTitle = line;
this.summaryBody = [];
return [''];
}
if (BANNER_TITLE.test(line)) {
this.mode = 'banner';
return ['', rule(line)];
}
if (this.mode === 'summary') {
this.summaryBody.push(line);
return [];
}
if (this.mode === 'banner') {
// The banner runs until the first streamed event.
if (!TIMESTAMP.test(line)) {
return [line.trim() ? ` ${field(line)}` : ''];
}
this.mode = 'stream';
}
// A blank line separates sections; the bar resumes on the next line of content.
if (!line.trim()) return [''];
if (PHASE_START.test(line)) {
// Two stops per phase, not one: the 256-colour tier collapses the seven stops into
// four xterm colours, and a single step would give consecutive phases the same bar.
// Seven is odd, so a stride of two still visits every stop before repeating.
this.stop += 2;
}
return [gutter(colorizeEvent(line), this.stop)];
}
private closeSummary(): string[] {
const title = this.summaryTitle;
const body = [...this.summaryBody];
while (body.length && !body[body.length - 1]?.trim()) body.pop();
while (body.length && !body[0]?.trim()) body.shift();
this.mode = 'stream';
this.summaryTitle = '';
this.summaryBody = [];
const rows = body.map((line) => (line.trim() ? field(line) : ''));
return [...panel(title, rows), ''];
}
}
+5
View File
@@ -24,6 +24,11 @@ export function isLocal(): boolean {
return getMode() === 'local';
}
/** The invocation prefix for the current mode, so help and hints point at a runnable command. */
export function commandPrefix(): string {
return getMode() === 'local' ? './shannon' : 'npx @keygraph/shannon';
}
export function isDevMode(): boolean {
return process.env.SHANNON_DEV === '1';
}
+28 -34
View File
@@ -1,13 +1,27 @@
/**
* Path resolution for --repo and --config arguments.
*
* Local mode supports bare repo names (e.g. "my-repo" → ./repos/my-repo).
* Both modes resolve relative paths against CWD.
* Both --repo and --config are filesystem paths, absolute or relative to CWD.
*/
import fs from 'node:fs';
import os from 'node:os';
import path from 'node:path';
import { isLocal } from './mode.js';
import { fail } from './errors.js';
/**
* Expand a leading `~` or `~/` to the home directory. The shell skips this in the
* `--flag=~/x` form (the tilde is not at the word start), so it must be done here.
*/
export function expandHome(inputPath: string): string {
if (inputPath === '~') {
return os.homedir();
}
if (inputPath.startsWith('~/')) {
return path.join(os.homedir(), inputPath.slice(2));
}
return inputPath;
}
export interface MountPair {
hostPath: string;
@@ -23,10 +37,10 @@ export interface MountPair {
export const INTERNAL_DIR = '.shannon';
/**
* Filename of the human-facing final report surfaced at the run directory root.
* Must match FINAL_REPORT_FILENAME in the worker package.
* Filename of the human-facing PDF report surfaced at the run directory root.
* Must match FINAL_REPORT_PDF_FILENAME in the worker package.
*/
export const FINAL_REPORT_FILENAME = 'Security-Assessment-Report.md';
export const FINAL_REPORT_PDF_FILENAME = 'Security-Assessment-Report.pdf';
/**
* Resolve a run-directory file (e.g. session.json, workflow.log), preferring the
@@ -47,36 +61,18 @@ export function resolveRunFile(runDir: string, filename: string): string {
}
/**
* Resolve --repo to absolute path and container mount.
* Dev mode: bare names (no / or . prefix) check ./repos/<name> first.
* Resolve --repo to an absolute path and container mount. The argument is a
* filesystem path, absolute or relative to CWD.
*/
export function resolveRepo(repoArg: string): MountPair {
let hostPath: string;
if (isLocal() && !repoArg.startsWith('/') && !repoArg.startsWith('.')) {
// Bare name — check ./repos/<name> for backward compatibility
const barePath = path.resolve('repos', repoArg);
if (fs.existsSync(barePath)) {
hostPath = barePath;
} else {
console.error(`ERROR: Repository not found at ./repos/${repoArg}`);
console.error('');
console.error('Place your target repository under the ./repos/ directory,');
console.error('or pass an absolute/relative path: -r /path/to/repo');
process.exit(1);
}
} else {
hostPath = path.resolve(repoArg);
}
const hostPath = path.resolve(expandHome(repoArg));
if (!fs.existsSync(hostPath)) {
console.error(`ERROR: Repository not found: ${hostPath}`);
process.exit(1);
fail(`Repository not found: ${hostPath}`);
}
if (!fs.statSync(hostPath).isDirectory()) {
console.error(`ERROR: Not a directory: ${hostPath}`);
process.exit(1);
fail(`Not a directory: ${hostPath}`);
}
const basename = path.basename(hostPath);
@@ -90,16 +86,14 @@ export function resolveRepo(repoArg: string): MountPair {
* Resolve --config to absolute path and container mount.
*/
export function resolveConfig(configArg: string): MountPair {
const hostPath = path.resolve(configArg);
const hostPath = path.resolve(expandHome(configArg));
if (!fs.existsSync(hostPath)) {
console.error(`ERROR: Config file not found: ${hostPath}`);
process.exit(1);
fail(`Config file not found: ${hostPath}`);
}
if (!fs.statSync(hostPath).isFile()) {
console.error(`ERROR: Not a file: ${hostPath}`);
process.exit(1);
fail(`Not a file: ${hostPath}`);
}
const basename = path.basename(hostPath);
+155
View File
@@ -0,0 +1,155 @@
/**
* Pure derivation of a scan's per-agent and per-phase state from its Temporal snapshot.
*
* This is the single source of truth for "what state is each agent in" — both the
* human progress tree (render.ts) and the machine-readable snapshot (status-json.ts)
* consume it, so the two views can never disagree about whether an agent is running,
* skipped, or still pending. No glyphs, no color, no formatting live here.
*/
import type { RunningAgent } from '../temporal-client.js';
import { agentClass, PIPELINE, type PipelineState } from './pipeline.js';
import type { RenderInput } from './render.js';
export type RunState = 'pending' | 'running' | 'completed' | 'failed' | 'skipped';
/** One agent's resolved state plus the raw metrics/timing a consumer needs to present it. Null metrics
* mean the value doesn't apply to the current state (e.g. duration only for completed agents). */
export interface DerivedAgent {
readonly name: string;
readonly label: string;
readonly state: RunState;
readonly durationMs: number | null;
readonly runningElapsedMs: number | null;
readonly attempt: number | null;
readonly error?: string;
}
export interface DerivedPhase {
readonly key: string;
readonly label: string;
readonly parallel: boolean;
readonly state: RunState;
readonly agents: readonly DerivedAgent[];
}
/** Terminal = anything other than an open, running execution. */
export function isTerminal(status: string): boolean {
return status !== 'RUNNING' && status !== 'UNSPECIFIED';
}
function isFailedAgent(name: string, state: PipelineState | null): boolean {
return !!state && (state.failedAgent === name || state.failedPipelines.some((f) => f.vulnType === agentClass(name)));
}
/** An agent has entered play once it is running, has metrics, or has failed. */
function isAgentActive(name: string, state: PipelineState | null, running: Set<string>): boolean {
return running.has(name) || !!state?.agentMetrics[name] || isFailedAgent(name, state);
}
/**
* Resolve one agent's state. "Ran" is signalled by a metrics entry, not by
* completedAgents — the workflow lists conditionally-skipped agents (e.g. exploit
* agents when there is nothing to exploit) as completed but records no metrics for
* them. `resolved` is true once we've moved past this agent's phase (the scan is
* terminal, or a later phase is already active), at which point a metric-less,
* non-running agent is skipped rather than still pending.
*/
function agentState(name: string, state: PipelineState | null, running: Set<string>, resolved: boolean): RunState {
if (running.has(name)) return 'running';
if (isFailedAgent(name, state)) return 'failed';
if (state?.agentMetrics[name]) return 'completed';
return resolved ? 'skipped' : 'pending';
}
function agentError(name: string, state: PipelineState | null, byAgent: Map<string, RunningAgent>): string | undefined {
const failed = state?.failedPipelines.find((f) => f.vulnType === agentClass(name));
return (
failed?.error ??
byAgent.get(name)?.lastFailure ??
(state?.failedAgent === name ? (state.error ?? undefined) : undefined)
);
}
/** Scan wall-clock elapsed ms: recorded duration for a closed scan, live elapsed for a running one. */
export function scanElapsedMs(input: RenderInput, now: number): number | undefined {
if (isTerminal(input.temporalStatus)) {
if (input.state?.summary) return input.state.summary.totalDurationMs;
if (input.endedAt !== undefined && input.startedAt !== undefined) return input.endedAt - input.startedAt;
return undefined;
}
return input.startedAt !== undefined ? now - input.startedAt : undefined;
}
/** Collapse a phase's agent states into a single state for the phase line. */
export function phaseGlyphState(states: readonly RunState[]): RunState {
if (states.some((s) => s === 'running')) return 'running';
if (states.some((s) => s === 'failed')) return 'failed';
if (states.every((s) => s === 'skipped')) return 'skipped';
if (states.every((s) => s === 'completed' || s === 'skipped')) return 'completed';
if (states.some((s) => s === 'completed')) return 'running';
return 'pending';
}
/**
* Compute each agent's RunState. This is the drift-prone part shared by every view.
*
* The pipeline is sequential across phases: the last phase with any active agent is the
* frontier. Earlier phases with nothing active were skipped (e.g. exploitation when no
* class had anything to exploit), not still pending.
*/
export function deriveAgentStates(input: RenderInput): Map<string, RunState> {
const runningSet = new Set(input.running.map((r) => r.agent));
const terminal = isTerminal(input.temporalStatus);
let frontier = -1;
PIPELINE.forEach((phase, idx) => {
if (phase.agents.some((a) => isAgentActive(a.name, input.state, runningSet))) frontier = idx;
});
const states = new Map<string, RunState>();
for (const [phaseIdx, phase] of PIPELINE.entries()) {
const resolved = terminal || phaseIdx < frontier;
for (const agent of phase.agents) {
states.set(agent.name, agentState(agent.name, input.state, runningSet, resolved));
}
}
return states;
}
/**
* Full structured view of the pipeline: every agent's state plus the raw
* metrics/timing needed to present it, and each phase's collapsed state.
*/
export function derivePipeline(input: RenderInput, now: number): DerivedPhase[] {
const states = deriveAgentStates(input);
const byAgent = new Map(input.running.map((r) => [r.agent, r]));
return PIPELINE.map((phase) => {
const agents = phase.agents.map((a): DerivedAgent => {
const state = states.get(a.name) ?? 'pending';
const metrics = input.state?.agentMetrics[a.name];
const runner = byAgent.get(a.name);
const error = agentError(a.name, input.state, byAgent);
return {
name: a.name,
label: a.label,
state,
durationMs: state === 'completed' && metrics ? metrics.durationMs : null,
runningElapsedMs: state === 'running' && runner?.startedAt !== undefined ? now - runner.startedAt : null,
attempt: state === 'running' && runner ? runner.attempt : null,
...(error !== undefined && { error }),
};
});
return {
key: phase.key,
label: phase.label,
parallel: phase.parallel,
state: phaseGlyphState(agents.map((ag) => ag.state)),
agents,
};
});
}
export { agentError };
+31
View File
@@ -0,0 +1,31 @@
/**
* Rendering for the worker's '|'-delimited failure string.
*
* `formatWorkflowError` in the worker joins error segments — phase context, error type,
* message, and remediation hint — with '|' as a delimiter. These helpers turn that raw
* string into readable output for the CLI's own surfaces.
*/
/**
* Split the failure string into trimmed, non-empty lines. Segments are delimited by '|', and a
* segment's own embedded newlines (e.g. a multi-line validation message) become their own lines so
* each aligns with the rest of the block.
*/
export function parseFailureSegments(message: string): string[] {
return message
.split(/[|\n]/)
.map((segment) => segment.trim())
.filter((segment) => segment.length > 0);
}
/** Multi-line block: one segment per indented line (the caller prints the header). */
export function indentFailureSegments(message: string, indent = ' '): string {
return parseFailureSegments(message)
.map((segment) => `${indent}${segment}`)
.join('\n');
}
/** Single-line summary for compact contexts like the status footer. */
export function inlineFailureReason(message: string): string {
return parseFailureSegments(message).join(' — ');
}
+123
View File
@@ -0,0 +1,123 @@
/**
* Static description of the Shannon scan pipeline, plus the worker types the CLI
* reads back from Temporal.
*
* The CLI cannot import from the worker package, so this mirrors it. Keep in sync with:
* - apps/worker/src/types/agents.ts (agent names / ordering)
* - apps/worker/src/session-manager.ts (phase membership)
* - apps/worker/src/temporal/activities.ts (the run*Agent activity names → `activityType`)
* - apps/worker/src/temporal/shared.ts (PipelineState / PipelineSummary)
* - apps/worker/src/types/metrics.ts (AgentMetrics)
*/
export interface AgentSpec {
/** Canonical agent name as it appears in PipelineState.completedAgents / agentMetrics. */
readonly name: string;
/** Short label for the progress tree. */
readonly label: string;
/** Temporal activity type name — how a running agent shows up in pendingActivities. */
readonly activityType: string;
}
export interface PhaseSpec {
readonly key: string;
readonly label: string;
readonly parallel: boolean;
readonly agents: readonly AgentSpec[];
}
/** The pipeline phases in execution order, each with its agents. */
export const PIPELINE: readonly PhaseSpec[] = [
{
// Preflight login check. Only authenticated scans record metrics here; a non-auth scan
// records none, so it renders as skipped — like Exploitation when nothing is exploitable.
key: 'auth-validation',
label: 'Authentication',
parallel: false,
agents: [{ name: 'validate-authentication', label: 'auth', activityType: 'runAuthenticationValidation' }],
},
{
key: 'pre-recon',
label: 'Pre-Recon',
parallel: false,
agents: [{ name: 'pre-recon', label: 'pre-recon', activityType: 'runPreReconAgent' }],
},
{
key: 'recon',
label: 'Recon',
parallel: false,
agents: [{ name: 'recon', label: 'recon', activityType: 'runReconAgent' }],
},
{
key: 'vulnerability-analysis',
label: 'Vulnerability Analysis',
parallel: true,
agents: [
{ name: 'injection-vuln', label: 'injection', activityType: 'runInjectionVulnAgent' },
{ name: 'xss-vuln', label: 'xss', activityType: 'runXssVulnAgent' },
{ name: 'auth-vuln', label: 'auth', activityType: 'runAuthVulnAgent' },
{ name: 'ssrf-vuln', label: 'ssrf', activityType: 'runSsrfVulnAgent' },
{ name: 'authz-vuln', label: 'authz', activityType: 'runAuthzVulnAgent' },
],
},
{
key: 'exploitation',
label: 'Exploitation',
parallel: true,
agents: [
{ name: 'injection-exploit', label: 'injection', activityType: 'runInjectionExploitAgent' },
{ name: 'xss-exploit', label: 'xss', activityType: 'runXssExploitAgent' },
{ name: 'auth-exploit', label: 'auth', activityType: 'runAuthExploitAgent' },
{ name: 'ssrf-exploit', label: 'ssrf', activityType: 'runSsrfExploitAgent' },
{ name: 'authz-exploit', label: 'authz', activityType: 'runAuthzExploitAgent' },
],
},
{
key: 'reporting',
label: 'Reporting',
parallel: false,
agents: [{ name: 'report', label: 'report', activityType: 'runReportAgent' }],
},
];
/** Temporal activity type name → canonical agent name, for mapping pendingActivities. */
export const ACTIVITY_TO_AGENT: Readonly<Record<string, string>> = Object.fromEntries(
PIPELINE.flatMap((phase) => phase.agents.map((agent) => [agent.activityType, agent.name])),
);
/** The vuln/exploit class of an agent (e.g. "authz-vuln" → "authz"), for failedPipelines matching. */
export function agentClass(name: string): string {
return name.replace(/-(vuln|exploit)$/, '');
}
// === Worker types read back from Temporal (mirror of shared.ts / metrics.ts) ===
export interface AgentMetrics {
readonly durationMs: number;
readonly costUsd: number | null;
readonly numTurns: number | null;
readonly model?: string;
readonly skipped?: boolean;
}
export interface PipelineSummary {
readonly totalCostUsd: number;
readonly totalDurationMs: number; // Wall-clock (end - start)
readonly totalTurns: number;
readonly agentCount: number;
}
export type PipelineStatus = 'running' | 'completed' | 'failed' | 'cancelled' | 'partial';
export interface PipelineState {
readonly status: PipelineStatus;
readonly currentPhase: string | null;
readonly currentAgent: string | null;
readonly completedAgents: string[];
readonly failedPipelines: { vulnType: string; error: string }[];
readonly failedAgent: string | null;
readonly error: string | null;
readonly startTime: number;
readonly agentMetrics: Record<string, AgentMetrics>;
readonly summary: PipelineSummary | null;
}
+250
View File
@@ -0,0 +1,250 @@
/**
* Renders a scan's Temporal state into the terminal progress tree.
*
* The same PipelineState drives both the live view (from the getProgress query) and
* the final view (from the workflow result); the running-agents overlay (from
* pendingActivities) supplies the in-flight set and retry counts the state lacks.
* Colors and Unicode glyphs are gated by the caller so the frame degrades off a TTY.
*/
import { BOLD, DIM, GOLD, paint, RED, YELLOW } from '../colors.js';
import { commandPrefix } from '../mode.js';
import type { RunningAgent } from '../temporal-client.js';
import { agentError, deriveAgentStates, isTerminal, phaseGlyphState, type RunState, scanElapsedMs } from './derive.js';
import { inlineFailureReason } from './failure.js';
import { PIPELINE, type PipelineState } from './pipeline.js';
export interface RenderInput {
readonly workspace: string;
/** Temporal workflow id backing this scan (differs from workspace on a resume); used for the dashboard link. */
readonly workflowId?: string;
/** Temporal WorkflowExecutionStatusName: RUNNING | COMPLETED | FAILED | CANCELLED | TERMINATED | … */
readonly temporalStatus: string;
/** Progress (live) or result (terminal). Null when unavailable, e.g. a hard failure with no result. */
readonly state: PipelineState | null;
readonly running: readonly RunningAgent[];
readonly startedAt?: number;
readonly endedAt?: number;
/** Failure text when a failed scan has no readable state. */
readonly failureMessage?: string;
}
export interface RenderOptions {
readonly now: number;
readonly color: boolean;
readonly unicode: boolean;
/** True for the live view (adds a watch footer); false for the final/one-shot frame. */
readonly live: boolean;
/** Animation tick — advances the running-agent spinner. Ignored for static frames. */
readonly frame: number;
}
const COLORS = {
red: RED,
gold: GOLD,
yellow: YELLOW,
dim: DIM,
bold: BOLD,
} as const;
// === Formatting ===
function formatDuration(ms: number): string {
const seconds = Math.max(0, Math.floor(ms / 1000));
const hours = Math.floor(seconds / 3600);
const minutes = Math.floor((seconds % 3600) / 60);
const secs = seconds % 60;
if (hours > 0) return `${hours}h ${minutes}m`;
if (minutes > 0) return `${minutes}m ${secs}s`;
return `${secs}s`;
}
function truncate(text: string, max: number): string {
const flat = text.replace(/\s+/g, ' ').trim();
return flat.length <= max ? flat : `${flat.slice(0, max - 1)}…`;
}
/** Temporal Web UI, published by compose on 8233; deep-links to the workflow when its id is known. */
function temporalDashboardUrl(workflowId: string | undefined): string {
const base = 'http://localhost:8233';
return workflowId ? `${base}/namespaces/default/workflows/${workflowId}` : base;
}
// === Glyphs & status ===
const GLYPH_UNICODE: Record<RunState, string> = {
pending: '○',
running: '⟳',
completed: '●',
failed: '✗',
skipped: '·',
};
const GLYPH_ASCII: Record<RunState, string> = {
pending: '.',
running: '>',
completed: '+',
failed: 'x',
skipped: '-',
};
const STATE_COLOR: Record<RunState, string> = {
pending: COLORS.dim,
running: COLORS.gold,
completed: COLORS.gold,
failed: COLORS.red,
skipped: COLORS.dim,
};
/** Braille spinner frames for running agents — the clack loader style. */
const SPINNER_FRAMES = ['⠋', '⠙', '⠹', '⠸', '⠼', '⠴', '⠦', '⠧', '⠇', '⠏'] as const;
function glyph(state: RunState, opts: RenderOptions): string {
if (state === 'running' && opts.unicode) {
const spin = SPINNER_FRAMES[opts.frame % SPINNER_FRAMES.length] ?? SPINNER_FRAMES[0];
return paint(spin, STATE_COLOR.running, opts.color);
}
const symbol = opts.unicode ? GLYPH_UNICODE[state] : GLYPH_ASCII[state];
return paint(symbol, STATE_COLOR[state], opts.color);
}
/** Badge text + color for the scan as a whole, preferring the workflow's own status when known. */
function statusBadge(input: RenderInput, opts: RenderOptions): string {
const workflowStatus = input.state?.status;
if (!isTerminal(input.temporalStatus)) return paint('running', COLORS.gold, opts.color);
if (workflowStatus === 'partial') return paint('partial', COLORS.yellow, opts.color);
if (input.temporalStatus === 'COMPLETED') return paint('completed', COLORS.gold, opts.color);
if (input.temporalStatus === 'TERMINATED') return paint('stopped', COLORS.yellow, opts.color);
if (input.temporalStatus === 'CANCELLED' || input.temporalStatus === 'CANCELED') {
return paint('cancelled', COLORS.yellow, opts.color);
}
if (input.temporalStatus === 'TIMED_OUT') return paint('timed out', COLORS.red, opts.color);
return paint('FAILED', COLORS.red, opts.color);
}
// === Line builders ===
function agentMeta(
state: RunState,
metrics: { durationMs: number } | undefined,
runner: RunningAgent | undefined,
error: string | undefined,
opts: RenderOptions,
): string {
if (state === 'completed') {
const duration = metrics?.durationMs != null ? formatDuration(metrics.durationMs) : 'done';
return paint(duration, COLORS.dim, opts.color);
}
if (state === 'running') {
const parts = ['running'];
if (runner?.startedAt !== undefined) parts.push(formatDuration(opts.now - runner.startedAt));
if (runner && runner.attempt > 1) parts.push(`retry ${runner.attempt}`);
return paint(parts.join(' · '), COLORS.gold, opts.color);
}
if (state === 'failed') {
const detail = error ? ` · ${truncate(error, 46)}` : '';
return paint(`failed${detail}`, COLORS.red, opts.color);
}
if (state === 'skipped') return paint('skipped', COLORS.dim, opts.color);
return paint('queued', COLORS.dim, opts.color);
}
function phaseMeta(states: readonly RunState[], inPlay: number, parallel: boolean, opts: RenderOptions): string {
if (states.every((s) => s === 'pending')) return paint('pending', COLORS.dim, opts.color);
if (states.every((s) => s === 'skipped')) return paint('skipped', COLORS.dim, opts.color);
if (states.some((s) => s === 'failed') && !states.some((s) => s === 'running')) {
return paint('failed', COLORS.red, opts.color);
}
if (!parallel) return '';
const done = states.filter((s) => s === 'completed').length;
const allDone = states.every((s) => s === 'completed' || s === 'skipped');
return paint(`${done}/${inPlay} done`, allDone ? COLORS.gold : COLORS.dim, opts.color);
}
/** Render the full progress frame as one string (no trailing newline). */
export function renderScan(input: RenderInput, opts: RenderOptions): string {
const byAgent = new Map(input.running.map((r) => [r.agent, r]));
const stateMap = deriveAgentStates(input);
const lines: string[] = ['', ...headerLines(input, opts), ''];
const metaFor = (name: string, state: RunState): string =>
agentMeta(state, input.state?.agentMetrics[name], byAgent.get(name), agentError(name, input.state, byAgent), opts);
// Only agents that have actually entered play are shown; pending/skipped ones stay hidden.
const inPlay = (s: RunState): boolean => s === 'running' || s === 'completed' || s === 'failed';
for (const phase of PIPELINE) {
const states = phase.agents.map((a) => stateMap.get(a.name) ?? 'pending');
const playing = states.filter(inPlay).length;
const phaseRunState: RunState = phaseGlyphState(states);
// A single-agent phase carries that agent's own duration/cost on the phase line once it
// starts; a parallel phase gets a "k/N done" summary over the agents in play.
const first = phase.agents[0];
const firstState = states[0];
const phaseMetaStr =
!phase.parallel && first && firstState && inPlay(firstState)
? metaFor(first.name, firstState)
: phaseMeta(states, playing, phase.parallel, opts);
lines.push(` ${glyph(phaseRunState, opts)} ${phase.label.padEnd(26)}${phaseMetaStr}`);
if (!phase.parallel) continue;
for (let i = 0; i < phase.agents.length; i++) {
const agent = phase.agents[i];
const state = states[i];
if (!agent || !state || !inPlay(state)) continue;
lines.push(` ${glyph(state, opts)} ${agent.label.padEnd(18)}${metaFor(agent.name, state)}`);
}
}
lines.push(...footerLines(input, opts));
return lines.join('\n');
}
function headerLines(input: RenderInput, opts: RenderOptions): string[] {
const elapsedMs = scanElapsedMs(input, opts.now);
const meta = [statusBadge(input, opts), elapsedMs !== undefined ? formatDuration(elapsedMs) : '—'].join(' · ');
return [` ${paint('Scan:', COLORS.bold, opts.color)} ${input.workspace.padEnd(22)} ${meta}`];
}
/** Aligned label column for the footer's Logs / Temporal rows. */
const FOOTER_LABEL_WIDTH = 12;
/** A thin rule that sets the footer apart from the phase list above it. */
function footerDivider(opts: RenderOptions): string {
return paint(` ${(opts.unicode ? '─' : '-').repeat(60)}`, COLORS.dim, opts.color);
}
/** One footer row: an accent-colored label in a fixed column, then its value in the default color. */
function footerRow(label: string, value: string, opts: RenderOptions): string {
return ` ${paint(label.padEnd(FOOTER_LABEL_WIDTH), COLORS.gold, opts.color)}${value}`;
}
function footerLines(input: RenderInput, opts: RenderOptions): string[] {
const prefix = commandPrefix();
if (isTerminal(input.temporalStatus) && input.state?.summary) {
const wall = formatDuration(input.state.summary.totalDurationMs);
return ['', ` Time Taken ${wall}`];
}
const logsValue = `${prefix} logs ${input.workspace}`;
const temporalValue = temporalDashboardUrl(input.workflowId);
if (isTerminal(input.temporalStatus)) {
const rawReason = input.failureMessage ?? input.state?.error;
const reason = rawReason ? inlineFailureReason(rawReason) : 'no result recorded';
return [
footerDivider(opts),
paint(
` ${input.temporalStatus === 'TERMINATED' ? 'Stopped' : 'Ended'} — ${truncate(reason, 240)}`,
COLORS.dim,
opts.color,
),
footerRow('Logs', logsValue, opts),
footerRow('Temporal', temporalValue, opts),
];
}
const lines = [footerDivider(opts), footerRow('Logs', logsValue, opts), footerRow('Temporal', temporalValue, opts)];
if (opts.live) lines.push('', paint(' Ctrl-C stops watching — the scan keeps running.', COLORS.dim, opts.color));
return lines;
}
+68
View File
@@ -0,0 +1,68 @@
/**
* Machine-readable snapshot of one scan, for `shannon status --json`.
*
* A point-in-time view built from the same derivation the human progress tree uses
* (derive.ts), so the JSON and the rendered tree can never disagree about an agent's
* state. One invocation is one snapshot — callers that want to track progress poll it.
*/
import type { DerivedPhase } from './derive.js';
import { derivePipeline, isTerminal, scanElapsedMs } from './derive.js';
import type { RenderInput } from './render.js';
/** Coarse scan status token, mirroring the human status badge in machine-friendly form. */
export type ScanStatus = 'running' | 'completed' | 'partial' | 'failed' | 'stopped' | 'cancelled' | 'timed_out';
export interface StatusJson {
readonly workspace: string;
/** Temporal workflow id backing this scan (differs from workspace on a resume). */
readonly workflowId?: string;
/** Coarse outcome: `running` until the scan closes, then its terminal status. */
readonly status: ScanStatus;
/** Raw Temporal WorkflowExecutionStatusName, for callers that need the source status. */
readonly temporalStatus: string;
/** Wall-clock elapsed ms (live for a running scan, final for a closed one), or null when unknown. */
readonly elapsedMs: number | null;
readonly startedAt?: string;
readonly endedAt?: string;
/** Failure text when a failed scan left no readable state. */
readonly failureMessage?: string;
readonly phases: readonly DerivedPhase[];
}
/** Map the raw Temporal status (and workflow status) onto the coarse machine token. */
function deriveStatus(input: RenderInput): ScanStatus {
if (!isTerminal(input.temporalStatus)) return 'running';
if (input.state?.status === 'partial') return 'partial';
switch (input.temporalStatus) {
case 'COMPLETED':
return 'completed';
case 'TERMINATED':
return 'stopped';
case 'CANCELLED':
case 'CANCELED':
return 'cancelled';
case 'TIMED_OUT':
return 'timed_out';
default:
return 'failed';
}
}
/** Build the JSON snapshot for a scan at instant `now`. */
export function toStatusJson(input: RenderInput, now: number): StatusJson {
const elapsedMs = scanElapsedMs(input, now);
return {
workspace: input.workspace,
...(input.workflowId !== undefined && { workflowId: input.workflowId }),
status: deriveStatus(input),
temporalStatus: input.temporalStatus,
elapsedMs: elapsedMs ?? null,
...(input.startedAt !== undefined && { startedAt: new Date(input.startedAt).toISOString() }),
...(input.endedAt !== undefined && { endedAt: new Date(input.endedAt).toISOString() }),
...(input.failureMessage !== undefined && { failureMessage: input.failureMessage }),
phases: derivePipeline(input, now),
};
}
+26
View File
@@ -0,0 +1,26 @@
/**
* Workspace → Temporal workflow-id resolution.
*
* A workspace name is not always its workflow id: a fresh scan's id equals the
* workspace name, but each resume spawns a new workflow (`<workspace>_resume_<ts>`).
* The workspace's session.json records the authoritative id — the latest resume
* attempt, or the original — so commands that query Temporal (status, stop) resolve
* through here instead of assuming the name is the id.
*/
import fs from 'node:fs';
import path from 'node:path';
import { getWorkspacesDir } from './home.js';
import { resolveRunFile } from './paths.js';
/** Latest workflow id recorded for a workspace: last resume attempt, else the original. */
export function resolveWorkflowId(workspace: string): string | undefined {
const sessionPath = resolveRunFile(path.join(getWorkspacesDir(), workspace), 'session.json');
try {
const session = JSON.parse(fs.readFileSync(sessionPath, 'utf-8'));
const resumeAttempts: { workflowId?: string }[] = session.session?.resumeAttempts ?? [];
return resumeAttempts.at(-1)?.workflowId ?? session.session?.originalWorkflowId ?? undefined;
} catch {
return undefined;
}
}
+53 -39
View File
@@ -3,52 +3,66 @@
* Color escapes are gated on terminal support; the Unicode art is always kept.
*/
import { supportsColor } from './tty.js';
import { palette } from './chrome.js';
/** SHANNON wordmark. Block glyphs take the row fill; box-drawing strokes take the deeper edge shade. */
const SHANNON = [
'███████╗██╗ ██╗ █████╗ ███╗ ██╗███╗ ██╗ ██████╗ ███╗ ██╗',
'██╔════╝██║ ██║██╔══██╗████╗ ██║████╗ ██║██╔═══██╗████╗ ██║',
'███████╗███████║███████║██╔██╗ ██║██╔██╗ ██║██║ ██║██╔██╗ ██║',
'╚════██║██╔══██║██╔══██║██║╚██╗██║██║╚██╗██║██║ ██║██║╚██╗██║',
'███████║██║ ██║██║ ██║██║ ╚████║██║ ╚████║╚██████╔╝██║ ╚████║',
'╚══════╝╚═╝ ╚═╝╚═╝ ╚═╝╚═╝ ╚═══╝╚═╝ ╚═══╝ ╚═════╝ ╚═╝ ╚═══╝',
];
export function displaySplash(version?: string): void {
const color = supportsColor();
const GOLD = color ? '\x1b[38;2;244;197;66m' : '';
const CYAN = color ? '\x1b[36;1m' : '';
const WHITE = color ? '\x1b[1;37m' : '';
const GRAY = color ? '\x1b[0;37m' : '';
const YELLOW = color ? '\x1b[1;33m' : '';
const RESET = color ? '\x1b[0m' : '';
const { color, RESET, WHITE, GRAY, DIM, ramp } = palette();
const B = `${CYAN}\u2551${RESET}`;
const S67 = ' '.repeat(67);
const HR = '\u2550'.repeat(67);
/** Color one wordmark row, emitting an escape only where the run changes. Spaces stay unpainted. */
const paint = (row: string, fill: string, edge: string): string => {
if (!color) return row;
let out = '';
let open = '';
for (const ch of row) {
const want = ch === ' ' ? '' : ch === '█' ? fill : edge;
if (want !== open) {
if (open) out += RESET;
out += want;
open = want;
}
out += ch;
}
return open ? out + RESET : out;
};
const lines = [
'',
` ${CYAN}\u2554${HR}\u2557${RESET}`,
` ${B}${S67}${B}`,
` ${B} ${GOLD}\u2588\u2588\u2588\u2588\u2588\u2588\u2588\u2557\u2588\u2588\u2557 \u2588\u2588\u2557 \u2588\u2588\u2588\u2588\u2588\u2557 \u2588\u2588\u2588\u2557 \u2588\u2588\u2557\u2588\u2588\u2588\u2557 \u2588\u2588\u2557 \u2588\u2588\u2588\u2588\u2588\u2588\u2557 \u2588\u2588\u2588\u2557 \u2588\u2588\u2557${RESET} ${B}`,
` ${B} ${GOLD}\u2588\u2588\u2554\u2550\u2550\u2550\u2550\u255D\u2588\u2588\u2551 \u2588\u2588\u2551\u2588\u2588\u2554\u2550\u2550\u2588\u2588\u2557\u2588\u2588\u2588\u2588\u2557 \u2588\u2588\u2551\u2588\u2588\u2588\u2588\u2557 \u2588\u2588\u2551\u2588\u2588\u2554\u2550\u2550\u2550\u2588\u2588\u2557\u2588\u2588\u2588\u2588\u2557 \u2588\u2588\u2551${RESET} ${B}`,
` ${B} ${GOLD}\u2588\u2588\u2588\u2588\u2588\u2588\u2588\u2557\u2588\u2588\u2588\u2588\u2588\u2588\u2588\u2551\u2588\u2588\u2588\u2588\u2588\u2588\u2588\u2551\u2588\u2588\u2554\u2588\u2588\u2557 \u2588\u2588\u2551\u2588\u2588\u2554\u2588\u2588\u2557 \u2588\u2588\u2551\u2588\u2588\u2551 \u2588\u2588\u2551\u2588\u2588\u2554\u2588\u2588\u2557 \u2588\u2588\u2551${RESET} ${B}`,
` ${B} ${GOLD}\u255A\u2550\u2550\u2550\u2550\u2588\u2588\u2551\u2588\u2588\u2554\u2550\u2550\u2588\u2588\u2551\u2588\u2588\u2554\u2550\u2550\u2588\u2588\u2551\u2588\u2588\u2551\u255A\u2588\u2588\u2557\u2588\u2588\u2551\u2588\u2588\u2551\u255A\u2588\u2588\u2557\u2588\u2588\u2551\u2588\u2588\u2551 \u2588\u2588\u2551\u2588\u2588\u2551\u255A\u2588\u2588\u2557\u2588\u2588\u2551${RESET} ${B}`,
` ${B} ${GOLD}\u2588\u2588\u2588\u2588\u2588\u2588\u2588\u2551\u2588\u2588\u2551 \u2588\u2588\u2551\u2588\u2588\u2551 \u2588\u2588\u2551\u2588\u2588\u2551 \u255A\u2588\u2588\u2588\u2588\u2551\u2588\u2588\u2551 \u255A\u2588\u2588\u2588\u2588\u2551\u255A\u2588\u2588\u2588\u2588\u2588\u2588\u2554\u255D\u2588\u2588\u2551 \u255A\u2588\u2588\u2588\u2588\u2551${RESET} ${B}`,
` ${B} ${GOLD}\u255A\u2550\u2550\u2550\u2550\u2550\u2550\u255D\u255A\u2550\u255D \u255A\u2550\u255D\u255A\u2550\u255D \u255A\u2550\u255D\u255A\u2550\u255D \u255A\u2550\u2550\u2550\u255D\u255A\u2550\u255D \u255A\u2550\u2550\u2550\u255D \u255A\u2550\u2550\u2550\u2550\u2550\u255D \u255A\u2550\u255D \u255A\u2550\u2550\u2550\u255D${RESET} ${B}`,
` ${B}${S67}${B}`,
` ${B} ${CYAN}\u2554\u2550\u2550\u2550\u2550\u2550\u2550\u2550\u2550\u2550\u2550\u2550\u2550\u2550\u2550\u2550\u2550\u2550\u2550\u2550\u2550\u2550\u2550\u2550\u2550\u2550\u2550\u2550\u2550\u2550\u2550\u2550\u2550\u2550\u2550\u2550\u2550\u2557${RESET} ${B}`,
` ${B} ${CYAN}\u2551${RESET} ${WHITE}AI Penetration Testing Framework${RESET} ${CYAN}\u2551${RESET} ${B}`,
` ${B} ${CYAN}\u255A\u2550\u2550\u2550\u2550\u2550\u2550\u2550\u2550\u2550\u2550\u2550\u2550\u2550\u2550\u2550\u2550\u2550\u2550\u2550\u2550\u2550\u2550\u2550\u2550\u2550\u2550\u2550\u2550\u2550\u2550\u2550\u2550\u2550\u2550\u2550\u2550\u255D${RESET} ${B}`,
` ${B}${S67}${B}`,
];
if (version) {
const verStr = `v${version}`;
const verPadLeft = Math.floor((67 - verStr.length) / 2);
const verPadRight = 67 - verStr.length - verPadLeft;
lines.push(` ${B}${' '.repeat(verPadLeft)}${GRAY}${verStr}${RESET}${' '.repeat(verPadRight)}${B}`);
}
lines.push(
` ${B}${S67}${B}`,
` ${B} ${YELLOW}\uD83D\uDD10 DEFENSIVE SECURITY ONLY \uD83D\uDD10${RESET} ${B}`,
` ${B}${S67}${B}`,
` ${CYAN}\u255A${HR}\u255D${RESET}`,
` ${WHITE}Keygraph${RESET}${version ? ` ${DIM}v${version}${RESET}` : ''}`,
'',
);
...SHANNON.map((row, i) => ` ${paint(row, ramp[i] ?? '', ramp[i + 1] ?? '')}`),
'',
` ${WHITE}AI Pentester for Web Apps and APIs${RESET}`,
'',
` ${GRAY}-Authorized Security Testing Only-${RESET}`,
'',
];
console.log(lines.join('\n'));
}
/** Matches the divider width the CI wrappers and the scan renderer already use. */
const RULE_WIDTH = 60;
/**
* Plain-text banner for non-terminal output (CI logs, pipes, redirects).
* Drops the wordmark but keeps the authorized-use notice, which a reader of
* someone else's pipeline log still needs to see.
*/
export function displayPlainBanner(version?: string): void {
const rule = '─'.repeat(RULE_WIDTH);
console.log(rule);
console.log(version ? ` Shannon v${version}` : ' Shannon');
console.log(' AI Pentester for Web Apps and APIs, by Keygraph');
console.log(' Authorized security testing only.');
console.log(rule);
}
+58
View File
@@ -0,0 +1,58 @@
/**
* "Did you mean?" suggestions for mistyped commands and flags.
*
* A single Levenshtein-based matcher powers both the unknown-command path in the
* dispatcher and the unknown-option path in `parseArgs`, so a typo like `statsu`
* or `--workspce` points the user at the closest real name instead of just failing.
*/
/** Levenshtein edit distance between two strings (insertions, deletions, substitutions). */
export function editDistance(a: string, b: string): number {
if (a.length === 0) return b.length;
if (b.length === 0) return a.length;
// Rolling single row; `diagonal` and `above` carry the two neighbours a full grid would.
const row = Array.from({ length: b.length + 1 }, (_, j) => j);
for (let i = 1; i <= a.length; i++) {
let diagonal = row[0] as number;
row[0] = i;
for (let j = 1; j <= b.length; j++) {
const above = row[j] as number;
const cost = a[i - 1] === b[j - 1] ? 0 : 1;
row[j] = Math.min(above + 1, (row[j - 1] as number) + 1, diagonal + cost);
diagonal = above;
}
}
return row[b.length] as number;
}
/**
* The candidate closest to `input`, or undefined if none is near enough.
*
* A prefix match ("stat" -> "status") wins first; otherwise the lowest edit
* distance within a length-scaled threshold, so unrelated words don't match.
*/
export function closestMatch(input: string, candidates: readonly string[]): string | undefined {
if (input.length >= 2) {
const prefix = candidates.find((candidate) => candidate.startsWith(input));
if (prefix) return prefix;
}
let best: string | undefined;
let bestDistance = Number.POSITIVE_INFINITY;
for (const candidate of candidates) {
if (candidate.length <= 3) continue;
const distance = editDistance(input, candidate);
if (distance < bestDistance) {
bestDistance = distance;
best = candidate;
}
}
if (best === undefined) return undefined;
const threshold = Math.max(2, Math.floor(best.length / 3));
return bestDistance <= threshold ? best : undefined;
}
+200
View File
@@ -0,0 +1,200 @@
/**
* Thin Temporal client for reading one scan's state.
*
* A running scan is queried live (getProgress) and read via pendingActivities for
* the in-flight agents; a closed scan is read once from its result. Everything goes
* straight to the frontend on 127.0.0.1:7233 — the gRPC port the compose file
* publishes — so this needs Temporal up, but no worker of its own.
*/
import { setTimeout as sleep } from 'node:timers/promises';
import { Client, Connection, WorkflowFailedError, WorkflowNotFoundError } from '@temporalio/client';
import { ACTIVITY_TO_AGENT, type PipelineState } from './scan/pipeline.js';
const ADDRESS = '127.0.0.1:7233';
const NAMESPACE = 'default';
// WorkflowExecutionStatusName values that mean the scan has closed. RUNNING (and the unused
// CONTINUED_AS_NEW) are the only non-terminal states.
const TERMINAL_STATUSES: ReadonlySet<string> = new Set(['COMPLETED', 'FAILED', 'CANCELLED', 'TERMINATED', 'TIMED_OUT']);
export interface RunningAgent {
readonly agent: string;
readonly attempt: number;
readonly startedAt?: number;
readonly lastFailure?: string;
}
/** Convert a proto ITimestamp (seconds is a Long) to epoch millis. */
function timestampMs(
ts: { seconds?: { toString(): string } | number | null; nanos?: number | null } | null,
): number | undefined {
const seconds = ts?.seconds;
if (seconds == null) return undefined;
const secNum = typeof seconds === 'number' ? seconds : Number(seconds.toString());
return secNum * 1000 + (ts?.nanos ?? 0) / 1e6;
}
export interface ScanDescription {
/** WorkflowExecutionStatusName: RUNNING | COMPLETED | FAILED | CANCELLED | TERMINATED | TIMED_OUT | … */
readonly status: string;
readonly startedAt?: number;
readonly closedAt?: number;
readonly runningAgents: readonly RunningAgent[];
}
export type TerminalOutcome =
| { readonly kind: 'success'; readonly state: PipelineState }
| { readonly kind: 'failed'; readonly message: string };
let clientPromise: Promise<Client> | null = null;
function getClient(): Promise<Client> {
if (!clientPromise) {
clientPromise = Connection.connect({ address: ADDRESS }).then(
(connection) => new Client({ connection, namespace: NAMESPACE }),
);
}
return clientPromise;
}
/** Describe a scan: status, timing, and the agents currently running (from pendingActivities). Null if not found. */
export async function describeScan(workflowId: string): Promise<ScanDescription | null> {
const client = await getClient();
try {
const desc = await client.workflow.getHandle(workflowId).describe();
const runningAgents: RunningAgent[] = [];
for (const pending of desc.raw.pendingActivities ?? []) {
const agent = ACTIVITY_TO_AGENT[pending.activityType?.name ?? ''];
if (!agent) continue;
const lastFailure = pending.lastFailure?.message;
const startedAt = timestampMs(pending.scheduledTime ?? pending.lastStartedTime ?? null);
runningAgents.push({
agent,
attempt: pending.attempt ?? 1,
...(startedAt !== undefined ? { startedAt } : {}),
...(lastFailure ? { lastFailure } : {}),
});
}
return {
status: desc.status.name,
runningAgents,
...(desc.startTime ? { startedAt: desc.startTime.getTime() } : {}),
...(desc.closeTime ? { closedAt: desc.closeTime.getTime() } : {}),
};
} catch (err) {
if (err instanceof WorkflowNotFoundError) return null;
throw err;
}
}
/** Live progress of a running scan via the getProgress query. Null if the query can't be served (no worker). */
export async function queryProgress(workflowId: string): Promise<PipelineState | null> {
const client = await getClient();
try {
return await client.workflow.getHandle(workflowId).query<PipelineState>('getProgress');
} catch {
// The query needs a live worker; a just-closed scan may have none. Caller falls back to the result.
return null;
}
}
/**
* Deepest message in a Temporal failure's cause chain — the real reason nested under generic
* wrappers (WorkflowFailedError → ActivityFailure → ApplicationFailure). Covers failed, cancelled,
* and terminated alike. Mirrors the SDK's `rootCause` (only exported from @temporalio/common).
*/
function rootFailureMessage(err: WorkflowFailedError): string {
let message = err.message;
let cause: unknown = err.cause;
while (cause instanceof Error && cause.message) {
message = cause.message;
cause = cause.cause;
}
return message;
}
/** How a {@link waitForWorkflowClose} watch ended. */
export type WatchEnd = { readonly reason: 'closed' } | { readonly reason: 'unreachable'; readonly lastError: string };
export interface WatchOptions {
/** Poll interval in ms (default 3000). */
readonly pollMs?: number;
/** Consecutive connection failures before giving up (default 10 → ~30s at the default interval). */
readonly maxConnectFailures?: number;
/** Consecutive connection failures before {@link onConnectionTrouble} fires once (default 3). */
readonly warnAfterFailures?: number;
/** Abort the watch (the caller stopped for another reason, e.g. Ctrl-C). */
readonly signal?: AbortSignal;
/** Called once when contact is first lost, so a live follower's log isn't silent during the outage. */
readonly onConnectionTrouble?: (lastError: string) => void;
/** Called once when contact is regained after {@link onConnectionTrouble} fired. */
readonly onReconnected?: () => void;
}
/**
* Resolve once the scan is no longer running, using the workflow's Temporal status as the
* completion signal. Ends on a terminal status, a not-found workflow (closed past retention), or
* maxConnectFailures consecutive unreachable polls (a scan can't progress while its Temporal is
* down, so sustained no-contact is a safe stop). Never rejects; connection errors surface via the
* callbacks and the returned {@link WatchEnd}.
*/
export async function waitForWorkflowClose(workflowId: string, opts: WatchOptions = {}): Promise<WatchEnd> {
const pollMs = opts.pollMs ?? 3000;
const maxConnectFailures = opts.maxConnectFailures ?? 10;
const warnAfterFailures = opts.warnAfterFailures ?? 3;
const signal = opts.signal;
let connectFailures = 0;
let lastError = '';
let warned = false;
while (!signal?.aborted) {
try {
const desc = await describeScan(workflowId);
if (desc === null || TERMINAL_STATUSES.has(desc.status)) {
return { reason: 'closed' };
}
// Reachable and still RUNNING — reset the failure streak and note any recovery.
if (warned) {
warned = false;
opts.onReconnected?.();
}
connectFailures = 0;
} catch (err) {
connectFailures++;
lastError = err instanceof Error ? err.message : String(err);
if (!warned && connectFailures >= warnAfterFailures) {
warned = true;
opts.onConnectionTrouble?.(lastError);
}
if (connectFailures >= maxConnectFailures) {
return { reason: 'unreachable', lastError };
}
}
try {
await sleep(pollMs, undefined, { signal });
} catch {
break; // Aborted mid-wait by the caller.
}
}
return { reason: 'closed' };
}
/** Final state of a closed scan: success carries the full PipelineState, failure carries the message. */
export async function getTerminalOutcome(workflowId: string): Promise<TerminalOutcome> {
const client = await getClient();
try {
const state = (await client.workflow.getHandle(workflowId).result()) as PipelineState;
return { kind: 'success', state };
} catch (err) {
if (err instanceof WorkflowFailedError) {
return { kind: 'failed', message: rootFailureMessage(err) };
}
throw err;
}
}
+3 -3
View File
@@ -3,6 +3,8 @@
* whether the user can be prompted interactively.
*/
import { fail } from './errors.js';
/** True when stdout is a real terminal — safe for color, cursor moves, and spinners. */
export function stdoutIsTerminal(): boolean {
return !!process.stdout.isTTY;
@@ -28,7 +30,5 @@ export function supportsColor(): boolean {
/** Exit with a clear error when an interactive-only command has no terminal, instead of hanging on a prompt. */
export function requireInteractive(command: string, alternative: string): void {
if (isInteractive()) return;
console.error(`ERROR: '${command}' needs an interactive terminal.`);
console.error(alternative);
process.exit(1);
fail(`'${command}' needs an interactive terminal.`, alternative);
}
+60
View File
@@ -0,0 +1,60 @@
/**
* Terminal status output for long-running steps.
*
* Commands are run with their output captured rather than inherited, so raw docker
* plumbing never floods the terminal. Progress is shown with a `@clack/prompts`
* spinner. On failure the captured output is printed so the error stays visible
* instead of being swallowed.
*/
import { spawn } from 'node:child_process';
import * as p from '@clack/prompts';
export interface StepResult {
ok: boolean;
output: string;
}
/**
* Run a command capturing stdout and stderr. Resolves the exit result and combined
* output; never rejects. Callers that want a spinner wrap this in one themselves.
*/
export function spawnCaptured(cmd: string, args: string[]): Promise<StepResult> {
return new Promise((resolve) => {
let output = '';
const child = spawn(cmd, args, { stdio: ['ignore', 'pipe', 'pipe'] });
child.stdout?.on('data', (chunk) => {
output += chunk.toString();
});
child.stderr?.on('data', (chunk) => {
output += chunk.toString();
});
child.on('close', (code) => resolve({ ok: code === 0, output }));
child.on('error', () => resolve({ ok: false, output }));
});
}
/** Print captured command output to stderr, so a failure is never swallowed. */
export function surfaceOutput(output: string): void {
const trimmed = output.trim();
if (trimmed) process.stderr.write(`${trimmed}\n`);
}
/**
* Run a command as a labeled step, with a spinner over it. On failure the captured
* output is surfaced. Returns the exit result and captured output.
*/
export async function runStep(label: string, cmd: string, args: string[]): Promise<StepResult> {
const spinner = p.spinner();
spinner.start(label);
const result = await spawnCaptured(cmd, args);
if (result.ok) {
spinner.stop(label);
} else {
spinner.error(label);
surfaceOutput(result.output);
}
return result;
}
+1 -1
View File
@@ -164,7 +164,7 @@
"sarif": {
"type": "string",
"enum": ["true", "false"],
"description": "Emit a SARIF 2.1.0 log (report.sarif) beside the report. Requires exploit=true; ignored otherwise."
"description": "Emit a SARIF 2.1.0 log (report.sarif) beside the report. On by default for exploit runs; set \"false\" to opt out. Ignored when exploit=false."
}
},
"additionalProperties": false
+3 -2
View File
@@ -96,8 +96,9 @@ rules:
# Report filters applied by the report agent when assembling the final report (optional).
# Example below is illustrative; edit, remove, or add sections as needed.
# report:
# # Emit a SARIF 2.1.0 log (report.sarif) beside the report. Requires exploit: "true".
# sarif: "true"
# # SARIF 2.1.0 log (report.sarif) beside the report. On by default for exploit runs;
# # set "false" to opt out. Ignored when exploit is "false".
# sarif: "false"
# min_severity: low
# min_confidence: low
# guidance: |
+3 -3
View File
@@ -19,9 +19,9 @@
"clean": "rm -rf dist"
},
"dependencies": {
"@earendil-works/pi-agent-core": "^0.82.1",
"@earendil-works/pi-ai": "^0.82.1",
"@earendil-works/pi-coding-agent": "^0.82.1",
"@earendil-works/pi-agent-core": "^0.84.2",
"@earendil-works/pi-ai": "^0.84.2",
"@earendil-works/pi-coding-agent": "^0.84.2",
"@gotgenes/pi-permission-system": "^10.9.0",
"@temporalio/activity": "^1.11.0",
"@temporalio/client": "^1.11.0",
+20 -1
View File
@@ -21,8 +21,10 @@
* built over an in-memory credential store primed from the environment.
*/
import { existsSync } from 'node:fs';
import path from 'node:path';
import type { Api, Credential, CredentialInfo, CredentialStore, Model } from '@earendil-works/pi-ai';
import { ModelRuntime } from '@earendil-works/pi-coding-agent';
import { getAgentDir, ModelRuntime } from '@earendil-works/pi-coding-agent';
/**
* Providers Shannon curates with their own credential variables, config sections,
@@ -203,12 +205,29 @@ class RuntimeCredentialStore implements CredentialStore {
}
}
/** The file pi reads credentials from: the agent dir's auth.json. */
function piAuthPath(): string {
return path.join(getAgentDir(), 'auth.json');
}
/** Whether the host's pi credentials are mounted (auth.json present in the agent dir). */
export function piAuthPresent(): boolean {
return existsSync(piAuthPath());
}
/**
* Build a ModelRuntime whose only credential is the one supplied. Model catalogs
* stay offline (`allowModelNetwork` defaults to false) so a scan never blocks on
* a catalog refresh.
*
* When the host's pi auth.json is present, the runtime reads it instead: pi's
* disk-backed store resolves the credential. The mount is writable so OAuth
* refreshes persist to the host for subsequent runs.
*/
export async function createModelRuntime(providerId: string, apiKey: string | undefined): Promise<ModelRuntime> {
if (piAuthPresent()) {
return ModelRuntime.create({ authPath: piAuthPath() });
}
return ModelRuntime.create({ credentials: new RuntimeCredentialStore(providerId, apiKey) });
}
+17
View File
@@ -290,6 +290,14 @@ export async function runPiPrompt(
// Declared out here so the catch can bill spend accrued before a failure.
let session: AgentSession | undefined;
// Abort the in-flight agent when the Temporal activity is cancelled (UI/CLI cancel).
// Without this the top-level session runs to startToCloseTimeout despite the cancel.
const onCancellation = (): void => {
void session?.abort().catch(() => {
// Best-effort — the session is torn down regardless once the prompt unwinds.
});
};
progress.start();
try {
@@ -307,6 +315,13 @@ export async function runPiPrompt(
resourceLoader,
}));
// Wire activity cancellation to the session now that it exists.
if (cancellationSignal?.aborted) {
onCancellation();
} else {
cancellationSignal?.addEventListener('abort', onCancellation, { once: true });
}
// 5. Map pi events to audit logging + progress + error capture.
session.subscribe((event: AgentSessionEvent) => {
switch (event.type) {
@@ -414,5 +429,7 @@ export async function runPiPrompt(
cacheWriteTokens: usage.cacheWriteTokens,
retryable: isRetryableFailure(err),
};
} finally {
cancellationSignal?.removeEventListener('abort', onCancellation);
}
}
+19 -13
View File
@@ -303,13 +303,17 @@ export class WorkflowLogger {
* Output: "Error: phase context\n ErrorType\n ..."
*/
private formatErrorBlock(errorString: string): string {
const segments = errorString.split('|');
const label = 'Error: ';
const indent = ' '.repeat(label.length);
const lines = segments.map((segment, i) => (i === 0 ? `${label}${segment.trim()}` : `${indent}${segment.trim()}`));
// Segments are delimited by '|'; a segment's own embedded newlines (e.g. a multi-line
// validation message) become their own lines so each aligns under the label.
const lines = errorString
.split(/[|\n]/)
.map((segment) => segment.trim())
.filter((segment) => segment.length > 0);
return `${lines.join('\n')}\n`;
return `${lines.map((line, i) => (i === 0 ? `${label}${line}` : `${indent}${line}`)).join('\n')}\n`;
}
/**
@@ -336,17 +340,19 @@ export class WorkflowLogger {
lines.push(this.formatErrorBlock(summary.error).trimEnd());
}
lines.push('');
lines.push('Agent Breakdown:');
if (summary.completedAgents.length > 0) {
lines.push('');
lines.push('Agent Breakdown:');
for (const agentName of summary.completedAgents) {
const metrics = summary.agentMetrics[agentName];
if (metrics) {
const duration = formatDuration(metrics.durationMs);
const cost = metrics.costUsd !== null ? `$${metrics.costUsd.toFixed(4)}` : 'N/A';
lines.push(` - ${agentName} (${duration}, ${cost})`);
} else {
lines.push(` - ${agentName}`);
for (const agentName of summary.completedAgents) {
const metrics = summary.agentMetrics[agentName];
if (metrics) {
const duration = formatDuration(metrics.durationMs);
const cost = metrics.costUsd !== null ? `$${metrics.costUsd.toFixed(4)}` : 'N/A';
lines.push(` - ${agentName} (${duration}, ${cost})`);
} else {
lines.push(` - ${agentName}`);
}
}
}
+18
View File
@@ -0,0 +1,18 @@
// Copyright (C) 2025 Keygraph, Inc.
/**
* Centralized brand strings for report deliverables.
*
* Kept in two parts because the two renderers join them differently: the Typst
* template splits its `brand` input on a pipe to set the cover's two lines
* (`report.typ:212`), while a human-read line takes an em dash.
*/
export const PRODUCT_NAME = 'Shannon';
export const PRODUCT_DESCRIPTOR = 'AI Pentester by Keygraph';
/** Cover wordmark for the Typst template, which parses the pipe. */
export const TYPST_BRAND = `${PRODUCT_NAME} | ${PRODUCT_DESCRIPTOR}`;
/** Attribution line for prose surfaces. */
export const BRAND_LOCKUP = `${PRODUCT_NAME} — ${PRODUCT_DESCRIPTOR}`;
+2 -1
View File
@@ -679,7 +679,8 @@ export const distributeConfig = (config: Config | null): DistributedConfig => {
const exploit = config?.exploit !== undefined ? config.exploit === 'true' : true;
const report = {
sarif: config?.report?.sarif === 'true',
// Default on; only an explicit "false" opts out.
sarif: config?.report?.sarif !== 'false',
...(config?.report?.min_severity && { min_severity: config.report.min_severity }),
...(config?.report?.min_confidence && { min_confidence: config.report.min_confidence }),
...(config?.report?.guidance && { guidance: config.report.guidance.trim() }),
+12 -3
View File
@@ -9,6 +9,9 @@ const WORKER_ROOT = path.resolve(import.meta.dirname, '..');
export const PROMPTS_DIR = path.join(WORKER_ROOT, 'prompts');
export const CONFIGS_DIR = path.join(WORKER_ROOT, 'configs');
/** Bundled Typst template that renders report.json into the PDF report. */
export const TYPST_TEMPLATE = path.join(WORKER_ROOT, 'templates', 'typst', 'report.typ');
/** Compiled pi extension dir that enforces bounded `bash` timeouts (resolved from dist/) */
export const BASH_TIMEOUT_EXTENSION_DIR = path.join(import.meta.dirname, 'ai', 'extensions', 'bash-timeout');
@@ -28,13 +31,19 @@ export const INTERNAL_DIR = '.shannon';
/** Filename of the assembled report inside the deliverables dir (internal, source of the surfaced copy) */
export const ASSEMBLED_REPORT_FILENAME = 'comprehensive_security_assessment_report.md';
/** Filename of the human-facing final report surfaced at the run directory root */
export const FINAL_REPORT_FILENAME = 'Security-Assessment-Report.md';
/** Filename of the compiled PDF report inside the deliverables dir (internal, source of the surfaced copy) */
export const ASSEMBLED_REPORT_PDF_FILENAME = 'comprehensive_security_assessment_report.pdf';
/** Filename of the human-facing PDF report surfaced at the run directory root */
export const FINAL_REPORT_PDF_FILENAME = 'Security-Assessment-Report.pdf';
/** Filename of the human-facing markdown report surfaced at the run directory root, alongside the PDF */
export const FINAL_REPORT_MD_FILENAME = 'Security-Assessment-Report.md';
/** Structured findings the report agent emits; the markdown report is rendered from it. */
export const REPORT_JSON_FILENAME = 'report.json';
/** SARIF 2.1.0 log, written only for exploit=true runs when report.sarif is enabled. */
/** SARIF 2.1.0 log, written for exploit=true runs unless report.sarif is set to false. */
export const SARIF_FILENAME = 'report.sarif';
/**
+1 -1
View File
@@ -29,7 +29,7 @@ export function getAgentGitPaths(agentName: AgentName): string[] {
paths.push(queueFilename);
}
// The report agent also emits the structured findings the markdown is rendered from, and the
// SARIF log when enabled. Listing the log unconditionally is harmless when it was not written,
// SARIF log when produced. Listing the log unconditionally is harmless when it was not written,
// and keeps a stale one from surviving the rollback of a failed attempt.
if (agentName === 'report') {
paths.push(REPORT_JSON_FILENAME);
+99
View File
@@ -0,0 +1,99 @@
// Copyright (C) 2025 Keygraph, Inc.
//
// This program is free software: you can redistribute it and/or modify
// it under the terms of the GNU Affero General Public License version 3
// as published by the Free Software Foundation.
/**
* Typst PDF renderer.
*
* Adapts the structured report.json into the Typst-shaped schema and compiles
* it to a PDF with the bundled report.typ template. Compilation runs in an
* isolated temp dir: the template is copied in and the adapted JSON is written
* beside it so `--root` can scope every file read to that dir, matching how the
* template resolves `--input data=/data.json`.
*
* The `typst` binary is installed in the worker image and resolved from PATH.
*/
import { execFile } from 'node:child_process';
import { existsSync } from 'node:fs';
import { copyFile, cp, mkdir, mkdtemp, rm, writeFile } from 'node:fs/promises';
import { tmpdir } from 'node:os';
import path from 'node:path';
import { promisify } from 'node:util';
import { TYPST_BRAND } from '../branding.js';
import { adaptReportToTypst } from './report-json-adapter.js';
import type { ReportData } from './report-renderer.js';
const execFileAsync = promisify(execFile);
const DEFAULT_TESTER = 'Shannon';
const DEFAULT_BRAND = TYPST_BRAND;
const DATA_FILENAME = 'data.json';
const TEMPLATE_FILENAME = 'report.typ';
const OUTPUT_FILENAME = 'report.pdf';
export interface RenderReportPdfOptions {
/** Structured report data (report.json contents), pre-assembly. */
readonly reportData: ReportData;
/** Absolute path to the bundled report.typ template. */
readonly templatePath: string;
/** Absolute path where the compiled PDF should be written. */
readonly outputPath: string;
/** Name shown on the cover/footer. Defaults to "Shannon". */
readonly tester?: string;
/** Wordmark shown on the cover. Defaults to "Shannon | AI Pentester by Keygraph". */
readonly brand?: string;
}
/**
* Compile the report to a PDF at `outputPath`.
*
* Throws if adaptation or `typst compile` fails; callers treat the PDF as a
* secondary artifact and should not let a failure here fail the run.
*/
export async function renderReportPdf(options: RenderReportPdfOptions): Promise<void> {
const { reportData, templatePath, outputPath } = options;
const tester = options.tester ?? DEFAULT_TESTER;
const brand = options.brand ?? DEFAULT_BRAND;
const typstData = adaptReportToTypst(reportData);
const workDir = await mkdtemp(path.join(tmpdir(), 'shannon-typst-'));
try {
const templateInWorkDir = path.join(workDir, TEMPLATE_FILENAME);
const dataInWorkDir = path.join(workDir, DATA_FILENAME);
const pdfInWorkDir = path.join(workDir, OUTPUT_FILENAME);
await copyFile(templatePath, templateInWorkDir);
// Ship the template's assets (e.g. the cover logo) so `--root`-scoped image reads resolve.
const assetsDir = path.join(path.dirname(templatePath), 'assets');
if (existsSync(assetsDir)) {
await cp(assetsDir, path.join(workDir, 'assets'), { recursive: true });
}
await writeFile(dataInWorkDir, JSON.stringify(typstData), 'utf-8');
await execFileAsync('typst', [
'compile',
'--root',
workDir,
'--input',
`data=/${DATA_FILENAME}`,
'--input',
`tester=${tester}`,
'--input',
`brand=${brand}`,
templateInWorkDir,
pdfInWorkDir,
]);
await mkdir(path.dirname(outputPath), { recursive: true });
await copyFile(pdfInWorkDir, outputPath);
} finally {
await rm(workDir, { recursive: true, force: true });
}
}
+6 -2
View File
@@ -42,6 +42,7 @@ import {
type ModelSpec,
type OpenAiFormat,
PI_CATALOG_URL,
piAuthPresent,
resolveGatewayFormat,
resolveModel,
resolveModelSpec,
@@ -296,6 +297,7 @@ function credentialHint(providerId: string): string {
/** Human-readable label for which credential path a run is using. */
function describeAuth(providerId: string, baseUrl: string | undefined): string {
if (baseUrl) return `custom endpoint (${baseUrl})`;
if (piAuthPresent()) return `${providerId} credentials from pi auth.json`;
if (providerId === 'amazon-bedrock') return 'Bedrock bearer token';
return `${providerId} API key`;
}
@@ -341,9 +343,11 @@ async function validateCredentials(logger: ActivityLogger): Promise<Result<void,
);
}
// With a mounted pi auth.json the env-var checks don't apply — step 5's probe validates it.
const isBedrock = spec.providerId === 'amazon-bedrock';
const missing = isBedrock ? ['AWS_REGION', 'AWS_BEARER_TOKEN_BEDROCK'].filter((n) => !process.env[n]) : [];
if (missing.length > 0 || (!isBedrock && !credentials.apiKey)) {
const missing =
isBedrock && !piAuthPresent() ? ['AWS_REGION', 'AWS_BEARER_TOKEN_BEDROCK'].filter((n) => !process.env[n]) : [];
if (!piAuthPresent() && (missing.length > 0 || (!isBedrock && !credentials.apiKey))) {
return err(
new PentestError(
`No credentials found for provider "${spec.providerId}". Set ${credentialHint(spec.providerId)} in .env.`,
@@ -0,0 +1,293 @@
// Copyright (C) 2025 Keygraph, Inc.
//
// This program is free software: you can redistribute it and/or modify
// it under the terms of the GNU Affero General Public License version 3
// as published by the Free Software Foundation.
/**
* Programmatic adapter: report.json → Typst ReportData JSON.
*
* Converts the renderer-neutral structured report output (produced by the
* finding-collector + set-report-meta CLI) into the Typst-specific schema that
* report.typ consumes.
*
* All Typst-specific concepts (PascalCase enums, computed aggregations,
* exploitedByType grouping) are confined to this file. The rest of the
* pipeline knows nothing about the Typst shape.
*/
import type { AddFindingInput, AdditionalSection, StepItem, StructuredStep } from '../collectors/finding-collector.js';
import type {
ExploitsReportData,
FindingsReportData,
TypstCategory,
TypstConfidence,
ReportData as TypstReportData,
TypstSeverity,
TypstStatus,
} from './report-output-schema.js';
import type { ReportData } from './report-renderer.js';
// ============================================================================
// CASING TRANSFORMS
// ============================================================================
const SEVERITY_MAP: Record<string, TypstSeverity> = {
critical: 'Critical',
high: 'High',
medium: 'Medium',
low: 'Low',
};
const STATUS_MAP: Record<string, TypstStatus> = {
exploited: 'Exploited',
out_of_scope: 'OutOfScope',
blocked_by_constraints: 'BlockedByConstraints',
false_positive: 'FalsePositive',
};
const CONFIDENCE_MAP: Record<string, TypstConfidence> = {
high: 'High',
medium: 'Medium',
low: 'Low',
};
const VALID_CATEGORIES = new Set<TypstCategory>([
'Authentication',
'Authorization',
'XSS',
'Injection',
'SSRF',
'Other',
]);
function toTypstSeverity(s: string): TypstSeverity {
return SEVERITY_MAP[s] ?? 'Low';
}
function toTypstStatus(s: string): TypstStatus {
return STATUS_MAP[s] ?? 'Exploited';
}
function toTypstConfidence(s: string): TypstConfidence {
return CONFIDENCE_MAP[s] ?? 'Medium';
}
function toTypstCategory(s: string): TypstCategory {
if (VALID_CATEGORIES.has(s as TypstCategory)) return s as TypstCategory;
return 'Other';
}
// ============================================================================
// STEP / ITEM TRANSFORMS
// ============================================================================
function adaptStepItem(item: StepItem): StepItem {
return item;
}
function adaptStep(step: StructuredStep, index: number): { number: number; title?: string; items: StepItem[] } {
return {
number: index + 1,
...(step.title && { title: step.title }),
items: step.items.map(adaptStepItem),
};
}
function adaptAdditionalSection(section: AdditionalSection): { heading: string; items: StepItem[] } {
return {
heading: section.heading,
items: section.items.map(adaptStepItem),
};
}
// ============================================================================
// AGGREGATION HELPERS
// ============================================================================
interface CategoryGroup {
category: TypstCategory;
findings: AddFindingInput[];
}
function groupByCategory(findings: readonly AddFindingInput[]): CategoryGroup[] {
const map = new Map<TypstCategory, AddFindingInput[]>();
for (const f of findings) {
const cat = toTypstCategory(f.category);
const list = map.get(cat) ?? [];
list.push(f);
map.set(cat, list);
}
return Array.from(map.entries()).map(([category, fs]) => ({ category, findings: fs }));
}
function countBySeverity(findings: readonly AddFindingInput[]): Record<TypstSeverity, number> {
const counts: Record<string, number> = {
Critical: 0,
High: 0,
Medium: 0,
Low: 0,
};
for (const f of findings) {
const sev = toTypstSeverity(f.severity);
counts[sev] = (counts[sev] ?? 0) + 1;
}
return counts as Record<TypstSeverity, number>;
}
// ============================================================================
// EXPLOIT MODE ADAPTER
// ============================================================================
function adaptExploitsMode(data: ReportData): ExploitsReportData {
const { report_meta, findings } = data;
const groups = groupByCategory(findings);
const sevCounts = countBySeverity(findings);
const statusCounts = { Exploited: 0, OutOfScope: 0, BlockedByConstraints: 0, FalsePositive: 0 };
for (const f of findings) {
const s = toTypstStatus(f.status ?? 'exploited');
statusCounts[s]++;
}
const exploitedFindings = findings.filter((f) => (f.status ?? 'exploited') === 'exploited');
return {
mode: 'exploits' as const,
meta: {
target: report_meta.target,
assessmentDate: report_meta.assessment_date,
classification: 'CONFIDENTIAL',
},
scope: report_meta.scope,
exploitedByType: groups.map((g) => {
const exploited = g.findings.filter((f) => (f.status ?? 'exploited') === 'exploited');
if (exploited.length === 0) {
return {
category: g.category,
narrative: `No ${g.category.toLowerCase()} vulnerabilities were successfully exploited during this assessment.`,
};
}
return {
category: g.category,
bullets: exploited.map((f) => ({ id: f.finding_id, description: f.title })),
};
}),
summary: {
totalIdentified: findings.length,
successfullyExploited: exploitedFindings.length,
exploitedBreakdown: groups
.map((g) => ({
category: g.category,
count: g.findings.filter((f) => (f.status ?? 'exploited') === 'exploited').length,
}))
.filter((e) => e.count > 0),
criticalFindings: findings.filter((f) => f.severity === 'critical').map((f) => `${f.finding_id}: ${f.title}`),
},
findings: findings.map((f) => ({
id: f.finding_id,
title: f.title,
category: toTypstCategory(f.category),
severity: toTypstSeverity(f.severity),
status: toTypstStatus(f.status ?? 'exploited'),
summary: {
vulnerableLocation: f.vulnerable_location,
overview: f.overview,
impact: f.impact,
},
// This branch only runs for an exploitative report, where the schema made these
// required. The fallbacks keep the superset type honest rather than assuming.
prerequisites: f.prerequisites ?? '',
exploitationSteps: (f.exploitation_steps ?? []).map(adaptStep),
proofOfImpact: (f.proof_of_impact ?? []).map(adaptStepItem),
...(f.notes && f.notes.length > 0 && { notes: f.notes.map(adaptStepItem) }),
...(f.additional_sections &&
f.additional_sections.length > 0 && {
additionalSections: f.additional_sections.map(adaptAdditionalSection),
}),
})),
derivedCounts: {
bySeverity: sevCounts,
byStatus: statusCounts,
},
};
}
// ============================================================================
// FINDINGS MODE ADAPTER
// ============================================================================
function adaptFindingsMode(data: ReportData): FindingsReportData {
const { report_meta, findings } = data;
const groups = groupByCategory(findings);
const sevCounts = countBySeverity(findings);
const confidenceCounts = { High: 0, Medium: 0, Low: 0 };
for (const f of findings) {
const c = toTypstConfidence(f.confidence ?? 'medium');
confidenceCounts[c]++;
}
return {
mode: 'findings' as const,
meta: {
target: report_meta.target,
assessmentDate: report_meta.assessment_date,
classification: 'CONFIDENTIAL',
},
scope: report_meta.scope,
identifiedByType: groups.map((g) => {
if (g.findings.length === 0) {
return {
category: g.category,
narrative: `No ${g.category.toLowerCase()} vulnerabilities were identified during this assessment.`,
};
}
return {
category: g.category,
bullets: g.findings.map((f) => ({ id: f.finding_id, description: f.title })),
};
}),
summary: {
totalIdentified: findings.length,
identifiedBreakdown: groups.map((g) => ({
category: g.category,
count: g.findings.length,
})),
criticalFindings: findings.filter((f) => f.severity === 'critical').map((f) => `${f.finding_id}: ${f.title}`),
},
findings: findings.map((f) => ({
id: f.finding_id,
title: f.title,
category: toTypstCategory(f.category),
severity: toTypstSeverity(f.severity),
confidence: toTypstConfidence(f.confidence ?? 'medium'),
summary: {
vulnerableLocation: f.vulnerable_location,
overview: f.overview,
impact: f.impact,
},
...(f.notes && f.notes.length > 0 && { notes: f.notes.map(adaptStepItem) }),
...(f.additional_sections &&
f.additional_sections.length > 0 && {
additionalSections: f.additional_sections.map(adaptAdditionalSection),
}),
})),
derivedCounts: {
bySeverity: sevCounts,
byConfidence: confidenceCounts,
},
};
}
// ============================================================================
// PUBLIC API
// ============================================================================
export function adaptReportToTypst(data: ReportData): TypstReportData {
const exploitEnabled = data.report_meta.exploit ?? true;
if (exploitEnabled) {
return adaptExploitsMode(data);
}
return adaptFindingsMode(data);
}
@@ -0,0 +1,157 @@
// Copyright (C) 2025 Keygraph, Inc.
//
// This program is free software: you can redistribute it and/or modify
// it under the terms of the GNU Affero General Public License version 3
// as published by the Free Software Foundation.
/**
* TypeScript types for the structured report the Typst template consumes, in two
* shapes keyed by a `mode` discriminator: `exploits` (exploit=true) and `findings`
* (exploit=false, analysis-only). Types only — the object is built programmatically
* in report-json-adapter.ts, so these exist to keep the adapter and report.typ in sync.
*/
// === Shared primitives ===
export type TypstSeverity = 'Critical' | 'High' | 'Medium' | 'Low';
export type TypstStatus = 'Exploited' | 'OutOfScope' | 'BlockedByConstraints' | 'FalsePositive';
export type TypstConfidence = 'High' | 'Medium' | 'Low';
export type TypstCategory = 'Authentication' | 'Authorization' | 'XSS' | 'Injection' | 'SSRF' | 'Other';
export interface CodeBlock {
readonly language: string;
readonly content: string;
}
export type StepItem =
| { readonly kind: 'prose'; readonly text: string }
| { readonly kind: 'code'; readonly block: CodeBlock };
export interface Step {
readonly number: number;
readonly title?: string;
readonly items: readonly StepItem[];
}
export interface AdditionalSection {
readonly heading: string;
readonly items: readonly StepItem[];
}
export interface FindingSummary {
readonly vulnerableLocation: string;
readonly overview: string;
readonly impact: string;
}
export interface Meta {
readonly target: string;
readonly assessmentDate: string;
readonly tester?: string;
readonly application?: string;
readonly classification: string;
}
export interface CategoryCount {
readonly category: TypstCategory;
readonly count: number;
readonly note?: string;
}
export type SeverityCounts = Record<TypstSeverity, number>;
export type StatusCounts = Record<TypstStatus, number>;
export type ConfidenceCounts = Record<TypstConfidence, number>;
export interface TypeEntryBullet {
readonly id: string;
readonly description: string;
}
// === Exploits-mode schema ===
export interface ExploitFinding {
readonly id: string;
readonly title: string;
readonly category: TypstCategory;
readonly severity: TypstSeverity;
readonly status: TypstStatus;
readonly summary: FindingSummary;
readonly prerequisites: string;
readonly exploitationSteps: readonly Step[];
readonly proofOfImpact: readonly StepItem[];
readonly notes?: readonly StepItem[];
readonly additionalSections?: readonly AdditionalSection[];
}
export interface ExploitedByTypeEntry {
readonly category: TypstCategory;
readonly bullets?: readonly TypeEntryBullet[];
readonly narrative?: string;
}
export interface ExploitsReportData {
readonly mode: 'exploits';
readonly meta: Meta;
readonly scope: string;
readonly exploitedByType: readonly ExploitedByTypeEntry[];
readonly summary: {
readonly totalIdentified: number;
readonly successfullyExploited: number;
readonly exploitedBreakdown: readonly CategoryCount[];
readonly outOfScope?: {
readonly total: number;
readonly breakdown?: readonly CategoryCount[];
readonly note?: string;
};
readonly blockedByConstraints?: {
readonly total: number;
readonly note?: string;
};
readonly criticalFindings: readonly string[];
};
readonly findings: readonly ExploitFinding[];
readonly derivedCounts: {
readonly bySeverity: SeverityCounts;
readonly byStatus: StatusCounts;
};
}
// === Findings-mode schema (analysis-only, exploit=false runs) ===
export interface AnalysisFinding {
readonly id: string;
readonly title: string;
readonly category: TypstCategory;
readonly severity: TypstSeverity;
readonly confidence: TypstConfidence;
readonly summary: FindingSummary;
readonly notes?: readonly StepItem[];
readonly additionalSections?: readonly AdditionalSection[];
}
export interface IdentifiedByTypeEntry {
readonly category: TypstCategory;
readonly bullets?: readonly TypeEntryBullet[];
readonly narrative?: string;
}
export interface FindingsReportData {
readonly mode: 'findings';
readonly meta: Meta;
readonly scope: string;
readonly identifiedByType: readonly IdentifiedByTypeEntry[];
readonly summary: {
readonly totalIdentified: number;
readonly identifiedBreakdown: readonly CategoryCount[];
readonly criticalFindings: readonly string[];
};
readonly findings: readonly AnalysisFinding[];
readonly derivedCounts: {
readonly bySeverity: SeverityCounts;
readonly byConfidence: ConfidenceCounts;
};
}
// === Discriminated union for downstream consumers that handle both ===
export type ReportData = ExploitsReportData | FindingsReportData;
@@ -12,6 +12,7 @@
* report agent previously wrote by hand. No LLM in the loop.
*/
import { BRAND_LOCKUP } from '../branding.js';
import type { AddFindingInput, AdditionalSection, StepItem, StructuredStep } from '../collectors/finding-collector.js';
import type { VulnClass } from '../types/config.js';
@@ -212,6 +213,8 @@ export function renderReport(data: ReportData): string {
// 1. Executive Summary
sections.push('# Security Assessment Report');
sections.push('');
sections.push(`*${BRAND_LOCKUP}*`);
sections.push('');
sections.push('## Executive Summary');
sections.push(`- Target: ${report_meta.target}`);
sections.push(`- Assessment Date: ${report_meta.assessment_date}`);
+21 -9
View File
@@ -7,8 +7,10 @@
import { fs, path } from 'zx';
import {
ASSEMBLED_REPORT_FILENAME,
ASSEMBLED_REPORT_PDF_FILENAME,
deliverablesDir,
FINAL_REPORT_FILENAME,
FINAL_REPORT_MD_FILENAME,
FINAL_REPORT_PDF_FILENAME,
resolveSessionJsonPath,
SARIF_FILENAME,
} from '../paths.js';
@@ -174,11 +176,12 @@ export async function injectModelIntoReport(
/**
* Surface the run's deliverables at the run directory's top level, so a customer opening the run
* folder sees the report without digging through internals. Sources stay in the deliverables dir
* (git-checkpointed, used by resume).
* (git-checkpointed, used by resume). Both the PDF and the markdown report are surfaced here as the
* customer-facing copies.
*
* The SARIF log is surfaced beside it when present, since a CI step consuming it needs a stable
* path and cannot be expected to reach into the internals directory. It is absent whenever the
* run was analysis-only or `report.sarif` was not enabled.
* run was analysis-only or `report.sarif` was set to false.
*/
export async function copyReportToRunRoot(
repoPath: string,
@@ -188,13 +191,22 @@ export async function copyReportToRunRoot(
): Promise<void> {
const dir = deliverablesDir(repoPath, deliverablesSubdir);
const source = path.join(dir, ASSEMBLED_REPORT_FILENAME);
if (await fs.pathExists(source)) {
const destination = path.join(runDir, FINAL_REPORT_FILENAME);
await fs.copy(source, destination, { overwrite: true });
logger.info(`Surfaced report at ${destination}`);
const pdfSource = path.join(dir, ASSEMBLED_REPORT_PDF_FILENAME);
if (await fs.pathExists(pdfSource)) {
const destination = path.join(runDir, FINAL_REPORT_PDF_FILENAME);
await fs.copy(pdfSource, destination, { overwrite: true });
logger.info(`Surfaced PDF report at ${destination}`);
} else {
logger.warn(`Final report not found, skipping ${FINAL_REPORT_FILENAME}`);
logger.warn(`PDF report not found, skipping ${FINAL_REPORT_PDF_FILENAME}`);
}
const markdownSource = path.join(dir, ASSEMBLED_REPORT_FILENAME);
if (await fs.pathExists(markdownSource)) {
const destination = path.join(runDir, FINAL_REPORT_MD_FILENAME);
await fs.copy(markdownSource, destination, { overwrite: true });
logger.info(`Surfaced markdown report at ${destination}`);
} else {
logger.warn(`Markdown report not found, skipping ${FINAL_REPORT_MD_FILENAME}`);
}
const sarifSource = path.join(dir, SARIF_FILENAME);
+1 -1
View File
@@ -154,7 +154,7 @@ function buildMessageMarkdown(finding: AddFindingInput): string {
parts.push('', '**Remediation**', '', finding.remediation);
// Exploitation steps and proof of impact are deliberately absent: SARIF has no structural home
// for them, and flattening them into prose would imply this file carries the evidence.
parts.push('', 'Full exploitation evidence: `Security-Assessment-Report.md`');
parts.push('', 'Full exploitation evidence: `Security-Assessment-Report.pdf`');
return parts.join('\n');
}
@@ -23,6 +23,7 @@ import type { ActivityLogger } from '../types/activity-logger.js';
import type { AgentEndResult } from '../types/audit.js';
import type { DistributedConfig } from '../types/config.js';
import { ErrorCode } from '../types/errors.js';
import type { AgentMetrics } from '../types/metrics.js';
import { err, ok, type Result } from '../types/result.js';
import { PentestError } from './error-handling.js';
import { loadPrompt } from './prompt-manager.js';
@@ -97,7 +98,9 @@ export interface ValidateAuthInput {
readonly cancellationSignal?: AbortSignal;
}
export async function validateAuthentication(input: ValidateAuthInput): Promise<Result<void, PentestError>> {
export async function validateAuthentication(
input: ValidateAuthInput,
): Promise<Result<AgentMetrics | null, PentestError>> {
const {
distributedConfig,
repoPath,
@@ -113,7 +116,7 @@ export async function validateAuthentication(input: ValidateAuthInput): Promise<
const authentication = distributedConfig.authentication;
if (!authentication) {
return ok(undefined);
return ok(null);
}
logger.info('Validating authentication credentials with live browser...', {
@@ -160,9 +163,10 @@ export async function validateAuthentication(input: ValidateAuthInput): Promise<
}
}
const durationMs = Date.now() - startTime;
const endResult: AgentEndResult = {
attemptNumber,
duration_ms: Date.now() - startTime,
duration_ms: durationMs,
cost_usd: result.cost || 0,
success: classification.ok,
...(result.model !== undefined && { model: result.model }),
@@ -170,7 +174,21 @@ export async function validateAuthentication(input: ValidateAuthInput): Promise<
};
await auditSession.endAgent(AGENT_NAME, endResult);
return classification;
if (!classification.ok) {
return err(classification.error);
}
const metrics: AgentMetrics = {
durationMs,
inputTokens: result.inputTokens ?? null,
outputTokens: result.outputTokens ?? null,
cacheReadTokens: result.cacheReadTokens ?? null,
cacheWriteTokens: result.cacheWriteTokens ?? null,
costUsd: result.cost ?? null,
numTurns: result.turns ?? null,
...(result.model !== undefined && { model: result.model }),
};
return ok(metrics);
}
async function verifySavedAuthState(stateFile: string, logger: ActivityLogger): Promise<Result<void, PentestError>> {
@@ -205,28 +223,32 @@ async function verifySavedAuthState(stateFile: string, logger: ActivityLogger):
);
}
const cookieCount = countStorageEntries(parsed, 'cookies');
const originCount = countStorageEntries(parsed, 'origins');
if (cookieCount === 0 && originCount === 0) {
const cookies = storageEntries(parsed, 'cookies');
const origins = storageEntries(parsed, 'origins');
if (!cookies || !origins) {
return err(
new PentestError(
`Preflight saved an authenticated session to ${stateFile}, but it contains no cookies or origins — the browser was not actually logged in.`,
`Preflight saved an authenticated session to ${stateFile}, but it is not a storage state — cookies and origins arrays are missing.`,
'validation',
true,
{ stateFile, cookieCount, originCount },
{ stateFile, hasCookies: !!cookies, hasOrigins: !!origins },
ErrorCode.AGENT_EXECUTION_FAILED,
),
);
}
logger.info('Preflight authenticated session saved', { stateFile, cookieCount, originCount });
logger.info('Preflight authenticated session saved', {
stateFile,
cookieCount: cookies.length,
originCount: origins.length,
});
return ok(undefined);
}
function countStorageEntries(parsed: unknown, key: 'cookies' | 'origins'): number {
if (typeof parsed !== 'object' || parsed === null) return 0;
function storageEntries(parsed: unknown, key: 'cookies' | 'origins'): unknown[] | null {
if (typeof parsed !== 'object' || parsed === null) return null;
const value = (parsed as Record<string, unknown>)[key];
return Array.isArray(value) ? value.length : 0;
return Array.isArray(value) ? value : null;
}
function classifyResult(
+53 -13
View File
@@ -27,11 +27,13 @@ import type { WorkflowSummary } from '../audit/workflow-logger.js';
import type { CheckpointContext } from '../interfaces/checkpoint-provider.js';
import {
ASSEMBLED_REPORT_FILENAME,
ASSEMBLED_REPORT_PDF_FILENAME,
DEFAULT_DELIVERABLES_SUBDIR,
deliverablesDir,
REPORT_JSON_FILENAME,
resolveSessionJsonPath,
SARIF_FILENAME,
TYPST_TEMPLATE,
} from '../paths.js';
import { getAgentGitPaths } from '../services/agent-git-paths.js';
import { getContainer, getOrCreateContainer, removeContainer } from '../services/container.js';
@@ -448,11 +450,14 @@ export async function runAuthzExploitAgent(input: ActivityInput): Promise<AgentM
}
/**
* Write report.sarif when the run is exploitative and the operator asked for it.
* Write report.sarif for exploitative runs unless the operator opted out with report.sarif: false.
*
* Skipped entirely for analysis-only runs: those findings carry no severity, so every
* `result.level` would be invented. Failures are logged and swallowed — the SARIF log is a
* secondary artifact and must not fail a run whose report is already written.
* On by default so a CI step consuming the log always finds one. Skipped for analysis-only runs.
* The original reason was that those findings carried no severity, so every `result.level` would
* have been invented; since severity is recorded in both modes an analysis run could now populate
* `level`, but it would report an assessed severity as a measured one, so the gate stays. Failures
* are logged and swallowed — the SARIF log is a secondary artifact and must not fail a run whose
* report is already written.
*/
async function writeSarifIfEnabled(
input: ActivityInput,
@@ -465,7 +470,8 @@ async function writeSarifIfEnabled(
const container = getOrCreateContainer(input.workflowId, buildSessionMetadata(input), buildContainerConfig(input));
const configResult = await container.configLoader.loadOptional(input.configPath, undefined, input.configYAML);
if (isErr(configResult) || configResult.value?.report?.sarif !== true) return;
// Only an explicit false opts out; a missing config keeps the default on.
if (isErr(configResult) || configResult.value?.report?.sarif === false) return;
try {
const { renderSarif } = await import('../services/sarif-renderer.js');
@@ -477,6 +483,30 @@ async function writeSarifIfEnabled(
}
}
/**
* Compile the PDF report from the assembled report data.
*
* Failures are logged and swallowed — the PDF is a secondary artifact and must not fail a run
* whose report is already written.
*/
async function writePdfReport(
reportData: ReportData,
deliverablesPath: string,
logger: ReturnType<typeof createActivityLogger>,
): Promise<void> {
try {
const { renderReportPdf } = await import('../services/pdf-renderer.js');
await renderReportPdf({
reportData,
templatePath: TYPST_TEMPLATE,
outputPath: path.join(deliverablesPath, ASSEMBLED_REPORT_PDF_FILENAME),
});
logger.info(`Wrote ${ASSEMBLED_REPORT_PDF_FILENAME}`);
} catch (error) {
logger.warn(`Failed to write ${ASSEMBLED_REPORT_PDF_FILENAME}: ${(error as Error).message}`);
}
}
export async function runReportAgent(input: ActivityInput, exploit: boolean): Promise<AgentMetrics> {
const { createFindingCollector } = await import('../collectors/finding-collector.js');
const { renderReport } = await import('../services/report-renderer.js');
@@ -532,6 +562,7 @@ export async function runReportAgent(input: ActivityInput, exploit: boolean): Pr
await atomicWrite(path.join(deliverablesPath, ASSEMBLED_REPORT_FILENAME), renderReport(reportData));
logger.info(`Wrote ${ASSEMBLED_REPORT_FILENAME} from structured data`);
await writePdfReport(reportData, deliverablesPath, logger);
await writeSarifIfEnabled(input, exploit, reportData, deliverablesPath, logger);
};
@@ -610,7 +641,7 @@ export async function runPreflightValidation(input: ActivityInput): Promise<void
* block; otherwise surfaces a classified failure (failurePoint +
* failureDetail in ApplicationFailure.details) on credential rejection.
*/
export async function runAuthenticationValidation(input: ActivityInput): Promise<void> {
export async function runAuthenticationValidation(input: ActivityInput): Promise<AgentMetrics | null> {
const startTime = Date.now();
const attemptNumber = Context.current().info.attempt;
@@ -628,13 +659,13 @@ export async function runAuthenticationValidation(input: ActivityInput): Promise
if (isErr(configResult)) {
// runPreflightValidation already validated parsing, so this is unexpected.
logger.warn(`runAuthenticationValidation: config load failed unexpectedly: ${configResult.error.message}`);
return;
return null;
}
const distributedConfig = configResult.value;
if (!distributedConfig?.authentication) {
logger.info('No authentication configured — skipping credential validation');
return;
return null;
}
const auditSession = new AuditSession(sessionMetadata);
@@ -673,6 +704,8 @@ export async function runAuthenticationValidation(input: ActivityInput): Promise
truncateStackTrace(failure);
throw failure;
}
return result.value;
} catch (error) {
if (error instanceof ApplicationFailure) {
throw error;
@@ -1111,9 +1144,19 @@ export async function restoreGitCheckpoint(
/**
* Record a resume attempt in session.json and write resume header to workflow.log.
*/
/**
* Register this resume's workflow id in session.json before loadResumeState (which can throw),
* so the CLI can resolve and follow the resume even when validation fails instead of timing out.
*/
export async function registerResumeAttempt(input: ActivityInput, terminatedWorkflows: string[]): Promise<void> {
const sessionMetadata = buildSessionMetadata(input);
const auditSession = new AuditSession(sessionMetadata);
await auditSession.initialize();
await auditSession.addResumeAttempt(input.workflowId, terminatedWorkflows);
}
export async function recordResumeAttempt(
input: ActivityInput,
terminatedWorkflows: string[],
checkpointHash: string,
previousWorkflowId: string,
completedAgents: string[],
@@ -1122,10 +1165,7 @@ export async function recordResumeAttempt(
const auditSession = new AuditSession(sessionMetadata);
await auditSession.initialize();
// Update session.json with resume attempt
await auditSession.addResumeAttempt(input.workflowId, terminatedWorkflows, checkpointHash);
// Write resume header to workflow.log
// session.json entry already added by registerResumeAttempt; here we only write the workflow.log header.
await auditSession.logResumeHeader({
previousWorkflowId,
newWorkflowId: input.workflowId,
-1
View File
@@ -14,5 +14,4 @@ export type {
ResumeState,
VulnExploitPipelineResult,
} from './shared.js';
export { PipelineExecutionError } from './shared.js';
export { pentestPipeline } from './workflows.js';
-15
View File
@@ -60,21 +60,6 @@ export interface PipelineState {
summary: PipelineSummary | null;
}
/**
* Thrown by pentestPipeline() when the run fails, carrying the fully-populated
* PipelineState (real agentMetrics, completedAgents, summary) so a consumer can
* report actual spend instead of synthesizing a zeroed failed state. `cause`
* preserves the original error for classification and Temporal failure reporting.
*/
export class PipelineExecutionError extends Error {
override name = 'PipelineExecutionError' as const;
readonly state: PipelineState;
constructor(message: string, state: PipelineState, options?: { cause?: unknown }) {
super(message, options);
this.state = state;
}
}
// Extended state returned by getProgress query (includes computed fields)
export interface PipelineProgress extends PipelineState {
workflowId: string;
+9 -4
View File
@@ -35,7 +35,12 @@ import { bundleWorkflowCode, NativeConnection, Worker } from '@temporalio/worker
import dotenv from 'dotenv';
import { sanitizeHostname } from '../audit/utils.js';
import { parseConfig } from '../config-parser.js';
import { ASSEMBLED_REPORT_FILENAME, deliverablesDir, FINAL_REPORT_FILENAME, resolveSessionJsonPath } from '../paths.js';
import {
ASSEMBLED_REPORT_PDF_FILENAME,
deliverablesDir,
FINAL_REPORT_PDF_FILENAME,
resolveSessionJsonPath,
} from '../paths.js';
import type { VulnClass } from '../types/config.js';
import { fileExists, readJson } from '../utils/file-io.js';
import * as activities from './activities.js';
@@ -389,9 +394,9 @@ function copyDeliverables(repoPath: string, outputPath: string): void {
}
// Surface the report under its human-facing name alongside the raw deliverables
const assembledReport = path.join(outputDir, ASSEMBLED_REPORT_FILENAME);
if (fs.existsSync(assembledReport)) {
fs.copyFileSync(assembledReport, path.join(outputPath, FINAL_REPORT_FILENAME));
const assembledPdf = path.join(outputDir, ASSEMBLED_REPORT_PDF_FILENAME);
if (fs.existsSync(assembledPdf)) {
fs.copyFileSync(assembledPdf, path.join(outputPath, FINAL_REPORT_PDF_FILENAME));
}
console.log(`Copied ${files.length} deliverable(s) to ${outputPath}`);
+20 -7
View File
@@ -24,6 +24,7 @@
*/
import {
ActivityCancellationType,
ApplicationFailure,
CancellationScope,
isCancellation,
@@ -40,7 +41,6 @@ import type { ActivityInput } from './activities.js';
import {
type AgentMetrics,
getProgress,
PipelineExecutionError,
type PipelineInput,
type PipelineProgress,
type PipelineState,
@@ -96,6 +96,8 @@ const acts = proxyActivities<typeof activities>({
startToCloseTimeout: '2 hours',
heartbeatTimeout: '60 minutes', // Extended for nested pi task execution
retry: PRODUCTION_RETRY,
// Cancel promptly instead of waiting out startToCloseTimeout; the agent aborts on the signal.
cancellationType: ActivityCancellationType.TRY_CANCEL,
});
// Activity proxy with testing retry configuration (fast)
@@ -103,6 +105,7 @@ const testActs = proxyActivities<typeof activities>({
startToCloseTimeout: '30 minutes',
heartbeatTimeout: '30 minutes', // Extended for sub-agent execution in testing
retry: TESTING_RETRY,
cancellationType: ActivityCancellationType.TRY_CANCEL,
});
// Retry configuration for preflight validation (short timeout, few retries)
@@ -119,6 +122,7 @@ const preflightActs = proxyActivities<typeof activities>({
startToCloseTimeout: '2 minutes',
heartbeatTimeout: '2 minutes',
retry: PREFLIGHT_RETRY,
cancellationType: ActivityCancellationType.TRY_CANCEL,
});
// Credential rejection is not retryable; transient provider errors get 3 attempts.
@@ -135,6 +139,7 @@ const authValidationActs = proxyActivities<typeof activities>({
startToCloseTimeout: '10 minutes',
heartbeatTimeout: '10 minutes',
retry: AUTH_VALIDATION_RETRY,
cancellationType: ActivityCancellationType.TRY_CANCEL,
});
/**
@@ -246,6 +251,10 @@ export async function pentestPipeline(input: PipelineInput): Promise<PipelineSta
let resumeState: ResumeState | null = null;
if (input.resumeFromWorkspace) {
// 0. Register the resume's workflow id in session.json before validation can fail, so the CLI
// can resolve and follow it instead of polling for an entry that never lands.
await a.registerResumeAttempt(activityInput, input.terminatedWorkflows || []);
// 1. Load resume state (validates workspace, cross-checks deliverables)
resumeState = await a.loadResumeState(
input.resumeFromWorkspace,
@@ -277,10 +286,9 @@ export async function pentestPipeline(input: PipelineInput): Promise<PipelineSta
return state;
}
// 4. Record this resume attempt in session.json and workflow.log
// 4. Write the resume header to workflow.log (the session.json entry was recorded in step 0)
await a.recordResumeAttempt(
activityInput,
input.terminatedWorkflows || [],
resumeState.checkpointHash,
resumeState.originalWorkflowId,
resumeState.completedAgents,
@@ -480,7 +488,11 @@ export async function pentestPipeline(input: PipelineInput): Promise<PipelineSta
// === Authentication Validation ===
state.currentPhase = 'auth-validation';
state.currentAgent = 'validate-authentication';
await authValidationActs.runAuthenticationValidation(activityInput);
const authMetrics = await authValidationActs.runAuthenticationValidation(activityInput);
// Null when no login ran (no-auth scan); left absent so status renders it skipped, not completed.
if (authMetrics) {
state.agentMetrics['validate-authentication'] = authMetrics;
}
state.currentAgent = null;
log.info('Authentication validation passed');
@@ -705,9 +717,10 @@ export async function pentestPipeline(input: PipelineInput): Promise<PipelineSta
});
}
// Carry the populated state so a consumer can report real spend instead of a zeroed
// failed state. The original error rides as `cause` for classification/reporting.
throw new PipelineExecutionError(state.error ?? 'Pipeline failed', state, { cause: error });
// Terminate the workflow in Temporal's FAILED state. WARNING: this must be an
// ApplicationFailure — any other thrown type becomes an unhandled workflow-task failure
// that Temporal retries indefinitely, leaving the run stuck in RUNNING.
throw ApplicationFailure.nonRetryable(state.error ?? 'Pipeline failed', 'PipelineExecutionError');
}
}
-174
View File
@@ -1,174 +0,0 @@
#!/usr/bin/env node
// Copyright (C) 2025 Keygraph, Inc.
//
// This program is free software: you can redistribute it and/or modify
// it under the terms of the GNU Affero General Public License version 3
// as published by the Free Software Foundation.
/**
* Workspace listing tool for Shannon.
*
* Reads workspaces/ directories, parses session.json files, and displays
* a formatted table of all workspaces with status, duration, and cost.
*
* Usage:
* node dist/temporal/workspaces.js
*
* Environment:
* WORKSPACES_DIR - Override workspaces directory (default: ./workspaces)
*/
import fs from 'node:fs/promises';
import path from 'node:path';
import { WORKSPACES_DIR as DEFAULT_WORKSPACES_DIR, resolveSessionJsonPath } from '../paths.js';
interface SessionJson {
session: {
id: string;
webUrl: string;
status: 'in-progress' | 'completed' | 'failed';
createdAt: string;
completedAt?: string;
};
metrics: {
total_cost_usd: number;
};
}
interface WorkspaceInfo {
name: string;
url: string;
status: 'in-progress' | 'completed' | 'failed';
createdAt: Date;
completedAt: Date | null;
costUsd: number;
}
function formatDuration(ms: number): string {
const seconds = Math.floor(ms / 1000);
const minutes = Math.floor(seconds / 60);
const hours = Math.floor(minutes / 60);
if (hours > 0) {
return `${hours}h ${minutes % 60}m`;
}
if (minutes > 0) {
return `${minutes}m`;
}
return `${seconds}s`;
}
function getStatusDisplay(status: string): string {
return status;
}
function truncate(str: string, maxLen: number): string {
if (str.length <= maxLen) return str;
return `${str.slice(0, maxLen - 1)}\u2026`;
}
async function listWorkspaces(): Promise<void> {
const workspacesDir = process.env.WORKSPACES_DIR || DEFAULT_WORKSPACES_DIR;
let entries: string[];
try {
entries = await fs.readdir(workspacesDir);
} catch {
console.log('No workspaces directory found.');
console.log(`Expected: ${workspacesDir}`);
return;
}
const workspaces: WorkspaceInfo[] = [];
for (const entry of entries) {
const sessionPath = resolveSessionJsonPath(path.join(workspacesDir, entry));
try {
const content = await fs.readFile(sessionPath, 'utf8');
const data = JSON.parse(content) as SessionJson;
workspaces.push({
name: entry,
url: data.session.webUrl,
status: data.session.status,
createdAt: new Date(data.session.createdAt),
completedAt: data.session.completedAt ? new Date(data.session.completedAt) : null,
costUsd: data.metrics.total_cost_usd,
});
} catch {
// Skip directories without valid session.json
}
}
if (workspaces.length === 0) {
console.log('\nNo workspaces found.');
console.log('Run a pipeline first: ./shannon start -u <url> -r <repo>');
return;
}
// Sort by creation date (most recent first)
workspaces.sort((a, b) => b.createdAt.getTime() - a.createdAt.getTime());
console.log('\n=== Shannon Workspaces ===\n');
// Column widths
const nameWidth = 30;
const urlWidth = 30;
const statusWidth = 14;
const durationWidth = 10;
const costWidth = 10;
// Header
console.log(
' ' +
'WORKSPACE'.padEnd(nameWidth) +
'URL'.padEnd(urlWidth) +
'STATUS'.padEnd(statusWidth) +
'DURATION'.padEnd(durationWidth) +
'COST'.padEnd(costWidth),
);
console.log(` ${'\u2500'.repeat(nameWidth + urlWidth + statusWidth + durationWidth + costWidth)}`);
let resumableCount = 0;
for (const ws of workspaces) {
const now = new Date();
const endTime = ws.completedAt || now;
const durationMs = endTime.getTime() - ws.createdAt.getTime();
const duration = formatDuration(durationMs);
const cost = `$${ws.costUsd.toFixed(2)}`;
const isResumable = ws.status !== 'completed';
if (isResumable) {
resumableCount++;
}
const resumeTag = isResumable ? ' (resumable)' : '';
console.log(
' ' +
truncate(ws.name, nameWidth - 2).padEnd(nameWidth) +
truncate(ws.url, urlWidth - 2).padEnd(urlWidth) +
getStatusDisplay(ws.status).padEnd(statusWidth) +
duration.padEnd(durationWidth) +
cost.padEnd(costWidth) +
resumeTag,
);
}
console.log();
const summary = `${workspaces.length} workspace${workspaces.length === 1 ? '' : 's'} found`;
const resumeSummary = resumableCount > 0 ? ` (${resumableCount} resumable)` : '';
console.log(`${summary}${resumeSummary}`);
if (resumableCount > 0) {
console.log('\nResume with: ./shannon start -u <url> -r <repo> -w <name>');
}
console.log();
}
listWorkspaces().catch((err) => {
console.error('Error listing workspaces:', err);
process.exit(1);
});
+4 -1
View File
@@ -32,7 +32,10 @@ export interface ReportConfig {
min_severity?: Severity;
min_confidence?: Confidence;
guidance?: string;
/** Emit report.sarif alongside the markdown report. Ignored when exploit is false. */
/**
* Emit report.sarif alongside the markdown report. On by default for exploit runs; set 'false'
* to opt out. Ignored when exploit is false.
*/
sarif?: 'true' | 'false';
}
Binary file not shown.

After

Width:  |  Height:  |  Size: 62 KiB

+535
View File
@@ -0,0 +1,535 @@
// =============================================================================
// Security Assessment Report — Typst template
// Invoke:
// typst compile --root <root> --input data=/data.json report.typ out.pdf
// Optional overrides:
// --input tester=<name> --input brand=<name>
// =============================================================================
#let data = json(sys.inputs.data)
// Top-level discriminator. Schema variants in report-output-schema.ts:
// exploits → ExploitsReportData (exploit=true runs, full reproduction)
// findings → FindingsReportData (exploit=false runs, analysis-only)
#let mode = data.at("mode", default: "exploits")
#let tester-override = sys.inputs.at("tester", default: "Shannon")
#let brand = sys.inputs.at("brand", default: "Shannon | AI Pentester by Keygraph")
// ---------- Palette ---------------------------------------------------------
// Kept distinct so Critical / High are not confused under monitor gamma.
#let sev-color(level) = {
if level == "Critical" { rgb("#DC2626") } // red-600
else if level == "High" { rgb("#EA580C") } // orange-600
else if level == "Medium" { rgb("#D97706") } // amber-600
else if level == "Low" { rgb("#2563EB") } // blue-600
else { rgb("#6B7280") }
}
#let confidence-color(c) = {
if c == "High" { rgb("#15803D") } // green-700
else if c == "Medium" { rgb("#D97706") } // amber-600
else if c == "Low" { rgb("#6B7280") } // gray-500
else { rgb("#6B7280") }
}
// Warm, editorial, high-contrast document palette.
#let ink = rgb("#141414") // warm near-black text
#let muted = rgb("#5C5850") // warm gray-brown labels
#let tertiary = rgb("#9A958D") // lightest muted
#let rule = rgb("#E6E1D9") // warm hair rules
#let rule-soft = rgb("#D9D3CA")
#let code-bg = rgb("#F6F1EB") // warm eggshell
#let alt-bg = rgb("#EBE6DF")
#let page-bg = white
// ---------- Page setup ------------------------------------------------------
#set document(title: "Security Assessment Report", author: brand)
#set page(
paper: "a4",
margin: (top: 2.2cm, bottom: 2.2cm, left: 2.2cm, right: 2.2cm),
fill: page-bg,
header: context {
if counter(page).get().first() > 1 [
#set text(size: 8.5pt, fill: muted)
#grid(columns: (1fr, auto),
[Security Assessment Report],
[CONFIDENTIAL],
)
#v(-4pt)
#line(length: 100%, stroke: 0.3pt + rule)
]
},
footer: context {
if counter(page).get().first() > 1 [
#set text(size: 8.5pt, fill: muted)
#line(length: 100%, stroke: 0.3pt + rule)
#v(2pt)
#grid(columns: (1fr, auto),
[#data.meta.assessmentDate],
[#counter(page).display() / #context counter(page).final().first()],
)
]
},
)
#set text(size: 10.5pt, fill: ink)
#set par(leading: 0.7em, justify: false)
#show heading.where(level: 1): it => [
#pagebreak(weak: true)
#v(4pt)
#set text(size: 24pt, weight: "bold", fill: ink)
#it.body
#v(4pt)
#line(length: 100%, stroke: 0.4pt + rule)
#v(10pt)
]
#show heading.where(level: 2): it => [
#v(10pt)
#set text(size: 14pt, weight: "semibold", fill: ink)
#it.body
#v(2pt)
]
#show heading.where(level: 3): it => [
#v(8pt)
#set text(size: 11.5pt, weight: "semibold", fill: ink)
#it.body
#v(-2pt)
]
#show raw: set text(size: 8.5pt)
#show raw.where(block: false): it => box(
fill: code-bg,
inset: (x: 3pt, y: 0pt),
outset: (y: 2pt),
radius: 2pt,
it,
)
#show raw.where(block: true): it => block(
fill: code-bg,
stroke: (left: 2pt + rule, rest: none),
inset: (x: 10pt, y: 8pt),
width: 100%,
breakable: true,
{
set par(leading: 0.5em, justify: false)
it
},
)
// ---------- Helpers ---------------------------------------------------------
#let chip(label, color) = box(
fill: color,
inset: (x: 6pt, y: 2pt),
radius: 2pt,
text(fill: white, weight: "bold", size: 7.5pt, tracking: 0.3pt, upper(label)),
)
#let categories-in-order = (
"Authentication",
"Authorization",
"XSS",
"Injection",
"SSRF",
"Other",
)
#let sev-chip(level) = chip(level, sev-color(level))
#let confidence-chip(c) = chip(c + " confidence", confidence-color(c))
// inline-code renders a string, turning backtick-wrapped spans into
// inline raw. Safe on odd counts — a trailing unclosed backtick is
// emitted as literal text so nothing gets swallowed.
#let inline-code(s) = {
if type(s) != str { return s }
let parts = s.split("`")
if parts.len() == 1 { return parts.at(0) }
let out = []
for (i, p) in parts.enumerate() {
if calc.even(i) {
out += [#p]
} else if i == parts.len() - 1 {
out += [#("`" + p)]
} else {
out += raw(p)
}
}
out
}
#let render-items(items) = {
for item in items {
if item.kind == "prose" [
#par(inline-code(item.text))
] else if item.kind == "code" [
#raw(item.block.content, lang: item.block.language, block: true)
]
}
}
// Render step items as a bulleted list; prose items become bullets,
// code items break the list and render as code blocks in between.
#let render-bulleted-items(items) = {
for item in items {
if item.kind == "prose" [
- #inline-code(item.text)
] else if item.kind == "code" [
#raw(item.block.content, lang: item.block.language, block: true)
]
}
}
// Render step items as a numbered list; prose items become enumerated,
// code items break the list and render as code blocks in between.
#let render-numbered-items(items) = {
for item in items {
if item.kind == "prose" [
+ #inline-code(item.text)
] else if item.kind == "code" [
#raw(item.block.content, lang: item.block.language, block: true)
]
}
}
// Render an array of strings as a bulleted list with inline-code support.
#let code-list(items) = list(..items.map(inline-code))
#let kv(label, value) = grid(
columns: (auto, 1fr),
column-gutter: 14pt,
row-gutter: 4pt,
text(fill: muted, size: 9.5pt)[#label],
value,
)
// ---------- COVER PAGE ------------------------------------------------------
#page(header: none, footer: none)[
#set align(left)
#v(3.2cm)
#let brand-parts = brand.split("|").map(p => p.trim())
#grid(
columns: (auto, 1fr),
column-gutter: 8pt,
align: (horizon, horizon),
image("/assets/keygraph-logo.png", width: 1.6cm),
{
set par(leading: 0.6em)
text(size: 11pt, fill: ink, weight: "semibold", tracking: 1.2pt)[
#upper(brand-parts.at(0))
]
if brand-parts.len() > 1 {
linebreak()
text(size: 9pt, fill: muted, weight: "regular")[
#brand-parts.slice(1).join(" ")
]
}
}
)
#set par(leading: 0.7em)
#v(1.6cm)
#set par(leading: 0.4em)
#text(size: 46pt, weight: "bold", fill: ink)[
Security\
Assessment\
Report
]
#set par(leading: 0.7em)
#v(1fr)
#line(length: 100%, stroke: 0.3pt + rule)
#v(0.6cm)
#grid(
columns: (1fr, 1fr),
column-gutter: 28pt,
row-gutter: 14pt,
grid(
columns: (auto, 1fr),
column-gutter: 18pt,
row-gutter: 14pt,
text(fill: muted, size: 9pt)[Target], text(size: 10pt)[#inline-code(data.meta.target)],
text(fill: muted, size: 9pt)[Date], text(size: 10pt)[#data.meta.assessmentDate],
..(if "application" in data.meta and data.meta.application != none {
(text(fill: muted, size: 9pt)[Application], text(size: 10pt)[#inline-code(data.meta.application)])
} else { () }),
),
grid(
columns: (auto, 1fr),
column-gutter: 18pt,
row-gutter: 14pt,
text(fill: muted, size: 9pt)[Tester], text(size: 10pt)[#tester-override],
text(fill: muted, size: 9pt)[Classification],
text(size: 10pt, weight: "semibold")[#data.meta.classification],
),
)
#v(0.8cm)
#text(size: 8pt, fill: muted)[
This document contains sensitive security findings.
Handle in accordance with your organization's data classification policy.
]
]
// ---------- TABLE OF CONTENTS -----------------------------------------------
#outline(title: [Contents], depth: 3, indent: auto)
// ---------- EXECUTIVE SUMMARY -----------------------------------------------
= Executive Summary
#grid(
columns: (auto, 1fr),
column-gutter: 20pt,
row-gutter: 12pt,
text(fill: muted, size: 10pt)[Target], text(size: 10.5pt)[#inline-code(data.meta.target)],
text(fill: muted, size: 10pt)[Date], text(size: 10.5pt)[#data.meta.assessmentDate],
..(if "application" in data.meta and data.meta.application != none {
(text(fill: muted, size: 10pt)[Application], text(size: 10.5pt)[#inline-code(data.meta.application)])
} else { () }),
text(fill: muted, size: 10pt)[Tester], text(size: 10.5pt)[#tester-override],
)
== Scope
#inline-code(data.scope)
// ---------- BY TYPE ---------------------------------------------------------
#let by-type-entries = if mode == "exploits" { data.exploitedByType } else { data.identifiedByType }
#if mode == "exploits" [
= Successfully Exploited Vulnerabilities by Type
] else [
= Identified Vulnerabilities by Type
]
#for entry in by-type-entries [
== #entry.category
#if "narrative" in entry and entry.narrative != none [
#inline-code(entry.narrative)
]
#if "bullets" in entry and entry.bullets != none [
#list(
..entry.bullets.map(b => [
#text(weight: "semibold")[#b.id] — #inline-code(b.description)
])
)
]
]
// ---------- SUMMARY ---------------------------------------------------------
= Summary
#let s = data.summary
#let sev = data.derivedCounts.bySeverity
#let severity-card(label, sev-key, n) = box(
fill: sev-color(sev-key),
inset: (x: 8pt, y: 12pt),
radius: 4pt,
width: 100%,
stack(
dir: ttb,
spacing: 6pt,
text(fill: white, weight: "bold", size: 20pt)[#n],
text(fill: white, size: 8pt, tracking: 0.5pt)[#upper(label)],
),
)
#grid(
columns: 4,
column-gutter: 8pt,
severity-card("Critical", "Critical", sev.Critical),
severity-card("High", "High", sev.High),
severity-card("Medium", "Medium", sev.Medium),
severity-card("Low", "Low", sev.Low),
)
#if mode == "findings" [
#v(18pt)
#let cf = data.derivedCounts.byConfidence
#let confidence-card(label, c-key, n) = box(
stroke: 0.6pt + confidence-color(c-key),
inset: (x: 8pt, y: 12pt),
radius: 4pt,
width: 100%,
stack(
dir: ttb,
spacing: 6pt,
text(fill: ink, weight: "bold", size: 20pt)[#n],
text(fill: confidence-color(c-key), size: 8pt, tracking: 0.5pt)[#upper(label + " confidence")],
),
)
#grid(
columns: 3,
column-gutter: 8pt,
confidence-card("High", "High", cf.High),
confidence-card("Medium", "Medium", cf.Medium),
confidence-card("Low", "Low", cf.Low),
)
]
#v(14pt)
#if mode == "exploits" [
#grid(
columns: (auto, 1fr),
column-gutter: 14pt,
row-gutter: 4pt,
text(fill: muted, size: 10pt)[Total identified],
text(weight: "semibold")[#s.totalIdentified],
text(fill: muted, size: 10pt)[Successfully exploited],
text(weight: "semibold")[#s.successfullyExploited],
)
] else [
#grid(
columns: (auto, 1fr),
column-gutter: 14pt,
row-gutter: 4pt,
text(fill: muted, size: 10pt)[Total identified],
text(weight: "semibold")[#s.totalIdentified],
)
]
#v(8pt)
#let breakdown = if mode == "exploits" { s.exploitedBreakdown } else { s.identifiedBreakdown }
#list(
..breakdown.map(c => [
#text(weight: "semibold")[#c.count] #c.category#if "note" in c and c.note != none [ — #inline-code(c.note)]
])
)
#if mode == "exploits" [
#if "outOfScope" in s and s.outOfScope != none [
#v(4pt)
#text(weight: "semibold")[Out of Scope#if "note" in s.outOfScope and s.outOfScope.note != none [ (#s.outOfScope.note)]:] #s.outOfScope.total vulnerabilities
#if "breakdown" in s.outOfScope and s.outOfScope.breakdown != none [
#list(
..s.outOfScope.breakdown.map(c => [
#text(weight: "semibold")[#c.count] #c.category#if "note" in c and c.note != none [ — #inline-code(c.note)]
])
)
]
]
#if "blockedByConstraints" in s and s.blockedByConstraints != none [
#v(4pt)
#text(weight: "semibold")[Blocked by Testing Constraints:] #s.blockedByConstraints.total#if "note" in s.blockedByConstraints and s.blockedByConstraints.note != none [ — #s.blockedByConstraints.note]
]
]
== Critical Findings
#enum(..s.criticalFindings.map(f => [#inline-code(f)]))
// ---------- FINDINGS OVERVIEW -----------------------------------------------
= Findings Overview
#let show-confidence-col = mode == "findings"
#table(
columns: if show-confidence-col { (auto, 1fr, auto, auto, auto) } else { (auto, 1fr, auto, auto) },
stroke: none,
inset: (x: 8pt, y: 7pt),
align: if show-confidence-col { (left, left, left, center, center) } else { (left, left, left, center) },
fill: (_, row) => if row == 0 { none } else if calc.even(row) { code-bg } else { none },
table.header(
text(size: 9.5pt, weight: "semibold")[ID],
text(size: 9.5pt, weight: "semibold")[Title],
text(size: 9.5pt, weight: "semibold")[Category],
text(size: 9.5pt, weight: "semibold")[Severity],
..(if show-confidence-col { (text(size: 9.5pt, weight: "semibold")[Confidence],) } else { () }),
),
..data.findings.map(f => (
text(weight: "semibold")[#f.id],
inline-code(f.title),
text(size: 9.5pt)[#f.category],
sev-chip(f.severity),
..(if show-confidence-col { (confidence-chip(f.confidence),) } else { () }),
)).flatten()
)
// ---------- FINDING RENDER --------------------------------------------------
#let render-finding-summary(f) = [
#v(8pt)
#grid(
columns: (auto, 1fr),
column-gutter: 18pt,
row-gutter: 12pt,
text(fill: muted, size: 9.5pt)[Location], text(size: 10pt)[#inline-code(f.summary.vulnerableLocation)],
text(fill: muted, size: 9.5pt)[Overview], text(size: 10pt)[#inline-code(f.summary.overview)],
text(fill: muted, size: 9.5pt)[Impact], text(size: 10pt)[#inline-code(f.summary.impact)],
)
]
#let render-finding-extras(f) = [
#if "notes" in f and f.notes != none and f.notes.len() > 0 [
#heading(level: 3, outlined: false)[Notes]
#render-bulleted-items(f.notes)
]
#if "additionalSections" in f and f.additionalSections != none [
#for extra in f.additionalSections [
#heading(level: 3, outlined: false)[#inline-code(extra.heading)]
#render-items(extra.items)
]
]
]
#let render-exploit(f) = [
== #f.id: #inline-code(f.title)
#sev-chip(f.severity)
#render-finding-summary(f)
=== Prerequisites
#inline-code(f.prerequisites)
=== Exploitation Steps
#for step in f.exploitationSteps [
#text(weight: "semibold")[Step #step.number#if "title" in step and step.title != none [ — #inline-code(step.title)]]
#render-items(step.items)
]
=== Proof of Impact
#render-numbered-items(f.proofOfImpact)
#render-finding-extras(f)
#v(16pt)
]
#let render-analysis(f) = [
== #f.id: #inline-code(f.title)
#sev-chip(f.severity) #h(4pt) #confidence-chip(f.confidence)
#render-finding-summary(f)
#render-finding-extras(f)
#v(16pt)
]
#let render-finding(f) = if mode == "exploits" { render-exploit(f) } else { render-analysis(f) }
// ---------- PER-CATEGORY -----------------------------------------------------
#let category-section-label(n) = if mode == "exploits" {
"Exploitation Evidence"
} else {
"Findings"
}
#for cat in categories-in-order {
let cat-findings = data.findings.filter(f => f.category == cat)
if cat-findings.len() > 0 [
= #cat #category-section-label(cat-findings.len()) (#cat-findings.len() #if cat-findings.len() == 1 [finding] else [findings])
#for f in cat-findings {
render-finding(f)
}
]
}
Binary file not shown.

Before

Width:  |  Height:  |  Size: 79 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 25 MiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 50 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 49 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 50 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 75 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 80 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 91 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 39 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 40 KiB

+65 -4
View File
@@ -25,7 +25,7 @@ Shannon forwards only the selected provider's credential into the scan container
### Any other provider
Shannon accepts any provider and model present in the Pi harness catalogue. Browse them at [pi.dev/models](https://pi.dev/models). These are technically supported but not recommended. Claude models are best-supported (see the note below).
Shannon accepts any provider and model present in the Pi harness catalogue. Browse them at [pi.dev/models](https://pi.dev/models).
```bash
export SHANNON_AI_API_KEY=your-api-key # the provider's API key
@@ -37,7 +37,7 @@ This path covers providers whose credential is a single API key. Providers that
`npx @keygraph/shannon setup` exposes this as the **Other provider** option.
> [!IMPORTANT]
> Claude models are the best-supported option. Shannon's evaluations, internal testing, and agent harness are tuned for Claude. Other models are permitted and validated against the harness catalogue, but may not follow Shannon's instructions or tool-use constraints as reliably. Use them at your own risk.
> Models are validated against the harness catalogue, but capability varies. A model that does not follow Shannon's instructions or tool-use constraints reliably will produce weaker pentests. Evaluate the model you choose against your own targets before depending on its results.
## Cyber safeguards (do this before your first scan)
@@ -58,7 +58,7 @@ These are the models `npx @keygraph/shannon setup` offers, best-first. They are
| --- | --- |
| `anthropic` | `claude-sonnet-4-6`, `claude-opus-4-8`, `claude-opus-4-7`, `claude-haiku-4-5-20251001` |
| `openai` | `gpt-5.6-sol`, `gpt-5.5`, `gpt-5.4` |
| `xai` | `grok-4.5` |
| `xai` | `grok-4.6`, `grok-4.5` |
| `amazon-bedrock` | `us.anthropic.claude-sonnet-4-6`, `us.anthropic.claude-opus-4-8`, `us.anthropic.claude-opus-4-7` |
Bedrock IDs are region-prefixed and must be enabled in your account, so the ID that works for you may differ from the one listed here.
@@ -144,12 +144,73 @@ The variable is rejected in preflight where it cannot take effect: with a non-`o
`npx @keygraph/shannon setup` covers this under **Custom Base URL**, which asks which API your gateway serves and configures the matching provider for you.
## OpenAI Codex (ChatGPT Plus/Pro subscription)
A ChatGPT Plus or Pro Codex subscription can run Shannon. Shannon reuses a login created by Pi.
Before running a pentest, review the [cyber safeguards requirements](#cyber-safeguards-do-this-before-your-first-scan).
1. Install Pi by following the instructions at [pi.dev](https://pi.dev).
2. Log in with your subscription using Pi's [subscription authentication guide](https://pi.dev/docs/latest/providers#subscriptions). This creates `~/.pi/agent/auth.json` with an `openai-codex` entry.
3. Select a Codex model and enable Pi authentication:
```bash
export SHANNON_USE_PI_AUTH=1
export SHANNON_AI_MODEL=openai-codex:gpt-5.5
```
4. In npx mode, run `npx @keygraph/shannon start ...` from the same shell. In source-build mode, add the two variables to `.env` and run `./shannon start ...`.
Supported Codex models are `gpt-5.6-sol`, `gpt-5.5`, and `gpt-5.4`.
## xAI (Grok subscription)
An xAI subscription can run Shannon. Shannon reuses a login created by Pi.
1. Install Pi by following the instructions at [pi.dev](https://pi.dev).
2. Log in with your subscription using Pi's [subscription authentication guide](https://pi.dev/docs/latest/providers#subscriptions). This creates `~/.pi/agent/auth.json` with an `xai` entry.
3. Select an xAI model and enable Pi authentication:
```bash
export SHANNON_USE_PI_AUTH=1
export SHANNON_AI_MODEL=xai:grok-4.6
```
4. In npx mode, run `npx @keygraph/shannon start ...` from the same shell. In source-build mode, add the two variables to `.env` and run `./shannon start ...`.
Suggested Grok models are `grok-4.6` and `grok-4.5`.
## Claude Code subscription
The latest version of Shannon does not support Claude Code subscriptions. The [`shannon-v1`](https://github.com/KeygraphHQ/shannon/tree/shannon-v1) branch is the final release built on the Claude Agent SDK and supports Claude Code OAuth.
Before running a pentest, review the [cyber safeguards requirements](#cyber-safeguards-do-this-before-your-first-scan).
1. Generate a Claude Code OAuth token:
```bash
claude setup-token
```
2. Run the setup flow for the final `shannon-v1` release:
```bash
npx @keygraph/shannon@1.9.0 setup
```
3. Select **OAuth Token** and enter the token generated by Claude Code.
4. Start the pentest with `npx @keygraph/shannon@1.9.0 start ...`.
These instructions apply only to `shannon-v1`.
## Validation
Checks run before a scan starts, so mistakes fail immediately rather than partway through a run:
- **Provider and model ID** — validated against the Pi harness catalogue. An unknown provider or model ID fails preflight with a pointer to [pi.dev/models](https://pi.dev/models). A custom base URL exempts the model ID, since a gateway may serve its own names.
- **Credential presence** — always validated for the selected provider.
- **Credential presence** — validated for the selected provider, or read from Pi when `SHANNON_USE_PI_AUTH=1`.
- **Credential validity** — one minimal request against the model the scan will use, so a rejected key, an exhausted quota, or a model the account cannot reach fails before any agent runs. Bedrock included: its bearer token and region go through the same probe.
## Migrating from the three-tier configuration
+7 -8
View File
@@ -99,33 +99,32 @@ rules:
# min_confidence: low
# guidance: |
# Drop findings about missing security headers and rate-limit gaps.
# sarif: "true"
# sarif: "false"
```
## Report Options
| Key | Effect |
| --- | --- |
| `min_severity` | Drops findings rated below this severity. Applies only when `exploit` is `"true"`. |
| `min_severity` | Drops findings rated below this severity. Applies in both exploitative and analysis-only runs. |
| `min_confidence` | Drops findings rated below this confidence. Applies only when `exploit` is `"false"`. |
| `guidance` | Free-text instruction to the report agent, such as which topics to exclude. |
| `sarif` | Emits a SARIF 2.1.0 log alongside the Markdown report. Requires `exploit: "true"`. |
| `sarif` | SARIF 2.1.0 log alongside the Markdown report. On by default for exploit runs; set `"false"` to opt out. Ignored when `exploit` is `"false"`. |
A finding carries one rating or the other, never both: an exploited finding is rated by severity, an analysis-only finding by confidence. Setting the threshold that does not apply to the run is ignored, and Shannon logs a warning naming the one to use instead.
Every finding carries a severity, but it does not mean the same thing in each mode: an exploitative run measures severity from what the exploit demonstrated, while an analysis-only run assesses it from the class of flaw and the impact it would have. An analysis-only finding carries a confidence rating alongside its severity, since nothing was proven. Setting `min_confidence` on an exploitative run is ignored, and Shannon logs a warning naming the threshold to use instead.
### SARIF Output
Set `sarif: "true"` to write `report.sarif` next to `Security-Assessment-Report.md` at the workspace root, for upload to GitHub code scanning or any other SARIF consumer.
On exploit-mode runs Shannon writes `report.sarif` next to `Security-Assessment-Report.pdf` at the workspace root by default, for upload to GitHub code scanning or any other SARIF consumer. No configuration is needed; set `sarif: "false"` to opt out.
```yaml
exploit: "true"
report:
sarif: "true"
sarif: "false"
```
Each finding becomes one SARIF result, filed under a rule per vulnerability class (`shannon/injection`, `shannon/xss`, `shannon/auth`, `shannon/authz`, `shannon/ssrf`) and tagged with its OWASP Top Ten 2025 category. Results are anchored to the code location the analysis phase recorded, falling back to the HTTP entry point when the finding names no file. Severity maps onto SARIF's three levels: `critical` and `high` become `error`, `medium` becomes `warning`, everything else becomes `note`.
The log is written only for exploitative runs. An analysis-only run rates findings by confidence and produces no severity, so there is nothing to populate `level` with; `sarif` is ignored when `exploit` is `"false"`.
The log is written only for exploitative runs. `sarif` is ignored when `exploit` is `"false"`.
Supported rule types include `url_path`, `subdomain`, `domain`, `method`, `header`, `parameter`, and `code_path`.
+20 -12
View File
@@ -58,7 +58,8 @@ Monitor progress:
```bash
npx @keygraph/shannon logs <workspace>
npx @keygraph/shannon status
npx @keygraph/shannon status <workspace>
npx @keygraph/shannon scans
npx @keygraph/shannon version
```
@@ -66,7 +67,8 @@ Source-build equivalents:
```bash
./shannon logs <workspace>
./shannon status
./shannon status <workspace>
./shannon scans
./shannon version
```
@@ -79,16 +81,17 @@ open http://localhost:8233
Stop Shannon:
```bash
npx @keygraph/shannon stop
npx @keygraph/shannon stop --clean # confirms first; add --yes (or -y) to skip
npx @keygraph/shannon uninstall # confirms first; add --yes (or -y) to skip
npx @keygraph/shannon stop <workspace> # stop one scan (confirms first; add --yes/-y to skip)
npx @keygraph/shannon stop --all # stop all scans (Temporal stays up)
npx @keygraph/shannon reset # stop everything and wipe all Temporal data (type 'confirm' to proceed; cannot be skipped)
```
Source-build equivalents:
```bash
./shannon stop
./shannon stop --clean # add --yes (or -y) to skip the confirmation
./shannon stop <workspace> # stop one scan (confirms first; add --yes/-y to skip)
./shannon stop --all # stop all scans (Temporal stays up)
./shannon reset # stop everything and wipe all Temporal data (type 'confirm' to proceed; cannot be skipped)
```
Usage examples:
@@ -106,8 +109,11 @@ npx @keygraph/shannon start -u https://example.com -r /path/to/repo -o ./my-repo
# Named workspace.
npx @keygraph/shannon start -u https://example.com -r /path/to/repo -w q1-audit
# List all workspaces.
npx @keygraph/shannon workspaces
# Stream the log until the scan finishes, then exit on its outcome (useful in CI).
npx @keygraph/shannon start -u https://example.com -r /path/to/repo --follow
# List completed scans.
npx @keygraph/shannon scans
```
Source-build examples:
@@ -117,7 +123,8 @@ Source-build examples:
./shannon start -u https://example.com -r /path/to/repo -c /path/to/my-config.yaml
./shannon start -u https://example.com -r /path/to/repo -o ./my-reports
./shannon start -u https://example.com -r /path/to/repo -w q1-audit
./shannon workspaces
./shannon start -u https://example.com -r /path/to/repo --follow
./shannon scans
# Rebuild the worker image.
./shannon build --no-cache
@@ -132,11 +139,12 @@ Results are saved to the workspaces directory:
Use `-o <path>` to copy deliverables to a custom output directory after a run completes.
Output structure — the run directory's top level holds only the final report; everything else is nested under a hidden `.shannon/` directory:
Output structure — the run directory's top level holds the final report, in PDF and Markdown; everything else is nested under a hidden `.shannon/` directory:
```text
workspaces/{hostname}_{sessionId}/
|-- Security-Assessment-Report.md # the final report (the deliverable)
|-- Security-Assessment-Report.pdf # the final report (PDF)
|-- Security-Assessment-Report.md # the final report (Markdown)
`-- .shannon/ # internals
|-- deliverables/ # report source, per-phase analysis, queues
|-- agents/ # per-agent logs
+1 -1
View File
@@ -4,7 +4,7 @@ This guide covers platform-specific notes and Docker networking behavior.
## Windows
Shannon on Windows is supported through WSL2. Native Windows, including Git Bash, is not supported.
Shannon on Windows is supported through WSL2, which behaves like Linux for everything below. Native Windows, including Git Bash, is community-supported: contributions are welcome, but Keygraph does not actively develop or test against it.
### Ensure WSL2
+1 -1
View File
@@ -28,7 +28,7 @@ For maximum isolation, run Shannon inside a disposable virtual machine.
## LLM and Automation Caveats
- **Verification is required**: Shannon uses a proof-by-exploitation methodology, but final reports can still contain weakly supported or incorrect details. Human review is essential.
- **Model support**: Shannon is officially supported only with Claude models. Alternative models may be incomplete, inaccurate, or unstable.
- **Model support**: results vary by model. A model that does not follow Shannon's instructions or tool-use constraints reliably may produce incomplete, inaccurate, or unstable runs.
- **Prompt injection risk**: Do not point Shannon at untrusted or adversarial codebases. AI-powered tools that read source code can be influenced by malicious repository content.
## Scope of Analysis
+4 -4
View File
@@ -11,7 +11,7 @@ Shannon uses workspaces to store scan state, logs, prompts, and deliverables. Wo
- Use `-w <name>` to give a run a custom name.
- To resume a run, pass the same workspace name with `-w`.
- Each agent's progress is checkpointed so resumed runs can skip completed work.
- The final report is surfaced at the workspace root as `Security-Assessment-Report.md`. Run internals — deliverables, logs, prompts, and session state — live under a hidden `.shannon/` directory.
- The final report is surfaced at the workspace root as `Security-Assessment-Report.pdf` and `Security-Assessment-Report.md`. Run internals — deliverables, logs, prompts, and session state — live under a hidden `.shannon/` directory.
> [!NOTE]
> The URL must match the original workspace URL when resuming. Shannon rejects mismatched URLs to prevent cross-target contamination.
@@ -36,10 +36,10 @@ Resume an auto-named workspace:
npx @keygraph/shannon start -u https://example.com -r /path/to/repo -w example-com_shannon-1771007534808
```
List all workspaces:
List completed scans:
```bash
npx @keygraph/shannon workspaces
npx @keygraph/shannon scans
```
Source-build equivalents:
@@ -47,5 +47,5 @@ Source-build equivalents:
```bash
./shannon start -u https://example.com -r /path/to/repo -w my-audit
./shannon start -u https://example.com -r /path/to/repo -w example-com_shannon-1771007534808
./shannon workspaces
./shannon scans
```
+1 -1
View File
@@ -12,7 +12,7 @@ if [ -n "$TARGET_UID" ] && [ "$TARGET_UID" != "$CURRENT_UID" ]; then
groupadd -g "$TARGET_GID" pentest
useradd -u "$TARGET_UID" -g pentest -s /bin/bash -M pentest
chown -R pentest:pentest /app/sessions /app/workspaces /tmp/.claude
chown -R pentest:pentest /app/sessions /app/workspaces /tmp/.claude /tmp/.pi
fi
exec su -m pentest -c "exec $*"
+162 -45
View File
@@ -8,27 +8,30 @@
# File: README.md
> [!NOTE]
> **[Shannon 2.0 now runs on the Pi harness](https://github.com/KeygraphHQ/shannon/discussions/393)**
> **[Shannon 2.0 is officially here](https://github.com/KeygraphHQ/shannon/discussions/405)**
<div align="center">
<img src="./assets/github-banner.png" alt="Shannon - AI Pentester by Keygraph" width="100%">
# Shannon - AI Pentester by Keygraph
<picture>
<source media="(prefers-color-scheme: dark)" srcset="./assets/github-banner-dark.png">
<source media="(prefers-color-scheme: light)" srcset="./assets/github-banner-light.png">
<img src="./assets/github-banner-light.png" alt="Shannon, AI Pentester for Web Apps and APIs, by Keygraph" width="100%">
</picture>
<a href="https://trendshift.io/repositories/15604" target="_blank"><img src="https://trendshift.io/api/badge/repositories/15604" alt="KeygraphHQ%2Fshannon | Trendshift" style="width: 250px; height: 55px;" width="250" height="55"/></a>
Shannon is an autonomous, white-box AI pentester for web applications and APIs. <br />
### Shannon is an autonomous, AI pentester for web applications and APIs.
It analyzes your source code, identifies attack paths, and executes real exploits to prove vulnerabilities before they reach production.
**This repository is Shannon Open Source: the full agent, run locally from your command line.**
---
<a href="https://discord.gg/9ZqQPuhJB7"><img src="./assets/discord.png" height="40" alt="Join Discord"></a>
<a href="https://keygraph.io/"><img src="./assets/Keygraph_Button.png" height="40" alt="Visit Keygraph.io"></a>
<a href="https://discord.gg/9ZqQPuhJB7"><picture><source media="(prefers-color-scheme: dark)" srcset="./assets/discord_button_dark.png"><source media="(prefers-color-scheme: light)" srcset="./assets/discord_button_light.png"><img src="./assets/discord_button_light.png" height="40" alt="Join Discord"></picture></a>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;<a href="https://keygraph.io/"><picture><source media="(prefers-color-scheme: dark)" srcset="./assets/keygraph_button_dark.png"><source media="(prefers-color-scheme: light)" srcset="./assets/keygraph_button_light.png"><img src="./assets/keygraph_button_light.png" height="40" alt="Visit Keygraph.io"></picture></a>
---
</div>
> [!TIP]
@@ -44,13 +47,14 @@ It analyzes your source code, identifies attack paths, and executes real exploit
- [Architecture](#architecture)
- [Documentation](#documentation)
- [Safety, Scope, and Limitations](#safety-scope-and-limitations)
- [License and Enterprise Licensing](#license-and-enterprise-licensing)
- [License](#license)
- [About Keygraph](#about-keygraph)
- [Community and Support](#community-and-support)
- [Common Questions](#common-questions)
## What is Shannon?
Shannon is an autonomous AI pentester developed by [Keygraph](https://keygraph.io). It performs white-box security testing of web applications and their underlying APIs by combining source-code analysis with live exploitation.
Shannon is an autonomous AI pentester developed by [Keygraph](https://keygraph.io). It performs security testing of web applications and their underlying APIs by combining source-code analysis with live exploitation.
Shannon analyzes your web application's source code to identify potential attack vectors, then uses browser automation and command-line tools to execute real exploits against the running application and its APIs. Only vulnerabilities with a working proof-of-concept are included in the final report.
@@ -62,6 +66,12 @@ Thanks to tools like Claude Code and Cursor, your team ships code non-stop. But
Shannon closes that gap by providing on-demand, automated penetration testing that can run against every build or release.
### Why "Shannon"?
It's named after Claude Shannon, the father of information theory. At its core, pentesting is an information problem: every probe reduces uncertainty about a system's state. The best tools maximize the signal gained from every request, turning those bits of knowledge into an exploit path.
Also, we wanted you to be able to say, "Hey Claude, run Shannon" to find all the security flaws in your vibe-coded app.
## Shannon in Action
<p align="center">
@@ -82,7 +92,7 @@ Sample penetration test reports from intentionally vulnerable applications, prod
- **Docker**: required for the worker container.
- **Node.js 18+**: required for the recommended `npx` workflow.
- **AI provider credentials**: Anthropic, OpenAI, xAI, or AWS Bedrock - or [any other provider](docs/ai-providers.md#any-other-provider). Claude models are recommended. Gateway and proxy setups are documented separately.
- **AI provider credentials**: Shannon runs on Anthropic, OpenAI, xAI, AWS Bedrock, [any other provider](docs/ai-providers.md#any-other-provider) in the harness catalogue, and any endpoint that speaks the Anthropic Messages API or the OpenAI Chat Completions or Responses API through a [custom base URL](docs/ai-providers.md#custom-base-url). You bring your own key, and Keygraph never proxies your model traffic. Shannon is provider-agnostic. See [AI providers](docs/ai-providers.md#suggested-models) for suggested model IDs.
- **Cyber safeguards cleared with your provider**: Anthropic and OpenAI apply real-time safeguards to cyber-security workloads, which can interrupt a scan mid-run. Complete their guidance for legitimate security testers before your first run - see [AI providers](docs/ai-providers.md#cyber-safeguards-do-this-before-your-first-scan).
### Run Shannon
@@ -103,7 +113,11 @@ Shannon pulls the worker image from Docker Hub, starts the required local infras
For source builds, authenticated scans, provider-specific setup, and platform notes, see [Documentation](#documentation).
> [!TIP]
> **Prefer to run on your Claude Code subscription instead of API credits?** The [`shannon-v1`](https://github.com/KeygraphHQ/shannon/tree/shannon-v1) branch is the last release built on the Claude Agent SDK, so it accepts a Claude Code OAuth token. Generate one with `claude setup-token`, then run `npx @keygraph/shannon@1.9.0 setup` and pick **OAuth Token**. Pentests then cost nothing beyond your existing subscription.
> **Prefer to use a subscription instead of API credits?**
>
> - **OpenAI Codex:** The latest version of Shannon supports ChatGPT Plus and Pro subscriptions. Follow the [OpenAI Codex subscription setup guide](docs/ai-providers.md#openai-codex-chatgpt-pluspro-subscription) to get started.
> - **xAI (Grok):** The latest version of Shannon supports xAI subscriptions. Follow the [xAI subscription setup guide](docs/ai-providers.md#xai-grok-subscription) to get started.
> - **Claude Code:** The latest version of Shannon does not support Claude Code subscriptions. Follow the [Claude Code subscription setup guide](docs/ai-providers.md#claude-code-subscription) to use version `1.9.0`, which is the final release built on the Claude Agent SDK.
## Key Capabilities
@@ -113,6 +127,8 @@ For source builds, authenticated scans, provider-specific setup, and platform no
- **Authenticated testing**: configuration files can describe login flows, test credentials, TOTP, email-based login flows, focus areas, and rules of engagement.
- **OWASP-focused coverage**: Shannon targets exploitable Injection, XSS, SSRF, Broken Authentication, and Broken Authorization issues.
- **Resumable workspaces**: Shannon can resume interrupted runs without re-running completed agents.
- **Machine-readable output**: Shannon emits findings as structured JSON, and as SARIF 2.1.0 by default on exploit-mode scans (opt out with `report.sarif: "false"`). SARIF is the OASIS standard for static analysis results, so findings flow into any code scanning service, vulnerability management platform, security dashboard, or CI/CD pipeline that reads it.
- **Bring your own key, provider-agnostic**: Shannon runs on Anthropic, OpenAI, xAI, AWS Bedrock, and any endpoint speaking the Anthropic Messages API or the OpenAI Chat Completions or Responses API, including self-hosted models served through Ollama, vLLM, or LM Studio and gateways such as OpenRouter and LiteLLM. You supply the credentials, so source code and model traffic stay inside your infrastructure. Local and self-hosted models are technically supported but not recommended: they may not follow Shannon's instructions or tool-use constraints as reliably as frontier models, so take that path only if you know how your chosen model behaves.
## Editions
@@ -199,7 +215,7 @@ Use these guides for operational detail:
| --- | --- |
| [Source build and CLI commands](docs/development.md) | Cloning, building, common commands, output paths, and local development. |
| [Configuration](docs/configuration.md) | Authenticated testing, login flows, rules of engagement, and report filters. |
| [AI providers](docs/ai-providers.md) | Selecting the model, the supported providers (Anthropic, OpenAI, xAI, AWS Bedrock), and custom gateways. |
| [AI providers](docs/ai-providers.md) | Selecting the model, the supported providers (Anthropic, OpenAI, xAI, AWS Bedrock, and any other Pi-supported provider), and custom gateways. |
| [Platforms and networking](docs/platforms.md) | Windows/WSL2, Linux, macOS, Docker networking, local apps, and custom hostnames. |
| [Workspaces and resuming](docs/workspaces.md) | Naming workspaces, resuming interrupted scans, and workspace storage. |
| [Safety and limitations](docs/safety.md) | Authorized-use requirements, non-production guidance, mutative effects, cost, and model caveats. |
@@ -216,13 +232,13 @@ Important limitations:
- Shannon Open Source focuses on actively exploitable issues such as Injection, XSS, SSRF, Broken Authentication, and Broken Authorization. Broader static-analysis coverage, including vulnerable dependencies and insecure configurations, is delivered through the Keygraph platform.
- Findings still require human review. LLM-generated reports can contain weakly supported or incorrect details.
- Shannon is officially supported with Claude models. Smaller, alternative, or proxied non-Claude models may be incomplete or unstable.
- Anthropic, OpenAI, xAI, and AWS Bedrock are built-in providers, and any Anthropic Messages API or OpenAI Chat Completions or Responses API endpoint works through a custom base URL. Model capability varies, and a model that does not follow Shannon's instructions or tool-use constraints reliably will produce weaker results.
- A full run can take roughly 1 to 1.5 hours and may incur LLM API costs depending on model pricing and application complexity.
- Do not scan untrusted or adversarial codebases. AI-powered tools that read source code can be exposed to prompt injection.
Read the full [Safety and limitations](docs/safety.md) guide before running Shannon in a new environment.
## License and Enterprise Licensing
## License
Shannon Open Source is licensed under the [GNU Affero General Public License v3.0](LICENSE).
@@ -255,6 +271,40 @@ Stay connected:
- [Twitter/X: @KeygraphHQ](https://twitter.com/KeygraphHQ)
- [LinkedIn: Keygraph](https://linkedin.com/company/keygraph)
## Common Questions
### Is Shannon free?
Yes. Shannon Open Source is free and licensed under AGPL-3.0. You run it yourself from the command line. Your only cost is the AI provider credits you supply.
### Can I self-host Shannon?
Yes. Shannon Open Source runs entirely on your own infrastructure in an ephemeral Docker container. Your source code is mounted read-only and never leaves your environment.
### Does Shannon support bring your own key (BYOK)?
Yes, always. You provide the LLM credentials Shannon uses to run a pentest, in every deployment, open source and commercial. Keygraph never proxies your model traffic.
### Does Shannon output SARIF?
Yes. Shannon emits SARIF 2.1.0, the OASIS standard format for static analysis results, alongside structured JSON. Any SARIF consumer reads it: code scanning services, vulnerability management platforms, security dashboards, and CI/CD pipelines. It is written by default on exploit-mode scans; set `report.sarif` to `"false"` in your configuration file to opt out.
### Which AI providers does Shannon support?
Anthropic, OpenAI, xAI, and AWS Bedrock are built in and configured directly by provider ID. Beyond those, Shannon runs on any endpoint that implements the Anthropic Messages API or the OpenAI Chat Completions or Responses API, reached through a custom base URL. The rule is the API format, not the vendor. Shannon uses a single unified model setting throughout a pentest.
### Can I run Shannon on a local or self-hosted model?
Technically yes, but it is not recommended. Shannon works with local models served through Ollama, vLLM, or LM Studio, which expose an OpenAI-compatible endpoint, as well as routers such as OpenRouter and gateways such as LiteLLM. Point Shannon at the endpoint with a custom base URL. Capability varies, and a model that does not follow Shannon's instructions or tool-use constraints reliably will produce weaker pentests than a frontier model, so take this path only if you know how your chosen model behaves. See [AI providers](docs/ai-providers.md#custom-base-url).
### Does Shannon actually exploit vulnerabilities, or just scan?
Shannon executes real exploits. It reports a finding only when it has produced a working proof-of-concept, and discards hypotheses it cannot prove. It is a pentester, not a scanner.
### Is Shannon free for startups and nonprofits?
Shannon Open Source is free for everyone. In addition, the Keygraph Community Program gives eligible nonprofits and early-stage startups free access to the commercial Keygraph platform. See [keygraph.io](https://keygraph.io).
<p align="center">
<b>Built by <a href="https://keygraph.io">Keygraph</a></b>
</p>
@@ -323,7 +373,8 @@ Monitor progress:
```bash
npx @keygraph/shannon logs <workspace>
npx @keygraph/shannon status
npx @keygraph/shannon status <workspace>
npx @keygraph/shannon scans
npx @keygraph/shannon version
```
@@ -331,7 +382,8 @@ Source-build equivalents:
```bash
./shannon logs <workspace>
./shannon status
./shannon status <workspace>
./shannon scans
./shannon version
```
@@ -344,16 +396,17 @@ open http://localhost:8233
Stop Shannon:
```bash
npx @keygraph/shannon stop
npx @keygraph/shannon stop --clean # confirms first; add --yes (or -y) to skip
npx @keygraph/shannon uninstall # confirms first; add --yes (or -y) to skip
npx @keygraph/shannon stop <workspace> # stop one scan (confirms first; add --yes/-y to skip)
npx @keygraph/shannon stop --all # stop all scans (Temporal stays up)
npx @keygraph/shannon reset # stop everything and wipe all Temporal data (type 'confirm' to proceed; cannot be skipped)
```
Source-build equivalents:
```bash
./shannon stop
./shannon stop --clean # add --yes (or -y) to skip the confirmation
./shannon stop <workspace> # stop one scan (confirms first; add --yes/-y to skip)
./shannon stop --all # stop all scans (Temporal stays up)
./shannon reset # stop everything and wipe all Temporal data (type 'confirm' to proceed; cannot be skipped)
```
Usage examples:
@@ -371,8 +424,11 @@ npx @keygraph/shannon start -u https://example.com -r /path/to/repo -o ./my-repo
# Named workspace.
npx @keygraph/shannon start -u https://example.com -r /path/to/repo -w q1-audit
# List all workspaces.
npx @keygraph/shannon workspaces
# Stream the log until the scan finishes, then exit on its outcome (useful in CI).
npx @keygraph/shannon start -u https://example.com -r /path/to/repo --follow
# List completed scans.
npx @keygraph/shannon scans
```
Source-build examples:
@@ -382,7 +438,8 @@ Source-build examples:
./shannon start -u https://example.com -r /path/to/repo -c /path/to/my-config.yaml
./shannon start -u https://example.com -r /path/to/repo -o ./my-reports
./shannon start -u https://example.com -r /path/to/repo -w q1-audit
./shannon workspaces
./shannon start -u https://example.com -r /path/to/repo --follow
./shannon scans
# Rebuild the worker image.
./shannon build --no-cache
@@ -397,11 +454,12 @@ Results are saved to the workspaces directory:
Use `-o <path>` to copy deliverables to a custom output directory after a run completes.
Output structure — the run directory's top level holds only the final report; everything else is nested under a hidden `.shannon/` directory:
Output structure — the run directory's top level holds the final report, in PDF and Markdown; everything else is nested under a hidden `.shannon/` directory:
```text
workspaces/{hostname}_{sessionId}/
|-- Security-Assessment-Report.md # the final report (the deliverable)
|-- Security-Assessment-Report.pdf # the final report (PDF)
|-- Security-Assessment-Report.md # the final report (Markdown)
`-- .shannon/ # internals
|-- deliverables/ # report source, per-phase analysis, queues
|-- agents/ # per-agent logs
@@ -516,33 +574,32 @@ rules:
# min_confidence: low
# guidance: |
# Drop findings about missing security headers and rate-limit gaps.
# sarif: "true"
# sarif: "false"
```
## Report Options
| Key | Effect |
| --- | --- |
| `min_severity` | Drops findings rated below this severity. Applies only when `exploit` is `"true"`. |
| `min_severity` | Drops findings rated below this severity. Applies in both exploitative and analysis-only runs. |
| `min_confidence` | Drops findings rated below this confidence. Applies only when `exploit` is `"false"`. |
| `guidance` | Free-text instruction to the report agent, such as which topics to exclude. |
| `sarif` | Emits a SARIF 2.1.0 log alongside the Markdown report. Requires `exploit: "true"`. |
| `sarif` | SARIF 2.1.0 log alongside the Markdown report. On by default for exploit runs; set `"false"` to opt out. Ignored when `exploit` is `"false"`. |
A finding carries one rating or the other, never both: an exploited finding is rated by severity, an analysis-only finding by confidence. Setting the threshold that does not apply to the run is ignored, and Shannon logs a warning naming the one to use instead.
Every finding carries a severity, but it does not mean the same thing in each mode: an exploitative run measures severity from what the exploit demonstrated, while an analysis-only run assesses it from the class of flaw and the impact it would have. An analysis-only finding carries a confidence rating alongside its severity, since nothing was proven. Setting `min_confidence` on an exploitative run is ignored, and Shannon logs a warning naming the threshold to use instead.
### SARIF Output
Set `sarif: "true"` to write `report.sarif` next to `Security-Assessment-Report.md` at the workspace root, for upload to GitHub code scanning or any other SARIF consumer.
On exploit-mode runs Shannon writes `report.sarif` next to `Security-Assessment-Report.pdf` at the workspace root by default, for upload to GitHub code scanning or any other SARIF consumer. No configuration is needed; set `sarif: "false"` to opt out.
```yaml
exploit: "true"
report:
sarif: "true"
sarif: "false"
```
Each finding becomes one SARIF result, filed under a rule per vulnerability class (`shannon/injection`, `shannon/xss`, `shannon/auth`, `shannon/authz`, `shannon/ssrf`) and tagged with its OWASP Top Ten 2025 category. Results are anchored to the code location the analysis phase recorded, falling back to the HTTP entry point when the finding names no file. Severity maps onto SARIF's three levels: `critical` and `high` become `error`, `medium` becomes `warning`, everything else becomes `note`.
The log is written only for exploitative runs. An analysis-only run rates findings by confidence and produces no severity, so there is nothing to populate `level` with; `sarif` is ignored when `exploit` is `"false"`.
The log is written only for exploitative runs. `sarif` is ignored when `exploit` is `"false"`.
Supported rule types include `url_path`, `subdomain`, `domain`, `method`, `header`, `parameter`, and `code_path`.
@@ -605,7 +662,7 @@ Shannon forwards only the selected provider's credential into the scan container
### Any other provider
Shannon accepts any provider and model present in the Pi harness catalogue. Browse them at [pi.dev/models](https://pi.dev/models). These are technically supported but not recommended. Claude models are best-supported (see the note below).
Shannon accepts any provider and model present in the Pi harness catalogue. Browse them at [pi.dev/models](https://pi.dev/models).
```bash
export SHANNON_AI_API_KEY=your-api-key # the provider's API key
@@ -617,7 +674,7 @@ This path covers providers whose credential is a single API key. Providers that
`npx @keygraph/shannon setup` exposes this as the **Other provider** option.
> [!IMPORTANT]
> Claude models are the best-supported option. Shannon's evaluations, internal testing, and agent harness are tuned for Claude. Other models are permitted and validated against the harness catalogue, but may not follow Shannon's instructions or tool-use constraints as reliably. Use them at your own risk.
> Models are validated against the harness catalogue, but capability varies. A model that does not follow Shannon's instructions or tool-use constraints reliably will produce weaker pentests. Evaluate the model you choose against your own targets before depending on its results.
## Cyber safeguards (do this before your first scan)
@@ -638,7 +695,7 @@ These are the models `npx @keygraph/shannon setup` offers, best-first. They are
| --- | --- |
| `anthropic` | `claude-sonnet-4-6`, `claude-opus-4-8`, `claude-opus-4-7`, `claude-haiku-4-5-20251001` |
| `openai` | `gpt-5.6-sol`, `gpt-5.5`, `gpt-5.4` |
| `xai` | `grok-4.5` |
| `xai` | `grok-4.6`, `grok-4.5` |
| `amazon-bedrock` | `us.anthropic.claude-sonnet-4-6`, `us.anthropic.claude-opus-4-8`, `us.anthropic.claude-opus-4-7` |
Bedrock IDs are region-prefixed and must be enabled in your account, so the ID that works for you may differ from the one listed here.
@@ -724,12 +781,73 @@ The variable is rejected in preflight where it cannot take effect: with a non-`o
`npx @keygraph/shannon setup` covers this under **Custom Base URL**, which asks which API your gateway serves and configures the matching provider for you.
## OpenAI Codex (ChatGPT Plus/Pro subscription)
A ChatGPT Plus or Pro Codex subscription can run Shannon. Shannon reuses a login created by Pi.
Before running a pentest, review the [cyber safeguards requirements](#cyber-safeguards-do-this-before-your-first-scan).
1. Install Pi by following the instructions at [pi.dev](https://pi.dev).
2. Log in with your subscription using Pi's [subscription authentication guide](https://pi.dev/docs/latest/providers#subscriptions). This creates `~/.pi/agent/auth.json` with an `openai-codex` entry.
3. Select a Codex model and enable Pi authentication:
```bash
export SHANNON_USE_PI_AUTH=1
export SHANNON_AI_MODEL=openai-codex:gpt-5.5
```
4. In npx mode, run `npx @keygraph/shannon start ...` from the same shell. In source-build mode, add the two variables to `.env` and run `./shannon start ...`.
Supported Codex models are `gpt-5.6-sol`, `gpt-5.5`, and `gpt-5.4`.
## xAI (Grok subscription)
An xAI subscription can run Shannon. Shannon reuses a login created by Pi.
1. Install Pi by following the instructions at [pi.dev](https://pi.dev).
2. Log in with your subscription using Pi's [subscription authentication guide](https://pi.dev/docs/latest/providers#subscriptions). This creates `~/.pi/agent/auth.json` with an `xai` entry.
3. Select an xAI model and enable Pi authentication:
```bash
export SHANNON_USE_PI_AUTH=1
export SHANNON_AI_MODEL=xai:grok-4.6
```
4. In npx mode, run `npx @keygraph/shannon start ...` from the same shell. In source-build mode, add the two variables to `.env` and run `./shannon start ...`.
Suggested Grok models are `grok-4.6` and `grok-4.5`.
## Claude Code subscription
The latest version of Shannon does not support Claude Code subscriptions. The [`shannon-v1`](https://github.com/KeygraphHQ/shannon/tree/shannon-v1) branch is the final release built on the Claude Agent SDK and supports Claude Code OAuth.
Before running a pentest, review the [cyber safeguards requirements](#cyber-safeguards-do-this-before-your-first-scan).
1. Generate a Claude Code OAuth token:
```bash
claude setup-token
```
2. Run the setup flow for the final `shannon-v1` release:
```bash
npx @keygraph/shannon@1.9.0 setup
```
3. Select **OAuth Token** and enter the token generated by Claude Code.
4. Start the pentest with `npx @keygraph/shannon@1.9.0 start ...`.
These instructions apply only to `shannon-v1`.
## Validation
Checks run before a scan starts, so mistakes fail immediately rather than partway through a run:
- **Provider and model ID** — validated against the Pi harness catalogue. An unknown provider or model ID fails preflight with a pointer to [pi.dev/models](https://pi.dev/models). A custom base URL exempts the model ID, since a gateway may serve its own names.
- **Credential presence** — always validated for the selected provider.
- **Credential presence** — validated for the selected provider, or read from Pi when `SHANNON_USE_PI_AUTH=1`.
- **Credential validity** — one minimal request against the model the scan will use, so a rejected key, an exhausted quota, or a model the account cannot reach fails before any agent runs. Bedrock included: its bearer token and region go through the same probe.
## Migrating from the three-tier configuration
@@ -765,7 +883,7 @@ This guide covers platform-specific notes and Docker networking behavior.
## Windows
Shannon on Windows is supported through WSL2. Native Windows, including Git Bash, is not supported.
Shannon on Windows is supported through WSL2, which behaves like Linux for everything below. Native Windows, including Git Bash, is community-supported: contributions are welcome, but Keygraph does not actively develop or test against it.
### Ensure WSL2
@@ -861,7 +979,7 @@ Shannon uses workspaces to store scan state, logs, prompts, and deliverables. Wo
- Use `-w <name>` to give a run a custom name.
- To resume a run, pass the same workspace name with `-w`.
- Each agent's progress is checkpointed so resumed runs can skip completed work.
- The final report is surfaced at the workspace root as `Security-Assessment-Report.md`. Run internals — deliverables, logs, prompts, and session state — live under a hidden `.shannon/` directory.
- The final report is surfaced at the workspace root as `Security-Assessment-Report.pdf` and `Security-Assessment-Report.md`. Run internals — deliverables, logs, prompts, and session state — live under a hidden `.shannon/` directory.
> [!NOTE]
> The URL must match the original workspace URL when resuming. Shannon rejects mismatched URLs to prevent cross-target contamination.
@@ -886,10 +1004,10 @@ Resume an auto-named workspace:
npx @keygraph/shannon start -u https://example.com -r /path/to/repo -w example-com_shannon-1771007534808
```
List all workspaces:
List completed scans:
```bash
npx @keygraph/shannon workspaces
npx @keygraph/shannon scans
```
Source-build equivalents:
@@ -897,7 +1015,7 @@ Source-build equivalents:
```bash
./shannon start -u https://example.com -r /path/to/repo -w my-audit
./shannon start -u https://example.com -r /path/to/repo -w example-com_shannon-1771007534808
./shannon workspaces
./shannon scans
```
---
@@ -934,7 +1052,7 @@ For maximum isolation, run Shannon inside a disposable virtual machine.
## LLM and Automation Caveats
- **Verification is required**: Shannon uses a proof-by-exploitation methodology, but final reports can still contain weakly supported or incorrect details. Human review is essential.
- **Model support**: Shannon is officially supported only with Claude models. Alternative models may be incomplete, inaccurate, or unstable.
- **Model support**: results vary by model. A model that does not follow Shannon's instructions or tool-use constraints reliably may produce incomplete, inaccurate, or unstable runs.
- **Prompt injection risk**: Do not point Shannon at untrusted or adversarial codebases. AI-powered tools that read source code can be influenced by malicious repository content.
## Scope of Analysis
@@ -955,7 +1073,6 @@ For broader coverage, the Keygraph platform adds black-box and white-box agentic
A full test run typically takes roughly 1 to 1.5 hours. LLM API costs vary by model pricing, target complexity, selected provider, and concurrency.
---
# File: docs/coverage-roadmap.md
+2 -2
View File
@@ -6,13 +6,13 @@ Use this file as the concise entry point for AI agents and LLMs reading this rep
## Start Here
- [README](README.md): Main project overview, editions, quick start, Shannon capabilities, Keygraph platform positioning, safety notes, licensing, and support links.
- [README](README.md): Main project overview, editions, quick start, Shannon capabilities, Keygraph platform positioning, common questions, safety notes, licensing, and support links.
- [Full Combined Context](llms-full.txt): README and documentation combined into one file for agents that need maximum local context.
## Shannon
- [Development](docs/development.md): Source-build workflow, common CLI commands, repository paths, and output locations.
- [Configuration](docs/configuration.md): Authenticated testing, login flows, rules of engagement, report filters, credential precedence, adaptive thinking, and rate-limit settings.
- [Configuration](docs/configuration.md): Authenticated testing, login flows, rules of engagement, report filters, credential precedence, and rate-limit settings.
- [AI Providers](docs/ai-providers.md): Anthropic, OpenAI, xAI, AWS Bedrock, any other Pi-supported provider, and custom gateway setup.
- [Platforms and Networking](docs/platforms.md): Windows/WSL2, Linux, macOS, Docker networking, local applications, and custom hostnames.
- [Workspaces and Resuming](docs/workspaces.md): Workspace storage, naming, resuming interrupted scans, and examples.
+81 -66
View File
@@ -26,6 +26,9 @@ importers:
'@clack/prompts':
specifier: ^1.1.0
version: 1.1.0
'@temporalio/client':
specifier: ^1.11.0
version: 1.15.0
chokidar:
specifier: ^5.0.0
version: 5.0.0
@@ -43,17 +46,17 @@ importers:
apps/worker:
dependencies:
'@earendil-works/pi-agent-core':
specifier: ^0.82.1
version: 0.82.1(@modelcontextprotocol/sdk@1.29.0(zod@4.3.6))(ws@8.21.0)(zod@4.3.6)
specifier: ^0.84.2
version: 0.84.2(@modelcontextprotocol/sdk@1.29.0(zod@4.3.6))(ws@8.21.0)(zod@4.3.6)
'@earendil-works/pi-ai':
specifier: ^0.82.1
version: 0.82.1(@modelcontextprotocol/sdk@1.29.0(zod@4.3.6))(ws@8.21.0)(zod@4.3.6)
specifier: ^0.84.2
version: 0.84.2(@modelcontextprotocol/sdk@1.29.0(zod@4.3.6))(ws@8.21.0)(zod@4.3.6)
'@earendil-works/pi-coding-agent':
specifier: ^0.82.1
version: 0.82.1(@modelcontextprotocol/sdk@1.29.0(zod@4.3.6))(ws@8.21.0)(zod@4.3.6)
specifier: ^0.84.2
version: 0.84.2(@modelcontextprotocol/sdk@1.29.0(zod@4.3.6))(ws@8.21.0)(zod@4.3.6)
'@gotgenes/pi-permission-system':
specifier: ^10.9.0
version: 10.9.0(@earendil-works/pi-coding-agent@0.82.1(@modelcontextprotocol/sdk@1.29.0(zod@4.3.6))(ws@8.21.0)(zod@4.3.6))(@earendil-works/pi-tui@0.82.1)
version: 10.9.0(@earendil-works/pi-coding-agent@0.84.2(@modelcontextprotocol/sdk@1.29.0(zod@4.3.6))(ws@8.21.0)(zod@4.3.6))(@earendil-works/pi-tui@0.84.2)
'@temporalio/activity':
specifier: ^1.11.0
version: 1.15.0
@@ -289,22 +292,34 @@ packages:
'@clack/prompts@1.1.0':
resolution: {integrity: sha512-pkqbPGtohJAvm4Dphs2M8xE29ggupihHdy1x84HNojZuMtFsHiUlRvqD24tM2+XmI+61LlfNceM3Wr7U5QES5g==}
'@earendil-works/pi-agent-core@0.82.1':
resolution: {integrity: sha512-Z3kloziJIE2dmrisRckZX8zDca/gIv9/YdFAzeoqpHiLV2wsni6bL4hInNSjVKLbqT+4kqLIkph2JQLKvSepjg==}
'@earendil-works/pi-agent-core@0.84.2':
resolution: {integrity: sha512-8Pn3wSCxj0cfo5I6jxQYVB/3uuQRmHhAlEclyjqpOuMEdQMIODHizRogv56FLdbU+dTiGnybeHQ2N+sV1/L2YA==}
engines: {node: '>=22.19.0'}
'@earendil-works/pi-ai@0.82.1':
resolution: {integrity: sha512-3WFYRhEp3lQB3444EhPMBcM7zSaEUE3eJgHOR7s4081NLqbw/FsWilIKWXSua0Gv3sRr7m9xMidR3pPDE7jI/A==}
'@earendil-works/pi-ai@0.84.2':
resolution: {integrity: sha512-6MzsrYIYNVlE7SfpbL2yYb67Qo58p/7Q+xWG1RZvoX1P80aRCHSod2/13aFpxkow1lPO2LEh3c495J0Gwmyjig==}
engines: {node: '>=22.19.0'}
hasBin: true
'@earendil-works/pi-coding-agent@0.82.1':
resolution: {integrity: sha512-zbkAhoIuDPMF3pKuja0ajZabrMWU29FUMV9A/XMXT/XC1yXs5xt6t6t13GogQFsDrDqbFP4DkZQO1w8rWRAzYA==}
'@earendil-works/pi-client@0.84.2':
resolution: {integrity: sha512-/RFSPhD/bZbpOp1oJj+UneSUFSgZhWxzcSENUY+8+8xhoBrWXMYI2t77XNx4Yf+c8YK2qTHquForhNcelYpXvg==}
engines: {node: '>=22.19.0'}
'@earendil-works/pi-coding-agent@0.84.2':
resolution: {integrity: sha512-l4E+B7hgXKWddRo8bC/eSue2aWZjEgJ9xIpf5p0Og+lq8a2TArCwJ0HCoCPCgaBP/tN4zbYH/wOwvx9pJpeLCA==}
engines: {node: '>=22.19.0'}
hasBin: true
'@earendil-works/pi-tui@0.82.1':
resolution: {integrity: sha512-9yN8hALfKaxZq7n54EMxqhFCWnMi6LHkraMJ/1YjHiATq75XrI6XDMVppn9EDtiK7Fks8hUe1SDXUTrIvwRWfQ==}
'@earendil-works/pi-protocol@0.84.2':
resolution: {integrity: sha512-jbBh03fkeckWEroHpcZBr4w5/Ibat8WwdXFlXHivYQImrQNFtLpDeL0t1cku4hmK0q3pceIRQHkw4fwbM4YILQ==}
engines: {node: '>=22.19.0'}
'@earendil-works/pi-telemetry@0.84.2':
resolution: {integrity: sha512-wg5caea7uIv1BHRBm2Y116RvFG4oSAiP5qk9tA2463PDGIr4K8M1Ceyyg5DOpF/shUUl0gk826yQJAeAcHYB9g==}
engines: {node: '>=22.19.0'}
'@earendil-works/pi-tui@0.84.2':
resolution: {integrity: sha512-ds2TLihOnM5sLJB3VpXV6y0uR5efVuHf4MN7yDpsty6hA2DUO/EDVzjp/0od0G2JslzVLMjT8T8zavtxVb+qbg==}
engines: {node: '>=22.19.0'}
'@emnapi/core@1.9.1':
@@ -554,14 +569,6 @@ packages:
resolution: {integrity: sha512-ABnA53mdfkGZwOFUdZNv2S0CWGO/EIuPj8Vv9xmBFmSYg/qFc7ihO6q5FcQjvoE67kZpWkEc4AhD6B/os04yuA==}
engines: {node: '>= 10'}
'@mistralai/mistralai@2.2.6':
resolution: {integrity: sha512-W8pX7zHxjJvMIpw8JMxeJEleapXX0Q9NPszdNzqkM3MIEoIGPObdodujj+WHteXEvGfaP/AMwlNyRfEzSY6dQQ==}
peerDependencies:
'@opentelemetry/api': ^1.9.0
peerDependenciesMeta:
'@opentelemetry/api':
optional: true
'@modelcontextprotocol/sdk@1.29.0':
resolution: {integrity: sha512-zo37mZA9hJWpULgkRpowewez1y6ML5GsXJPY8FI0tBBCd77HEvza4jDqRKOXgHNn867PVGCyTdzqpz0izu5ZjQ==}
engines: {node: '>=18'}
@@ -582,10 +589,6 @@ packages:
resolution: {integrity: sha512-3giAOQvZiH5F9bMlMiv8+GSPMeqg0dbaeo58/0SlA9sxSqZhnUtxzX9/2FzyhS9sWQf5S0GJE0AKBrFqjpeYcg==}
engines: {node: '>=8.0.0'}
'@opentelemetry/semantic-conventions@1.43.0':
resolution: {integrity: sha512-eSYWTm620tTk45EKSedaUL8MFYI8hW164hIXsgIHyxu3VobUB3fFCu5t0hQby6OoWRPsG1KkKUG2M5UadiLiVg==}
engines: {node: '>=14'}
'@oxc-project/types@0.122.0':
resolution: {integrity: sha512-oLAl5kBpV4w69UtFZ9xqcmTi+GENWOcPF7FCrczTiBbmC0ibXxCwyvZGbO39rCVEuLGAZM84DH0pUIyyv/YJzA==}
@@ -1373,6 +1376,10 @@ packages:
graceful-fs@4.2.11:
resolution: {integrity: sha512-RbJ5/jmFcNNCcDV5o9eTnBLJ/HszWV0P73bc+Ff4nS/rJj+YaS6IGyiOL0VoBYX+l1Wrl3k63h/KrH+nhJ0XvQ==}
grok-mermaid@0.2.2:
resolution: {integrity: sha512-XcJEP5dDC8liHBh52mlLjU18fNvu1ckFsu0QpIG3+APZ270fsj9wxpiA6cOURmbUEuoMVgjbC2+UYgTdCqqgzA==}
engines: {node: '>=18'}
has-flag@4.0.0:
resolution: {integrity: sha512-EykJT/Q1KjTWctppgIAgfSO0tKVuZUjhgMr17kqTumMl6Afv3EISleU7qZUzoXDFTAHTDC4NOoG/ZxU3EvlMPQ==}
engines: {node: '>=8'}
@@ -1617,9 +1624,8 @@ packages:
once@1.4.0:
resolution: {integrity: sha512-lNaJgI+2Q5URQBkccEKHTQOPaXdUxnZZElQTZY0MFUAuaEqe1E+Nyvgdz/aIyNi6Z9MzO5dv1H8n58/GELp3+w==}
openai@6.26.0:
resolution: {integrity: sha512-zd23dbWTjiJ6sSAX6s0HrCZi41JwTA1bQVs0wLQPZ2/5o2gxOJA5wh7yOAUgwYybfhDXyhwlpeQf7Mlgx8EOCA==}
hasBin: true
openai@6.40.0:
resolution: {integrity: sha512-MWtTjd/gQt4jpbji61NTgFWJLoY/PdRJ6wG9/ZDRMYNMlBKrCrSlkLI+KgHP1vR1qT6LKSAyAqIxno6lcK9JiA==}
peerDependencies:
ws: ^8.18.0
zod: ^3.25 || ^4.0
@@ -2000,6 +2006,9 @@ packages:
typebox@1.1.38:
resolution: {integrity: sha512-pZ0aQPmMmXoUvSbeuWf/Hzsc+avNw/Zd6VeE8CFgkVGWyuHPJvqeJJDeJqLve+K70LvjYIoleGcoJHPT17cWoA==}
typebox@1.3.7:
resolution: {integrity: sha512-meKuifc33Pccx0O6PdIzYMq3Og8zvP4TIi/a+Bw3AEMZMxOD0+RHGQvpglEe6Zdy3wZ8nqn/j95h8LUZLk/6Hg==}
typescript@5.9.3:
resolution: {integrity: sha512-jl1vZzPDinLr9eUt3J/t7V6FgNEw9QjvBPdysz9KfQDD41fQrC2Y4vKQdiaUpFT4bXlb1RHhLpp8wtm6M5TgSw==}
engines: {node: '>=14.17'}
@@ -2011,8 +2020,8 @@ packages:
undici-types@7.18.2:
resolution: {integrity: sha512-AsuCzffGHJybSaRrmr5eHr81mwJU3kjw6M+uprWvCXiNeN9SOGwQ3Jn8jb8m3Z6izVgknn1R0FTCEAP2QrLY/w==}
undici@8.5.0:
resolution: {integrity: sha512-xamtWoB1EshgjpmlXd7GGm2VfdDtw1+rD8uhry8pSNW3If6S8E0m2T2+orSKeZXEn/aPJMviCpDBA65WJt8zhg==}
undici@8.9.0:
resolution: {integrity: sha512-aWZpUj7XoGonMClx4gdDRfgBjqeA+F473aDmROQQbM9n6PRfK/u1q/a0X4wMTgcHfT8H6fpbt98PFuDUwFg2YA==}
engines: {node: '>=22.19.0'}
unionfs@4.6.0:
@@ -2428,12 +2437,13 @@ snapshots:
'@clack/core': 1.1.0
sisteransi: 1.0.5
'@earendil-works/pi-agent-core@0.82.1(@modelcontextprotocol/sdk@1.29.0(zod@4.3.6))(ws@8.21.0)(zod@4.3.6)':
'@earendil-works/pi-agent-core@0.84.2(@modelcontextprotocol/sdk@1.29.0(zod@4.3.6))(ws@8.21.0)(zod@4.3.6)':
dependencies:
'@earendil-works/pi-ai': 0.82.1(@modelcontextprotocol/sdk@1.29.0(zod@4.3.6))(ws@8.21.0)(zod@4.3.6)
'@earendil-works/pi-ai': 0.84.2(@modelcontextprotocol/sdk@1.29.0(zod@4.3.6))(ws@8.21.0)(zod@4.3.6)
'@earendil-works/pi-telemetry': 0.84.2
diff: 8.0.4
ignore: 7.0.5
typebox: 1.1.38
typebox: 1.3.7
yaml: 2.9.0
transitivePeerDependencies:
- '@modelcontextprotocol/sdk'
@@ -2443,19 +2453,19 @@ snapshots:
- ws
- zod
'@earendil-works/pi-ai@0.82.1(@modelcontextprotocol/sdk@1.29.0(zod@4.3.6))(ws@8.21.0)(zod@4.3.6)':
'@earendil-works/pi-ai@0.84.2(@modelcontextprotocol/sdk@1.29.0(zod@4.3.6))(ws@8.21.0)(zod@4.3.6)':
dependencies:
'@anthropic-ai/sdk': 0.91.1(zod@4.3.6)
'@aws-sdk/client-bedrock-runtime': 3.1048.0
'@earendil-works/pi-telemetry': 0.84.2
'@google/genai': 1.52.0(@modelcontextprotocol/sdk@1.29.0(zod@4.3.6))
'@mistralai/mistralai': 2.2.6(@opentelemetry/api@1.9.0)
'@opentelemetry/api': 1.9.0
'@smithy/node-http-handler': 4.7.3
http-proxy-agent: 7.0.2
https-proxy-agent: 7.0.6
openai: 6.26.0(ws@8.21.0)(zod@4.3.6)
openai: 6.40.0(ws@8.21.0)(zod@4.3.6)
partial-json: 0.1.7
typebox: 1.1.38
typebox: 1.3.7
transitivePeerDependencies:
- '@modelcontextprotocol/sdk'
- bufferutil
@@ -2464,16 +2474,23 @@ snapshots:
- ws
- zod
'@earendil-works/pi-coding-agent@0.82.1(@modelcontextprotocol/sdk@1.29.0(zod@4.3.6))(ws@8.21.0)(zod@4.3.6)':
'@earendil-works/pi-client@0.84.2':
dependencies:
'@earendil-works/pi-agent-core': 0.82.1(@modelcontextprotocol/sdk@1.29.0(zod@4.3.6))(ws@8.21.0)(zod@4.3.6)
'@earendil-works/pi-ai': 0.82.1(@modelcontextprotocol/sdk@1.29.0(zod@4.3.6))(ws@8.21.0)(zod@4.3.6)
'@earendil-works/pi-tui': 0.82.1
'@earendil-works/pi-protocol': 0.84.2
'@earendil-works/pi-coding-agent@0.84.2(@modelcontextprotocol/sdk@1.29.0(zod@4.3.6))(ws@8.21.0)(zod@4.3.6)':
dependencies:
'@earendil-works/pi-agent-core': 0.84.2(@modelcontextprotocol/sdk@1.29.0(zod@4.3.6))(ws@8.21.0)(zod@4.3.6)
'@earendil-works/pi-ai': 0.84.2(@modelcontextprotocol/sdk@1.29.0(zod@4.3.6))(ws@8.21.0)(zod@4.3.6)
'@earendil-works/pi-client': 0.84.2
'@earendil-works/pi-protocol': 0.84.2
'@earendil-works/pi-tui': 0.84.2
'@silvia-odwyer/photon-node': 0.3.4
chalk: 5.6.2
cross-spawn: 7.0.6
diff: 8.0.4
glob: 13.0.6
grok-mermaid: 0.2.2
highlight.js: 10.7.3
hosted-git-info: 9.0.3
ignore: 7.0.5
@@ -2481,8 +2498,8 @@ snapshots:
minimatch: 10.2.5
proper-lockfile: 4.1.2
semver: 7.8.0
typebox: 1.1.38
undici: 8.5.0
typebox: 1.3.7
undici: 8.9.0
yaml: 2.9.0
optionalDependencies:
'@mariozechner/clipboard': 0.3.9
@@ -2494,7 +2511,13 @@ snapshots:
- ws
- zod
'@earendil-works/pi-tui@0.82.1':
'@earendil-works/pi-protocol@0.84.2':
dependencies:
typebox: 1.3.7
'@earendil-works/pi-telemetry@0.84.2': {}
'@earendil-works/pi-tui@0.84.2':
dependencies:
get-east-asian-width: 1.6.0
marked: 18.0.5
@@ -2528,10 +2551,10 @@ snapshots:
- supports-color
- utf-8-validate
'@gotgenes/pi-permission-system@10.9.0(@earendil-works/pi-coding-agent@0.82.1(@modelcontextprotocol/sdk@1.29.0(zod@4.3.6))(ws@8.21.0)(zod@4.3.6))(@earendil-works/pi-tui@0.82.1)':
'@gotgenes/pi-permission-system@10.9.0(@earendil-works/pi-coding-agent@0.84.2(@modelcontextprotocol/sdk@1.29.0(zod@4.3.6))(ws@8.21.0)(zod@4.3.6))(@earendil-works/pi-tui@0.84.2)':
dependencies:
'@earendil-works/pi-coding-agent': 0.82.1(@modelcontextprotocol/sdk@1.29.0(zod@4.3.6))(ws@8.21.0)(zod@4.3.6)
'@earendil-works/pi-tui': 0.82.1
'@earendil-works/pi-coding-agent': 0.84.2(@modelcontextprotocol/sdk@1.29.0(zod@4.3.6))(ws@8.21.0)(zod@4.3.6)
'@earendil-works/pi-tui': 0.84.2
tree-sitter-bash: 0.25.1
web-tree-sitter: 0.26.9
transitivePeerDependencies:
@@ -2746,18 +2769,6 @@ snapshots:
'@mariozechner/clipboard-win32-x64-msvc': 0.3.9
optional: true
'@mistralai/mistralai@2.2.6(@opentelemetry/api@1.9.0)':
dependencies:
'@opentelemetry/semantic-conventions': 1.43.0
ws: 8.21.0
zod: 4.3.6
zod-to-json-schema: 3.25.2(zod@4.3.6)
optionalDependencies:
'@opentelemetry/api': 1.9.0
transitivePeerDependencies:
- bufferutil
- utf-8-validate
'@modelcontextprotocol/sdk@1.29.0(zod@4.3.6)':
dependencies:
'@hono/node-server': 1.19.14(hono@4.12.14)
@@ -2792,8 +2803,6 @@ snapshots:
'@opentelemetry/api@1.9.0': {}
'@opentelemetry/semantic-conventions@1.43.0': {}
'@oxc-project/types@0.122.0': {}
'@protobufjs/aspromise@1.1.2': {}
@@ -3595,6 +3604,8 @@ snapshots:
graceful-fs@4.2.11: {}
grok-mermaid@0.2.2: {}
has-flag@4.0.0: {}
has-symbols@1.1.0:
@@ -3817,7 +3828,7 @@ snapshots:
wrappy: 1.0.2
optional: true
openai@6.26.0(ws@8.21.0)(zod@4.3.6):
openai@6.40.0(ws@8.21.0)(zod@4.3.6):
optionalDependencies:
ws: 8.21.0
zod: 4.3.6
@@ -4209,6 +4220,8 @@ snapshots:
typebox@1.1.38: {}
typebox@1.3.7: {}
typescript@5.9.3: {}
unconfig-core@7.5.0:
@@ -4218,7 +4231,7 @@ snapshots:
undici-types@7.18.2: {}
undici@8.5.0: {}
undici@8.9.0: {}
unionfs@4.6.0:
dependencies:
@@ -4321,7 +4334,9 @@ snapshots:
zod-to-json-schema@3.25.2(zod@4.3.6):
dependencies:
zod: 4.3.6
optional: true
zod@4.3.6: {}
zod@4.3.6:
optional: true
zx@8.8.5: {}