Compare commits

..
Author SHA1 Message Date
ezl-keygraph 24862ddc9b chore: mark the pi harness migration as a breaking change
BREAKING CHANGE: Google Vertex AI is no longer a supported provider. The
CLAUDE_CODE_USE_VERTEX, ANTHROPIC_VERTEX_PROJECT, CLOUD_ML_REGION, and
GOOGLE_APPLICATION_CREDENTIALS environment variables, along with the
use_vertex, vertex_project, and cloud_ml_region config.toml keys, are
removed. Vertex users must switch to Anthropic, AWS Bedrock, or a custom
Anthropic-compatible base URL.

The CLAUDE_CODE_MAX_OUTPUT_TOKENS environment variable and the
max_output_tokens config.toml key are also removed.
2026-07-16 19:07:21 +05:30
ezl-keygraph b24caf8ae2 style(cli): collapse usage hint now that the beta tag is gone 2026-07-16 18:58:27 +05:30
ezl-keygraph 934a0ddd79 Merge remote-tracking branch 'origin/main' into feat/pi-harness-migration
# Conflicts:
#	CLAUDE.md
#	README.md
#	apps/cli/src/commands/start.ts
#	apps/cli/src/commands/uninstall.ts
#	apps/worker/src/services/agent-execution.ts
#	apps/worker/src/services/preflight.ts
#	apps/worker/src/session-manager.ts
#	apps/worker/src/temporal/activities.ts
#	apps/worker/src/temporal/shared.ts
#	apps/worker/src/temporal/workflows.ts
#	docs/ai-providers.md
#	llms-full.txt
2026-07-16 18:58:18 +05:30
ezl-keygraph d0b0ec3378 refactor(worker): converge shared core with shannon-oss (#388)
* fix(worker): port keygraph shared-core correctness fixes

* refactor(worker): adopt collectors/ and ai/pi/ layout; add task budget cap and cancellation

* refactor(worker): drop inconsistent Collector "Server" suffix

* refactor(worker): drop unused providerConfig/apiKey seams, resolve credentials from env only

* refactor(worker): port oss code_path pattern expansion + external_directory allow

* fix(worker): preserve dotfile paths in code_path avoid patterns (.env no longer stripped to env)

* feat(worker): render Unprocessed Vulnerabilities section in exploit deliverable (align with oss)

* feat(worker): request set_blind_spots for all vuln classes (align auth/ssrf with production prompts)

* refactor(worker): adopt unified permissionSystem* naming and helper layout

* refactor(worker): inline blind_spots into vuln deliverable section array

* chore(worker): drop unused zod dependency (tree is typebox-native)

* fix(worker): normalize base32 TOTP secret to accept padding and whitespace

* refactor(worker): adopt shared toolResult helper and flatSchema naming in collectors

* refactor(worker): use undefined over null in queue-schema builders

* docs(worker): converge renderer/collector doc comments to current pi terminology

* refactor(worker): adopt schema.ts cleanInput/stringEnum helpers in collectors

* feat(worker): converge exploit-collector/renderer with vendored; capture and render overview for blocked findings

* refactor(worker): converge session-tools/pipeline/exploitation-checker with vendored

* refactor(worker): converge task-tool usage reporting with vendored onUsage callback

* refactor(worker): converge structured output onto a submitTool executor channel

* docs(worker): expand exploit-renderer docstring to match shannon-oss

* docs(worker): adopt richer vuln-renderer docstring from shannon-oss

* docs(worker): neutralize billing-detection wording for shannon-oss parity

* fix(worker): verify checkpoint hash in the deliverables clone being reset

* fix(worker): fail fast on malformed exploitation queue JSON

* fix(worker): honor retryable flag when classifying exploitation-queue check failures

* fix(worker): fail fast on corrupted session.json in run-scope validation

* feat(worker): propagate Temporal cancellation signal into agent and auth pi sessions

* fix(worker): mark exploit agent complete when exploitation is skipped so resume skips it

* prompts: drop scan description from executive report prompt

* refactor(worker): add createGenericSubmitTool for raw JSON-schema submit tools

* refactor(worker): gate playwright-cli skill to browser agents via skillsOverride (adopt shannon-oss mechanism)

* docs(worker): correct formatLogTime comment to UTC to match toISOString

* refactor(worker): converge queue-schemas with shannon-oss (guarded count, decl order)

* refactor(worker): converge task-tool with shannon-oss (byte-identical; modelRegistry optional)

* fix(worker): use replaceLiteral for all prompt value insertions to prevent $-mangling

* fix(worker): classify agent execution failures by error type instead of hardcoding validation

* fix(worker): cap auth-failure detail at 250 chars to match shannon-oss

* style(worker): apply biome formatting

* refactor(worker): remove per-session task delegation cap from task tool
2026-07-16 14:09:02 +05:30
ezl-keygraph b845936aa5 feat(cli): restructure run folder and improve terminal UX (#383)
* feat: surface report at run root and nest run internals under .shannon

* feat: use plain-language wording in user-facing terminal messages

* feat(cli): guide users to watch scan progress and surface report path on start

* docs: sync run-folder layout and CLI wording across docs and comments

* feat(cli): add version command reporting package version or git SHA

* feat(cli): detect TTY for interactive prompts, color, and progress output

* docs: document --yes flag, version command, and tty module

* fix(cli): FORCE_COLOR precedence and plain uninstall --yes output

* fix(cli): respect empty NO_COLOR

* fix(cli): let NO_COLOR take precedence over FORCE_COLOR

* docs: mark claude-code-router integration as removed
2026-07-04 21:14:00 +05:30
ezl-keygraph 5596411bd3 fix: render agent deliverables before the success commit so resume preserves them (#377) 2026-06-23 14:25:17 +05:30
ezl-keygraph 6a86b6c4c3 fix(cli): pin npx command hints to beta tag 2026-06-17 18:30:02 +05:30
ezl-keygraph fb14a0170a ci: bump the beta release line to 2.0.0 (#356) 2026-06-17 18:09:27 +05:30
ezl-keygraph cf396fb9c7 feat(worker): enforce bounded bash timeouts via pi extension 2026-06-16 12:48:32 +05:30
ezl-keygraph f97afb482e refactor(worker): unify provider precedence between preflight and executor 2026-06-15 23:06:48 +05:30
ezl-keygraph c2bceba95c docs(worker): update stale sdk comments 2026-06-15 22:50:44 +05:30
ezl-keygraph 7c20384991 docs: remove vertex references from llms context 2026-06-15 22:48:59 +05:30
ezl-keygraph 0bc004a583 build: drop @anthropic-ai/claude-code from worker image 2026-06-15 22:42:12 +05:30
ezl-keygraph d3beea504a refactor(cli): remove CLAUDE_CODE_MAX_OUTPUT_TOKENS config 2026-06-15 22:40:50 +05:30
ezl-keygraph f46243a35a feat(worker): load playwright-cli skill via pi resource loader 2026-06-15 22:37:36 +05:30
ezl-keygraph 09e11b3ad9 fix(worker): restore minLength/minItems on pre-recon and exploit collector schemas 2026-06-15 21:06:29 +05:30
ezl-keygraph e16dcba13f refactor(prompts): drop collector server names from deliverable instructions 2026-06-15 20:21:22 +05:30
ezl-keygraph 5547afa73f refactor(prompts): drop stale MCP terminology for collector tools 2026-06-15 20:18:53 +05:30
ezl-keygraph 667e6ac4b0 refactor(prompts): use pi tool names (task, todo_write, read, bash, glob) 2026-06-15 20:03:26 +05:30
ezl-keygraph d18e928a6a feat(worker): add glob custom tool and route code_path globs to it 2026-06-15 20:03:26 +05:30
ezl-keygraph 58d0defea7 feat(worker): give task sub-agent write+bash, align tool descriptions 2026-06-15 19:54:20 +05:30
ezl-keygraph 9e845159b3 fix(worker): restore minLength/minItems on vuln-collector schemas 2026-06-15 18:42:53 +05:30
ezl-keygraph 0fd2f6bbe4 fix(worker): gate adaptive thinking to Opus models, drop CLAUDE_THINKING_LEVEL 2026-06-15 18:11:46 +05:30
ezl-keygraph 575465a741 feat(worker): pi-event-driven output formatting 2026-06-15 16:16:46 +05:30
ezl-keygraph 263b18e98a refactor(worker): rename claude-executor to pi-executor 2026-06-15 16:05:31 +05:30
ezl-keygraph 56241625a4 fix(worker): count sub-agent cost and surface compaction failures 2026-06-15 15:59:55 +05:30
ezl-keygraph 79fb49c159 feat(prompts): instruct agents to call submit_exploitation_queue and submit_auth_result 2026-06-15 15:49:02 +05:30
ezl-keygraph c275b27a6c fix(worker): route Bedrock and custom-base-URL providers from env 2026-06-15 15:36:14 +05:30
ezl-keygraph a9e966026c feat: remove Google Vertex AI provider support 2026-06-15 12:49:40 +05:30
ezl-keygraph 1908156525 feat(worker): migrate agent runtime from Claude Agent SDK to pi harness 2026-06-15 12:05:32 +05:30
103 changed files with 2337 additions and 7304 deletions

No files matched your search

+42 -51
View File
@@ -1,55 +1,46 @@
# Copy to .env and uncomment one provider block.
# SHANNON_AI_MODEL is <provider>:<model-id>, split on the first colon.
# Defaults to anthropic:claude-sonnet-4-6.
# Shannon Environment Configuration
# Copy this file to .env and fill in your credentials
# --- Anthropic ---------------------------------------------------------------
SHANNON_AI_API_KEY=your-api-key-here
SHANNON_AI_MODEL=anthropic:claude-sonnet-4-6
# CLAUDE_CODE_OAUTH_TOKEN=your-oauth-token-here
# Adaptive thinking is enabled automatically on Opus 4.6/4.7/4.8. Set to false to disable.
# CLAUDE_ADAPTIVE_THINKING=false
# --- OpenAI ------------------------------------------------------------------
# SHANNON_AI_API_KEY=your-api-key-here
# SHANNON_AI_MODEL=openai:gpt-5.5
# --- xAI ---------------------------------------------------------------------
# SHANNON_AI_API_KEY=your-api-key-here
# SHANNON_AI_MODEL=xai:grok-4.5
# --- AWS Bedrock -------------------------------------------------------------
# Bearer token only; model must be enabled in your region.
# AWS_REGION=us-east-1
# AWS_BEARER_TOKEN_BEDROCK=your-bearer-token
# SHANNON_AI_MODEL=amazon-bedrock:us.anthropic.claude-opus-4-8
# --- Custom Base URL ---------------------------------------------------------
# Route through a proxy or gateway (LiteLLM, an internal endpoint).
# Pick the block matching the API dialect your gateway speaks, and uncomment all
# three lines. The provider prefix picks the dialect; the model id is whatever
# name your gateway serves it under.
# Anthropic compatible - Anthropic Messages:
# SHANNON_AI_API_KEY=your-gateway-key-here
# SHANNON_AI_BASE_URL=https://llm-gateway.example.com
# SHANNON_AI_MODEL=anthropic:claude-sonnet-4-6
# OpenAI compatible - Chat Completions (default) or Responses:
# SHANNON_AI_API_KEY=your-gateway-key-here
# SHANNON_AI_BASE_URL=https://llm-gateway.example.com/v1
# SHANNON_AI_MODEL=openai:gpt-5.5
# SHANNON_AI_OPENAI_FORMAT=responses
# --- Other provider ----------------------------------------------------------
# Any other provider the Pi harness supports. Name it in SHANNON_AI_MODEL and
# supply the key via the generic SHANNON_AI_API_KEY. Pi validates the provider
# and model at preflight.
# SHANNON_AI_MODEL=openrouter:moonshotai/kimi-k3
# SHANNON_AI_API_KEY=your-api-key-here
# --- Misc --------------------------------------------------------------------
# Forward /etc/hosts entries into the worker container.
# Shannon forwards your machine's /etc/hosts entries into the worker container. Set to false to disable.
# SHANNON_FORWARD_HOSTS=false
# See the guide below to use an OpenAI subscription
# https://github.com/KeygraphHQ/shannon/blob/main/docs/ai-providers.md#openai-codex-chatgpt-pluspro-subscription
# SHANNON_USE_PI_AUTH=1
# SHANNON_AI_MODEL=openai-codex:gpt-5.5
# =============================================================================
# OPTION 1: Direct Anthropic
# =============================================================================
ANTHROPIC_API_KEY=your-api-key-here
# OR use OAuth token instead
# CLAUDE_CODE_OAUTH_TOKEN=your-oauth-token-here
# =============================================================================
# OPTION 2: Custom Base URL (compatible proxies, gateways, etc.)
# =============================================================================
# Point the SDK at an alternative Anthropic-compatible endpoint.
# ANTHROPIC_BASE_URL=https://your-proxy.example.com
# ANTHROPIC_AUTH_TOKEN=your-auth-token # Auth token for the custom endpoint
# =============================================================================
# Model Tier Overrides (Anthropic API / OAuth / Custom Base URL / Bedrock)
# =============================================================================
# Override which model is used for each tier. Defaults are used if not set.
# Optional for direct Anthropic and custom base URL modes. Required for Bedrock.
# ANTHROPIC_SMALL_MODEL=... # Small tier (default: claude-haiku-4-5-20251001)
# ANTHROPIC_MEDIUM_MODEL=... # Medium tier (default: claude-sonnet-4-6)
# ANTHROPIC_LARGE_MODEL=... # Large tier (default: claude-opus-4-8)
# =============================================================================
# OPTION 3: AWS Bedrock
# =============================================================================
# https://aws.amazon.com/blogs/machine-learning/accelerate-ai-development-with-amazon-bedrock-api-keys/
# Requires the model tier overrides above to be set with Bedrock-specific model IDs.
# Example Bedrock model IDs for us-east-1:
# ANTHROPIC_SMALL_MODEL=us.anthropic.claude-haiku-4-5-20251001-v1:0
# ANTHROPIC_MEDIUM_MODEL=us.anthropic.claude-sonnet-4-6
# ANTHROPIC_LARGE_MODEL=us.anthropic.claude-opus-4-8
# CLAUDE_CODE_USE_BEDROCK=1
# AWS_REGION=us-east-1
# AWS_BEARER_TOKEN_BEDROCK=your-bearer-token
+2 -5
View File
@@ -117,12 +117,9 @@ body:
options:
- "Anthropic (API key)"
- "Anthropic (OAuth token)"
- "OpenAI"
- "xAI"
- "Custom base URL (proxy/gateway)"
- "AWS Bedrock"
- "Custom base URL - Anthropic Messages"
- "Custom base URL - OpenAI Chat Completions"
- "Custom base URL - OpenAI Responses"
- "Google Vertex AI"
validations:
required: true
+22 -25
View File
@@ -44,8 +44,8 @@ echo "ANTHROPIC_API_KEY=your-key" > .env
./shannon build
# Run
./shannon start -u <url> -r ./my-repo
./shannon start -u <url> -r ./my-repo -c ./apps/worker/configs/my-config.yaml
./shannon start -u <url> -r my-repo
./shannon start -u <url> -r my-repo -c ./apps/worker/configs/my-config.yaml
./shannon start -u <url> -r /any/path/to/repo
```
@@ -56,24 +56,25 @@ echo "ANTHROPIC_API_KEY=your-key" > .env
npx @keygraph/shannon setup
# Workspaces & Resume
./shannon start -u <url> -r ./my-repo -w my-audit # New named workspace
./shannon start -u <url> -r ./my-repo -w my-audit # Resume (same command)
./shannon start -u <url> -r my-repo -w my-audit # New named workspace
./shannon start -u <url> -r my-repo -w my-audit # Resume (same command)
./shannon workspaces # List all workspaces
# Monitor
./shannon logs <workspace> # Show a scan's live log
./shannon status <workspace> # Live phase/agent progress of one scan, read from Temporal (redraws, then exits)
./shannon status # Show running scans
# Dashboard: http://localhost:8233
# Stop
./shannon stop <workspace> # Stop one scan (confirms first; --yes/-y to skip)
./shannon stop --all # Stop all running scans (Temporal stays up; confirms first)
./shannon reset # Stop everything and wipe all Temporal data + volumes (type 'confirm' to proceed; cannot be skipped)
./shannon stop # Preserves scan data
./shannon stop --clean # Full cleanup including volumes (confirms first; --yes/-y to skip)
# Version
./shannon version # npx: package version; local: git SHA
# Image management
./shannon build [--no-cache] # Local mode: build worker image
npx @keygraph/shannon uninstall # npx mode: remove ~/.shannon/ (confirms first; --yes/-y to skip)
# Build TypeScript (development)
pnpm run build # Build all packages via Turborepo
@@ -84,7 +85,7 @@ pnpm biome:fix # Auto-fix lint, format, and import sorting
**Monorepo tooling:** pnpm workspaces, Turborepo for task orchestration, Biome for linting/formatting. TypeScript compiler options shared via `tsconfig.base.json` at the root. All packages extend it, overriding only `rootDir` and `outDir`. Shared devDependencies (`typescript`, `@types/node`, `turbo`, `@biomejs/biome`) are hoisted to the root workspace.
**Options:** `-c <file>` (YAML config), `-o <path>` (output directory), `-w <name>` (named workspace; auto-resumes if exists), `--pipeline-testing` (minimal prompts, 10s retries), `--keep-container` (preserve worker container after exit for log inspection), `--yes`/`-y` (skip the confirmation prompt on `stop`; required for non-interactive use; `reset` requires a typed `confirm` and cannot be skipped)
**Options:** `-c <file>` (YAML config), `-o <path>` (output directory), `-w <name>` (named workspace; auto-resumes if exists), `--pipeline-testing` (minimal prompts, 10s retries), `--debug` (preserve worker container after exit for log inspection), `--yes`/`-y` (skip the confirmation prompt on `stop --clean`/`uninstall`; required for non-interactive use)
## Architecture
@@ -96,20 +97,17 @@ apps/worker/ — @shannon/worker (private, Temporal worker + pipeline logic)
```
### CLI Package (`apps/cli/`)
Published as `@keygraph/shannon` on npm. Contains Docker orchestration logic plus a read-only `@temporalio/client` reader (for `status`); no worker/pipeline business logic or prompts. Bundled with tsdown for single-file ESM output (deps stay external).
Published as `@keygraph/shannon` on npm. Contains only Docker orchestration logic — no Temporal SDK, business logic, or prompts. Bundled with tsdown for single-file ESM output.
- `apps/cli/src/index.ts` — CLI dispatcher (`setup`, `start`, `stop`, `reset`, `logs`, `status`, `build`, `version`)
- `apps/cli/src/temporal-client.ts` — `@temporalio/client` reader for `status`: connects to the frontend on `127.0.0.1:7233` (published by compose), `describeScan` (status + `pendingActivities` → running agents), `queryProgress` (live `getProgress` query → `PipelineState`), `getTerminalOutcome` (workflow `result()`). No worker of its own; scans are visible only within Temporal's ~24h retention (namespace default, unset in compose)
- `apps/cli/src/scan/` — `status` rendering: `pipeline.ts` (static phase/agent plan + `run*Agent` activity-type→agent map + mirrored `PipelineState`/`AgentMetrics` types; keep in sync with the worker), `render.ts` (one renderer for both the live query state and the terminal result)
- `apps/cli/src/index.ts` — CLI dispatcher (`setup`, `start`, `stop`, `logs`, `workspaces`, `status`, `build`, `uninstall`, `version`)
- `apps/cli/src/mode.ts` — Auto-detection: local mode if `SHANNON_LOCAL=1` env var is set
- `apps/cli/src/docker.ts` — Compose lifecycle, image pull/build, ephemeral `docker run` worker spawning
- `apps/cli/src/home.ts` — State directory management (`~/.shannon/` for npx, `./` for local)
- `apps/cli/src/env.ts` — `.env` loading, TOML fallback (npx only) via `apps/cli/src/config/resolver.ts`, credential validation, provider-scoped env flag building
- `apps/cli/src/model-spec.ts` — `SHANNON_AI_MODEL` (`<provider>:<model-id>`) parsing; mirrors `apps/worker/src/ai/models.ts`
- `apps/cli/src/env.ts` — `.env` loading, TOML fallback (npx only) via `apps/cli/src/config/resolver.ts`, credential validation, env flag building
- `apps/cli/src/config/resolver.ts` — Cascading config (npx only): env vars → `~/.shannon/config.toml` (parsed with `smol-toml`)
- `apps/cli/src/config/writer.ts` — TOML serialization and secure file persistence (0o600)
- `apps/cli/src/commands/setup.ts` — Interactive TUI wizard (`@clack/prompts`) for provider credential setup (npx only)
- `apps/cli/src/paths.ts` — Repo/config path resolution (any absolute or relative path)
- `apps/cli/src/paths.ts` — Repo/config path resolution (bare name → `./repos/<name>`, or any absolute/relative path)
- `apps/cli/src/version.ts` — Version reporting (npx: `package.json` version; local: `git-<sha>`)
- `apps/cli/src/tty.ts` — Terminal capability detection: `requireInteractive` guard (fails fast off-TTY instead of hanging on a prompt), `supportsColor` color gating (`NO_COLOR`/`FORCE_COLOR`), and `stdoutIsTerminal` for spinner/cursor output
- `apps/cli/src/commands/` — Command handlers
@@ -129,7 +127,7 @@ Infra (Temporal) runs via `docker-compose.yml`. Workers are ephemeral `docker ru
- `apps/worker/src/paths.ts` — Centralized path constants (`PROMPTS_DIR`, `CONFIGS_DIR`, `WORKSPACES_DIR`)
- `apps/worker/src/session-manager.ts` — Agent definitions (`AGENTS` record). Agent types in `apps/worker/src/types/agents.ts`
- `apps/worker/src/config-parser.ts` — YAML config parsing with JSON Schema validation
- `apps/worker/src/ai/pi/pi-executor.ts` — pi harness integration (agent-level retry disabled so Temporal owns restarts; provider-level retry on, see `apps/worker/src/ai/pi/retry-settings.ts`)
- `apps/worker/src/ai/pi-executor.ts` — pi harness integration (retry disabled; Temporal owns retry)
- `apps/worker/src/services/` — Business logic layer (Temporal-agnostic). Activities delegate here. Key: `agent-execution.ts`, `error-handling.ts`, `container.ts`
- `apps/worker/src/types/` — Consolidated types: `Result<T,E>`, `ErrorCode`, `AgentName`, `ActivityLogger`, etc.
- `apps/worker/src/utils/` — Shared utilities (file I/O, formatting, concurrency)
@@ -152,13 +150,12 @@ Durable workflow orchestration with crash recovery, queryable progress, intellig
5. **Reporting** (`report`) — Executive-level security report
### Supporting Systems
- **Configuration** — YAML configs in `apps/worker/configs/` with JSON Schema validation (`config-schema.json`). Supports auth settings (MFA/TOTP), URL/code rule scoping (`rules.avoid`/`rules.focus`), run-scope steering (`vuln_classes`, `exploit`), free-form `rules_of_engagement`, and post-hoc `report` options (`min_severity`, `min_confidence`, `guidance`, and `sarif` to emit a SARIF 2.1.0 log via `apps/worker/src/services/sarif-renderer.ts`; exploit-only). `code_path` avoid rules are enforced via the `@gotgenes/pi-permission-system` extension: `apps/worker/src/temporal/activities.ts:syncCodePathDenyRules` writes a global `path` deny config once per workflow (`apps/worker/src/ai/pi/permission-system.ts:syncPermissionSystemConfig`), and the executor loads the extension when that config is present (`apps/worker/src/ai/pi/pi-executor.ts`), so denies fire across every tool and child `task` session. `vuln_classes`/`exploit` scope is locked into `session.json` on first run; resumes with a different scope fail fast (`persistOrValidateRunScope`). Credential resolution — local mode: env vars → `./.env`; npx mode: env vars → `~/.shannon/config.toml` (via `npx @keygraph/shannon setup`)
- **Configuration** — YAML configs in `apps/worker/configs/` with JSON Schema validation (`config-schema.json`). Supports auth settings (MFA/TOTP), URL/code rule scoping (`rules.avoid`/`rules.focus`), run-scope steering (`vuln_classes`, `exploit`), free-form `rules_of_engagement`, and post-hoc `report` filters (`min_severity`, `min_confidence`, `guidance`). `code_path` avoid rules are enforced via the `@gotgenes/pi-permission-system` extension: `apps/worker/src/temporal/activities.ts:syncCodePathDenyRules` writes a global `path` deny config once per workflow (`apps/worker/src/ai/settings-writer.ts:writeCodePathPermissionConfig`), and the executor loads the extension when that config is present (`apps/worker/src/ai/pi-executor.ts`), so denies fire across every tool and child `task` session. `vuln_classes`/`exploit` scope is locked into `session.json` on first run; resumes with a different scope fail fast (`persistOrValidateRunScope`). Credential resolution — local mode: env vars → `./.env`; npx mode: env vars → `~/.shannon/config.toml` (via `npx @keygraph/shannon setup`)
- **Prompts** — Per-phase templates in `apps/worker/prompts/` with variable substitution (`{{TARGET_URL}}`, `{{CONFIG_CONTEXT}}`). Shared partials in `apps/worker/prompts/shared/` via `apps/worker/src/services/prompt-manager.ts`, including `_code-path-rules.txt` (focus/avoid `[FILE]`/`[GLOB]` routing) and `_rules-of-engagement.txt` (free-text engagement rules). When `exploit: false`, `apps/worker/src/services/findings-renderer.ts` deterministically converts each `*_exploitation_queue.json` into a `*_findings.md` for report assembly — no LLM in the loop
- **Agent Harness (pi)** — Uses the **pi harness** (`@earendil-works/pi-coding-agent`, requires Node ≥ 22.19) via `apps/worker/src/ai/pi/pi-executor.ts` (`runPiPrompt` → `createAgentSession`). Retry is split in `apps/worker/src/ai/pi/retry-settings.ts`: pi's agent-level loop is off so Temporal owns agent restarts, while `provider.maxRetries` stays on — pi reads the `provider` block independently of the `enabled` flag — so transport faults are absorbed in-session rather than costing a full agent re-run. `maxRetryDelayMs` is left at pi's 60s default. One model runs every phase, named by `SHANNON_AI_MODEL=<provider>:<model-id>` (default `anthropic:claude-sonnet-4-6`). `apps/worker/src/ai/models.ts` parses the spec — splitting on the **first** colon only, so Bedrock IDs keep theirs — and resolves it through pi's `ModelRuntime`. pi ships the `CredentialStore` interface but no in-memory implementation (its own reads `auth.json` from disk), so `RuntimeCredentialStore` in that file supplies one: credentials arrive as env vars in an ephemeral container and must never touch disk. `createModelRuntime(providerId, apiKey)` builds the runtime; `allowModelNetwork` stays at its default `false` so a scan never blocks on a catalog refresh. `resolveModelSelection()` is **async** because `ModelRuntime.create()` is. Any pi-ai provider id is accepted — `parseModelSpec` no longer rejects against a hardcoded list, so pi's registry is the authority (an unknown provider/model surfaces as a clear "not found in pi registry" error at preflight, which points to the browsable catalogue at `pi.dev/models` — `PI_CATALOG_URL` in `apps/worker/src/ai/models.ts`, appended to the not-found errors and shown in the setup wizard's "Other provider" hint). Four providers are **curated** (`CURATED_PROVIDERS`: `anthropic`, `openai`, `xai`, `amazon-bedrock`) with their own credential variables, config sections, and setup flows; each provider's API key env var is declared once in `PROVIDER_API_KEY_ENV` — Shannon uses each vendor's own variable name (`OPENAI_API_KEY`, `XAI_API_KEY`, …), never an invented one; Bedrock's entry is `AWS_BEARER_TOKEN_BEDROCK`, paired with `AWS_REGION`, which preflight requires separately as provider config rather than a credential. Any other provider uses the **generic** credential path: `SHANNON_AI_API_KEY` (`GENERIC_API_KEY_ENV`) supplies the key for any provider whose credential is a plain API key. Curated providers' own variables take precedence over it, and it also works as a fallback for them — Bedrock is the sole exception (it authenticates through its AWS_ variables, so the generic key never stands in for it). The CLI forwards `SHANNON_AI_API_KEY` in `COMMON_FORWARD_VARS` (it is provider-neutral, binding to whatever `SHANNON_AI_MODEL` names, so the "only one provider configured" guard counts only named credentials), and stores it under a generic `[provider]` config.toml section (`provider.api_key`). `npx @keygraph/shannon setup` exposes this as the "Other provider" option: free-text provider id + model id + key (a curated provider id is rejected there, since it has its own option). `SHANNON_AI_BASE_URL` overrides the endpoint for any provider (proxies/gateways); the credential is unchanged. `pointAtGateway` (`apps/worker/src/ai/models.ts`) applies the one dialect change: behind a base URL, `openai` follows `SHANNON_AI_OPENAI_FORMAT` (`chat-completions` default, or `responses`). On `chat-completions` it switches the API to `openai-completions` and drops the catalogue's Responses-shaped `compat` block so pi's `detectCompat` derives completions settings; on `responses` the descriptor is unchanged but for the endpoint. `resolveGatewayFormat` rejects the variable when the provider is not `openai` or no base URL is set, since it cannot take effect there. All other providers keep their API. The CLI mirrors the accepted values in `apps/cli/src/model-spec.ts`, forwards the variable in `COMMON_FORWARD_VARS`, and maps it to `openai.format` in config.toml. `buildEnvFlags` forwards only the selected provider's credential into the worker container. The CLI mirrors the parse rule and the provider/credential tables in `apps/cli/src/model-spec.ts` (it cannot import from the worker package); the two must stay in sync. pi ships no JSON-schema output or `Task`/`TodoWrite` built-ins, so structured queues are captured via a `submit_exploitation_queue` custom tool (`apps/worker/src/ai/queue-schemas.ts`), and `task` (child sessions scoped to `read`, `grep`, `find`, `ls`, `write`, and `bash` — no nested `task` or collector tools; `CHILD_TOOLS` in `apps/worker/src/ai/pi/task-tool.ts`) + `todo_write` (`apps/worker/src/ai/pi/session-tools.ts`) are provided as custom tools; the per-phase collectors are pi custom tools (TypeBox `defineTool` in `apps/worker/src/collectors/`). Shannon sets no thinking configuration at all — no `thinkingLevel` is passed to any `createAgentSession` call, so pi's own default applies. There Line truncated
- **Pi Credential Reuse** — `SHANNON_USE_PI_AUTH=1` opts into reusing the host's Pi login, including an `openai-codex` ChatGPT Plus/Pro subscription selected with `SHANNON_AI_MODEL=openai-codex:<model-id>`. `apps/cli/src/env.ts` requires `~/.pi/agent/auth.json`; `start.ts` passes its path to `spawnWorker`, which mounts only that file read-write at `/tmp/.pi/agent/auth.json`. The flag itself is not forwarded: the worker detects the file with `piAuthPresent()` and passes its path to `ModelRuntime.create`. CLI and worker API-key presence checks are skipped on this path, but the normal preflight model probe still validates the credential. The image and UID-remapping entrypoint keep `/tmp/.pi/agent` owned by `pentest` so adjacent Pi/Shannon configuration remains writable. Refreshed OAuth state is persisted to the host for subsequent scans.
- **Audit System** — Crash-safe append-only logging in `workspaces/{hostname}_{sessionId}/`. The run directory's top level holds the human-facing report in both formats (`Security-Assessment-Report.pdf` and `Security-Assessment-Report.md`, `FINAL_REPORT_PDF_FILENAME`/`FINAL_REPORT_MD_FILENAME` in `apps/worker/src/paths.ts`); everything else — deliverables, per-agent logs, prompts, `session.json`, `workflow.log`, and browser artifacts — is nested under a hidden `.shannon/` internals dir (`INTERNAL_DIR`) so a customer sees only the report. Audit path helpers route through `generateInternalPath` (`apps/worker/src/audit/utils.ts`); the CLI nests the overlay backing dirs under the same `.shannon/` (`apps/cli/src/docker.ts`, `start.ts`). `session.json`/`workflow.log` reads use dual-read resolvers (`resolveSessionJsonPath`, `resolveRunFile`) that prefer `.shannon/` and fall back to the legacy run-root layout, so pre-restructure workspaces stay listable (`workspaces`/`logs`) without migration. Resuming a pre-restructure workspace upgrades it in place first: `migrateLegacyWorkspaceLayout` (`apps/cli/src/commands/start.ts`) renames the flat deliverables/logs/session entries into `.shannon/` (carrying the deliverables `.git` along) before the overlay dirs are mounted, so resume finds the old checkpoints instead of re-running every agent. The report agent writes structured findings to `report.json`, from which `report-renderer.ts` renders the assembled markdown and `report-json-adapter.ts` produces the Typst-shaped JSON that `pdf-renderer.ts` compiles into `comprehensive_security_assessment_report.pdf` using the bundled `apps/worker/templates/typst/report.typ` template (the `typst` binary is installed in the worker image). `copyReportToRunRoot` (`apps/worker/src/services/reporting.ts`) surfaces both the PDF and the markdown to the run root as `Security-Assessment-Report.pdf` and `Security-Assessment-Report.md`; the deliverables-dir copies remain as the git-checkpointed sources. PDF compilation is best-effort — a failure is logged and the run still completes. WorkflowLogger (`apps/worker/src/audit/workflow-logger.ts`) provides unified human-readable per-workflow logs, backed by LogStream (`apps/worker/src/audit/log-stream.ts`) shared stream primitive
- **Agent Harness (pi)** — Uses the **pi harness** (`@earendil-works/pi-coding-agent`, requires Node ≥ 22.19) via `apps/worker/src/ai/pi-executor.ts` (`runPiPrompt` → `createAgentSession`, retry disabled so Temporal owns retry). Models resolve through pi-ai in `apps/worker/src/ai/models.ts` (Anthropic / Bedrock / custom base URL via `ModelRegistry`+`AuthStorage`). pi ships no JSON-schema output or `Task`/`TodoWrite` built-ins, so structured queues are captured via a `submit_exploitation_queue` custom tool (`apps/worker/src/ai/queue-schemas.ts`), and `task` (read-only child sessions) + `todo_write` are provided as custom tools (`apps/worker/src/ai/tools.ts`); the per-phase MCP collectors are pi custom tools (TypeBox `defineTool` in `apps/worker/src/mcp-server/`). Adaptive thinking (pi's `medium` level) is enabled only on Opus 4.6/4.7/4.8 (`supportsAdaptiveThinking`); every other model runs with thinking `off`. Disable per-scan via `CLAUDE_ADAPTIVE_THINKING=false` (→ `off`) / `core.adaptive_thinking = false` (npx TOML). Browser automation via `playwright-cli` with session isolation (`-s=<session>`). TOTP generation via `generate-totp` CLI tool. Login flow template at `apps/worker/prompts/shared/login-instructions.txt` supports form, SSO, API, and basic auth. On authenticated whitebox scans, the `validate-authentication` preflight performs the single real login and saves the browser session to `auth-state.json` in the per-session audit directory (path from `authStateFile()` in `apps/worker/src/audit/utils.ts`, derived from `generateAuditPath()`). The validation activity (`apps/worker/src/services/validate-authentication.ts`) removes any stale file from a prior run before the agent runs and verifies the file parses and contains cookies or storage before the preflight is marked complete; `logWorkflowComplete` deletes it when the workflow ends so authenticated cookies don't sit on disk between scans. Agent prompts opt in to session reuse by `@include(shared/_shared-session.txt)` before their `<login_instructions>` block — the partial restores the session and falls through to the full login flow if verification fails. `vuln-auth`/`exploit-auth` omit the include and own their own login
- **Audit System** — Crash-safe append-only logging in `workspaces/{hostname}_{sessionId}/`. The run directory's top level holds only the human-facing report (`Security-Assessment-Report.md`, `FINAL_REPORT_FILENAME` in `apps/worker/src/paths.ts`); everything else — deliverables, per-agent logs, prompts, `session.json`, `workflow.log`, and browser artifacts — is nested under a hidden `.shannon/` internals dir (`INTERNAL_DIR`) so a customer sees only the report. Audit path helpers route through `generateInternalPath` (`apps/worker/src/audit/utils.ts`); the CLI nests the overlay backing dirs under the same `.shannon/` (`apps/cli/src/docker.ts`, `start.ts`). `session.json`/`workflow.log` reads use dual-read resolvers (`resolveSessionJsonPath`, `resolveRunFile`) that prefer `.shannon/` and fall back to the legacy run-root layout, so pre-restructure workspaces stay listable (`workspaces`/`logs`) without migration. Resuming a pre-restructure workspace upgrades it in place first: `migrateLegacyWorkspaceLayout` (`apps/cli/src/commands/start.ts`) renames the flat deliverables/logs/session entries into `.shannon/` (carrying the deliverables `.git` along) before the overlay dirs are mounted, so resume finds the old checkpoints instead of re-running every agent. The report is surfaced by copying the assembled `comprehensive_security_assessment_report.md` from the deliverables dir to the run root (`copyReportToRunRoot` in `apps/worker/src/services/reporting.ts`). WorkflowLogger (`apps/worker/src/audit/workflow-logger.ts`) provides unified human-readable per-workflow logs, backed by LogStream (`apps/worker/src/audit/log-stream.ts`) shared stream primitive
- **Deliverables** — Saved to `.shannon/deliverables/` in the target repo via the `save-deliverable` CLI script (`apps/worker/src/scripts/save-deliverable.ts`)
- **Workspaces & Resume** — Named workspaces via `-w <name>` or auto-named from URL+timestamp. Resume detects completed agents via `session.json`. `loadResumeState()` in `apps/worker/src/temporal/activities.ts` validates deliverable existence, restores git checkpoints, and cleans up incomplete deliverables
- **Workspaces & Resume** — Named workspaces via `-w <name>` or auto-named from URL+timestamp. Resume detects completed agents via `session.json`. `loadResumeState()` in `apps/worker/src/temporal/activities.ts` validates deliverable existence, restores git checkpoints, and cleans up incomplete deliverables. Workspace listing via `apps/worker/src/temporal/workspaces.ts`
## Development Notes
@@ -236,7 +233,7 @@ Comments must be **timeless** — no references to this conversation, refactorin
**Entry Points:** `apps/worker/src/temporal/workflows.ts`, `apps/worker/src/temporal/activities.ts`, `apps/worker/src/temporal/worker.ts`
**Core Logic:** `apps/worker/src/session-manager.ts`, `apps/worker/src/ai/pi/pi-executor.ts`, `apps/worker/src/ai/pi/permission-system.ts` (writes `code_path` deny rules to the `@gotgenes/pi-permission-system` global config), `apps/worker/src/config-parser.ts`, `apps/worker/src/services/` (incl. `preflight.ts`, `findings-renderer.ts`, `reporting.ts`), `apps/worker/src/audit/`
**Core Logic:** `apps/worker/src/session-manager.ts`, `apps/worker/src/ai/pi-executor.ts`, `apps/worker/src/ai/settings-writer.ts` (writes `code_path` deny rules to the `@gotgenes/pi-permission-system` global config), `apps/worker/src/config-parser.ts`, `apps/worker/src/services/` (incl. `preflight.ts`, `findings-renderer.ts`, `reporting.ts`), `apps/worker/src/audit/`
**Config:** `docker-compose.yml`, `apps/cli/infra/compose.yml`, `apps/worker/configs/`, `apps/worker/prompts/`, `tsconfig.base.json` (shared compiler options), `turbo.json`, `biome.json`
@@ -248,9 +245,9 @@ Package managers are configured with a minimum release age (7 days). Requires pn
## Troubleshooting
- **"Repository not found"** — Pass a path to the target repo (`-r /path/to/repo` or `-r ./my-repo`)
- **"Repository not found"** — Pass a bare name (`-r my-repo`) for `./repos/my-repo`, or a path (`-r /path/to/repo`) for any directory
- **"Temporal not ready"** — Wait for health check or `docker compose logs temporal`
- **Worker not processing** — Check `docker ps --filter "name=shannon-worker-"`
- **Reset state** — `./shannon reset`
- **Reset state** — `./shannon stop --clean`
- **Local apps unreachable** — Use `host.docker.internal` instead of `localhost`
- **Container permissions** — On Linux, may need `sudo` for docker commands
+3 -23
View File
@@ -52,8 +52,6 @@ RUN apk update && apk add --no-cache \
curl \
ca-certificates \
shadow \
# Typst tarball decompression
xz \
# Language runtimes (minimal)
nodejs-22 \
npm \
@@ -75,22 +73,6 @@ RUN apk update && apk add --no-cache \
# Font rendering
fontconfig
# Install Typst (report PDF compilation)
ARG TYPST_VERSION=0.14.2
RUN case "$(uname -m)" in \
x86_64) TYPST_ARCH=x86_64-unknown-linux-musl ;; \
aarch64) TYPST_ARCH=aarch64-unknown-linux-musl ;; \
*) echo "unsupported arch $(uname -m)" && exit 1 ;; \
esac && \
mkdir -p /tmp/typst-dl /usr/local/bin && cd /tmp/typst-dl && \
curl -fsSL "https://github.com/typst/typst/releases/download/v${TYPST_VERSION}/typst-${TYPST_ARCH}.tar.xz" -o typst.tar.xz && \
xz -d typst.tar.xz && \
tar -xf typst.tar && \
mv "typst-${TYPST_ARCH}/typst" /usr/local/bin/typst && \
chmod +x /usr/local/bin/typst && \
cd / && rm -rf /tmp/typst-dl && \
typst --version
# Create non-root user
RUN addgroup -g 1001 pentest && \
adduser -u 1001 -G pentest -s /bin/bash -D pentest
@@ -119,18 +101,16 @@ RUN mkdir -p /tmp/.claude/skills && \
RUN ln -s /app/apps/worker/dist/scripts/save-deliverable.js /usr/local/bin/save-deliverable && \
chmod +x /app/apps/worker/dist/scripts/save-deliverable.js && \
ln -s /app/apps/worker/dist/scripts/generate-totp.js /usr/local/bin/generate-totp && \
chmod +x /app/apps/worker/dist/scripts/generate-totp.js && \
ln -s /app/apps/worker/dist/scripts/set-report-meta.js /usr/local/bin/set-report-meta && \
chmod +x /app/apps/worker/dist/scripts/set-report-meta.js
chmod +x /app/apps/worker/dist/scripts/generate-totp.js
# Create directories for session data and ensure proper permissions
RUN mkdir -p /app/sessions /app/repos /app/workspaces && \
mkdir -p /tmp/.cache /tmp/.config /tmp/.npm /tmp/.pi/agent && \
mkdir -p /tmp/.cache /tmp/.config /tmp/.npm && \
chmod 777 /app && \
chmod 777 /tmp/.cache && \
chmod 777 /tmp/.config && \
chmod 777 /tmp/.npm && \
chown -R pentest:pentest /app /tmp/.claude /tmp/.pi
chown -R pentest:pentest /app /tmp/.claude
COPY entrypoint.sh /app/entrypoint.sh
RUN chmod +x /app/entrypoint.sh
+8 -15
View File
@@ -1,5 +1,5 @@
> [!NOTE]
> **[Shannon 2.0 is officially here](https://github.com/KeygraphHQ/shannon/discussions/405)**
> **[Shannon Now Runs on the Pi Harness (Beta) - run it today with `npx @keygraph/shannon@beta`](https://github.com/KeygraphHQ/shannon/discussions/358)**
<div align="center">
@@ -9,7 +9,7 @@
<a href="https://trendshift.io/repositories/15604" target="_blank"><img src="https://trendshift.io/api/badge/repositories/15604" alt="KeygraphHQ%2Fshannon | Trendshift" style="width: 250px; height: 55px;" width="250" height="55"/></a>
Shannon is an autonomous, AI pentester for web applications and APIs. <br />
Shannon is an autonomous, white-box AI pentester for web applications and APIs. <br />
It analyzes your source code, identifies attack paths, and executes real exploits to prove vulnerabilities before they reach production.
**This repository is Shannon Open Source: the full agent, run locally from your command line.**
@@ -35,13 +35,13 @@ It analyzes your source code, identifies attack paths, and executes real exploit
- [Architecture](#architecture)
- [Documentation](#documentation)
- [Safety, Scope, and Limitations](#safety-scope-and-limitations)
- [License](#license)
- [License and Enterprise Licensing](#license-and-enterprise-licensing)
- [About Keygraph](#about-keygraph)
- [Community and Support](#community-and-support)
## What is Shannon?
Shannon is an autonomous AI pentester developed by [Keygraph](https://keygraph.io). It performs security testing of web applications and their underlying APIs by combining source-code analysis with live exploitation.
Shannon is an autonomous AI pentester developed by [Keygraph](https://keygraph.io). It performs white-box security testing of web applications and their underlying APIs by combining source-code analysis with live exploitation.
Shannon analyzes your web application's source code to identify potential attack vectors, then uses browser automation and command-line tools to execute real exploits against the running application and its APIs. Only vulnerabilities with a working proof-of-concept are included in the final report.
@@ -73,8 +73,7 @@ Sample penetration test reports from intentionally vulnerable applications, prod
- **Docker**: required for the worker container.
- **Node.js 18+**: required for the recommended `npx` workflow.
- **AI provider credentials**: Anthropic, OpenAI, xAI, or AWS Bedrock - or [any other provider](docs/ai-providers.md#any-other-provider). Claude models are recommended. For suggested model IDs per provider, plus gateways and custom base URLs, see [AI providers](docs/ai-providers.md#suggested-models).
- **Cyber safeguards cleared with your provider**: Anthropic and OpenAI apply real-time safeguards to cyber-security workloads, which can interrupt a scan mid-run. Complete their guidance for legitimate security testers before your first run - see [AI providers](docs/ai-providers.md#cyber-safeguards-do-this-before-your-first-scan).
- **AI provider credentials**: Anthropic is recommended. AWS Bedrock and compatible proxy setups are documented separately.
### Run Shannon
@@ -93,12 +92,6 @@ Shannon pulls the worker image from Docker Hub, starts the required local infras
For source builds, authenticated scans, provider-specific setup, and platform notes, see [Documentation](#documentation).
> [!TIP]
> **Prefer to use a subscription instead of API credits?**
>
> - **OpenAI Codex:** The latest version of Shannon supports ChatGPT Plus and Pro subscriptions. Follow the [OpenAI Codex subscription setup guide](docs/ai-providers.md#openai-codex-chatgpt-pluspro-subscription) to get started.
> - **Claude Code:** The latest version of Shannon does not support Claude Code subscriptions. Follow the [Claude Code subscription setup guide](docs/ai-providers.md#claude-code-subscription) to use version `1.9.0`, which is the final release built on the Claude Agent SDK.
## Key Capabilities
- **Proof-by-exploitation reports**: Shannon reports validated findings with reproducible proof-of-concept steps instead of speculative warnings.
@@ -192,8 +185,8 @@ Use these guides for operational detail:
| Guide | Use it for |
| --- | --- |
| [Source build and CLI commands](docs/development.md) | Cloning, building, common commands, output paths, and local development. |
| [Configuration](docs/configuration.md) | Authenticated testing, login flows, rules of engagement, and report filters. |
| [AI providers](docs/ai-providers.md) | Selecting the model, the supported providers (Anthropic, OpenAI, xAI, AWS Bedrock, and any other Pi-supported provider), and custom gateways. |
| [Configuration](docs/configuration.md) | Authenticated testing, login flows, rules of engagement, report filters, and rate-limit settings. |
| [AI providers](docs/ai-providers.md) | Anthropic, AWS Bedrock, and custom Anthropic-compatible endpoints. |
| [Platforms and networking](docs/platforms.md) | Windows/WSL2, Linux, macOS, Docker networking, local apps, and custom hostnames. |
| [Workspaces and resuming](docs/workspaces.md) | Naming workspaces, resuming interrupted scans, and workspace storage. |
| [Safety and limitations](docs/safety.md) | Authorized-use requirements, non-production guidance, mutative effects, cost, and model caveats. |
@@ -216,7 +209,7 @@ Important limitations:
Read the full [Safety and limitations](docs/safety.md) guide before running Shannon in a new environment.
## License
## License and Enterprise Licensing
Shannon Open Source is licensed under the [GNU Affero General Public License v3.0](LICENSE).
-1
View File
@@ -18,7 +18,6 @@
},
"dependencies": {
"@clack/prompts": "^1.1.0",
"@temporalio/client": "^1.11.0",
"chokidar": "^5.0.0",
"dotenv": "^17.3.1",
"smol-toml": "^1.6.1"
-106
View File
@@ -1,106 +0,0 @@
/**
* Shared argument parsing for CLI commands.
*
* Every command declares which boolean flags, value options, and positionals it
* accepts; `parseArgs` resolves aliases, rejects anything unrecognized, and hands
* back a typed result. This centralizes the common flags (notably `--yes`/`-y`) so
* each command no longer re-hardcodes `args.includes('--yes')`, and it makes
* unknown flags and stray arguments fail loudly instead of being silently ignored.
*/
import { closestMatch } from './suggest.js';
/** Thrown when argv does not match a command's schema. The dispatcher formats it. */
export class ArgError extends Error {}
/** Tokens that set the "skip confirmation" flag, declared once for every command. */
export const YES_FLAGS = ['--yes', '-y'] as const;
export interface ArgSchema {
/** Boolean flags: result key -> accepted tokens (canonical plus any aliases). */
readonly booleans?: Record<string, readonly string[]>;
/** Value-taking options: result key -> accepted tokens. */
readonly values?: Record<string, readonly string[]>;
/** Maximum positional arguments allowed. Defaults to 0. */
readonly maxPositionals?: number;
/** Extra guidance appended to the error when too many positionals are given. */
readonly positionalHint?: string;
}
export interface ParsedArgs {
readonly flags: Record<string, boolean>;
readonly values: Record<string, string>;
readonly positionals: readonly string[];
}
/** Build a token -> result-key lookup from a schema section. */
function indexTokens(section: Record<string, readonly string[]>): Map<string, string> {
const byToken = new Map<string, string>();
for (const [key, tokens] of Object.entries(section)) {
for (const token of tokens) {
byToken.set(token, key);
}
}
return byToken;
}
export function parseArgs(argv: readonly string[], schema: ArgSchema): ParsedArgs {
const booleanByToken = indexTokens(schema.booleans ?? {});
const valueByToken = indexTokens(schema.values ?? {});
const maxPositionals = schema.maxPositionals ?? 0;
const flags: Record<string, boolean> = {};
const values: Record<string, string> = {};
const positionals: string[] = [];
for (let i = 0; i < argv.length; i++) {
const arg = argv[i];
if (arg === undefined) {
continue;
}
const equalsIndex = arg.startsWith('--') ? arg.indexOf('=') : -1;
const token = equalsIndex === -1 ? arg : arg.slice(0, equalsIndex);
const inlineValue = equalsIndex === -1 ? undefined : arg.slice(equalsIndex + 1);
const booleanKey = booleanByToken.get(token);
if (booleanKey !== undefined) {
if (inlineValue !== undefined) {
throw new ArgError(`Flag ${token} does not take a value`);
}
flags[booleanKey] = true;
continue;
}
const valueKey = valueByToken.get(token);
if (valueKey !== undefined) {
if (inlineValue !== undefined) {
values[valueKey] = inlineValue;
continue;
}
const next = argv[i + 1];
if (next === undefined || next.startsWith('-')) {
throw new ArgError(`Option ${token} requires a value`);
}
values[valueKey] = next;
i++;
continue;
}
if (arg.startsWith('-')) {
const suggestion = closestMatch(token, [...booleanByToken.keys(), ...valueByToken.keys()]);
const hint = suggestion ? `\nDid you mean '${suggestion}'?` : '';
throw new ArgError(`Unknown option: ${token}${hint}`);
}
positionals.push(arg);
}
if (positionals.length > maxPositionals) {
const extra = positionals[maxPositionals];
const hint = schema.positionalHint ? `\n${schema.positionalHint}` : '';
throw new ArgError(`Unexpected argument: ${extra}${hint}`);
}
return { flags, values, positionals };
}
-34
View File
@@ -1,34 +0,0 @@
/**
* ANSI color and style escapes — the single source for the CLI's palette.
*
* Codes are plain constants; callers decide whether to emit them via `paint`
* (wrap-and-reset) or `gate` (prefix-or-empty), gating on `supportsColor()` from
* `tty.ts`. Cursor-control escapes live with their sole consumer, not here — this
* module is color only.
*/
export const RESET = '\x1b[0m';
/** Shannon brand gold — the running/completed accent, shared with the splash logo. */
export const GOLD = '\x1b[38;2;244;197;66m';
export const BOLD = '\x1b[1m';
export const RED = '\x1b[31m';
export const YELLOW = '\x1b[33m';
export const DIM = '\x1b[90m';
// The splash logo uses bolder variants of cyan/white/yellow than the progress tree.
export const CYAN = '\x1b[36;1m';
export const WHITE = '\x1b[1;37m';
export const GRAY = '\x1b[0;37m';
export const BOLD_YELLOW = '\x1b[1;33m';
/** Wrap `text` in `code` and reset, or return it unchanged when color is off. */
export function paint(text: string, code: string, enabled: boolean): string {
return enabled ? `${code}${text}${RESET}` : text;
}
/** A style code when color is on, or an empty string when off — for templates that interleave prefixes directly. */
export function gate(code: string, enabled: boolean): string {
return enabled ? code : '';
}
+12 -13
View File
@@ -1,20 +1,19 @@
/**
* `shannon build` command — build the worker Docker image from the repository.
* Requires a clone (Dockerfile in the working directory).
* `shannon build` command — build the worker Docker image locally.
* Only available in local mode (running from cloned repository).
*/
import { buildImage, canBuildImage, ensureDocker } from '../docker.js';
import { fail } from '../errors.js';
import { buildImage } from '../docker.js';
import { isLocal } from '../mode.js';
export function build(noCache: boolean, version: string): void {
ensureDocker();
if (!canBuildImage()) {
fail(
'Build is only available when running from the Shannon repository',
' (Dockerfile not found in current directory)',
);
export function build(noCache: boolean): void {
if (!isLocal()) {
console.error('ERROR: Build is only available when running from the Shannon repository');
console.error(' (Dockerfile not found in current directory)');
console.error('');
console.error('For npx usage, run: shannon update');
process.exit(1);
}
buildImage(noCache, version);
buildImage(noCache);
}
+52 -66
View File
@@ -8,10 +8,8 @@
import fs from 'node:fs';
import path from 'node:path';
import { watch } from 'chokidar';
import { fail } from '../errors.js';
import { getWorkspacesDir } from '../home.js';
import { resolveRunFile } from '../paths.js';
import { stdoutIsTerminal } from '../tty.js';
// Match the exact line the worker writes — anchored to prevent false positives from agent output
const COMPLETION_PATTERN = /^Scan (COMPLETED|FAILED)$/m;
@@ -30,7 +28,7 @@ function readRange(filePath: string, start: number, end: number): string {
}
/** Resolve a workspace ID to its workflow.log path, or exit with an error. */
export function resolveLogFile(workspaceId: string): string {
function resolveLogFile(workspaceId: string): string {
const workspacesDir = getWorkspacesDir();
// 1. Direct match
@@ -51,71 +49,59 @@ export function resolveLogFile(workspaceId: string): string {
if (fs.existsSync(namedPath)) return namedPath;
}
fail(
`No scan found named: ${workspaceId}`,
'',
'Possible causes:',
" - The scan hasn't started yet",
' - The workspace name is incorrect',
'',
'Check the dashboard at http://localhost:8233 for scan details',
);
}
/**
* Tail a scan's log until it reports completion, resolving when the completion marker appears
* (or the file is gone, or Ctrl-C stops it). Never exits the process, so the caller decides what
* happens next: plain `logs` exits 0; `start --follow` reads the workflow outcome first.
*/
export function tailUntilComplete(logFile: string): Promise<void> {
return new Promise((resolve) => {
let position = 0;
/**
* Output any new content appended since the last read.
* Returns true when the workflow completion marker is detected.
*/
function flush(): boolean {
try {
const { size } = fs.statSync(logFile);
if (size <= position) return false;
const data = readRange(logFile, position, size);
process.stdout.write(data);
position = size;
return COMPLETION_PATTERN.test(data);
} catch {
// File deleted or unreadable — treat as done
return true;
}
}
// 1. Output existing content
if (flush()) {
resolve();
return;
}
// 2. Watch for appended content via chokidar
const watcher = watch(logFile, { persistent: true });
const stop = (): void => {
watcher.close().finally(() => resolve());
// Safety net — resolve anyway if watcher.close() stalls
setTimeout(() => resolve(), 1000).unref();
};
watcher.on('change', () => {
if (flush()) stop();
});
process.on('SIGINT', stop);
});
console.error(`ERROR: No scan found named: ${workspaceId}`);
console.error('');
console.error('Possible causes:');
console.error(" - The scan hasn't started yet");
console.error(' - The workspace name is incorrect');
console.error('');
console.error('Check the dashboard at http://localhost:8233 for scan details');
process.exit(1);
}
export function logs(workspaceId: string): void {
const logFile = resolveLogFile(workspaceId);
console.error(stdoutIsTerminal() ? `Tailing scan log: ${logFile}` : 'Tailing scan log');
tailUntilComplete(logFile).finally(() => process.exit(0));
let position = 0;
/**
* Output any new content appended since the last read.
* Returns true when the workflow completion marker is detected.
*/
function flush(): boolean {
try {
const { size } = fs.statSync(logFile);
if (size <= position) return false;
const data = readRange(logFile, position, size);
process.stdout.write(data);
position = size;
return COMPLETION_PATTERN.test(data);
} catch {
// File deleted or unreadable — treat as done
return true;
}
}
console.log(`Tailing scan log: ${logFile}`);
// 1. Output existing content
if (flush()) {
process.exit(0);
}
// 2. Watch for appended content via chokidar
const watcher = watch(logFile, { persistent: true });
const shutdown = (): void => {
watcher.close().finally(() => process.exit(0));
// Safety net — force exit if watcher.close() stalls
setTimeout(() => process.exit(0), 1000).unref();
};
watcher.on('change', () => {
if (flush()) shutdown();
});
process.on('SIGINT', shutdown);
}
-26
View File
@@ -1,26 +0,0 @@
/**
* `shannon reset` command — stop everything and wipe all Temporal data and volumes,
* returning the machine to a clean slate. The destructive counterpart to `stop`.
*/
import * as p from '@clack/prompts';
import { confirmByTyping } from '../confirm.js';
import { ensureDocker, runningContainers, stopContainers, stopInfra, WORKER_FILTER } from '../docker.js';
export async function reset(): Promise<void> {
ensureDocker();
console.log('This will stop all running scans and permanently remove all Temporal data and volumes.');
await confirmByTyping('reset', 'confirm');
const spinner = p.spinner();
spinner.start('Stopping scans');
const running = runningContainers(WORKER_FILTER);
await stopContainers(running);
spinner.stop(
running.length > 0 ? `Stopped ${running.length} scan${running.length === 1 ? '' : 's'}` : 'No scans running',
);
await stopInfra(true);
console.log('Reset complete.');
}
-197
View File
@@ -1,197 +0,0 @@
/**
* `shannon scans` command — list completed scans and where each report lives.
*
* A scan counts as completed when it produced a report. The report can live in any of a
* few locations depending on the version that ran it, so `findReport` probes them in order
* and the first hit is both the completion signal and the link target behind the workspace
* name. The date and wall-clock duration come from the run's session.json
* (createdAt/completedAt), with the report file's mtime as the date fallback for
* runs that lack a recorded time.
*
* Human-readable by default; `--json` emits the same rows as raw machine values on stdout.
*
* Filesystem-only (local ./workspaces/ or npx ~/.shannon/workspaces/ via getWorkspacesDir);
* no Temporal dependency.
*/
import fs from 'node:fs';
import path from 'node:path';
import { pathToFileURL } from 'node:url';
import { BOLD, GOLD, paint } from '../colors.js';
import { getWorkspacesDir } from '../home.js';
import { commandPrefix } from '../mode.js';
import { FINAL_REPORT_PDF_FILENAME, INTERNAL_DIR, resolveRunFile } from '../paths.js';
import { stdoutIsTerminal, supportsColor } from '../tty.js';
/** Assembled report in the deliverables dir. Must match ASSEMBLED_REPORT_FILENAME in the worker package. */
const ASSEMBLED_REPORT_FILENAME = 'comprehensive_security_assessment_report.md';
/** Run-root markdown surfaced by older versions, before the PDF. Kept so those runs still list. */
const FINAL_REPORT_MD_FILENAME = 'Security-Assessment-Report.md';
const DELIVERABLES_SUBDIR = 'deliverables';
/** One completed scan; raw values so the table and --json render from one source. */
interface ScanRow {
readonly workspace: string;
/** Completion time in ms — sort key and date source. */
readonly finishedMs: number;
/** Wall-clock duration (completedAt − createdAt) in ms, or null when unknown. */
readonly durationMs: number | null;
/** Absolute path to the report file — the link target behind the workspace name. */
readonly report: string;
}
/** The --json row shape: raw machine values, one per completed scan. */
interface JsonRow {
readonly workspace: string;
readonly finishedAt: string;
readonly durationMs: number | null;
readonly reportPath: string;
}
/** Compact wall-clock duration from milliseconds: "47s", "1m 32s", "1h 47m". */
function formatDuration(ms: number): string {
const totalSeconds = Math.round(ms / 1000);
if (totalSeconds < 60) {
return `${totalSeconds}s`;
}
const totalMinutes = Math.floor(totalSeconds / 60);
if (totalMinutes < 60) {
return `${totalMinutes}m ${totalSeconds % 60}s`;
}
return `${Math.floor(totalMinutes / 60)}h ${totalMinutes % 60}m`;
}
/**
* Wrap `text` in an OSC 8 hyperlink to `url` so a supporting terminal opens it on click,
* or return `text` unchanged. Terminals without OSC 8 simply show the text.
*/
function hyperlink(text: string, url: string): string {
return `\x1b]8;;${url}\x1b\\${text}\x1b]8;;\x1b\\`;
}
/** First existing report path for a run (newest-surfaced first), or null if it has none. */
function findReport(runDir: string): string | null {
const candidates = [
path.join(runDir, FINAL_REPORT_PDF_FILENAME),
path.join(runDir, FINAL_REPORT_MD_FILENAME),
path.join(runDir, INTERNAL_DIR, DELIVERABLES_SUBDIR, ASSEMBLED_REPORT_FILENAME),
path.join(runDir, DELIVERABLES_SUBDIR, ASSEMBLED_REPORT_FILENAME),
];
for (const candidate of candidates) {
if (fs.existsSync(candidate)) {
return candidate;
}
}
return null;
}
interface SessionData {
readonly session: { readonly createdAt?: string; readonly completedAt?: string };
}
/** Read a run's session.json (dual-read across layouts). Missing or unreadable → empty shape. */
function readSession(runDir: string): SessionData {
try {
const parsed = JSON.parse(fs.readFileSync(resolveRunFile(runDir, 'session.json'), 'utf8'));
return { session: parsed?.session ?? {} };
} catch {
return { session: {} };
}
}
/** Gather every workspace that has a report, one row each. */
function collectCompletedScans(workspacesDir: string): ScanRow[] {
let entries: fs.Dirent[];
try {
entries = fs.readdirSync(workspacesDir, { withFileTypes: true });
} catch {
// Workspaces directory does not exist yet — no scans have ever run.
return [];
}
const rows: ScanRow[] = [];
for (const entry of entries) {
if (!entry.isDirectory()) {
continue;
}
const runDir = path.join(workspacesDir, entry.name);
const reportPath = findReport(runDir);
if (!reportPath) {
continue;
}
const { session } = readSession(runDir);
const completedMs = Date.parse(session.completedAt ?? '');
const createdMs = Date.parse(session.createdAt ?? '');
const finishedMs = Number.isNaN(completedMs) ? fs.statSync(reportPath).mtimeMs : completedMs;
const durationMs = Number.isNaN(completedMs) || Number.isNaN(createdMs) ? null : completedMs - createdMs;
rows.push({ workspace: entry.name, finishedMs, durationMs, report: reportPath });
}
return rows;
}
function toJsonRow(row: ScanRow): JsonRow {
return {
workspace: row.workspace,
finishedAt: new Date(row.finishedMs).toISOString(),
durationMs: row.durationMs,
reportPath: row.report,
};
}
/** Print the completed scans as an aligned table with the workspace name linked to its report. */
function printTable(workspacesDir: string, rows: readonly ScanRow[]): void {
if (rows.length === 0) {
const prefix = commandPrefix();
console.log(`No completed scans yet. Run '${prefix} start -u <url> -r <path>' to begin.`);
return;
}
const color = supportsColor();
// On a terminal the workspace name is an OSC 8 hyperlink that opens its report; when
// piped there is nothing to click, so it prints as plain text.
const linkable = stdoutIsTerminal();
const table = rows.map((row) => ({
finished: new Date(row.finishedMs).toISOString().slice(0, 10),
duration: row.durationMs === null ? '—' : formatDuration(row.durationMs),
workspace: row.workspace,
report: row.report,
}));
const dateWidth = Math.max('FINISHED'.length, 'YYYY-MM-DD'.length);
const durationWidth = Math.max('DURATION'.length, ...table.map((row) => row.duration.length));
console.log(`\nCompleted scans in ${workspacesDir}:\n`);
const header = `${'FINISHED'.padEnd(dateWidth)} ${'DURATION'.padEnd(durationWidth)} WORKSPACE`;
console.log(paint(header, BOLD, color));
for (const row of table) {
const finished = row.finished.padEnd(dateWidth);
const duration = row.duration.padEnd(durationWidth);
const name = paint(row.workspace, GOLD, color);
const workspace = linkable ? hyperlink(name, pathToFileURL(row.report).href) : name;
console.log(`${finished} ${duration} ${workspace}`);
}
console.log('');
}
export function scans(opts: { readonly json: boolean }): void {
const workspacesDir = getWorkspacesDir();
const rows = collectCompletedScans(workspacesDir);
// Latest on top.
rows.sort((a, b) => b.finishedMs - a.finishedMs);
if (opts.json) {
console.log(JSON.stringify(rows.map(toJsonRow), null, 2));
return;
}
printTable(workspacesDir, rows);
}
+154 -232
View File
@@ -1,162 +1,57 @@
/**
* `npx @keygraph/shannon setup` — interactive TUI wizard for one-time credential configuration.
*
* Walks the user through selecting a provider, entering credentials, and naming
* the model that runs the whole scan, then persists everything to
* ~/.shannon/config.toml with 0o600 permissions.
* Walks the user through selecting a provider and entering credentials,
* then persists everything to ~/.shannon/config.toml with 0o600 permissions.
*/
import os from 'node:os';
import path from 'node:path';
import * as p from '@clack/prompts';
import { type ShannonConfig, saveConfig } from '../config/writer.js';
import { CURATED_PROVIDERS, type CuratedProviderId, isCuratedProvider, type OpenAiFormat } from '../model-spec.js';
import { requireInteractive } from '../tty.js';
const SHANNON_HOME = path.join(os.homedir(), '.shannon');
const CUSTOM_MODEL = '__custom__';
const CUSTOM_BASE_URL = '__custom_base_url__';
const OTHER_PROVIDER = '__other_provider__';
/**
* Wire formats reachable through the gateway route. The format picks the provider
* that supplies the credential, and for OpenAI it also picks which of the two
* OpenAI APIs Shannon calls.
*/
const GATEWAY_DIALECTS: readonly {
value: string;
label: string;
provider: 'anthropic' | 'openai';
format?: OpenAiFormat;
}[] = [
{ value: 'anthropic', label: 'Anthropic Messages', provider: 'anthropic' },
{
value: 'openai-chat-completions',
label: 'OpenAI Chat Completions',
provider: 'openai',
format: 'chat-completions',
},
{ value: 'openai-responses', label: 'OpenAI Responses', provider: 'openai', format: 'responses' },
];
/** Suggested models per curated provider, best-first. Free-text entry accepts any model in the provider's catalogue. */
const MODEL_SUGGESTIONS: Readonly<Record<CuratedProviderId, readonly string[]>> = {
anthropic: ['claude-sonnet-4-6', 'claude-opus-4-8', 'claude-opus-4-7', 'claude-haiku-4-5-20251001'],
openai: ['gpt-5.6-sol', 'gpt-5.5', 'gpt-5.4'],
xai: ['grok-4.5'],
'amazon-bedrock': ['us.anthropic.claude-sonnet-4-6', 'us.anthropic.claude-opus-4-8', 'us.anthropic.claude-opus-4-7'],
};
/** Placeholder shown in the free-text model ID prompt, per curated provider. */
const MODEL_ID_PLACEHOLDER: Readonly<Record<CuratedProviderId, string>> = {
anthropic: 'claude-sonnet-4-6',
openai: 'gpt-5.6-sol',
xai: 'grok-4.5',
'amazon-bedrock': 'us.anthropic.claude-opus-4-8',
};
/** Model ID placeholder for a provider, absent when the provider is not curated. */
function modelIdPlaceholder(provider: string): string | undefined {
return isCuratedProvider(provider) ? MODEL_ID_PLACEHOLDER[provider] : undefined;
}
type Provider = 'anthropic' | 'custom_base_url' | 'bedrock';
export async function setup(): Promise<void> {
requireInteractive('setup', 'For non-interactive use, export credentials as env vars (e.g. ANTHROPIC_API_KEY).');
p.intro('Shannon Setup');
// 1. Select provider. "Custom Base URL" is a route, not a provider — it asks
// which API dialect the gateway speaks and configures that provider. "Other
// provider" reaches any pi-supported provider Shannon does not curate.
const selected = await p.select({
// 1. Select provider
const provider = await p.select({
message: 'Select your AI provider',
options: [
{ value: 'anthropic' as const, label: 'Anthropic', hint: 'Claude models - recommended' },
{ value: 'openai' as const, label: 'OpenAI', hint: 'GPT models' },
{ value: 'xai' as const, label: 'xAI', hint: 'Grok models' },
{ value: 'amazon-bedrock' as const, label: 'AWS Bedrock', hint: 'Claude models via AWS' },
{ value: CUSTOM_BASE_URL as typeof CUSTOM_BASE_URL, label: 'Custom Base URL', hint: 'your own proxy or gateway' },
{
value: OTHER_PROVIDER as typeof OTHER_PROVIDER,
label: 'Other provider',
hint: 'any other Pi-supported provider',
},
{ value: 'anthropic' as const, label: 'Claude Direct', hint: 'recommended' },
{ value: 'custom_base_url' as const, label: 'Custom Base URL', hint: 'proxies, gateways' },
{ value: 'bedrock' as const, label: 'Claude via AWS Bedrock' },
],
});
if (p.isCancel(selected)) return cancelAndExit();
// 2. Credentials — and, on the gateway route, the endpoint and its dialect.
const { provider, config, gateway } = await setupSelection(selected);
// 3. The model that runs every phase.
const modelId = await promptModel(provider);
config.core = { ...config.core, model: `${provider}:${modelId}` };
if (gateway) config.core = { ...config.core, base_url: gateway.baseUrl };
saveConfig(config);
const configPath = path.join(SHANNON_HOME, 'config.toml');
const summary = [`Provider ${provider}`, `Model ${modelId}`];
if (gateway) summary.push(`Endpoint ${gateway.baseUrl}`);
if (gateway?.format) summary.push(`API ${gateway.format}`);
p.log.success(`Configuration saved to ${configPath}`);
p.log.info(summary.join('\n'));
p.outro('Run `npx @keygraph/shannon start` to begin a scan.');
}
interface Selection {
provider: string;
config: ShannonConfig;
gateway?: GatewaySetup;
}
/** Resolve the provider selection into a provider id and its credential config. */
async function setupSelection(
selected: CuratedProviderId | typeof CUSTOM_BASE_URL | typeof OTHER_PROVIDER,
): Promise<Selection> {
if (selected === CUSTOM_BASE_URL) {
const gateway = await setupGateway();
return { provider: gateway.provider, config: gateway.config, gateway };
}
if (selected === OTHER_PROVIDER) {
return setupOtherProvider();
}
return { provider: selected, config: await setupProvider(selected) };
}
async function setupProvider(provider: CuratedProviderId): Promise<ShannonConfig> {
switch (provider) {
case 'amazon-bedrock':
return setupBedrock();
case 'anthropic':
return setupAnthropic();
case 'openai':
return { openai: { api_key: await promptSecret('Enter your OpenAI API key') } };
case 'xai':
return { xai: { api_key: await promptSecret('Enter your xAI API key') } };
}
}
/**
* Any pi provider Shannon does not curate. The id is free text — the worker's
* preflight validates it — and the key is stored generically as SHANNON_AI_API_KEY.
*/
async function setupOtherProvider(): Promise<Selection> {
p.log.info('Browse supported providers and models at https://pi.dev/models');
const provider = await p.text({
message: 'Provider ID',
validate: (value) => {
const id = value?.trim();
if (!id) return 'Provider ID is required';
if (isCuratedProvider(id)) return `${id} has its own option.`;
return undefined;
},
});
if (p.isCancel(provider)) return cancelAndExit();
const apiKey = await promptSecret('Enter the API key');
return { provider: provider.trim(), config: { provider: { api_key: apiKey } } };
const config = await setupProvider(provider as Provider);
// 2. Adaptive thinking
await maybePromptAdaptiveThinking(config);
// 3. Save config
saveConfig(config);
const configPath = path.join(SHANNON_HOME, 'config.toml');
p.log.success(`Configuration saved to ${configPath}`);
p.outro('Run `npx @keygraph/shannon start` to begin a scan.');
}
async function setupProvider(provider: Provider): Promise<ShannonConfig> {
switch (provider) {
case 'anthropic':
return setupAnthropic();
case 'custom_base_url':
return setupCustomBaseUrl();
case 'bedrock':
return setupBedrock();
}
}
// === Provider Setup Flows ===
@@ -171,54 +66,58 @@ async function setupAnthropic(): Promise<ShannonConfig> {
});
if (p.isCancel(authMethod)) return cancelAndExit();
const config: ShannonConfig = {};
if (authMethod === 'oauth') {
const token = await promptSecret('Enter your OAuth token');
return { anthropic: { oauth_token: token } };
config.anthropic = { oauth_token: token };
} else {
const apiKey = await promptSecret('Enter your Anthropic API key');
config.anthropic = { api_key: apiKey };
}
const apiKey = await promptSecret('Enter your Anthropic API key');
return { anthropic: { api_key: apiKey } };
}
async function setupBedrock(): Promise<ShannonConfig> {
const region = await p.text({
message: 'AWS Region',
placeholder: 'us-east-1',
validate: required('AWS Region is required'),
const customizeModels = await p.confirm({
message:
'Do you want to change the default models?\n' +
' Small - claude-haiku-4-5-20251001\n' +
' Medium - claude-sonnet-4-6\n' +
' Large - claude-opus-4-8',
initialValue: false,
});
if (p.isCancel(region)) return cancelAndExit();
if (p.isCancel(customizeModels)) return cancelAndExit();
const token = await promptSecret('Enter your AWS Bearer Token');
if (customizeModels) {
const small = await p.text({
message: 'Small model ID',
initialValue: 'claude-haiku-4-5-20251001',
validate: required('Small model ID is required'),
});
if (p.isCancel(small)) return cancelAndExit();
return { bedrock: { region, token } };
const medium = await p.text({
message: 'Medium model ID',
initialValue: 'claude-sonnet-4-6',
validate: required('Medium model ID is required'),
});
if (p.isCancel(medium)) return cancelAndExit();
const large = await p.text({
message: 'Large model ID',
initialValue: 'claude-opus-4-8',
validate: required('Large model ID is required'),
});
if (p.isCancel(large)) return cancelAndExit();
config.models = { small, medium, large };
}
return config;
}
interface GatewaySetup {
provider: CuratedProviderId;
config: ShannonConfig;
baseUrl: string;
format?: OpenAiFormat;
}
/**
* Gateway route: the endpoint decides where requests go, but the format still
* picks a real provider, because that is what supplies the credential and the
* wire protocol.
*/
async function setupGateway(): Promise<GatewaySetup> {
const choice = await p.select({
message: 'API format',
options: GATEWAY_DIALECTS.map(({ value, label }) => ({ value, label })),
});
if (p.isCancel(choice)) return cancelAndExit();
const dialect = GATEWAY_DIALECTS.find((entry) => entry.value === choice);
if (!dialect) return cancelAndExit();
const provider = dialect.provider;
async function setupCustomBaseUrl(): Promise<ShannonConfig> {
const baseUrl = await p.text({
message: 'Endpoint URL',
placeholder: 'https://llm-gateway.example.com',
placeholder: 'https://your-proxy.example.com',
validate: (value) => {
if (!value) return 'Endpoint URL is required';
try {
@@ -231,80 +130,103 @@ async function setupGateway(): Promise<GatewaySetup> {
});
if (p.isCancel(baseUrl)) return cancelAndExit();
const authToken = await promptSecret('Enter the auth token for the endpoint');
const config: ShannonConfig =
provider === 'anthropic'
? { anthropic: { api_key: authToken } }
: { openai: { api_key: authToken, ...(dialect.format && { format: dialect.format }) } };
const authToken = await promptSecret('Enter the auth token for the custom endpoint');
return { provider, config, baseUrl, ...(dialect.format && { format: dialect.format }) };
}
const config: ShannonConfig = {
custom_base_url: { base_url: baseUrl, auth_token: authToken },
};
// === Model Selection ===
const customizeModels = await p.confirm({
message:
'Do you want to change the default models?\n' +
' Small - claude-haiku-4-5-20251001\n' +
' Medium - claude-sonnet-4-6\n' +
' Large - claude-opus-4-8',
initialValue: false,
});
if (p.isCancel(customizeModels)) return cancelAndExit();
/**
* Ask for the one model that runs every phase. Providers with suggestions offer a
* pick list with a free-text escape hatch; the rest go straight to free text.
*/
async function promptModel(provider: string): Promise<string> {
const suggestions = isCuratedProvider(provider) ? MODEL_SUGGESTIONS[provider] : [];
if (customizeModels) {
const small = await p.text({
message: 'Small model ID',
initialValue: 'claude-haiku-4-5-20251001',
validate: required('Small model ID is required'),
});
if (p.isCancel(small)) return cancelAndExit();
if (suggestions.length === 0) {
return promptModelId(provider, modelIdPlaceholder(provider));
const medium = await p.text({
message: 'Medium model ID',
initialValue: 'claude-sonnet-4-6',
validate: required('Medium model ID is required'),
});
if (p.isCancel(medium)) return cancelAndExit();
const large = await p.text({
message: 'Large model ID',
initialValue: 'claude-opus-4-8',
validate: required('Large model ID is required'),
});
if (p.isCancel(large)) return cancelAndExit();
config.models = { small, medium, large };
}
const choice = await p.select({
message: 'Model',
options: [
...suggestions.map((model) => ({ value: model, label: model })),
{ value: CUSTOM_MODEL, label: 'Enter a model ID…' },
],
});
if (p.isCancel(choice)) return cancelAndExit();
if (choice === CUSTOM_MODEL) {
return promptModelId(provider, modelIdPlaceholder(provider));
}
return choice as string;
return config;
}
/**
* A leading `<provider>:` naming a supported provider other than the selected
* one. Bedrock model IDs carry their own colons (`…-v1:0`), so only a genuine
* provider id counts as a prefix.
*/
function conflictingProviderPrefix(provider: string, value: string): string | undefined {
const separator = value.indexOf(':');
if (separator === -1) return undefined;
const head = value.slice(0, separator);
if (head === provider) return undefined;
return (CURATED_PROVIDERS as readonly string[]).includes(head) ? head : undefined;
}
/**
* Ask for a model ID. The provider is already chosen, so this takes the bare ID
* and the caller pairs it with the provider — pasting a full `<provider>:<model>`
* spec just has its redundant prefix dropped.
*/
async function promptModelId(provider: string, placeholder?: string): Promise<string> {
const modelId = await p.text({
message: 'Model ID',
...(placeholder && { placeholder }),
validate: (value) => {
if (!value) return 'Model ID is required';
const conflicting = conflictingProviderPrefix(provider, value);
if (conflicting) return `That model ID is for ${conflicting}, but you selected ${provider}.`;
return undefined;
},
async function setupBedrock(): Promise<ShannonConfig> {
const region = await p.text({
message: 'AWS Region',
placeholder: 'us-east-1',
validate: required('AWS Region is required'),
});
if (p.isCancel(modelId)) return cancelAndExit();
if (p.isCancel(region)) return cancelAndExit();
return modelId.startsWith(`${provider}:`) ? modelId.slice(provider.length + 1) : modelId;
const token = await promptSecret('Enter your AWS Bearer Token');
const small = await p.text({
message: 'Small model ID',
placeholder: 'us.anthropic.claude-haiku-4-5-20251001-v1:0',
validate: required('Small model ID is required'),
});
if (p.isCancel(small)) return cancelAndExit();
const medium = await p.text({
message: 'Medium model ID',
placeholder: 'us.anthropic.claude-sonnet-4-6',
validate: required('Medium model ID is required'),
});
if (p.isCancel(medium)) return cancelAndExit();
const large = await p.text({
message: 'Large model ID',
placeholder: 'us.anthropic.claude-opus-4-8',
validate: required('Large model ID is required'),
});
if (p.isCancel(large)) return cancelAndExit();
return {
bedrock: { use: true, region, token },
models: { small, medium, large },
};
}
// === Helpers ===
async function maybePromptAdaptiveThinking(config: ShannonConfig): Promise<void> {
const m = config.models;
const hasAdaptiveModel = !m || [m.small, m.medium, m.large].some((v) => v && /opus-4-[678]/.test(v));
if (!hasAdaptiveModel) return;
const enable = await p.confirm({
message: 'Enable adaptive thinking on Opus 4.6/4.7/4.8? Claude decides when and how deeply to reason.',
initialValue: true,
});
if (p.isCancel(enable)) return cancelAndExit();
config.core = { ...config.core, adaptive_thinking: enable };
}
async function promptSecret(message: string): Promise<string> {
const value = await p.password({
message,
+99 -130
View File
@@ -8,27 +8,13 @@
import { execFileSync } from 'node:child_process';
import fs from 'node:fs';
import path from 'node:path';
import { setTimeout as sleep } from 'node:timers/promises';
import * as p from '@clack/prompts';
import { ensureDocker, ensureImage, ensureInfra, randomSuffix, spawnWorker } from '../docker.js';
import { buildEnvFlags, loadEnv, resolveHostPiAuthPath, shouldUsePiAuth, validateCredentials } from '../env.js';
import { fail } from '../errors.js';
import { ensureImage, ensureInfra, randomSuffix, spawnWorker } from '../docker.js';
import { buildEnvFlags, loadEnv, validateCredentials } from '../env.js';
import { getWorkspacesDir, initHome } from '../home.js';
import { commandPrefix, isLocal } from '../mode.js';
import { resolveModelSpec } from '../model-spec.js';
import {
expandHome,
FINAL_REPORT_PDF_FILENAME,
INTERNAL_DIR,
resolveConfig,
resolveRepo,
resolveRunFile,
} from '../paths.js';
import { resolveWorkflowId } from '../session.js';
import { isLocal } from '../mode.js';
import { FINAL_REPORT_FILENAME, INTERNAL_DIR, resolveConfig, resolveRepo, resolveRunFile } from '../paths.js';
import { displaySplash } from '../splash.js';
import { getTerminalOutcome } from '../temporal-client.js';
import { stdoutIsTerminal } from '../tty.js';
import { tailUntilComplete } from './logs.js';
export interface StartArgs {
url: string;
@@ -37,8 +23,7 @@ export interface StartArgs {
workspace?: string;
output?: string;
pipelineTesting: boolean;
keepContainer: boolean;
follow: boolean;
debug: boolean;
version: string;
}
@@ -73,29 +58,22 @@ export async function start(args: StartArgs): Promise<void> {
// 2. Validate credentials
const creds = validateCredentials();
if (!creds.valid) {
fail(creds.error ?? 'Invalid credentials');
console.error(`ERROR: ${creds.error}`);
process.exit(1);
}
// 3. Resolve paths
const repo = resolveRepo(args.repo);
const config = args.config ? resolveConfig(args.config) : undefined;
// Inputs are valid — show the splash before the Docker/Temporal setup work.
displaySplash(isLocal() ? undefined : args.version);
// 4. Ensure workspaces dir is writable by container user (UID 1001)
const workspacesDir = getWorkspacesDir();
fs.mkdirSync(workspacesDir, { recursive: true });
fs.chmodSync(workspacesDir, 0o777);
// 5. Ensure Docker and the worker image are available (pull/build prints its own progress).
ensureDocker();
// 5. Ensure image (auto-build in dev, pull in npx) and start infra
ensureImage(args.version);
// One spinner spans the whole launch: bringing up Temporal and registering the worker.
const spinner = p.spinner();
spinner.start('Starting scan');
await ensureInfra(spinner);
await ensureInfra();
// 6. Generate unique task queue and container name
const suffix = randomSuffix();
@@ -130,7 +108,7 @@ export async function start(args: StartArgs): Promise<void> {
fs.mkdirSync(path.join(repo.hostPath, '.playwright'), { recursive: true });
// 10. Resolve output directory
const outputDir = args.output ? path.resolve(expandHome(args.output)) : undefined;
const outputDir = args.output ? path.resolve(args.output) : undefined;
if (outputDir) {
fs.mkdirSync(outputDir, { recursive: true });
}
@@ -138,7 +116,10 @@ export async function start(args: StartArgs): Promise<void> {
// 11. Resolve prompts directory (local mode only)
const promptsDir = isLocal() ? path.resolve('apps/worker/prompts') : undefined;
// 12. Spawn worker container
// 12. Display splash screen
displaySplash(isLocal() ? undefined : args.version);
// 13. Spawn worker container
const proc = spawnWorker({
version: args.version,
url: args.url,
@@ -152,18 +133,19 @@ export async function start(args: StartArgs): Promise<void> {
...(outputDir && { outputDir }),
workspace,
...(args.pipelineTesting && { pipelineTesting: true }),
...(args.keepContainer && { keepContainer: true }),
...(shouldUsePiAuth() && { piAuthHostPath: resolveHostPiAuthPath() }),
...(args.debug && { debug: true }),
});
// Bail if `docker run -d` itself fails (mount error, image missing, etc.)
// 14. Bail if `docker run -d` itself fails (mount error, image missing, etc.)
const dockerExitCode = await new Promise<number>((resolve) => {
proc.once('exit', (code) => resolve(code ?? 1));
proc.once('error', () => resolve(1));
proc.once('error', (err) => {
console.error(`Failed to start the scan: ${err.message}`);
resolve(1);
});
});
if (dockerExitCode !== 0) {
spinner.error('Could not start the scan');
process.exit(1);
}
@@ -180,23 +162,64 @@ export async function start(args: StartArgs): Promise<void> {
}
}
// Poll for workflow to register in session.json. Off-TTY, skip the dots and
// clear-line escape so redirected logs stay clean.
const animate = stdoutIsTerminal();
process.stdout.write('Waiting for the scan to start...');
let workflowId = '';
let started = false;
let attempts = 0;
const pollInterval = setInterval(() => {
attempts++;
if (attempts > 60) {
clearInterval(pollInterval);
process.stdout.write('\n');
console.error('Timed out waiting for the scan to start');
process.exit(1);
}
// Stop the worker only if the scan hasn't registered yet (e.g. Ctrl-C mid-startup).
try {
const session = JSON.parse(fs.readFileSync(sessionJson, 'utf-8'));
const resumeAttempts: { workflowId: string }[] = session.session?.resumeAttempts ?? [];
// Fresh: session.json appears with originalWorkflowId. Resume: new resumeAttempts entry.
const ready = isResume ? resumeAttempts.length > initialResumeCount : !!session.session?.originalWorkflowId;
if (ready) {
clearInterval(pollInterval);
started = true;
// Latest workflow ID: last resume attempt, or originalWorkflowId for fresh scans
workflowId = resumeAttempts.at(-1)?.workflowId ?? session.session?.originalWorkflowId ?? '';
// Clear the waiting line, or just break it off-TTY
process.stdout.write(animate ? '\r\x1b[K' : '\n');
printInfo(args, workspace, workflowId, repo.hostPath, workspacesDir);
return;
}
} catch {
// File doesn't exist yet
}
if (animate) process.stdout.write('.');
}, 2000);
// Stop the worker container only if it hasn't started yet
let cleaned = false;
const cleanup = (): void => {
if (cleaned || started) return;
cleaned = true;
spinner.stop('Stopping scan');
clearInterval(pollInterval);
console.log('\nStopping scan...');
try {
execFileSync('docker', ['stop', containerName], { stdio: 'pipe' });
} catch {
// Container may have already exited
}
if (args.keepContainer) {
printPreservedContainerHint(containerName);
if (args.debug) {
printDebugHint(containerName);
}
};
process.on('SIGINT', () => {
cleanup();
process.exit(0);
@@ -206,69 +229,9 @@ export async function start(args: StartArgs): Promise<void> {
process.exit(0);
});
process.on('exit', cleanup);
// Poll for the workflow to register in session.json; the spinner resolves once it does.
spinner.message('Waiting for the scan to start');
for (let attempts = 0; attempts < 60; attempts++) {
try {
const session = JSON.parse(fs.readFileSync(sessionJson, 'utf-8'));
const resumeAttempts: { workflowId: string }[] = session.session?.resumeAttempts ?? [];
// Fresh: session.json appears with originalWorkflowId. Resume: new resumeAttempts entry.
const ready = isResume ? resumeAttempts.length > initialResumeCount : !!session.session?.originalWorkflowId;
if (ready) {
started = true;
spinner.stop(`Scan started — ${workspace}`);
printInfo(args, workspace, repo.hostPath, workspacesDir);
if (args.follow) {
await followScan(workspace, workspacesDir);
}
return;
}
} catch {
// File doesn't exist yet
}
await sleep(2000);
}
spinner.error('Timed out waiting for the scan to start');
process.exit(1);
}
/**
* Follow a just-started scan (for `--follow`, aimed at CI): stream its log to completion, then
* exit on the workflow outcome — 0 if the assessment ran, 1 if the scan failed. That tracks
* whether the pipeline ran, not whether vulnerabilities were found.
*/
async function followScan(workspace: string, workspacesDir: string): Promise<never> {
const logFile = resolveRunFile(path.join(workspacesDir, workspace), 'workflow.log');
// The worker creates workflow.log as it starts; wait briefly so the first read doesn't
// mistake a not-yet-created file for an already-finished scan.
for (let attempts = 0; attempts < 30 && !fs.existsSync(logFile); attempts++) {
await sleep(1000);
}
if (stdoutIsTerminal()) {
console.error('\n Following scan log (Ctrl-C to stop watching):\n');
}
await tailUntilComplete(logFile);
const workflowId = resolveWorkflowId(workspace);
if (!workflowId) {
fail('Scan finished but its workflow id could not be resolved from session.json.');
}
try {
const outcome = await getTerminalOutcome(workflowId);
process.exit(outcome.kind === 'success' ? 0 : 1);
} catch {
fail('Could not reach Temporal at 127.0.0.1:7233 to read the scan outcome.');
}
}
function printPreservedContainerHint(containerName: string): void {
function printDebugHint(containerName: string): void {
console.log('');
console.log(` Worker container preserved: ${containerName}`);
console.log(` Inspect logs: docker logs ${containerName}`);
@@ -276,45 +239,51 @@ function printPreservedContainerHint(containerName: string): void {
console.log('');
}
function printInfo(args: StartArgs, workspace: string, repoPath: string, workspacesDir: string): void {
const interactive = stdoutIsTerminal();
if (interactive && !args.follow) {
console.log(' It runs in the background — you can close this terminal.');
console.log('');
}
function printInfo(
args: StartArgs,
workspace: string,
workflowId: string,
repoPath: string,
workspacesDir: string,
): void {
const logsCmd = isLocal() ? `./shannon logs ${workspace}` : `npx @keygraph/shannon logs ${workspace}`;
const reportPath = path.join(workspacesDir, workspace, FINAL_REPORT_FILENAME);
console.log(' Scan started — it runs in the background, so you can close this terminal.');
console.log('');
console.log(` Target: ${args.url}`);
console.log(` Repository: ${interactive ? repoPath : path.basename(repoPath)}`);
console.log(` Repository: ${repoPath}`);
console.log(` Workspace: ${workspace}`);
if (args.config) {
console.log(` Config: ${interactive ? path.resolve(args.config) : path.basename(args.config)}`);
console.log(` Config: ${path.resolve(args.config)}`);
}
if (args.pipelineTesting) {
console.log(' Mode: Pipeline Testing');
}
const spec = resolveModelSpec();
if (typeof spec !== 'string') {
console.log(` Model: ${spec.providerId}:${spec.modelId}`);
// Surface Fable usage: its safety classifiers route cybersecurity tasks to
// Opus 4.8, so those phases run on Opus 4.8 regardless of the tier setting.
const fableTiers = (
[
['small', process.env.ANTHROPIC_SMALL_MODEL],
['medium', process.env.ANTHROPIC_MEDIUM_MODEL],
['large', process.env.ANTHROPIC_LARGE_MODEL],
] as const
).filter(([, model]) => model && /fable/i.test(model));
if (fableTiers.length > 0) {
const tierList = fableTiers.map(([tier, model]) => `${tier} (${model})`).join(', ');
console.log(` Note: ${tierList} set to a Fable model. Fable's safety classifiers`);
console.log(' route cybersecurity tasks to Opus 4.8, so those phases run on Opus 4.8.');
}
if (!interactive) {
return;
console.log('');
console.log(' Watch scan progress:');
console.log(` Live logs: ${logsCmd}`);
if (workflowId) {
console.log(` Dashboard: http://localhost:8233/namespaces/default/workflows/${workflowId}`);
} else {
console.log(' Dashboard: http://localhost:8233');
}
const reportPath = path.join(workspacesDir, workspace, FINAL_REPORT_PDF_FILENAME);
// When following, the scan log streams inline next, so the "run these to watch it" hints
// would only contradict that.
if (!args.follow) {
const prefix = commandPrefix();
console.log('');
console.log(' Watch scan progress:');
console.log(` Live logs: ${prefix} logs ${workspace}`);
console.log(` Progress: ${prefix} status ${workspace}`);
}
console.log('');
console.log(' Report (when the scan finishes):');
console.log(` ${reportPath}`);
+16 -188
View File
@@ -1,196 +1,24 @@
/**
* `shannon status <workspace>` — one scan's live progress from Temporal.
*
* While the scan runs, polls Temporal and redraws the phase/agent tree on a
* terminal (a pipe or a finished scan gets a single frame). When the scan reaches
* a terminal state, prints the overall result and exits. Reads Temporal directly —
* no worker, no session files — so it needs Temporal up and shows scans within its
* ~24h retention window.
* `shannon status` command — show running scans and Temporal health.
*/
import { setTimeout as sleep } from 'node:timers/promises';
import { fail } from '../errors.js';
import { isLocal } from '../mode.js';
import { type RenderInput, renderScan } from '../scan/render.js';
import { toStatusJson } from '../scan/status-json.js';
import { resolveWorkflowId } from '../session.js';
import { displaySplash } from '../splash.js';
import { describeScan, getTerminalOutcome, queryProgress, type ScanDescription } from '../temporal-client.js';
import { stdoutIsTerminal, supportsColor } from '../tty.js';
import { getVersion } from '../version.js';
import { isTemporalReady, listRunningWorkers } from '../docker.js';
const HIDE_CURSOR = '\x1b[?25l';
const SHOW_CURSOR = '\x1b[?25h';
/** Redraw cadence for the spinner animation; data is refreshed on the slower poll. */
const RENDER_MS = 120;
const POLL_MS = 1200;
/** Terminal = anything other than an open, running execution. */
function isTerminalStatus(status: string): boolean {
return status !== 'RUNNING' && status !== 'UNSPECIFIED';
}
// Match SGR color escapes (ESC[…m) so a line's on-screen width excludes them. Built from the ESC
// char code so the source carries no literal control character.
const ANSI_PATTERN = new RegExp(`${String.fromCharCode(27)}\\[[0-9;]*m`, 'g');
/**
* Physical terminal rows a frame occupies, so the live redraw moves the cursor up by the right
* amount. A line wider than the terminal wraps onto extra rows, so counting logical lines alone
* undercounts and the redraw drifts downward. Color escapes don't take screen columns, so strip them.
*/
function physicalRows(frame: string): number {
const columns = process.stdout.columns || 80;
return frame.split('\n').reduce((rows, line) => {
const width = line.replace(ANSI_PATTERN, '').length;
return rows + Math.max(1, Math.ceil(width / columns));
}, 0);
}
function exitCodeFor(input: RenderInput): number {
if (input.temporalStatus === 'FAILED' || input.temporalStatus === 'TIMED_OUT') return 1;
if (input.state?.status === 'failed') return 1;
return 0;
}
/** Live view of a running scan: its progress query plus the in-flight agents from describe. */
async function buildRunningInput(workspace: string, workflowId: string, desc: ScanDescription): Promise<RenderInput> {
const state = await queryProgress(workflowId);
return {
workspace,
workflowId,
temporalStatus: desc.status,
state,
running: desc.runningAgents,
...(desc.startedAt !== undefined && { startedAt: desc.startedAt }),
};
}
/** Final view of a closed scan: its result (or the failure) plus timing from describe. */
async function buildTerminalInput(workspace: string, workflowId: string, desc: ScanDescription): Promise<RenderInput> {
const outcome = await getTerminalOutcome(workflowId);
const timing = {
...(desc.startedAt !== undefined && { startedAt: desc.startedAt }),
...(desc.closedAt !== undefined && { endedAt: desc.closedAt }),
};
if (outcome.kind === 'success') {
return { workspace, workflowId, temporalStatus: desc.status, state: outcome.state, running: [], ...timing };
export function status(): void {
// 1. Temporal health
const temporalUp = isTemporalReady();
console.log(`Temporal: ${temporalUp ? 'running' : 'not running'}`);
if (temporalUp) {
console.log(' Dashboard: http://localhost:8233');
}
return {
workspace,
workflowId,
temporalStatus: desc.status,
state: null,
running: [],
failureMessage: outcome.message,
...timing,
};
}
console.log('');
function printFrame(input: RenderInput): void {
const frame = renderScan(input, {
now: Date.now(),
color: supportsColor(),
unicode: stdoutIsTerminal(),
live: false,
frame: 0,
});
process.stdout.write(`${frame}\n`);
}
/**
* Poll Temporal and redraw until the scan reaches a terminal state, then print the
* final frame and exit. A fast ticker animates the running spinner off the cached
* snapshot; the network poll refreshes that snapshot on a slower cadence.
*/
async function watch(workspace: string, workflowId: string): Promise<never> {
let prevRows = 0;
let frame = 0;
let cached: RenderInput | null = null;
const draw = (input: RenderInput, live: boolean): void => {
const out = renderScan(input, { now: Date.now(), color: supportsColor(), unicode: true, live, frame });
if (prevRows > 0) process.stdout.write(`\x1b[${prevRows}A\x1b[0J`);
process.stdout.write(`${out}\n`);
prevRows = physicalRows(out);
};
process.on('exit', () => process.stdout.write(SHOW_CURSOR));
process.on('SIGINT', () => {
process.stdout.write('\n');
process.exit(0);
});
process.stdout.write(HIDE_CURSOR);
const ticker = setInterval(() => {
frame++;
if (cached) draw(cached, true);
}, RENDER_MS);
for (;;) {
const desc = await describeScan(workflowId);
if (!desc) {
clearInterval(ticker);
fail(`Scan "${workspace}" is no longer in Temporal.`);
}
if (isTerminalStatus(desc.status)) {
clearInterval(ticker);
const input = await buildTerminalInput(workspace, workflowId, desc);
draw(input, false);
process.exit(exitCodeFor(input));
}
cached = await buildRunningInput(workspace, workflowId, desc);
await sleep(POLL_MS);
// 2. Running scans
const workers = listRunningWorkers();
if (workers) {
console.log('Running scans:');
console.log(workers);
} else {
console.log('No scans running');
}
}
/** Read one point-in-time snapshot from Temporal: the terminal result if closed, else live progress. */
async function snapshot(workspace: string, workflowId: string, desc: ScanDescription): Promise<RenderInput> {
return isTerminalStatus(desc.status)
? buildTerminalInput(workspace, workflowId, desc)
: buildRunningInput(workspace, workflowId, desc);
}
export async function status(workspace: string, opts: { readonly json: boolean }): Promise<void> {
// A resume spawns a new workflow id (recorded in session.json); resolve through there so status
// follows the current resume, not the superseded original. Fresh scans: the name is the id.
const workflowId = resolveWorkflowId(workspace) ?? workspace;
let desc: ScanDescription | null;
try {
desc = await describeScan(workflowId);
} catch {
fail('Could not reach Temporal at 127.0.0.1:7233.', 'Start Temporal (it comes up with a scan) and try again.');
}
if (!desc) {
fail(
`No scan found for "${workspace}".`,
'',
'Scans are visible while running and for ~24h after they finish (Temporal retention).',
);
}
// --json is always a single snapshot then exit, even on a TTY — it never enters the live watch loop.
if (opts.json) {
const input = await snapshot(workspace, workflowId, desc);
process.stdout.write(`${JSON.stringify(toStatusJson(input, Date.now()), null, 2)}\n`);
process.exit(exitCodeFor(input));
}
// Human-facing views open with the splash; skip it off a real terminal so piped output stays clean.
if (stdoutIsTerminal()) {
displaySplash(isLocal() ? undefined : getVersion());
}
// A finished scan, or output that isn't a live terminal, gets a single frame.
if (isTerminalStatus(desc.status) || !stdoutIsTerminal()) {
const input = await snapshot(workspace, workflowId, desc);
printFrame(input);
process.exit(exitCodeFor(input));
}
await watch(workspace, workflowId);
}
+14 -118
View File
@@ -1,127 +1,23 @@
/**
* `shannon stop` command — stop one scan by workspace, or every scan with --all.
* Never touches infra or data; to wipe Temporal state entirely, use `shannon reset`.
* `shannon stop` command — stop workers and infrastructure.
*/
import * as p from '@clack/prompts';
import { confirmOrExit } from '../confirm.js';
import {
anyRunningScanWorkflow,
ensureDocker,
isTemporalReady,
isWorkflowRunning,
runningContainers,
scanFilter,
stopContainers,
terminateAllWorkflows,
terminateWorkflow,
WORKER_FILTER,
} from '../docker.js';
import { fail, failUsage, warn } from '../errors.js';
import { commandPrefix } from '../mode.js';
import { resolveWorkflowId } from '../session.js';
import { stopInfra, stopWorkers } from '../docker.js';
import { requireInteractive } from '../tty.js';
export interface StopOptions {
all: boolean;
yes: boolean;
workspace?: string;
}
/**
* Stop a single scan. Terminating the workflow both clears Temporal's record and
* brings the container down (the worker waits on the workflow result), so that runs
* first; `docker stop` is the fallback for the pre-registration window and an
* unreachable Temporal. The stop is then verified rather than assumed.
*/
async function stopSingleScan(workspace: string, yes: boolean): Promise<void> {
const workflowId = resolveWorkflowId(workspace);
const filter = scanFilter(workspace);
const temporalUp = isTemporalReady();
const initialContainers = runningContainers(filter);
const workflowRunning = Boolean(workflowId && temporalUp && isWorkflowRunning(workflowId));
// Resolve what is running before prompting, so we never confirm a no-op.
if (initialContainers.length === 0 && !workflowRunning) {
if (!workflowId) {
fail(`No scan found for workspace: ${workspace}`);
export async function stop(clean: boolean, yes: boolean): Promise<void> {
if (clean && !yes) {
requireInteractive('stop --clean', 'Re-run with --yes to skip this confirmation.');
const confirmed = await p.confirm({
message: 'This will stop all running scans and remove the Temporal data. Continue?',
});
if (p.isCancel(confirmed) || !confirmed) {
p.cancel('Aborted.');
process.exit(0);
}
console.log(`Nothing was running for ${workspace}.`);
return;
}
await confirmOrExit('stop', `Stop the scan "${workspace}"?`, yes);
const spinner = p.spinner();
spinner.start(`Stopping scan ${workspace}`);
if (workflowId && workflowRunning) {
terminateWorkflow(workflowId, `Stopped via shannon stop ${workspace}`);
}
await stopContainers(runningContainers(filter));
const stillRunning = runningContainers(filter);
if (stillRunning.length > 0) {
spinner.error(`Scan ${workspace} may still be running`);
console.error(`${stillRunning.length} container(s) did not stop. Retry: ${commandPrefix()} stop ${workspace}`);
process.exit(1);
}
spinner.stop(`Stopped scan ${workspace}`);
if (workflowId && temporalUp && isWorkflowRunning(workflowId)) {
warn(`scan ${workspace} stopped, but its workflow is still Running in Temporal.`);
}
}
async function stopAllScans(yes: boolean): Promise<void> {
const temporalUp = isTemporalReady();
const initial = runningContainers(WORKER_FILTER);
// Resolve what is running before prompting, so we never confirm a no-op.
if (initial.length === 0) {
console.log('No running scans to stop.');
return;
}
await confirmOrExit('stop', 'This will stop all running scans. Continue?', yes);
const spinner = p.spinner();
spinner.start('Stopping all scans');
if (temporalUp) {
terminateAllWorkflows('Stopped via shannon stop --all');
}
await stopContainers(runningContainers(WORKER_FILTER));
const stillRunning = runningContainers(WORKER_FILTER);
if (stillRunning.length > 0) {
spinner.error(`Stopped ${initial.length - stillRunning.length} of ${initial.length} scans`);
console.error(`${stillRunning.length} container(s) did not stop. Retry: ${commandPrefix()} stop --all`);
process.exit(1);
}
spinner.stop(`Stopped ${initial.length} scan${initial.length === 1 ? '' : 's'}`);
if (temporalUp && anyRunningScanWorkflow()) {
warn('some scan workflows are still Running in Temporal — check http://localhost:8233');
}
}
export async function stop(opts: StopOptions): Promise<void> {
ensureDocker();
// Validate the target: exactly one of <workspace> or --all.
if (opts.all && opts.workspace) {
failUsage('Pass a workspace name or --all, not both.');
}
if (!opts.all && !opts.workspace) {
failUsage('Specify which scan to stop: `stop <workspace>`, or `stop --all` to stop every scan.');
}
if (opts.workspace) {
await stopSingleScan(opts.workspace, opts.yes);
} else {
await stopAllScans(opts.yes);
}
stopWorkers();
stopInfra(clean);
}
+55
View File
@@ -0,0 +1,55 @@
/**
* `npx @keygraph/shannon uninstall` command — remove ~/.shannon/ after confirmation (npx only).
*/
import fs from 'node:fs';
import os from 'node:os';
import path from 'node:path';
import * as p from '@clack/prompts';
import { stopInfra, stopWorkers } from '../docker.js';
import { requireInteractive } from '../tty.js';
const SHANNON_HOME = path.join(os.homedir(), '.shannon');
export async function uninstall(yes: boolean): Promise<void> {
const interactive = !yes;
if (interactive) p.intro('Shannon Uninstall');
if (!fs.existsSync(SHANNON_HOME)) {
const message = 'Nothing to remove. Shannon is not configured on this machine.';
if (interactive) {
p.log.info(message);
p.outro('Done.');
} else {
console.log(message);
}
return;
}
if (interactive) {
requireInteractive('uninstall', 'Re-run with --yes to skip this confirmation.');
const confirmed = await p.confirm({
message: 'This will permanently remove all past scan data, saved configurations, and API keys. Continue?',
});
if (p.isCancel(confirmed) || !confirmed) {
p.cancel('Aborted.');
process.exit(0);
}
}
// Stop any running containers first
stopWorkers();
stopInfra(false);
fs.rmSync(SHANNON_HOME, { recursive: true, force: true });
const done = 'All Shannon data has been removed.';
const hint = 'Shannon has been uninstalled. Run `npx @keygraph/shannon setup` to start fresh.';
if (interactive) {
p.log.success(done);
p.outro(hint);
} else {
console.log(done);
console.log(hint);
}
}
+35
View File
@@ -0,0 +1,35 @@
/**
* `shannon workspaces` command — list all workspaces.
*/
import { execFileSync } from 'node:child_process';
import os from 'node:os';
import { getWorkerImage } from '../docker.js';
import { getWorkspacesDir } from '../home.js';
export function workspaces(version: string): void {
const workspacesDir = getWorkspacesDir();
const image = getWorkerImage(version);
try {
execFileSync(
'docker',
[
'run',
'--rm',
'-v',
`${workspacesDir}:/app/workspaces`,
'-e',
'WORKSPACES_DIR=/app/workspaces',
image,
'node',
'apps/worker/dist/temporal/workspaces.js',
],
{ stdio: 'inherit', ...(os.platform() === 'win32' && { env: { ...process.env, MSYS_NO_PATHCONV: '1' } }) },
);
} catch {
console.error('ERROR: Failed to list workspaces. Is the Docker image available?');
console.error(` Run: docker pull ${image}`);
process.exit(1);
}
}
+80 -82
View File
@@ -7,16 +7,8 @@
import fs from 'node:fs';
import { parse as parseTOML } from 'smol-toml';
import { fail } from '../errors.js';
import { getConfigFile } from '../home.js';
import { getMode } from '../mode.js';
import {
type CuratedProviderId,
DEFAULT_MODEL_SPEC,
GENERIC_API_KEY_ENV,
isCuratedProvider,
parseModelSpec,
} from '../model-spec.js';
// === TOML ↔ Env Mapping ===
@@ -31,40 +23,28 @@ interface ConfigMapping {
/** Maps every supported env var to its TOML path (section.key) and expected type. */
const CONFIG_MAP: readonly ConfigMapping[] = [
// Core — base_url points any provider at a proxy or gateway
{ env: 'SHANNON_AI_MODEL', toml: 'core.model', type: 'string' },
{ env: 'SHANNON_AI_BASE_URL', toml: 'core.base_url', type: 'string' },
// Core
{ env: 'CLAUDE_ADAPTIVE_THINKING', toml: 'core.adaptive_thinking', type: 'boolean', boolFormat: 'literal' },
// Anthropic
{ env: 'ANTHROPIC_API_KEY', toml: 'anthropic.api_key', type: 'string' },
{ env: 'CLAUDE_CODE_OAUTH_TOKEN', toml: 'anthropic.oauth_token', type: 'string' },
// OpenAI — format picks the wire API a gateway serves
{ env: 'OPENAI_API_KEY', toml: 'openai.api_key', type: 'string' },
{ env: 'SHANNON_AI_OPENAI_FORMAT', toml: 'openai.format', type: 'string' },
// xAI
{ env: 'XAI_API_KEY', toml: 'xai.api_key', type: 'string' },
// Bedrock
{ env: 'CLAUDE_CODE_USE_BEDROCK', toml: 'bedrock.use', type: 'boolean' },
{ env: 'AWS_REGION', toml: 'bedrock.region', type: 'string' },
{ env: 'AWS_BEARER_TOKEN_BEDROCK', toml: 'bedrock.token', type: 'string' },
// Generic — credential for any provider Shannon does not curate
{ env: GENERIC_API_KEY_ENV, toml: 'provider.api_key', type: 'string' },
// Custom Base URL
{ env: 'ANTHROPIC_BASE_URL', toml: 'custom_base_url.base_url', type: 'string' },
{ env: 'ANTHROPIC_AUTH_TOKEN', toml: 'custom_base_url.auth_token', type: 'string' },
// Model tiers
{ env: 'ANTHROPIC_SMALL_MODEL', toml: 'models.small', type: 'string' },
{ env: 'ANTHROPIC_MEDIUM_MODEL', toml: 'models.medium', type: 'string' },
{ env: 'ANTHROPIC_LARGE_MODEL', toml: 'models.large', type: 'string' },
] as const;
/** TOML section holding each curated provider's credentials, keyed by provider id. */
const PROVIDER_SECTIONS: Readonly<Record<CuratedProviderId, string>> = {
anthropic: 'anthropic',
openai: 'openai',
xai: 'xai',
'amazon-bedrock': 'bedrock',
};
/** TOML section holding the generic credential for uncurated providers. */
const GENERIC_PROVIDER_SECTION = 'provider';
// === TOML Parsing ===
type TOMLValue = string | number | boolean;
@@ -101,9 +81,10 @@ function loadTOML(): TOMLConfig | null {
const mode = fs.statSync(configPath).mode;
if (mode & 0o077) {
const actual = (mode & 0o777).toString(8).padStart(3, '0');
fail(
`Your config file is readable by other users on this machine (${actual}). Lock it down: chmod 600 ${configPath}`,
console.error(
`\nYour config file is readable by other users on this machine (${actual}). Lock it down: chmod 600 ${configPath}\n`,
);
process.exit(1);
}
}
@@ -112,7 +93,9 @@ function loadTOML(): TOMLConfig | null {
return parseTOML(content) as TOMLConfig;
} catch (err) {
const message = err instanceof Error ? err.message : String(err);
fail(`Failed to parse ${configPath}: ${message}`, `Run 'npx @keygraph/shannon setup' to reconfigure.`);
console.error(`\nFailed to parse ${configPath}: ${message}`);
console.error(`\nRun 'npx @keygraph/shannon setup' to reconfigure.\n`);
process.exit(1);
}
}
@@ -135,42 +118,52 @@ function buildSchema(): Map<string, Map<string, TOMLType>> {
return schema;
}
/**
* Check that the section backing the selected provider carries a usable
* credential. `core.model` names the provider, so only that section is required;
* other providers' sections are ignored and never forwarded. An uncurated
* provider draws its credential from the generic [provider] section.
*/
function validateProviderFields(config: TOMLConfig, providerId: string, errors: string[]): void {
if (!isCuratedProvider(providerId)) {
const section = config[GENERIC_PROVIDER_SECTION] as Record<string, unknown> | undefined;
if (!section || !Object.keys(section).includes('api_key')) {
errors.push(`[${GENERIC_PROVIDER_SECTION}] requires api_key for provider "${providerId}"`);
/** Check that a provider section has all required fields and dependencies. */
function validateProviderFields(config: TOMLConfig, provider: string, errors: string[]): void {
const section = config[provider] as Record<string, unknown> | undefined;
if (!section) return;
const keys = Object.keys(section);
switch (provider) {
case 'anthropic':
if (!keys.includes('api_key') && !keys.includes('oauth_token')) {
errors.push('[anthropic] requires either api_key or oauth_token');
}
break;
case 'custom_base_url': {
const required = ['base_url', 'auth_token'];
const missing = required.filter((k) => !keys.includes(k));
if (missing.length > 0) {
errors.push(`[custom_base_url] missing required keys: ${missing.join(', ')}`);
}
break;
}
case 'bedrock': {
const required = ['use', 'region', 'token'];
const missing = required.filter((k) => !keys.includes(k));
if (missing.length > 0) {
errors.push(`[bedrock] missing required keys: ${missing.join(', ')}`);
}
validateModelTiers(config, 'bedrock', errors);
break;
}
}
}
/** Bedrock requires a [models] section with all three tiers. */
function validateModelTiers(config: TOMLConfig, provider: string, errors: string[]): void {
const models = config.models as Record<string, unknown> | undefined;
if (!models || typeof models !== 'object') {
errors.push(`[${provider}] requires a [models] section with small, medium, and large`);
return;
}
const sectionName = PROVIDER_SECTIONS[providerId];
const section = config[sectionName] as Record<string, unknown> | undefined;
const keys = section ? Object.keys(section) : [];
if (providerId === 'amazon-bedrock') {
const missing = ['region', 'token'].filter((k) => !keys.includes(k));
if (missing.length > 0) {
errors.push(`[bedrock] missing required keys: ${missing.join(', ')}`);
}
return;
}
if (providerId === 'anthropic') {
if (!keys.includes('api_key') && !keys.includes('oauth_token')) {
errors.push('[anthropic] requires either api_key or oauth_token');
}
return;
}
if (!keys.includes('api_key')) {
errors.push(`[${sectionName}] requires api_key`);
const required = ['small', 'medium', 'large'];
const missing = required.filter((k) => !Object.keys(models).includes(k));
if (missing.length > 0) {
errors.push(`[models] missing required keys for ${provider}: ${missing.join(', ')}`);
}
}
@@ -218,19 +211,23 @@ function validateConfig(config: TOMLConfig): string[] {
}
}
// 4. core.model must parse and name a supported provider
const modelValue = config.core?.model;
if (modelValue !== undefined && typeof modelValue !== 'string') {
return errors;
}
const spec = parseModelSpec(modelValue || DEFAULT_MODEL_SPEC);
if (typeof spec === 'string') {
errors.push(`[core].model — ${spec}`);
return errors;
// 4. Only one provider section allowed (ignore empty sections)
const PROVIDER_SECTIONS = ['anthropic', 'custom_base_url', 'bedrock'] as const;
const present = PROVIDER_SECTIONS.filter((s) => {
const section = config[s];
return section && typeof section === 'object' && Object.keys(section).length > 0;
});
if (present.length > 1) {
errors.push(
`Multiple providers configured: [${present.join('], [')}]. Only one provider section is allowed at a time`,
);
}
// 5. The selected provider's section must carry a credential
validateProviderFields(config, spec.providerId, errors);
// 5. Required fields per provider
const singleProvider = present.length === 1 ? present[0] : undefined;
if (singleProvider) {
validateProviderFields(config, singleProvider, errors);
}
return errors;
}
@@ -254,11 +251,12 @@ export function resolveConfig(): void {
// Validate before injecting
const errors = validateConfig(toml);
if (errors.length > 0) {
fail(
'Invalid configuration:',
...errors.map((err) => ` - ${err}`),
`Run 'npx @keygraph/shannon setup' to reconfigure.`,
);
console.error('\nInvalid configuration:');
for (const err of errors) {
console.error(` - ${err}`);
}
console.error(`\nRun 'npx @keygraph/shannon setup' to reconfigure.\n`);
process.exit(1);
}
for (const mapping of CONFIG_MAP) {
+4 -6
View File
@@ -8,13 +8,11 @@ import { getConfigFile } from '../home.js';
// === Types ===
export interface ShannonConfig {
core?: { model?: string; base_url?: string };
core?: { adaptive_thinking?: boolean };
anthropic?: { api_key?: string; oauth_token?: string };
openai?: { api_key?: string; format?: string };
xai?: { api_key?: string };
bedrock?: { region?: string; token?: string };
/** Generic credential for any provider Shannon does not curate. Maps to SHANNON_AI_API_KEY. */
provider?: { api_key?: string };
custom_base_url?: { base_url?: string; auth_token?: string };
bedrock?: { use?: boolean; region?: string; token?: string };
models?: { small?: string; medium?: string; large?: string };
}
// === File Operations ===
-43
View File
@@ -1,43 +0,0 @@
/**
* Shared confirmation prompt for destructive or batch commands.
*
* `stop` and `reset` gate their action behind the same "confirm unless --yes"
* flow. Centralizing it here keeps the behavior identical across commands and
* impossible to change in only one place by accident.
*/
import * as p from '@clack/prompts';
import { requireInteractive } from './tty.js';
/**
* Ask the user to confirm an action, unless `yes` was passed. Off a TTY without
* `--yes`, fails fast rather than hanging on a prompt. Exits 0 if the user declines.
*/
export async function confirmOrExit(command: string, message: string, yes: boolean): Promise<void> {
if (yes) {
return;
}
requireInteractive(command, 'Re-run with --yes to skip this confirmation.');
const confirmed = await p.confirm({ message });
if (p.isCancel(confirmed) || !confirmed) {
p.cancel('Aborted.');
process.exit(0);
}
}
/**
* Severe-tier confirmation: the user must type `word` exactly to proceed. Unlike
* `confirmOrExit` there is no `--yes` bypass. Off a TTY it fails fast; exits 0 if declined.
*/
export async function confirmByTyping(command: string, word: string): Promise<void> {
requireInteractive(command, `'${command}' cannot be run non-interactively.`);
const typed = await p.text({
message: `Type ${word} to confirm — this cannot be undone:`,
validate: (value) => (value === word ? undefined : `Type ${word} to proceed, or press Ctrl-C to abort.`),
});
if (p.isCancel(typed) || typed !== word) {
p.cancel('Aborted.');
process.exit(0);
}
}
+58 -168
View File
@@ -12,36 +12,18 @@ import os from 'node:os';
import path from 'node:path';
import { setTimeout as sleep } from 'node:timers/promises';
import { fileURLToPath } from 'node:url';
import type { SpinnerResult } from '@clack/prompts';
import { envBool, PI_AUTH_CONTAINER_PATH } from './env.js';
import { fail } from './errors.js';
import { getMode, isDevMode } from './mode.js';
import { getMode } from './mode.js';
import { INTERNAL_DIR } from './paths.js';
import { runStep, spawnCaptured, surfaceOutput } from './ui.js';
const __dirname = path.dirname(fileURLToPath(import.meta.url));
const NPX_IMAGE_REPO = 'keygraph/shannon';
const DEV_IMAGE = 'shannon-worker';
/** Docker label stamped on each worker container, mapping it back to its workspace so a single scan can be stopped by name. */
const WORKSPACE_LABEL = 'shannon.workspace';
export function getWorkerImage(version: string): string {
return getMode() === 'local' ? DEV_IMAGE : `${NPX_IMAGE_REPO}:${version}`;
}
/** True when the working directory supplies a Dockerfile and build context. */
export function canBuildImage(): boolean {
if (getMode() === 'local') return true;
if (!isDevMode()) return false;
const hasDockerfile = fs.existsSync(path.resolve('Dockerfile'));
const hasCompose = fs.existsSync(path.resolve('docker-compose.yml'));
return hasDockerfile && hasCompose;
}
function getComposeFile(): string {
return getMode() === 'local'
? path.resolve('docker-compose.yml')
@@ -72,116 +54,80 @@ function runOutput(cmd: string, args: string[]): string {
}
}
/** Run a command asynchronously, resolving true on success. Never rejects. */
function spawnQuiet(cmd: string, args: string[]): Promise<boolean> {
return new Promise((resolve) => {
const child = spawn(cmd, args, { stdio: 'ignore' });
child.on('close', (code) => resolve(code === 0));
child.on('error', () => resolve(false));
});
}
const TEMPORAL_CONTAINER = 'shannon-temporal';
const TEMPORAL_ADDRESS = 'localhost:7233';
/** Query matching every running pentest scan workflow. */
const RUNNING_SCAN_QUERY = "ExecutionStatus = 'Running' AND WorkflowType = 'pentestPipelineWorkflow'";
/** Build `docker exec` args for a `temporal` CLI command run inside the Temporal container. */
function temporalCmd(...args: string[]): string[] {
return ['exec', TEMPORAL_CONTAINER, 'temporal', ...args, '--address', TEMPORAL_ADDRESS];
}
/**
* Verify Docker is installed and its daemon is running, exiting otherwise.
* `docker info` succeeds only when both are true. Call this before any command
* that shells out to Docker.
*/
export function ensureDocker(): void {
try {
execFileSync('docker', ['info'], { stdio: 'pipe' });
} catch {
fail(
'Docker must be installed and running. Start Docker and try again.',
'Install Docker: https://docs.docker.com/get-docker/',
);
}
}
/**
* Check if Temporal is running and healthy.
*/
export function isTemporalReady(): boolean {
const output = runOutput('docker', temporalCmd('operator', 'cluster', 'health'));
const output = runOutput('docker', [
'exec',
'shannon-temporal',
'temporal',
'operator',
'cluster',
'health',
'--address',
'localhost:7233',
]);
return output.includes('SERVING');
}
/**
* Ensure Temporal is running via compose.
*/
export async function ensureInfra(spinner: SpinnerResult): Promise<void> {
export async function ensureInfra(): Promise<void> {
if (isTemporalReady()) {
return;
}
// Drive the caller's spinner — the whole "start" flow is one spinner, not several.
spinner.message('Starting Temporal');
const composeFile = getComposeFile();
const result = await spawnCaptured('docker', ['compose', '-f', composeFile, 'up', '-d']);
if (!result.ok) {
spinner.error('Could not start Temporal');
surfaceOutput(result.output);
process.exit(1);
}
console.log('Starting Shannon infrastructure...');
execFileSync('docker', ['compose', '-f', composeFile, 'up', '-d'], { stdio: 'inherit' });
spinner.message('Waiting for Temporal to be ready');
console.log('Waiting for Temporal to be ready...');
for (let i = 0; i < 30; i++) {
if (isTemporalReady()) {
console.log('Temporal is ready!');
return;
}
await sleep(2000);
}
spinner.error('Temporal did not become ready in time');
console.error('Timeout waiting for Temporal');
process.exit(1);
}
/**
* Build the worker image from the repository, tagged with the name this mode
* resolves at run time.
* Build the worker image locally (local mode only).
*/
export function buildImage(noCache: boolean, version: string): void {
const image = getWorkerImage(version);
console.log(`Building ${image}...`);
export function buildImage(noCache: boolean): void {
console.log(`Building ${DEV_IMAGE}...`);
const args = ['build'];
if (noCache) args.push('--no-cache');
args.push('-t', image, '.');
args.push('-t', DEV_IMAGE, '.');
execFileSync('docker', args, { stdio: 'inherit' });
console.log(`Build complete: ${image}`);
console.log(`Build complete: ${DEV_IMAGE}`);
}
/**
* Ensure the worker image is available.
* Buildable checkout: auto-builds if missing. Otherwise: pulls from Docker Hub.
* Local mode: auto-builds if missing. NPX mode: pulls from Docker Hub.
*/
export function ensureImage(version: string): void {
const image = getWorkerImage(version);
const exists = runQuiet('docker', ['image', 'inspect', image]);
if (exists) return;
if (canBuildImage()) {
if (getMode() === 'local') {
console.log('Shannon image not found, building...');
buildImage(false, version);
buildImage(false);
} else {
console.log(`Pulling ${image}...`);
try {
execFileSync('docker', ['pull', image], { stdio: 'inherit' });
} catch {
fail(
`Failed to pull ${image}`,
'The image may not be available for your platform yet.',
'Check https://hub.docker.com/r/keygraph/shannon for available tags.',
);
console.error(`\nERROR: Failed to pull ${image}`);
console.error('The image may not be available for your platform yet.');
console.error('Check https://hub.docker.com/r/keygraph/shannon for available tags.');
process.exit(1);
}
pruneOldImages(version);
}
@@ -244,7 +190,7 @@ function shouldSkipHostsName(name: string, hostname: string): boolean {
* `host-gateway` so they target the host's loopback instead of the container's.
*/
function forwardEtcHostsFlags(): string[] {
if (!envBool('SHANNON_FORWARD_HOSTS', true)) return [];
if (process.env.SHANNON_FORWARD_HOSTS === 'false') return [];
if (os.platform() === 'win32') return [];
let content: string;
@@ -295,24 +241,20 @@ export interface WorkerOptions {
outputDir?: string;
workspace: string;
pipelineTesting?: boolean;
keepContainer?: boolean;
piAuthHostPath?: string;
debug?: boolean;
}
/**
* Spawn the worker container in detached mode and return the process.
* When `opts.keepContainer` is true, omits `--rm` so the container persists for log inspection.
* When `opts.debug` is true, omits `--rm` so the container persists for log inspection.
*/
export function spawnWorker(opts: WorkerOptions): ChildProcess {
const args = ['run', '-d'];
if (!opts.keepContainer) {
if (!opts.debug) {
args.push('--rm');
}
args.push('--name', opts.containerName, '--network', 'shannon-net');
// Tag with the workspace so `stop <workspace>` can target this scan's container
args.push('--label', `${WORKSPACE_LABEL}=${opts.workspace}`);
// Add host flag for Linux
args.push(...addHostFlag());
@@ -350,11 +292,6 @@ export function spawnWorker(opts: WorkerOptions): ChildProcess {
args.push('-v', `${opts.outputDir}:/app/output`);
}
// Reuse the host's pi credentials: mount only the auth file, allowing token refreshes to persist.
if (opts.piAuthHostPath) {
args.push('-v', `${opts.piAuthHostPath}:${PI_AUTH_CONTAINER_PATH}`);
}
// Environment
args.push(...opts.envFlags);
@@ -387,86 +324,26 @@ export function spawnWorker(opts: WorkerOptions): ChildProcess {
});
}
/** `docker ps --filter` args matching every running worker container. */
export const WORKER_FILTER: readonly string[] = ['--filter', 'name=shannon-worker-'];
/**
* Stop all running shannon-worker-* containers.
*/
export function stopWorkers(): void {
const workers = runOutput('docker', ['ps', '-q', '--filter', 'name=shannon-worker-']);
if (!workers) return;
/** `docker ps --filter` args matching one scan's worker container(s), by workspace label. */
export function scanFilter(workspace: string): readonly string[] {
return ['--filter', `label=${WORKSPACE_LABEL}=${workspace}`];
const ids = workers.split('\n').filter(Boolean);
console.log('Stopping running scans...');
execFileSync('docker', ['stop', ...ids], { stdio: 'inherit' });
}
/**
* IDs of running containers matching the filter. Re-querying this after a stop is
* the authoritative check for whether containers actually stopped — `docker stop`'s
* exit code can't distinguish "already gone" from "failed to stop".
* Tear down the compose stack.
*/
export function runningContainers(filter: readonly string[]): string[] {
const output = runOutput('docker', ['ps', '-q', ...filter]);
return output.split('\n').filter(Boolean);
}
/**
* Stop containers by ID, tolerating any that vanished between being listed and
* stopped (a `--rm` worker exiting is success, not an error). Async so a spinner
* can animate during docker's graceful-shutdown wait.
*/
export async function stopContainers(ids: string[]): Promise<void> {
await Promise.all(ids.map((id) => spawnQuiet('docker', ['stop', id])));
}
/**
* Terminate a Temporal workflow so a stopped scan doesn't linger as a running
* workflow with no worker. Best-effort: returns false if Temporal is unreachable
* or the workflow already closed. Requires Temporal to be up (guard with isTemporalReady).
*/
export function terminateWorkflow(workflowId: string, reason: string): boolean {
return runQuiet('docker', temporalCmd('workflow', 'terminate', '--workflow-id', workflowId, '--reason', reason));
}
/**
* Terminate every running pentest workflow in one batch, so `stop --all` doesn't
* leave workflows running with no worker. Best-effort: returns false if Temporal
* is unreachable. Requires Temporal to be up (guard with isTemporalReady).
*/
export function terminateAllWorkflows(reason: string): boolean {
return runQuiet(
'docker',
temporalCmd('workflow', 'terminate', '--query', RUNNING_SCAN_QUERY, '--reason', reason, '--yes'),
);
}
/**
* Whether a specific workflow is still in the Running state. Re-querying this after
* a terminate verifies it actually took effect, rather than trusting the terminate
* command's exit code. Requires Temporal to be up (guard with isTemporalReady).
*/
export function isWorkflowRunning(workflowId: string): boolean {
const query = `WorkflowId = '${workflowId}' AND ExecutionStatus = 'Running'`;
const output = runOutput('docker', temporalCmd('workflow', 'list', '--query', query));
return output.includes(workflowId);
}
/**
* Whether any pentest scan workflow is still Running — the `stop --all` counterpart
* to isWorkflowRunning. Requires Temporal to be up (guard with isTemporalReady).
*/
export function anyRunningScanWorkflow(): boolean {
const output = runOutput('docker', temporalCmd('workflow', 'list', '--query', RUNNING_SCAN_QUERY));
return output.includes('pentestPipelineWorkflow');
}
/**
* Tear down the compose stack. When `clean` is set, volumes are removed too.
*/
export async function stopInfra(clean: boolean): Promise<void> {
export function stopInfra(clean: boolean): void {
const composeFile = getComposeFile();
const args = ['compose', '-f', composeFile, 'down'];
if (clean) args.push('-v');
const label = clean ? 'Removing Temporal data and volumes' : 'Stopping Temporal';
const step = await runStep(label, 'docker', args);
if (!step.ok) {
fail(`${label} failed. See the output above.`);
}
execFileSync('docker', args, { stdio: 'inherit' });
}
/**
@@ -482,3 +359,16 @@ function pruneOldImages(currentVersion: string): void {
runQuiet('docker', ['rmi', `${NPX_IMAGE_REPO}:${tag}`]);
}
}
/**
* List running worker containers.
*/
export function listRunningWorkers(): string {
return runOutput('docker', [
'ps',
'--filter',
'name=shannon-worker-',
'--format',
'table {{.Names}}\t{{.Status}}\t{{.RunningFor}}',
]);
}
+68 -140
View File
@@ -5,73 +5,25 @@
* NPX mode: fills gaps from ~/.shannon/config.toml (no .env).
*/
import fs from 'node:fs';
import os from 'node:os';
import path from 'node:path';
import dotenv from 'dotenv';
import { resolveConfig } from './config/resolver.js';
import { getMode } from './mode.js';
import {
CURATED_PROVIDERS,
type CuratedProviderId,
GENERIC_API_KEY_ENV,
isCuratedProvider,
PROVIDER_API_KEY_ENV,
PROVIDER_CREDENTIAL_HINT,
PROVIDER_EXTRA_ENV,
resolveModelSpec,
} from './model-spec.js';
/**
* Variables forwarded to every worker container regardless of provider. Each is
* forwarded only when set, so an unused one never appears in the container.
* SHANNON_AI_API_KEY rides along because it is provider-neutral.
*/
const COMMON_FORWARD_VARS = [
'SHANNON_AI_MODEL',
'SHANNON_AI_BASE_URL',
'SHANNON_AI_OPENAI_FORMAT',
GENERIC_API_KEY_ENV,
/** Environment variables forwarded to worker containers. */
const FORWARD_VARS = [
'ANTHROPIC_API_KEY',
'ANTHROPIC_BASE_URL',
'ANTHROPIC_AUTH_TOKEN',
'CLAUDE_CODE_OAUTH_TOKEN',
'CLAUDE_CODE_USE_BEDROCK',
'AWS_REGION',
'AWS_BEARER_TOKEN_BEDROCK',
'ANTHROPIC_SMALL_MODEL',
'ANTHROPIC_MEDIUM_MODEL',
'ANTHROPIC_LARGE_MODEL',
'CLAUDE_ADAPTIVE_THINKING',
] as const;
/**
* Credential variables for one provider. Only the selected provider's entries are
* forwarded, so a key for an unused provider never enters the scan container. An
* uncurated provider has none — it relies on the common SHANNON_AI_API_KEY.
*/
function providerForwardVars(providerId: string): readonly string[] {
if (!isCuratedProvider(providerId)) return [];
return [...PROVIDER_API_KEY_ENV[providerId], ...PROVIDER_EXTRA_ENV[providerId]];
}
/** Parse a user-facing boolean env var: `1`/`true` (any case) true, `0`/`false`/empty false, else the default. */
export function envBool(name: string, defaultValue: boolean): boolean {
const raw = process.env[name]?.trim().toLowerCase();
if (raw === undefined || raw === '') return defaultValue;
if (raw === '1' || raw === 'true') return true;
if (raw === '0' || raw === 'false') return false;
return defaultValue;
}
const USE_PI_AUTH_ENV = 'SHANNON_USE_PI_AUTH';
/** Where the host's auth.json is mounted: pi's standard location (worker HOME is /tmp), read natively. */
export const PI_AUTH_CONTAINER_PATH = '/tmp/.pi/agent/auth.json';
/** Host path to pi's credential file. */
export function resolveHostPiAuthPath(): string {
return path.join(os.homedir(), '.pi', 'agent', 'auth.json');
}
export function piAuthFlagEnabled(): boolean {
return envBool(USE_PI_AUTH_ENV, false);
}
/** Opted into pi auth via the flag, and the auth file exists to mount. */
export function shouldUsePiAuth(): boolean {
return piAuthFlagEnabled() && fs.existsSync(resolveHostPiAuthPath());
}
/**
* Load credentials into process.env.
* Local mode: loads ./.env via dotenv.
@@ -87,19 +39,15 @@ export function loadEnv(): void {
}
/**
* Build `-e` flags for docker run. Forwards the common vars plus only the
* selected provider's credentials, passed by name (`-e KEY`) so secret values
* stay out of the `docker run` argv; docker inherits them from this process's env.
* Build `-e KEY=VALUE` flags for docker run, only for set variables.
*/
export function buildEnvFlags(): string[] {
const flags: string[] = ['-e', 'TEMPORAL_ADDRESS=shannon-temporal:7233'];
const spec = resolveModelSpec();
const providerVars = typeof spec === 'string' ? [] : providerForwardVars(spec.providerId);
for (const key of [...COMMON_FORWARD_VARS, ...providerVars]) {
if (process.env[key]) {
flags.push('-e', key);
for (const key of FORWARD_VARS) {
const value = process.env[key];
if (value) {
flags.push('-e', `${key}=${value}`);
}
}
@@ -109,91 +57,71 @@ export function buildEnvFlags(): string[] {
interface CredentialValidation {
valid: boolean;
error?: string;
mode: 'api-key' | 'oauth' | 'custom-base-url' | 'bedrock';
}
/** Whether a curated provider has its own named credential set (API key plus any extra var). */
function hasNamedCredential(providerId: CuratedProviderId): boolean {
const apiKeys = PROVIDER_API_KEY_ENV[providerId];
if (!apiKeys.some((name) => Boolean(process.env[name]))) return false;
return PROVIDER_EXTRA_ENV[providerId].every((name) => Boolean(process.env[name]));
/** Check if a custom Anthropic-compatible base URL is configured. */
function isCustomBaseUrlConfigured(): boolean {
return !!(process.env.ANTHROPIC_BASE_URL && process.env.ANTHROPIC_AUTH_TOKEN);
}
/** Whether the selected provider has a credential. Bedrock needs its AWS_ vars; the generic key never stands in for it. */
function hasCredential(providerId: string): boolean {
if (providerId === 'amazon-bedrock') return hasNamedCredential('amazon-bedrock');
if (isCuratedProvider(providerId) && hasNamedCredential(providerId)) return true;
return Boolean(process.env[GENERIC_API_KEY_ENV]);
}
/** Curated providers with a named credential. The generic key is neutral, so it never counts toward ambiguity. */
function configuredProviders(): CuratedProviderId[] {
return CURATED_PROVIDERS.filter((providerId) => hasNamedCredential(providerId));
/** Detect which providers are configured via environment variables. */
function detectProviders(): string[] {
const providers: string[] = [];
if (process.env.ANTHROPIC_API_KEY) providers.push('Anthropic API key');
if (process.env.CLAUDE_CODE_OAUTH_TOKEN) providers.push('Anthropic OAuth');
if (isCustomBaseUrlConfigured()) providers.push('Custom Base URL');
if (process.env.CLAUDE_CODE_USE_BEDROCK === '1') providers.push('AWS Bedrock');
return providers;
}
/**
* Validate that the model selection parses and its provider has a credential.
* Runs before any Docker work so mistakes fail immediately.
* Validate that exactly one authentication method is configured.
*/
export function validateCredentials(): CredentialValidation {
// 1. Model selection must parse into a provider and model id
const spec = resolveModelSpec();
if (typeof spec === 'string') {
return { valid: false, error: spec };
}
// Pi-auth: skip the API-key checks, but the host auth file must exist to mount.
if (piAuthFlagEnabled()) {
const authPath = resolveHostPiAuthPath();
if (!fs.existsSync(authPath)) {
return {
valid: false,
error: `${USE_PI_AUTH_ENV} is set but no pi credentials were found at ${authPath}. Authenticate with pi first.`,
};
}
return { valid: true };
}
// 2. The selected provider must have a credential
if (!hasCredential(spec.providerId)) {
const requirement = isCuratedProvider(spec.providerId)
? PROVIDER_CREDENTIAL_HINT[spec.providerId]
: GENERIC_API_KEY_ENV;
const hint =
getMode() === 'local'
? `Set ${requirement} in .env or export it.`
: `Export the variables or run 'npx @keygraph/shannon setup'.`;
// Reject multiple providers
const providers = detectProviders();
if (providers.length > 1) {
return {
valid: false,
error: `No credentials found for provider "${spec.providerId}". ${hint}`,
mode: 'api-key',
error: `Multiple providers detected: ${providers.join(', ')}. Only one provider can be active at a time.`,
};
}
// 3. Exactly one provider may be configured. Several complete credentials make
// the scan's provider depend on SHANNON_AI_MODEL alone, which is too easy to
// misread as "both are in play" and too easy to redirect by editing one line.
const configured = configuredProviders();
if (configured.length > 1) {
const setKeys = (id: CuratedProviderId): string[] =>
PROVIDER_API_KEY_ENV[id].filter((name) => Boolean(process.env[name]));
const list = configured.map((id) => `${id} (${setKeys(id).join(', ')})`).join(' and ');
const others = configured.filter((id) => id !== spec.providerId);
const extraVars = others.flatMap(setKeys);
const dropHint =
getMode() === 'local'
? 'remove them from .env or unset them in your shell:'
: "unset them in your shell, or reconfigure with 'npx @keygraph/shannon setup':";
const lines = [`Credentials for more than one provider are set: ${list}.`];
if (extraVars.length > 0) {
lines.push(
`Shannon runs one provider per scan, selected by SHANNON_AI_MODEL ("${spec.providerId}:...").`,
`Keep ${spec.providerId} and drop the rest — ${dropHint}`,
` unset ${extraVars.join(' ')}`,
);
if (process.env.ANTHROPIC_API_KEY) {
return { valid: true, mode: 'api-key' };
}
if (process.env.CLAUDE_CODE_OAUTH_TOKEN) {
return { valid: true, mode: 'oauth' };
}
if (isCustomBaseUrlConfigured()) {
return { valid: true, mode: 'custom-base-url' };
}
if (process.env.CLAUDE_CODE_USE_BEDROCK === '1') {
const missing: string[] = [];
if (!process.env.AWS_REGION) missing.push('AWS_REGION');
if (!process.env.AWS_BEARER_TOKEN_BEDROCK) missing.push('AWS_BEARER_TOKEN_BEDROCK');
if (!process.env.ANTHROPIC_SMALL_MODEL) missing.push('ANTHROPIC_SMALL_MODEL');
if (!process.env.ANTHROPIC_MEDIUM_MODEL) missing.push('ANTHROPIC_MEDIUM_MODEL');
if (!process.env.ANTHROPIC_LARGE_MODEL) missing.push('ANTHROPIC_LARGE_MODEL');
if (missing.length > 0) {
return {
valid: false,
mode: 'bedrock',
error: `Bedrock mode requires: ${missing.join(', ')}`,
};
}
return { valid: false, error: lines.join('\n') };
return { valid: true, mode: 'bedrock' };
}
return { valid: true };
const hint =
getMode() === 'local'
? `No credentials found. Set ANTHROPIC_API_KEY in .env or export it.`
: `Authentication not configured. Export variables or run 'npx @keygraph/shannon setup'.`;
return {
valid: false,
mode: 'api-key',
error: hint,
};
}
-70
View File
@@ -1,70 +0,0 @@
/**
* Centralized error reporting.
*
* `fail` — an expected, user-fixable error (bad input, missing prerequisite):
* a clean message on stderr and a non-zero exit, never a stack trace.
* `failUsage` — a malformed invocation (unknown command, bad or missing
* arguments): the same clean message, but a distinct exit code so callers can
* tell a usage mistake from an operational failure.
* `crash` — an unexpected error (a bug): a brief message, the full stack written
* to a log file for a bug report, and a pointer to the issue tracker.
*/
import fs from 'node:fs';
import os from 'node:os';
import path from 'node:path';
const ISSUES_URL = 'https://github.com/KeygraphHQ/shannon/issues';
/** Report an expected, user-fixable error (with optional extra lines) and exit non-zero. */
export function fail(message: string, ...hints: string[]): never {
console.error(`ERROR: ${message}`);
for (const hint of hints) {
console.error(hint);
}
process.exit(1);
}
/** Report a usage/argument error (with optional extra lines) and exit 2. */
export function failUsage(message: string, ...hints: string[]): never {
console.error(`ERROR: ${message}`);
for (const hint of hints) {
console.error(hint);
}
process.exit(2);
}
/** Report a non-fatal warning on stderr (with optional extra lines) without exiting. */
export function warn(message: string, ...hints: string[]): void {
console.error(`WARNING: ${message}`);
for (const hint of hints) {
console.error(hint);
}
}
/** Report an unexpected error: brief message, full stack to a log file, plus the issue link. */
export function crash(error: unknown): never {
console.error(`ERROR: ${error instanceof Error ? error.message : String(error)}`);
if (process.env.DEBUG) {
console.error(error instanceof Error ? error.stack : String(error));
}
const logPath = writeCrashLog(error);
if (logPath) {
console.error(`Details written to ${logPath}`);
}
console.error(`If this looks like a bug, please report it: ${ISSUES_URL}`);
process.exit(1);
}
/** Write the full error and stack to a log file; return its path, or null if it can't be written. */
function writeCrashLog(error: unknown): string | null {
try {
const logPath = path.join(os.tmpdir(), 'shannon-error.log');
const detail = error instanceof Error && error.stack ? error.stack : String(error);
fs.writeFileSync(logPath, `${new Date().toISOString()}\n${detail}\n`);
return logPath;
} catch {
return null;
}
}
-145
View File
@@ -1,145 +0,0 @@
/**
* Per-command help text.
*
* `shannon <command> --help`, `shannon <command> -h`, and `shannon help <command>`
* all render the matching command's usage, so a user can discover a command's
* flags without scanning the global help. The global help lives in index.ts.
*/
import { commandPrefix, getMode } from './mode.js';
interface CommandHelp {
readonly usage: readonly string[];
readonly description: string;
readonly options?: readonly (readonly [string, string])[];
readonly examples?: readonly string[];
}
const YES_OPTION: readonly [string, string] = [
'-y, --yes',
'Skip the confirmation prompt (required for non-interactive use)',
];
const HELP_OPTION: readonly [string, string] = ['-h, --help', 'Show this help'];
/**
* `start`'s flags, the single source rendered by both the per-command help here
* and the global help in index.ts, so the two can never drift.
*/
export const START_OPTIONS: readonly (readonly [string, string])[] = [
['-u, --url <url>', 'Target URL (required)'],
['-r, --repo <path>', 'Repository path (required)'],
['-c, --config <path>', 'Configuration file (YAML)'],
['-o, --output <path>', 'Copy deliverables to this directory after the run'],
['-w, --workspace <name>', 'Named workspace (auto-resumes if it exists)'],
['-f, --follow', 'Stream the scan log until it finishes'],
['--pipeline-testing', 'Use minimal prompts for fast testing'],
['--keep-container', 'Preserve the worker container after exit for log inspection'],
];
const COMMAND_HELP: Readonly<Record<string, CommandHelp>> = {
start: {
usage: ['start -u <url> -r <path> [options]'],
description: 'Start a pentest scan.',
examples: [
'start -u https://example.com -r ./my-repo',
'start -u https://example.com -r /path/to/repo -c config.yaml -w q1-audit',
'start -u https://example.com -r ./my-repo --follow',
],
},
stop: {
usage: ['stop <workspace> [--yes]', 'stop --all [--yes]'],
description: 'Stop one scan by workspace, or every scan with --all (Temporal stays up).',
options: [['--all', 'Stop all running scans'], YES_OPTION],
examples: ['stop q1-audit', 'stop --all'],
},
reset: {
usage: ['reset'],
description: 'Stop everything and permanently remove all Temporal data and volumes.',
},
logs: {
usage: ['logs <workspace>'],
description: "Tail a scan's live log until it completes.",
examples: ['logs q1-audit'],
},
status: {
usage: ['status <workspace> [--json]'],
description:
"Show one scan's phase-by-phase progress, read live from Temporal. Watches and redraws until the scan finishes on a terminal; prints one frame when piped or already finished. With --json, prints a single machine-readable snapshot and exits.",
options: [['--json', 'Output a point-in-time snapshot as JSON, then exit']],
examples: ['status q1-audit', 'status q1-audit --json'],
},
scans: {
usage: ['scans [--json]'],
description: 'List completed scans and where each report lives.',
options: [['--json', 'Output the scan list as JSON']],
examples: ['scans', 'scans --json'],
},
build: {
usage: ['build [--no-cache]'],
description: 'Build the worker Docker image (local mode only).',
options: [['--no-cache', 'Build without using the Docker layer cache']],
},
setup: {
usage: ['setup'],
description: 'Configure provider credentials interactively (npx mode only).',
},
version: {
usage: ['version [--json]'],
description: 'Show the version. With --json, prints the version and mode as a machine-readable object.',
options: [['--json', 'Output the version and mode as JSON']],
examples: ['version', 'version --json'],
},
};
/** Commands that only exist in one mode; everything else is available in both. */
const MODE_ONLY: Readonly<Record<string, 'local' | 'npx'>> = {
build: 'local',
setup: 'npx',
};
/** Whether a command has its own help page (and so responds to `--help`/`-h`). */
export function isHelpableCommand(command: string): boolean {
return command in COMMAND_HELP;
}
/**
* User-facing command names available in the current mode, for "did you mean?"
* suggestions. Derived from the same table that backs per-command help, so the
* suggestion set can never drift from the commands that actually exist.
*/
export function availableCommands(): readonly string[] {
const mode = getMode();
const commands = Object.keys(COMMAND_HELP).filter((command) => (MODE_ONLY[command] ?? mode) === mode);
return [...commands, 'help'];
}
/** Print the help page for one command. No-op if the command has no page. */
export function printCommandHelp(command: string): void {
const help = COMMAND_HELP[command];
if (!help) return;
const prefix = commandPrefix();
const baseOptions = command === 'start' ? START_OPTIONS : (help.options ?? []);
const options = [...baseOptions, HELP_OPTION];
const flagWidth = Math.max(...options.map(([flag]) => flag.length));
const lines: string[] = ['', help.description, '', 'USAGE'];
for (const line of help.usage) {
lines.push(` ${prefix} ${line}`);
}
lines.push('', 'OPTIONS');
for (const [flag, desc] of options) {
lines.push(` ${flag.padEnd(flagWidth)} ${desc}`);
}
if (help.examples && help.examples.length > 0) {
lines.push('', 'EXAMPLES');
for (const example of help.examples) {
lines.push(` ${prefix} ${example}`);
}
}
lines.push('');
console.log(lines.join('\n'));
}
+153 -179
View File
@@ -9,19 +9,15 @@
* in the current working directory.
*/
import { ArgError, parseArgs, YES_FLAGS } from './args.js';
import { build } from './commands/build.js';
import { logs } from './commands/logs.js';
import { reset } from './commands/reset.js';
import { scans } from './commands/scans.js';
import { setup } from './commands/setup.js';
import { start } from './commands/start.js';
import { status } from './commands/status.js';
import { stop } from './commands/stop.js';
import { crash, fail, failUsage } from './errors.js';
import { availableCommands, isHelpableCommand, printCommandHelp, START_OPTIONS } from './help.js';
import { commandPrefix, getMode } from './mode.js';
import { closestMatch } from './suggest.js';
import { uninstall } from './commands/uninstall.js';
import { workspaces } from './commands/workspaces.js';
import { getMode } from './mode.js';
import { getVersion, getVersionLine } from './version.js';
function blockSudo(): void {
@@ -29,30 +25,23 @@ function blockSudo(): void {
const isRoot = process.geteuid?.() === 0;
if (!isSudo && !isRoot) return;
const linuxHints =
process.platform === 'linux'
? ['Configure Docker to run without sudo first:', 'https://docs.docker.com/engine/install/linux-postinstall']
: [];
if (isSudo) {
fail('Shannon must not be run with sudo.', 'Re-run this command as your normal user.', ...linuxHints);
console.error('ERROR: Shannon must not be run with sudo.');
console.error('Re-run this command as your normal user.');
} else {
console.error('ERROR: Shannon must not be run as the root user.');
console.error('Switch to a regular user account and re-run this command.');
}
fail(
'Shannon must not be run as the root user.',
'Switch to a regular user account and re-run this command.',
...linuxHints,
);
}
/** Render `start`'s flags for the global help, from the same source as `start --help`. */
function renderStartOptions(): string {
const flagWidth = Math.max(...START_OPTIONS.map(([flag]) => flag.length));
return START_OPTIONS.map(([flag, desc]) => ` ${flag.padEnd(flagWidth)} ${desc}`).join('\n');
if (process.platform === 'linux') {
console.error('Configure Docker to run without sudo first:');
console.error('https://docs.docker.com/engine/install/linux-postinstall');
}
process.exit(1);
}
function showHelp(): void {
const mode = getMode();
const prefix = commandPrefix();
const prefix = mode === 'local' ? './shannon' : 'npx @keygraph/shannon';
console.log(`
Shannon - AI Penetration Testing Framework
@@ -64,31 +53,33 @@ Usage:${
${prefix} setup Configure credentials`
}
${prefix} start --url <url> --repo <path> [options] Start a pentest scan
${prefix} stop <workspace> [--yes] Stop one scan
${prefix} stop --all [--yes] Stop all scans (Temporal stays up)
${prefix} reset Stop everything and wipe all Temporal data
${prefix} stop [--clean] [--yes] Stop all running scans
${prefix} workspaces List all workspaces
${prefix} logs <workspace> Show a scan's live log
${prefix} status <workspace> [--json] Live phase/agent progress of one scan
${prefix} scans [--json] List completed scans and their reports${
${prefix} status Show running scans${
mode === 'local'
? `
${prefix} build [--no-cache] Build worker image`
: ''
: `
${prefix} uninstall [--yes] Remove ~/.shannon/ and all data`
}
${prefix} version [--json] Show version
${prefix} version Show version
${prefix} help Show this help
Options for 'start':
${renderStartOptions()}
-u, --url <url> Target URL (required)
-r, --repo <path> Repository path${mode === 'local' ? ' or bare name' : ''} (required)
-c, --config <path> Configuration file (YAML)
-o, --output <path> Copy deliverables to this directory after run
-w, --workspace <name> Named workspace (auto-resumes if exists)
--pipeline-testing Use minimal prompts for fast testing
--debug Preserve worker container after exit for log inspection
Examples:
${prefix} start -u https://example.com -r ./my-repo
${prefix} start -u https://example.com -r ${mode === 'local' ? 'my-repo' : './my-repo'}
${prefix} start -u https://example.com -r /path/to/repo -c config.yaml -w q1-audit
${prefix} logs q1-audit
${prefix} stop q1-audit
${prefix} reset
Run '${prefix} <command> --help' for help on a specific command.
${prefix} stop --clean
${
mode === 'local'
? `
@@ -97,7 +88,6 @@ State directory: ./workspaces/`
State directory: ~/.shannon/`
}
Monitor scans at http://localhost:8233
Docs & source: https://github.com/KeygraphHQ/shannon
`);
}
@@ -108,166 +98,150 @@ interface ParsedStartArgs {
workspace?: string;
output?: string;
pipelineTesting: boolean;
keepContainer: boolean;
follow: boolean;
debug: boolean;
}
function parseStartArgs(argv: string[]): ParsedStartArgs {
const { flags, values } = parseArgs(argv, {
values: {
url: ['-u', '--url'],
repo: ['-r', '--repo'],
config: ['-c', '--config'],
output: ['-o', '--output'],
workspace: ['-w', '--workspace'],
},
booleans: {
pipelineTesting: ['--pipeline-testing'],
keepContainer: ['--keep-container'],
follow: ['-f', '--follow'],
},
});
let url = '';
let repo = '';
let config: string | undefined;
let workspace: string | undefined;
let output: string | undefined;
let pipelineTesting = false;
let debug = false;
const url = values.url ?? '';
const repo = values.repo ?? '';
if (!url || !repo) {
failUsage('--url and --repo are required', `Usage: ${commandPrefix()} start -u <url> -r <path>`);
for (let i = 0; i < argv.length; i++) {
const arg = argv[i];
const next = argv[i + 1];
switch (arg) {
case '-u':
case '--url':
if (next && !next.startsWith('-')) {
url = next;
i++;
}
break;
case '-r':
case '--repo':
if (next && !next.startsWith('-')) {
repo = next;
i++;
}
break;
case '-c':
case '--config':
if (next && !next.startsWith('-')) {
config = next;
i++;
}
break;
case '-w':
case '--workspace':
if (next && !next.startsWith('-')) {
workspace = next;
i++;
}
break;
case '-o':
case '--output':
if (next && !next.startsWith('-')) {
output = next;
i++;
}
break;
case '--pipeline-testing':
pipelineTesting = true;
break;
case '--debug':
debug = true;
break;
default:
console.error(`Unknown option: ${arg}`);
console.error(`Run "${getMode() === 'local' ? './shannon' : 'npx @keygraph/shannon'} help" for usage`);
process.exit(1);
}
}
try {
new URL(url);
} catch {
failUsage(`invalid --url: ${url}`);
if (!url || !repo) {
console.error('ERROR: --url and --repo are required');
console.error(`Usage: ${getMode() === 'local' ? './shannon' : 'npx @keygraph/shannon'} start -u <url> -r <path>`);
process.exit(1);
}
return {
url,
repo,
pipelineTesting: !!flags.pipelineTesting,
keepContainer: !!flags.keepContainer,
follow: !!flags.follow,
...(values.config && { config: values.config }),
...(values.workspace && { workspace: values.workspace }),
...(values.output && { output: values.output }),
pipelineTesting,
debug,
...(config && { config }),
...(workspace && { workspace }),
...(output && { output }),
};
}
// === Main Dispatch ===
async function main(): Promise<void> {
// A reader that closes early (e.g. `shannon logs my-scan | head`) makes writes
// to stdout raise EPIPE. That's normal for a piped CLI, not a crash — exit quietly
// instead of letting Node dump an unhandled-error stack trace.
process.stdout.on('error', (err: NodeJS.ErrnoException) => {
if (err.code === 'EPIPE') process.exit(0);
throw err;
});
blockSudo();
blockSudo();
const args = process.argv.slice(2);
const command = args[0];
const args = process.argv.slice(2);
const command = args[0];
const rest = args.slice(1);
if (command === undefined || command === 'help' || command === '--help' || command === '-h') {
const topic = rest[0];
if (topic && isHelpableCommand(topic)) {
printCommandHelp(topic);
} else {
showHelp();
}
return;
switch (command) {
case 'start': {
const parsed = parseStartArgs(args.slice(1));
await start({ ...parsed, version: getVersion() });
break;
}
// Reachable from any invocation: `-h`/`--help` anywhere wins over the rest of the line.
if (isHelpableCommand(command) && (rest.includes('-h') || rest.includes('--help'))) {
printCommandHelp(command);
return;
case 'stop':
stop(args.includes('--clean'), args.includes('--yes') || args.includes('-y'));
break;
case 'logs': {
const workspaceId = args[1];
if (!workspaceId) {
console.error('ERROR: Workspace ID is required');
console.error(`Usage: ${getMode() === 'local' ? './shannon' : 'npx @keygraph/shannon'} logs <workspace>`);
process.exit(1);
}
logs(workspaceId);
break;
}
switch (command) {
case 'start': {
const parsed = parseStartArgs(rest);
await start({ ...parsed, version: getVersion() });
break;
case 'workspaces':
workspaces(getVersion());
break;
case 'status':
status();
break;
case 'setup':
if (getMode() === 'local') {
console.error('ERROR: setup is only available in npx mode. In local mode, use .env');
process.exit(1);
}
case 'stop': {
const { flags, positionals } = parseArgs(rest, {
booleans: { all: ['--all'], yes: YES_FLAGS },
maxPositionals: 1,
});
await stop({ all: !!flags.all, yes: !!flags.yes, ...(positionals[0] && { workspace: positionals[0] }) });
break;
setup();
break;
case 'build':
build(args.includes('--no-cache'));
break;
case 'uninstall':
if (getMode() === 'local') {
console.error('ERROR: uninstall is only available in npx mode.');
process.exit(1);
}
case 'reset': {
// reset is all-or-nothing; a stray name likely means the user wanted `stop <name>`.
parseArgs(rest, {
positionalHint: 'reset takes no workspace argument. To stop one scan, use: stop <name>',
});
await reset();
break;
}
case 'logs': {
const { positionals } = parseArgs(rest, { maxPositionals: 1 });
const workspaceId = positionals[0];
if (!workspaceId) {
failUsage('Workspace ID is required', `Usage: ${commandPrefix()} logs <workspace>`);
}
logs(workspaceId);
break;
}
case 'status': {
const { flags, positionals } = parseArgs(rest, { booleans: { json: ['--json'] }, maxPositionals: 1 });
const workspaceId = positionals[0];
if (!workspaceId) {
failUsage('Workspace is required', `Usage: ${commandPrefix()} status <workspace> [--json]`);
}
await status(workspaceId, { json: !!flags.json });
break;
}
case 'scans': {
const { flags } = parseArgs(rest, { booleans: { json: ['--json'] } });
scans({ json: !!flags.json });
break;
}
case 'setup':
if (getMode() === 'local') {
fail('setup is only available in npx mode. In local mode, use .env');
}
parseArgs(rest, {});
await setup();
break;
case 'build': {
const { flags } = parseArgs(rest, { booleans: { noCache: ['--no-cache'] } });
build(!!flags.noCache, getVersion());
break;
}
case 'version':
case '--version':
case '-v': {
const { flags } = parseArgs(rest, { booleans: { json: ['--json'] } });
if (flags.json) {
console.log(JSON.stringify({ version: getVersion(), mode: getMode() }, null, 2));
} else {
console.log(getVersionLine());
}
break;
}
default: {
const prefix = commandPrefix();
const suggestion = closestMatch(command, availableCommands());
const hints = [
...(suggestion ? [`Did you mean '${suggestion}'?`] : []),
`Run '${prefix} help' to see available commands.`,
];
failUsage(`Unknown command: ${command}`, ...hints);
}
}
uninstall(args.includes('--yes') || args.includes('-y'));
break;
case 'version':
case '--version':
case '-v':
console.log(getVersionLine());
break;
case 'help':
case '--help':
case '-h':
case undefined:
showHelp();
break;
default:
console.error(`Unknown command: ${command}`);
showHelp();
process.exit(1);
}
main().catch((err) => {
if (err instanceof ArgError) {
failUsage(err.message, `Run "${commandPrefix()} help" for usage`);
}
crash(err);
});
-9
View File
@@ -23,12 +23,3 @@ export function setMode(mode: Mode): void {
export function isLocal(): boolean {
return getMode() === 'local';
}
/** The invocation prefix for the current mode, so help and hints point at a runnable command. */
export function commandPrefix(): string {
return getMode() === 'local' ? './shannon' : 'npx @keygraph/shannon';
}
export function isDevMode(): boolean {
return process.env.SHANNON_DEV === '1';
}
-91
View File
@@ -1,91 +0,0 @@
/**
* Parsing for the single model setting, `SHANNON_AI_MODEL=<provider>:<model-id>`.
*
* Mirrors apps/worker/src/ai/models.ts. The CLI cannot import from the worker
* package (it ships as a standalone bundle), so the provider list and the parse
* rule are duplicated here deliberately and must stay in sync.
*/
/**
* Providers Shannon curates with their own credential variables, config sections,
* and setup flows. Any other pi provider is reachable via the generic credential
* path. Mirrors CURATED_PROVIDERS in apps/worker/src/ai/models.ts.
*/
export const CURATED_PROVIDERS = ['anthropic', 'openai', 'xai', 'amazon-bedrock'] as const;
export type CuratedProviderId = (typeof CURATED_PROVIDERS)[number];
export function isCuratedProvider(value: string): value is CuratedProviderId {
return (CURATED_PROVIDERS as readonly string[]).includes(value);
}
/** Generic API key, honored for any provider Shannon does not curate. Mirrors the worker. */
export const GENERIC_API_KEY_ENV = 'SHANNON_AI_API_KEY';
/**
* Env vars carrying each curated provider's API key, in precedence order. Any one of
* them satisfies the provider. Mirrors PROVIDER_API_KEY_ENV in apps/worker/src/ai/models.ts.
*/
export const PROVIDER_API_KEY_ENV: Readonly<Record<CuratedProviderId, readonly string[]>> = {
anthropic: ['ANTHROPIC_API_KEY', 'CLAUDE_CODE_OAUTH_TOKEN'],
openai: ['OPENAI_API_KEY'],
xai: ['XAI_API_KEY'],
'amazon-bedrock': ['AWS_BEARER_TOKEN_BEDROCK'],
};
/** Additional env vars a curated provider requires beyond its API key. All must be set. */
export const PROVIDER_EXTRA_ENV: Readonly<Record<CuratedProviderId, readonly string[]>> = {
anthropic: [],
openai: [],
xai: [],
'amazon-bedrock': ['AWS_REGION'],
};
/** Human-readable credential requirement, used in "nothing configured" errors. */
export const PROVIDER_CREDENTIAL_HINT: Readonly<Record<CuratedProviderId, string>> = {
anthropic: 'ANTHROPIC_API_KEY (or CLAUDE_CODE_OAUTH_TOKEN)',
openai: 'OPENAI_API_KEY',
xai: 'XAI_API_KEY',
'amazon-bedrock': 'AWS_REGION and AWS_BEARER_TOKEN_BEDROCK',
};
/** Model used when SHANNON_AI_MODEL is unset. */
export const DEFAULT_MODEL_SPEC = 'anthropic:claude-sonnet-4-6';
/**
* Values SHANNON_AI_OPENAI_FORMAT accepts, selecting the wire format an
* OpenAI-compatible gateway serves. Mirrors OPENAI_FORMATS in
* apps/worker/src/ai/models.ts; the worker validates and applies it.
*/
export const OPENAI_FORMATS = ['chat-completions', 'responses'] as const;
export type OpenAiFormat = (typeof OPENAI_FORMATS)[number];
export interface ModelSpec {
providerId: string;
modelId: string;
}
/**
* Parse a `<provider>:<model-id>` spec. Splits on the first colon only, so colons
* inside a model ID survive (`amazon-bedrock:us.anthropic.claude-opus-4-5-20251101-v1:0`).
* The provider id is passed through as given — the worker's preflight validates it
* against pi. Returns an error string rather than throwing, for the CLI's flow.
*/
export function parseModelSpec(spec: string): ModelSpec | string {
const trimmed = spec.trim();
const separator = trimmed.indexOf(':');
const malformed = `SHANNON_AI_MODEL must be "<provider>:<model-id>", got "${trimmed}". Example: ${DEFAULT_MODEL_SPEC}`;
if (separator === -1) return malformed;
const providerId = trimmed.slice(0, separator).trim();
const modelId = trimmed.slice(separator + 1).trim();
if (!providerId || !modelId) return malformed;
return { providerId, modelId };
}
/** Resolve the run's model spec from the environment, or an error string. */
export function resolveModelSpec(): ModelSpec | string {
return parseModelSpec(process.env.SHANNON_AI_MODEL || DEFAULT_MODEL_SPEC);
}
+34 -28
View File
@@ -1,27 +1,13 @@
/**
* Path resolution for --repo and --config arguments.
*
* Both --repo and --config are filesystem paths, absolute or relative to CWD.
* Local mode supports bare repo names (e.g. "my-repo" → ./repos/my-repo).
* Both modes resolve relative paths against CWD.
*/
import fs from 'node:fs';
import os from 'node:os';
import path from 'node:path';
import { fail } from './errors.js';
/**
* Expand a leading `~` or `~/` to the home directory. The shell skips this in the
* `--flag=~/x` form (the tilde is not at the word start), so it must be done here.
*/
export function expandHome(inputPath: string): string {
if (inputPath === '~') {
return os.homedir();
}
if (inputPath.startsWith('~/')) {
return path.join(os.homedir(), inputPath.slice(2));
}
return inputPath;
}
import { isLocal } from './mode.js';
export interface MountPair {
hostPath: string;
@@ -37,10 +23,10 @@ export interface MountPair {
export const INTERNAL_DIR = '.shannon';
/**
* Filename of the human-facing PDF report surfaced at the run directory root.
* Must match FINAL_REPORT_PDF_FILENAME in the worker package.
* Filename of the human-facing final report surfaced at the run directory root.
* Must match FINAL_REPORT_FILENAME in the worker package.
*/
export const FINAL_REPORT_PDF_FILENAME = 'Security-Assessment-Report.pdf';
export const FINAL_REPORT_FILENAME = 'Security-Assessment-Report.md';
/**
* Resolve a run-directory file (e.g. session.json, workflow.log), preferring the
@@ -61,18 +47,36 @@ export function resolveRunFile(runDir: string, filename: string): string {
}
/**
* Resolve --repo to an absolute path and container mount. The argument is a
* filesystem path, absolute or relative to CWD.
* Resolve --repo to absolute path and container mount.
* Dev mode: bare names (no / or . prefix) check ./repos/<name> first.
*/
export function resolveRepo(repoArg: string): MountPair {
const hostPath = path.resolve(expandHome(repoArg));
let hostPath: string;
if (isLocal() && !repoArg.startsWith('/') && !repoArg.startsWith('.')) {
// Bare name — check ./repos/<name> for backward compatibility
const barePath = path.resolve('repos', repoArg);
if (fs.existsSync(barePath)) {
hostPath = barePath;
} else {
console.error(`ERROR: Repository not found at ./repos/${repoArg}`);
console.error('');
console.error('Place your target repository under the ./repos/ directory,');
console.error('or pass an absolute/relative path: -r /path/to/repo');
process.exit(1);
}
} else {
hostPath = path.resolve(repoArg);
}
if (!fs.existsSync(hostPath)) {
fail(`Repository not found: ${hostPath}`);
console.error(`ERROR: Repository not found: ${hostPath}`);
process.exit(1);
}
if (!fs.statSync(hostPath).isDirectory()) {
fail(`Not a directory: ${hostPath}`);
console.error(`ERROR: Not a directory: ${hostPath}`);
process.exit(1);
}
const basename = path.basename(hostPath);
@@ -86,14 +90,16 @@ export function resolveRepo(repoArg: string): MountPair {
* Resolve --config to absolute path and container mount.
*/
export function resolveConfig(configArg: string): MountPair {
const hostPath = path.resolve(expandHome(configArg));
const hostPath = path.resolve(configArg);
if (!fs.existsSync(hostPath)) {
fail(`Config file not found: ${hostPath}`);
console.error(`ERROR: Config file not found: ${hostPath}`);
process.exit(1);
}
if (!fs.statSync(hostPath).isFile()) {
fail(`Not a file: ${hostPath}`);
console.error(`ERROR: Not a file: ${hostPath}`);
process.exit(1);
}
const basename = path.basename(hostPath);
-155
View File
@@ -1,155 +0,0 @@
/**
* Pure derivation of a scan's per-agent and per-phase state from its Temporal snapshot.
*
* This is the single source of truth for "what state is each agent in" — both the
* human progress tree (render.ts) and the machine-readable snapshot (status-json.ts)
* consume it, so the two views can never disagree about whether an agent is running,
* skipped, or still pending. No glyphs, no color, no formatting live here.
*/
import type { RunningAgent } from '../temporal-client.js';
import { agentClass, PIPELINE, type PipelineState } from './pipeline.js';
import type { RenderInput } from './render.js';
export type RunState = 'pending' | 'running' | 'completed' | 'failed' | 'skipped';
/** One agent's resolved state plus the raw metrics/timing a consumer needs to present it. Null metrics
* mean the value doesn't apply to the current state (e.g. duration only for completed agents). */
export interface DerivedAgent {
readonly name: string;
readonly label: string;
readonly state: RunState;
readonly durationMs: number | null;
readonly runningElapsedMs: number | null;
readonly attempt: number | null;
readonly error?: string;
}
export interface DerivedPhase {
readonly key: string;
readonly label: string;
readonly parallel: boolean;
readonly state: RunState;
readonly agents: readonly DerivedAgent[];
}
/** Terminal = anything other than an open, running execution. */
export function isTerminal(status: string): boolean {
return status !== 'RUNNING' && status !== 'UNSPECIFIED';
}
function isFailedAgent(name: string, state: PipelineState | null): boolean {
return !!state && (state.failedAgent === name || state.failedPipelines.some((f) => f.vulnType === agentClass(name)));
}
/** An agent has entered play once it is running, has metrics, or has failed. */
function isAgentActive(name: string, state: PipelineState | null, running: Set<string>): boolean {
return running.has(name) || !!state?.agentMetrics[name] || isFailedAgent(name, state);
}
/**
* Resolve one agent's state. "Ran" is signalled by a metrics entry, not by
* completedAgents — the workflow lists conditionally-skipped agents (e.g. exploit
* agents when there is nothing to exploit) as completed but records no metrics for
* them. `resolved` is true once we've moved past this agent's phase (the scan is
* terminal, or a later phase is already active), at which point a metric-less,
* non-running agent is skipped rather than still pending.
*/
function agentState(name: string, state: PipelineState | null, running: Set<string>, resolved: boolean): RunState {
if (running.has(name)) return 'running';
if (isFailedAgent(name, state)) return 'failed';
if (state?.agentMetrics[name]) return 'completed';
return resolved ? 'skipped' : 'pending';
}
function agentError(name: string, state: PipelineState | null, byAgent: Map<string, RunningAgent>): string | undefined {
const failed = state?.failedPipelines.find((f) => f.vulnType === agentClass(name));
return (
failed?.error ??
byAgent.get(name)?.lastFailure ??
(state?.failedAgent === name ? (state.error ?? undefined) : undefined)
);
}
/** Scan wall-clock elapsed ms: recorded duration for a closed scan, live elapsed for a running one. */
export function scanElapsedMs(input: RenderInput, now: number): number | undefined {
if (isTerminal(input.temporalStatus)) {
if (input.state?.summary) return input.state.summary.totalDurationMs;
if (input.endedAt !== undefined && input.startedAt !== undefined) return input.endedAt - input.startedAt;
return undefined;
}
return input.startedAt !== undefined ? now - input.startedAt : undefined;
}
/** Collapse a phase's agent states into a single state for the phase line. */
export function phaseGlyphState(states: readonly RunState[]): RunState {
if (states.some((s) => s === 'running')) return 'running';
if (states.some((s) => s === 'failed')) return 'failed';
if (states.every((s) => s === 'skipped')) return 'skipped';
if (states.every((s) => s === 'completed' || s === 'skipped')) return 'completed';
if (states.some((s) => s === 'completed')) return 'running';
return 'pending';
}
/**
* Compute each agent's RunState. This is the drift-prone part shared by every view.
*
* The pipeline is sequential across phases: the last phase with any active agent is the
* frontier. Earlier phases with nothing active were skipped (e.g. exploitation when no
* class had anything to exploit), not still pending.
*/
export function deriveAgentStates(input: RenderInput): Map<string, RunState> {
const runningSet = new Set(input.running.map((r) => r.agent));
const terminal = isTerminal(input.temporalStatus);
let frontier = -1;
PIPELINE.forEach((phase, idx) => {
if (phase.agents.some((a) => isAgentActive(a.name, input.state, runningSet))) frontier = idx;
});
const states = new Map<string, RunState>();
for (const [phaseIdx, phase] of PIPELINE.entries()) {
const resolved = terminal || phaseIdx < frontier;
for (const agent of phase.agents) {
states.set(agent.name, agentState(agent.name, input.state, runningSet, resolved));
}
}
return states;
}
/**
* Full structured view of the pipeline: every agent's state plus the raw
* metrics/timing needed to present it, and each phase's collapsed state.
*/
export function derivePipeline(input: RenderInput, now: number): DerivedPhase[] {
const states = deriveAgentStates(input);
const byAgent = new Map(input.running.map((r) => [r.agent, r]));
return PIPELINE.map((phase) => {
const agents = phase.agents.map((a): DerivedAgent => {
const state = states.get(a.name) ?? 'pending';
const metrics = input.state?.agentMetrics[a.name];
const runner = byAgent.get(a.name);
const error = agentError(a.name, input.state, byAgent);
return {
name: a.name,
label: a.label,
state,
durationMs: state === 'completed' && metrics ? metrics.durationMs : null,
runningElapsedMs: state === 'running' && runner?.startedAt !== undefined ? now - runner.startedAt : null,
attempt: state === 'running' && runner ? runner.attempt : null,
...(error !== undefined && { error }),
};
});
return {
key: phase.key,
label: phase.label,
parallel: phase.parallel,
state: phaseGlyphState(agents.map((ag) => ag.state)),
agents,
};
});
}
export { agentError };
-123
View File
@@ -1,123 +0,0 @@
/**
* Static description of the Shannon scan pipeline, plus the worker types the CLI
* reads back from Temporal.
*
* The CLI cannot import from the worker package, so this mirrors it. Keep in sync with:
* - apps/worker/src/types/agents.ts (agent names / ordering)
* - apps/worker/src/session-manager.ts (phase membership)
* - apps/worker/src/temporal/activities.ts (the run*Agent activity names → `activityType`)
* - apps/worker/src/temporal/shared.ts (PipelineState / PipelineSummary)
* - apps/worker/src/types/metrics.ts (AgentMetrics)
*/
export interface AgentSpec {
/** Canonical agent name as it appears in PipelineState.completedAgents / agentMetrics. */
readonly name: string;
/** Short label for the progress tree. */
readonly label: string;
/** Temporal activity type name — how a running agent shows up in pendingActivities. */
readonly activityType: string;
}
export interface PhaseSpec {
readonly key: string;
readonly label: string;
readonly parallel: boolean;
readonly agents: readonly AgentSpec[];
}
/** The pipeline phases in execution order, each with its agents. */
export const PIPELINE: readonly PhaseSpec[] = [
{
// Preflight login check. Only authenticated scans record metrics here; a non-auth scan
// records none, so it renders as skipped — like Exploitation when nothing is exploitable.
key: 'auth-validation',
label: 'Authentication',
parallel: false,
agents: [{ name: 'validate-authentication', label: 'auth', activityType: 'runAuthenticationValidation' }],
},
{
key: 'pre-recon',
label: 'Pre-Recon',
parallel: false,
agents: [{ name: 'pre-recon', label: 'pre-recon', activityType: 'runPreReconAgent' }],
},
{
key: 'recon',
label: 'Recon',
parallel: false,
agents: [{ name: 'recon', label: 'recon', activityType: 'runReconAgent' }],
},
{
key: 'vulnerability-analysis',
label: 'Vulnerability Analysis',
parallel: true,
agents: [
{ name: 'injection-vuln', label: 'injection', activityType: 'runInjectionVulnAgent' },
{ name: 'xss-vuln', label: 'xss', activityType: 'runXssVulnAgent' },
{ name: 'auth-vuln', label: 'auth', activityType: 'runAuthVulnAgent' },
{ name: 'ssrf-vuln', label: 'ssrf', activityType: 'runSsrfVulnAgent' },
{ name: 'authz-vuln', label: 'authz', activityType: 'runAuthzVulnAgent' },
],
},
{
key: 'exploitation',
label: 'Exploitation',
parallel: true,
agents: [
{ name: 'injection-exploit', label: 'injection', activityType: 'runInjectionExploitAgent' },
{ name: 'xss-exploit', label: 'xss', activityType: 'runXssExploitAgent' },
{ name: 'auth-exploit', label: 'auth', activityType: 'runAuthExploitAgent' },
{ name: 'ssrf-exploit', label: 'ssrf', activityType: 'runSsrfExploitAgent' },
{ name: 'authz-exploit', label: 'authz', activityType: 'runAuthzExploitAgent' },
],
},
{
key: 'reporting',
label: 'Reporting',
parallel: false,
agents: [{ name: 'report', label: 'report', activityType: 'runReportAgent' }],
},
];
/** Temporal activity type name → canonical agent name, for mapping pendingActivities. */
export const ACTIVITY_TO_AGENT: Readonly<Record<string, string>> = Object.fromEntries(
PIPELINE.flatMap((phase) => phase.agents.map((agent) => [agent.activityType, agent.name])),
);
/** The vuln/exploit class of an agent (e.g. "authz-vuln" → "authz"), for failedPipelines matching. */
export function agentClass(name: string): string {
return name.replace(/-(vuln|exploit)$/, '');
}
// === Worker types read back from Temporal (mirror of shared.ts / metrics.ts) ===
export interface AgentMetrics {
readonly durationMs: number;
readonly costUsd: number | null;
readonly numTurns: number | null;
readonly model?: string;
readonly skipped?: boolean;
}
export interface PipelineSummary {
readonly totalCostUsd: number;
readonly totalDurationMs: number; // Wall-clock (end - start)
readonly totalTurns: number;
readonly agentCount: number;
}
export type PipelineStatus = 'running' | 'completed' | 'failed' | 'cancelled' | 'partial';
export interface PipelineState {
readonly status: PipelineStatus;
readonly currentPhase: string | null;
readonly currentAgent: string | null;
readonly completedAgents: string[];
readonly failedPipelines: { vulnType: string; error: string }[];
readonly failedAgent: string | null;
readonly error: string | null;
readonly startTime: number;
readonly agentMetrics: Record<string, AgentMetrics>;
readonly summary: PipelineSummary | null;
}
-248
View File
@@ -1,248 +0,0 @@
/**
* Renders a scan's Temporal state into the terminal progress tree.
*
* The same PipelineState drives both the live view (from the getProgress query) and
* the final view (from the workflow result); the running-agents overlay (from
* pendingActivities) supplies the in-flight set and retry counts the state lacks.
* Colors and Unicode glyphs are gated by the caller so the frame degrades off a TTY.
*/
import { BOLD, DIM, GOLD, paint, RED, YELLOW } from '../colors.js';
import { commandPrefix } from '../mode.js';
import type { RunningAgent } from '../temporal-client.js';
import { agentError, deriveAgentStates, isTerminal, phaseGlyphState, type RunState, scanElapsedMs } from './derive.js';
import { PIPELINE, type PipelineState } from './pipeline.js';
export interface RenderInput {
readonly workspace: string;
/** Temporal workflow id backing this scan (differs from workspace on a resume); used for the dashboard link. */
readonly workflowId?: string;
/** Temporal WorkflowExecutionStatusName: RUNNING | COMPLETED | FAILED | CANCELLED | TERMINATED | … */
readonly temporalStatus: string;
/** Progress (live) or result (terminal). Null when unavailable, e.g. a hard failure with no result. */
readonly state: PipelineState | null;
readonly running: readonly RunningAgent[];
readonly startedAt?: number;
readonly endedAt?: number;
/** Failure text when a failed scan has no readable state. */
readonly failureMessage?: string;
}
export interface RenderOptions {
readonly now: number;
readonly color: boolean;
readonly unicode: boolean;
/** True for the live view (adds a watch footer); false for the final/one-shot frame. */
readonly live: boolean;
/** Animation tick — advances the running-agent spinner. Ignored for static frames. */
readonly frame: number;
}
const COLORS = {
red: RED,
gold: GOLD,
yellow: YELLOW,
dim: DIM,
bold: BOLD,
} as const;
// === Formatting ===
function formatDuration(ms: number): string {
const seconds = Math.max(0, Math.floor(ms / 1000));
const hours = Math.floor(seconds / 3600);
const minutes = Math.floor((seconds % 3600) / 60);
const secs = seconds % 60;
if (hours > 0) return `${hours}h ${minutes}m`;
if (minutes > 0) return `${minutes}m ${secs}s`;
return `${secs}s`;
}
function truncate(text: string, max: number): string {
const flat = text.replace(/\s+/g, ' ').trim();
return flat.length <= max ? flat : `${flat.slice(0, max - 1)}…`;
}
/** Temporal Web UI, published by compose on 8233; deep-links to the workflow when its id is known. */
function temporalDashboardUrl(workflowId: string | undefined): string {
const base = 'http://localhost:8233';
return workflowId ? `${base}/namespaces/default/workflows/${workflowId}` : base;
}
// === Glyphs & status ===
const GLYPH_UNICODE: Record<RunState, string> = {
pending: '○',
running: '⟳',
completed: '●',
failed: '✗',
skipped: '·',
};
const GLYPH_ASCII: Record<RunState, string> = {
pending: '.',
running: '>',
completed: '+',
failed: 'x',
skipped: '-',
};
const STATE_COLOR: Record<RunState, string> = {
pending: COLORS.dim,
running: COLORS.gold,
completed: COLORS.gold,
failed: COLORS.red,
skipped: COLORS.dim,
};
/** Braille spinner frames for running agents — the clack loader style. */
const SPINNER_FRAMES = ['⠋', '⠙', '⠹', '⠸', '⠼', '⠴', '⠦', '⠧', '⠇', '⠏'] as const;
function glyph(state: RunState, opts: RenderOptions): string {
if (state === 'running' && opts.unicode) {
const spin = SPINNER_FRAMES[opts.frame % SPINNER_FRAMES.length] ?? SPINNER_FRAMES[0];
return paint(spin, STATE_COLOR.running, opts.color);
}
const symbol = opts.unicode ? GLYPH_UNICODE[state] : GLYPH_ASCII[state];
return paint(symbol, STATE_COLOR[state], opts.color);
}
/** Badge text + color for the scan as a whole, preferring the workflow's own status when known. */
function statusBadge(input: RenderInput, opts: RenderOptions): string {
const workflowStatus = input.state?.status;
if (!isTerminal(input.temporalStatus)) return paint('running', COLORS.gold, opts.color);
if (workflowStatus === 'partial') return paint('partial', COLORS.yellow, opts.color);
if (input.temporalStatus === 'COMPLETED') return paint('completed', COLORS.gold, opts.color);
if (input.temporalStatus === 'TERMINATED') return paint('stopped', COLORS.yellow, opts.color);
if (input.temporalStatus === 'CANCELLED' || input.temporalStatus === 'CANCELED') {
return paint('cancelled', COLORS.yellow, opts.color);
}
if (input.temporalStatus === 'TIMED_OUT') return paint('timed out', COLORS.red, opts.color);
return paint('FAILED', COLORS.red, opts.color);
}
// === Line builders ===
function agentMeta(
state: RunState,
metrics: { durationMs: number } | undefined,
runner: RunningAgent | undefined,
error: string | undefined,
opts: RenderOptions,
): string {
if (state === 'completed') {
const duration = metrics?.durationMs != null ? formatDuration(metrics.durationMs) : 'done';
return paint(duration, COLORS.dim, opts.color);
}
if (state === 'running') {
const parts = ['running'];
if (runner?.startedAt !== undefined) parts.push(formatDuration(opts.now - runner.startedAt));
if (runner && runner.attempt > 1) parts.push(`retry ${runner.attempt}`);
return paint(parts.join(' · '), COLORS.gold, opts.color);
}
if (state === 'failed') {
const detail = error ? ` · ${truncate(error, 46)}` : '';
return paint(`failed${detail}`, COLORS.red, opts.color);
}
if (state === 'skipped') return paint('skipped', COLORS.dim, opts.color);
return paint('queued', COLORS.dim, opts.color);
}
function phaseMeta(states: readonly RunState[], inPlay: number, parallel: boolean, opts: RenderOptions): string {
if (states.every((s) => s === 'pending')) return paint('pending', COLORS.dim, opts.color);
if (states.every((s) => s === 'skipped')) return paint('skipped', COLORS.dim, opts.color);
if (states.some((s) => s === 'failed') && !states.some((s) => s === 'running')) {
return paint('failed', COLORS.red, opts.color);
}
if (!parallel) return '';
const done = states.filter((s) => s === 'completed').length;
const allDone = states.every((s) => s === 'completed' || s === 'skipped');
return paint(`${done}/${inPlay} done`, allDone ? COLORS.gold : COLORS.dim, opts.color);
}
/** Render the full progress frame as one string (no trailing newline). */
export function renderScan(input: RenderInput, opts: RenderOptions): string {
const byAgent = new Map(input.running.map((r) => [r.agent, r]));
const stateMap = deriveAgentStates(input);
const lines: string[] = ['', ...headerLines(input, opts), ''];
const metaFor = (name: string, state: RunState): string =>
agentMeta(state, input.state?.agentMetrics[name], byAgent.get(name), agentError(name, input.state, byAgent), opts);
// Only agents that have actually entered play are shown; pending/skipped ones stay hidden.
const inPlay = (s: RunState): boolean => s === 'running' || s === 'completed' || s === 'failed';
for (const phase of PIPELINE) {
const states = phase.agents.map((a) => stateMap.get(a.name) ?? 'pending');
const playing = states.filter(inPlay).length;
const phaseRunState: RunState = phaseGlyphState(states);
// A single-agent phase carries that agent's own duration/cost on the phase line once it
// starts; a parallel phase gets a "k/N done" summary over the agents in play.
const first = phase.agents[0];
const firstState = states[0];
const phaseMetaStr =
!phase.parallel && first && firstState && inPlay(firstState)
? metaFor(first.name, firstState)
: phaseMeta(states, playing, phase.parallel, opts);
lines.push(` ${glyph(phaseRunState, opts)} ${phase.label.padEnd(26)}${phaseMetaStr}`);
if (!phase.parallel) continue;
for (let i = 0; i < phase.agents.length; i++) {
const agent = phase.agents[i];
const state = states[i];
if (!agent || !state || !inPlay(state)) continue;
lines.push(` ${glyph(state, opts)} ${agent.label.padEnd(18)}${metaFor(agent.name, state)}`);
}
}
lines.push(...footerLines(input, opts));
return lines.join('\n');
}
function headerLines(input: RenderInput, opts: RenderOptions): string[] {
const elapsedMs = scanElapsedMs(input, opts.now);
const meta = [statusBadge(input, opts), elapsedMs !== undefined ? formatDuration(elapsedMs) : '—'].join(' · ');
return [` ${paint('Scan:', COLORS.bold, opts.color)} ${input.workspace.padEnd(22)} ${meta}`];
}
/** Aligned label column for the footer's Logs / Temporal rows. */
const FOOTER_LABEL_WIDTH = 12;
/** A thin rule that sets the footer apart from the phase list above it. */
function footerDivider(opts: RenderOptions): string {
return paint(` ${(opts.unicode ? '─' : '-').repeat(60)}`, COLORS.dim, opts.color);
}
/** One footer row: an accent-colored label in a fixed column, then its value in the default color. */
function footerRow(label: string, value: string, opts: RenderOptions): string {
return ` ${paint(label.padEnd(FOOTER_LABEL_WIDTH), COLORS.gold, opts.color)}${value}`;
}
function footerLines(input: RenderInput, opts: RenderOptions): string[] {
const prefix = commandPrefix();
if (isTerminal(input.temporalStatus) && input.state?.summary) {
const wall = formatDuration(input.state.summary.totalDurationMs);
return ['', ` Time Taken ${wall}`];
}
const logsValue = `${prefix} logs ${input.workspace}`;
const temporalValue = temporalDashboardUrl(input.workflowId);
if (isTerminal(input.temporalStatus)) {
const reason = input.failureMessage ?? input.state?.error ?? 'no result recorded';
return [
footerDivider(opts),
paint(
` ${input.temporalStatus === 'TERMINATED' ? 'Stopped' : 'Ended'} — ${truncate(reason, 240)}`,
COLORS.dim,
opts.color,
),
footerRow('Logs', logsValue, opts),
footerRow('Temporal', temporalValue, opts),
];
}
const lines = [footerDivider(opts), footerRow('Logs', logsValue, opts), footerRow('Temporal', temporalValue, opts)];
if (opts.live) lines.push('', paint(' Ctrl-C stops watching — the scan keeps running.', COLORS.dim, opts.color));
return lines;
}
-68
View File
@@ -1,68 +0,0 @@
/**
* Machine-readable snapshot of one scan, for `shannon status --json`.
*
* A point-in-time view built from the same derivation the human progress tree uses
* (derive.ts), so the JSON and the rendered tree can never disagree about an agent's
* state. One invocation is one snapshot — callers that want to track progress poll it.
*/
import type { DerivedPhase } from './derive.js';
import { derivePipeline, isTerminal, scanElapsedMs } from './derive.js';
import type { RenderInput } from './render.js';
/** Coarse scan status token, mirroring the human status badge in machine-friendly form. */
export type ScanStatus = 'running' | 'completed' | 'partial' | 'failed' | 'stopped' | 'cancelled' | 'timed_out';
export interface StatusJson {
readonly workspace: string;
/** Temporal workflow id backing this scan (differs from workspace on a resume). */
readonly workflowId?: string;
/** Coarse outcome: `running` until the scan closes, then its terminal status. */
readonly status: ScanStatus;
/** Raw Temporal WorkflowExecutionStatusName, for callers that need the source status. */
readonly temporalStatus: string;
/** Wall-clock elapsed ms (live for a running scan, final for a closed one), or null when unknown. */
readonly elapsedMs: number | null;
readonly startedAt?: string;
readonly endedAt?: string;
/** Failure text when a failed scan left no readable state. */
readonly failureMessage?: string;
readonly phases: readonly DerivedPhase[];
}
/** Map the raw Temporal status (and workflow status) onto the coarse machine token. */
function deriveStatus(input: RenderInput): ScanStatus {
if (!isTerminal(input.temporalStatus)) return 'running';
if (input.state?.status === 'partial') return 'partial';
switch (input.temporalStatus) {
case 'COMPLETED':
return 'completed';
case 'TERMINATED':
return 'stopped';
case 'CANCELLED':
case 'CANCELED':
return 'cancelled';
case 'TIMED_OUT':
return 'timed_out';
default:
return 'failed';
}
}
/** Build the JSON snapshot for a scan at instant `now`. */
export function toStatusJson(input: RenderInput, now: number): StatusJson {
const elapsedMs = scanElapsedMs(input, now);
return {
workspace: input.workspace,
...(input.workflowId !== undefined && { workflowId: input.workflowId }),
status: deriveStatus(input),
temporalStatus: input.temporalStatus,
elapsedMs: elapsedMs ?? null,
...(input.startedAt !== undefined && { startedAt: new Date(input.startedAt).toISOString() }),
...(input.endedAt !== undefined && { endedAt: new Date(input.endedAt).toISOString() }),
...(input.failureMessage !== undefined && { failureMessage: input.failureMessage }),
phases: derivePipeline(input, now),
};
}
-26
View File
@@ -1,26 +0,0 @@
/**
* Workspace → Temporal workflow-id resolution.
*
* A workspace name is not always its workflow id: a fresh scan's id equals the
* workspace name, but each resume spawns a new workflow (`<workspace>_resume_<ts>`).
* The workspace's session.json records the authoritative id — the latest resume
* attempt, or the original — so commands that query Temporal (status, stop) resolve
* through here instead of assuming the name is the id.
*/
import fs from 'node:fs';
import path from 'node:path';
import { getWorkspacesDir } from './home.js';
import { resolveRunFile } from './paths.js';
/** Latest workflow id recorded for a workspace: last resume attempt, else the original. */
export function resolveWorkflowId(workspace: string): string | undefined {
const sessionPath = resolveRunFile(path.join(getWorkspacesDir(), workspace), 'session.json');
try {
const session = JSON.parse(fs.readFileSync(sessionPath, 'utf-8'));
const resumeAttempts: { workflowId?: string }[] = session.session?.resumeAttempts ?? [];
return resumeAttempts.at(-1)?.workflowId ?? session.session?.originalWorkflowId ?? undefined;
} catch {
return undefined;
}
}
+36 -59
View File
@@ -5,73 +5,50 @@
import { supportsColor } from './tty.js';
/** SHANNON wordmark. Block glyphs take the row fill; box-drawing strokes take the deeper edge shade. */
const SHANNON = [
'███████╗██╗ ██╗ █████╗ ███╗ ██╗███╗ ██╗ ██████╗ ███╗ ██╗',
'██╔════╝██║ ██║██╔══██╗████╗ ██║████╗ ██║██╔═══██╗████╗ ██║',
'███████╗███████║███████║██╔██╗ ██║██╔██╗ ██║██║ ██║██╔██╗ ██║',
'╚════██║██╔══██║██╔══██║██║╚██╗██║██║╚██╗██║██║ ██║██║╚██╗██║',
'███████║██║ ██║██║ ██║██║ ╚████║██║ ╚████║╚██████╔╝██║ ╚████║',
'╚══════╝╚═╝ ╚═╝╚═╝ ╚═╝╚═╝ ╚═══╝╚═╝ ╚═══╝ ╚═════╝ ╚═╝ ╚═══╝',
];
/**
* Sunset ramp, yellow at the top row down to burnt orange at the base.
* Wordmark row i is filled with stop i and edged with stop i + 1, so the
* box-drawing strokes read as a shadow one shade deeper than their row.
* `xterm` is the 256-color approximation for terminals without 24-bit color.
*/
const SUNSET: ReadonlyArray<{ rgb: readonly [number, number, number]; xterm: number }> = [
{ rgb: [247, 203, 45], xterm: 220 },
{ rgb: [246, 182, 38], xterm: 220 },
{ rgb: [245, 160, 32], xterm: 214 },
{ rgb: [242, 141, 28], xterm: 214 },
{ rgb: [238, 121, 24], xterm: 208 },
{ rgb: [231, 100, 21], xterm: 208 },
{ rgb: [222, 82, 19], xterm: 202 },
];
export function displaySplash(version?: string): void {
const color = supportsColor();
const truecolor = color && /truecolor|24bit/i.test(process.env.COLORTERM ?? '');
const RESET = color ? '\x1b[0m' : '';
const WHITE = color ? '\x1b[1;97m' : '';
const GOLD = color ? '\x1b[38;2;244;197;66m' : '';
const CYAN = color ? '\x1b[36;1m' : '';
const WHITE = color ? '\x1b[1;37m' : '';
const GRAY = color ? '\x1b[0;37m' : '';
const DIM = color ? '\x1b[90m' : '';
const YELLOW = color ? '\x1b[1;33m' : '';
const RESET = color ? '\x1b[0m' : '';
const ramp = SUNSET.map(({ rgb: [r, g, b], xterm }) => {
if (!color) return '';
return truecolor ? `\x1b[38;2;${r};${g};${b}m` : `\x1b[38;5;${xterm}m`;
});
/** Color one wordmark row, emitting an escape only where the run changes. Spaces stay unpainted. */
const paint = (row: string, fill: string, edge: string): string => {
if (!color) return row;
let out = '';
let open = '';
for (const ch of row) {
const want = ch === ' ' ? '' : ch === '█' ? fill : edge;
if (want !== open) {
if (open) out += RESET;
out += want;
open = want;
}
out += ch;
}
return open ? out + RESET : out;
};
const B = `${CYAN}\u2551${RESET}`;
const S67 = ' '.repeat(67);
const HR = '\u2550'.repeat(67);
const lines = [
'',
` ${WHITE}Keygraph${RESET}${version ? ` ${DIM}v${version}${RESET}` : ''}`,
'',
...SHANNON.map((row, i) => ` ${paint(row, ramp[i] ?? '', ramp[i + 1] ?? '')}`),
'',
` ${WHITE}AI Pentester for Web Apps and APIs${RESET}`,
'',
` ${GRAY}-Authorized Security Testing Only-${RESET}`,
'',
` ${CYAN}\u2554${HR}\u2557${RESET}`,
` ${B}${S67}${B}`,
` ${B} ${GOLD}\u2588\u2588\u2588\u2588\u2588\u2588\u2588\u2557\u2588\u2588\u2557 \u2588\u2588\u2557 \u2588\u2588\u2588\u2588\u2588\u2557 \u2588\u2588\u2588\u2557 \u2588\u2588\u2557\u2588\u2588\u2588\u2557 \u2588\u2588\u2557 \u2588\u2588\u2588\u2588\u2588\u2588\u2557 \u2588\u2588\u2588\u2557 \u2588\u2588\u2557${RESET} ${B}`,
` ${B} ${GOLD}\u2588\u2588\u2554\u2550\u2550\u2550\u2550\u255D\u2588\u2588\u2551 \u2588\u2588\u2551\u2588\u2588\u2554\u2550\u2550\u2588\u2588\u2557\u2588\u2588\u2588\u2588\u2557 \u2588\u2588\u2551\u2588\u2588\u2588\u2588\u2557 \u2588\u2588\u2551\u2588\u2588\u2554\u2550\u2550\u2550\u2588\u2588\u2557\u2588\u2588\u2588\u2588\u2557 \u2588\u2588\u2551${RESET} ${B}`,
` ${B} ${GOLD}\u2588\u2588\u2588\u2588\u2588\u2588\u2588\u2557\u2588\u2588\u2588\u2588\u2588\u2588\u2588\u2551\u2588\u2588\u2588\u2588\u2588\u2588\u2588\u2551\u2588\u2588\u2554\u2588\u2588\u2557 \u2588\u2588\u2551\u2588\u2588\u2554\u2588\u2588\u2557 \u2588\u2588\u2551\u2588\u2588\u2551 \u2588\u2588\u2551\u2588\u2588\u2554\u2588\u2588\u2557 \u2588\u2588\u2551${RESET} ${B}`,
` ${B} ${GOLD}\u255A\u2550\u2550\u2550\u2550\u2588\u2588\u2551\u2588\u2588\u2554\u2550\u2550\u2588\u2588\u2551\u2588\u2588\u2554\u2550\u2550\u2588\u2588\u2551\u2588\u2588\u2551\u255A\u2588\u2588\u2557\u2588\u2588\u2551\u2588\u2588\u2551\u255A\u2588\u2588\u2557\u2588\u2588\u2551\u2588\u2588\u2551 \u2588\u2588\u2551\u2588\u2588\u2551\u255A\u2588\u2588\u2557\u2588\u2588\u2551${RESET} ${B}`,
` ${B} ${GOLD}\u2588\u2588\u2588\u2588\u2588\u2588\u2588\u2551\u2588\u2588\u2551 \u2588\u2588\u2551\u2588\u2588\u2551 \u2588\u2588\u2551\u2588\u2588\u2551 \u255A\u2588\u2588\u2588\u2588\u2551\u2588\u2588\u2551 \u255A\u2588\u2588\u2588\u2588\u2551\u255A\u2588\u2588\u2588\u2588\u2588\u2588\u2554\u255D\u2588\u2588\u2551 \u255A\u2588\u2588\u2588\u2588\u2551${RESET} ${B}`,
` ${B} ${GOLD}\u255A\u2550\u2550\u2550\u2550\u2550\u2550\u255D\u255A\u2550\u255D \u255A\u2550\u255D\u255A\u2550\u255D \u255A\u2550\u255D\u255A\u2550\u255D \u255A\u2550\u2550\u2550\u255D\u255A\u2550\u255D \u255A\u2550\u2550\u2550\u255D \u255A\u2550\u2550\u2550\u2550\u2550\u255D \u255A\u2550\u255D \u255A\u2550\u2550\u2550\u255D${RESET} ${B}`,
` ${B}${S67}${B}`,
` ${B} ${CYAN}\u2554\u2550\u2550\u2550\u2550\u2550\u2550\u2550\u2550\u2550\u2550\u2550\u2550\u2550\u2550\u2550\u2550\u2550\u2550\u2550\u2550\u2550\u2550\u2550\u2550\u2550\u2550\u2550\u2550\u2550\u2550\u2550\u2550\u2550\u2550\u2550\u2550\u2557${RESET} ${B}`,
` ${B} ${CYAN}\u2551${RESET} ${WHITE}AI Penetration Testing Framework${RESET} ${CYAN}\u2551${RESET} ${B}`,
` ${B} ${CYAN}\u255A\u2550\u2550\u2550\u2550\u2550\u2550\u2550\u2550\u2550\u2550\u2550\u2550\u2550\u2550\u2550\u2550\u2550\u2550\u2550\u2550\u2550\u2550\u2550\u2550\u2550\u2550\u2550\u2550\u2550\u2550\u2550\u2550\u2550\u2550\u2550\u2550\u255D${RESET} ${B}`,
` ${B}${S67}${B}`,
];
if (version) {
const verStr = `v${version}`;
const verPadLeft = Math.floor((67 - verStr.length) / 2);
const verPadRight = 67 - verStr.length - verPadLeft;
lines.push(` ${B}${' '.repeat(verPadLeft)}${GRAY}${verStr}${RESET}${' '.repeat(verPadRight)}${B}`);
}
lines.push(
` ${B}${S67}${B}`,
` ${B} ${YELLOW}\uD83D\uDD10 DEFENSIVE SECURITY ONLY \uD83D\uDD10${RESET} ${B}`,
` ${B}${S67}${B}`,
` ${CYAN}\u255A${HR}\u255D${RESET}`,
'',
);
console.log(lines.join('\n'));
}
-58
View File
@@ -1,58 +0,0 @@
/**
* "Did you mean?" suggestions for mistyped commands and flags.
*
* A single Levenshtein-based matcher powers both the unknown-command path in the
* dispatcher and the unknown-option path in `parseArgs`, so a typo like `statsu`
* or `--workspce` points the user at the closest real name instead of just failing.
*/
/** Levenshtein edit distance between two strings (insertions, deletions, substitutions). */
export function editDistance(a: string, b: string): number {
if (a.length === 0) return b.length;
if (b.length === 0) return a.length;
// Rolling single row; `diagonal` and `above` carry the two neighbours a full grid would.
const row = Array.from({ length: b.length + 1 }, (_, j) => j);
for (let i = 1; i <= a.length; i++) {
let diagonal = row[0] as number;
row[0] = i;
for (let j = 1; j <= b.length; j++) {
const above = row[j] as number;
const cost = a[i - 1] === b[j - 1] ? 0 : 1;
row[j] = Math.min(above + 1, (row[j - 1] as number) + 1, diagonal + cost);
diagonal = above;
}
}
return row[b.length] as number;
}
/**
* The candidate closest to `input`, or undefined if none is near enough.
*
* A prefix match ("stat" -> "status") wins first; otherwise the lowest edit
* distance within a length-scaled threshold, so unrelated words don't match.
*/
export function closestMatch(input: string, candidates: readonly string[]): string | undefined {
if (input.length >= 2) {
const prefix = candidates.find((candidate) => candidate.startsWith(input));
if (prefix) return prefix;
}
let best: string | undefined;
let bestDistance = Number.POSITIVE_INFINITY;
for (const candidate of candidates) {
if (candidate.length <= 3) continue;
const distance = editDistance(input, candidate);
if (distance < bestDistance) {
bestDistance = distance;
best = candidate;
}
}
if (best === undefined) return undefined;
const threshold = Math.max(2, Math.floor(best.length / 3));
return bestDistance <= threshold ? best : undefined;
}
-126
View File
@@ -1,126 +0,0 @@
/**
* Thin Temporal client for reading one scan's state.
*
* A running scan is queried live (getProgress) and read via pendingActivities for
* the in-flight agents; a closed scan is read once from its result. Everything goes
* straight to the frontend on 127.0.0.1:7233 — the gRPC port the compose file
* publishes — so this needs Temporal up, but no worker of its own.
*/
import { Client, Connection, WorkflowFailedError, WorkflowNotFoundError } from '@temporalio/client';
import { ACTIVITY_TO_AGENT, type PipelineState } from './scan/pipeline.js';
const ADDRESS = '127.0.0.1:7233';
const NAMESPACE = 'default';
export interface RunningAgent {
readonly agent: string;
readonly attempt: number;
readonly startedAt?: number;
readonly lastFailure?: string;
}
/** Convert a proto ITimestamp (seconds is a Long) to epoch millis. */
function timestampMs(
ts: { seconds?: { toString(): string } | number | null; nanos?: number | null } | null,
): number | undefined {
const seconds = ts?.seconds;
if (seconds == null) return undefined;
const secNum = typeof seconds === 'number' ? seconds : Number(seconds.toString());
return secNum * 1000 + (ts?.nanos ?? 0) / 1e6;
}
export interface ScanDescription {
/** WorkflowExecutionStatusName: RUNNING | COMPLETED | FAILED | CANCELLED | TERMINATED | TIMED_OUT | … */
readonly status: string;
readonly startedAt?: number;
readonly closedAt?: number;
readonly runningAgents: readonly RunningAgent[];
}
export type TerminalOutcome =
| { readonly kind: 'success'; readonly state: PipelineState }
| { readonly kind: 'failed'; readonly message: string };
let clientPromise: Promise<Client> | null = null;
function getClient(): Promise<Client> {
if (!clientPromise) {
clientPromise = Connection.connect({ address: ADDRESS }).then(
(connection) => new Client({ connection, namespace: NAMESPACE }),
);
}
return clientPromise;
}
/** Describe a scan: status, timing, and the agents currently running (from pendingActivities). Null if not found. */
export async function describeScan(workflowId: string): Promise<ScanDescription | null> {
const client = await getClient();
try {
const desc = await client.workflow.getHandle(workflowId).describe();
const runningAgents: RunningAgent[] = [];
for (const pending of desc.raw.pendingActivities ?? []) {
const agent = ACTIVITY_TO_AGENT[pending.activityType?.name ?? ''];
if (!agent) continue;
const lastFailure = pending.lastFailure?.message;
const startedAt = timestampMs(pending.scheduledTime ?? pending.lastStartedTime ?? null);
runningAgents.push({
agent,
attempt: pending.attempt ?? 1,
...(startedAt !== undefined ? { startedAt } : {}),
...(lastFailure ? { lastFailure } : {}),
});
}
return {
status: desc.status.name,
runningAgents,
...(desc.startTime ? { startedAt: desc.startTime.getTime() } : {}),
...(desc.closeTime ? { closedAt: desc.closeTime.getTime() } : {}),
};
} catch (err) {
if (err instanceof WorkflowNotFoundError) return null;
throw err;
}
}
/** Live progress of a running scan via the getProgress query. Null if the query can't be served (no worker). */
export async function queryProgress(workflowId: string): Promise<PipelineState | null> {
const client = await getClient();
try {
return await client.workflow.getHandle(workflowId).query<PipelineState>('getProgress');
} catch {
// The query needs a live worker; a just-closed scan may have none. Caller falls back to the result.
return null;
}
}
/**
* Deepest message in a Temporal failure's cause chain — the real reason nested under generic
* wrappers (WorkflowFailedError → ActivityFailure → ApplicationFailure). Covers failed, cancelled,
* and terminated alike. Mirrors the SDK's `rootCause` (only exported from @temporalio/common).
*/
function rootFailureMessage(err: WorkflowFailedError): string {
let message = err.message;
let cause: unknown = err.cause;
while (cause instanceof Error && cause.message) {
message = cause.message;
cause = cause.cause;
}
return message;
}
/** Final state of a closed scan: success carries the full PipelineState, failure carries the message. */
export async function getTerminalOutcome(workflowId: string): Promise<TerminalOutcome> {
const client = await getClient();
try {
const state = (await client.workflow.getHandle(workflowId).result()) as PipelineState;
return { kind: 'success', state };
} catch (err) {
if (err instanceof WorkflowFailedError) {
return { kind: 'failed', message: rootFailureMessage(err) };
}
throw err;
}
}
+3 -3
View File
@@ -3,8 +3,6 @@
* whether the user can be prompted interactively.
*/
import { fail } from './errors.js';
/** True when stdout is a real terminal — safe for color, cursor moves, and spinners. */
export function stdoutIsTerminal(): boolean {
return !!process.stdout.isTTY;
@@ -30,5 +28,7 @@ export function supportsColor(): boolean {
/** Exit with a clear error when an interactive-only command has no terminal, instead of hanging on a prompt. */
export function requireInteractive(command: string, alternative: string): void {
if (isInteractive()) return;
fail(`'${command}' needs an interactive terminal.`, alternative);
console.error(`ERROR: '${command}' needs an interactive terminal.`);
console.error(alternative);
process.exit(1);
}
-60
View File
@@ -1,60 +0,0 @@
/**
* Terminal status output for long-running steps.
*
* Commands are run with their output captured rather than inherited, so raw docker
* plumbing never floods the terminal. Progress is shown with a `@clack/prompts`
* spinner. On failure the captured output is printed so the error stays visible
* instead of being swallowed.
*/
import { spawn } from 'node:child_process';
import * as p from '@clack/prompts';
export interface StepResult {
ok: boolean;
output: string;
}
/**
* Run a command capturing stdout and stderr. Resolves the exit result and combined
* output; never rejects. Callers that want a spinner wrap this in one themselves.
*/
export function spawnCaptured(cmd: string, args: string[]): Promise<StepResult> {
return new Promise((resolve) => {
let output = '';
const child = spawn(cmd, args, { stdio: ['ignore', 'pipe', 'pipe'] });
child.stdout?.on('data', (chunk) => {
output += chunk.toString();
});
child.stderr?.on('data', (chunk) => {
output += chunk.toString();
});
child.on('close', (code) => resolve({ ok: code === 0, output }));
child.on('error', () => resolve({ ok: false, output }));
});
}
/** Print captured command output to stderr, so a failure is never swallowed. */
export function surfaceOutput(output: string): void {
const trimmed = output.trim();
if (trimmed) process.stderr.write(`${trimmed}\n`);
}
/**
* Run a command as a labeled step, with a spinner over it. On failure the captured
* output is surfaced. Returns the exit result and captured output.
*/
export async function runStep(label: string, cmd: string, args: string[]): Promise<StepResult> {
const spinner = p.spinner();
spinner.start(label);
const result = await spawnCaptured(cmd, args);
if (result.ok) {
spinner.stop(label);
} else {
spinner.error(label);
surfaceOutput(result.output);
}
return result;
}
+19 -6
View File
@@ -102,6 +102,23 @@
"required": ["login_type", "login_url", "credentials", "success_condition"],
"additionalProperties": false
},
"pipeline": {
"type": "object",
"description": "Pipeline execution settings for retry behavior and concurrency",
"properties": {
"retry_preset": {
"type": "string",
"enum": ["default", "subscription"],
"description": "Retry preset. 'subscription' extends timeouts for Anthropic subscription rate limit windows (5h+)."
},
"max_concurrent_pipelines": {
"type": "string",
"pattern": "^[1-5]$",
"description": "Max concurrent vulnerability pipelines (1-5, default: 5)"
}
},
"additionalProperties": false
},
"rules": {
"type": "object",
"description": "Testing rules that define what to focus on or avoid during penetration testing",
@@ -160,11 +177,6 @@
"minLength": 1,
"maxLength": 500,
"description": "Free-text guidance to the report agent (e.g., 'Drop findings about missing security headers')."
},
"sarif": {
"type": "string",
"enum": ["true", "false"],
"description": "Emit a SARIF 2.1.0 log (report.sarif) beside the report. Requires exploit=true; ignored otherwise."
}
},
"additionalProperties": false
@@ -206,6 +218,7 @@
"properties": {
"description": {
"type": "string",
"minLength": 1,
"maxLength": 200,
"description": "Human-readable description of the rule"
},
@@ -221,7 +234,7 @@
"description": "Value to match"
}
},
"required": ["type", "value"],
"required": ["description", "type", "value"],
"additionalProperties": false
}
}
+5 -2
View File
@@ -96,10 +96,13 @@ rules:
# Report filters applied by the report agent when assembling the final report (optional).
# Example below is illustrative; edit, remove, or add sections as needed.
# report:
# # Emit a SARIF 2.1.0 log (report.sarif) beside the report. Requires exploit: "true".
# sarif: "true"
# min_severity: low
# min_confidence: low
# guidance: |
# Drop findings about missing security headers and rate-limit gaps.
# ...
# Pipeline execution settings (optional)
# pipeline:
# retry_preset: subscription # 'default' or 'subscription' (6h max retry for rate limit recovery)
# max_concurrent_pipelines: 2 # 1-5, default: 5 (reduce to lower API usage spikes)
+3 -3
View File
@@ -19,9 +19,9 @@
"clean": "rm -rf dist"
},
"dependencies": {
"@earendil-works/pi-agent-core": "^0.82.1",
"@earendil-works/pi-ai": "^0.82.1",
"@earendil-works/pi-coding-agent": "^0.82.1",
"@earendil-works/pi-agent-core": "^0.79.1",
"@earendil-works/pi-ai": "^0.79.1",
"@earendil-works/pi-coding-agent": "^0.79.1",
"@gotgenes/pi-permission-system": "^10.9.0",
"@temporalio/activity": "^1.11.0",
"@temporalio/client": "^1.11.0",
+94 -180
View File
@@ -1,198 +1,112 @@
<role>
<exploit_mode_role>
You are the Security Report Writer for a multi-agent security assessment pipeline. Upstream agents have already explored the target application, generated security hypotheses, and verified them by exploitation. Your job is to synthesize the verified findings into structured data that downstream renderers will use to produce reports and persist to the database.
</exploit_mode_role>
<analysis_mode_role>
You are the Security Report Writer for a multi-agent security assessment pipeline. Upstream agents have explored the target application, generated security hypotheses, and assessed them against the source code. Your job is to synthesize those findings into structured data that downstream renderers will use to produce reports and persist to the database.
</analysis_mode_role>
You are an Executive Summary Writer and Report Cleaner for security assessments. Your job is to:
1. MODIFY the existing concatenated report by adding an executive summary at the top
2. CLEAN UP hallucinated or extraneous sections throughout the report
</role>
<task>
Record all findings as structured data using the `add_finding` tool. You do NOT write a markdown report — a downstream renderer produces the report from your structured output.
<audience>
Technical leadership (CTOs, CISOs, Engineering VPs) who need both technical accuracy and executive brevity.
</audience>
1. **Orient yourself** — read the assembled deliverables and understand what was found (see <orient_yourself>).
2. **Filter and clean** — identify real findings, remove noise, rewrite weak titles (see <filter_and_clean>).
3. **Record report metadata** — run `set-report-meta` once (see <record_report_meta>).
4. **Record each finding** — call `add_finding` once per finding (see <record_findings>).
</task>
<objective>
The orchestrator has already concatenated all per-class deliverables into `comprehensive_security_assessment_report.md`. Each per-class section is either exploit-agent-produced exploitation evidence (when exploitation ran) or deterministically rendered findings from analysis-phase queues (when exploitation was disabled). The cleanup rules below apply uniformly to either source.
Your task is to:
1. Read this existing concatenated report
2. Add an Executive Summary (vulnerability overview) at the top
3. Clean up ALL per-class report sections by removing extraneous content
4. Save the modified version back to the same file
<tools_reference>
You have two tools for recording findings:
IMPORTANT: You are MODIFYING an existing file, not creating a new one.
</objective>
- **set-report-meta** (CLI via `bash`) — Write top-level report metadata. Call once before recording findings.
`set-report-meta --target "https://..." --assessment-date "YYYY-MM-DD" --scope "..." --executive-summary "..."`
Returns: `{"status":"success"}`
Shell quoting: wrap flag values in double quotes. Escape any literal double quotes as \", dollar signs as \$, and backticks as \`.
<target>
URL: {{WEB_URL}}
- **add_finding** (tool) — Record a single finding as structured data. Call once per finding. Rejects duplicate finding_ids. The tool schema describes all required and optional fields — fill them in directly.
</tools_reference>
Filesystem:
- {{REPO_PATH}}/ (read only)
- {{REPO_PATH}}/.shannon/deliverables/ (read-write)
- {{REPO_PATH}}/.shannon/scratchpad/ (read-write) - screenshots, scripts, scratch work, etc.
</target>
<orient_yourself>
Before recording anything, read and understand your inputs.
### Your goal
<exploit_mode_orient>
You are the final agent in the pipeline. Upstream agents have already performed reconnaissance, analyzed vulnerabilities, and exploited them. Their evidence has been assembled into a concatenated report. Your job is to read that report, identify the real findings, and emit each one as structured data via the `add_finding` tool.
</exploit_mode_orient>
<analysis_mode_orient>
You are the final agent in the pipeline. Upstream agents have performed reconnaissance and analyzed vulnerabilities in the source code. **No exploitation phase ran** — nothing was executed against the target and no vulnerability was confirmed by attack. Their analysis has been assembled into a concatenated report. Your job is to read that report, identify the real findings, and emit each one as structured data via the `add_finding` tool.
</analysis_mode_orient>
### Your inputs
Read these files:
- `.shannon/deliverables/comprehensive_security_assessment_report.md` — The concatenated per-class deliverables. This is your primary input. Each per-class section contains vulnerability entries with IDs.
- `.shannon/deliverables/pre_recon_deliverable.md` — Initial reconnaissance and technology stack (for executive summary context).
- `.shannon/deliverables/recon_deliverable.md` — Attack surface mapping and endpoint discovery (for executive summary context).
### Vulnerability ID patterns
Findings have IDs matching `[TYPE]-VULN-[NUMBER]` (e.g., INJ-VULN-01, AUTH-VULN-03).
### Context
Target URL: {{WEB_URL}}
Vulnerability classes tested: {{VULN_CLASSES_TESTED}}
Exploitation: {{EXPLOITATION}}
<context>
Authentication Context:
{{AUTH_CONTEXT}}
</orient_yourself>
{{NOT_ASSESSED_CLASSES}}
Vulnerability classes tested in this run: {{VULN_CLASSES_TESTED}}
Exploitation: {{EXPLOITATION}}
</context>
{{REPORT_FILTERS_BLOCK}}
<filter_and_clean>
Read through the concatenated report and identify which vulnerability entries to record. Apply these rules:
<input_files>
You will analyze the following deliverable files:
- `.shannon/deliverables/pre_recon_deliverable.md` - Initial reconnaissance and technology stack
- `.shannon/deliverables/recon_deliverable.md` - Attack surface mapping and endpoint discovery
- `.shannon/deliverables/comprehensive_security_assessment_report.md` - The already-concatenated report that you will modify
</input_files>
### KEEP — these are real findings to record via `add_finding`
- Vulnerability entries under `## {{REPORT_VULN_SUBHEADING}}` sections with IDs matching `### [TYPE]-VULN-[NUMBER]`
<deliverable_instructions>
MODIFY the existing file `.shannon/deliverables/comprehensive_security_assessment_report.md` by:
1. ADDING these sections at the top:
# Security Assessment Report
## Executive Summary
- Target: {{WEB_URL}}
- Assessment Date: {current date}
- Scope: {{VULN_CLASSES_TESTED}}
- Exploitation: {{EXPLOITATION}}
## Summary by Vulnerability Type
{For each vulnerability type below, examine the comprehensive_security_assessment_report.md sections and provide a summary. If no section exists for that type or no vulnerabilities are listed, explicitly state that none were found.}
{{VULN_SUMMARY_SUBSECTIONS}}
2. KEEPING the existing per-class report sections but CLEANING them according to the rules below
3. The final structure should be:
- Executive Summary (new)
- All existing per-class report sections (cleaned)
IMPORTANT: Do NOT reorder the existing per-class report sections. Maintain the exact order they appear in the concatenated report. Only remove sections that do not match the defined criteria above.
</deliverable_instructions>
<instructions>
1. Read the pre_recon and recon deliverable files to gather security-relevant information:
- Skip basic information such as technology stack information (the team knows their own stack)
- Use technical leadership tone - precise but concise
- Use the current date for the assessment date
2. Create the Executive Summary content:
- Executive Summary: Technical overview with actionable findings for engineering leaders
3. Clean the per-class report sections in `.shannon/deliverables/comprehensive_security_assessment_report.md` by applying these rules:
- KEEP these specific section headings:
NOTE: these sections will contain vulnerability lists with IDs matching pattern `### [TYPE]-VULN-[NUMBER]`
* `# [Type] {{REPORT_VULN_HEADING}}`
* `## {{REPORT_VULN_SUBHEADING}}`
{{REPORT_FILTER_RULES}}
- REMOVE ANY OTHER SECTIONS (even if they contain vulnerability IDs), such as:
* `## Potential Vulnerabilities (Validation Blocked)` (All agents)
* Standalone "Recommendations" sections
* "Conclusion" sections
* "Summary" sections
* "Next Steps" sections
* "Additional Analysis" sections
* Any other meta-commentary sections without vulnerability IDs
* False positives sections
* any intros in the sections
* any counts in the sections
- Preserve exact vulnerability IDs (`### [TYPE]-VULN-NN:`); if the title after the colon is only a short category label rather than a descriptive phrase, rewrite it to a concise human-readable descriptor derived from the finding's Vulnerable location and Overview.
### SKIP — do not record these
<exploit_mode_skip>
- `## Potential Vulnerabilities (Validation Blocked)` entries
</exploit_mode_skip>
- Standalone "Recommendations", "Conclusion", "Summary", "Next Steps", "Additional Analysis" sections
- False positives sections
- Introductory text, vulnerability counts, or meta-commentary without vulnerability IDs
- Any section that does not contain a finding with a valid vulnerability ID
4. Combine the content:
- Place the Executive Summary and Network Reconnaissance sections at the top
- Follow with the cleaned per-class report sections
- Save as the modified `.shannon/deliverables/comprehensive_security_assessment_report.md`
### Title cleanup
If a finding's title (the text after the colon in `### TYPE-VULN-NN: Title`) is only a short category label rather than a descriptive phrase, rewrite it to a concise descriptor derived from the finding's "Vulnerable location" and "Overview" fields. Use the improved title when calling `add_finding`.
</filter_and_clean>
CRITICAL: You are modifying the existing concatenated report at `.shannon/deliverables/comprehensive_security_assessment_report.md` IN-PLACE, not creating a separate file.
</instructions>
<record_report_meta>
Run `set-report-meta` once before recording any individual findings (see <tools_reference> for usage).
Fields:
- `target`: `{{WEB_URL}}`
- `assessment_date`: Use the current date in ISO format (YYYY-MM-DD)
- `scope`: `{{VULN_CLASSES_TESTED}}`
<exploit_mode_summary>
- `executive_summary`: 2-3 sentences summarizing the security posture for technical leadership (CTOs, CISOs, Engineering VPs). Must include the target URL and assessment date. Provide a high-level characterization based on the findings — severity distribution, most critical issues, and overall risk demonstrated by exploitation. If no vulnerabilities were confirmed in the assessed classes, state that scope clearly. A clean report is valid only when no <not_assessed_classes> block is present. If that block is present, explicitly say the listed classes were not assessed and do not assert they are free of vulnerabilities.
</exploit_mode_summary>
<analysis_mode_summary>
- `executive_summary`: 2-3 sentences summarizing the security posture for technical leadership (CTOs, CISOs, Engineering VPs). Must include the target URL and assessment date. Provide a high-level characterization based on the findings — severity and confidence distribution, the most serious weaknesses identified, and overall risk. State plainly that this was an analysis-only assessment and that no finding was confirmed by exploitation; do not describe risk as demonstrated or proven, and present severity as assessed rather than measured. If no vulnerabilities were identified in the assessed classes, state that scope clearly. A clean report is valid only when no <not_assessed_classes> block is present. If that block is present, explicitly say the listed classes were not assessed and do not assert they are free of vulnerabilities.
</analysis_mode_summary>
</record_report_meta>
<record_findings>
For each finding identified in <filter_and_clean>, call `add_finding` once.
Record findings in the order they appear in the concatenated report (which groups by vulnerability class: injection, xss, auth, ssrf, authz).
Each `finding_id` may only be recorded once — duplicate calls are rejected.
### How to fill in each field
Map the finding's content from the per-class deliverable sections to `add_finding` fields:
- `finding_id`: The vulnerability ID exactly as it appears (e.g., `"INJ-VULN-01"`, `"AUTH-VULN-07"`)
- `title`: The cleaned-up title (see title cleanup rules in <filter_and_clean>)
- `category`: Derived from the finding type prefix — `INJ` → `"Injection"`, `XSS` → `"XSS"`, `AUTH` → `"Authentication"`, `AUTHZ` → `"Authorization"`, `SSRF` → `"SSRF"`
<exploit_mode_fields>
- `severity`: From the finding's "Severity" field. Use as-is; do not reassess.
</exploit_mode_fields>
<analysis_mode_fields>
- `confidence`: From the finding's "Confidence" field. Use as-is; do not reassess.
- `severity`: The analysis deliverables carry no severity field — no exploit ran to measure impact. Assess it from the vulnerability class and the impact you describe. It is an assessed rating, not a measured one.
</analysis_mode_fields>
- `owasp_category`: Map to the appropriate OWASP Top 10 (2025) category:
- `"A01:2025 — Broken Access Control"`
- `"A02:2025 — Security Misconfiguration"`
- `"A03:2025 — Software Supply Chain Failures"`
- `"A04:2025 — Cryptographic Failures"`
- `"A05:2025 — Injection"`
- `"A06:2025 — Insecure Design"`
- `"A07:2025 — Authentication Failures"`
- `"A08:2025 — Software or Data Integrity Failures"`
- `"A09:2025 — Security Logging and Alerting Failures"`
- `"A10:2025 — Mishandling of Exceptional Conditions"`
- `vulnerable_location`: From the finding's "Vulnerable location" field
- `http_location`: The HTTP request the finding is reached through, when the deliverable names one (e.g. `"GET /api/products?id="` gives `method: "GET"`, `url: "{{WEB_URL}}/api/products"`, `parameter: "id"`). Omit for findings with no network entry point.
- `overview`: Synthesize from the finding's "Overview" field into professional prose. Do not paste verbatim.
- `remediation`: Specific, actionable fix guidance from the finding. Code-level or configuration-level. Avoid generic advice.
<exploit_mode_fields>
- `impact`: From the finding's "Impact" field if present, otherwise derive from the overview and proof of impact
- `auth_state`: From the finding's authentication context or prerequisites
- `prerequisites`: From the finding's "Prerequisites" field, or `"None"` if not specified
- `exploitation_steps`: From the finding's exploitation steps or proof-of-concept. Each step gets a title and ordered prose/code items. Use `"bash"` for shell commands, `"http"` for raw HTTP, `"json"` for response bodies.
- `proof_of_impact`: From the finding's "Proof of Impact" or evidence section. What the exploit demonstrably achieved.
- `status`: Optional. Use `"exploited"` for confirmed exploits.
</exploit_mode_fields>
<analysis_mode_fields>
- `impact`: What an attacker could achieve if this vulnerability were exploited. Derive it from the finding's "Impact" and "Overview" fields. Write it as assessed, never as achieved.
This run had no exploitation phase. Nothing was executed against the target, nothing was demonstrated, and no exploit evidence exists. Accordingly `auth_state`, `prerequisites`, `exploitation_steps`, `proof_of_impact` and `status` are **not** part of your tool schema — the deliverables contain no source for any of them. `confidence` is the deliverable's own rating and carries over verbatim; `severity` is yours to assess, since nothing measured it. Do not compensate for the missing fields by describing attack execution in `overview`, `impact` or `notes`. Report the weakness and how to fix it; that is the whole deliverable for this run.
</analysis_mode_fields>
**Optional fields:**
- `notes`: From the finding's "Notes" section if present
- `additional_sections`: Any extra subsections on the finding that don't fit the fields above
### Zero findings
If no valid findings exist after filtering, do not call `add_finding` at all. The `set-report-meta` executive summary should state that no vulnerabilities were identified in the assessed classes. If a <not_assessed_classes> block is present, it must also state that those listed classes were not assessed.
</record_findings>
<constraints>
<exploit_mode_constraints>
- **No Fabrications:** Do not invent exploitation steps, evidence, or impact. Every piece of data must come from the deliverable files. If a finding has incomplete data, include it but note the gap in `overview`.
- **No Severity Changes:** Use the severity from the deliverable as-is. Do not inflate or deflate.
</exploit_mode_constraints>
<analysis_mode_constraints>
- **No Fabrications:** Every piece of data must come from the deliverable files. If a finding has incomplete data, include it but note the gap in `overview`.
- **Nothing Was Demonstrated:** No exploit ran. Do not write that a vulnerability was confirmed, proven, exploited, or verified against the running target, and do not describe payloads, requests, or responses as having been sent.
- **No Confidence Changes:** Use the confidence from the deliverable as-is. Do not raise or lower it.
- **Severity Is Assessed:** Rate severity from the vulnerability class and the impact you describe. Never present it as measured or demonstrated.
</analysis_mode_constraints>
- **No Speculation:** Only record findings that appear in the deliverables with valid vulnerability IDs. Do not add your own assessments.
- **OWASP 2025:** Map all findings to OWASP Top 10 (2025) categories.
- **Remediation Quality:** Provide specific, actionable remediation — code-level or configuration-level fixes. Avoid generic advice like "validate input" or "follow best practices".
</constraints>
<self_check>
Before finalizing, verify:
- [ ] Did I run `set-report-meta` exactly once with target, assessment_date, scope, and executive_summary?
- [ ] For each valid finding in the deliverables, did I call `add_finding` exactly once with the correct `finding_id`?
<exploit_mode_skipcheck>
- [ ] Did I skip all entries from "Potential Vulnerabilities (Validation Blocked)", false positives, and meta-commentary sections?
</exploit_mode_skipcheck>
<analysis_mode_skipcheck>
- [ ] Did I skip false positives and meta-commentary sections?
</analysis_mode_skipcheck>
- [ ] Does every finding have a valid `owasp_category` using the 2025 categorization?
- [ ] Does every finding have `overview`, `impact`, and `remediation`?
<exploit_mode_checks>
- [ ] Does every finding have `auth_state` and `prerequisites`?
- [ ] Does every finding have `exploitation_steps` with prose/code items?
- [ ] Does every finding have `proof_of_impact`?
- [ ] Are severity ratings unchanged from the source deliverables?
</exploit_mode_checks>
<analysis_mode_checks>
- [ ] Does every finding have `confidence` carried over unchanged from the deliverable?
- [ ] Is every `severity` assessed from the impact I described, with no claim that it was measured?
- [ ] Is every `impact` phrased as assessed rather than demonstrated, with no claim that anything was executed?
</analysis_mode_checks>
- [ ] Are remediation recommendations specific and actionable (not generic)?
If any answer is NO, fix it before finalizing.
</self_check>
+1 -5
View File
@@ -100,11 +100,7 @@ An **exploitable vulnerability** is a logical flaw in the code that represents a
"missing_defense": "Concise description of the core problem (e.g., 'No rate limit on POST /login', 'Session ID not rotated after login').",
"exploitation_hypothesis": "A clear statement describing the expected outcome of a successful exploit. This should describe an active attack outcome, not just vulnerability confirmation (e.g., 'An attacker can successfully log into an account by guessing a simple password,' 'An attacker can hijack a user's session by replaying a stolen cookie').",
"suggested_exploit_technique": "The specific attack pattern to attempt, derived from the methodology. The exploitation agent should actively execute this attack, not just confirm it's possible (e.g., 'brute_force_login', 'credential_stuffing', 'session_hijacking', 'session_fixation').",
"confidence": "high | medium | low",
"code_locations": [
{ "file": "lib/insecurity.ts", "start_line": 21, "role": "sink", "symbol": "verify" },
{ "file": "routes/login.ts", "start_line": 34, "role": "guard" }
],
"confidence": "High | Medium | Low",
"notes": "Relevant details about required session state, applicable roles, observed headers, or links to related findings."
}
</exploitation_queue_format>
+2 -6
View File
@@ -107,11 +107,7 @@ An **exploitable vulnerability** is a logical flaw in the code that represents a
"side_effect": "specific unauthorized action possible (e.g., 'read other users profiles', 'delete any file', 'skip payment')",
"reason": "1-2 lines explaining why this is vulnerable",
"minimal_witness": "brief description of how to trigger (e.g., 'change user_id param to another user', 'call admin endpoint as regular user')",
"confidence": "high | medium | low",
"code_locations": [
{ "file": "server.ts", "start_line": 365, "role": "sink", "symbol": "getUserById" },
{ "file": "lib/insecurity.ts", "start_line": 92, "role": "guard" }
],
"confidence": "high | med | low",
"notes": "any assumptions, related findings, or special conditions"
}
</exploitation_queue_format>
@@ -224,7 +220,7 @@ An **exploitable vulnerability** is a logical flaw in the code that represents a
- `guard_evidence` (missing/misplaced),
- `side_effect` observed,
- `reason` (1–2 lines: e.g., "ownership check absent"),
- `confidence` (high/medium/low),
- `confidence` (high/med/low),
- `minimal_witness` (sketch for exploit agent).
---
+2 -6
View File
@@ -111,11 +111,7 @@ An **exploitable vulnerability** is a confirmed source-to-sink path where the en
"verdict": "safe | vulnerable.",
"mismatch_reason": "if vulnerable, 1–2 lines in plain language.",
"witness_payload": "minimal input you'd use later to show structure influence (e.g., ' for SQLi, ; ls -la for Command Injection, ../../../../etc/passwd for LFI, {{7*7}} for SSTI).",
"confidence": "high | medium | low",
"code_locations": [
{ "file": "routes/search.ts", "start_line": 42, "role": "sink", "symbol": "searchProducts" },
{ "file": "routes/search.ts", "start_line": 38, "role": "source" }
],
"confidence": "high | med | low.",
"notes": "assumptions, untraversed branches, anything unusual."
}
</exploitation_queue_format>
@@ -168,7 +164,7 @@ An **exploitable vulnerability** is a confirmed source-to-sink path where the en
- `verdict` (`safe` / `vulnerable`)
- `mismatch_reason` (plain-language, 1–2 lines)
- `witness_payload` (minimal input to demonstrate structure influence — **for later exploit phase**)
- `confidence` (`high` / `medium` / `low`)
- `confidence` (`high` / `med` / `low`)
- `notes` (assumptions, untraversed branches, unusual conditions)
- **7) Score confidence**
- **High:** binds on value/like/numeric; strict casts; whitelists for all syntax slots; **no** post-sanitization concat.
+1 -5
View File
@@ -102,11 +102,7 @@ An **exploitable vulnerability** is a data flow where user-controlled input infl
"missing_defense": "Concise description of the core problem (e.g., 'No URL allowlist validation', 'Missing protocol restriction').",
"exploitation_hypothesis": "A clear statement describing the expected outcome of a successful exploit (e.g., 'An attacker can access internal services by providing internal URLs', 'An attacker can retrieve cloud metadata through URL manipulation').",
"suggested_exploit_technique": "The specific attack pattern to attempt, derived from the methodology (e.g., 'internal_service_access', 'cloud_metadata_retrieval', 'port_scanning', 'webhook_abuse').",
"confidence": "high | medium | low",
"code_locations": [
{ "file": "routes/profileImageUrlUpload.ts", "start_line": 24, "role": "sink", "symbol": "profileImageUrlUpload" },
{ "file": "routes/profileImageUrlUpload.ts", "start_line": 18, "role": "source" }
],
"confidence": "High | Medium | Low",
"notes": "Relevant details about request format, authentication requirements, observed timeouts, or links to related findings."
}
</exploitation_queue_format>
+1 -5
View File
@@ -108,11 +108,7 @@ Structure: The vulnerability JSON object MUST follow this exact format:
"verdict": "vulnerable | safe.",
"mismatch_reason": "If vulnerable, explain why the observed encoding is wrong for the render context (e.g., 'URL encoding used in an HTML attribute context, allowing event handler injection.').",
"witness_payload": "A minimal, non-malicious payload that proves context control (e.g., '><script>alert(1)</script>', '" onmouseover=alert(1) ').",
"confidence": "high | medium | low",
"code_locations": [
{ "file": "frontend/src/app/search-result/search-result.component.ts", "start_line": 121, "role": "sink", "symbol": "filterTable" },
{ "file": "frontend/src/app/search-result/search-result.component.ts", "start_line": 115, "role": "source" }
],
"confidence": "high | med | low.",
"notes": "Relevant CSP, HttpOnly flags, WAF behavior, or other environmental factors."
}
</exploitation_queue_format>
@@ -23,7 +23,7 @@ function evaluateBashTimeout(timeout: number | undefined): ToolCallEventResult |
if (!hasValidTimeout) {
return {
block: true,
reason: `A timeout in seconds is required for the bash tool. The bash tool was not executed. Use the default of ${DEFAULT_TIMEOUT_SECONDS} seconds, or up to a maximum of ${MAX_TIMEOUT_SECONDS} seconds.`,
reason: `Set bash 'timeout' (seconds). Default ${DEFAULT_TIMEOUT_SECONDS}s, max ${MAX_TIMEOUT_SECONDS}s.`,
};
}
+111 -292
View File
@@ -5,338 +5,157 @@
// as published by the Free Software Foundation.
/**
* Model selection and resolution for the pi harness.
* Model tier definitions and resolution for the pi harness.
*
* One model runs the entire workflow. Users name it with a single setting:
* Three tiers mapped to capability levels:
* - "small" (Haiku — summarization, structured extraction)
* - "medium" (Sonnet — tool use, general analysis)
* - "large" (Opus — deep reasoning, complex analysis)
*
* SHANNON_AI_MODEL=<provider>:<model-id>
* Users override per tier via ANTHROPIC_SMALL_MODEL / ANTHROPIC_MEDIUM_MODEL /
* ANTHROPIC_LARGE_MODEL, which works across all providers (Anthropic, Bedrock,
* custom base URL).
*
* The provider half decides the endpoint, the credential, and the API dialect;
* the model half is passed to pi's registry as-is. The separator is a colon
* because model IDs routinely contain slashes, and it is the *first* colon that
* splits, because Bedrock model IDs contain colons of their own
* (`amazon-bedrock:us.anthropic.claude-opus-4-5-20251101-v1:0`).
*
* Resolution returns a pi `Model` plus the `ModelRuntime` that owns its auth,
* built over an in-memory credential store primed from the environment.
* The active provider is chosen from the env-var contract the CLI forwards
* (`CLAUDE_CODE_USE_BEDROCK`, `ANTHROPIC_BASE_URL`+`ANTHROPIC_AUTH_TOKEN`, else
* direct Anthropic). Resolution returns a pi `Model` via `ModelRegistry.find`, the
* `thinkingLevel`, and an `AuthStorage` primed with the right credential. Bedrock
* authenticates from the AWS_ env vars via pi-ai.
*/
import { existsSync } from 'node:fs';
import path from 'node:path';
import type { Api, Credential, CredentialInfo, CredentialStore, Model } from '@earendil-works/pi-ai';
import { getAgentDir, ModelRuntime } from '@earendil-works/pi-coding-agent';
import type { ThinkingLevel } from '@earendil-works/pi-agent-core';
import type { Api, Model } from '@earendil-works/pi-ai';
import { AuthStorage, type ModelRegistry } from '@earendil-works/pi-coding-agent';
/**
* Providers Shannon curates with their own credential variables, config sections,
* and setup flows. Each is a pi-ai provider id; any other pi provider is still
* reachable through the generic credential path below.
*/
export const CURATED_PROVIDERS = ['anthropic', 'openai', 'xai', 'amazon-bedrock'] as const;
export type ModelTier = 'small' | 'medium' | 'large';
export type CuratedProviderId = (typeof CURATED_PROVIDERS)[number];
function isCuratedProvider(value: string): value is CuratedProviderId {
return (CURATED_PROVIDERS as readonly string[]).includes(value);
}
/** Generic API key, honored for any provider Shannon does not curate. */
export const GENERIC_API_KEY_ENV = 'SHANNON_AI_API_KEY';
/**
* Env vars carrying each curated provider's API key, in precedence order. Shannon
* does not invent credential names — these are the variables each provider's own
* tooling uses. Bedrock pairs its bearer token with AWS_REGION, which is provider
* config rather than a credential.
*/
export const PROVIDER_API_KEY_ENV: Readonly<Record<CuratedProviderId, readonly string[]>> = {
anthropic: ['ANTHROPIC_API_KEY', 'CLAUDE_CODE_OAUTH_TOKEN'],
openai: ['OPENAI_API_KEY'],
xai: ['XAI_API_KEY'],
'amazon-bedrock': ['AWS_BEARER_TOKEN_BEDROCK'],
const DEFAULT_MODELS: Readonly<Record<ModelTier, string>> = {
small: 'claude-haiku-4-5-20251001',
medium: 'claude-sonnet-4-6',
large: 'claude-opus-4-8',
};
/** Model used when SHANNON_AI_MODEL is unset. */
export const DEFAULT_MODEL_SPEC = 'anthropic:claude-sonnet-4-6';
/** Browsable pi model catalogue — the source of valid `<provider>:<model-id>` ids. */
export const PI_CATALOG_URL = 'https://pi.dev/models';
/**
* Wire formats an OpenAI-compatible gateway may serve, named by
* SHANNON_AI_OPENAI_FORMAT. Only `openai` offers a choice: every other supported
* provider has exactly one API in pi's registry.
*/
export const OPENAI_FORMATS = {
'chat-completions': 'openai-completions',
responses: 'openai-responses',
} as const;
export type OpenAiFormat = keyof typeof OPENAI_FORMATS;
/** Format assumed when a gateway is configured but no format is named. */
export const DEFAULT_OPENAI_FORMAT: OpenAiFormat = 'chat-completions';
function isOpenAiFormat(value: string): value is OpenAiFormat {
return value in OPENAI_FORMATS;
}
/**
* Read SHANNON_AI_OPENAI_FORMAT. Unset returns undefined, which lets the caller
* distinguish "not configured" from an explicit choice and reject the variable
* where it has no effect.
*/
export function resolveOpenAiFormat(): OpenAiFormat | undefined {
const raw = process.env.SHANNON_AI_OPENAI_FORMAT?.trim();
if (!raw) return undefined;
if (!isOpenAiFormat(raw)) {
throw new Error(
`SHANNON_AI_OPENAI_FORMAT must be one of: ${Object.keys(OPENAI_FORMATS).join(', ')}. Got "${raw}".`,
);
}
return raw;
}
export interface ModelSpec {
export interface EffectiveProvider {
/** pi-ai provider id: 'anthropic' or 'amazon-bedrock'. */
providerId: string;
modelId: string;
}
/**
* Parse a `<provider>:<model-id>` spec. Splits on the first colon only, so colons
* inside a model ID survive. The provider id is passed through as given — pi's
* registry validates it later — so this throws only on a malformed spec.
*/
export function parseModelSpec(spec: string): ModelSpec {
const trimmed = spec.trim();
const separator = trimmed.indexOf(':');
if (separator === -1) {
throw new Error(
`SHANNON_AI_MODEL must be "<provider>:<model-id>", got "${trimmed}". Example: ${DEFAULT_MODEL_SPEC}`,
);
}
const providerId = trimmed.slice(0, separator).trim();
const modelId = trimmed.slice(separator + 1).trim();
if (!providerId || !modelId) {
throw new Error(
`SHANNON_AI_MODEL must be "<provider>:<model-id>", got "${trimmed}". Example: ${DEFAULT_MODEL_SPEC}`,
);
}
return { providerId, modelId };
}
/** Resolve the run's model from SHANNON_AI_MODEL, falling back to the default. */
export function resolveModelSpec(): ModelSpec {
return parseModelSpec(process.env.SHANNON_AI_MODEL || DEFAULT_MODEL_SPEC);
}
export interface ProviderCredentials {
/** Endpoint override, applied whatever the provider (proxies, gateways). */
/** Custom-base-URL override applied to the resolved anthropic model. */
baseUrl?: string;
/** Runtime API key primed into the ModelRuntime's credential store. */
apiKey?: string;
/** Runtime credential to prime on AuthStorage for the 'anthropic' provider. */
anthropicToken?: string;
}
/**
* Collect the API key and optional endpoint override for a provider. A curated
* provider's own variables win, then the generic SHANNON_AI_API_KEY. Bedrock is
* excluded — it authenticates through its AWS_ variables, which pi reads directly.
* Determine the active provider + auth from the env-var contract the CLI forwards:
* `CLAUDE_CODE_USE_BEDROCK` → Bedrock; `ANTHROPIC_BASE_URL`+`ANTHROPIC_AUTH_TOKEN`
* → custom base URL; else direct Anthropic (`ANTHROPIC_API_KEY`, or
* `CLAUDE_CODE_OAUTH_TOKEN`). Bedrock authenticates from the AWS_ env vars via
* pi-ai, so it needs no anthropic token.
*/
export function resolveProviderCredentials(providerId: string): ProviderCredentials {
const credentials: ProviderCredentials = {};
const namedVars = isCuratedProvider(providerId) ? PROVIDER_API_KEY_ENV[providerId] : [];
for (const name of namedVars) {
const value = process.env[name];
if (value) {
credentials.apiKey = value;
break;
}
export function resolveEffectiveProvider(): EffectiveProvider {
// Bedrock — env flag.
if (process.env.CLAUDE_CODE_USE_BEDROCK === '1') {
return { providerId: 'amazon-bedrock' };
}
if (!credentials.apiKey && providerId !== 'amazon-bedrock' && process.env[GENERIC_API_KEY_ENV]) {
credentials.apiKey = process.env[GENERIC_API_KEY_ENV];
}
if (process.env.SHANNON_AI_BASE_URL) credentials.baseUrl = process.env.SHANNON_AI_BASE_URL;
return credentials;
// Custom base URL — env contract.
if (process.env.ANTHROPIC_BASE_URL && process.env.ANTHROPIC_AUTH_TOKEN) {
return {
providerId: 'anthropic',
baseUrl: process.env.ANTHROPIC_BASE_URL,
anthropicToken: process.env.ANTHROPIC_AUTH_TOKEN,
};
}
// Direct Anthropic (API key, or OAuth token).
const eff: EffectiveProvider = { providerId: 'anthropic' };
const token = process.env.ANTHROPIC_API_KEY ?? process.env.CLAUDE_CODE_OAUTH_TOKEN;
if (token) eff.anthropicToken = token;
return eff;
}
/** Resolve a model tier to a concrete model ID (env override → default). */
export function resolveModelId(tier: ModelTier = 'medium'): string {
switch (tier) {
case 'small':
return process.env.ANTHROPIC_SMALL_MODEL || DEFAULT_MODELS.small;
case 'large':
return process.env.ANTHROPIC_LARGE_MODEL || DEFAULT_MODELS.large;
default:
return process.env.ANTHROPIC_MEDIUM_MODEL || DEFAULT_MODELS.medium;
}
}
/** Whether a model supports adaptive thinking. Opus 4.6, 4.7, and 4.8 only. */
export function supportsAdaptiveThinking(model: string): boolean {
return /opus-4-[678]/.test(model);
}
/**
* In-memory credential store holding the selected provider's API key.
* Resolve the thinking level for a run.
*
* pi ships the `CredentialStore` interface but no in-memory implementation — its
* own store reads `auth.json` from disk. Shannon's credentials arrive as env vars
* in an ephemeral container, so nothing may be read from or written to disk.
* Adaptive thinking is enabled only on capable models (Opus 4.6/4.7/4.8), mapped to
* pi's 'medium' level; every other model runs with thinking 'off'. The
* CLAUDE_ADAPTIVE_THINKING=false kill switch forces 'off' regardless of model.
*/
class RuntimeCredentialStore implements CredentialStore {
private readonly credentials = new Map<string, Credential>();
constructor(providerId: string, apiKey: string | undefined) {
if (apiKey) {
this.credentials.set(providerId, { type: 'api_key', key: apiKey });
}
}
async read(providerId: string): Promise<Credential | undefined> {
return this.credentials.get(providerId);
}
async list(): Promise<readonly CredentialInfo[]> {
return [...this.credentials].map(([providerId, credential]) => ({ providerId, type: credential.type }));
}
/** Serialized read-modify-write. `fn` returning undefined leaves the entry alone. */
async modify(
providerId: string,
fn: (current: Credential | undefined) => Promise<Credential | undefined>,
): Promise<Credential | undefined> {
const next = await fn(this.credentials.get(providerId));
if (next !== undefined) {
this.credentials.set(providerId, next);
}
return this.credentials.get(providerId);
}
async delete(providerId: string): Promise<void> {
this.credentials.delete(providerId);
}
}
/** The file pi reads credentials from: the agent dir's auth.json. */
function piAuthPath(): string {
return path.join(getAgentDir(), 'auth.json');
}
/** Whether the host's pi credentials are mounted (auth.json present in the agent dir). */
export function piAuthPresent(): boolean {
return existsSync(piAuthPath());
}
/**
* Build a ModelRuntime whose only credential is the one supplied. Model catalogs
* stay offline (`allowModelNetwork` defaults to false) so a scan never blocks on
* a catalog refresh.
*
* When the host's pi auth.json is present, the runtime reads it instead: pi's
* disk-backed store resolves the credential. The mount is writable so OAuth
* refreshes persist to the host for subsequent runs.
*/
export async function createModelRuntime(providerId: string, apiKey: string | undefined): Promise<ModelRuntime> {
if (piAuthPresent()) {
return ModelRuntime.create({ authPath: piAuthPath() });
}
return ModelRuntime.create({ credentials: new RuntimeCredentialStore(providerId, apiKey) });
export function resolveThinkingLevel(modelId: string): ThinkingLevel {
if (process.env.CLAUDE_ADAPTIVE_THINKING === 'false') return 'off';
return supportsAdaptiveThinking(modelId) ? 'medium' : 'off';
}
export interface ModelSelection {
model: Model<Api>;
modelRuntime: ModelRuntime;
thinkingLevel: ThinkingLevel;
authStorage: AuthStorage;
modelId: string;
providerId: string;
}
/**
* Point a model descriptor at a gateway.
*
* An OpenAI gateway may serve either wire format, named by
* SHANNON_AI_OPENAI_FORMAT and defaulting to chat completions, which is what
* most gateway software exposes. Switching to completions also drops the stored
* `compat` block: the catalogue's block describes Responses, and an explicit
* entry outranks pi's `detectCompat`, so leaving it would apply Responses
* settings to a completions request. Staying on Responses keeps it, since it
* then describes the format in use. Every other provider has one API and only
* changes address.
* Resolve the active provider (see resolveEffectiveProvider), prime an AuthStorage
* with its credential, and resolve the tier's model from a fresh ModelRegistry.
* Anthropic / custom-base-URL use a runtime anthropic key; Bedrock authenticates
* from the AWS_ env vars (bearer token primed explicitly as a belt-and-suspenders).
*/
function pointAtGateway(model: Model<Api>, providerId: string, baseUrl: string, format: OpenAiFormat): Model<Api> {
if (providerId !== 'openai') return { ...model, baseUrl };
if (format === 'responses') return { ...model, baseUrl, api: OPENAI_FORMATS.responses };
export function resolveModelSelection(
registryFactory: (authStorage: AuthStorage) => ModelRegistry,
modelTier: ModelTier,
): ModelSelection {
const eff = resolveEffectiveProvider();
const modelId = resolveModelId(modelTier);
const { compat: _responsesCompat, ...withoutCompat } = model;
return { ...withoutCompat, baseUrl, api: OPENAI_FORMATS['chat-completions'] };
}
/**
* Resolve a model against a runtime.
*
* Direct to a provider, the model must exist in the catalogue. Behind a custom
* endpoint it need not: a gateway may serve models under its own names, so an
* unknown id is passed through on a descriptor borrowed from the provider's
* catalogue for its API dialect. Cost and context window on such a descriptor
* are the reference model's, so spend figures are approximate there.
*
* Returns undefined when the id is unresolvable — unknown with no endpoint
* override, or a provider carrying no models at all.
*/
export function resolveModel(
modelRuntime: ModelRuntime,
providerId: string,
modelId: string,
baseUrl: string | undefined,
format: OpenAiFormat = DEFAULT_OPENAI_FORMAT,
): Model<Api> | undefined {
const found = modelRuntime.getModel(providerId, modelId);
if (found) {
return baseUrl ? pointAtGateway(found, providerId, baseUrl, format) : found;
const authStorage = AuthStorage.inMemory();
if (eff.providerId === 'anthropic' && eff.anthropicToken) {
authStorage.setRuntimeApiKey('anthropic', eff.anthropicToken);
}
if (!baseUrl) return undefined;
const reference = modelRuntime.getModels(providerId)[0];
if (!reference) return undefined;
return pointAtGateway({ ...reference, id: modelId, name: modelId }, providerId, baseUrl, format);
}
/**
* Validate SHANNON_AI_OPENAI_FORMAT against the rest of the configuration and
* return the format a gateway run should use.
*
* The variable only reaches a request when both an OpenAI model and a gateway
* are configured, so it is rejected outside that combination rather than
* silently ignored.
*/
export function resolveGatewayFormat(providerId: string, baseUrl: string | undefined): OpenAiFormat {
const configured = resolveOpenAiFormat();
if (!configured) return DEFAULT_OPENAI_FORMAT;
if (providerId !== 'openai') {
throw new Error(
`SHANNON_AI_OPENAI_FORMAT applies to openai models only, but SHANNON_AI_MODEL selects "${providerId}". ` +
`${providerId} serves a single API, so there is no format to choose.`,
);
// Bedrock auth flows from the AWS_ env vars; prime the bearer token explicitly so
// it resolves via AuthStorage in addition to pi-ai's own env fallback.
if (eff.providerId === 'amazon-bedrock' && process.env.AWS_BEARER_TOKEN_BEDROCK) {
authStorage.setRuntimeApiKey('amazon-bedrock', process.env.AWS_BEARER_TOKEN_BEDROCK);
}
if (!baseUrl) {
throw new Error(
'SHANNON_AI_OPENAI_FORMAT applies to gateway runs only. Set SHANNON_AI_BASE_URL, or unset the format to call OpenAI directly.',
);
const registry = registryFactory(authStorage);
const found = registry.find(eff.providerId, modelId);
if (!found) {
throw new Error(`Model not found in pi registry: provider="${eff.providerId}" model="${modelId}"`);
}
return configured;
}
/**
* Resolve SHANNON_AI_MODEL, build a ModelRuntime primed with the provider's
* credential, and look the model up in it.
*/
export async function resolveModelSelection(): Promise<ModelSelection> {
const { providerId, modelId } = resolveModelSpec();
const credentials = resolveProviderCredentials(providerId);
const format = resolveGatewayFormat(providerId, credentials.baseUrl);
const modelRuntime = await createModelRuntime(providerId, credentials.apiKey);
const model = resolveModel(modelRuntime, providerId, modelId, credentials.baseUrl, format);
if (!model) {
throw new Error(
`Model not found in pi registry: provider="${providerId}" model="${modelId}". Browse valid providers and models at ${PI_CATALOG_URL}.`,
);
}
// Custom base URL: override the resolved model's endpoint.
const model: Model<Api> = eff.baseUrl ? { ...found, baseUrl: eff.baseUrl } : found;
return {
model,
modelRuntime,
thinkingLevel: resolveThinkingLevel(modelId),
authStorage,
modelId,
providerId,
providerId: eff.providerId,
};
}
/**
* Whether a model is in the Fable family. Fable's safety classifiers flag
* cybersecurity tasks and route them to Opus 4.8, so a security scan on Fable
* largely runs on Opus 4.8 anyway.
*/
export function isFableModel(model: string): boolean {
return /fable/i.test(model);
}
+69 -86
View File
@@ -9,11 +9,11 @@
import os from 'node:os';
import type { AgentMessage } from '@earendil-works/pi-agent-core';
import {
type AgentSession,
type AgentSessionEvent,
createAgentSession,
DefaultResourceLoader,
getAgentDir,
ModelRegistry,
type ResourceLoader,
SessionManager,
SettingsManager,
@@ -23,14 +23,16 @@ import {
import { fs, path } from 'zx';
import type { AuditSession } from '../../audit/index.js';
import { BASH_TIMEOUT_EXTENSION_DIR, deliverablesDir } from '../../paths.js';
import { isRetryableFailure, PentestError } from '../../services/error-handling.js';
import { isRetryableError, PentestError } from '../../services/error-handling.js';
import { AGENT_VALIDATORS } from '../../session-manager.js';
import type { ActivityLogger } from '../../types/activity-logger.js';
import { ErrorCode } from '../../types/errors.js';
import { isSpendingCapBehavior, matchesBillingTextPattern } from '../../utils/billing-detection.js';
import { isBrowserAgent } from '../../utils/browser-agents.js';
import { formatTimestamp } from '../../utils/formatting.js';
import { Timer } from '../../utils/metrics.js';
import { createAuditLogger } from '../audit-logger.js';
import { resolveModelSelection } from '../models.js';
import { type ModelTier, resolveModelSelection } from '../models.js';
import {
detectExecutionContext,
formatAssistantOutput,
@@ -41,10 +43,8 @@ import {
import { createProgressManager } from '../progress-manager.js';
import type { CapturedSubmitTool } from '../submit-tool.js';
import { permissionSystemConfigExists, permissionSystemPackageDir } from './permission-system.js';
import { PI_RETRY_SETTINGS } from './retry-settings.js';
import { createGlobTool, createTodoWriteTool } from './session-tools.js';
import { createTaskTool } from './task-tool.js';
import { providerTurnError } from './turn-error.js';
declare global {
var SHANNON_DISABLE_LOADER: boolean | undefined;
@@ -105,41 +105,15 @@ async function buildResourceLoader(
return loader;
}
interface ChildUsage {
cost: number;
inputTokens: number;
outputTokens: number;
cacheReadTokens: number;
cacheWriteTokens: number;
}
/**
* Usage for one agent: the parent session plus every `task` sub-session it
* spawned. Sub-sessions keep their own stats, so their spend is accumulated
* separately and added here.
*/
function totalUsage(session: AgentSession | undefined, childUsage: ChildUsage) {
const stats = session?.getSessionStats();
return {
cost: (stats?.cost ?? 0) + childUsage.cost,
inputTokens: (stats?.tokens.input ?? 0) + childUsage.inputTokens,
outputTokens: (stats?.tokens.output ?? 0) + childUsage.outputTokens,
cacheReadTokens: (stats?.tokens.cacheRead ?? 0) + childUsage.cacheReadTokens,
cacheWriteTokens: (stats?.tokens.cacheWrite ?? 0) + childUsage.cacheWriteTokens,
};
}
export interface PiPromptResult {
result?: string | null | undefined;
success: boolean;
duration: number;
turns?: number | undefined;
cost: number;
inputTokens?: number | undefined;
outputTokens?: number | undefined;
cacheReadTokens?: number | undefined;
cacheWriteTokens?: number | undefined;
model?: string | undefined;
partialCost?: number | undefined;
apiErrorDetected?: boolean | undefined;
error?: string | undefined;
errorType?: string | undefined;
prompt?: string | undefined;
@@ -164,7 +138,7 @@ async function writeErrorLog(
timestamp: formatTimestamp(),
agent: 'pi-executor',
error: { name: err.constructor.name, message: err.message, code: err.code, status: err.status, stack: err.stack },
context: { sourceDir, prompt: `${fullPrompt.slice(0, 200)}...`, retryable: isRetryableFailure(err) },
context: { sourceDir, prompt: `${fullPrompt.slice(0, 200)}...`, retryable: isRetryableError(err) },
duration,
};
const logPath = path.join(deliverablesDir(sourceDir), 'error.log');
@@ -216,6 +190,28 @@ function extractAssistantText(message: AgentMessage): string {
.join('\n');
}
/**
* Classify error-bearing text into a PentestError, mirroring the prior provider error
* handling. Spending-cap / billing text is retryable (Temporal backs off and
* recovers when the cap resets); session limit is permanent.
*/
function classifyErrorText(content: string): PentestError | null {
if (!content) return null;
if (matchesBillingTextPattern(content)) {
return new PentestError(
`Billing limit reached: ${content.slice(0, 100)}`,
'billing',
true,
{},
ErrorCode.SPENDING_CAP_REACHED,
);
}
if (content.toLowerCase().includes('session limit reached')) {
return new PentestError('Session limit reached', 'billing', false);
}
return null;
}
// Low-level pi execution. Drives one agent session to completion with progress and
// audit logging. Exported for Temporal activities to call single-attempt execution.
export async function runPiPrompt(
@@ -226,6 +222,7 @@ export async function runPiPrompt(
agentName: string | null = null,
auditSession: AuditSession | null = null,
logger: ActivityLogger,
modelTier: ModelTier = 'medium',
callerTools?: ToolDefinition[],
deliverablesSubdir?: string,
cancellationSignal?: AbortSignal,
@@ -257,22 +254,21 @@ export async function runPiPrompt(
// 4. Resolve model + auth, then assemble the tool set (universal task/todo tools
// plus any caller-supplied collector/submit tools).
const selection = await resolveModelSelection();
const selection = resolveModelSelection((auth) => ModelRegistry.create(auth), modelTier);
const resourceLoader = await buildResourceLoader(sourceDir, logger, agentName);
// Accumulates usage from in-process `task` child sessions so the parent's reported
// cost includes sub-agent spend (their getSessionStats is separate from ours).
const childUsage: ChildUsage = { cost: 0, inputTokens: 0, outputTokens: 0, cacheReadTokens: 0, cacheWriteTokens: 0 };
const childUsage = { cost: 0, inputTokens: 0, outputTokens: 0 };
const customTools: ToolDefinition[] = [
createTaskTool({
model: selection.model,
modelRuntime: selection.modelRuntime,
thinkingLevel: selection.thinkingLevel,
authStorage: selection.authStorage,
cwd: sourceDir,
onUsage: (usage) => {
childUsage.cost += usage.cost;
childUsage.inputTokens += usage.inputTokens;
childUsage.outputTokens += usage.outputTokens;
childUsage.cacheReadTokens += usage.cacheReadTokens;
childUsage.cacheWriteTokens += usage.cacheWriteTokens;
},
resourceLoader,
...(cancellationSignal && { cancellationSignal }),
@@ -287,40 +283,24 @@ export async function runPiPrompt(
let turnCount = 0;
let pendingError: PentestError | null = null;
// Declared out here so the catch can bill spend accrued before a failure.
let session: AgentSession | undefined;
// Abort the in-flight agent when the Temporal activity is cancelled (UI/CLI cancel).
// Without this the top-level session runs to startToCloseTimeout despite the cancel.
const onCancellation = (): void => {
void session?.abort().catch(() => {
// Best-effort — the session is torn down regardless once the prompt unwinds.
});
};
let apiErrorDetected = false;
progress.start();
try {
({ session } = await createAgentSession({
const { session } = await createAgentSession({
cwd: sourceDir,
model: selection.model,
thinkingLevel: selection.thinkingLevel,
tools,
customTools,
modelRuntime: selection.modelRuntime,
authStorage: selection.authStorage,
sessionManager: SessionManager.inMemory(),
// Temporal owns agent restarts, pi absorbs transport faults (see
// PI_RETRY_SETTINGS); compaction stays on to guard against context overflow
// on long agent runs.
settingsManager: SettingsManager.inMemory({ retry: PI_RETRY_SETTINGS, compaction: { enabled: true } }),
// Temporal owns retry; pi compaction stays on (no analog previously, guards
// against context overflow on long agent runs).
settingsManager: SettingsManager.inMemory({ retry: { enabled: false }, compaction: { enabled: true } }),
resourceLoader,
}));
// Wire activity cancellation to the session now that it exists.
if (cancellationSignal?.aborted) {
onCancellation();
} else {
cancellationSignal?.addEventListener('abort', onCancellation, { once: true });
}
});
// 5. Map pi events to audit logging + progress + error capture.
session.subscribe((event: AgentSessionEvent) => {
@@ -334,9 +314,15 @@ export async function runPiPrompt(
progress.stop();
outputLines(formatAssistantOutput(text, execContext, turnCount, description));
progress.start();
const billing = classifyErrorText(text);
if (billing) pendingError = billing;
}
if (msg.role === 'assistant' && msg.stopReason === 'error') {
pendingError = pendingError ?? providerTurnError(msg, 'Agent error', selection.model.contextWindow);
apiErrorDetected = true;
pendingError =
pendingError ??
classifyErrorText(msg.errorMessage ?? '') ??
new PentestError(`Agent error: ${(msg.errorMessage ?? 'unknown').slice(0, 200)}`, 'unknown', true);
}
break;
}
@@ -362,6 +348,7 @@ export async function runPiPrompt(
if (!event.aborted && !event.willRetry && event.errorMessage) {
pendingError =
pendingError ??
classifyErrorText(event.errorMessage) ??
new PentestError(`Context compaction failed: ${event.errorMessage.slice(0, 200)}`, 'unknown', true);
}
break;
@@ -378,9 +365,19 @@ export async function runPiPrompt(
if (pendingError) throw pendingError;
// 8. Read usage/cost and final text.
const usage = totalUsage(session, childUsage);
const stats = session.getSessionStats();
const totalCost = stats.cost + childUsage.cost;
const result = session.getLastAssistantText() ?? null;
// 9. Defense-in-depth: detect a spending cap that produced an empty/cheap run.
if (isSpendingCapBehavior(turnCount, totalCost, result || '')) {
throw new PentestError(
`Spending cap likely reached (turns=${turnCount}, cost=$0): ${result?.slice(0, 100)}`,
'billing',
true,
);
}
const duration = timer.stop();
progress.finish(formatCompletionMessage(execContext, description, turnCount, duration));
@@ -393,12 +390,10 @@ export async function runPiPrompt(
success: true,
duration,
turns: turnCount,
cost: usage.cost,
inputTokens: usage.inputTokens,
outputTokens: usage.outputTokens,
cacheReadTokens: usage.cacheReadTokens,
cacheWriteTokens: usage.cacheWriteTokens,
cost: totalCost,
model: selection.model.id,
partialCost: totalCost,
apiErrorDetected,
...(structuredOutput !== undefined && { structuredOutput }),
};
} catch (error) {
@@ -407,29 +402,17 @@ export async function runPiPrompt(
const err = error as Error & { code?: string; status?: number };
await auditLogger.logError(err, duration, turnCount);
progress.stop();
outputLines(formatErrorOutput(err, execContext, description, duration, sourceDir, isRetryableFailure(err)));
outputLines(formatErrorOutput(err, execContext, description, duration, sourceDir, isRetryableError(err)));
await writeErrorLog(err, sourceDir, fullPrompt, duration);
// A failed agent still spent money — on its own turns and, since Shannon's
// prompts delegate the heavy work, mostly on `task` sub-agents. Both count
// toward the run's usage.
const usage = totalUsage(session, childUsage);
return {
error: err.message,
errorType: err instanceof PentestError && err.code ? err.code : err.constructor.name,
errorType: err.constructor.name,
prompt: `${fullPrompt.slice(0, 100)}...`,
success: false,
duration,
turns: turnCount,
cost: usage.cost,
inputTokens: usage.inputTokens,
outputTokens: usage.outputTokens,
cacheReadTokens: usage.cacheReadTokens,
cacheWriteTokens: usage.cacheWriteTokens,
retryable: isRetryableFailure(err),
cost: 0,
retryable: isRetryableError(err),
};
} finally {
cancellationSignal?.removeEventListener('abort', onCancellation);
}
}
-28
View File
@@ -1,28 +0,0 @@
// Copyright (C) 2025 Keygraph, Inc.
//
// This program is free software: you can redistribute it and/or modify
// it under the terms of the GNU Affero General Public License version 3
// as published by the Free Software Foundation.
/**
* Retry split between the two layers that can restart work.
*
* `enabled: false` turns off pi's own agent-level retry loop — Temporal owns
* agent restarts, and both retrying the same turn would compound. `provider`
* settings are read independently of that flag, so transport faults
* (408/409/429/5xx) are still absorbed inside the session, which is far cheaper
* than a Temporal retry that re-runs the agent and respends its tokens.
*
* `maxRetries` is handed to the selected vendor's SDK, which owns the backoff, so
* the schedule varies by provider rather than following one formula.
*
* NOTE: pi recommends keeping this at 0, since SDK-level retries consume
* out-of-usage-limit responses before pi's classifier can mark them terminal.
* Shannon accepts that trade for the transport-fault coverage. `maxRetryDelayMs`
* is left at pi's 60s default so a server asking for a longer wait fails fast
* instead of parking the activity.
*/
export const PI_RETRY_SETTINGS = {
enabled: false,
provider: { maxRetries: 8 },
} as const;
+17 -20
View File
@@ -16,25 +16,28 @@
* resource loader, and a fixed child tool surface.
*/
import type { ThinkingLevel } from '@earendil-works/pi-agent-core';
import { type AssistantMessage, type Model, Type } from '@earendil-works/pi-ai';
import {
type AuthStorage,
createAgentSession,
defineTool,
getAgentDir,
type ModelRuntime,
type ModelRegistry,
type ResourceLoader,
SessionManager,
SettingsManager,
type ToolDefinition,
} from '@earendil-works/pi-coding-agent';
import { PI_RETRY_SETTINGS } from './retry-settings.js';
export interface TaskToolContext {
cwd: string;
// eslint-disable-next-line @typescript-eslint/no-explicit-any
model: Model<any>;
/** Parent's model/auth runtime, reused so sub-agents share its resolved credential. */
modelRuntime: ModelRuntime;
thinkingLevel?: ThinkingLevel;
authStorage: AuthStorage;
/** Explicit model registry for sub-session resolution. Omit to inherit the parent's default. */
modelRegistry?: ModelRegistry;
resourceLoader: ResourceLoader;
cancellationSignal?: AbortSignal | undefined;
/**
@@ -43,13 +46,7 @@ export interface TaskToolContext {
* so without this their spend (the bulk of a whitebox run, since Shannon
* prompts delegate the heavy work) is invisible to billing.
*/
onUsage?: (usage: {
cost: number;
inputTokens: number;
outputTokens: number;
cacheReadTokens: number;
cacheWriteTokens: number;
}) => void;
onUsage?: (usage: { cost: number; inputTokens: number; outputTokens: number }) => void;
}
const CHILD_TOOLS = ['read', 'grep', 'find', 'ls', 'write', 'bash'];
@@ -86,11 +83,13 @@ export function createTaskTool(config: TaskToolContext): ToolDefinition {
agentDir,
resourceLoader: config.resourceLoader,
model: config.model,
...(config.thinkingLevel && { thinkingLevel: config.thinkingLevel }),
tools: CHILD_TOOLS,
modelRuntime: config.modelRuntime,
authStorage: config.authStorage,
...(config.modelRegistry && { modelRegistry: config.modelRegistry }),
sessionManager: SessionManager.inMemory(config.cwd),
settingsManager: SettingsManager.inMemory({
retry: PI_RETRY_SETTINGS,
retry: { enabled: false },
compaction: { enabled: true },
}),
});
@@ -110,6 +109,8 @@ export function createTaskTool(config: TaskToolContext): ToolDefinition {
let resultText = '';
let subCost = 0;
let subInputTokens = 0;
let subOutputTokens = 0;
subSession.subscribe((event) => {
if (event.type === 'turn_end') {
const msg = event.message as AssistantMessage | undefined;
@@ -119,6 +120,8 @@ export function createTaskTool(config: TaskToolContext): ToolDefinition {
}
}
if (msg?.usage?.cost?.total != null) subCost += msg.usage.cost.total;
subInputTokens += msg?.usage?.input ?? 0;
subOutputTokens += msg?.usage?.output ?? 0;
}
});
@@ -135,13 +138,7 @@ export function createTaskTool(config: TaskToolContext): ToolDefinition {
// Read stats before dispose; reconcile cost the same way the parent does.
const subStats = subSession.getSessionStats();
if (subStats.cost > subCost) subCost = subStats.cost;
config.onUsage?.({
cost: subCost,
inputTokens: subStats.tokens.input,
outputTokens: subStats.tokens.output,
cacheReadTokens: subStats.tokens.cacheRead,
cacheWriteTokens: subStats.tokens.cacheWrite,
});
config.onUsage?.({ cost: subCost, inputTokens: subInputTokens, outputTokens: subOutputTokens });
} finally {
config.cancellationSignal?.removeEventListener('abort', onCancellation);
subSession.dispose();
-44
View File
@@ -1,44 +0,0 @@
// Copyright (C) 2025 Keygraph, Inc.
//
// This program is free software: you can redistribute it and/or modify
// it under the terms of the GNU Affero General Public License version 3
// as published by the Free Software Foundation.
import { type AssistantMessage, isContextOverflow, isRetryableAssistantError } from '@earendil-works/pi-ai';
import { PentestError } from '../../services/error-handling.js';
import { ErrorCode } from '../../types/errors.js';
/**
* Wrap a failed assistant turn, taking the verdict from pi.
*
* Overflow is separated first, as pi's retry contract requires: it means the
* request was too large, not that the provider faltered, so an identical retry
* would overflow again. Everything else goes to pi's classifier, which treats
* quota, billing, and auth exhaustion as terminal and load, throttling, and
* transport faults as transient — those were already retried in-session, so
* reaching here means the attempts were exhausted.
*
* `contextWindow` is omitted where overflow cannot apply, such as a one-word
* credential probe.
*/
export function providerTurnError(message: AssistantMessage, label: string, contextWindow?: number): PentestError {
const detail = (message.errorMessage ?? 'unknown provider error').slice(0, 300);
if (contextWindow !== undefined && isContextOverflow(message, contextWindow)) {
return new PentestError(
`${label}: context window exceeded after compaction: ${detail}`,
'unknown',
false,
{ contextWindow },
ErrorCode.AGENT_EXECUTION_FAILED,
);
}
return new PentestError(
`${label}: ${detail}`,
'unknown',
isRetryableAssistantError(message),
{},
ErrorCode.AGENT_EXECUTION_FAILED,
);
}
+2 -30
View File
@@ -14,7 +14,6 @@
import { defineTool } from '@earendil-works/pi-coding-agent';
import { type Static, type TObject, Type } from 'typebox';
import { stringEnum } from '../collectors/schema.js';
import type { AgentName } from '../types/agents.js';
import type { CapturedSubmitTool } from './submit-tool.js';
@@ -24,38 +23,13 @@ function optStr(description?: string) {
return Type.Optional(Type.String(description === undefined ? {} : { description }));
}
/**
* Base fields shared by every queue entry. `notes` gains guidance in analysis mode.
*
* `confidence` is enumerated so it reaches the report agent in the same casing the report
* schema accepts — an analysis-only run carries it through verbatim as its only rating.
*/
/** Base fields shared by every queue entry. `notes` gains guidance in analysis mode. */
function baseFields(exploit: boolean) {
return {
ID: Type.String(),
vulnerability_type: Type.String(),
externally_exploitable: Type.Boolean(),
confidence: stringEnum(['high', 'medium', 'low'], {
description: 'Confidence that this is a real, reachable vulnerability.',
}),
code_locations: Type.Optional(
Type.Array(
Type.Object({
file: Type.String({ description: 'Repository-relative path, no leading slash.' }),
start_line: Type.Optional(Type.Integer({ minimum: 1 })),
end_line: Type.Optional(Type.Integer({ minimum: 1, description: 'Set when the flaw spans a range.' })),
role: stringEnum(['sink', 'source', 'guard'], {
description:
'sink where the flaw manifests, source where untrusted input enters, guard for a check ' +
'that is missing or misplaced.',
}),
symbol: Type.Optional(
Type.String({ description: 'Enclosing function or method, named as written in the code.' }),
),
}),
{ description: 'Every code site this finding touches, sink first.' },
),
),
confidence: Type.String(),
notes: exploit ? optStr() : optStr(ANALYSIS_NOTES_DESCRIPTION),
};
}
@@ -120,8 +94,6 @@ const authEntry = () => Type.Object({ ...baseFields(true), ...authFields });
const ssrfEntry = () => Type.Object({ ...baseFields(true), ...ssrfFields });
const authzEntry = () => Type.Object({ ...baseFields(true), ...authzFields });
export type QueueCodeLocation = NonNullable<Static<ReturnType<typeof injectionEntry>>['code_locations']>[number];
export type InjectionFinding = Static<ReturnType<typeof injectionEntry>>;
export type XssFinding = Static<ReturnType<typeof xssEntry>>;
export type AuthFinding = Static<ReturnType<typeof authEntry>>;
+1 -23
View File
@@ -23,11 +23,6 @@ interface AttemptData {
attempt_number: number;
duration_ms: number;
cost_usd: number;
input_tokens?: number | undefined;
output_tokens?: number | undefined;
cache_read_tokens?: number | undefined;
cache_write_tokens?: number | undefined;
turns?: number | undefined;
success: boolean;
timestamp: string;
model?: string | undefined;
@@ -39,10 +34,6 @@ interface AgentAuditMetrics {
attempts: AttemptData[];
final_duration_ms: number;
total_cost_usd: number;
total_input_tokens: number;
total_output_tokens: number;
total_cache_read_tokens: number;
total_cache_write_tokens: number;
model?: string | undefined;
checkpoint?: string | undefined;
}
@@ -183,10 +174,6 @@ export class MetricsTracker {
attempts: [],
final_duration_ms: 0,
total_cost_usd: 0,
total_input_tokens: 0,
total_output_tokens: 0,
total_cache_read_tokens: 0,
total_cache_write_tokens: 0,
};
this.data.metrics.agents[agentName] = agent;
@@ -197,11 +184,6 @@ export class MetricsTracker {
cost_usd: result.cost_usd,
success: result.success,
timestamp: formatTimestamp(),
...(result.input_tokens !== undefined && { input_tokens: result.input_tokens }),
...(result.output_tokens !== undefined && { output_tokens: result.output_tokens }),
...(result.cache_read_tokens !== undefined && { cache_read_tokens: result.cache_read_tokens }),
...(result.cache_write_tokens !== undefined && { cache_write_tokens: result.cache_write_tokens }),
...(result.turns !== undefined && { turns: result.turns }),
};
if (result.model) {
@@ -215,12 +197,8 @@ export class MetricsTracker {
// 3. Append attempt to history
agent.attempts.push(attempt);
// 4. Recalculate totals across all attempts (includes failures)
// 4. Recalculate total cost across all attempts (includes failures)
agent.total_cost_usd = agent.attempts.reduce((sum, a) => sum + a.cost_usd, 0);
agent.total_input_tokens = agent.attempts.reduce((sum, a) => sum + (a.input_tokens ?? 0), 0);
agent.total_output_tokens = agent.attempts.reduce((sum, a) => sum + (a.output_tokens ?? 0), 0);
agent.total_cache_read_tokens = agent.attempts.reduce((sum, a) => sum + (a.cache_read_tokens ?? 0), 0);
agent.total_cache_write_tokens = agent.attempts.reduce((sum, a) => sum + (a.cache_write_tokens ?? 0), 0);
// 5. Update agent status based on outcome
if (result.success) {
+14
View File
@@ -12,6 +12,7 @@
*/
import fs from 'node:fs/promises';
import { isFableModel, resolveModelId } from '../ai/models.js';
import { formatDuration, formatTimestamp } from '../utils/formatting.js';
import { LogStream } from './log-stream.js';
import { generateWorkflowLogPath, type SessionMetadata } from './utils.js';
@@ -86,6 +87,19 @@ export class WorkflowLogger {
`Started: ${formatTimestamp()}`,
];
// Surface Fable usage: its safety classifiers route cybersecurity tasks to
// Opus 4.8, so those phases run on Opus 4.8 regardless of the tier setting.
const fableTiers = (['small', 'medium', 'large'] as const)
.map((tier) => ({ tier, model: resolveModelId(tier) }))
.filter(({ model }) => isFableModel(model));
if (fableTiers.length > 0) {
const tierList = fableTiers.map(({ tier, model }) => `${tier} (${model})`).join(', ');
lines.push(
`Note: ${tierList} set to a Fable model. Fable's safety classifiers`,
` route cybersecurity tasks to Opus 4.8, so those phases run on Opus 4.8.`,
);
}
lines.push(`================================================================================`, ``);
return this.logStream.write(lines.join('\n'));
@@ -122,7 +122,8 @@ export function buildSchemas(validIds: ReadonlySet<string>) {
const vulnerableLocationField = Type.String({
minLength: 1,
description:
'Endpoint or mechanism where the vulnerability exists (e.g. "GET /api/products?id=", ' + '"POST /login").',
'Endpoint or mechanism where the vulnerability exists (e.g. "GET /api/products?id=", ' +
'"POST /login", or a code location like "controllers/userController.js:42").',
});
const overviewField = Type.String({
@@ -1,336 +0,0 @@
// Copyright (C) 2025 Keygraph, Inc.
//
// This program is free software: you can redistribute it and/or modify
// it under the terms of the GNU Affero General Public License version 3
// as published by the Free Software Foundation.
/**
* Finding Collector tools
*
* Collects structured findings from the report agent via a pi tool. The agent
* calls `add_finding` once per finding with TypeBox-validated parameters. After
* the agent finishes, the caller retrieves collected findings via `getAll()`
* for downstream rendering (markdown, PDF, DB).
*
* The tool schema is mode-dependent: fields describing a demonstrated attack have no source in
* an analysis-only run, and offering them would only make the agent invent them.
*/
import { defineTool, type ToolDefinition } from '@earendil-works/pi-coding-agent';
import { type Static, Type } from 'typebox';
import { cleanInput, stringEnum } from './schema.js';
// ============================================================================
// SCHEMA
// ============================================================================
const OWASP_CATEGORY_VALUES = [
'A01:2025 — Broken Access Control',
'A02:2025 — Security Misconfiguration',
'A03:2025 — Software Supply Chain Failures',
'A04:2025 — Cryptographic Failures',
'A05:2025 — Injection',
'A06:2025 — Insecure Design',
'A07:2025 — Authentication Failures',
'A08:2025 — Software or Data Integrity Failures',
'A09:2025 — Security Logging and Alerting Failures',
'A10:2025 — Mishandling of Exceptional Conditions',
] as const;
const SEVERITY_VALUES = ['critical', 'high', 'medium', 'low'] as const;
const STATUS_VALUES = ['exploited', 'out_of_scope', 'blocked_by_constraints', 'false_positive'] as const;
const CONFIDENCE_VALUES = ['high', 'medium', 'low'] as const;
const StepItemSchema = Type.Union([
Type.Object({
kind: Type.Literal('prose'),
text: Type.String({ minLength: 1, description: 'Narrative prose for this item.' }),
}),
Type.Object({
kind: Type.Literal('code'),
block: Type.Object({
language: Type.String({
description: 'Language identifier for syntax highlighting (e.g., "bash", "http", "json").',
}),
content: Type.String({ minLength: 1, description: 'The code content.' }),
}),
}),
]);
const StructuredStepSchema = Type.Object({
title: Type.Optional(
Type.Union([Type.String(), Type.Null()], {
description: 'Optional title for this step (e.g., "Send malicious payload").',
}),
),
items: Type.Array(StepItemSchema, {
minItems: 1,
description: 'Ordered list of prose and code items that make up this step.',
}),
});
const CodeLocationSchema = Type.Object({
file: Type.String({
minLength: 1,
description: 'Repository-relative path, no leading slash (e.g., "routes/search.ts").',
}),
start_line: Type.Optional(
Type.Union([Type.Integer({ minimum: 1 }), Type.Null()], {
description: '1-indexed line number. Omit when the deliverable gives only a file.',
}),
),
end_line: Type.Optional(
Type.Union([Type.Integer({ minimum: 1 }), Type.Null()], {
description: 'End of the range, when the finding spans multiple lines.',
}),
),
role: stringEnum(['sink', 'source', 'guard'], {
description:
'What this location is in the data flow. `sink` is where the vulnerability manifests, `source` ' +
'where untrusted input enters, `guard` a check that is missing or misplaced.',
}),
symbol: Type.Optional(
Type.Union([Type.String(), Type.Null()], {
description: 'Enclosing function or method name, when known.',
}),
),
});
const HttpLocationSchema = Type.Object({
method: Type.String({ minLength: 1, description: 'HTTP method (e.g., "GET", "POST").' }),
url: Type.String({ minLength: 1, description: 'Full URL of the affected endpoint.' }),
parameter: Type.Optional(
Type.Union([Type.String(), Type.Null()], {
description: 'The specific parameter carrying the payload, when the finding names one.',
}),
),
});
const AdditionalSectionSchema = Type.Object({
heading: Type.String({
minLength: 1,
description: 'Section heading (e.g., "Real-World Attack Scenario").',
}),
items: Type.Array(StepItemSchema, {
minItems: 1,
description: 'Ordered prose and code items for this section.',
}),
});
/**
* `severity` is recorded in both modes, but it does not mean the same thing in each: an exploit
* run measures it from what the exploit demonstrated, an analysis run assesses it from the class
* of flaw. The description says which, so the agent never presents an assessment as a measurement.
*/
function identityFields(exploit: boolean) {
const severityDescription = exploit
? 'Severity of the finding, based on the impact the exploit demonstrated.'
: 'Severity of the finding, assessed from the vulnerability class and the impact it would have.';
return {
severity: stringEnum(SEVERITY_VALUES, { description: severityDescription }),
finding_id: Type.String({
minLength: 1,
description: 'Finding identifier (e.g., "AUTH-VULN-07", "INJ-VULN-03"). Must be unique per report.',
}),
title: Type.String({
minLength: 1,
description:
'Descriptive name (e.g., "SQL Injection — User Search", "IDOR — Unauthorized Access to User Orders").',
}),
category: stringEnum(['Injection', 'XSS', 'Authentication', 'Authorization', 'SSRF'], {
description:
'From the finding_id prefix: INJ-VULN-xxx Injection, ' +
'XSS-VULN-xxx XSS, AUTH-VULN-xxx Authentication, AUTHZ-VULN-xxx Authorization, ' +
'SSRF-VULN-xxx SSRF.',
}),
owasp_category: stringEnum(OWASP_CATEGORY_VALUES, {
description: 'OWASP Top Ten 2025 category.',
}),
};
}
function locationFields() {
return {
vulnerable_location: Type.String({
minLength: 1,
description: 'Endpoint or code location where the vulnerability exists.',
}),
http_location: Type.Optional(
Type.Union([HttpLocationSchema, Type.Null()], {
description:
'The HTTP request the finding is reached through, when the deliverable names one. Omit for ' +
'findings with no network entry point.',
}),
),
};
}
/** `impact` is described per mode: an analysis run demonstrated nothing, and implying otherwise invites fabrication. */
function narrativeFields(exploit: boolean) {
const impactDescription = exploit
? 'What the exploit demonstrably achieved.'
: 'What an attacker could achieve if this were exploited. State it as assessed, not demonstrated.';
return {
overview: Type.String({
minLength: 1,
description: 'What the vulnerability is and why it matters. 2-3 sentences of professional prose.',
}),
impact: Type.String({ minLength: 1, description: impactDescription }),
remediation: Type.String({
minLength: 1,
description: 'Specific, actionable fix guidance. Code-level or configuration-level.',
}),
};
}
/** Fields that only mean something once an exploit has run. Absent from the analysis schema. */
function exploitOnlyFields() {
return {
auth_state: Type.String({
minLength: 1,
description: 'Authentication state during testing (e.g., "Unauthenticated", "Any authenticated user").',
}),
prerequisites: Type.String({
minLength: 1,
description: 'What is needed to exploit the vulnerability (or "None").',
}),
exploitation_steps: Type.Array(StructuredStepSchema, {
minItems: 1,
description: 'Ordered exploitation steps. Each step has an optional title and prose/code items.',
}),
proof_of_impact: Type.Array(StepItemSchema, {
minItems: 1,
description: 'Evidence of what the exploit achieved — prose and code items.',
}),
status: Type.Optional(
Type.Union([stringEnum(STATUS_VALUES), Type.Null()], {
description: 'Finding status. Use "exploited" for confirmed exploits.',
}),
),
};
}
/** Accompanies `severity` when nothing was exploited — the rating the analysis deliverable itself carries. */
function analysisOnlyFields() {
return {
confidence: stringEnum(CONFIDENCE_VALUES, {
description:
'Confidence that this is a real, reachable vulnerability. Carry it over from the analysis ' +
'deliverable rather than reassessing.',
}),
};
}
function sharedOptionalFields() {
return {
notes: Type.Optional(
Type.Union([Type.Array(StepItemSchema), Type.Null()], {
description: 'Additional context as prose/code items.',
}),
),
additional_sections: Type.Optional(
Type.Union([Type.Array(AdditionalSectionSchema), Type.Null()], {
description: 'Extra report sections that do not fit into other fields (e.g., "Real-World Attack Scenario").',
}),
),
};
}
export function buildAddFindingSchema(exploit: boolean) {
return Type.Object({
...identityFields(exploit),
...(exploit ? exploitOnlyFields() : analysisOnlyFields()),
...locationFields(),
...narrativeFields(exploit),
...sharedOptionalFields(),
});
}
/**
* Superset of both modes, for typing only. Consumers must check presence rather than assume:
* `report.json` from an analysis run has no `exploitation_steps` key at all. `severity` is the
* exception — both modes record it, so it is required here too.
*/
const AddFindingSupersetSchema = Type.Object({
...identityFields(true),
code_locations: Type.Optional(Type.Array(CodeLocationSchema)),
auth_state: Type.Optional(Type.String()),
prerequisites: Type.Optional(Type.String()),
exploitation_steps: Type.Optional(Type.Array(StructuredStepSchema)),
proof_of_impact: Type.Optional(Type.Array(StepItemSchema)),
status: Type.Optional(Type.Union([stringEnum(STATUS_VALUES), Type.Null()])),
confidence: Type.Optional(Type.Union([stringEnum(CONFIDENCE_VALUES), Type.Null()])),
...locationFields(),
...narrativeFields(true),
...sharedOptionalFields(),
});
export type AddFindingInput = Static<typeof AddFindingSupersetSchema>;
// Re-export schema types for downstream consumers
export type CodeLocation = Static<typeof CodeLocationSchema>;
export type HttpLocation = Static<typeof HttpLocationSchema>;
export type StepItem = Static<typeof StepItemSchema>;
export type StructuredStep = Static<typeof StructuredStepSchema>;
export type AdditionalSection = Static<typeof AdditionalSectionSchema>;
// ============================================================================
// RESPONSE HELPERS
// ============================================================================
function toolResult(payload: Record<string, unknown>) {
return {
content: [{ type: 'text' as const, text: JSON.stringify(payload, null, 2) }],
details: undefined,
};
}
function successResult(data: Record<string, unknown>) {
return toolResult({ status: 'success', ...data });
}
function errorResult(message: string, errorType = 'ValidationError', retryable = true) {
return toolResult({ status: 'error', message, errorType, retryable });
}
// ============================================================================
// COLLECTOR FACTORY
// ============================================================================
export interface FindingCollector {
tools: ToolDefinition[];
getAll(): AddFindingInput[];
}
export function createFindingCollector(exploit: boolean): FindingCollector {
const findings: AddFindingInput[] = [];
const schema = buildAddFindingSchema(exploit);
const addFindingTool = defineTool({
name: 'add_finding',
label: 'Add Finding',
description:
'Record a single finding as structured data for report rendering and DB persistence. Call once per finding after grouping/dedup. Duplicate finding_ids are rejected.',
parameters: schema,
async execute(_toolCallId, input) {
const existing = findings.find((f) => f.finding_id === input.finding_id);
if (existing) {
return errorResult(
`Finding ${input.finding_id} has already been recorded. Each finding may only be added once.`,
'DuplicateError',
false,
);
}
const typed = cleanInput(schema, input) as AddFindingInput;
findings.push(typed);
return successResult({ added: [typed.finding_id] });
},
});
return {
tools: [addFindingTool],
getAll: (): AddFindingInput[] => [...findings],
};
}
+5 -12
View File
@@ -514,7 +514,7 @@ const validateRulesSecurity = (rules: Rule[] | undefined, ruleType: string): voi
ErrorCode.CONFIG_VALIDATION_FAILED,
);
}
if (rule.description !== undefined && pattern.test(rule.description)) {
if (pattern.test(rule.description)) {
throw new PentestError(
`rules.${ruleType}[${index}].description contains potentially dangerous pattern: ${pattern.source}`,
'config',
@@ -656,15 +656,11 @@ const checkForConflicts = (avoidRules: Rule[] = [], focusRules: Rule[] = []): vo
};
const sanitizeRule = (rule: Rule): Rule => {
const sanitized: Rule = {
return {
description: rule.description.trim(),
type: rule.type.toLowerCase().trim() as Rule['type'],
value: rule.value.trim(),
};
const description = rule.description?.trim();
if (description) {
sanitized.description = description;
}
return sanitized;
};
export const distributeConfig = (config: Config | null): DistributedConfig => {
@@ -679,7 +675,6 @@ export const distributeConfig = (config: Config | null): DistributedConfig => {
const exploit = config?.exploit !== undefined ? config.exploit === 'true' : true;
const report = {
sarif: config?.report?.sarif === 'true',
...(config?.report?.min_severity && { min_severity: config.report.min_severity }),
...(config?.report?.min_confidence && { min_confidence: config.report.min_confidence }),
...(config?.report?.guidance && { guidance: config.report.guidance.trim() }),
@@ -706,15 +701,13 @@ const sanitizeAuthentication = (auth: Authentication): Authentication => {
credentials: {
username: auth.credentials.username.trim(),
...(auth.credentials.password && { password: auth.credentials.password }),
...(auth.credentials.totp_secret && {
totp_secret: auth.credentials.totp_secret.replace(/\s/g, ''),
}),
...(auth.credentials.totp_secret && { totp_secret: auth.credentials.totp_secret.trim() }),
...(auth.credentials.email_login && {
email_login: {
address: auth.credentials.email_login.address.trim(),
password: auth.credentials.email_login.password,
...(auth.credentials.email_login.totp_secret && {
totp_secret: auth.credentials.email_login.totp_secret.replace(/\s/g, ''),
totp_secret: auth.credentials.email_login.totp_secret.trim(),
}),
},
}),
+2 -17
View File
@@ -9,9 +9,6 @@ const WORKER_ROOT = path.resolve(import.meta.dirname, '..');
export const PROMPTS_DIR = path.join(WORKER_ROOT, 'prompts');
export const CONFIGS_DIR = path.join(WORKER_ROOT, 'configs');
/** Bundled Typst template that renders report.json into the PDF report. */
export const TYPST_TEMPLATE = path.join(WORKER_ROOT, 'templates', 'typst', 'report.typ');
/** Compiled pi extension dir that enforces bounded `bash` timeouts (resolved from dist/) */
export const BASH_TIMEOUT_EXTENSION_DIR = path.join(import.meta.dirname, 'ai', 'extensions', 'bash-timeout');
@@ -31,20 +28,8 @@ export const INTERNAL_DIR = '.shannon';
/** Filename of the assembled report inside the deliverables dir (internal, source of the surfaced copy) */
export const ASSEMBLED_REPORT_FILENAME = 'comprehensive_security_assessment_report.md';
/** Filename of the compiled PDF report inside the deliverables dir (internal, source of the surfaced copy) */
export const ASSEMBLED_REPORT_PDF_FILENAME = 'comprehensive_security_assessment_report.pdf';
/** Filename of the human-facing PDF report surfaced at the run directory root */
export const FINAL_REPORT_PDF_FILENAME = 'Security-Assessment-Report.pdf';
/** Filename of the human-facing markdown report surfaced at the run directory root, alongside the PDF */
export const FINAL_REPORT_MD_FILENAME = 'Security-Assessment-Report.md';
/** Structured findings the report agent emits; the markdown report is rendered from it. */
export const REPORT_JSON_FILENAME = 'report.json';
/** SARIF 2.1.0 log, written only for exploit=true runs when report.sarif is enabled. */
export const SARIF_FILENAME = 'report.sarif';
/** Filename of the human-facing final report surfaced at the run directory root */
export const FINAL_REPORT_FILENAME = 'Security-Assessment-Report.md';
/**
* Resolve the session.json path for a run directory, preferring the current
-139
View File
@@ -1,139 +0,0 @@
#!/usr/bin/env node
// Copyright (C) 2025 Keygraph, Inc.
//
// This program is free software: you can redistribute it and/or modify
// it under the terms of the GNU Affero General Public License version 3
// as published by the Free Software Foundation.
/**
* set-report-meta CLI
*
* Writes top-level report metadata to report.json.
* Called once by the report agent before recording individual findings.
* Overwrites any existing report_meta — idempotent.
*
* Usage:
* set-report-meta --target "https://example.com" --assessment-date "2026-05-07" \
* --scope "injection, xss, auth, authz, ssrf" --executive-summary "..."
*
* Output (JSON to stdout):
* { "status": "success" }
* { "status": "error", "message": "...", "retryable": true }
*/
import { existsSync, mkdirSync, readFileSync, renameSync, unlinkSync, writeFileSync } from 'node:fs';
import { resolve } from 'node:path';
const REPORT_FILENAME = 'report.json';
interface ReportMeta {
target: string;
assessment_date: string;
scope: string;
executive_summary: string;
}
interface ReportFile {
report_meta?: ReportMeta;
findings: Array<Record<string, unknown>>;
}
const HELP = `set-report-meta — write top-level report metadata to report.json
Usage:
set-report-meta --target "https://example.com" --assessment-date "2026-05-07" \\
--scope "injection, xss, auth" --executive-summary "..."
Required flags: --target, --assessment-date, --scope, --executive-summary
Output: JSON to stdout with status "success" or "error".`;
function getFlag(argv: string[], flag: string): string | undefined {
for (let i = 2; i < argv.length; i++) {
if (argv[i] === flag && argv[i + 1] && !argv[i + 1]!.startsWith('--')) {
return argv[i + 1]!;
}
}
return undefined;
}
function readReportFile(filePath: string): ReportFile {
if (!existsSync(filePath)) {
return { findings: [] };
}
const raw = readFileSync(filePath, 'utf-8');
return JSON.parse(raw) as ReportFile;
}
function writeReportFile(filePath: string, data: ReportFile): void {
const tmpPath = `${filePath}.tmp`;
const payload = JSON.stringify(data, null, 2);
try {
writeFileSync(tmpPath, payload, 'utf-8');
renameSync(tmpPath, filePath);
} catch (err) {
try {
unlinkSync(tmpPath);
} catch {
/* best-effort */
}
throw err;
}
}
function main(): void {
if (process.argv[2] === '--help' || process.argv[2] === '-h') {
console.log(HELP);
return;
}
const target = getFlag(process.argv, '--target');
const assessmentDate = getFlag(process.argv, '--assessment-date');
const scope = getFlag(process.argv, '--scope');
const executiveSummary = getFlag(process.argv, '--executive-summary');
if (!target) {
console.log(JSON.stringify({ status: 'error', message: 'Missing required --target flag', retryable: true }));
process.exit(1);
}
if (!assessmentDate) {
console.log(
JSON.stringify({ status: 'error', message: 'Missing required --assessment-date flag', retryable: true }),
);
process.exit(1);
}
if (!scope) {
console.log(JSON.stringify({ status: 'error', message: 'Missing required --scope flag', retryable: true }));
process.exit(1);
}
if (!executiveSummary) {
console.log(
JSON.stringify({ status: 'error', message: 'Missing required --executive-summary flag', retryable: true }),
);
process.exit(1);
}
const subdir = process.env.SHANNON_DELIVERABLES_SUBDIR || '.shannon/deliverables';
const deliverablesDir = resolve(process.cwd(), ...subdir.split('/'));
mkdirSync(deliverablesDir, { recursive: true });
const filePath = resolve(deliverablesDir, REPORT_FILENAME);
const data = readReportFile(filePath);
data.report_meta = {
target,
assessment_date: assessmentDate,
scope,
executive_summary: executiveSummary,
};
writeReportFile(filePath, data);
console.log(JSON.stringify({ status: 'success' }));
}
try {
main();
} catch (error) {
const message = error instanceof Error ? error.message : String(error);
console.log(JSON.stringify({ status: 'error', message, retryable: true }));
process.exit(1);
}
+55 -63
View File
@@ -13,6 +13,7 @@
* - Create git checkpoint
* - Start audit logging
* - Invoke the pi agent via runPiPrompt
* - Spending cap check using isSpendingCapBehavior
* - Handle failure (rollback, audit)
* - Validate output using AGENTS[agentName].deliverableFilename
* - Render the deliverable to disk via the writeDeliverable hook (if provided)
@@ -33,6 +34,7 @@ import type { AgentEndResult } from '../types/audit.js';
import { ErrorCode, type PentestErrorType } from '../types/errors.js';
import type { AgentMetrics } from '../types/metrics.js';
import { err, isErr, ok, type Result } from '../types/result.js';
import { isSpendingCapBehavior } from '../utils/billing-detection.js';
import { getAgentGitPaths } from './agent-git-paths.js';
import type { ConfigLoaderService } from './config-loader.js';
import { PentestError } from './error-handling.js';
@@ -53,7 +55,6 @@ export interface AgentExecutionInput {
attemptNumber: number;
promptDir?: string | undefined;
customTools?: import('@earendil-works/pi-coding-agent').ToolDefinition[];
failedClasses?: readonly import('../types/config.js').VulnClass[] | undefined;
// Renders the deliverable to disk; invoked after validation, before the success commit.
writeDeliverable?: (deliverablesPath: string) => Promise<void>;
cancellationSignal?: AbortSignal | undefined;
@@ -79,6 +80,11 @@ function errorCodeFromResult(result: PiPromptResult): ErrorCode {
function categoryForErrorCode(code: ErrorCode): PentestErrorType {
switch (code) {
case ErrorCode.SPENDING_CAP_REACHED:
case ErrorCode.INSUFFICIENT_CREDITS:
case ErrorCode.BILLING_ERROR:
case ErrorCode.API_RATE_LIMITED:
return 'billing';
case ErrorCode.GIT_CHECKPOINT_FAILED:
case ErrorCode.GIT_ROLLBACK_FAILED:
return 'filesystem';
@@ -147,7 +153,6 @@ export class AgentExecutionService {
attemptNumber,
promptDir,
customTools,
failedClasses,
writeDeliverable,
cancellationSignal,
} = input;
@@ -166,12 +171,7 @@ export class AgentExecutionService {
try {
prompt = await loadPrompt(
promptTemplate,
{
webUrl,
repoPath,
AUTH_STATE_FILE: authStateFile(auditSession.sessionMetadata),
...(failedClasses !== undefined && { failedClasses }),
},
{ webUrl, repoPath, AUTH_STATE_FILE: authStateFile(auditSession.sessionMetadata) },
distributedConfig,
pipelineTestingMode,
logger,
@@ -227,13 +227,31 @@ export class AgentExecutionService {
agentName,
auditSession,
logger,
AGENTS[agentName].modelTier,
customTools,
path.relative(repoPath, deliverablesPath),
cancellationSignal,
submitTool,
);
// 6. Handle execution failure
// 6. Spending cap check - defense-in-depth
if (result.success && (result.turns ?? 0) <= 2 && (result.cost || 0) === 0) {
const resultText = result.result || '';
if (isSpendingCapBehavior(result.turns ?? 0, result.cost || 0, resultText)) {
return this.failAgent(agentName, deliverablesPath, auditSession, logger, {
attemptNumber,
result,
rollbackReason: 'spending cap detected',
errorMessage: `Spending cap likely reached: ${resultText.slice(0, 100)}`,
errorCode: ErrorCode.SPENDING_CAP_REACHED,
category: 'billing',
retryable: true,
context: { agentName, turns: result.turns, cost: result.cost },
});
}
}
// 7. Handle execution failure
if (!result.success) {
const errorCode = errorCodeFromResult(result);
return this.failAgent(agentName, deliverablesPath, auditSession, logger, {
@@ -252,53 +270,39 @@ export class AgentExecutionService {
// the write→validate→commit sequence is atomic against concurrent sibling agents.
let commitHash: string | undefined;
const finalizationError = await withGitRepoLock(async (): Promise<PentestError | null> => {
// Every step below must surface as a returned error rather than a throw: only the
// returned path rolls the workspace back and records the failed attempt.
try {
// 8. Write structured output to disk (vuln agents only) from the executor's capture
const queueFilename = getQueueFilename(agentName);
if (submitTool && queueFilename && result.structuredOutput !== undefined) {
await fs.ensureDir(deliverablesPath);
const queuePath = path.join(deliverablesPath, queueFilename);
await fs.writeFile(queuePath, JSON.stringify(result.structuredOutput, null, 2), 'utf8');
logger.info(`Wrote structured output queue to ${queueFilename}`);
}
// 8. Write structured output to disk (vuln agents only) from the executor's capture
const queueFilename = getQueueFilename(agentName);
if (submitTool && queueFilename && result.structuredOutput !== undefined) {
await fs.ensureDir(deliverablesPath);
const queuePath = path.join(deliverablesPath, queueFilename);
await fs.writeFile(queuePath, JSON.stringify(result.structuredOutput, null, 2), 'utf8');
logger.info(`Wrote structured output queue to ${queueFilename}`);
}
// 9. Validate output
const validationPassed = await validateAgentOutput(result, agentName, deliverablesPath, logger);
if (!validationPassed) {
return new PentestError(
`Agent ${agentName} failed output validation`,
'validation',
true,
{ agentName, deliverableFilename: AGENTS[agentName].deliverableFilename },
ErrorCode.OUTPUT_VALIDATION_FAILED,
);
}
// 10. Render the deliverable to disk so the success commit below stages it
if (writeDeliverable) {
await writeDeliverable(deliverablesPath);
}
// 11. Success - commit deliverables (scoped) and capture the checkpoint hash
const commitResult = await commitGitSuccess(deliverablesPath, agentName, logger, gitPaths);
if (!commitResult.success) {
return gitFailureForAgent(agentName, 'commit successful results', commitResult.error);
}
commitHash = commitResult.commitHash;
return null;
} catch (error) {
if (error instanceof PentestError) return error;
const errorMessage = error instanceof Error ? error.message : String(error);
// 9. Validate output
const validationPassed = await validateAgentOutput(result, agentName, deliverablesPath, logger);
if (!validationPassed) {
return new PentestError(
`Agent ${agentName} post-processing failed: ${errorMessage}`,
`Agent ${agentName} failed output validation`,
'validation',
true,
{ agentName, originalError: errorMessage },
{ agentName, deliverableFilename: AGENTS[agentName].deliverableFilename },
ErrorCode.OUTPUT_VALIDATION_FAILED,
);
}
// 10. Render the deliverable to disk so the success commit below stages it
if (writeDeliverable) {
await writeDeliverable(deliverablesPath);
}
// 11. Success - commit deliverables (scoped) and capture the checkpoint hash
const commitResult = await commitGitSuccess(deliverablesPath, agentName, logger, gitPaths);
if (!commitResult.success) {
return gitFailureForAgent(agentName, 'commit successful results', commitResult.error);
}
commitHash = commitResult.commitHash;
return null;
});
if (finalizationError) {
@@ -322,11 +326,6 @@ export class AgentExecutionService {
attemptNumber,
duration_ms: result.duration,
cost_usd: result.cost || 0,
input_tokens: result.inputTokens,
output_tokens: result.outputTokens,
cache_read_tokens: result.cacheReadTokens,
cache_write_tokens: result.cacheWriteTokens,
turns: result.turns,
success: true,
model: result.model,
...(commitHash && { checkpoint: commitHash }),
@@ -354,11 +353,6 @@ export class AgentExecutionService {
attemptNumber: opts.attemptNumber,
duration_ms: opts.result.duration,
cost_usd: opts.result.cost || 0,
input_tokens: opts.result.inputTokens,
output_tokens: opts.result.outputTokens,
cache_read_tokens: opts.result.cacheReadTokens,
cache_write_tokens: opts.result.cacheWriteTokens,
turns: opts.result.turns,
success: false,
model: opts.result.model,
error: opts.errorMessage,
@@ -412,10 +406,8 @@ export class AgentExecutionService {
static toMetrics(endResult: AgentEndResult, result: PiPromptResult): AgentMetrics {
return {
durationMs: endResult.duration_ms,
inputTokens: result.inputTokens ?? null,
outputTokens: result.outputTokens ?? null,
cacheReadTokens: result.cacheReadTokens ?? null,
cacheWriteTokens: result.cacheWriteTokens ?? null,
inputTokens: null, // Not currently exposed by the pi executor
outputTokens: null,
costUsd: endResult.cost_usd,
numTurns: result.turns ?? null,
model: result.model,
@@ -14,7 +14,6 @@
*/
import { getQueueFilename } from '../ai/queue-schemas.js';
import { REPORT_JSON_FILENAME, SARIF_FILENAME } from '../paths.js';
import { AGENTS } from '../session-manager.js';
import type { AgentName } from '../types/agents.js';
@@ -28,12 +27,5 @@ export function getAgentGitPaths(agentName: AgentName): string[] {
if (queueFilename) {
paths.push(queueFilename);
}
// The report agent also emits the structured findings the markdown is rendered from, and the
// SARIF log when enabled. Listing the log unconditionally is harmless when it was not written,
// and keeps a stale one from surviving the rollback of a failed attempt.
if (agentName === 'report') {
paths.push(REPORT_JSON_FILENAME);
paths.push(SARIF_FILENAME);
}
return [...new Set(paths)];
}
@@ -1,78 +0,0 @@
// Copyright (C) 2025 Keygraph, Inc.
//
// This program is free software: you can redistribute it and/or modify
// it under the terms of the GNU Affero General Public License version 3
// as published by the Free Software Foundation.
/**
* Attach vuln-queue code locations to collected findings.
*
* The vuln agent authors `code_locations` once, into its queue. Every stage after that used to
* re-transcribe them — the exploit agent into its evidence, the report agent into `add_finding` —
* and each hop lost some: 100% in the queue, 98% in the evidence, 42-63% by the report. Nothing
* about the copy is a judgement call, and `finding_id` matches the queue `ID` exactly, so the
* locations are joined here instead of being asked for again.
*/
import { fs, path } from 'zx';
import type { QueueCodeLocation } from '../ai/queue-schemas.js';
import type { AddFindingInput } from '../collectors/finding-collector.js';
import type { ActivityLogger } from '../types/activity-logger.js';
import { ALL_VULN_CLASSES } from '../types/config.js';
interface QueueEntry {
ID?: string;
code_locations?: QueueCodeLocation[];
}
/** Read every per-class queue in the deliverables dir into an ID-to-locations map. */
async function loadQueueLocations(
deliverablesPath: string,
logger: ActivityLogger,
): Promise<Map<string, QueueCodeLocation[]>> {
const locations = new Map<string, QueueCodeLocation[]>();
for (const vulnClass of ALL_VULN_CLASSES) {
const queuePath = path.join(deliverablesPath, `${vulnClass}_exploitation_queue.json`);
if (!(await fs.pathExists(queuePath))) continue;
try {
const doc = (await fs.readJson(queuePath)) as { vulnerabilities?: QueueEntry[] };
for (const entry of doc.vulnerabilities ?? []) {
if (entry.ID && entry.code_locations && entry.code_locations.length > 0) {
locations.set(entry.ID, entry.code_locations);
}
}
} catch (error) {
logger.warn(`Could not read ${vulnClass} queue for code locations: ${(error as Error).message}`);
}
}
return locations;
}
/**
* Return the findings with `code_locations` filled in from the queue.
*
* A finding with no matching queue entry keeps none — the join never invents one. Findings are
* copied rather than mutated so the collector's own state stays untouched.
*/
export async function attachQueueCodeLocations(
findings: readonly AddFindingInput[],
deliverablesPath: string,
logger: ActivityLogger,
): Promise<AddFindingInput[]> {
const byId = await loadQueueLocations(deliverablesPath, logger);
if (byId.size === 0) return [...findings];
let matched = 0;
const joined = findings.map((finding) => {
const locations = byId.get(finding.finding_id);
if (!locations) return finding;
matched += 1;
return { ...finding, code_locations: locations };
});
logger.info(`Attached code locations to ${matched}/${findings.length} finding(s) from the vuln queues`);
return joined;
}
+7 -3
View File
@@ -38,9 +38,13 @@ export class ConfigLoaderService {
} catch (error) {
const errorMessage = error instanceof Error ? error.message : String(error);
// parseConfig throws PentestErrors that already name the failure; anything
// else reaching here is a parse-time fault.
const errorCode = error instanceof PentestError && error.code ? error.code : ErrorCode.CONFIG_PARSE_ERROR;
// Determine appropriate error code based on error message
let errorCode = ErrorCode.CONFIG_PARSE_ERROR;
if (errorMessage.includes('not found') || errorMessage.includes('ENOENT')) {
errorCode = ErrorCode.CONFIG_NOT_FOUND;
} else if (errorMessage.includes('validation failed')) {
errorCode = ErrorCode.CONFIG_VALIDATION_FAILED;
}
return err(
new PentestError(
+157 -28
View File
@@ -4,8 +4,8 @@
// it under the terms of the GNU Affero General Public License version 3
// as published by the Free Software Foundation.
import { type AssistantMessage, isRetryableAssistantError } from '@earendil-works/pi-ai';
import { ErrorCode, type PentestErrorContext, type PentestErrorType, type PromptErrorResult } from '../types/errors.js';
import { matchesBillingApiPattern, matchesBillingTextPattern } from '../utils/billing-detection.js';
export class PentestError extends Error {
override name = 'PentestError' as const;
@@ -44,23 +44,53 @@ export function handlePromptError(promptName: string, error: Error): PromptError
};
}
/**
* Whether a failed agent attempt is worth retrying.
*
* A PentestError already carries a verdict — for provider turns that verdict
* comes from pi — so it is taken as given. Anything else is raw text, judged by
* pi's classifier: transient for load, throttling, and transport failures,
* terminal for quota, billing, and auth. Unrecognised errors are not retried, so
* a permanent fault fails fast.
*/
export function isRetryableFailure(error: Error): boolean {
if (error instanceof PentestError) return error.retryable;
const RETRYABLE_PATTERNS = [
// Network and connection errors
'network',
'connection',
'timeout',
'econnreset',
'enotfound',
'econnrefused',
// Rate limiting
'rate limit',
'429',
'too many requests',
// Server errors
'server error',
'5xx',
'internal server error',
'service unavailable',
'bad gateway',
// Provider API errors
'model unavailable',
'service temporarily unavailable',
'api error',
'terminated',
// Max turns
'max turns',
'maximum turns',
];
return isRetryableAssistantError({
role: 'assistant',
stopReason: 'error',
errorMessage: error.message,
} as AssistantMessage);
// Patterns that indicate non-retryable errors (checked before default)
const NON_RETRYABLE_PATTERNS = [
'authentication',
'invalid prompt',
'out of memory',
'permission denied',
'session limit reached',
'invalid api key',
];
// Conservative retry classification - unknown errors don't retry (fail-safe default)
export function isRetryableError(error: Error): boolean {
const message = error.message.toLowerCase();
if (NON_RETRYABLE_PATTERNS.some((pattern) => message.includes(pattern))) {
return false;
}
return RETRYABLE_PATTERNS.some((pattern) => message.includes(pattern));
}
/**
@@ -69,6 +99,14 @@ export function isRetryableFailure(error: Error): boolean {
*/
function classifyByErrorCode(code: ErrorCode, retryableFromError: boolean): { type: string; retryable: boolean } {
switch (code) {
// Billing errors - retryable (wait for cap reset or credits added)
case ErrorCode.SPENDING_CAP_REACHED:
case ErrorCode.INSUFFICIENT_CREDITS:
return { type: 'BillingError', retryable: true };
case ErrorCode.API_RATE_LIMITED:
return { type: 'RateLimitError', retryable: true };
// Config errors - non-retryable (need manual fix)
case ErrorCode.CONFIG_NOT_FOUND:
case ErrorCode.CONFIG_VALIDATION_FAILED:
@@ -105,10 +143,11 @@ function classifyByErrorCode(code: ErrorCode, retryableFromError: boolean): { ty
case ErrorCode.AUTH_LOGIN_FAILED:
return { type: 'AuthLoginFailedError', retryable: false };
case ErrorCode.TARGET_UNREACHABLE:
return { type: 'InvalidTargetError', retryable: false };
case ErrorCode.BILLING_ERROR:
return { type: 'BillingError', retryable: true };
default:
// Unknown code - fall through to string matching
return { type: 'UnknownError', retryable: retryableFromError };
}
}
@@ -122,8 +161,8 @@ function classifyByErrorCode(code: ErrorCode, retryableFromError: boolean): { ty
* - Non-retryable errors: Temporal fails immediately
*
* Classification priority:
* 1. A PentestError carrying an ErrorCode is classified by that code.
* 2. Anything else falls through to isRetryableFailure.
* 1. If error is PentestError with ErrorCode, classify by code (reliable)
* 2. Fall through to string matching for external errors (provider, network, etc.)
*/
export function classifyErrorForTemporal(error: unknown): { type: string; retryable: boolean } {
// === CODE-BASED CLASSIFICATION (Preferred for internal errors) ===
@@ -131,11 +170,101 @@ export function classifyErrorForTemporal(error: unknown): { type: string; retrya
return classifyByErrorCode(error.code, error.retryable);
}
// === FALLBACK ===
// Everything else is a raw throw: a library error, or a PentestError carrying no
// code. isRetryableFailure decides — pi's classifier for provider text, the
// error's own verdict when it has one, and no retry for anything unrecognised.
const err = error instanceof Error ? error : new Error(String(error));
const retryable = isRetryableFailure(err);
return { type: retryable ? 'TransientError' : 'PermanentError', retryable };
// === STRING-BASED CLASSIFICATION (Fallback for external errors) ===
const message = (error instanceof Error ? error.message : String(error)).toLowerCase();
// === BILLING ERRORS (Retryable with long backoff) ===
// Anthropic returns billing as 400 invalid_request_error
// Human can add credits OR wait for spending cap to reset (5-30 min backoff)
// Check both API patterns and text patterns for comprehensive detection
if (matchesBillingApiPattern(message) || matchesBillingTextPattern(message)) {
return { type: 'BillingError', retryable: true };
}
// === PERMANENT ERRORS (Non-retryable) ===
// Authentication (401) - bad API key won't fix itself
if (
message.includes('authentication') ||
message.includes('api key') ||
message.includes('401') ||
message.includes('authentication_error')
) {
return { type: 'AuthenticationError', retryable: false };
}
// Permission (403) - access won't be granted
if (message.includes('permission') || message.includes('forbidden') || message.includes('403')) {
return { type: 'PermissionError', retryable: false };
}
// Out of memory - deterministic resource exhaustion, retrying won't help
if (message.includes('out of memory')) {
return { type: 'OutOfMemoryError', retryable: false };
}
// Invalid prompt - malformed/rejected prompt content won't fix itself on retry
if (message.includes('invalid prompt')) {
return { type: 'InvalidPromptError', retryable: false };
}
// Session limit reached - distinct from billing/rate-limit; needs manual intervention
if (message.includes('session limit reached')) {
return { type: 'SessionLimitError', retryable: false };
}
// Overloaded - provider's own error-type token is authoritative regardless of the
// HTTP status it arrives under (seen in production under 400, not just 529)
if (message.includes('overloaded_error') || message.includes('overloaded')) {
return { type: 'OverloadedError', retryable: true };
}
// === OUTPUT VALIDATION ERRORS (Retryable) ===
// Agent didn't produce expected deliverables - retry may succeed
// IMPORTANT: Must come BEFORE generic 'validation' check below
if (message.includes('failed output validation') || message.includes('output validation failed')) {
return { type: 'OutputValidationError', retryable: true };
}
// Invalid Request (400) - malformed request is permanent
// Note: Checked AFTER billing and AFTER output validation
if (message.includes('invalid_request_error') || message.includes('malformed') || message.includes('validation')) {
return { type: 'InvalidRequestError', retryable: false };
}
// Request Too Large (413) - won't fit no matter how many retries
if (message.includes('request_too_large') || message.includes('too large') || message.includes('413')) {
return { type: 'RequestTooLargeError', retryable: false };
}
// Configuration errors - missing files need manual fix
if (message.includes('enoent') || message.includes('no such file') || message.includes('cli not installed')) {
return { type: 'ConfigurationError', retryable: false };
}
// Execution limits - max turns/budget reached
if (
message.includes('max turns') ||
message.includes('budget') ||
message.includes('execution limit') ||
message.includes('error_max_turns') ||
message.includes('error_max_budget')
) {
return { type: 'ExecutionLimitError', retryable: false };
}
// Invalid target URL - bad URL format won't fix itself
if (
message.includes('invalid url') ||
message.includes('invalid target') ||
message.includes('malformed url') ||
message.includes('invalid uri')
) {
return { type: 'InvalidTargetError', retryable: false };
}
// === TRANSIENT ERRORS (Retryable) ===
// Rate limits (429), server errors (5xx), network issues
// Let Temporal retry with configured backoff
return { type: 'TransientError', retryable: true };
}
@@ -53,15 +53,9 @@ function formatLocation(endpoint: string | undefined, codeLocation: string | und
return endpoint ?? codeLocation ?? '';
}
/** The analysis queue carries no severity, so confidence is the only rating. */
interface CommonEntryFields {
readonly confidence: string;
}
function buildEntry(
id: string,
title: string,
common: CommonEntryFields,
summaryRows: ReadonlyArray<string | null>,
notes: string | undefined,
): string {
@@ -69,7 +63,6 @@ function buildEntry(
lines.push(`### ${id}: ${title}`);
lines.push('');
lines.push('**Summary:**');
lines.push(`- **Confidence:** ${common.confidence}`);
for (const row of summaryRows) {
if (row !== null) lines.push(row);
}
@@ -86,7 +79,6 @@ function renderAuthEntry(e: AuthFinding): string {
return buildEntry(
e.ID,
e.vulnerability_type,
{ confidence: e.confidence },
[
summaryRow('Vulnerable location', formatLocation(e.source_endpoint, e.vulnerable_code_location)),
summaryRow('Overview', e.missing_defense),
@@ -100,7 +92,6 @@ function renderSsrfEntry(e: SsrfFinding): string {
return buildEntry(
e.ID,
e.vulnerability_type,
{ confidence: e.confidence },
[
summaryRow('Vulnerable location', formatLocation(e.source_endpoint, e.vulnerable_code_location)),
summaryRow('Overview', e.missing_defense),
@@ -114,7 +105,6 @@ function renderAuthzEntry(e: AuthzFinding): string {
return buildEntry(
e.ID,
e.vulnerability_type,
{ confidence: e.confidence },
[
summaryRow('Vulnerable location', formatLocation(e.endpoint, e.vulnerable_code_location)),
summaryRow('Overview', e.guard_evidence),
@@ -129,7 +119,6 @@ function renderInjectionEntry(e: InjectionFinding): string {
return buildEntry(
e.ID,
e.vulnerability_type,
{ confidence: e.confidence },
[summaryRow('Vulnerable location', location), summaryRow('Overview', e.mismatch_reason)],
e.notes,
);
@@ -140,7 +129,6 @@ function renderXssEntry(e: XssFinding): string {
return buildEntry(
e.ID,
e.vulnerability_type,
{ confidence: e.confidence },
[summaryRow('Vulnerable location', location), summaryRow('Overview', e.mismatch_reason)],
e.notes,
);
-2
View File
@@ -20,6 +20,4 @@ export type { ContainerDependencies } from './container.js';
export { Container, getContainer, getOrCreateContainer, removeContainer, setContainerFactory } from './container.js';
export { ExploitationCheckerService } from './exploitation-checker.js';
export { loadPrompt } from './prompt-manager.js';
export type { ReportData, ReportMeta } from './report-renderer.js';
export { renderReport } from './report-renderer.js';
export { assembleFinalReport, copyReportToRunRoot, injectModelIntoReport } from './reporting.js';
-98
View File
@@ -1,98 +0,0 @@
// Copyright (C) 2025 Keygraph, Inc.
//
// This program is free software: you can redistribute it and/or modify
// it under the terms of the GNU Affero General Public License version 3
// as published by the Free Software Foundation.
/**
* Typst PDF renderer.
*
* Adapts the structured report.json into the Typst-shaped schema and compiles
* it to a PDF with the bundled report.typ template. Compilation runs in an
* isolated temp dir: the template is copied in and the adapted JSON is written
* beside it so `--root` can scope every file read to that dir, matching how the
* template resolves `--input data=/data.json`.
*
* The `typst` binary is installed in the worker image and resolved from PATH.
*/
import { execFile } from 'node:child_process';
import { existsSync } from 'node:fs';
import { copyFile, cp, mkdir, mkdtemp, rm, writeFile } from 'node:fs/promises';
import { tmpdir } from 'node:os';
import path from 'node:path';
import { promisify } from 'node:util';
import { adaptReportToTypst } from './report-json-adapter.js';
import type { ReportData } from './report-renderer.js';
const execFileAsync = promisify(execFile);
const DEFAULT_TESTER = 'Shannon';
const DEFAULT_BRAND = 'Shannon | AI Pentester by Keygraph';
const DATA_FILENAME = 'data.json';
const TEMPLATE_FILENAME = 'report.typ';
const OUTPUT_FILENAME = 'report.pdf';
export interface RenderReportPdfOptions {
/** Structured report data (report.json contents), pre-assembly. */
readonly reportData: ReportData;
/** Absolute path to the bundled report.typ template. */
readonly templatePath: string;
/** Absolute path where the compiled PDF should be written. */
readonly outputPath: string;
/** Name shown on the cover/footer. Defaults to "Shannon". */
readonly tester?: string;
/** Wordmark shown on the cover. Defaults to "Shannon | AI Pentester by Keygraph". */
readonly brand?: string;
}
/**
* Compile the report to a PDF at `outputPath`.
*
* Throws if adaptation or `typst compile` fails; callers treat the PDF as a
* secondary artifact and should not let a failure here fail the run.
*/
export async function renderReportPdf(options: RenderReportPdfOptions): Promise<void> {
const { reportData, templatePath, outputPath } = options;
const tester = options.tester ?? DEFAULT_TESTER;
const brand = options.brand ?? DEFAULT_BRAND;
const typstData = adaptReportToTypst(reportData);
const workDir = await mkdtemp(path.join(tmpdir(), 'shannon-typst-'));
try {
const templateInWorkDir = path.join(workDir, TEMPLATE_FILENAME);
const dataInWorkDir = path.join(workDir, DATA_FILENAME);
const pdfInWorkDir = path.join(workDir, OUTPUT_FILENAME);
await copyFile(templatePath, templateInWorkDir);
// Ship the template's assets (e.g. the cover logo) so `--root`-scoped image reads resolve.
const assetsDir = path.join(path.dirname(templatePath), 'assets');
if (existsSync(assetsDir)) {
await cp(assetsDir, path.join(workDir, 'assets'), { recursive: true });
}
await writeFile(dataInWorkDir, JSON.stringify(typstData), 'utf-8');
await execFileAsync('typst', [
'compile',
'--root',
workDir,
'--input',
`data=/${DATA_FILENAME}`,
'--input',
`tester=${tester}`,
'--input',
`brand=${brand}`,
templateInWorkDir,
pdfInWorkDir,
]);
await mkdir(path.dirname(outputPath), { recursive: true });
await copyFile(pdfInWorkDir, outputPath);
} finally {
await rm(workDir, { recursive: true, force: true });
}
}
+144 -158
View File
@@ -15,7 +15,7 @@
* 1. Repository path exists and is a directory
* 2. Config file parses and validates (if provided)
* 3. code_path rules match real entries in the repo (filesystem only)
* 4. Credentials validate via a minimal pi session against the run's own model
* 4. Credentials validate via a minimal pi session (API key, OAuth, or Bedrock)
* 5. Target URL resolves, is not link-local (cloud metadata), and is reachable (DNS + HTTP)
*/
@@ -26,36 +26,22 @@ import http from 'node:http';
import https from 'node:https';
import net, { type LookupFunction } from 'node:net';
import os from 'node:os';
import type { Api, AssistantMessage, Model } from '@earendil-works/pi-ai';
import {
type AgentSession,
AuthStorage,
createAgentSession,
type ModelRuntime,
ModelRegistry,
SessionManager,
SettingsManager,
} from '@earendil-works/pi-coding-agent';
import { glob } from 'zx';
import {
type CuratedProviderId,
createModelRuntime,
GENERIC_API_KEY_ENV,
type ModelSpec,
type OpenAiFormat,
PI_CATALOG_URL,
piAuthPresent,
resolveGatewayFormat,
resolveModel,
resolveModelSpec,
resolveProviderCredentials,
} from '../ai/models.js';
import { PI_RETRY_SETTINGS } from '../ai/pi/retry-settings.js';
import { providerTurnError } from '../ai/pi/turn-error.js';
import { resolveEffectiveProvider, resolveModelId } from '../ai/models.js';
import { parseConfig } from '../config-parser.js';
import type { ActivityLogger } from '../types/activity-logger.js';
import type { Config, Rule } from '../types/config.js';
import { ErrorCode } from '../types/errors.js';
import { err, isErr, ok, type Result } from '../types/result.js';
import { isRetryableFailure, PentestError } from './error-handling.js';
import { matchesBillingTextPattern } from '../utils/billing-detection.js';
import { PentestError } from './error-handling.js';
const TARGET_URL_TIMEOUT_MS = 10_000;
@@ -182,7 +168,7 @@ type RuleKind = 'avoid' | 'focus';
interface MissingCodePath {
kind: RuleKind;
value: string;
description?: string;
description: string;
}
async function validateCodePathsExist(
@@ -205,16 +191,12 @@ async function validateCodePathsExist(
const missing: MissingCodePath[] = [];
for (const { kind, rule } of tagged) {
if (!(await patternMatchesAny(repoPath, rule.value))) {
const entry: MissingCodePath = { kind, value: rule.value };
if (rule.description) {
entry.description = rule.description;
}
missing.push(entry);
missing.push({ kind, value: rule.value, description: rule.description });
}
}
if (missing.length > 0) {
const lines = missing.map((m) => `[${m.kind}] '${m.value}'${m.description ? ` - ${m.description}` : ''}`);
const lines = missing.map((m) => `[${m.kind}] '${m.value}' — ${m.description}`);
return err(
new PentestError(
`code_path rules don't match any file or directory in the repo:\n - ${lines.join('\n - ')}\n` +
@@ -233,86 +215,76 @@ async function validateCodePathsExist(
// === Credential Validation ===
/**
* Minimal pi session probe against the model the scan will use, so credentials the
* account cannot use fail here rather than partway through the run. The descriptor
* already carries the run's endpoint and wire format, so the probe exercises the
* same path the scan will.
*/
async function probeCredentialsWithPi(
model: Model<Api>,
modelRuntime: ModelRuntime,
authType: string,
): Promise<Result<void, PentestError>> {
let failedTurn: AssistantMessage | undefined;
let session: AgentSession | undefined;
try {
({ session } = await createAgentSession({
cwd: os.tmpdir(),
model,
noTools: 'all',
modelRuntime,
sessionManager: SessionManager.inMemory(),
settingsManager: SettingsManager.inMemory({ retry: PI_RETRY_SETTINGS, compaction: { enabled: false } }),
}));
session.subscribe((e) => {
if (e.type === 'turn_end' && e.message.role === 'assistant' && e.message.stopReason === 'error') {
failedTurn = e.message;
}
});
await session.prompt('hi');
} catch (error) {
const thrown = error instanceof Error ? error : new Error(String(error));
/** Map provider error text to a human-readable preflight PentestError. */
/** Classify a provider error message (thrown or from a failed turn) into a PentestError. */
function classifyCredentialError(text: string, authType: string): Result<void, PentestError> {
const lower = text.toLowerCase();
if (matchesBillingTextPattern(text)) {
return err(
new PentestError(
`${authType} validation failed: ${thrown.message.slice(0, 300)}`,
'unknown',
isRetryableFailure(thrown),
`Anthropic account has a billing or rate-limit issue during ${authType} validation. Add credits or wait and retry.`,
'billing',
true,
{ authType },
ErrorCode.AGENT_EXECUTION_FAILED,
ErrorCode.BILLING_ERROR,
),
);
} finally {
session?.dispose();
}
if (failedTurn) return err(providerTurnError(failedTurn, `${authType} validation failed`));
return ok(undefined);
}
/** Credential env var a curated provider reads, for "credential missing" messages. */
const PROVIDER_CREDENTIAL_HINT: Readonly<Record<CuratedProviderId, string>> = {
anthropic: 'ANTHROPIC_API_KEY (or CLAUDE_CODE_OAUTH_TOKEN)',
openai: 'OPENAI_API_KEY',
xai: 'XAI_API_KEY',
'amazon-bedrock': 'AWS_BEARER_TOKEN_BEDROCK and AWS_REGION',
};
/** Which variable to set when a provider's credential is missing. */
function credentialHint(providerId: string): string {
const curated = (PROVIDER_CREDENTIAL_HINT as Record<string, string | undefined>)[providerId];
return curated ?? GENERIC_API_KEY_ENV;
}
/** Human-readable label for which credential path a run is using. */
function describeAuth(providerId: string, baseUrl: string | undefined): string {
if (baseUrl) return `custom endpoint (${baseUrl})`;
if (piAuthPresent()) return `${providerId} credentials from pi auth.json`;
if (providerId === 'amazon-bedrock') return 'Bedrock bearer token';
return `${providerId} API key`;
}
/** Validate the model selection and its credentials via a minimal pi session. */
async function validateCredentials(logger: ActivityLogger): Promise<Result<void, PentestError>> {
// 1. Resolve the run's model. A malformed spec or unknown provider fails here,
// before any scan work begins.
let spec: ModelSpec;
try {
spec = resolveModelSpec();
} catch (error) {
if (/401|403|invalid[ _-]?api[ _-]?key|unauthorized|authentication|forbidden|not allowed|x-api-key/.test(lower)) {
return err(
new PentestError(
error instanceof Error ? error.message : String(error),
`Invalid ${authType}. Check your credentials in .env and try again.`,
'config',
false,
{ authType },
ErrorCode.AUTH_FAILED,
),
);
}
if (/model/.test(lower) && /not found|not available|unknown/.test(lower)) {
return err(
new PentestError(
`Configured model is not available for this account. Check ANTHROPIC_*_MODEL in .env.`,
'config',
false,
{ authType },
),
);
}
if (
/network|timeout|enotfound|econnrefused|fetch failed|getaddrinfo|socket|overloaded|unavailable|50\d/.test(lower)
) {
return err(
new PentestError(`Anthropic API unreachable or temporarily unavailable. Try again shortly.`, 'network', true, {
authType,
}),
);
}
return err(
new PentestError(
`${authType} validation failed: ${text.slice(0, 150)}`,
'config',
false,
{ authType },
ErrorCode.AUTH_FAILED,
),
);
}
/** Minimal pi session probe to validate credentials. An optional baseUrl overrides the endpoint. */
async function probeCredentialsWithPi(
authType: string,
token?: string,
baseUrl?: string,
): Promise<Result<void, PentestError>> {
const authStorage = AuthStorage.inMemory();
if (token) authStorage.setRuntimeApiKey('anthropic', token);
const baseModel = ModelRegistry.create(authStorage).find('anthropic', resolveModelId('small'));
if (!baseModel) {
return err(
new PentestError(
`Model not found in pi registry: ${resolveModelId('small')}`,
'config',
false,
{},
@@ -320,78 +292,92 @@ async function validateCredentials(logger: ActivityLogger): Promise<Result<void,
),
);
}
logger.info(`Model: ${spec.providerId}:${spec.modelId}`);
const model = baseUrl ? { ...baseModel, baseUrl } : baseModel;
// 2. Credential presence. Bedrock needs both AWS_ vars; every other provider
// needs one API key.
const credentials = resolveProviderCredentials(spec.providerId);
// 3. Wire format for an OpenAI gateway. Rejects a format named where it cannot
// take effect, rather than letting the run proceed on the wrong API.
let format: OpenAiFormat;
let errText: string | undefined;
try {
format = resolveGatewayFormat(spec.providerId, credentials.baseUrl);
const { session } = await createAgentSession({
cwd: os.tmpdir(),
model,
thinkingLevel: 'off',
noTools: 'all',
authStorage,
sessionManager: SessionManager.inMemory(),
settingsManager: SettingsManager.inMemory({ retry: { enabled: false }, compaction: { enabled: false } }),
});
session.subscribe((e) => {
if (e.type === 'turn_end' && e.message.role === 'assistant' && e.message.stopReason === 'error') {
errText = e.message.errorMessage ?? 'unknown provider error';
}
});
await session.prompt('hi');
session.dispose();
} catch (error) {
errText = error instanceof Error ? error.message : String(error);
}
if (errText) return classifyCredentialError(errText, authType);
return ok(undefined);
}
/** Validate credentials via a minimal pi session. */
async function validateCredentials(logger: ActivityLogger): Promise<Result<void, PentestError>> {
// Resolve the active provider through the same precedence the executor uses, so
// preflight validates exactly the credentials the run will use (no drift).
const eff = resolveEffectiveProvider();
// 1. Bedrock mode — validate required AWS credentials are present (pi-ai owns the
// live AWS auth, so there is no cheap session probe here)
if (eff.providerId === 'amazon-bedrock') {
const required = [
'AWS_REGION',
'AWS_BEARER_TOKEN_BEDROCK',
'ANTHROPIC_SMALL_MODEL',
'ANTHROPIC_MEDIUM_MODEL',
'ANTHROPIC_LARGE_MODEL',
];
const missing = required.filter((v) => !process.env[v]);
if (missing.length > 0) {
return err(
new PentestError(
`Bedrock mode requires the following env vars in .env: ${missing.join(', ')}`,
'config',
false,
{ missing },
ErrorCode.AUTH_FAILED,
),
);
}
logger.info('Bedrock credentials OK');
return ok(undefined);
}
// 2. Custom base URL — validate the endpoint via a minimal pi session
if (eff.baseUrl) {
logger.info('Validating custom base URL');
const probe = await probeCredentialsWithPi(`custom endpoint (${eff.baseUrl})`, eff.anthropicToken, eff.baseUrl);
if (isErr(probe)) return probe;
logger.info('Custom base URL OK');
return ok(undefined);
}
// 3. Direct Anthropic — require a credential, then validate via a minimal pi session
if (!eff.anthropicToken) {
return err(
new PentestError(
error instanceof Error ? error.message : String(error),
'No API credentials found. Set ANTHROPIC_API_KEY or CLAUDE_CODE_OAUTH_TOKEN in .env (or use CLAUDE_CODE_USE_BEDROCK=1 for AWS Bedrock)',
'config',
false,
{ providerId: spec.providerId },
{},
ErrorCode.AUTH_FAILED,
),
);
}
// With a mounted pi auth.json the env-var checks don't apply — step 5's probe validates it.
const isBedrock = spec.providerId === 'amazon-bedrock';
const missing =
isBedrock && !piAuthPresent() ? ['AWS_REGION', 'AWS_BEARER_TOKEN_BEDROCK'].filter((n) => !process.env[n]) : [];
if (!piAuthPresent() && (missing.length > 0 || (!isBedrock && !credentials.apiKey))) {
return err(
new PentestError(
`No credentials found for provider "${spec.providerId}". Set ${credentialHint(spec.providerId)} in .env.`,
'config',
false,
{ providerId: spec.providerId, ...(missing.length > 0 && { missing }) },
ErrorCode.AUTH_FAILED,
),
);
}
// 4. Model must exist in the registry, for every provider — Bedrock IDs are the
// easiest to get wrong, since region prefixes and version suffixes differ per
// model (`us.anthropic.claude-opus-5` exists, bare `anthropic.` does not).
// A custom endpoint is exempt: it may serve models under its own names.
const modelRuntime = await createModelRuntime(spec.providerId, credentials.apiKey);
const baseModel = resolveModel(modelRuntime, spec.providerId, spec.modelId, credentials.baseUrl, format);
if (!baseModel) {
return err(
new PentestError(
`Model not found in pi registry: provider="${spec.providerId}" model="${spec.modelId}". Check SHANNON_AI_MODEL — browse valid providers and models at ${PI_CATALOG_URL}.`,
'config',
false,
{ providerId: spec.providerId, modelId: spec.modelId },
ErrorCode.AUTH_FAILED,
),
);
}
if (!modelRuntime.getModel(spec.providerId, spec.modelId)) {
logger.warn(
`Model "${spec.modelId}" is not in the ${spec.providerId} catalogue; passing it to the custom endpoint as given. Cost figures will be approximate.`,
);
}
if (credentials.baseUrl && spec.providerId === 'openai') {
logger.info(`Gateway API: ${format} (${baseModel.api})`);
}
// 5. One real request, so a credential the account cannot use fails here
// rather than partway through the run. Bedrock included: pi resolves the
// bearer token from the primed credential and the region from AWS_REGION,
// so the probe exercises the same auth path the scan will.
const authType = describeAuth(spec.providerId, credentials.baseUrl);
const usingApiKey = Boolean(process.env.ANTHROPIC_API_KEY);
const authType = usingApiKey ? 'API key' : 'OAuth token';
logger.info(`Validating ${authType} via pi...`);
const probe = await probeCredentialsWithPi(baseModel, modelRuntime, authType);
const probe = await probeCredentialsWithPi(authType, eff.anthropicToken);
if (isErr(probe)) return probe;
logger.info(`${authType} OK`);
return ok(undefined);
+58 -104
View File
@@ -8,81 +8,63 @@ import { fs, path } from 'zx';
import { PROMPTS_DIR } from '../paths.js';
import { PLAYWRIGHT_SESSION_MAPPING } from '../session-manager.js';
import type { ActivityLogger } from '../types/activity-logger.js';
import type { Authentication, DistributedConfig, DistributedReportConfig, Rule, VulnClass } from '../types/config.js';
import type { Authentication, DistributedConfig, ReportConfig, Rule, VulnClass } from '../types/config.js';
import { isGlobPattern } from '../utils/glob.js';
import { handlePromptError, PentestError } from './error-handling.js';
function renderRuleLine(tag: string, value: string, description?: string): string {
const base = `- ${tag} ${value}`;
return description ? `${base} - ${description}` : base;
}
function renderUrlRules(rules: Rule[]): string {
if (rules.length === 0) return 'None';
return rules.map((r) => renderRuleLine(`[${r.type.toUpperCase()}]`, r.value, r.description)).join('\n');
}
function renderCodePathRules(rules: Rule[]): string {
const filtered = rules.filter((r) => r.type === 'code_path');
if (filtered.length === 0) return 'None';
return filtered
.map((r) => renderRuleLine(isGlobPattern(r.value) ? '[GLOB]' : '[FILE]', r.value, r.description))
.map((r) => {
const kind = isGlobPattern(r.value) ? '[GLOB]' : '[FILE]';
return `- ${r.value} ${kind} — ${r.description}`;
})
.join('\n');
}
const VULN_CLASS_HEADINGS: Record<VulnClass, string> = {
auth: 'Authentication Vulnerabilities',
authz: 'Authorization Vulnerabilities',
xss: 'Cross-Site Scripting (XSS) Vulnerabilities',
injection: 'SQL/Command Injection Vulnerabilities',
ssrf: 'Server-Side Request Forgery (SSRF) Vulnerabilities',
};
/**
* Renders the <not_assessed_classes> block. Empty when every class completed.
*
* A class whose analysis failed was never assessed, so the report must not present its
* absence of findings as a clean result. The block is authoritative for that caveat.
*/
function renderNotAssessedClassesBlock(failed: readonly VulnClass[] = []): string {
if (failed.length === 0) {
return '';
}
const classes = [...new Set(failed)];
const lines: string[] = [
'<not_assessed_classes>',
'The following vulnerability classes did not complete and were NOT assessed in this run. Treat this list as authoritative for completeness caveats.',
'',
];
for (const cls of classes) {
lines.push(
`- ${VULN_CLASS_HEADINGS[cls]}: analysis did not complete; this class was NOT assessed. Absence of findings here does not indicate the class is clean.`,
);
}
lines.push(
'',
'When writing report_meta.executive_summary, scope any no-findings statement to the classes that were assessed and mention these not-assessed classes. Do not state or imply that the target is clean for these classes.',
'</not_assessed_classes>',
);
return lines.join('\n');
interface VulnSummarySpec {
readonly heading: string;
readonly evidenceSection: string;
readonly noneFoundLabel: string;
}
/**
* Which configured filters this run can actually enforce.
*
* Every finding carries `severity` (see ../collectors/finding-collector.ts), so a severity
* threshold always applies. `confidence` exists only on an analysed finding — handing an
* exploit run a confidence threshold is a directive it cannot honor.
*/
function applicableFilters(report: DistributedReportConfig | undefined, exploitEnabled: boolean) {
return {
severity: Boolean(report?.min_severity),
confidence: Boolean(report?.min_confidence) && !exploitEnabled,
guidance: Boolean(report?.guidance?.trim()),
};
const VULN_SUMMARY_SPECS: Record<VulnClass, VulnSummarySpec> = {
auth: {
heading: 'Authentication Vulnerabilities',
evidenceSection: 'Authentication Exploitation Evidence',
noneFoundLabel: 'authentication',
},
authz: {
heading: 'Authorization Vulnerabilities',
evidenceSection: 'Authorization Exploitation Evidence',
noneFoundLabel: 'authorization',
},
xss: {
heading: 'Cross-Site Scripting (XSS) Vulnerabilities',
evidenceSection: 'XSS Exploitation Evidence',
noneFoundLabel: 'XSS',
},
injection: {
heading: 'SQL/Command Injection Vulnerabilities',
evidenceSection: 'Injection Exploitation Evidence',
noneFoundLabel: 'SQL or command injection',
},
ssrf: {
heading: 'Server-Side Request Forgery (SSRF) Vulnerabilities',
evidenceSection: 'SSRF Exploitation Evidence',
noneFoundLabel: 'SSRF',
},
};
function renderVulnSummarySubsections(selected: readonly VulnClass[]): string {
const classes = selected.length > 0 ? selected : (Object.keys(VULN_SUMMARY_SPECS) as VulnClass[]);
return classes
.map((cls) => {
const spec = VULN_SUMMARY_SPECS[cls];
return `**${spec.heading}:**\n{Check for "${spec.evidenceSection}" section. Include actually exploited vulnerabilities and those blocked by security controls. Exclude theoretical vulnerabilities requiring internal network access. If vulnerabilities exist, summarize their impact and severity. If section is missing or empty, state: "No ${spec.noneFoundLabel} vulnerabilities were found."}`;
})
.join('\n\n');
}
/**
@@ -90,23 +72,22 @@ function applicableFilters(report: DistributedReportConfig | undefined, exploitE
* each filter is included only when the operator configured it, so the agent
* never sees `none` placeholders or instructions for filters that don't apply.
*/
function renderReportFiltersBlock(report: DistributedReportConfig | undefined, exploitEnabled: boolean): string {
function renderReportFiltersBlock(report: ReportConfig | undefined): string {
if (!report) return '';
const guidance = report.guidance?.trim();
const applies = applicableFilters(report, exploitEnabled);
if (!applies.severity && !applies.confidence && !applies.guidance) return '';
if (!report.min_severity && !report.min_confidence && !guidance) return '';
const lines: string[] = [
'<report_filters>',
'The filters below are user-supplied and binding for this assessment. Honor each strictly when assembling the final report.',
'',
];
if (applies.severity) {
if (report.min_severity) {
lines.push(
`- Minimum severity: ${report.min_severity} — keep only findings rated this severity or higher (scale: low < medium < high < critical).`,
);
}
if (applies.confidence) {
if (report.min_confidence) {
lines.push(
`- Minimum confidence: ${report.min_confidence} — keep only findings rated this confidence or higher (scale: low < medium < high).`,
);
@@ -125,11 +106,10 @@ function renderReportFiltersBlock(report: DistributedReportConfig | undefined, e
* confidence inline as concrete thresholds; guidance is referenced by pointer
* so the actual text only lives in <report_filters>, avoiding double-statement.
*/
function renderReportFilterRules(report: DistributedReportConfig | undefined, exploitEnabled: boolean): string {
const applies = applicableFilters(report, exploitEnabled);
function renderReportFilterRules(report: ReportConfig | undefined): string {
const drops: string[] = [];
if (applies.severity) drops.push(`* severity is below ${report?.min_severity}`);
if (applies.confidence) drops.push(`* confidence is below ${report?.min_confidence}`);
if (report?.min_severity) drops.push(`* severity is below ${report.min_severity}`);
if (report?.min_confidence) drops.push(`* confidence is below ${report.min_confidence}`);
if (report?.guidance?.trim()) drops.push('* topic matches an exclusion in the user guidance');
if (drops.length === 0) return '';
return [' - DROP any `### [TYPE]-VULN-[NUMBER]` finding whose:', ...drops.map((d) => ` ${d}`)].join('\n');
@@ -138,8 +118,6 @@ function renderReportFilterRules(report: DistributedReportConfig | undefined, ex
interface PromptVariables {
webUrl: string;
repoPath: string;
/** Classes whose analysis did not complete, so the report can mark them not assessed. */
failedClasses?: readonly VulnClass[];
AUTH_STATE_FILE: string;
PLAYWRIGHT_SESSION?: string;
}
@@ -345,8 +323,8 @@ async function interpolateVariables(
if (avoidUrlRules.length === 0 && focusUrlRules.length === 0) {
result = result.replace(/<rules>[\s\S]*?<\/rules>\s*/g, '');
} else {
const avoidStr = renderUrlRules(avoidUrlRules);
const focusStr = renderUrlRules(focusUrlRules);
const avoidStr = avoidUrlRules.length > 0 ? avoidUrlRules.map((r) => `- ${r.description}`).join('\n') : 'None';
const focusStr = focusUrlRules.length > 0 ? focusUrlRules.map((r) => `- ${r.description}`).join('\n') : 'None';
result = replaceLiteral(result, /{{RULES_AVOID}}/g, avoidStr);
result = replaceLiteral(result, /{{RULES_FOCUS}}/g, focusStr);
}
@@ -386,43 +364,19 @@ async function interpolateVariables(
/{{VULN_CLASSES_TESTED}}/g,
vulnClasses.length > 0 ? vulnClasses.join(', ') : 'injection, xss, auth, authz, ssrf',
);
result = replaceLiteral(
result,
/{{NOT_ASSESSED_CLASSES}}/g,
renderNotAssessedClassesBlock(variables.failedClasses ?? []),
);
result = replaceLiteral(result, /{{VULN_SUMMARY_SUBSECTIONS}}/g, renderVulnSummarySubsections(vulnClasses));
const exploitEnabled = config?.exploit ?? true;
// Drop every block belonging to the mode this run is not in, so the prompt never documents
// a field the tool would reject. The backreference pins each match to a closed pair.
const droppedMode = exploitEnabled ? 'analysis' : 'exploit';
result = result.replace(new RegExp(`<(${droppedMode}_mode_[a-z_]+)>[\\s\\S]*?</\\1>\\n?`, 'g'), '');
result = result.replace(/<\/?(?:exploit|analysis)_mode_[a-z_]+>\n?/g, '');
result = replaceLiteral(result, /{{EXPLOITATION}}/g, exploitEnabled ? 'enabled' : 'disabled');
result = replaceLiteral(result, /{{REPORT_VULN_HEADING}}/g, exploitEnabled ? 'Exploitation Evidence' : 'Findings');
result = replaceLiteral(
result,
/{{REPORT_VULN_SUBHEADING}}/g,
exploitEnabled ? 'Successfully Exploited Vulnerabilities' : 'Identified Vulnerabilities',
);
if (config?.report?.min_confidence && exploitEnabled) {
logger.warn(
`report.min_confidence="${config.report.min_confidence}" is ignored when exploit=true: an ` +
'exploited finding is rated by severity, not confidence. Use report.min_severity.',
);
}
result = replaceLiteral(
result,
/{{REPORT_FILTERS_BLOCK}}/g,
renderReportFiltersBlock(config?.report, exploitEnabled),
);
result = replaceLiteral(
result,
/{{REPORT_FILTER_RULES}}/g,
renderReportFilterRules(config?.report, exploitEnabled),
);
result = replaceLiteral(result, /{{REPORT_FILTERS_BLOCK}}/g, renderReportFiltersBlock(config?.report));
result = replaceLiteral(result, /{{REPORT_FILTER_RULES}}/g, renderReportFilterRules(config?.report));
// Collapse runs of 3+ newlines (left behind by tag-strip and empty-fragment substitutions).
result = result.replace(/\n{3,}/g, '\n\n');
@@ -1,293 +0,0 @@
// Copyright (C) 2025 Keygraph, Inc.
//
// This program is free software: you can redistribute it and/or modify
// it under the terms of the GNU Affero General Public License version 3
// as published by the Free Software Foundation.
/**
* Programmatic adapter: report.json → Typst ReportData JSON.
*
* Converts the renderer-neutral structured report output (produced by the
* finding-collector + set-report-meta CLI) into the Typst-specific schema that
* report.typ consumes.
*
* All Typst-specific concepts (PascalCase enums, computed aggregations,
* exploitedByType grouping) are confined to this file. The rest of the
* pipeline knows nothing about the Typst shape.
*/
import type { AddFindingInput, AdditionalSection, StepItem, StructuredStep } from '../collectors/finding-collector.js';
import type {
ExploitsReportData,
FindingsReportData,
TypstCategory,
TypstConfidence,
ReportData as TypstReportData,
TypstSeverity,
TypstStatus,
} from './report-output-schema.js';
import type { ReportData } from './report-renderer.js';
// ============================================================================
// CASING TRANSFORMS
// ============================================================================
const SEVERITY_MAP: Record<string, TypstSeverity> = {
critical: 'Critical',
high: 'High',
medium: 'Medium',
low: 'Low',
};
const STATUS_MAP: Record<string, TypstStatus> = {
exploited: 'Exploited',
out_of_scope: 'OutOfScope',
blocked_by_constraints: 'BlockedByConstraints',
false_positive: 'FalsePositive',
};
const CONFIDENCE_MAP: Record<string, TypstConfidence> = {
high: 'High',
medium: 'Medium',
low: 'Low',
};
const VALID_CATEGORIES = new Set<TypstCategory>([
'Authentication',
'Authorization',
'XSS',
'Injection',
'SSRF',
'Other',
]);
function toTypstSeverity(s: string): TypstSeverity {
return SEVERITY_MAP[s] ?? 'Low';
}
function toTypstStatus(s: string): TypstStatus {
return STATUS_MAP[s] ?? 'Exploited';
}
function toTypstConfidence(s: string): TypstConfidence {
return CONFIDENCE_MAP[s] ?? 'Medium';
}
function toTypstCategory(s: string): TypstCategory {
if (VALID_CATEGORIES.has(s as TypstCategory)) return s as TypstCategory;
return 'Other';
}
// ============================================================================
// STEP / ITEM TRANSFORMS
// ============================================================================
function adaptStepItem(item: StepItem): StepItem {
return item;
}
function adaptStep(step: StructuredStep, index: number): { number: number; title?: string; items: StepItem[] } {
return {
number: index + 1,
...(step.title && { title: step.title }),
items: step.items.map(adaptStepItem),
};
}
function adaptAdditionalSection(section: AdditionalSection): { heading: string; items: StepItem[] } {
return {
heading: section.heading,
items: section.items.map(adaptStepItem),
};
}
// ============================================================================
// AGGREGATION HELPERS
// ============================================================================
interface CategoryGroup {
category: TypstCategory;
findings: AddFindingInput[];
}
function groupByCategory(findings: readonly AddFindingInput[]): CategoryGroup[] {
const map = new Map<TypstCategory, AddFindingInput[]>();
for (const f of findings) {
const cat = toTypstCategory(f.category);
const list = map.get(cat) ?? [];
list.push(f);
map.set(cat, list);
}
return Array.from(map.entries()).map(([category, fs]) => ({ category, findings: fs }));
}
function countBySeverity(findings: readonly AddFindingInput[]): Record<TypstSeverity, number> {
const counts: Record<string, number> = {
Critical: 0,
High: 0,
Medium: 0,
Low: 0,
};
for (const f of findings) {
const sev = toTypstSeverity(f.severity);
counts[sev] = (counts[sev] ?? 0) + 1;
}
return counts as Record<TypstSeverity, number>;
}
// ============================================================================
// EXPLOIT MODE ADAPTER
// ============================================================================
function adaptExploitsMode(data: ReportData): ExploitsReportData {
const { report_meta, findings } = data;
const groups = groupByCategory(findings);
const sevCounts = countBySeverity(findings);
const statusCounts = { Exploited: 0, OutOfScope: 0, BlockedByConstraints: 0, FalsePositive: 0 };
for (const f of findings) {
const s = toTypstStatus(f.status ?? 'exploited');
statusCounts[s]++;
}
const exploitedFindings = findings.filter((f) => (f.status ?? 'exploited') === 'exploited');
return {
mode: 'exploits' as const,
meta: {
target: report_meta.target,
assessmentDate: report_meta.assessment_date,
classification: 'CONFIDENTIAL',
},
scope: report_meta.scope,
exploitedByType: groups.map((g) => {
const exploited = g.findings.filter((f) => (f.status ?? 'exploited') === 'exploited');
if (exploited.length === 0) {
return {
category: g.category,
narrative: `No ${g.category.toLowerCase()} vulnerabilities were successfully exploited during this assessment.`,
};
}
return {
category: g.category,
bullets: exploited.map((f) => ({ id: f.finding_id, description: f.title })),
};
}),
summary: {
totalIdentified: findings.length,
successfullyExploited: exploitedFindings.length,
exploitedBreakdown: groups
.map((g) => ({
category: g.category,
count: g.findings.filter((f) => (f.status ?? 'exploited') === 'exploited').length,
}))
.filter((e) => e.count > 0),
criticalFindings: findings.filter((f) => f.severity === 'critical').map((f) => `${f.finding_id}: ${f.title}`),
},
findings: findings.map((f) => ({
id: f.finding_id,
title: f.title,
category: toTypstCategory(f.category),
severity: toTypstSeverity(f.severity),
status: toTypstStatus(f.status ?? 'exploited'),
summary: {
vulnerableLocation: f.vulnerable_location,
overview: f.overview,
impact: f.impact,
},
// This branch only runs for an exploitative report, where the schema made these
// required. The fallbacks keep the superset type honest rather than assuming.
prerequisites: f.prerequisites ?? '',
exploitationSteps: (f.exploitation_steps ?? []).map(adaptStep),
proofOfImpact: (f.proof_of_impact ?? []).map(adaptStepItem),
...(f.notes && f.notes.length > 0 && { notes: f.notes.map(adaptStepItem) }),
...(f.additional_sections &&
f.additional_sections.length > 0 && {
additionalSections: f.additional_sections.map(adaptAdditionalSection),
}),
})),
derivedCounts: {
bySeverity: sevCounts,
byStatus: statusCounts,
},
};
}
// ============================================================================
// FINDINGS MODE ADAPTER
// ============================================================================
function adaptFindingsMode(data: ReportData): FindingsReportData {
const { report_meta, findings } = data;
const groups = groupByCategory(findings);
const sevCounts = countBySeverity(findings);
const confidenceCounts = { High: 0, Medium: 0, Low: 0 };
for (const f of findings) {
const c = toTypstConfidence(f.confidence ?? 'medium');
confidenceCounts[c]++;
}
return {
mode: 'findings' as const,
meta: {
target: report_meta.target,
assessmentDate: report_meta.assessment_date,
classification: 'CONFIDENTIAL',
},
scope: report_meta.scope,
identifiedByType: groups.map((g) => {
if (g.findings.length === 0) {
return {
category: g.category,
narrative: `No ${g.category.toLowerCase()} vulnerabilities were identified during this assessment.`,
};
}
return {
category: g.category,
bullets: g.findings.map((f) => ({ id: f.finding_id, description: f.title })),
};
}),
summary: {
totalIdentified: findings.length,
identifiedBreakdown: groups.map((g) => ({
category: g.category,
count: g.findings.length,
})),
criticalFindings: findings.filter((f) => f.severity === 'critical').map((f) => `${f.finding_id}: ${f.title}`),
},
findings: findings.map((f) => ({
id: f.finding_id,
title: f.title,
category: toTypstCategory(f.category),
severity: toTypstSeverity(f.severity),
confidence: toTypstConfidence(f.confidence ?? 'medium'),
summary: {
vulnerableLocation: f.vulnerable_location,
overview: f.overview,
impact: f.impact,
},
...(f.notes && f.notes.length > 0 && { notes: f.notes.map(adaptStepItem) }),
...(f.additional_sections &&
f.additional_sections.length > 0 && {
additionalSections: f.additional_sections.map(adaptAdditionalSection),
}),
})),
derivedCounts: {
bySeverity: sevCounts,
byConfidence: confidenceCounts,
},
};
}
// ============================================================================
// PUBLIC API
// ============================================================================
export function adaptReportToTypst(data: ReportData): TypstReportData {
const exploitEnabled = data.report_meta.exploit ?? true;
if (exploitEnabled) {
return adaptExploitsMode(data);
}
return adaptFindingsMode(data);
}
@@ -1,157 +0,0 @@
// Copyright (C) 2025 Keygraph, Inc.
//
// This program is free software: you can redistribute it and/or modify
// it under the terms of the GNU Affero General Public License version 3
// as published by the Free Software Foundation.
/**
* TypeScript types for the structured report the Typst template consumes, in two
* shapes keyed by a `mode` discriminator: `exploits` (exploit=true) and `findings`
* (exploit=false, analysis-only). Types only — the object is built programmatically
* in report-json-adapter.ts, so these exist to keep the adapter and report.typ in sync.
*/
// === Shared primitives ===
export type TypstSeverity = 'Critical' | 'High' | 'Medium' | 'Low';
export type TypstStatus = 'Exploited' | 'OutOfScope' | 'BlockedByConstraints' | 'FalsePositive';
export type TypstConfidence = 'High' | 'Medium' | 'Low';
export type TypstCategory = 'Authentication' | 'Authorization' | 'XSS' | 'Injection' | 'SSRF' | 'Other';
export interface CodeBlock {
readonly language: string;
readonly content: string;
}
export type StepItem =
| { readonly kind: 'prose'; readonly text: string }
| { readonly kind: 'code'; readonly block: CodeBlock };
export interface Step {
readonly number: number;
readonly title?: string;
readonly items: readonly StepItem[];
}
export interface AdditionalSection {
readonly heading: string;
readonly items: readonly StepItem[];
}
export interface FindingSummary {
readonly vulnerableLocation: string;
readonly overview: string;
readonly impact: string;
}
export interface Meta {
readonly target: string;
readonly assessmentDate: string;
readonly tester?: string;
readonly application?: string;
readonly classification: string;
}
export interface CategoryCount {
readonly category: TypstCategory;
readonly count: number;
readonly note?: string;
}
export type SeverityCounts = Record<TypstSeverity, number>;
export type StatusCounts = Record<TypstStatus, number>;
export type ConfidenceCounts = Record<TypstConfidence, number>;
export interface TypeEntryBullet {
readonly id: string;
readonly description: string;
}
// === Exploits-mode schema ===
export interface ExploitFinding {
readonly id: string;
readonly title: string;
readonly category: TypstCategory;
readonly severity: TypstSeverity;
readonly status: TypstStatus;
readonly summary: FindingSummary;
readonly prerequisites: string;
readonly exploitationSteps: readonly Step[];
readonly proofOfImpact: readonly StepItem[];
readonly notes?: readonly StepItem[];
readonly additionalSections?: readonly AdditionalSection[];
}
export interface ExploitedByTypeEntry {
readonly category: TypstCategory;
readonly bullets?: readonly TypeEntryBullet[];
readonly narrative?: string;
}
export interface ExploitsReportData {
readonly mode: 'exploits';
readonly meta: Meta;
readonly scope: string;
readonly exploitedByType: readonly ExploitedByTypeEntry[];
readonly summary: {
readonly totalIdentified: number;
readonly successfullyExploited: number;
readonly exploitedBreakdown: readonly CategoryCount[];
readonly outOfScope?: {
readonly total: number;
readonly breakdown?: readonly CategoryCount[];
readonly note?: string;
};
readonly blockedByConstraints?: {
readonly total: number;
readonly note?: string;
};
readonly criticalFindings: readonly string[];
};
readonly findings: readonly ExploitFinding[];
readonly derivedCounts: {
readonly bySeverity: SeverityCounts;
readonly byStatus: StatusCounts;
};
}
// === Findings-mode schema (analysis-only, exploit=false runs) ===
export interface AnalysisFinding {
readonly id: string;
readonly title: string;
readonly category: TypstCategory;
readonly severity: TypstSeverity;
readonly confidence: TypstConfidence;
readonly summary: FindingSummary;
readonly notes?: readonly StepItem[];
readonly additionalSections?: readonly AdditionalSection[];
}
export interface IdentifiedByTypeEntry {
readonly category: TypstCategory;
readonly bullets?: readonly TypeEntryBullet[];
readonly narrative?: string;
}
export interface FindingsReportData {
readonly mode: 'findings';
readonly meta: Meta;
readonly scope: string;
readonly identifiedByType: readonly IdentifiedByTypeEntry[];
readonly summary: {
readonly totalIdentified: number;
readonly identifiedBreakdown: readonly CategoryCount[];
readonly criticalFindings: readonly string[];
};
readonly findings: readonly AnalysisFinding[];
readonly derivedCounts: {
readonly bySeverity: SeverityCounts;
readonly byConfidence: ConfidenceCounts;
};
}
// === Discriminated union for downstream consumers that handle both ===
export type ReportData = ExploitsReportData | FindingsReportData;
-298
View File
@@ -1,298 +0,0 @@
// Copyright (C) 2025 Keygraph, Inc.
//
// This program is free software: you can redistribute it and/or modify
// it under the terms of the GNU Affero General Public License version 3
// as published by the Free Software Foundation.
/**
* Deterministic report.json → markdown renderer.
*
* Converts the structured report output (produced by the finding-collector
* tool + set-report-meta CLI) into the same markdown format that the
* report agent previously wrote by hand. No LLM in the loop.
*/
import type { AddFindingInput, AdditionalSection, StepItem, StructuredStep } from '../collectors/finding-collector.js';
import type { VulnClass } from '../types/config.js';
// ============================================================================
// TYPES
// ============================================================================
export interface ReportMeta {
readonly target: string;
readonly assessment_date: string;
readonly scope: string;
readonly executive_summary: string;
readonly exploit?: boolean;
readonly model?: string;
}
export interface ReportData {
readonly report_meta: ReportMeta;
readonly findings: readonly AddFindingInput[];
// Vuln classes whose pipeline failed and were not assessed this run. Rendered as an explicit
// caveat so an un-assessed class is never presented as a clean result.
readonly not_assessed?: readonly VulnClass[];
}
// Without this, an analysis-only report reads as though the impact was demonstrated.
const ANALYSIS_ONLY_DISCLAIMER = [
'> Exploitation was not run for this assessment. Each finding documents a vulnerability',
'> identified through analysis; impact is assessed rather than demonstrated, and no live',
'> exploitation steps or proof of impact are included.',
].join('\n');
const NOT_ASSESSED_LABELS: Record<VulnClass, string> = {
auth: 'Authentication',
authz: 'Authorization',
xss: 'Cross-Site Scripting (XSS)',
injection: 'SQL/Command Injection',
ssrf: 'Server-Side Request Forgery (SSRF)',
};
function renderNotAssessedSection(notAssessed: readonly VulnClass[]): string {
const lines: string[] = ['## Not Assessed', ''];
lines.push(
'The following vulnerability classes were NOT assessed in this run because their analysis did ' +
'not complete. Absence of findings for these classes does not indicate they are clean — re-run ' +
'to assess them:',
);
lines.push('');
for (const cls of notAssessed) {
lines.push(`- ${NOT_ASSESSED_LABELS[cls]} — analysis did not complete; not assessed.`);
}
return lines.join('\n');
}
// ============================================================================
// STEP ITEM RENDERING
// ============================================================================
function renderStepItem(item: StepItem): string {
if (item.kind === 'prose') {
return item.text;
}
const lang = item.block.language || '';
return `\`\`\`${lang}\n${item.block.content}\n\`\`\``;
}
function renderStepItems(items: readonly StepItem[]): string {
return items.map(renderStepItem).join('\n\n');
}
function renderStructuredStep(step: StructuredStep, index: number): string {
const lines: string[] = [];
const title = step.title ? `**Step ${index + 1}: ${step.title}**` : `**Step ${index + 1}**`;
lines.push(title);
lines.push('');
lines.push(renderStepItems(step.items));
return lines.join('\n');
}
function renderAdditionalSection(section: AdditionalSection): string {
const lines: string[] = [];
lines.push(`#### ${section.heading}`);
lines.push('');
lines.push(renderStepItems(section.items));
return lines.join('\n');
}
// ============================================================================
// FINDING RENDERING
// ============================================================================
function titleCase(s: string): string {
return s.charAt(0).toUpperCase() + s.slice(1);
}
function renderFinding(finding: AddFindingInput, exploitEnabled: boolean): string {
const lines: string[] = [];
// Heading
lines.push(`### ${finding.finding_id}: ${finding.title}`);
lines.push('');
// Each row is emitted only when the mode that produced the finding supplied its field.
lines.push('**Summary:**');
if (finding.severity) {
lines.push(`- **Severity:** ${titleCase(finding.severity)}`);
}
if (finding.confidence) {
lines.push(`- **Confidence:** ${titleCase(finding.confidence)}`);
}
lines.push(`- **OWASP:** ${finding.owasp_category}`);
lines.push(`- **Vulnerable location:** ${finding.vulnerable_location}`);
if (finding.auth_state) {
lines.push(`- **Auth state:** ${finding.auth_state}`);
}
if (exploitEnabled && finding.status) {
lines.push(`- **Status:** ${titleCase(finding.status)}`);
}
if (finding.prerequisites) {
lines.push(`- **Prerequisites:** ${finding.prerequisites}`);
}
lines.push('');
// Overview
lines.push('**Overview:**');
lines.push(finding.overview);
lines.push('');
// Impact
lines.push('**Impact:**');
lines.push(finding.impact);
lines.push('');
if (finding.exploitation_steps && finding.exploitation_steps.length > 0) {
lines.push('**Exploitation Steps:**');
lines.push('');
for (let i = 0; i < finding.exploitation_steps.length; i++) {
lines.push(renderStructuredStep(finding.exploitation_steps[i]!, i));
lines.push('');
}
}
if (finding.proof_of_impact && finding.proof_of_impact.length > 0) {
lines.push('**Proof of Impact:**');
lines.push('');
lines.push(renderStepItems(finding.proof_of_impact));
lines.push('');
}
// Remediation
lines.push('**Remediation:**');
lines.push(finding.remediation);
lines.push('');
// Notes
if (finding.notes && finding.notes.length > 0) {
lines.push('**Notes:**');
lines.push('');
lines.push(renderStepItems(finding.notes));
lines.push('');
}
// Additional sections
if (finding.additional_sections && finding.additional_sections.length > 0) {
for (const section of finding.additional_sections) {
lines.push(renderAdditionalSection(section));
lines.push('');
}
}
return lines.join('\n').trimEnd();
}
// ============================================================================
// CATEGORY GROUPING
// ============================================================================
const CATEGORY_ORDER: readonly string[] = ['Injection', 'XSS', 'Authentication', 'SSRF', 'Authorization'];
function categorySort(a: string, b: string): number {
const ai = CATEGORY_ORDER.indexOf(a);
const bi = CATEGORY_ORDER.indexOf(b);
if (ai !== -1 && bi !== -1) return ai - bi;
if (ai !== -1) return -1;
if (bi !== -1) return 1;
return a.localeCompare(b);
}
// ============================================================================
// REPORT RENDERING
// ============================================================================
export function renderReport(data: ReportData): string {
const { report_meta, findings, not_assessed = [] } = data;
const notAssessedClasses = [...new Set(not_assessed)];
const exploitEnabled = report_meta.exploit ?? true;
const sections: string[] = [];
// 1. Executive Summary
sections.push('# Security Assessment Report');
sections.push('');
sections.push('## Executive Summary');
sections.push(`- Target: ${report_meta.target}`);
sections.push(`- Assessment Date: ${report_meta.assessment_date}`);
sections.push(`- Scope: ${report_meta.scope}`);
sections.push(`- Exploitation: ${exploitEnabled ? 'enabled' : 'disabled'}`);
if (report_meta.model) {
sections.push(`- Model: ${report_meta.model}`);
}
sections.push('');
sections.push(report_meta.executive_summary);
sections.push('');
if (!exploitEnabled) {
sections.push(ANALYSIS_ONLY_DISCLAIMER);
sections.push('');
}
if (findings.length === 0) {
if (notAssessedClasses.length > 0) {
// Some classes were not assessed — a blanket "no vulnerabilities" statement would be a false
// clean bill of health. Scope the clean statement to assessed classes and list the gaps.
sections.push('No vulnerabilities were identified in the classes that were assessed.');
sections.push('');
sections.push(renderNotAssessedSection(notAssessedClasses));
} else {
sections.push('No vulnerabilities were identified during this assessment.');
}
return sections.join('\n').trimEnd() + '\n';
}
if (notAssessedClasses.length > 0) {
sections.push(renderNotAssessedSection(notAssessedClasses));
sections.push('');
}
// 2. Summary by Vulnerability Type
const byCategory = new Map<string, AddFindingInput[]>();
for (const f of findings) {
const list = byCategory.get(f.category) ?? [];
list.push(f);
byCategory.set(f.category, list);
}
const sortedCategories = [...byCategory.keys()].sort(categorySort);
sections.push('## Summary by Vulnerability Type');
sections.push('');
for (const cat of sortedCategories) {
const catFindings = byCategory.get(cat)!;
sections.push(`### ${cat}`);
sections.push('');
for (const f of catFindings) {
// Both ratings when the mode produced both. Confidence is labelled so it is never
// read as a severity in the position where a severity usually sits.
const ratings: string[] = [];
if (f.severity) {
ratings.push(titleCase(f.severity));
}
if (f.confidence) {
ratings.push(`${titleCase(f.confidence)} confidence`);
}
const suffix = ratings.length > 0 ? ` (${ratings.join(', ')})` : '';
sections.push(`- **${f.finding_id}:** ${f.title}${suffix}`);
}
sections.push('');
}
// 3. Per-category finding sections
const subheading = exploitEnabled ? 'Successfully Exploited Vulnerabilities' : 'Identified Vulnerabilities';
const heading = exploitEnabled ? 'Exploitation Evidence' : 'Findings';
for (const cat of sortedCategories) {
const catFindings = byCategory.get(cat)!;
sections.push(`# ${cat} ${heading}`);
sections.push('');
sections.push(`## ${subheading}`);
sections.push('');
for (const f of catFindings) {
sections.push(renderFinding(f, exploitEnabled));
sections.push('');
}
}
return sections.join('\n').trimEnd() + '\n';
}
+11 -40
View File
@@ -5,15 +5,7 @@
// as published by the Free Software Foundation.
import { fs, path } from 'zx';
import {
ASSEMBLED_REPORT_FILENAME,
ASSEMBLED_REPORT_PDF_FILENAME,
deliverablesDir,
FINAL_REPORT_MD_FILENAME,
FINAL_REPORT_PDF_FILENAME,
resolveSessionJsonPath,
SARIF_FILENAME,
} from '../paths.js';
import { ASSEMBLED_REPORT_FILENAME, deliverablesDir, FINAL_REPORT_FILENAME, resolveSessionJsonPath } from '../paths.js';
import type { ActivityLogger } from '../types/activity-logger.js';
import { ErrorCode } from '../types/errors.js';
import { PentestError } from './error-handling.js';
@@ -174,14 +166,9 @@ export async function injectModelIntoReport(
}
/**
* Surface the run's deliverables at the run directory's top level, so a customer opening the run
* folder sees the report without digging through internals. Sources stay in the deliverables dir
* (git-checkpointed, used by resume). Both the PDF and the markdown report are surfaced here as the
* customer-facing copies.
*
* The SARIF log is surfaced beside it when present, since a CI step consuming it needs a stable
* path and cannot be expected to reach into the internals directory. It is absent whenever the
* run was analysis-only or `report.sarif` was not enabled.
* Surface the assembled report at the run directory's top level as the single
* human-facing deliverable, so a customer opening the run folder sees only the
* report. The source stays in the deliverables dir (git-checkpointed, used by resume).
*/
export async function copyReportToRunRoot(
repoPath: string,
@@ -189,30 +176,14 @@ export async function copyReportToRunRoot(
runDir: string,
logger: ActivityLogger,
): Promise<void> {
const dir = deliverablesDir(repoPath, deliverablesSubdir);
const source = path.join(deliverablesDir(repoPath, deliverablesSubdir), ASSEMBLED_REPORT_FILENAME);
const pdfSource = path.join(dir, ASSEMBLED_REPORT_PDF_FILENAME);
if (await fs.pathExists(pdfSource)) {
const destination = path.join(runDir, FINAL_REPORT_PDF_FILENAME);
await fs.copy(pdfSource, destination, { overwrite: true });
logger.info(`Surfaced PDF report at ${destination}`);
} else {
logger.warn(`PDF report not found, skipping ${FINAL_REPORT_PDF_FILENAME}`);
if (!(await fs.pathExists(source))) {
logger.warn(`Final report not found, skipping ${FINAL_REPORT_FILENAME}`);
return;
}
const markdownSource = path.join(dir, ASSEMBLED_REPORT_FILENAME);
if (await fs.pathExists(markdownSource)) {
const destination = path.join(runDir, FINAL_REPORT_MD_FILENAME);
await fs.copy(markdownSource, destination, { overwrite: true });
logger.info(`Surfaced markdown report at ${destination}`);
} else {
logger.warn(`Markdown report not found, skipping ${FINAL_REPORT_MD_FILENAME}`);
}
const sarifSource = path.join(dir, SARIF_FILENAME);
if (await fs.pathExists(sarifSource)) {
const sarifDestination = path.join(runDir, SARIF_FILENAME);
await fs.copy(sarifSource, sarifDestination, { overwrite: true });
logger.info(`Surfaced SARIF log at ${sarifDestination}`);
}
const destination = path.join(runDir, FINAL_REPORT_FILENAME);
await fs.copy(source, destination, { overwrite: true });
logger.info(`Surfaced report at ${destination}`);
}
-293
View File
@@ -1,293 +0,0 @@
// Copyright (C) 2025 Keygraph, Inc.
//
// This program is free software: you can redistribute it and/or modify
// it under the terms of the GNU Affero General Public License version 3
// as published by the Free Software Foundation.
/** Deterministic report.json to SARIF 2.1.0 renderer, for `exploit=true` runs only. */
import type { AddFindingInput, CodeLocation } from '../collectors/finding-collector.js';
import type { ReportData } from './report-renderer.js';
export interface SarifOptions {
readonly workspaceName: string;
}
interface SarifRule {
readonly id: string;
readonly name: string;
readonly shortDescription: { text: string };
readonly fullDescription: { text: string };
readonly help: { text: string };
readonly properties: { tags: string[] };
}
const TOOL_NAME = 'Shannon';
const TOOL_URI = 'https://github.com/KeygraphHQ/shannon';
/** Taxonomy identity. A reference resolves the component by name, so this must not be reworded. */
const OWASP_TAXONOMY_NAME = 'OWASP Top Ten 2025';
/**
* One rule per vulnerability class, keyed by `finding.category`.
*
* Rule IDs are the unit of alert grouping: renaming one detaches every alert filed under it.
* `fullDescription` and `help` describe the class, never the instance, and GitHub requires the
* `text` of both.
*/
const RULES: Record<string, SarifRule> = {
Injection: {
id: 'shannon/injection',
name: 'Injection',
shortDescription: { text: 'Injection' },
fullDescription: {
text: 'Untrusted input reaches an interpreter sink (SQL, OS command, template, file path or deserializer) at a position where it can alter the structure of the statement rather than only supply data.',
},
help: {
text: 'Separate code from data at the sink: bind SQL parameters, pass command arguments as an array, and allowlist file paths. Escaping is a weaker control than parameterisation and breaks whenever the sink context changes.',
},
properties: { tags: ['security', 'shannon'] },
},
XSS: {
id: 'shannon/xss',
name: 'Cross-Site Scripting',
shortDescription: { text: 'Cross-Site Scripting' },
fullDescription: {
text: 'Untrusted input reaches a browser rendering context without the encoding that context requires.',
},
help: {
text: 'Encode at the point of output for the specific context (HTML body, attribute, URL, script or style); no single encoder is correct for all of them. Prefer APIs that treat input as text, such as textContent over innerHTML.',
},
properties: { tags: ['security', 'shannon'] },
},
Authentication: {
id: 'shannon/auth',
name: 'Authentication',
shortDescription: { text: 'Authentication' },
fullDescription: {
text: 'A weakness in credential verification or session lifecycle that lets an attacker assume another identity or retain access they should have lost.',
},
help: {
text: 'Issue a fresh session identifier on every privilege change, set HttpOnly, Secure and SameSite on session cookies, rate-limit credential endpoints, and verify the signature and algorithm of externally issued tokens.',
},
properties: { tags: ['security', 'shannon'] },
},
Authorization: {
id: 'shannon/authz',
name: 'Authorization',
shortDescription: { text: 'Authorization' },
fullDescription: {
text: 'An access control decision is missing, evaluated in the client, or applied at the wrong layer, letting a caller act on resources they do not own.',
},
help: {
text: 'Check ownership and role on the server for every object reference, and enforce it in the data-access layer rather than per route, denying by default. An unguessable identifier is not an access control.',
},
properties: { tags: ['security', 'shannon'] },
},
SSRF: {
id: 'shannon/ssrf',
name: 'Server-Side Request Forgery',
shortDescription: { text: 'Server-Side Request Forgery' },
fullDescription: {
text: 'A server-side request takes its destination from untrusted input, letting an attacker reach hosts the server can see but they cannot.',
},
help: {
text: 'Allowlist destination hosts and schemes, resolve DNS before validating the address so rebinding cannot slip through, and block loopback, private and link-local ranges including cloud metadata. Do not follow redirects.',
},
properties: { tags: ['security', 'shannon'] },
},
};
const CATEGORY_ORDER: readonly string[] = ['Injection', 'XSS', 'Authentication', 'SSRF', 'Authorization'];
/**
* Five severities collapse into SARIF's three usable levels, so `critical` and `high` are
* indistinguishable. `security-severity` would separate them but lives on the rule, which would
* flatten every finding of a class to one score instead.
*/
function severityToLevel(severity: string | undefined): string {
switch (severity) {
case 'critical':
case 'high':
return 'error';
case 'medium':
return 'warning';
default:
return 'note';
}
}
function toPhysicalLocation(location: CodeLocation) {
const region: Record<string, number> = {};
if (location.start_line) region.startLine = location.start_line;
if (location.end_line) region.endLine = location.end_line;
return {
physicalLocation: {
artifactLocation: { uri: location.file },
...(Object.keys(region).length > 0 && { region }),
},
...(location.symbol && { logicalLocations: [{ name: location.symbol, kind: 'function' }] }),
message: { text: location.role },
};
}
/**
* Fall back to the HTTP entry point when a finding names no file: a result with no location is
* silently discarded downstream. No `uriBaseId`, since the path does not resolve in the repo.
*/
function syntheticLocationFromHttp(finding: AddFindingInput) {
if (!finding.http_location) return undefined;
let uri = finding.http_location.url;
try {
const parsed = new URL(finding.http_location.url);
uri = `${parsed.pathname}${parsed.hash}`;
} catch {}
return {
physicalLocation: { artifactLocation: { uri } },
message: { text: `${finding.http_location.method} ${finding.http_location.url}` },
};
}
function buildMessageMarkdown(finding: AddFindingInput): string {
const parts = [`**${finding.title}**`, '', finding.overview, '', '**Impact**', '', finding.impact];
parts.push('', '**Remediation**', '', finding.remediation);
// Exploitation steps and proof of impact are deliberately absent: SARIF has no structural home
// for them, and flattening them into prose would imply this file carries the evidence.
parts.push('', 'Full exploitation evidence: `Security-Assessment-Report.pdf`');
return parts.join('\n');
}
/**
* `owasp_category` is one label, `A05:2025 <separator> Injection`; SARIF wants the id and the name
* as separate fields. The enum in ../collectors/finding-collector.ts fixes the shape, so the
* separator is dropped by position rather than matched.
*/
function splitOwaspCategory(label: string): { id: string; name: string } {
const [id, , ...nameParts] = label.split(' ');
return { id: id ?? label, name: nameParts.join(' ') };
}
interface RenderedResult {
readonly result: Record<string, unknown>;
readonly category: string;
readonly owaspId: string;
}
function renderResult(finding: AddFindingInput, ruleId: string): RenderedResult | null {
const codeLocations = finding.code_locations ?? [];
const sinks = codeLocations.filter((l) => l.role === 'sink');
const related = codeLocations.filter((l) => l.role !== 'sink');
const primary = sinks[0] ?? codeLocations[0];
const locations = primary ? [toPhysicalLocation(primary)] : [syntheticLocationFromHttp(finding)].filter(Boolean);
if (locations.length === 0) return null;
const properties: Record<string, unknown> = { findingId: finding.finding_id };
if (finding.http_location?.parameter) properties.parameter = finding.http_location.parameter;
if (finding.status) properties.status = finding.status;
if (finding.auth_state) properties.authState = finding.auth_state;
if (finding.prerequisites) properties.prerequisites = finding.prerequisites;
const owaspId = splitOwaspCategory(finding.owasp_category).id;
return {
category: finding.category,
owaspId,
result: {
ruleId,
level: severityToLevel(finding.severity),
message: {
text: `${finding.title}. ${finding.overview}`,
markdown: buildMessageMarkdown(finding),
},
locations,
...(related.length > 0 && {
relatedLocations: related.map((l, i) => ({ id: i + 1, ...toPhysicalLocation(l) })),
}),
...(finding.http_location && {
// No `parameters`: SARIF wants a name-to-value map and the deliverable names only the
// parameter, so any value here would be invented. It travels in `properties` instead.
webRequest: { method: finding.http_location.method, target: finding.http_location.url },
}),
taxa: [
{
id: owaspId,
toolComponent: { name: OWASP_TAXONOMY_NAME },
},
],
properties,
},
};
}
/** Render a SARIF 2.1.0 log from the structured report. Findings with no location are omitted. */
export function renderSarif(data: ReportData, options: SarifOptions): string {
const { report_meta, findings, not_assessed = [] } = data;
const rendered: RenderedResult[] = [];
for (const finding of findings) {
const rule = RULES[finding.category];
if (!rule) continue;
const result = renderResult(finding, rule.id);
if (result !== null) rendered.push(result);
}
// Only classes that produced a result are declared, and `ruleIndex` is the position in this list.
const usedRules = CATEGORY_ORDER.flatMap((category) => {
const rule = RULES[category];
if (!rule || !rendered.some((r) => r.category === category)) return [];
return [{ category, rule }];
});
const rules = usedRules.map((u) => u.rule);
const results: Record<string, unknown>[] = usedRules.flatMap(({ category }, ruleIndex) =>
rendered.filter((r) => r.category === category).map((r) => ({ ...r.result, ruleIndex })),
);
const owaspCategories = [...new Set(findings.map((f) => f.owasp_category))]
.map(splitOwaspCategory)
.filter((c) => rendered.some((r) => r.owaspId === c.id))
.sort((a, b) => a.id.localeCompare(b.id));
const log = {
$schema: 'https://json.schemastore.org/sarif-2.1.0.json',
version: '2.1.0',
runs: [
{
tool: {
driver: {
name: TOOL_NAME,
informationUri: TOOL_URI,
rules,
},
},
// Scoped to the exploit pipeline: an analysis run of the same target has a different
// finding population, which would read as alerts resolved.
automationDetails: { id: `shannon/exploit/${options.workspaceName}` },
invocations: [
{
// A failed class produced no results; reporting success would read as resolved alerts.
executionSuccessful: not_assessed.length === 0,
},
],
...(owaspCategories.length > 0 && {
taxonomies: [
{
name: OWASP_TAXONOMY_NAME,
organization: 'OWASP',
informationUri: 'https://owasp.org/Top10/',
shortDescription: { text: 'OWASP Top Ten 2025 categories.' },
taxa: owaspCategories.map((c) => ({ id: c.id, name: c.name })),
},
],
}),
results,
properties: { target: report_meta.target, assessmentDate: report_meta.assessment_date },
},
],
};
return `${JSON.stringify(log, null, 2)}\n`;
}
@@ -23,7 +23,6 @@ import type { ActivityLogger } from '../types/activity-logger.js';
import type { AgentEndResult } from '../types/audit.js';
import type { DistributedConfig } from '../types/config.js';
import { ErrorCode } from '../types/errors.js';
import type { AgentMetrics } from '../types/metrics.js';
import { err, ok, type Result } from '../types/result.js';
import { PentestError } from './error-handling.js';
import { loadPrompt } from './prompt-manager.js';
@@ -98,9 +97,7 @@ export interface ValidateAuthInput {
readonly cancellationSignal?: AbortSignal;
}
export async function validateAuthentication(
input: ValidateAuthInput,
): Promise<Result<AgentMetrics | null, PentestError>> {
export async function validateAuthentication(input: ValidateAuthInput): Promise<Result<void, PentestError>> {
const {
distributedConfig,
repoPath,
@@ -116,7 +113,7 @@ export async function validateAuthentication(
const authentication = distributedConfig.authentication;
if (!authentication) {
return ok(null);
return ok(undefined);
}
logger.info('Validating authentication credentials with live browser...', {
@@ -148,6 +145,7 @@ export async function validateAuthentication(
AGENT_NAME,
auditSession,
logger,
'medium',
undefined, // callerTools
deliverablesSubdir,
cancellationSignal,
@@ -163,10 +161,9 @@ export async function validateAuthentication(
}
}
const durationMs = Date.now() - startTime;
const endResult: AgentEndResult = {
attemptNumber,
duration_ms: durationMs,
duration_ms: Date.now() - startTime,
cost_usd: result.cost || 0,
success: classification.ok,
...(result.model !== undefined && { model: result.model }),
@@ -174,21 +171,7 @@ export async function validateAuthentication(
};
await auditSession.endAgent(AGENT_NAME, endResult);
if (!classification.ok) {
return err(classification.error);
}
const metrics: AgentMetrics = {
durationMs,
inputTokens: result.inputTokens ?? null,
outputTokens: result.outputTokens ?? null,
cacheReadTokens: result.cacheReadTokens ?? null,
cacheWriteTokens: result.cacheWriteTokens ?? null,
costUsd: result.cost ?? null,
numTurns: result.turns ?? null,
...(result.model !== undefined && { model: result.model }),
};
return ok(metrics);
return classification;
}
async function verifySavedAuthState(stateFile: string, logger: ActivityLogger): Promise<Result<void, PentestError>> {
@@ -223,32 +206,28 @@ async function verifySavedAuthState(stateFile: string, logger: ActivityLogger):
);
}
const cookies = storageEntries(parsed, 'cookies');
const origins = storageEntries(parsed, 'origins');
if (!cookies || !origins) {
const cookieCount = countStorageEntries(parsed, 'cookies');
const originCount = countStorageEntries(parsed, 'origins');
if (cookieCount === 0 && originCount === 0) {
return err(
new PentestError(
`Preflight saved an authenticated session to ${stateFile}, but it is not a storage state — cookies and origins arrays are missing.`,
`Preflight saved an authenticated session to ${stateFile}, but it contains no cookies or origins — the browser was not actually logged in.`,
'validation',
true,
{ stateFile, hasCookies: !!cookies, hasOrigins: !!origins },
{ stateFile, cookieCount, originCount },
ErrorCode.AGENT_EXECUTION_FAILED,
),
);
}
logger.info('Preflight authenticated session saved', {
stateFile,
cookieCount: cookies.length,
originCount: origins.length,
});
logger.info('Preflight authenticated session saved', { stateFile, cookieCount, originCount });
return ok(undefined);
}
function storageEntries(parsed: unknown, key: 'cookies' | 'origins'): unknown[] | null {
if (typeof parsed !== 'object' || parsed === null) return null;
function countStorageEntries(parsed: unknown, key: 'cookies' | 'origins'): number {
if (typeof parsed !== 'object' || parsed === null) return 0;
const value = (parsed as Record<string, unknown>)[key];
return Array.isArray(value) ? value : null;
return Array.isArray(value) ? value.length : 0;
}
function classifyResult(
+1
View File
@@ -17,6 +17,7 @@ export const AGENTS: Readonly<Record<AgentName, AgentDefinition>> = Object.freez
prerequisites: [],
promptTemplate: 'pre-recon-code',
deliverableFilename: 'pre_recon_deliverable.md',
modelTier: 'large',
},
recon: {
name: 'recon',
+14 -152
View File
@@ -25,16 +25,7 @@ import type { ResumeAttempt } from '../audit/metrics-tracker.js';
import { authStateFile, generateAuditPath, generateSessionJsonPath, type SessionMetadata } from '../audit/utils.js';
import type { WorkflowSummary } from '../audit/workflow-logger.js';
import type { CheckpointContext } from '../interfaces/checkpoint-provider.js';
import {
ASSEMBLED_REPORT_FILENAME,
ASSEMBLED_REPORT_PDF_FILENAME,
DEFAULT_DELIVERABLES_SUBDIR,
deliverablesDir,
REPORT_JSON_FILENAME,
resolveSessionJsonPath,
SARIF_FILENAME,
TYPST_TEMPLATE,
} from '../paths.js';
import { DEFAULT_DELIVERABLES_SUBDIR, deliverablesDir, resolveSessionJsonPath } from '../paths.js';
import { getAgentGitPaths } from '../services/agent-git-paths.js';
import { getContainer, getOrCreateContainer, removeContainer } from '../services/container.js';
import { classifyErrorForTemporal, PentestError } from '../services/error-handling.js';
@@ -43,7 +34,6 @@ import { renderFindingsFromQueues } from '../services/findings-renderer.js';
import { executeGitCommandWithRetry } from '../services/git-manager.js';
import { runPreflightChecks } from '../services/preflight.js';
import type { ExploitationDecision, VulnType } from '../services/queue-validation.js';
import type { ReportData, ReportMeta } from '../services/report-renderer.js';
import { assembleFinalReport, copyReportToRunRoot, injectModelIntoReport } from '../services/reporting.js';
import { validateAuthentication } from '../services/validate-authentication.js';
import { AGENTS } from '../session-manager.js';
@@ -86,10 +76,6 @@ export interface ActivityInput {
auditDir?: string;
promptDir?: string;
sastSarifPath?: string;
// Vuln classes whose pipeline failed. Set before the report stage on a partial run so the
// report marks them "not assessed" instead of asserting no findings were present.
failedClasses?: VulnClass[];
}
/**
@@ -201,7 +187,6 @@ async function runAgentActivity(
attemptNumber,
...(input.promptDir !== undefined && { promptDir: input.promptDir }),
...(input.configYAML !== undefined && { configYAML: input.configYAML }),
...(input.failedClasses !== undefined && { failedClasses: input.failedClasses }),
...(customTools && { customTools }),
...(writeDeliverable && { writeDeliverable }),
cancellationSignal: Context.current().cancellationSignal,
@@ -213,12 +198,10 @@ async function runAgentActivity(
// 4. Return metrics
return {
durationMs: Date.now() - startTime,
inputTokens: endResult.input_tokens ?? null,
outputTokens: endResult.output_tokens ?? null,
cacheReadTokens: endResult.cache_read_tokens ?? null,
cacheWriteTokens: endResult.cache_write_tokens ?? null,
inputTokens: null,
outputTokens: null,
costUsd: endResult.cost_usd,
numTurns: endResult.turns ?? null,
numTurns: null,
model: endResult.model,
};
} catch (error) {
@@ -449,120 +432,8 @@ export async function runAuthzExploitAgent(input: ActivityInput): Promise<AgentM
return runExploitAgentWithCollector('authz-exploit', 'authz', input);
}
/**
* Write report.sarif when the run is exploitative and the operator asked for it.
*
* Skipped entirely for analysis-only runs: those findings carry no severity, so every
* `result.level` would be invented. Failures are logged and swallowed — the SARIF log is a
* secondary artifact and must not fail a run whose report is already written.
*/
async function writeSarifIfEnabled(
input: ActivityInput,
exploit: boolean,
reportData: ReportData,
deliverablesPath: string,
logger: ReturnType<typeof createActivityLogger>,
): Promise<void> {
if (!exploit) return;
const container = getOrCreateContainer(input.workflowId, buildSessionMetadata(input), buildContainerConfig(input));
const configResult = await container.configLoader.loadOptional(input.configPath, undefined, input.configYAML);
if (isErr(configResult) || configResult.value?.report?.sarif !== true) return;
try {
const { renderSarif } = await import('../services/sarif-renderer.js');
const sarif = renderSarif(reportData, { workspaceName: input.sessionId });
await atomicWrite(path.join(deliverablesPath, SARIF_FILENAME), sarif);
logger.info(`Wrote ${SARIF_FILENAME}`);
} catch (error) {
logger.warn(`Failed to write ${SARIF_FILENAME}: ${(error as Error).message}`);
}
}
/**
* Compile the PDF report from the assembled report data.
*
* Failures are logged and swallowed — the PDF is a secondary artifact and must not fail a run
* whose report is already written.
*/
async function writePdfReport(
reportData: ReportData,
deliverablesPath: string,
logger: ReturnType<typeof createActivityLogger>,
): Promise<void> {
try {
const { renderReportPdf } = await import('../services/pdf-renderer.js');
await renderReportPdf({
reportData,
templatePath: TYPST_TEMPLATE,
outputPath: path.join(deliverablesPath, ASSEMBLED_REPORT_PDF_FILENAME),
});
logger.info(`Wrote ${ASSEMBLED_REPORT_PDF_FILENAME}`);
} catch (error) {
logger.warn(`Failed to write ${ASSEMBLED_REPORT_PDF_FILENAME}: ${(error as Error).message}`);
}
}
export async function runReportAgent(input: ActivityInput, exploit: boolean): Promise<AgentMetrics> {
const { createFindingCollector } = await import('../collectors/finding-collector.js');
const { renderReport } = await import('../services/report-renderer.js');
const collector = createFindingCollector(exploit);
const writeDeliverable = async (deliverablesPath: string): Promise<void> => {
const logger = createActivityLogger();
const { attachQueueCodeLocations } = await import('../services/code-location-join.js');
const collected = collector.getAll();
logger.info(`Collected ${collected.length} finding(s) from report agent`);
const findings = await attachQueueCodeLocations(collected, deliverablesPath, logger);
// report_meta is written by the set-report-meta CLI while the agent runs; read it back so
// the two halves of report.json end up in one document.
const reportJsonPath = path.join(deliverablesPath, REPORT_JSON_FILENAME);
let reportMeta: ReportMeta = {
target: input.webUrl,
assessment_date: new Date().toISOString().split('T')[0]!,
scope: '',
executive_summary: '',
exploit,
};
if (await fileExists(reportJsonPath)) {
try {
const existing = await readJson<{ report_meta?: Record<string, unknown> }>(reportJsonPath);
if (existing.report_meta) {
reportMeta = {
target: String(existing.report_meta.target ?? input.webUrl),
assessment_date: String(existing.report_meta.assessment_date ?? reportMeta.assessment_date),
scope: String(existing.report_meta.scope ?? ''),
executive_summary: String(existing.report_meta.executive_summary ?? ''),
// Run scope, not agent output — keeps the rendered report and the schema the agent
// was given in agreement.
exploit,
...(existing.report_meta.model !== undefined && { model: String(existing.report_meta.model) }),
};
}
} catch {
logger.warn('Failed to read report_meta from report.json, using defaults');
}
}
const reportData: ReportData = {
report_meta: reportMeta,
findings,
...(input.failedClasses && input.failedClasses.length > 0 && { not_assessed: input.failedClasses }),
};
await atomicWrite(reportJsonPath, JSON.stringify(reportData, null, 2));
logger.info(`Wrote ${REPORT_JSON_FILENAME} with ${findings.length} finding(s)`);
await atomicWrite(path.join(deliverablesPath, ASSEMBLED_REPORT_FILENAME), renderReport(reportData));
logger.info(`Wrote ${ASSEMBLED_REPORT_FILENAME} from structured data`);
await writePdfReport(reportData, deliverablesPath, logger);
await writeSarifIfEnabled(input, exploit, reportData, deliverablesPath, logger);
};
return runAgentActivity('report', input, collector.tools, writeDeliverable);
export async function runReportAgent(input: ActivityInput): Promise<AgentMetrics> {
return runAgentActivity('report', input);
}
/**
@@ -637,7 +508,7 @@ export async function runPreflightValidation(input: ActivityInput): Promise<void
* block; otherwise surfaces a classified failure (failurePoint +
* failureDetail in ApplicationFailure.details) on credential rejection.
*/
export async function runAuthenticationValidation(input: ActivityInput): Promise<AgentMetrics | null> {
export async function runAuthenticationValidation(input: ActivityInput): Promise<void> {
const startTime = Date.now();
const attemptNumber = Context.current().info.attempt;
@@ -655,13 +526,13 @@ export async function runAuthenticationValidation(input: ActivityInput): Promise
if (isErr(configResult)) {
// runPreflightValidation already validated parsing, so this is unexpected.
logger.warn(`runAuthenticationValidation: config load failed unexpectedly: ${configResult.error.message}`);
return null;
return;
}
const distributedConfig = configResult.value;
if (!distributedConfig?.authentication) {
logger.info('No authentication configured — skipping credential validation');
return null;
return;
}
const auditSession = new AuditSession(sessionMetadata);
@@ -700,8 +571,6 @@ export async function runAuthenticationValidation(input: ActivityInput): Promise
truncateStackTrace(failure);
throw failure;
}
return result.value;
} catch (error) {
if (error instanceof ApplicationFailure) {
throw error;
@@ -1140,19 +1009,9 @@ export async function restoreGitCheckpoint(
/**
* Record a resume attempt in session.json and write resume header to workflow.log.
*/
/**
* Register this resume's workflow id in session.json before loadResumeState (which can throw),
* so the CLI can resolve and follow the resume even when validation fails instead of timing out.
*/
export async function registerResumeAttempt(input: ActivityInput, terminatedWorkflows: string[]): Promise<void> {
const sessionMetadata = buildSessionMetadata(input);
const auditSession = new AuditSession(sessionMetadata);
await auditSession.initialize();
await auditSession.addResumeAttempt(input.workflowId, terminatedWorkflows);
}
export async function recordResumeAttempt(
input: ActivityInput,
terminatedWorkflows: string[],
checkpointHash: string,
previousWorkflowId: string,
completedAgents: string[],
@@ -1161,7 +1020,10 @@ export async function recordResumeAttempt(
const auditSession = new AuditSession(sessionMetadata);
await auditSession.initialize();
// session.json entry already added by registerResumeAttempt; here we only write the workflow.log header.
// Update session.json with resume attempt
await auditSession.addResumeAttempt(input.workflowId, terminatedWorkflows, checkpointHash);
// Write resume header to workflow.log
await auditSession.logResumeHeader({
previousWorkflowId,
newWorkflowId: input.workflowId,
+2 -1
View File
@@ -2,7 +2,7 @@ import { defineQuery } from '@temporalio/workflow';
export type { AgentMetrics } from '../types/metrics.js';
import type { DistributedConfig, VulnClass } from '../types/config.js';
import type { DistributedConfig, PipelineConfig, VulnClass } from '../types/config.js';
import type { ErrorCode } from '../types/errors.js';
import type { AgentMetrics } from '../types/metrics.js';
@@ -12,6 +12,7 @@ export interface PipelineInput {
configPath?: string;
outputPath?: string;
pipelineTestingMode?: boolean;
pipelineConfig?: PipelineConfig;
workflowId?: string; // Used for audit correlation
sessionId?: string; // Workspace directory name (distinct from workflowId for named workspaces)
resumeFromWorkspace?: string; // Workspace name to resume from
+17 -11
View File
@@ -35,13 +35,8 @@ import { bundleWorkflowCode, NativeConnection, Worker } from '@temporalio/worker
import dotenv from 'dotenv';
import { sanitizeHostname } from '../audit/utils.js';
import { parseConfig } from '../config-parser.js';
import {
ASSEMBLED_REPORT_PDF_FILENAME,
deliverablesDir,
FINAL_REPORT_PDF_FILENAME,
resolveSessionJsonPath,
} from '../paths.js';
import type { VulnClass } from '../types/config.js';
import { ASSEMBLED_REPORT_FILENAME, deliverablesDir, FINAL_REPORT_FILENAME, resolveSessionJsonPath } from '../paths.js';
import type { PipelineConfig, VulnClass } from '../types/config.js';
import { fileExists, readJson } from '../utils/file-io.js';
import * as activities from './activities.js';
import type { PipelineInput, PipelineProgress, PipelineState } from './shared.js';
@@ -281,16 +276,26 @@ async function resolveWorkspace(client: Client, args: CliArgs): Promise<Workspac
// === Pipeline Input Construction ===
interface OrchestrationConfig {
pipelineConfig: PipelineConfig;
vulnClasses?: VulnClass[];
exploit?: boolean;
}
async function loadOrchestrationConfig(configPath: string | undefined): Promise<OrchestrationConfig> {
if (!configPath) return {};
if (!configPath) return { pipelineConfig: {} };
try {
const config = await parseConfig(configPath);
const pipelineConfig: PipelineConfig = {};
if (config.pipeline?.retry_preset !== undefined) {
pipelineConfig.retry_preset = config.pipeline.retry_preset;
}
if (config.pipeline?.max_concurrent_pipelines !== undefined) {
pipelineConfig.max_concurrent_pipelines = Number(config.pipeline.max_concurrent_pipelines);
}
return {
pipelineConfig,
...(config.vuln_classes && config.vuln_classes.length > 0 && { vulnClasses: [...config.vuln_classes] }),
...(config.exploit !== undefined && { exploit: config.exploit === 'true' }),
};
@@ -317,6 +322,7 @@ function buildPipelineInput(
...(args.pipelineTestingMode && { pipelineTestingMode: args.pipelineTestingMode }),
...(workspace.isResume && args.resumeFromWorkspace && { resumeFromWorkspace: args.resumeFromWorkspace }),
...(workspace.terminatedWorkflows.length > 0 && { terminatedWorkflows: workspace.terminatedWorkflows }),
...(Object.keys(orchestration.pipelineConfig).length > 0 && { pipelineConfig: orchestration.pipelineConfig }),
...(orchestration.vulnClasses && { vulnClasses: orchestration.vulnClasses }),
...(orchestration.exploit !== undefined && { exploit: orchestration.exploit }),
};
@@ -394,9 +400,9 @@ function copyDeliverables(repoPath: string, outputPath: string): void {
}
// Surface the report under its human-facing name alongside the raw deliverables
const assembledPdf = path.join(outputDir, ASSEMBLED_REPORT_PDF_FILENAME);
if (fs.existsSync(assembledPdf)) {
fs.copyFileSync(assembledPdf, path.join(outputPath, FINAL_REPORT_PDF_FILENAME));
const assembledReport = path.join(outputDir, ASSEMBLED_REPORT_FILENAME);
if (fs.existsSync(assembledReport)) {
fs.copyFileSync(assembledReport, path.join(outputPath, FINAL_REPORT_FILENAME));
}
console.log(`Copied ${files.length} deliverable(s) to ${outputPath}`);
+6 -1
View File
@@ -21,6 +21,8 @@ import { ErrorCode } from '../types/errors.js';
*/
const ERROR_TYPE_TO_CODE: Record<string, ErrorCode> = {
AuthenticationError: ErrorCode.AUTH_FAILED,
BillingError: ErrorCode.BILLING_ERROR,
RateLimitError: ErrorCode.API_RATE_LIMITED,
ConfigurationError: ErrorCode.CONFIG_VALIDATION_FAILED,
OutputValidationError: ErrorCode.OUTPUT_VALIDATION_FAILED,
AgentExecutionError: ErrorCode.AGENT_EXECUTION_FAILED,
@@ -42,10 +44,13 @@ export function classifyErrorCode(error: unknown): ErrorCode | undefined {
/** Maps Temporal error type strings to actionable remediation hints. */
const REMEDIATION_HINTS: Record<string, string> = {
AuthenticationError: "Verify the selected provider's API key is valid and not expired.",
AuthenticationError: 'Verify ANTHROPIC_API_KEY or CLAUDE_CODE_OAUTH_TOKEN in .env is valid and not expired.',
ConfigurationError: 'Check your CONFIG file path and contents.',
BillingError: 'Check your Anthropic billing dashboard. Add credits or wait for spending cap reset.',
GitError: 'Check repository path and git state.',
InvalidTargetError: 'Verify the target URL is correct and accessible.',
PermissionError: 'Check file and network permissions.',
ExecutionLimitError: 'Agent exceeded maximum turns or budget. Review prompt complexity.',
};
/**
+37 -34
View File
@@ -17,14 +17,13 @@
*
* Features:
* - Queryable state via getProgress
* - Automatic retry with backoff for transient errors
* - Automatic retry with backoff for transient/billing errors
* - Non-retryable classification for permanent errors
* - Audit correlation via workflowId
* - Graceful failure handling: pipelines continue if one fails
*/
import {
ActivityCancellationType,
ApplicationFailure,
CancellationScope,
isCancellation,
@@ -65,21 +64,21 @@ function computeExpectedAgents(vulnClasses: readonly VulnClass[], exploit: boole
return expected;
}
// Retry configuration for production (long intervals so a rate-limit window can clear)
// Retry configuration for production (long intervals for billing recovery)
const PRODUCTION_RETRY = {
initialInterval: '5 minutes',
maximumInterval: '30 minutes',
backoffCoefficient: 2,
maximumAttempts: 50,
// Belt-and-braces: activities already throw non-retryable ApplicationFailures for
// these. Only types that are always permanent belong here — GitError and
// AgentExecutionError carry a per-error verdict and must not be listed.
nonRetryableErrorTypes: [
'AuthenticationError',
'PermissionError',
'InvalidRequestError',
'RequestTooLargeError',
'ConfigurationError',
'InvalidTargetError',
'ExecutionLimitError',
'AuthLoginFailedError',
'PermanentError',
],
};
@@ -97,8 +96,6 @@ const acts = proxyActivities<typeof activities>({
startToCloseTimeout: '2 hours',
heartbeatTimeout: '60 minutes', // Extended for nested pi task execution
retry: PRODUCTION_RETRY,
// Cancel promptly instead of waiting out startToCloseTimeout; the agent aborts on the signal.
cancellationType: ActivityCancellationType.TRY_CANCEL,
});
// Activity proxy with testing retry configuration (fast)
@@ -106,7 +103,22 @@ const testActs = proxyActivities<typeof activities>({
startToCloseTimeout: '30 minutes',
heartbeatTimeout: '30 minutes', // Extended for sub-agent execution in testing
retry: TESTING_RETRY,
cancellationType: ActivityCancellationType.TRY_CANCEL,
});
// Retry configuration for subscription plans (5h+ rolling rate limit windows)
const SUBSCRIPTION_RETRY = {
initialInterval: '5 minutes',
maximumInterval: '6 hours',
backoffCoefficient: 2,
maximumAttempts: 100,
nonRetryableErrorTypes: PRODUCTION_RETRY.nonRetryableErrorTypes,
};
// Activity proxy for subscription plan recovery (extended timeouts)
const subscriptionActs = proxyActivities<typeof activities>({
startToCloseTimeout: '8 hours',
heartbeatTimeout: '2 hours',
retry: SUBSCRIPTION_RETRY,
});
// Retry configuration for preflight validation (short timeout, few retries)
@@ -123,7 +135,6 @@ const preflightActs = proxyActivities<typeof activities>({
startToCloseTimeout: '2 minutes',
heartbeatTimeout: '2 minutes',
retry: PREFLIGHT_RETRY,
cancellationType: ActivityCancellationType.TRY_CANCEL,
});
// Credential rejection is not retryable; transient provider errors get 3 attempts.
@@ -140,7 +151,6 @@ const authValidationActs = proxyActivities<typeof activities>({
startToCloseTimeout: '10 minutes',
heartbeatTimeout: '10 minutes',
retry: AUTH_VALIDATION_RETRY,
cancellationType: ActivityCancellationType.TRY_CANCEL,
});
/**
@@ -157,9 +167,6 @@ function computeSummary(state: PipelineState): PipelineSummary {
};
}
/** One pipeline per vulnerability class, all five in flight together. */
const MAX_CONCURRENT_PIPELINES = 5;
const MAX_PIPELINE_ERROR_MESSAGE_LENGTH = 2000;
function truncatePipelineErrorMessage(message: string): string {
@@ -193,7 +200,14 @@ export async function pentestPipeline(input: PipelineInput): Promise<PipelineSta
const { workflowId } = workflowInfo();
const a = input.pipelineTestingMode ? testActs : acts;
// Select activity proxy based on mode: testing (fast), subscription (extended), or default
function selectActivityProxy(pipelineInput: PipelineInput) {
if (pipelineInput.pipelineTestingMode) return testActs;
if (pipelineInput.pipelineConfig?.retry_preset === 'subscription') return subscriptionActs;
return acts;
}
const a = selectActivityProxy(input);
const state: PipelineState = {
status: 'running',
@@ -252,10 +266,6 @@ export async function pentestPipeline(input: PipelineInput): Promise<PipelineSta
let resumeState: ResumeState | null = null;
if (input.resumeFromWorkspace) {
// 0. Register the resume's workflow id in session.json before validation can fail, so the CLI
// can resolve and follow it instead of polling for an entry that never lands.
await a.registerResumeAttempt(activityInput, input.terminatedWorkflows || []);
// 1. Load resume state (validates workspace, cross-checks deliverables)
resumeState = await a.loadResumeState(
input.resumeFromWorkspace,
@@ -287,9 +297,10 @@ export async function pentestPipeline(input: PipelineInput): Promise<PipelineSta
return state;
}
// 4. Write the resume header to workflow.log (the session.json entry was recorded in step 0)
// 4. Record this resume attempt in session.json and workflow.log
await a.recordResumeAttempt(
activityInput,
input.terminatedWorkflows || [],
resumeState.checkpointHash,
resumeState.originalWorkflowId,
resumeState.completedAgents,
@@ -489,11 +500,7 @@ export async function pentestPipeline(input: PipelineInput): Promise<PipelineSta
// === Authentication Validation ===
state.currentPhase = 'auth-validation';
state.currentAgent = 'validate-authentication';
const authMetrics = await authValidationActs.runAuthenticationValidation(activityInput);
// Null when no login ran (no-auth scan); left absent so status renders it skipped, not completed.
if (authMetrics) {
state.agentMetrics['validate-authentication'] = authMetrics;
}
await authValidationActs.runAuthenticationValidation(activityInput);
state.currentAgent = null;
log.info('Authentication validation passed');
@@ -604,6 +611,8 @@ export async function pentestPipeline(input: PipelineInput): Promise<PipelineSta
}
}
const maxConcurrent = input.pipelineConfig?.max_concurrent_pipelines ?? 5;
const pipelineConfigs = buildPipelineConfigs();
const pipelineThunks: Array<() => Promise<VulnExploitPipelineResult>> = [];
let alreadyCompletedPipelineCount = 0;
@@ -623,15 +632,9 @@ export async function pentestPipeline(input: PipelineInput): Promise<PipelineSta
}
}
const pipelineResults = await runWithConcurrencyLimit(pipelineThunks, MAX_CONCURRENT_PIPELINES);
const pipelineResults = await runWithConcurrencyLimit(pipelineThunks, maxConcurrent);
aggregatePipelineResults(pipelineResults, alreadyCompletedPipelineCount);
// Surface the not-assessed classes to the report stage so a failed class renders as
// "analysis did not complete" rather than the absence assertion "no findings".
if (state.failedPipelines.length > 0) {
activityInput.failedClasses = state.failedPipelines.map((f) => f.vulnType);
}
state.currentPhase = 'exploitation';
state.currentAgent = null;
await a.logPhaseTransition(activityInput, 'vulnerability-exploitation', 'complete');
@@ -646,7 +649,7 @@ export async function pentestPipeline(input: PipelineInput): Promise<PipelineSta
await a.assembleReportActivity(activityInput, exploit);
// Then run the report agent to add executive summary and clean up
state.agentMetrics.report = await a.runReportAgent(activityInput, exploit);
state.agentMetrics.report = await a.runReportAgent(activityInput);
state.completedAgents.push('report');
if (input.checkpointsEnabled) {
await a.saveCheckpoint(activityInput, 'report', 'reporting', state);
+174
View File
@@ -0,0 +1,174 @@
#!/usr/bin/env node
// Copyright (C) 2025 Keygraph, Inc.
//
// This program is free software: you can redistribute it and/or modify
// it under the terms of the GNU Affero General Public License version 3
// as published by the Free Software Foundation.
/**
* Workspace listing tool for Shannon.
*
* Reads workspaces/ directories, parses session.json files, and displays
* a formatted table of all workspaces with status, duration, and cost.
*
* Usage:
* node dist/temporal/workspaces.js
*
* Environment:
* WORKSPACES_DIR - Override workspaces directory (default: ./workspaces)
*/
import fs from 'node:fs/promises';
import path from 'node:path';
import { WORKSPACES_DIR as DEFAULT_WORKSPACES_DIR, resolveSessionJsonPath } from '../paths.js';
interface SessionJson {
session: {
id: string;
webUrl: string;
status: 'in-progress' | 'completed' | 'failed';
createdAt: string;
completedAt?: string;
};
metrics: {
total_cost_usd: number;
};
}
interface WorkspaceInfo {
name: string;
url: string;
status: 'in-progress' | 'completed' | 'failed';
createdAt: Date;
completedAt: Date | null;
costUsd: number;
}
function formatDuration(ms: number): string {
const seconds = Math.floor(ms / 1000);
const minutes = Math.floor(seconds / 60);
const hours = Math.floor(minutes / 60);
if (hours > 0) {
return `${hours}h ${minutes % 60}m`;
}
if (minutes > 0) {
return `${minutes}m`;
}
return `${seconds}s`;
}
function getStatusDisplay(status: string): string {
return status;
}
function truncate(str: string, maxLen: number): string {
if (str.length <= maxLen) return str;
return `${str.slice(0, maxLen - 1)}\u2026`;
}
async function listWorkspaces(): Promise<void> {
const workspacesDir = process.env.WORKSPACES_DIR || DEFAULT_WORKSPACES_DIR;
let entries: string[];
try {
entries = await fs.readdir(workspacesDir);
} catch {
console.log('No workspaces directory found.');
console.log(`Expected: ${workspacesDir}`);
return;
}
const workspaces: WorkspaceInfo[] = [];
for (const entry of entries) {
const sessionPath = resolveSessionJsonPath(path.join(workspacesDir, entry));
try {
const content = await fs.readFile(sessionPath, 'utf8');
const data = JSON.parse(content) as SessionJson;
workspaces.push({
name: entry,
url: data.session.webUrl,
status: data.session.status,
createdAt: new Date(data.session.createdAt),
completedAt: data.session.completedAt ? new Date(data.session.completedAt) : null,
costUsd: data.metrics.total_cost_usd,
});
} catch {
// Skip directories without valid session.json
}
}
if (workspaces.length === 0) {
console.log('\nNo workspaces found.');
console.log('Run a pipeline first: ./shannon start -u <url> -r <repo>');
return;
}
// Sort by creation date (most recent first)
workspaces.sort((a, b) => b.createdAt.getTime() - a.createdAt.getTime());
console.log('\n=== Shannon Workspaces ===\n');
// Column widths
const nameWidth = 30;
const urlWidth = 30;
const statusWidth = 14;
const durationWidth = 10;
const costWidth = 10;
// Header
console.log(
' ' +
'WORKSPACE'.padEnd(nameWidth) +
'URL'.padEnd(urlWidth) +
'STATUS'.padEnd(statusWidth) +
'DURATION'.padEnd(durationWidth) +
'COST'.padEnd(costWidth),
);
console.log(` ${'\u2500'.repeat(nameWidth + urlWidth + statusWidth + durationWidth + costWidth)}`);
let resumableCount = 0;
for (const ws of workspaces) {
const now = new Date();
const endTime = ws.completedAt || now;
const durationMs = endTime.getTime() - ws.createdAt.getTime();
const duration = formatDuration(durationMs);
const cost = `$${ws.costUsd.toFixed(2)}`;
const isResumable = ws.status !== 'completed';
if (isResumable) {
resumableCount++;
}
const resumeTag = isResumable ? ' (resumable)' : '';
console.log(
' ' +
truncate(ws.name, nameWidth - 2).padEnd(nameWidth) +
truncate(ws.url, urlWidth - 2).padEnd(urlWidth) +
getStatusDisplay(ws.status).padEnd(statusWidth) +
duration.padEnd(durationWidth) +
cost.padEnd(costWidth) +
resumeTag,
);
}
console.log();
const summary = `${workspaces.length} workspace${workspaces.length === 1 ? '' : 's'} found`;
const resumeSummary = resumableCount > 0 ? ` (${resumableCount} resumable)` : '';
console.log(`${summary}${resumeSummary}`);
if (resumableCount > 0) {
console.log('\nResume with: ./shannon start -u <url> -r <repo> -w <name>');
}
console.log();
}
listWorkspaces().catch((err) => {
console.error('Error listing workspaces:', err);
process.exit(1);
});
+1
View File
@@ -48,6 +48,7 @@ export interface AgentDefinition {
prerequisites: AgentName[];
promptTemplate: string;
deliverableFilename: string;
modelTier?: 'small' | 'medium' | 'large';
}
/**
-5
View File
@@ -27,11 +27,6 @@ export interface AgentEndResult {
attemptNumber: number;
duration_ms: number;
cost_usd: number;
input_tokens?: number | undefined;
output_tokens?: number | undefined;
cache_read_tokens?: number | undefined;
cache_write_tokens?: number | undefined;
turns?: number | undefined;
success: boolean;
model?: string | undefined;
error?: string | undefined;
+9 -6
View File
@@ -11,7 +11,7 @@
export type RuleType = 'url_path' | 'subdomain' | 'domain' | 'method' | 'header' | 'parameter' | 'code_path';
export interface Rule {
description?: string;
description: string;
type: RuleType;
value: string;
}
@@ -32,8 +32,6 @@ export interface ReportConfig {
min_severity?: Severity;
min_confidence?: Confidence;
guidance?: string;
/** Emit report.sarif alongside the markdown report. Ignored when exploit is false. */
sarif?: 'true' | 'false';
}
export type LoginType = 'form' | 'sso' | 'api' | 'basic';
@@ -67,6 +65,7 @@ export interface Authentication {
export interface Config {
rules?: Rules;
authentication?: Authentication;
pipeline?: PipelineConfig;
description?: string;
vuln_classes?: VulnClass[];
exploit?: 'true' | 'false';
@@ -74,8 +73,12 @@ export interface Config {
rules_of_engagement?: string;
}
/** Report config after coercion. The YAML form of `sarif` is a string (see ReportConfig). */
export type DistributedReportConfig = Omit<ReportConfig, 'sarif'> & { sarif: boolean };
export type RetryPreset = 'default' | 'subscription';
export interface PipelineConfig {
retry_preset?: RetryPreset;
max_concurrent_pipelines?: number;
}
export interface DistributedConfig {
avoid: Rule[];
@@ -84,7 +87,7 @@ export interface DistributedConfig {
description: string;
vuln_classes: VulnClass[];
exploit: boolean;
report: DistributedReportConfig;
report: ReportConfig;
rules_of_engagement: string;
}
+7 -1
View File
@@ -25,6 +25,11 @@ export enum ErrorCode {
AGENT_EXECUTION_FAILED = 'AGENT_EXECUTION_FAILED',
OUTPUT_VALIDATION_FAILED = 'OUTPUT_VALIDATION_FAILED',
// Billing errors (PentestErrorType: 'billing')
API_RATE_LIMITED = 'API_RATE_LIMITED',
SPENDING_CAP_REACHED = 'SPENDING_CAP_REACHED',
INSUFFICIENT_CREDITS = 'INSUFFICIENT_CREDITS',
// Git errors (PentestErrorType: 'filesystem')
GIT_CHECKPOINT_FAILED = 'GIT_CHECKPOINT_FAILED',
GIT_ROLLBACK_FAILED = 'GIT_ROLLBACK_FAILED',
@@ -40,9 +45,10 @@ export enum ErrorCode {
TARGET_UNREACHABLE = 'TARGET_UNREACHABLE',
AUTH_FAILED = 'AUTH_FAILED',
AUTH_LOGIN_FAILED = 'AUTH_LOGIN_FAILED',
BILLING_ERROR = 'BILLING_ERROR',
}
export type PentestErrorType = 'config' | 'network' | 'prompt' | 'filesystem' | 'validation' | 'unknown';
export type PentestErrorType = 'config' | 'network' | 'prompt' | 'filesystem' | 'validation' | 'billing' | 'unknown';
export interface PentestErrorContext {
[key: string]: unknown;
-2
View File
@@ -13,8 +13,6 @@ export interface AgentMetrics {
durationMs: number;
inputTokens: number | null;
outputTokens: number | null;
cacheReadTokens: number | null;
cacheWriteTokens: number | null;
costUsd: number | null;
numTurns: number | null;
model?: string | undefined;
@@ -0,0 +1,90 @@
// Copyright (C) 2025 Keygraph, Inc.
//
// This program is free software: you can redistribute it and/or modify
// it under the terms of the GNU Affero General Public License version 3
// as published by the Free Software Foundation.
/**
* Consolidated billing/spending cap detection utilities.
*
* Anthropic's spending cap behavior is inconsistent:
* - Sometimes a proper provider error (billing_error)
* - Sometimes the model responds with text about the cap
* - Sometimes partial billing before cutoff
*
* This module provides defense-in-depth detection with shared pattern lists
* to prevent drift between detection points.
*/
/**
* Text patterns for model-output sniffing (what the model says).
* Used by the pi executor and the behavioral heuristic.
*/
export const BILLING_TEXT_PATTERNS = [
'spending cap',
'spending limit',
'cap reached',
'budget exceeded',
'usage limit',
] as const;
/**
* API patterns for error message classification (what the API returns).
* Used by classifyErrorForTemporal in error-handling.ts.
*/
export const BILLING_API_PATTERNS = [
'billing_error',
'credit balance is too low',
'insufficient credits',
'usage is blocked due to insufficient credits',
'please visit plans & billing',
'please visit plans and billing',
'usage limit reached',
'quota exceeded',
'daily rate limit',
'limit will reset',
'billing limit reached',
] as const;
/**
* Checks if text matches any billing text pattern.
* Used for sniffing model output content for spending cap messages.
*/
export function matchesBillingTextPattern(text: string): boolean {
const lowerText = text.toLowerCase();
return BILLING_TEXT_PATTERNS.some((pattern) => lowerText.includes(pattern));
}
/**
* Checks if an error message matches any billing API pattern.
* Used for classifying API error messages.
*/
export function matchesBillingApiPattern(message: string): boolean {
const lowerMessage = message.toLowerCase();
return BILLING_API_PATTERNS.some((pattern) => lowerMessage.includes(pattern));
}
/**
* Behavioral heuristic for detecting spending cap.
*
* When the model hits a spending cap, it often returns a short message
* with $0 cost. Legitimate agent work NEVER costs $0 with only 1-2 turns.
*
* This combines three signals:
* 1. Very low turn count (<=2)
* 2. Zero cost ($0)
* 3. Text matches billing patterns
*
* @param turns - Number of turns the agent took
* @param cost - Total cost in USD
* @param resultText - The result text from the agent
* @returns true if this looks like a spending cap hit
*/
export function isSpendingCapBehavior(turns: number, cost: number, resultText: string): boolean {
// Only check if turns <= 2 AND cost is exactly 0
if (turns > 2 || cost !== 0) {
return false;
}
return matchesBillingTextPattern(resultText);
}
Binary file not shown.

Before

Width:  |  Height:  |  Size: 62 KiB

-535
View File
@@ -1,535 +0,0 @@
// =============================================================================
// Security Assessment Report — Typst template
// Invoke:
// typst compile --root <root> --input data=/data.json report.typ out.pdf
// Optional overrides:
// --input tester=<name> --input brand=<name>
// =============================================================================
#let data = json(sys.inputs.data)
// Top-level discriminator. Schema variants in report-output-schema.ts:
// exploits → ExploitsReportData (exploit=true runs, full reproduction)
// findings → FindingsReportData (exploit=false runs, analysis-only)
#let mode = data.at("mode", default: "exploits")
#let tester-override = sys.inputs.at("tester", default: "Shannon")
#let brand = sys.inputs.at("brand", default: "Shannon | AI Pentester by Keygraph")
// ---------- Palette ---------------------------------------------------------
// Kept distinct so Critical / High are not confused under monitor gamma.
#let sev-color(level) = {
if level == "Critical" { rgb("#DC2626") } // red-600
else if level == "High" { rgb("#EA580C") } // orange-600
else if level == "Medium" { rgb("#D97706") } // amber-600
else if level == "Low" { rgb("#2563EB") } // blue-600
else { rgb("#6B7280") }
}
#let confidence-color(c) = {
if c == "High" { rgb("#15803D") } // green-700
else if c == "Medium" { rgb("#D97706") } // amber-600
else if c == "Low" { rgb("#6B7280") } // gray-500
else { rgb("#6B7280") }
}
// Warm, editorial, high-contrast document palette.
#let ink = rgb("#141414") // warm near-black text
#let muted = rgb("#5C5850") // warm gray-brown labels
#let tertiary = rgb("#9A958D") // lightest muted
#let rule = rgb("#E6E1D9") // warm hair rules
#let rule-soft = rgb("#D9D3CA")
#let code-bg = rgb("#F6F1EB") // warm eggshell
#let alt-bg = rgb("#EBE6DF")
#let page-bg = white
// ---------- Page setup ------------------------------------------------------
#set document(title: "Security Assessment Report", author: brand)
#set page(
paper: "a4",
margin: (top: 2.2cm, bottom: 2.2cm, left: 2.2cm, right: 2.2cm),
fill: page-bg,
header: context {
if counter(page).get().first() > 1 [
#set text(size: 8.5pt, fill: muted)
#grid(columns: (1fr, auto),
[Security Assessment Report],
[CONFIDENTIAL],
)
#v(-4pt)
#line(length: 100%, stroke: 0.3pt + rule)
]
},
footer: context {
if counter(page).get().first() > 1 [
#set text(size: 8.5pt, fill: muted)
#line(length: 100%, stroke: 0.3pt + rule)
#v(2pt)
#grid(columns: (1fr, auto),
[#data.meta.assessmentDate],
[#counter(page).display() / #context counter(page).final().first()],
)
]
},
)
#set text(size: 10.5pt, fill: ink)
#set par(leading: 0.7em, justify: false)
#show heading.where(level: 1): it => [
#pagebreak(weak: true)
#v(4pt)
#set text(size: 24pt, weight: "bold", fill: ink)
#it.body
#v(4pt)
#line(length: 100%, stroke: 0.4pt + rule)
#v(10pt)
]
#show heading.where(level: 2): it => [
#v(10pt)
#set text(size: 14pt, weight: "semibold", fill: ink)
#it.body
#v(2pt)
]
#show heading.where(level: 3): it => [
#v(8pt)
#set text(size: 11.5pt, weight: "semibold", fill: ink)
#it.body
#v(-2pt)
]
#show raw: set text(size: 8.5pt)
#show raw.where(block: false): it => box(
fill: code-bg,
inset: (x: 3pt, y: 0pt),
outset: (y: 2pt),
radius: 2pt,
it,
)
#show raw.where(block: true): it => block(
fill: code-bg,
stroke: (left: 2pt + rule, rest: none),
inset: (x: 10pt, y: 8pt),
width: 100%,
breakable: true,
{
set par(leading: 0.5em, justify: false)
it
},
)
// ---------- Helpers ---------------------------------------------------------
#let chip(label, color) = box(
fill: color,
inset: (x: 6pt, y: 2pt),
radius: 2pt,
text(fill: white, weight: "bold", size: 7.5pt, tracking: 0.3pt, upper(label)),
)
#let categories-in-order = (
"Authentication",
"Authorization",
"XSS",
"Injection",
"SSRF",
"Other",
)
#let sev-chip(level) = chip(level, sev-color(level))
#let confidence-chip(c) = chip(c + " confidence", confidence-color(c))
// inline-code renders a string, turning backtick-wrapped spans into
// inline raw. Safe on odd counts — a trailing unclosed backtick is
// emitted as literal text so nothing gets swallowed.
#let inline-code(s) = {
if type(s) != str { return s }
let parts = s.split("`")
if parts.len() == 1 { return parts.at(0) }
let out = []
for (i, p) in parts.enumerate() {
if calc.even(i) {
out += [#p]
} else if i == parts.len() - 1 {
out += [#("`" + p)]
} else {
out += raw(p)
}
}
out
}
#let render-items(items) = {
for item in items {
if item.kind == "prose" [
#par(inline-code(item.text))
] else if item.kind == "code" [
#raw(item.block.content, lang: item.block.language, block: true)
]
}
}
// Render step items as a bulleted list; prose items become bullets,
// code items break the list and render as code blocks in between.
#let render-bulleted-items(items) = {
for item in items {
if item.kind == "prose" [
- #inline-code(item.text)
] else if item.kind == "code" [
#raw(item.block.content, lang: item.block.language, block: true)
]
}
}
// Render step items as a numbered list; prose items become enumerated,
// code items break the list and render as code blocks in between.
#let render-numbered-items(items) = {
for item in items {
if item.kind == "prose" [
+ #inline-code(item.text)
] else if item.kind == "code" [
#raw(item.block.content, lang: item.block.language, block: true)
]
}
}
// Render an array of strings as a bulleted list with inline-code support.
#let code-list(items) = list(..items.map(inline-code))
#let kv(label, value) = grid(
columns: (auto, 1fr),
column-gutter: 14pt,
row-gutter: 4pt,
text(fill: muted, size: 9.5pt)[#label],
value,
)
// ---------- COVER PAGE ------------------------------------------------------
#page(header: none, footer: none)[
#set align(left)
#v(3.2cm)
#let brand-parts = brand.split("|").map(p => p.trim())
#grid(
columns: (auto, 1fr),
column-gutter: 8pt,
align: (horizon, horizon),
image("/assets/keygraph-logo.png", width: 1.6cm),
{
set par(leading: 0.6em)
text(size: 11pt, fill: ink, weight: "semibold", tracking: 1.2pt)[
#upper(brand-parts.at(0))
]
if brand-parts.len() > 1 {
linebreak()
text(size: 9pt, fill: muted, weight: "regular")[
#brand-parts.slice(1).join(" ")
]
}
}
)
#set par(leading: 0.7em)
#v(1.6cm)
#set par(leading: 0.4em)
#text(size: 46pt, weight: "bold", fill: ink)[
Security\
Assessment\
Report
]
#set par(leading: 0.7em)
#v(1fr)
#line(length: 100%, stroke: 0.3pt + rule)
#v(0.6cm)
#grid(
columns: (1fr, 1fr),
column-gutter: 28pt,
row-gutter: 14pt,
grid(
columns: (auto, 1fr),
column-gutter: 18pt,
row-gutter: 14pt,
text(fill: muted, size: 9pt)[Target], text(size: 10pt)[#inline-code(data.meta.target)],
text(fill: muted, size: 9pt)[Date], text(size: 10pt)[#data.meta.assessmentDate],
..(if "application" in data.meta and data.meta.application != none {
(text(fill: muted, size: 9pt)[Application], text(size: 10pt)[#inline-code(data.meta.application)])
} else { () }),
),
grid(
columns: (auto, 1fr),
column-gutter: 18pt,
row-gutter: 14pt,
text(fill: muted, size: 9pt)[Tester], text(size: 10pt)[#tester-override],
text(fill: muted, size: 9pt)[Classification],
text(size: 10pt, weight: "semibold")[#data.meta.classification],
),
)
#v(0.8cm)
#text(size: 8pt, fill: muted)[
This document contains sensitive security findings.
Handle in accordance with your organization's data classification policy.
]
]
// ---------- TABLE OF CONTENTS -----------------------------------------------
#outline(title: [Contents], depth: 3, indent: auto)
// ---------- EXECUTIVE SUMMARY -----------------------------------------------
= Executive Summary
#grid(
columns: (auto, 1fr),
column-gutter: 20pt,
row-gutter: 12pt,
text(fill: muted, size: 10pt)[Target], text(size: 10.5pt)[#inline-code(data.meta.target)],
text(fill: muted, size: 10pt)[Date], text(size: 10.5pt)[#data.meta.assessmentDate],
..(if "application" in data.meta and data.meta.application != none {
(text(fill: muted, size: 10pt)[Application], text(size: 10.5pt)[#inline-code(data.meta.application)])
} else { () }),
text(fill: muted, size: 10pt)[Tester], text(size: 10.5pt)[#tester-override],
)
== Scope
#inline-code(data.scope)
// ---------- BY TYPE ---------------------------------------------------------
#let by-type-entries = if mode == "exploits" { data.exploitedByType } else { data.identifiedByType }
#if mode == "exploits" [
= Successfully Exploited Vulnerabilities by Type
] else [
= Identified Vulnerabilities by Type
]
#for entry in by-type-entries [
== #entry.category
#if "narrative" in entry and entry.narrative != none [
#inline-code(entry.narrative)
]
#if "bullets" in entry and entry.bullets != none [
#list(
..entry.bullets.map(b => [
#text(weight: "semibold")[#b.id] — #inline-code(b.description)
])
)
]
]
// ---------- SUMMARY ---------------------------------------------------------
= Summary
#let s = data.summary
#let sev = data.derivedCounts.bySeverity
#let severity-card(label, sev-key, n) = box(
fill: sev-color(sev-key),
inset: (x: 8pt, y: 12pt),
radius: 4pt,
width: 100%,
stack(
dir: ttb,
spacing: 6pt,
text(fill: white, weight: "bold", size: 20pt)[#n],
text(fill: white, size: 8pt, tracking: 0.5pt)[#upper(label)],
),
)
#grid(
columns: 4,
column-gutter: 8pt,
severity-card("Critical", "Critical", sev.Critical),
severity-card("High", "High", sev.High),
severity-card("Medium", "Medium", sev.Medium),
severity-card("Low", "Low", sev.Low),
)
#if mode == "findings" [
#v(18pt)
#let cf = data.derivedCounts.byConfidence
#let confidence-card(label, c-key, n) = box(
stroke: 0.6pt + confidence-color(c-key),
inset: (x: 8pt, y: 12pt),
radius: 4pt,
width: 100%,
stack(
dir: ttb,
spacing: 6pt,
text(fill: ink, weight: "bold", size: 20pt)[#n],
text(fill: confidence-color(c-key), size: 8pt, tracking: 0.5pt)[#upper(label + " confidence")],
),
)
#grid(
columns: 3,
column-gutter: 8pt,
confidence-card("High", "High", cf.High),
confidence-card("Medium", "Medium", cf.Medium),
confidence-card("Low", "Low", cf.Low),
)
]
#v(14pt)
#if mode == "exploits" [
#grid(
columns: (auto, 1fr),
column-gutter: 14pt,
row-gutter: 4pt,
text(fill: muted, size: 10pt)[Total identified],
text(weight: "semibold")[#s.totalIdentified],
text(fill: muted, size: 10pt)[Successfully exploited],
text(weight: "semibold")[#s.successfullyExploited],
)
] else [
#grid(
columns: (auto, 1fr),
column-gutter: 14pt,
row-gutter: 4pt,
text(fill: muted, size: 10pt)[Total identified],
text(weight: "semibold")[#s.totalIdentified],
)
]
#v(8pt)
#let breakdown = if mode == "exploits" { s.exploitedBreakdown } else { s.identifiedBreakdown }
#list(
..breakdown.map(c => [
#text(weight: "semibold")[#c.count] #c.category#if "note" in c and c.note != none [ — #inline-code(c.note)]
])
)
#if mode == "exploits" [
#if "outOfScope" in s and s.outOfScope != none [
#v(4pt)
#text(weight: "semibold")[Out of Scope#if "note" in s.outOfScope and s.outOfScope.note != none [ (#s.outOfScope.note)]:] #s.outOfScope.total vulnerabilities
#if "breakdown" in s.outOfScope and s.outOfScope.breakdown != none [
#list(
..s.outOfScope.breakdown.map(c => [
#text(weight: "semibold")[#c.count] #c.category#if "note" in c and c.note != none [ — #inline-code(c.note)]
])
)
]
]
#if "blockedByConstraints" in s and s.blockedByConstraints != none [
#v(4pt)
#text(weight: "semibold")[Blocked by Testing Constraints:] #s.blockedByConstraints.total#if "note" in s.blockedByConstraints and s.blockedByConstraints.note != none [ — #s.blockedByConstraints.note]
]
]
== Critical Findings
#enum(..s.criticalFindings.map(f => [#inline-code(f)]))
// ---------- FINDINGS OVERVIEW -----------------------------------------------
= Findings Overview
#let show-confidence-col = mode == "findings"
#table(
columns: if show-confidence-col { (auto, 1fr, auto, auto, auto) } else { (auto, 1fr, auto, auto) },
stroke: none,
inset: (x: 8pt, y: 7pt),
align: if show-confidence-col { (left, left, left, center, center) } else { (left, left, left, center) },
fill: (_, row) => if row == 0 { none } else if calc.even(row) { code-bg } else { none },
table.header(
text(size: 9.5pt, weight: "semibold")[ID],
text(size: 9.5pt, weight: "semibold")[Title],
text(size: 9.5pt, weight: "semibold")[Category],
text(size: 9.5pt, weight: "semibold")[Severity],
..(if show-confidence-col { (text(size: 9.5pt, weight: "semibold")[Confidence],) } else { () }),
),
..data.findings.map(f => (
text(weight: "semibold")[#f.id],
inline-code(f.title),
text(size: 9.5pt)[#f.category],
sev-chip(f.severity),
..(if show-confidence-col { (confidence-chip(f.confidence),) } else { () }),
)).flatten()
)
// ---------- FINDING RENDER --------------------------------------------------
#let render-finding-summary(f) = [
#v(8pt)
#grid(
columns: (auto, 1fr),
column-gutter: 18pt,
row-gutter: 12pt,
text(fill: muted, size: 9.5pt)[Location], text(size: 10pt)[#inline-code(f.summary.vulnerableLocation)],
text(fill: muted, size: 9.5pt)[Overview], text(size: 10pt)[#inline-code(f.summary.overview)],
text(fill: muted, size: 9.5pt)[Impact], text(size: 10pt)[#inline-code(f.summary.impact)],
)
]
#let render-finding-extras(f) = [
#if "notes" in f and f.notes != none and f.notes.len() > 0 [
#heading(level: 3, outlined: false)[Notes]
#render-bulleted-items(f.notes)
]
#if "additionalSections" in f and f.additionalSections != none [
#for extra in f.additionalSections [
#heading(level: 3, outlined: false)[#inline-code(extra.heading)]
#render-items(extra.items)
]
]
]
#let render-exploit(f) = [
== #f.id: #inline-code(f.title)
#sev-chip(f.severity)
#render-finding-summary(f)
=== Prerequisites
#inline-code(f.prerequisites)
=== Exploitation Steps
#for step in f.exploitationSteps [
#text(weight: "semibold")[Step #step.number#if "title" in step and step.title != none [ — #inline-code(step.title)]]
#render-items(step.items)
]
=== Proof of Impact
#render-numbered-items(f.proofOfImpact)
#render-finding-extras(f)
#v(16pt)
]
#let render-analysis(f) = [
== #f.id: #inline-code(f.title)
#sev-chip(f.severity) #h(4pt) #confidence-chip(f.confidence)
#render-finding-summary(f)
#render-finding-extras(f)
#v(16pt)
]
#let render-finding(f) = if mode == "exploits" { render-exploit(f) } else { render-analysis(f) }
// ---------- PER-CATEGORY -----------------------------------------------------
#let category-section-label(n) = if mode == "exploits" {
"Exploitation Evidence"
} else {
"Findings"
}
#for cat in categories-in-order {
let cat-findings = data.findings.filter(f => f.category == cat)
if cat-findings.len() > 0 [
= #cat #category-section-label(cat-findings.len()) (#cat-findings.len() #if cat-findings.len() == 1 [finding] else [findings])
#for f in cat-findings {
render-finding(f)
}
]
}
+54 -185
View File
@@ -1,219 +1,88 @@
# AI Providers
One model runs the entire scan — pre-recon, recon, vulnerability analysis, exploitation, and reporting. A single setting names both the provider and the model:
Shannon works best with Claude models. Anthropic API keys are recommended for most users, and Shannon also supports AWS Bedrock and custom Anthropic-compatible endpoints.
## Anthropic
Run the setup wizard:
```bash
export SHANNON_AI_MODEL=<provider>:<model-id>
npx @keygraph/shannon setup
```
The provider half decides where the request goes, which credential is used, and which API dialect is spoken. You never configure those separately.
## Supported providers
| Provider | Value | Credential |
| --- | --- | --- |
| Anthropic | `anthropic` | `SHANNON_AI_API_KEY` (or `CLAUDE_CODE_OAUTH_TOKEN`) |
| OpenAI | `openai` | `SHANNON_AI_API_KEY` |
| xAI | `xai` | `SHANNON_AI_API_KEY` |
| AWS Bedrock | `amazon-bedrock` | `AWS_REGION` and `AWS_BEARER_TOKEN_BEDROCK` |
`SHANNON_AI_API_KEY` holds the key for whichever provider `SHANNON_AI_MODEL` names. Bedrock is the exception — it authenticates through its `AWS_` variables only. If `SHANNON_AI_MODEL` is unset, Shannon uses `anthropic:claude-sonnet-4-6`.
Anthropic, OpenAI, and xAI also accept their native variables (`ANTHROPIC_API_KEY`, `OPENAI_API_KEY`, `XAI_API_KEY`); if one of those is set, it is used instead of `SHANNON_AI_API_KEY`.
Shannon forwards only the selected provider's credential into the scan container. Keys for other providers stay on your machine.
### Any other provider
Shannon accepts any provider and model present in the Pi harness catalogue. Browse them at [pi.dev/models](https://pi.dev/models). These are technically supported but not recommended. Claude models are best-supported (see the note below).
Or export an API key directly:
```bash
export SHANNON_AI_API_KEY=your-api-key # the provider's API key
export SHANNON_AI_MODEL=openrouter:moonshotai/kimi-k3 # <provider>:<model-id>
export ANTHROPIC_API_KEY=your-api-key
```
This path covers providers whose credential is a single API key. Providers that need more than that are not currently supported.
`npx @keygraph/shannon setup` exposes this as the **Other provider** option.
> [!IMPORTANT]
> Claude models are the best-supported option. Shannon's evaluations, internal testing, and agent harness are tuned for Claude. Other models are permitted and validated against the harness catalogue, but may not follow Shannon's instructions or tool-use constraints as reliably. Use them at your own risk.
## Cyber safeguards (do this before your first scan)
Anthropic and OpenAI both apply real-time safeguards to cyber-security workloads. Shannon is exactly such a workload. If a safeguard engages mid-run, the model can refuse, and the scan fails partway through rather than at the start.
Review each vendor's guidance and complete the verification or enrollment they ask of legitimate security testers before running Shannon:
- Anthropic - [Real-time cyber safeguards on Claude Opus and Sonnet](https://support.claude.com/en/articles/14604842-real-time-cyber-safeguards-on-claude-opus-and-sonnet)
- OpenAI - [Cyber](https://chatgpt.com/cyber)
This applies to the Anthropic and OpenAI providers, including when either is reached through a gateway. Bedrock serves Claude models and is subject to Anthropic's safeguards as well.
## Suggested models
These are the models `npx @keygraph/shannon setup` offers, best-first. They are suggestions: the wizard also takes a typed model ID, and `SHANNON_AI_MODEL` accepts any model in the provider's catalogue.
| Provider | Suggested model IDs |
| --- | --- |
| `anthropic` | `claude-sonnet-4-6`, `claude-opus-4-8`, `claude-opus-4-7`, `claude-haiku-4-5-20251001` |
| `openai` | `gpt-5.6-sol`, `gpt-5.5`, `gpt-5.4` |
| `xai` | `grok-4.5` |
| `amazon-bedrock` | `us.anthropic.claude-sonnet-4-6`, `us.anthropic.claude-opus-4-8`, `us.anthropic.claude-opus-4-7` |
Bedrock IDs are region-prefixed and must be enabled in your account, so the ID that works for you may differ from the one listed here.
## Switching provider
The pattern is learned once: export the provider's key, name the model. Two lines change, nothing else.
Anthropic (default):
Source-build mode can use a `.env` file:
```bash
export SHANNON_AI_API_KEY=sk-ant-...
export SHANNON_AI_MODEL=anthropic:claude-sonnet-4-6
ANTHROPIC_API_KEY=your-api-key
```
OpenAI:
```bash
export SHANNON_AI_API_KEY=sk-...
export SHANNON_AI_MODEL=openai:gpt-5.6-sol
```
xAI:
```bash
export SHANNON_AI_API_KEY=xai-...
export SHANNON_AI_MODEL=xai:grok-4.5
```
Source-build mode reads the same variables from a `.env` file.
Each tier can be pointed at any Claude model via `ANTHROPIC_SMALL_MODEL` / `ANTHROPIC_MEDIUM_MODEL` / `ANTHROPIC_LARGE_MODEL` (or the setup wizard). If you set a tier to `claude-fable-5`, note that Fable's safety classifiers route cybersecurity tasks to Opus 4.8, so those phases run on Opus 4.8 regardless.
## AWS Bedrock
Run `npx @keygraph/shannon setup` and select **AWS Bedrock**, or export directly:
Run `npx @keygraph/shannon setup` and select **AWS Bedrock**. The wizard prompts for region, bearer token, and model IDs.
Or export environment variables directly:
```bash
export CLAUDE_CODE_USE_BEDROCK=1
export AWS_REGION=us-east-1
export AWS_BEARER_TOKEN_BEDROCK=your-bearer-token
export SHANNON_AI_MODEL=amazon-bedrock:us.anthropic.claude-opus-4-8
export ANTHROPIC_SMALL_MODEL=us.anthropic.claude-haiku-4-5-20251001-v1:0
export ANTHROPIC_MEDIUM_MODEL=us.anthropic.claude-sonnet-4-6
export ANTHROPIC_LARGE_MODEL=us.anthropic.claude-opus-4-8
```
Bedrock uses bearer-token authentication only. IAM access keys, session tokens, assumed roles, and instance profiles are not supported. The model must be enabled in your region.
## Custom base URL
To route model traffic through your own infrastructure — a corporate proxy, an LLM gateway such as LiteLLM, or a regional endpoint — set a base URL alongside your normal model selection. The provider half of `SHANNON_AI_MODEL` decides which key is sent and which API Shannon speaks, so pick the one your gateway serves:
| Gateway serves | Model prefix | API key |
| --- | --- | --- |
| Anthropic Messages | `anthropic:` | `SHANNON_AI_API_KEY` |
| OpenAI Chat Completions | `openai:` | `SHANNON_AI_API_KEY` |
| OpenAI Responses | `openai:` + `SHANNON_AI_OPENAI_FORMAT=responses` | `SHANNON_AI_API_KEY` |
The model ID is whatever name your gateway serves it under; it does not have to exist in Shannon's catalogue.
Anthropic Messages:
Source-build `.env` equivalent:
```bash
export SHANNON_AI_API_KEY=sk-ant-...
export SHANNON_AI_MODEL=anthropic:claude-sonnet-4-6
export SHANNON_AI_BASE_URL=https://llm-gateway.example.com
CLAUDE_CODE_USE_BEDROCK=1
AWS_REGION=us-east-1
AWS_BEARER_TOKEN_BEDROCK=your-bearer-token
ANTHROPIC_SMALL_MODEL=us.anthropic.claude-haiku-4-5-20251001-v1:0
ANTHROPIC_MEDIUM_MODEL=us.anthropic.claude-sonnet-4-6
ANTHROPIC_LARGE_MODEL=us.anthropic.claude-opus-4-8
```
OpenAI Chat Completions:
Shannon uses three model tiers:
- **small** for summarization
- **medium** for security analysis
- **large** for deep reasoning
Set `ANTHROPIC_SMALL_MODEL`, `ANTHROPIC_MEDIUM_MODEL`, and `ANTHROPIC_LARGE_MODEL` to Bedrock model IDs available in your region.
## Custom Base URL
Shannon supports pointing the SDK at an Anthropic-compatible endpoint with `ANTHROPIC_BASE_URL`. For proxy-based routing, use an LLM proxy such as LiteLLM configured to expose an Anthropic-compatible endpoint.
> [!IMPORTANT]
> Only Claude models are officially supported. Shannon's evaluations, internal testing, and agent harness are optimized for Claude. Smaller or alternative models, including non-Claude models routed through a proxy, may not reliably follow Shannon's instructions or tool-use constraints. Use them at your own risk.
The experimental `claude-code-router` integration has been removed. If you previously relied on it, migrate to an Anthropic-compatible proxy such as LiteLLM.
Run `npx @keygraph/shannon setup` and select **Custom Base URL**, or export variables directly:
```bash
export SHANNON_AI_API_KEY=sk-...
export SHANNON_AI_MODEL=openai:gpt-5.6-sol
export SHANNON_AI_BASE_URL=https://llm-gateway.example.com/v1
export ANTHROPIC_BASE_URL=https://your-proxy.example.com
export ANTHROPIC_AUTH_TOKEN=your-auth-token
export ANTHROPIC_SMALL_MODEL=claude-haiku-4-5-20251001
export ANTHROPIC_MEDIUM_MODEL=claude-sonnet-4-6
export ANTHROPIC_LARGE_MODEL=claude-opus-4-8
```
`SHANNON_AI_MODEL` is always `<provider>:<model-id>`, gateway or not.
OpenAI is the one provider serving two APIs, so a gateway run picks one:
Source-build `.env` equivalent:
```bash
export SHANNON_AI_OPENAI_FORMAT=responses # default: chat-completions
ANTHROPIC_BASE_URL=https://your-proxy.example.com
ANTHROPIC_AUTH_TOKEN=your-auth-token
ANTHROPIC_SMALL_MODEL=claude-haiku-4-5-20251001
ANTHROPIC_MEDIUM_MODEL=claude-sonnet-4-6
ANTHROPIC_LARGE_MODEL=claude-opus-4-8
```
Chat Completions is the default because that is what most gateway software exposes. Set `responses` for a gateway that passes the Responses API through — it preserves reasoning state between turns, which Chat Completions cannot. `openai:gpt-5` with no base URL always calls OpenAI's Responses API directly.
The variable is rejected in preflight where it cannot take effect: with a non-`openai` model, since Anthropic, xAI, and Bedrock each serve one API, and with no `SHANNON_AI_BASE_URL`, since a direct OpenAI run is always Responses.
`npx @keygraph/shannon setup` covers this under **Custom Base URL**, which asks which API your gateway serves and configures the matching provider for you.
## OpenAI Codex (ChatGPT Plus/Pro subscription)
A ChatGPT Plus or Pro Codex subscription can run Shannon. Shannon reuses a login created by Pi.
Before running a pentest, review the [cyber safeguards requirements](#cyber-safeguards-do-this-before-your-first-scan).
1. Install Pi by following the instructions at [pi.dev](https://pi.dev).
2. Log in with your subscription using Pi's [subscription authentication guide](https://pi.dev/docs/latest/providers#subscriptions). This creates `~/.pi/agent/auth.json` with an `openai-codex` entry.
3. Select a Codex model and enable Pi authentication:
```bash
export SHANNON_USE_PI_AUTH=1
export SHANNON_AI_MODEL=openai-codex:gpt-5.5
```
4. In npx mode, run `npx @keygraph/shannon start ...` from the same shell. In source-build mode, add the two variables to `.env` and run `./shannon start ...`.
Supported Codex models are `gpt-5.6-sol`, `gpt-5.5`, and `gpt-5.4`.
## Claude Code subscription
The latest version of Shannon does not support Claude Code subscriptions. The [`shannon-v1`](https://github.com/KeygraphHQ/shannon/tree/shannon-v1) branch is the final release built on the Claude Agent SDK and supports Claude Code OAuth.
Before running a pentest, review the [cyber safeguards requirements](#cyber-safeguards-do-this-before-your-first-scan).
1. Generate a Claude Code OAuth token:
```bash
claude setup-token
```
2. Run the setup flow for the final `shannon-v1` release:
```bash
npx @keygraph/shannon@1.9.0 setup
```
3. Select **OAuth Token** and enter the token generated by Claude Code.
4. Start the pentest with `npx @keygraph/shannon@1.9.0 start ...`.
These instructions apply only to `shannon-v1`.
## Validation
Checks run before a scan starts, so mistakes fail immediately rather than partway through a run:
- **Provider and model ID** — validated against the Pi harness catalogue. An unknown provider or model ID fails preflight with a pointer to [pi.dev/models](https://pi.dev/models). A custom base URL exempts the model ID, since a gateway may serve its own names.
- **Credential presence** — validated for the selected provider, or read from Pi when `SHANNON_USE_PI_AUTH=1`.
- **Credential validity** — one minimal request against the model the scan will use, so a rejected key, an exhausted quota, or a model the account cannot reach fails before any agent runs. Bedrock included: its bearer token and region go through the same probe.
## Migrating from the three-tier configuration
Earlier versions took three model variables. They no longer do anything — replace them with `SHANNON_AI_MODEL`.
| Before | Now |
| --- | --- |
| `ANTHROPIC_SMALL_MODEL`, `ANTHROPIC_MEDIUM_MODEL`, `ANTHROPIC_LARGE_MODEL` | a single `SHANNON_AI_MODEL` |
| `CLAUDE_CODE_USE_BEDROCK=1` plus three Bedrock model IDs | `SHANNON_AI_MODEL=amazon-bedrock:<model-id>` |
| `ANTHROPIC_BASE_URL` + `ANTHROPIC_AUTH_TOKEN` selected a provider | `SHANNON_AI_BASE_URL` overrides the endpoint; `SHANNON_AI_MODEL` selects the provider |
In `~/.shannon/config.toml`, the `[models]` section and `bedrock.use` are gone, each provider has its own section, and the model lives at `core.model`:
```toml
[core]
model = "anthropic:claude-sonnet-4-6"
# base_url = "https://llm-gateway.example.com"
[anthropic]
api_key = "your-api-key"
```
Re-run `npx @keygraph/shannon setup` to regenerate the file.
+21 -28
View File
@@ -1,6 +1,6 @@
# Configuration
Shannon can run without a configuration file, but configuration enables authenticated testing, scope guidance, rules of engagement, and report filtering.
Shannon can run without a configuration file, but configuration enables authenticated testing, scope guidance, rules of engagement, report filtering, and rate-limit tuning.
## Credential Precedence
@@ -93,40 +93,14 @@ rules:
type: url_path
value: "/api"
# Report options applied when assembling the final report.
# Filters applied by the report agent when assembling the final report.
# report:
# min_severity: low
# min_confidence: low
# guidance: |
# Drop findings about missing security headers and rate-limit gaps.
# sarif: "true"
```
## Report Options
| Key | Effect |
| --- | --- |
| `min_severity` | Drops findings rated below this severity. Applies only when `exploit` is `"true"`. |
| `min_confidence` | Drops findings rated below this confidence. Applies only when `exploit` is `"false"`. |
| `guidance` | Free-text instruction to the report agent, such as which topics to exclude. |
| `sarif` | Emits a SARIF 2.1.0 log alongside the Markdown report. Requires `exploit: "true"`. |
A finding carries one rating or the other, never both: an exploited finding is rated by severity, an analysis-only finding by confidence. Setting the threshold that does not apply to the run is ignored, and Shannon logs a warning naming the one to use instead.
### SARIF Output
Set `sarif: "true"` to write `report.sarif` next to `Security-Assessment-Report.pdf` at the workspace root, for upload to GitHub code scanning or any other SARIF consumer.
```yaml
exploit: "true"
report:
sarif: "true"
```
Each finding becomes one SARIF result, filed under a rule per vulnerability class (`shannon/injection`, `shannon/xss`, `shannon/auth`, `shannon/authz`, `shannon/ssrf`) and tagged with its OWASP Top Ten 2025 category. Results are anchored to the code location the analysis phase recorded, falling back to the HTTP entry point when the finding names no file. Severity maps onto SARIF's three levels: `critical` and `high` become `error`, `medium` becomes `warning`, everything else becomes `note`.
The log is written only for exploitative runs. An analysis-only run rates findings by confidence and produces no severity, so there is nothing to populate `level` with; `sarif` is ignored when `exploit` is `"false"`.
Supported rule types include `url_path`, `subdomain`, `domain`, `method`, `header`, `parameter`, and `code_path`.
## Writing Login Flow
@@ -156,3 +130,22 @@ login_flow:
- "If prompted for 2FA, type $totp in <exact code field label or placeholder>"
- "Click <exact button text>"
```
## Adaptive Thinking
Claude decides when and how deeply to reason on Opus 4.6, 4.7, and 4.8. This is enabled by default whenever a tier resolves to one of these models.
- `npx` mode: `npx @keygraph/shannon setup` prompts you during the wizard.
- Source-build mode: set `CLAUDE_ADAPTIVE_THINKING=false` in `.env` or export it in your shell.
## Subscription Plan Rate Limits
Anthropic subscription plans reset usage on a rolling 5-hour window. The default retry strategy may exhaust retries before the window resets. Add this to your config:
```yaml
pipeline:
retry_preset: subscription
max_concurrent_pipelines: 2
```
`max_concurrent_pipelines` controls how many vulnerability pipelines run simultaneously. Supported values are 1-5, with a default of 5. Lower values reduce burst API usage but increase wall-clock time.
+12 -20
View File
@@ -58,8 +58,7 @@ Monitor progress:
```bash
npx @keygraph/shannon logs <workspace>
npx @keygraph/shannon status <workspace>
npx @keygraph/shannon scans
npx @keygraph/shannon status
npx @keygraph/shannon version
```
@@ -67,8 +66,7 @@ Source-build equivalents:
```bash
./shannon logs <workspace>
./shannon status <workspace>
./shannon scans
./shannon status
./shannon version
```
@@ -81,17 +79,16 @@ open http://localhost:8233
Stop Shannon:
```bash
npx @keygraph/shannon stop <workspace> # stop one scan (confirms first; add --yes/-y to skip)
npx @keygraph/shannon stop --all # stop all scans (Temporal stays up)
npx @keygraph/shannon reset # stop everything and wipe all Temporal data (type 'confirm' to proceed; cannot be skipped)
npx @keygraph/shannon stop
npx @keygraph/shannon stop --clean # confirms first; add --yes (or -y) to skip
npx @keygraph/shannon uninstall # confirms first; add --yes (or -y) to skip
```
Source-build equivalents:
```bash
./shannon stop <workspace> # stop one scan (confirms first; add --yes/-y to skip)
./shannon stop --all # stop all scans (Temporal stays up)
./shannon reset # stop everything and wipe all Temporal data (type 'confirm' to proceed; cannot be skipped)
./shannon stop
./shannon stop --clean # add --yes (or -y) to skip the confirmation
```
Usage examples:
@@ -109,11 +106,8 @@ npx @keygraph/shannon start -u https://example.com -r /path/to/repo -o ./my-repo
# Named workspace.
npx @keygraph/shannon start -u https://example.com -r /path/to/repo -w q1-audit
# Stream the log until the scan finishes, then exit on its outcome (useful in CI).
npx @keygraph/shannon start -u https://example.com -r /path/to/repo --follow
# List completed scans.
npx @keygraph/shannon scans
# List all workspaces.
npx @keygraph/shannon workspaces
```
Source-build examples:
@@ -123,8 +117,7 @@ Source-build examples:
./shannon start -u https://example.com -r /path/to/repo -c /path/to/my-config.yaml
./shannon start -u https://example.com -r /path/to/repo -o ./my-reports
./shannon start -u https://example.com -r /path/to/repo -w q1-audit
./shannon start -u https://example.com -r /path/to/repo --follow
./shannon scans
./shannon workspaces
# Rebuild the worker image.
./shannon build --no-cache
@@ -139,12 +132,11 @@ Results are saved to the workspaces directory:
Use `-o <path>` to copy deliverables to a custom output directory after a run completes.
Output structure — the run directory's top level holds the final report, in PDF and Markdown; everything else is nested under a hidden `.shannon/` directory:
Output structure — the run directory's top level holds only the final report; everything else is nested under a hidden `.shannon/` directory:
```text
workspaces/{hostname}_{sessionId}/
|-- Security-Assessment-Report.pdf # the final report (PDF)
|-- Security-Assessment-Report.md # the final report (Markdown)
|-- Security-Assessment-Report.md # the final report (the deliverable)
`-- .shannon/ # internals
|-- deliverables/ # report source, per-phase analysis, queues
|-- agents/ # per-agent logs
+1
View File
@@ -49,3 +49,4 @@ For broader coverage, the Keygraph platform adds black-box and white-box agentic
A full test run typically takes roughly 1 to 1.5 hours. LLM API costs vary by model pricing, target complexity, selected provider, and concurrency.
If you use subscription-based model access, consider the rate-limit guidance in [Configuration](configuration.md).
+4 -4
View File
@@ -11,7 +11,7 @@ Shannon uses workspaces to store scan state, logs, prompts, and deliverables. Wo
- Use `-w <name>` to give a run a custom name.
- To resume a run, pass the same workspace name with `-w`.
- Each agent's progress is checkpointed so resumed runs can skip completed work.
- The final report is surfaced at the workspace root as `Security-Assessment-Report.pdf` and `Security-Assessment-Report.md`. Run internals — deliverables, logs, prompts, and session state — live under a hidden `.shannon/` directory.
- The final report is surfaced at the workspace root as `Security-Assessment-Report.md`. Run internals — deliverables, logs, prompts, and session state — live under a hidden `.shannon/` directory.
> [!NOTE]
> The URL must match the original workspace URL when resuming. Shannon rejects mismatched URLs to prevent cross-target contamination.
@@ -36,10 +36,10 @@ Resume an auto-named workspace:
npx @keygraph/shannon start -u https://example.com -r /path/to/repo -w example-com_shannon-1771007534808
```
List completed scans:
List all workspaces:
```bash
npx @keygraph/shannon scans
npx @keygraph/shannon workspaces
```
Source-build equivalents:
@@ -47,5 +47,5 @@ Source-build equivalents:
```bash
./shannon start -u https://example.com -r /path/to/repo -w my-audit
./shannon start -u https://example.com -r /path/to/repo -w example-com_shannon-1771007534808
./shannon scans
./shannon workspaces
```
+1 -1
View File
@@ -12,7 +12,7 @@ if [ -n "$TARGET_UID" ] && [ "$TARGET_UID" != "$CURRENT_UID" ]; then
groupadd -g "$TARGET_GID" pentest
useradd -u "$TARGET_UID" -g pentest -s /bin/bash -M pentest
chown -R pentest:pentest /app/sessions /app/workspaces /tmp/.claude /tmp/.pi
chown -R pentest:pentest /app/sessions /app/workspaces /tmp/.claude
fi
exec su -m pentest -c "exec $*"
+96 -202
View File
@@ -8,7 +8,7 @@
# File: README.md
> [!NOTE]
> **[Shannon 2.0 now runs on the Pi harness](https://github.com/KeygraphHQ/shannon/discussions/393)**
> **[Shannon Now Runs on the Pi Harness (Beta) - run it today with `npx @keygraph/shannon@beta`](https://github.com/KeygraphHQ/shannon/discussions/358)**
<div align="center">
@@ -82,8 +82,7 @@ Sample penetration test reports from intentionally vulnerable applications, prod
- **Docker**: required for the worker container.
- **Node.js 18+**: required for the recommended `npx` workflow.
- **AI provider credentials**: Anthropic, OpenAI, xAI, or AWS Bedrock - or [any other provider](docs/ai-providers.md#any-other-provider). Claude models are recommended. Gateway and proxy setups are documented separately.
- **Cyber safeguards cleared with your provider**: Anthropic and OpenAI apply real-time safeguards to cyber-security workloads, which can interrupt a scan mid-run. Complete their guidance for legitimate security testers before your first run - see [AI providers](docs/ai-providers.md#cyber-safeguards-do-this-before-your-first-scan).
- **AI provider credentials**: Anthropic is recommended. AWS Bedrock and compatible proxy setups are documented separately.
### Run Shannon
@@ -102,9 +101,6 @@ Shannon pulls the worker image from Docker Hub, starts the required local infras
For source builds, authenticated scans, provider-specific setup, and platform notes, see [Documentation](#documentation).
> [!TIP]
> **Prefer to run on your Claude Code subscription instead of API credits?** The [`shannon-v1`](https://github.com/KeygraphHQ/shannon/tree/shannon-v1) branch is the last release built on the Claude Agent SDK, so it accepts a Claude Code OAuth token. Generate one with `claude setup-token`, then run `npx @keygraph/shannon@1.9.0 setup` and pick **OAuth Token**. Pentests then cost nothing beyond your existing subscription.
## Key Capabilities
- **Proof-by-exploitation reports**: Shannon reports validated findings with reproducible proof-of-concept steps instead of speculative warnings.
@@ -198,8 +194,8 @@ Use these guides for operational detail:
| Guide | Use it for |
| --- | --- |
| [Source build and CLI commands](docs/development.md) | Cloning, building, common commands, output paths, and local development. |
| [Configuration](docs/configuration.md) | Authenticated testing, login flows, rules of engagement, and report filters. |
| [AI providers](docs/ai-providers.md) | Selecting the model, the supported providers (Anthropic, OpenAI, xAI, AWS Bedrock), and custom gateways. |
| [Configuration](docs/configuration.md) | Authenticated testing, login flows, rules of engagement, report filters, and rate-limit settings. |
| [AI providers](docs/ai-providers.md) | Anthropic, AWS Bedrock, and custom Anthropic-compatible endpoints. |
| [Platforms and networking](docs/platforms.md) | Windows/WSL2, Linux, macOS, Docker networking, local apps, and custom hostnames. |
| [Workspaces and resuming](docs/workspaces.md) | Naming workspaces, resuming interrupted scans, and workspace storage. |
| [Safety and limitations](docs/safety.md) | Authorized-use requirements, non-production guidance, mutative effects, cost, and model caveats. |
@@ -323,8 +319,7 @@ Monitor progress:
```bash
npx @keygraph/shannon logs <workspace>
npx @keygraph/shannon status <workspace>
npx @keygraph/shannon scans
npx @keygraph/shannon status
npx @keygraph/shannon version
```
@@ -332,8 +327,7 @@ Source-build equivalents:
```bash
./shannon logs <workspace>
./shannon status <workspace>
./shannon scans
./shannon status
./shannon version
```
@@ -346,17 +340,16 @@ open http://localhost:8233
Stop Shannon:
```bash
npx @keygraph/shannon stop <workspace> # stop one scan (confirms first; add --yes/-y to skip)
npx @keygraph/shannon stop --all # stop all scans (Temporal stays up)
npx @keygraph/shannon reset # stop everything and wipe all Temporal data (type 'confirm' to proceed; cannot be skipped)
npx @keygraph/shannon stop
npx @keygraph/shannon stop --clean # confirms first; add --yes (or -y) to skip
npx @keygraph/shannon uninstall # confirms first; add --yes (or -y) to skip
```
Source-build equivalents:
```bash
./shannon stop <workspace> # stop one scan (confirms first; add --yes/-y to skip)
./shannon stop --all # stop all scans (Temporal stays up)
./shannon reset # stop everything and wipe all Temporal data (type 'confirm' to proceed; cannot be skipped)
./shannon stop
./shannon stop --clean # add --yes (or -y) to skip the confirmation
```
Usage examples:
@@ -374,11 +367,8 @@ npx @keygraph/shannon start -u https://example.com -r /path/to/repo -o ./my-repo
# Named workspace.
npx @keygraph/shannon start -u https://example.com -r /path/to/repo -w q1-audit
# Stream the log until the scan finishes, then exit on its outcome (useful in CI).
npx @keygraph/shannon start -u https://example.com -r /path/to/repo --follow
# List completed scans.
npx @keygraph/shannon scans
# List all workspaces.
npx @keygraph/shannon workspaces
```
Source-build examples:
@@ -388,8 +378,7 @@ Source-build examples:
./shannon start -u https://example.com -r /path/to/repo -c /path/to/my-config.yaml
./shannon start -u https://example.com -r /path/to/repo -o ./my-reports
./shannon start -u https://example.com -r /path/to/repo -w q1-audit
./shannon start -u https://example.com -r /path/to/repo --follow
./shannon scans
./shannon workspaces
# Rebuild the worker image.
./shannon build --no-cache
@@ -404,12 +393,11 @@ Results are saved to the workspaces directory:
Use `-o <path>` to copy deliverables to a custom output directory after a run completes.
Output structure — the run directory's top level holds the final report, in PDF and Markdown; everything else is nested under a hidden `.shannon/` directory:
Output structure — the run directory's top level holds only the final report; everything else is nested under a hidden `.shannon/` directory:
```text
workspaces/{hostname}_{sessionId}/
|-- Security-Assessment-Report.pdf # the final report (PDF)
|-- Security-Assessment-Report.md # the final report (Markdown)
|-- Security-Assessment-Report.md # the final report (the deliverable)
`-- .shannon/ # internals
|-- deliverables/ # report source, per-phase analysis, queues
|-- agents/ # per-agent logs
@@ -425,7 +413,7 @@ workspaces/{hostname}_{sessionId}/
# Configuration
Shannon can run without a configuration file, but configuration enables authenticated testing, scope guidance, rules of engagement, and report filtering.
Shannon can run without a configuration file, but configuration enables authenticated testing, scope guidance, rules of engagement, report filtering, and rate-limit tuning.
## Credential Precedence
@@ -518,40 +506,14 @@ rules:
type: url_path
value: "/api"
# Report options applied when assembling the final report.
# Filters applied by the report agent when assembling the final report.
# report:
# min_severity: low
# min_confidence: low
# guidance: |
# Drop findings about missing security headers and rate-limit gaps.
# sarif: "true"
```
## Report Options
| Key | Effect |
| --- | --- |
| `min_severity` | Drops findings rated below this severity. Applies only when `exploit` is `"true"`. |
| `min_confidence` | Drops findings rated below this confidence. Applies only when `exploit` is `"false"`. |
| `guidance` | Free-text instruction to the report agent, such as which topics to exclude. |
| `sarif` | Emits a SARIF 2.1.0 log alongside the Markdown report. Requires `exploit: "true"`. |
A finding carries one rating or the other, never both: an exploited finding is rated by severity, an analysis-only finding by confidence. Setting the threshold that does not apply to the run is ignored, and Shannon logs a warning naming the one to use instead.
### SARIF Output
Set `sarif: "true"` to write `report.sarif` next to `Security-Assessment-Report.pdf` at the workspace root, for upload to GitHub code scanning or any other SARIF consumer.
```yaml
exploit: "true"
report:
sarif: "true"
```
Each finding becomes one SARIF result, filed under a rule per vulnerability class (`shannon/injection`, `shannon/xss`, `shannon/auth`, `shannon/authz`, `shannon/ssrf`) and tagged with its OWASP Top Ten 2025 category. Results are anchored to the code location the analysis phase recorded, falling back to the HTTP entry point when the finding names no file. Severity maps onto SARIF's three levels: `critical` and `high` become `error`, `medium` becomes `warning`, everything else becomes `note`.
The log is written only for exploitative runs. An analysis-only run rates findings by confidence and produces no severity, so there is nothing to populate `level` with; `sarif` is ignored when `exploit` is `"false"`.
Supported rule types include `url_path`, `subdomain`, `domain`, `method`, `header`, `parameter`, and `code_path`.
## Writing Login Flow
@@ -582,187 +544,118 @@ login_flow:
- "Click <exact button text>"
```
## Adaptive Thinking
Claude decides when and how deeply to reason on Opus 4.6, 4.7, and 4.8. This is enabled by default whenever a tier resolves to one of these models.
- `npx` mode: `npx @keygraph/shannon setup` prompts you during the wizard.
- Source-build mode: set `CLAUDE_ADAPTIVE_THINKING=false` in `.env` or export it in your shell.
## Subscription Plan Rate Limits
Anthropic subscription plans reset usage on a rolling 5-hour window. The default retry strategy may exhaust retries before the window resets. Add this to your config:
```yaml
pipeline:
retry_preset: subscription
max_concurrent_pipelines: 2
```
`max_concurrent_pipelines` controls how many vulnerability pipelines run simultaneously. Supported values are 1-5, with a default of 5. Lower values reduce burst API usage but increase wall-clock time.
---
# File: docs/ai-providers.md
# AI Providers
One model runs the entire scan — pre-recon, recon, vulnerability analysis, exploitation, and reporting. A single setting names both the provider and the model:
Shannon works best with Claude models. Anthropic API keys are recommended for most users, and Shannon also supports AWS Bedrock and custom Anthropic-compatible endpoints.
## Anthropic
Run the setup wizard:
```bash
export SHANNON_AI_MODEL=<provider>:<model-id>
npx @keygraph/shannon setup
```
The provider half decides where the request goes, which credential is used, and which API dialect is spoken. You never configure those separately.
## Supported providers
| Provider | Value | Credential |
| --- | --- | --- |
| Anthropic | `anthropic` | `SHANNON_AI_API_KEY` (or `CLAUDE_CODE_OAUTH_TOKEN`) |
| OpenAI | `openai` | `SHANNON_AI_API_KEY` |
| xAI | `xai` | `SHANNON_AI_API_KEY` |
| AWS Bedrock | `amazon-bedrock` | `AWS_REGION` and `AWS_BEARER_TOKEN_BEDROCK` |
`SHANNON_AI_API_KEY` holds the key for whichever provider `SHANNON_AI_MODEL` names. Bedrock is the exception — it authenticates through its `AWS_` variables only. If `SHANNON_AI_MODEL` is unset, Shannon uses `anthropic:claude-sonnet-4-6`.
Anthropic, OpenAI, and xAI also accept their native variables (`ANTHROPIC_API_KEY`, `OPENAI_API_KEY`, `XAI_API_KEY`); if one of those is set, it is used instead of `SHANNON_AI_API_KEY`.
Shannon forwards only the selected provider's credential into the scan container. Keys for other providers stay on your machine.
### Any other provider
Shannon accepts any provider and model present in the Pi harness catalogue. Browse them at [pi.dev/models](https://pi.dev/models). These are technically supported but not recommended. Claude models are best-supported (see the note below).
Or export an API key directly:
```bash
export SHANNON_AI_API_KEY=your-api-key # the provider's API key
export SHANNON_AI_MODEL=openrouter:moonshotai/kimi-k3 # <provider>:<model-id>
export ANTHROPIC_API_KEY=your-api-key
```
This path covers providers whose credential is a single API key. Providers that need more than that are not currently supported.
`npx @keygraph/shannon setup` exposes this as the **Other provider** option.
> [!IMPORTANT]
> Claude models are the best-supported option. Shannon's evaluations, internal testing, and agent harness are tuned for Claude. Other models are permitted and validated against the harness catalogue, but may not follow Shannon's instructions or tool-use constraints as reliably. Use them at your own risk.
## Cyber safeguards (do this before your first scan)
Anthropic and OpenAI both apply real-time safeguards to cyber-security workloads. Shannon is exactly such a workload. If a safeguard engages mid-run, the model can refuse, and the scan fails partway through rather than at the start.
Review each vendor's guidance and complete the verification or enrollment they ask of legitimate security testers before running Shannon:
- Anthropic - [Real-time cyber safeguards on Claude Opus and Sonnet](https://support.claude.com/en/articles/14604842-real-time-cyber-safeguards-on-claude-opus-and-sonnet)
- OpenAI - [Cyber](https://chatgpt.com/cyber)
This applies to the Anthropic and OpenAI providers, including when either is reached through a gateway. Bedrock serves Claude models and is subject to Anthropic's safeguards as well.
## Suggested models
These are the models `npx @keygraph/shannon setup` offers, best-first. They are suggestions: the wizard also takes a typed model ID, and `SHANNON_AI_MODEL` accepts any model in the provider's catalogue.
| Provider | Suggested model IDs |
| --- | --- |
| `anthropic` | `claude-sonnet-4-6`, `claude-opus-4-8`, `claude-opus-4-7`, `claude-haiku-4-5-20251001` |
| `openai` | `gpt-5.6-sol`, `gpt-5.5`, `gpt-5.4` |
| `xai` | `grok-4.5` |
| `amazon-bedrock` | `us.anthropic.claude-sonnet-4-6`, `us.anthropic.claude-opus-4-8`, `us.anthropic.claude-opus-4-7` |
Bedrock IDs are region-prefixed and must be enabled in your account, so the ID that works for you may differ from the one listed here.
## Switching provider
The pattern is learned once: export the provider's key, name the model. Two lines change, nothing else.
Anthropic (default):
Source-build mode can use a `.env` file:
```bash
export SHANNON_AI_API_KEY=sk-ant-...
export SHANNON_AI_MODEL=anthropic:claude-sonnet-4-6
ANTHROPIC_API_KEY=your-api-key
```
OpenAI:
```bash
export SHANNON_AI_API_KEY=sk-...
export SHANNON_AI_MODEL=openai:gpt-5.6-sol
```
xAI:
```bash
export SHANNON_AI_API_KEY=xai-...
export SHANNON_AI_MODEL=xai:grok-4.5
```
Source-build mode reads the same variables from a `.env` file.
Each tier can be pointed at any Claude model via `ANTHROPIC_SMALL_MODEL` / `ANTHROPIC_MEDIUM_MODEL` / `ANTHROPIC_LARGE_MODEL` (or the setup wizard). If you set a tier to `claude-fable-5`, note that Fable's safety classifiers route cybersecurity tasks to Opus 4.8, so those phases run on Opus 4.8 regardless.
## AWS Bedrock
Run `npx @keygraph/shannon setup` and select **AWS Bedrock**, or export directly:
Run `npx @keygraph/shannon setup` and select **AWS Bedrock**. The wizard prompts for region, bearer token, and model IDs.
Or export environment variables directly:
```bash
export CLAUDE_CODE_USE_BEDROCK=1
export AWS_REGION=us-east-1
export AWS_BEARER_TOKEN_BEDROCK=your-bearer-token
export SHANNON_AI_MODEL=amazon-bedrock:us.anthropic.claude-opus-4-8
export ANTHROPIC_SMALL_MODEL=us.anthropic.claude-haiku-4-5-20251001-v1:0
export ANTHROPIC_MEDIUM_MODEL=us.anthropic.claude-sonnet-4-6
export ANTHROPIC_LARGE_MODEL=us.anthropic.claude-opus-4-8
```
Bedrock uses bearer-token authentication only. IAM access keys, session tokens, assumed roles, and instance profiles are not supported. The model must be enabled in your region.
## Custom base URL
To route model traffic through your own infrastructure — a corporate proxy, an LLM gateway such as LiteLLM, or a regional endpoint — set a base URL alongside your normal model selection. The provider half of `SHANNON_AI_MODEL` decides which key is sent and which API Shannon speaks, so pick the one your gateway serves:
| Gateway serves | Model prefix | API key |
| --- | --- | --- |
| Anthropic Messages | `anthropic:` | `SHANNON_AI_API_KEY` |
| OpenAI Chat Completions | `openai:` | `SHANNON_AI_API_KEY` |
| OpenAI Responses | `openai:` + `SHANNON_AI_OPENAI_FORMAT=responses` | `SHANNON_AI_API_KEY` |
The model ID is whatever name your gateway serves it under; it does not have to exist in Shannon's catalogue.
Anthropic Messages:
Source-build `.env` equivalent:
```bash
export SHANNON_AI_API_KEY=sk-ant-...
export SHANNON_AI_MODEL=anthropic:claude-sonnet-4-6
export SHANNON_AI_BASE_URL=https://llm-gateway.example.com
CLAUDE_CODE_USE_BEDROCK=1
AWS_REGION=us-east-1
AWS_BEARER_TOKEN_BEDROCK=your-bearer-token
ANTHROPIC_SMALL_MODEL=us.anthropic.claude-haiku-4-5-20251001-v1:0
ANTHROPIC_MEDIUM_MODEL=us.anthropic.claude-sonnet-4-6
ANTHROPIC_LARGE_MODEL=us.anthropic.claude-opus-4-8
```
OpenAI Chat Completions:
Shannon uses three model tiers:
- **small** for summarization
- **medium** for security analysis
- **large** for deep reasoning
Set `ANTHROPIC_SMALL_MODEL`, `ANTHROPIC_MEDIUM_MODEL`, and `ANTHROPIC_LARGE_MODEL` to Bedrock model IDs available in your region.
## Custom Base URL
Shannon supports pointing the SDK at an Anthropic-compatible endpoint with `ANTHROPIC_BASE_URL`. For proxy-based routing, use an LLM proxy such as LiteLLM configured to expose an Anthropic-compatible endpoint.
> [!IMPORTANT]
> Only Claude models are officially supported. Shannon's evaluations, internal testing, and agent harness are optimized for Claude. Smaller or alternative models, including non-Claude models routed through a proxy, may not reliably follow Shannon's instructions or tool-use constraints. Use them at your own risk.
The experimental `claude-code-router` integration has been removed. If you previously relied on it, migrate to an Anthropic-compatible proxy such as LiteLLM.
Run `npx @keygraph/shannon setup` and select **Custom Base URL**, or export variables directly:
```bash
export SHANNON_AI_API_KEY=sk-...
export SHANNON_AI_MODEL=openai:gpt-5.6-sol
export SHANNON_AI_BASE_URL=https://llm-gateway.example.com/v1
export ANTHROPIC_BASE_URL=https://your-proxy.example.com
export ANTHROPIC_AUTH_TOKEN=your-auth-token
export ANTHROPIC_SMALL_MODEL=claude-haiku-4-5-20251001
export ANTHROPIC_MEDIUM_MODEL=claude-sonnet-4-6
export ANTHROPIC_LARGE_MODEL=claude-opus-4-8
```
`SHANNON_AI_MODEL` is always `<provider>:<model-id>`, gateway or not.
OpenAI is the one provider serving two APIs, so a gateway run picks one:
Source-build `.env` equivalent:
```bash
export SHANNON_AI_OPENAI_FORMAT=responses # default: chat-completions
ANTHROPIC_BASE_URL=https://your-proxy.example.com
ANTHROPIC_AUTH_TOKEN=your-auth-token
ANTHROPIC_SMALL_MODEL=claude-haiku-4-5-20251001
ANTHROPIC_MEDIUM_MODEL=claude-sonnet-4-6
ANTHROPIC_LARGE_MODEL=claude-opus-4-8
```
Chat Completions is the default because that is what most gateway software exposes. Set `responses` for a gateway that passes the Responses API through — it preserves reasoning state between turns, which Chat Completions cannot. `openai:gpt-5` with no base URL always calls OpenAI's Responses API directly.
The variable is rejected in preflight where it cannot take effect: with a non-`openai` model, since Anthropic, xAI, and Bedrock each serve one API, and with no `SHANNON_AI_BASE_URL`, since a direct OpenAI run is always Responses.
`npx @keygraph/shannon setup` covers this under **Custom Base URL**, which asks which API your gateway serves and configures the matching provider for you.
## Validation
Checks run before a scan starts, so mistakes fail immediately rather than partway through a run:
- **Provider and model ID** — validated against the Pi harness catalogue. An unknown provider or model ID fails preflight with a pointer to [pi.dev/models](https://pi.dev/models). A custom base URL exempts the model ID, since a gateway may serve its own names.
- **Credential presence** — always validated for the selected provider.
- **Credential validity** — one minimal request against the model the scan will use, so a rejected key, an exhausted quota, or a model the account cannot reach fails before any agent runs. Bedrock included: its bearer token and region go through the same probe.
## Migrating from the three-tier configuration
Earlier versions took three model variables. They no longer do anything — replace them with `SHANNON_AI_MODEL`.
| Before | Now |
| --- | --- |
| `ANTHROPIC_SMALL_MODEL`, `ANTHROPIC_MEDIUM_MODEL`, `ANTHROPIC_LARGE_MODEL` | a single `SHANNON_AI_MODEL` |
| `CLAUDE_CODE_USE_BEDROCK=1` plus three Bedrock model IDs | `SHANNON_AI_MODEL=amazon-bedrock:<model-id>` |
| `ANTHROPIC_BASE_URL` + `ANTHROPIC_AUTH_TOKEN` selected a provider | `SHANNON_AI_BASE_URL` overrides the endpoint; `SHANNON_AI_MODEL` selects the provider |
In `~/.shannon/config.toml`, the `[models]` section and `bedrock.use` are gone, each provider has its own section, and the model lives at `core.model`:
```toml
[core]
model = "anthropic:claude-sonnet-4-6"
# base_url = "https://llm-gateway.example.com"
[anthropic]
api_key = "your-api-key"
```
Re-run `npx @keygraph/shannon setup` to regenerate the file.
---
# File: docs/platforms.md
@@ -869,7 +762,7 @@ Shannon uses workspaces to store scan state, logs, prompts, and deliverables. Wo
- Use `-w <name>` to give a run a custom name.
- To resume a run, pass the same workspace name with `-w`.
- Each agent's progress is checkpointed so resumed runs can skip completed work.
- The final report is surfaced at the workspace root as `Security-Assessment-Report.pdf` and `Security-Assessment-Report.md`. Run internals — deliverables, logs, prompts, and session state — live under a hidden `.shannon/` directory.
- The final report is surfaced at the workspace root as `Security-Assessment-Report.md`. Run internals — deliverables, logs, prompts, and session state — live under a hidden `.shannon/` directory.
> [!NOTE]
> The URL must match the original workspace URL when resuming. Shannon rejects mismatched URLs to prevent cross-target contamination.
@@ -894,10 +787,10 @@ Resume an auto-named workspace:
npx @keygraph/shannon start -u https://example.com -r /path/to/repo -w example-com_shannon-1771007534808
```
List completed scans:
List all workspaces:
```bash
npx @keygraph/shannon scans
npx @keygraph/shannon workspaces
```
Source-build equivalents:
@@ -905,7 +798,7 @@ Source-build equivalents:
```bash
./shannon start -u https://example.com -r /path/to/repo -w my-audit
./shannon start -u https://example.com -r /path/to/repo -w example-com_shannon-1771007534808
./shannon scans
./shannon workspaces
```
---
@@ -963,6 +856,7 @@ For broader coverage, the Keygraph platform adds black-box and white-box agentic
A full test run typically takes roughly 1 to 1.5 hours. LLM API costs vary by model pricing, target complexity, selected provider, and concurrency.
If you use subscription-based model access, consider the rate-limit guidance in [Configuration](configuration.md).
---
Loaded 100 of 103 files, more files were not shown because too many files have changed in this diff. Show more