mirror of
https://github.com/KeygraphHQ/shannon.git
synced 2026-10-07 16:56:51 +02:00
Compare commits
12
Commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
f5054dc885 | ||
|
|
921d702141 | ||
|
|
c2a44b8100 | ||
|
|
45062a8f9e | ||
|
|
f6a770d119 | ||
|
|
685dc6c756 | ||
|
|
84212f376d | ||
|
|
0ab7c0b41b | ||
|
|
a14c7944d8 | ||
|
|
57c511ff8e | ||
|
|
327c10fd90 | ||
|
|
22b093aac5 |
No files matched your search
+5
-5
@@ -9,11 +9,11 @@ SHANNON_AI_MODEL=anthropic:claude-sonnet-4-6
|
||||
|
||||
# --- OpenAI ------------------------------------------------------------------
|
||||
# SHANNON_AI_API_KEY=your-api-key-here
|
||||
# SHANNON_AI_MODEL=openai:gpt-5.5
|
||||
# SHANNON_AI_MODEL=openai:gpt-6-sol
|
||||
|
||||
# --- xAI ---------------------------------------------------------------------
|
||||
# SHANNON_AI_API_KEY=your-api-key-here
|
||||
# SHANNON_AI_MODEL=xai:grok-4.5
|
||||
# SHANNON_AI_MODEL=xai:grok-4.7
|
||||
|
||||
# --- AWS Bedrock -------------------------------------------------------------
|
||||
# Bearer token only; model must be enabled in your region.
|
||||
@@ -30,7 +30,7 @@ SHANNON_AI_MODEL=anthropic:claude-sonnet-4-6
|
||||
# OpenAI Responses API:
|
||||
# SHANNON_AI_API_KEY=your-gateway-key-here
|
||||
# SHANNON_AI_BASE_URL=https://llm-gateway.example.com/v1
|
||||
# SHANNON_AI_MODEL=openai:gpt-5.5
|
||||
# SHANNON_AI_MODEL=openai:gpt-6-sol
|
||||
|
||||
# --- Other provider ----------------------------------------------------------
|
||||
# Any other provider the Pi harness supports. Name it in SHANNON_AI_MODEL and
|
||||
@@ -48,9 +48,9 @@ SHANNON_AI_MODEL=anthropic:claude-sonnet-4-6
|
||||
# See the guide below to use an OpenAI subscription
|
||||
# https://github.com/KeygraphHQ/shannon/blob/main/docs/ai-providers.md#openai-codex-chatgpt-pluspro-subscription
|
||||
# SHANNON_USE_PI_AUTH=1
|
||||
# SHANNON_AI_MODEL=openai-codex:gpt-5.5
|
||||
# SHANNON_AI_MODEL=openai-codex:gpt-6-sol
|
||||
|
||||
# Or the guide below to use an xAI subscription
|
||||
# https://github.com/KeygraphHQ/shannon/blob/main/docs/ai-providers.md#xai-grok-subscription
|
||||
# SHANNON_USE_PI_AUTH=1
|
||||
# SHANNON_AI_MODEL=xai:grok-4.6
|
||||
# SHANNON_AI_MODEL=xai:grok-4.7
|
||||
@@ -87,7 +87,7 @@ pnpm biome:fix # Auto-fix lint, format, and import sorting
|
||||
|
||||
**Monorepo tooling:** pnpm workspaces, Turborepo for task orchestration, Biome for linting/formatting. TypeScript compiler options shared via `tsconfig.base.json` at the root. All packages extend it, overriding only `rootDir` and `outDir`. Shared devDependencies (`typescript`, `@types/node`, `turbo`, `@biomejs/biome`) are hoisted to the root workspace.
|
||||
|
||||
**Options:** `-c <file>` (YAML config), `--models-config <file>` (pi `models.json` defining models pi's catalogue lacks), `-o <path>` (output directory), `-w <name>` (named workspace; auto-resumes if exists), `--pipeline-testing` (minimal prompts, 10s retries), `--keep-container` (preserve worker container after exit for log inspection), `--yes`/`-y` (skip the confirmation prompt on `stop`; required for non-interactive use; `reset` requires a typed `confirm` and cannot be skipped)
|
||||
**Options:** `-c <file>` (YAML config), `--models-config <file>` (pi `models.json` defining models pi's catalogue lacks), `-o <path>` (output directory), `-w <name>` (named workspace; auto-resumes if exists), `--validate-auth` (run preflight and auth validation only, then stop; no pentest or report; requires a fresh workspace and an `authentication` block in the config), `--validate-model` (run the preflight model checks only — credential/registry resolution plus the cyber-access verification for cyber-gated providers — then stop; no pentest or report; requires a fresh workspace; needs no config; mutually exclusive with `--validate-auth`), `--pipeline-testing` (minimal prompts, 10s retries), `--keep-container` (preserve worker container after exit for log inspection), `--yes`/`-y` (skip the confirmation prompt on `stop`; required for non-interactive use; `reset` requires a typed `confirm` and cannot be skipped)
|
||||
|
||||
## Architecture
|
||||
|
||||
@@ -126,7 +126,7 @@ Infra (Temporal) runs via `docker-compose.yml`. Workers are ephemeral `docker ru
|
||||
- `docker-compose.yml` — Infra only: `shannon-temporal` (port 7233/8233). Network: `shannon-net`
|
||||
- `Dockerfile` — 2-stage build (builder + Chainguard Wolfi runtime). Uses pnpm. Entrypoint: `CMD ["node", "apps/worker/dist/temporal/worker.js"]`
|
||||
- No `docker-compose.docker.yml` — host gateway handled via `--add-host` flag in CLI
|
||||
- `/etc/hosts` forwarding — at worker spawn, `forwardEtcHostsFlags` in `apps/cli/src/docker.ts` reads the host's `/etc/hosts` and emits one `--add-host` flag per valid user-added entry. Loopback IPs (`127.x`, `::1`) are rewritten to `host-gateway`; IPv6 addresses are bracketed. Disable per-scan via `SHANNON_FORWARD_HOSTS=false`. No-op on Windows native (WSL2 reads its own `/etc/hosts` via the Linux path).
|
||||
- `/etc/hosts` forwarding — at worker spawn, `forwardEtcHostsFlags` in `apps/cli/src/docker.ts` reads the host's `/etc/hosts` and emits one `--add-host` flag per valid user-added entry. Loopback IPs (`127.x`, `::1`) are rewritten to `host-gateway`; IPv6 addresses are bracketed. Disable per-scan via `SHANNON_FORWARD_HOSTS=false`. Native Windows is refused at startup (`blockNativeWindows` in `apps/cli/src/index.ts`, pointing to the WSL2 guide in `docs/platforms.md`); WSL2 reads its own `/etc/hosts` via the Linux path.
|
||||
|
||||
### Worker Package (`apps/worker/`)
|
||||
- `apps/worker/src/paths.ts` — Centralized path constants (`PROMPTS_DIR`, `CONFIGS_DIR`, `WORKSPACES_DIR`)
|
||||
@@ -165,7 +165,7 @@ Around those phases:
|
||||
- **Configuration** — YAML configs in `apps/worker/configs/` use the closed JSON Schema in `config-schema.json`. Every fresh scan runs the fixed five analysis classes; there is no public class selector. `agentic_sast.enabled` is the only public agentic-SAST setting. Finding reconciliation runs on every scan and has no public setting of its own. Config also supports authentication (MFA/TOTP), URL/code rule scoping (`rules.avoid`/`rules.focus`), `exploit`, free-form `rules_of_engagement`, and post-hoc `report` options (`min_severity`, `min_confidence`, `guidance`, and exploit-only `sarif` output via `apps/worker/src/services/sarif-renderer.ts`, on by default for exploit runs and opt out with `report.sarif: "false"`). `code_path` avoid rules are enforced via the `@gotgenes/pi-permission-system` extension: `apps/worker/src/temporal/activities.ts:syncCodePathDenyRules` writes a global `path` deny config once per workflow (`apps/worker/src/ai/pi/permission-system.ts:syncPermissionSystemConfig`), and the executor loads the extension when that config is present (`apps/worker/src/ai/pi/pi-executor.ts`), so denies fire across every tool and child `task` session. Credential resolution — local mode: env vars → `./.env`; npx mode: env vars → `~/.shannon/config.toml` (via `npx @keygraph/shannon setup`)
|
||||
- **Agentic SAST progress** — Capella runs as a child workflow, so its activities are absent from the parent's `pendingActivities` and invisible to the CLI. The child signals each stage boundary up via `capellaStageProgress` (`apps/worker/src/temporal/shared.ts`); the parent's handler validates the payload and writes the child-supplied `startedAt` and `durationMs` directly to `operationalStages['agentic-sast:<stage>']`, so both the live `getProgress` query and the terminal result carry per-stage rows. Signalling is best-effort and every failure is swallowed — a closed or unreachable parent must never fail a SAST run. `CAPELLA_STAGE_LABELS` in `apps/worker/src/ai/sast/types.ts` is the one label table, shared by the scan log and the status tree; `CAPELLA_PROGRESS_STAGES` omits `export`, which runs no model and so never becomes a row. Scans predating the signal keep the aggregate `agentic-sast` span and render as a bare phase line
|
||||
- **Prompts** — Per-phase templates in `apps/worker/prompts/` with variable substitution (`{{TARGET_URL}}`, `{{CONFIG_CONTEXT}}`). Shared partials in `apps/worker/prompts/shared/` via `apps/worker/src/services/prompt-manager.ts`, including `_code-path-rules.txt` (focus/avoid `[FILE]`/`[GLOB]` routing) and `_rules-of-engagement.txt` (free-text engagement rules). When `exploit: false`, `apps/worker/src/services/findings-renderer.ts` deterministically converts each `*_exploitation_queue.json` into a `*_findings.md` for report assembly — no LLM in the loop
|
||||
- **Agent Harness (pi)** — Uses the **pi harness** (`@earendil-works/pi-coding-agent`, requires Node ≥ 22.19) via `apps/worker/src/ai/pi/pi-executor.ts` (`runPiPrompt` → `createAgentSession`). Retry is split in `apps/worker/src/ai/pi/retry-settings.ts`: pi's agent-level loop is off so Temporal owns agent restarts, while `provider.maxRetries` stays on — pi reads the `provider` block independently of the `enabled` flag — so transport faults are absorbed in-session rather than costing a full agent re-run. `maxRetryDelayMs` is left at pi's 60s default. One model runs every phase, named by `SHANNON_AI_MODEL=<provider>:<model-id>` (default `anthropic:claude-sonnet-4-6`). `apps/worker/src/ai/models.ts` parses the spec — splitting on the **first** colon only, so Bedrock IDs keep theirs — and resolves it through pi's `ModelRuntime`. pi ships the `CredentialStore` interface but no in-memory implementation (its own reads `auth.json` from disk), so `RuntimeCredentialStore` in that file supplies one: credentials arrive as env vars in an ephemeral container and must never touch disk. `createModelRuntime(providerId, apiKey)` builds the runtime; `allowModelNetwork` stays at its default `false` so a scan never blocks on a catalog refresh. `resolveModelSelection()` is **async** because `ModelRuntime.create()` is. Any pi-ai provider id is accepted — `parseModelSpec` no longer rejects against a hardcoded list, so pi's registry is the authority (an unknown provider/model surfaces as a clear "not found in pi registry" error at preflight, which points to the browsable catalogue at `pi.dev/models` — `PI_CATALOG_URL` in `apps/worker/src/ai/models.ts`, appended to the not-found errors and shown in the setup wizard's "Other provider" hint). Four providers are **curated** (`CURATED_PROVIDERS`: `anthropic`, `openai`, `xai`, `amazon-bedrock`) with their own credential variables, config sections, and setup flows; each provider's API key env var is declared once in `PROVIDER_API_KEY_ENV` — Shannon uses each vendor's own variable name (`OPENAI_API_KEY`, `XAI_API_KEY`, …), never an invented one; Bedrock's entry is `AWS_BEARER_TOKEN_BEDROCK`, paired with `AWS_REGION`, which preflight requires separately as provider config rather than a credential. Any other provider uses the **generic** credential path: `SHANNON_AI_API_KEY` (`GENERIC_API_KEY_ENV`) supplies the key for any provider whose credential is a plain API key. Curated providers' own variables take precedence over it, and it also works as a fallback for them — Bedrock is the sole exception (it authenticates through its AWS_ variables, so the generic key never stands in for it). The CLI forwards `SHANNON_AI_API_KEY` in `COMMON_FORWARD_VARS` (it is provider-neutral, binding to whatever `SHANNON_AI_MODEL` names, so the "only one provider configured" guard counts only named credentials), and stores it under a generic `[provider]` config.toml section (`provider.api_key`). `npx @keygraph/shannon setup` exposes this as the "Other provider" option: free-text provider id + model id + key (a curated provider id is rejected there, since it has its own option). A model too new for the pinned pi release does not require an SDK bump: `--models-config <file>` mounts a pi `models.json` read-only at `/app/models.json`. The mount is the entire CLI→worker protocol: nothing is forwarded through the environment, and `modelsConfigPath()` detects the file at that fixed path, exactly as `piAuthPresent()` detects the pi auth mount whose flag is likewise not forwarded (`MODELS_CONFIG_CONTAINER_PATH` in the CLI and `MODELS_CONFIG_PATH` in `apps/worker/src/paths.ts` must stay in sync). `createModelRuntime` always names `modelsPath` explicitly — the mounted path, or **`null` when no config was supplied**, which switches models.json off outright. It is never left to pi's default of `<agent dir>/models.json`, because that dir is shared with the pi auth mount, so a file landing there must not silently contribute model definitions to a scan that did not ask for one. `modelsStorePath` is pinned to the agent dir alongside it, since pi otherwise derives it from `dirname(modelsPath)` and would try to write beside a read-only mount. Custom definitions merge over the built-in catalogue: a matching model id replaces the built-in entry, a new id is added alongside, and `modelOverrides` adjusts a built-in without replacing the provider's list. Omitted fields take pi's defaults (`contextWindow` 128000, `maxTokens` 16384), so a large-context model needs them stated. Credentials are unaffected: `RuntimeCredentialStore` outranks any `apiKey` the file carries, so a snippet pasted from `pi.dev/models` can keep its placeholder key while the real secret stays in the environment — and pi's `!command` config-value form is never reached through `apiKey`. Preflight reports `modelRuntime.getError()` **before** the "model not found" check, because a file that fails to parse, fails schema validation, or composeLine truncated
|
||||
- **Agent Harness (pi)** — Uses the **pi harness** (`@earendil-works/pi-coding-agent`, requires Node ≥ 22.19) via `apps/worker/src/ai/pi/pi-executor.ts` (`runPiPrompt` → `createAgentSession`). Retry is split in `apps/worker/src/ai/pi/retry-settings.ts`: pi's agent-level loop is off so Temporal owns agent restarts, while `provider.maxRetries` stays on — pi reads the `provider` block independently of the `enabled` flag — so transport faults are absorbed in-session rather than costing a full agent re-run. `maxRetryDelayMs` is left at pi's 60s default. One model runs every phase, named by `SHANNON_AI_MODEL=<provider>:<model-id>` (default `anthropic:claude-sonnet-4-6`). `apps/worker/src/ai/models.ts` parses the spec — splitting on the **first** colon only, so Bedrock IDs keep theirs — and resolves it through pi's `ModelRuntime`. pi ships the `CredentialStore` interface but no in-memory implementation (its own reads `auth.json` from disk), so `RuntimeCredentialStore` in that file supplies one: credentials arrive as env vars in an ephemeral container and must never touch disk. `createModelRuntime(providerId, apiKey)` builds the runtime with `allowModelNetwork: true`, so `ModelRuntime.create()` refreshes the model catalogue over the network at scan start and a freshly released model resolves without a `--models-config` file. The fetch is bounded (10s) and falls back to the static catalogue on timeout, so an unreachable catalogue endpoint cannot hang the scan. The refresh does not override a `--models-config`: pi reloads and re-applies that file as a config overlay on every refresh (it reloads `this.config` at the top of `refresh()`), so custom definitions still win over the fetched catalogue; the merge semantics below are unchanged, just layered over a fresher base. `resolveModelSelection()` is **async** because `ModelRuntime.create()` is. Any pi-ai provider id is accepted — `parseModelSpec` no longer rejects against a hardcoded list, so pi's registry is the authority (an unknown provider/model surfaces as a clear "not found in pi registry" error at preflight, which points to the browsable catalogue at `pi.dev/models` — `PI_CATALOG_URL` in `apps/worker/src/ai/models.ts`, appended to the not-found errors and shown in the setup wizard's "Other provider" hint). Four providers are **curated** (`CURATED_PROVIDERS`: `anthropic`, `openai`, `xai`, `amazon-bedrock`) with their own credential variables, config sections, and setup flows; each provider's API key env var is declared once in `PROVIDER_API_KEY_ENV` — Shannon uses each vendor's own variable name (`OPENAI_API_KEY`, `XAI_API_KEY`, …), never an invented one; Bedrock's entry is `AWS_BEARER_TOKEN_BEDROCK`, paired with `AWS_REGION`, which preflight requires separately as provider config rather than a credential. Any other provider uses the **generic** credential path: `SHANNON_AI_API_KEY` (`GENERIC_API_KEY_ENV`) supplies the key for any provider whose credential is a plain API key. Curated providers' own variables take precedence over it, and it also works as a fallback for them — Bedrock is the sole exception (it authenticates through its AWS_ variables, so the generic key never stands in for it). The CLI forwards `SHANNON_AI_API_KEY` in `COMMON_FORWARD_VARS` (it is provider-neutral, binding to whatever `SHANNON_AI_MODEL` names, so the "only one provider configured" guard counts only named credentials), and stores it under a generic `[provider]` config.toml section (`provider.api_key`). `npx @keygraph/shannon setup` exposes this as the "Other provider" option: free-text provider id + model id + key (a curated provider id is rejected there, since it has its own option). A model pi's catalogue does not carry, such as a self-hosted model, is reachable without an SDK bump: `--models-config <file>` mounts a pi `models.json` read-only at `/app/models.json`. The mount is the entire CLI→worker protocol: nothing is forwarded through the environment, and `modelsConfigPath()` detects the file at that fixed path, exactly as `piAuthPresent()` detects the pi auth mount whose flag is likewise not forwarded (`MODELS_CONFIG_CONTAINER_PATH` in the CLI and `MODELS_CONFIG_PATH` in `apps/worker/src/paths.ts` must stay in sync). `createModelRuntime` always names `modelsPath` explicitly — the mounted path, or **`null` when no config was supplied**, which switches models.json off outright. It is never left to pi's default of `<agent dir>/models.json`, because that dir is shared with the pi auth mount, so a file landing there must not silently contribute model definitions to a scan that did not ask for one. `modelsStorePath` is pinned to the agent dir alongside it, since pi otherwise derives it from `dirname(modelsPath)` and would try to write beside a read-only mount. Custom definitions merge over the built-in catalogue: a matching model id replaces the built-in entry, a new id is added alongside, and `modelOverrides` adjusts a built-in without replacing the proviLine truncated
|
||||
- **Pi Credential Reuse** — `SHANNON_USE_PI_AUTH=1` opts into reusing the host's Pi login, including an `openai-codex` ChatGPT Plus/Pro subscription (`SHANNON_AI_MODEL=openai-codex:<model-id>`) or an `xai` Grok subscription (`SHANNON_AI_MODEL=xai:<model-id>`); the mechanism is provider-agnostic and works for any Pi login. `apps/cli/src/env.ts` requires `~/.pi/agent/auth.json`; `start.ts` passes its path to `spawnWorker`, which mounts only that file read-write at `/tmp/.pi/agent/auth.json`. The flag itself is not forwarded: the worker detects the file with `piAuthPresent()` and passes its path to `ModelRuntime.create`. CLI and worker API-key presence checks are skipped on this path, but the normal preflight model probe still validates the credential. The image and UID-remapping entrypoint keep `/tmp/.pi/agent` owned by `pentest` so adjacent Pi/Shannon configuration remains writable. Refreshed OAuth state is persisted to the host for subsequent scans.
|
||||
- **Audit System** — Crash-safe append-only logging in `workspaces/{hostname}_{sessionId}/`. The run directory's top level holds the human-facing report in both formats (`Security-Assessment-Report.pdf` and `Security-Assessment-Report.md`, `FINAL_REPORT_PDF_FILENAME`/`FINAL_REPORT_MD_FILENAME` in `apps/worker/src/paths.ts`); everything else — deliverables, per-agent logs, prompts, `session.json`, `workflow.log`, and browser artifacts — is nested under a hidden `.shannon/` internals dir (`INTERNAL_DIR`) so a customer sees only the report. Audit path helpers route through `generateInternalPath` (`apps/worker/src/audit/utils.ts`); the CLI nests the overlay backing dirs under the same `.shannon/` (`apps/cli/src/docker.ts`, `start.ts`). `session.json`/`workflow.log` reads use dual-read resolvers (`resolveSessionJsonPath`, `resolveRunFile`) that prefer `.shannon/` and fall back to the legacy run-root layout, so pre-restructure workspaces stay listable (`scans`/`logs`) without migration. A pre-restructure workspace cannot be resumed: `classifyWorkspaceLaunch` (`apps/cli/src/commands/start.ts`) requires `.shannon/launch.json`, and its absence fails the launch as "created by an earlier version of Shannon" before anything on disk is touched. There is no in-place migration — the workspace's files and report are left untouched, and the operator starts a new scan under a different `-w` name. The report agent writes structured findings to `report.json`, from which `report-renderer.ts` renders the assembled markdown and `report-json-adapter.ts` produces the Typst-shaped JSON that `pdf-renderer.ts` compiles into `comprehensive_security_assessment_report.pdf` using the bundled `apps/worker/templates/typst/report.typ` template (the `typst` binary is installed in the worker image). `copyReportToRunRoot` (`apps/worker/src/services/reporting.ts`) surfaces both the PDF and the markdown to the run root as `Security-Assessment-Report.pdf` and `Security-Assessment-Report.md`; the deliverables-dir copies remain as the git-checkpointed sources. PDF compilation is best-effort — a failure is logged and the run still completes. WorkflowLogger (`apps/worker/src/audit/workflow-logger.ts`) provides unified human-readable per-workflow logs, backed by LogStream (`apps/worker/src/audit/log-stream.ts`) shared stream primitive. Every combined-log line is also projected into a per-agent file under `.shannon/agents/<slug>.log` (one per pipeline agent, one per Capella stage; subagents fold into the parent's file, and a stage's concurrent sessions share its file with an inline session label). The projection boundary is `apps/worker/src/audit/actor-projection.ts` (`projectActor` maps a `TraceActor` to its combined prefix and owning file slug — slugs come only from closed fields); fan-out is best-effort and never blocks the canonical combined log. A lifecycle owner holds a `LogStream` lease per agent file (the pipeline agent's `logAgent` span, or a Capella stage activity's `try/finally`) so per-line writes ride the reference count; `CapellaStageTrace.drain()` flushes a stage's trace queue before its activity returns. The CLI tails one file with `shannon logs --agent <name>` (`--list-agents` to enumerate); the default `shannon logs` path is unchanged
|
||||
- **Deliverables** — Saved to `.shannon/deliverables/` in the target repo via the `save-deliverable` CLI script (`apps/worker/src/scripts/save-deliverable.ts`)
|
||||
|
||||
@@ -1,5 +1,5 @@
|
||||
> [!NOTE]
|
||||
> **[Shannon 3.0 is live](https://github.com/KeygraphHQ/shannon/discussions/439):** deeper security code analysis, more thoroughly vetted findings, a rebuilt CLI, native CI/CD, professional PDF reports, and SARIF.
|
||||
> **[Cyber verification for Anthropic and OpenAI models](https://github.com/KeygraphHQ/shannon/discussions/483):** complete your provider's cyber verification program to prevent model failures and refusals during cyber workloads.
|
||||
|
||||
<div align="center">
|
||||
|
||||
@@ -76,7 +76,7 @@ Shannon is an autonomous AI pentester developed by [Keygraph](https://keygraph.i
|
||||
|
||||
Shannon analyzes your web application's source code to identify potential attack vectors, then uses browser automation and command-line tools to execute real exploits against the running application and its APIs. Only vulnerabilities with a working proof-of-concept are included in the final report.
|
||||
|
||||
Shannon is the agent. This repository is Shannon Open Source, the standalone pentester you run yourself. The same Shannon also powers the [Keygraph platform](https://keygraph.io), Keygraph's commercial pentesting product. See [Editions](#editions) for how the two compare.
|
||||
Shannon is the agent. This repository is Shannon Open Source, the standalone pentester you run yourself. The same Shannon also powers the [Keygraph platform](https://keygraph.io), Keygraph's commercial pentesting product. See [Editions](#editions) for how Shannon Open Source compares with the platform's Community Program, Pro, and Enterprise editions.
|
||||
|
||||
<a id="why-shannon-exists"></a>
|
||||
<details>
|
||||
@@ -131,8 +131,8 @@ These reports are from Shannon Open Source scans of Photoview 2.4.0, one of the
|
||||
|
||||
- **Docker**: required for the worker container.
|
||||
- **Node.js 18+**: required for the recommended `npx` workflow.
|
||||
- **AI provider credentials**: Shannon runs on Anthropic, OpenAI, xAI, AWS Bedrock, and [any other provider](docs/ai-providers.md#any-other-provider) in the harness catalogue — each of which you can point at a proxy or LLM gateway through a [custom base URL](docs/ai-providers.md#custom-base-url), and a model the catalogue does not yet carry can be described with a [custom model configuration](docs/ai-providers.md#custom-model-configuration). You bring your own key, and Keygraph never proxies your model traffic. Shannon is provider-agnostic. See [AI providers](docs/ai-providers.md#suggested-models) for suggested model IDs.
|
||||
- **Cyber safeguards cleared with your provider**: Anthropic and OpenAI apply real-time safeguards to cyber-security workloads, which can interrupt a scan mid-run. Complete their guidance for legitimate security testers before your first run - see [AI providers](docs/ai-providers.md#cyber-safeguards-do-this-before-your-first-scan).
|
||||
- **AI provider credentials**: Shannon runs on Anthropic, OpenAI, xAI, AWS Bedrock, and [any other provider](docs/ai-providers.md#any-other-provider) in the harness catalogue — each of which you can point at a proxy or LLM gateway through a [custom base URL](docs/ai-providers.md#custom-base-url), and a model the catalogue does not carry can be described with a [custom model configuration](docs/ai-providers.md#custom-model-configuration). You bring your own key, and Keygraph never proxies your model traffic. Shannon is provider-agnostic. See [AI providers](docs/ai-providers.md#suggested-models) for suggested model IDs.
|
||||
- **Cyber safeguards cleared with your provider**: Anthropic and OpenAI apply real-time safeguards to cyber-security workloads, which can interrupt a scan mid-run. Complete their guidance for legitimate security testers before your first run - see [AI providers](docs/ai-providers.md#cyber-safeguards-do-this-before-your-first-scan) and the [cyber verification announcement](https://github.com/KeygraphHQ/shannon/discussions/483).
|
||||
|
||||
|
||||
|
||||
@@ -235,11 +235,28 @@ See the [Shannon GitHub Action documentation](https://github.com/KeygraphHQ/shan
|
||||
|
||||
## Editions
|
||||
|
||||
**Shannon Open Source** is a complete autonomous pentester, especially well suited to individual developers and small teams running focused security tests locally or in CI/CD.
|
||||
Shannon Lite is now Shannon Open Source. Shannon Pro is now the Keygraph platform.
|
||||
|
||||
**Keygraph Enterprise Platform** is for organizations that need a shared platform for continuous agentic pentesting/AppSec across many teams, repositories, and environments. It centralizes deeper analysis, vulnerability management, remediation, verification, governance, and reporting so teams do not have to assemble and maintain those workflows themselves.
|
||||
| Edition | What it is | Price |
|
||||
| --- | --- | --- |
|
||||
| **Shannon Open Source** (Shannon OSS) | This repository. A complete autonomous pentester you run yourself, locally or in CI/CD, against an application whose source code you have. Well suited to individual developers and small teams. | Open source under AGPL-3.0. You pay only your own model costs. |
|
||||
| **Community Program** | The full Keygraph platform, cloud-hosted, for U.S.-based 501(c)(3) nonprofits and for seed or pre-Series-A startups with 20 or fewer active developers. | $0 in cloud service fees while you qualify. |
|
||||
| **Pro** | The full Keygraph platform, cloud-hosted and managed by Keygraph, with every module included. | $50 per active developer per month. |
|
||||
| **Enterprise** | Everything in Pro, self-hosted in your own environment or fully air-gapped. | Custom. |
|
||||
|
||||
[Learn about the Keygraph Enterprise Platform and compare editions →](docs/keygraph-platform.md)
|
||||
See [keygraph.io/pricing](https://keygraph.io/pricing) for current prices and the full feature table.
|
||||
|
||||
The **Keygraph platform** runs an enterprise-hardened fork of Shannon. The Community Program, Pro, and Enterprise all add:
|
||||
|
||||
- **Black-box pentesting**: tests the running application from the outside, with no source code needed.
|
||||
- **Dependency (SCA) and secret checks**: SCA with reachability, and secrets scanning that includes repository history.
|
||||
- **Findings management**: one record per vulnerability per repository across scans and scanners, with status history, owners, SLAs, and dashboards.
|
||||
- **Fix pull requests and retests**: reviewable fix pull requests, with each fix verified by re-analysis and exploit replay, without a full rescan.
|
||||
- **Jira and access control**: two-way Jira sync, SSO, RBAC, and audit logs.
|
||||
|
||||
No source code? Use the [Blackbox Pentester](https://keygraph.io/agentic-blackbox-pentester).
|
||||
|
||||
[Learn about the Keygraph platform and compare editions →](docs/keygraph-platform.md)
|
||||
|
||||
## Architecture
|
||||
|
||||
@@ -291,7 +308,7 @@ Use these guides for operational detail:
|
||||
| [Workspaces and resuming](docs/workspaces.md) | Naming workspaces, resuming interrupted scans, and workspace storage. |
|
||||
| [Safety and limitations](docs/safety.md) | Authorized-use requirements, non-production guidance, mutative effects, cost, and model caveats. |
|
||||
| [Coverage and roadmap](docs/coverage-roadmap.md) | Current vulnerability coverage and planned work. |
|
||||
| [Keygraph Enterprise Platform](docs/keygraph-platform.md) | Exhaustive agentic SAST, continuous pentesting, full-lifecycle finding management, remediation, targeted verification, enterprise governance, and on-premises deployment. |
|
||||
| [Keygraph platform](docs/keygraph-platform.md) | Shannon Open Source compared with the Community Program, Pro, and Enterprise editions: black-box pentesting, SCA and secrets, findings management, fix pull requests, retests, Jira, and deployment. |
|
||||
|
||||
|
||||
|
||||
@@ -304,7 +321,7 @@ You are responsible for using Shannon legally and ethically. Do not point Shanno
|
||||
|
||||
Important limitations:
|
||||
|
||||
- Shannon Open Source is tuned for fast, code-informed pentesting in everyday development and CI/CD. Exhaustive agentic SAST, broader scanner coverage, centralized governance, and full-lifecycle vulnerability management are delivered through the Keygraph Enterprise Platform.
|
||||
- Shannon Open Source is tuned for fast, code-informed pentesting in everyday development and CI/CD. Exhaustive agentic SAST, broader scanner coverage, centralized governance, and full-lifecycle vulnerability management are delivered through the Keygraph platform, in its Community Program, Pro, and Enterprise editions.
|
||||
- Findings still require human review. LLM-generated reports can contain weakly supported or incorrect details.
|
||||
- Anthropic, OpenAI, xAI, and AWS Bedrock are built-in providers, and any other provider in the harness catalogue works too — each reachable through a custom base URL that points it at a proxy or LLM gateway. Model capability varies, and a model that does not follow Shannon's instructions or tool-use constraints reliably will produce weaker results.
|
||||
- A full run can take roughly 1 to 1.5 hours and may incur LLM API costs depending on model pricing and application complexity.
|
||||
@@ -375,7 +392,7 @@ Yes. Shannon emits SARIF 2.1.0, the OASIS standard format for static analysis re
|
||||
|
||||
### Which AI providers does Shannon support?
|
||||
|
||||
Anthropic, OpenAI, xAI, and AWS Bedrock are built in and configured directly by provider ID. Beyond those, Shannon runs on any provider in the Pi harness catalogue, named the same `<provider>:<model-id>` way. Any provider can be pointed at a proxy or LLM gateway through a custom base URL, which overrides only the endpoint and keeps that provider's API dialect. A model the catalogue does not yet carry, such as one released after Shannon's pinned harness version, runs without waiting for a Shannon release. Describe it in a [custom model configuration](docs/ai-providers.md#custom-model-configuration) file and pass it with `--models-config`. Shannon uses a single unified model setting throughout a pentest.
|
||||
Anthropic, OpenAI, xAI, and AWS Bedrock are built in and configured directly by provider ID. Beyond those, Shannon runs on any provider in the Pi harness catalogue, named the same `<provider>:<model-id>` way. Any provider can be pointed at a proxy or LLM gateway through a custom base URL, which overrides only the endpoint and keeps that provider's API dialect. A model the catalogue does not carry, such as one a router or gateway serves under its own ID, or a self-hosted model, is described in a [custom model configuration](docs/ai-providers.md#custom-model-configuration) file and passed with `--models-config`. Shannon uses a single unified model setting throughout a pentest.
|
||||
|
||||
### Can I run Shannon on a local or self-hosted model?
|
||||
|
||||
|
||||
@@ -21,7 +21,15 @@ import { resolveWorkflowId } from '../session.js';
|
||||
import { waitForWorkflowClose } from '../temporal-client.js';
|
||||
import { stdoutIsTerminal } from '../tty.js';
|
||||
|
||||
const TERMINAL_HEADINGS = new Set(['Scan COMPLETED', 'Scan PARTIAL', 'Scan FAILED', 'Scan CANCELLED']);
|
||||
const TERMINAL_HEADINGS = new Set([
|
||||
'Scan COMPLETED',
|
||||
'Scan PARTIAL',
|
||||
'Scan FAILED',
|
||||
'Scan CANCELLED',
|
||||
'Validation COMPLETED',
|
||||
'Validation FAILED',
|
||||
'Validation CANCELLED',
|
||||
]);
|
||||
|
||||
// The combined log resets completion on the bare `RESUMED` heading; a per-agent file carries the
|
||||
// distinct `--- RESUMED (<workflow id>) ---` boundary that WorkflowLogger.logResumeBoundary writes
|
||||
@@ -48,7 +56,7 @@ export class LogCompletionState {
|
||||
this.failureIsLastMarker = false;
|
||||
} else if (TERMINAL_HEADINGS.has(line)) {
|
||||
this.terminalIsLastMarker = true;
|
||||
this.failureIsLastMarker = line === 'Scan FAILED';
|
||||
this.failureIsLastMarker = line.endsWith('FAILED');
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
@@ -36,17 +36,24 @@ const GATEWAY_DIALECTS: readonly {
|
||||
|
||||
/** Suggested models per curated provider, best-first. Free-text entry accepts any model in the provider's catalogue. */
|
||||
const MODEL_SUGGESTIONS: Readonly<Record<CuratedProviderId, readonly string[]>> = {
|
||||
anthropic: ['claude-sonnet-4-6', 'claude-opus-4-8', 'claude-opus-4-7', 'claude-haiku-4-5-20251001'],
|
||||
openai: ['gpt-5.6-sol', 'gpt-5.5', 'gpt-5.4'],
|
||||
xai: ['grok-4.5'],
|
||||
anthropic: [
|
||||
'claude-sonnet-5',
|
||||
'claude-opus-5',
|
||||
'claude-sonnet-4-6',
|
||||
'claude-opus-4-8',
|
||||
'claude-opus-4-7',
|
||||
'claude-haiku-4-5-20251001',
|
||||
],
|
||||
openai: ['gpt-6-sol', 'gpt-5.6-sol', 'gpt-5.5', 'gpt-5.4'],
|
||||
xai: ['grok-4.7'],
|
||||
'amazon-bedrock': ['us.anthropic.claude-sonnet-4-6', 'us.anthropic.claude-opus-4-8', 'us.anthropic.claude-opus-4-7'],
|
||||
};
|
||||
|
||||
/** Placeholder shown in the free-text model ID prompt, per curated provider. */
|
||||
const MODEL_ID_PLACEHOLDER: Readonly<Record<CuratedProviderId, string>> = {
|
||||
anthropic: 'claude-sonnet-4-6',
|
||||
openai: 'gpt-5.6-sol',
|
||||
xai: 'grok-4.5',
|
||||
openai: 'gpt-6-sol',
|
||||
xai: 'grok-4.7',
|
||||
'amazon-bedrock': 'us.anthropic.claude-opus-4-8',
|
||||
};
|
||||
|
||||
|
||||
+144
-28
@@ -31,7 +31,12 @@ import { clearPendingWorkflowIdentity, writePendingWorkflowIdentity } from '../p
|
||||
import { indentFailureSegments, parseFailureSegments } from '../scan/failure.js';
|
||||
import { resolveWorkflowId } from '../session.js';
|
||||
import { displayPlainBanner, displaySplash } from '../splash.js';
|
||||
import { describeWorkflowLifecycle, getTerminalOutcome, queryProgress } from '../temporal-client.js';
|
||||
import {
|
||||
describeWorkflowLifecycle,
|
||||
getTerminalOutcome,
|
||||
queryProgress,
|
||||
runningActivityTypes,
|
||||
} from '../temporal-client.js';
|
||||
import { stdoutIsTerminal } from '../tty.js';
|
||||
import { tailUntilComplete } from './logs.js';
|
||||
|
||||
@@ -45,6 +50,8 @@ export interface StartArgs {
|
||||
pipelineTesting: boolean;
|
||||
keepContainer: boolean;
|
||||
follow: boolean;
|
||||
authOnly: boolean;
|
||||
validateModel: boolean;
|
||||
version: string;
|
||||
}
|
||||
|
||||
@@ -60,6 +67,10 @@ const FIXED_CLASSES = ['injection', 'xss', 'auth', 'authz', 'ssrf'] as const;
|
||||
interface LaunchState {
|
||||
readonly schema_version: typeof LAUNCH_STATE_SCHEMA_VERSION;
|
||||
readonly customer_output_path?: string;
|
||||
/** True when the workspace was created by an auth-validation run; such a workspace is not a scan. */
|
||||
readonly auth_only?: boolean;
|
||||
/** True when the workspace was created by a model-validation run; such a workspace is not a scan. */
|
||||
readonly model_only?: boolean;
|
||||
}
|
||||
|
||||
export interface WorkspaceLaunchDecision {
|
||||
@@ -124,17 +135,31 @@ function readLaunchState(filePath: string): LaunchState {
|
||||
if (!isRecord(value)) fail(NEWER_RELEASE_MESSAGE);
|
||||
// Unknown keys mean a newer release wrote this workspace; refuse rather than half-read it.
|
||||
const keys = Object.keys(value).sort();
|
||||
const keysAreValid = keys.every((key) => key === 'customer_output_path' || key === 'schema_version');
|
||||
const keysAreValid = keys.every(
|
||||
(key) => key === 'auth_only' || key === 'model_only' || key === 'customer_output_path' || key === 'schema_version',
|
||||
);
|
||||
const customerPath = value.customer_output_path;
|
||||
const pathIsValid =
|
||||
customerPath === undefined ||
|
||||
(typeof customerPath === 'string' && path.isAbsolute(customerPath) && path.resolve(customerPath) === customerPath);
|
||||
if (value.schema_version !== LAUNCH_STATE_SCHEMA_VERSION || !keysAreValid || !pathIsValid) {
|
||||
const authOnly = value.auth_only;
|
||||
const authOnlyIsValid = authOnly === undefined || typeof authOnly === 'boolean';
|
||||
const modelOnly = value.model_only;
|
||||
const modelOnlyIsValid = modelOnly === undefined || typeof modelOnly === 'boolean';
|
||||
if (
|
||||
value.schema_version !== LAUNCH_STATE_SCHEMA_VERSION ||
|
||||
!keysAreValid ||
|
||||
!pathIsValid ||
|
||||
!authOnlyIsValid ||
|
||||
!modelOnlyIsValid
|
||||
) {
|
||||
fail(NEWER_RELEASE_MESSAGE);
|
||||
}
|
||||
return {
|
||||
schema_version: LAUNCH_STATE_SCHEMA_VERSION,
|
||||
...(typeof customerPath === 'string' && { customer_output_path: customerPath }),
|
||||
...(authOnly === true && { auth_only: true }),
|
||||
...(modelOnly === true && { model_only: true }),
|
||||
};
|
||||
}
|
||||
|
||||
@@ -149,6 +174,8 @@ export function classifyWorkspaceLaunch(
|
||||
workspacePath: string,
|
||||
expectedUrl: string,
|
||||
requestedOutputDir: string | undefined,
|
||||
requestedAuthOnly: boolean,
|
||||
requestedModelOnly: boolean,
|
||||
): WorkspaceLaunchDecision {
|
||||
const sessionPath = resolveRunFile(workspacePath, 'session.json');
|
||||
const sessionExists = fs.existsSync(sessionPath);
|
||||
@@ -163,6 +190,16 @@ export function classifyWorkspaceLaunch(
|
||||
|
||||
const launchPath = path.join(workspacePath, INTERNAL_DIR, LAUNCH_STATE_FILENAME);
|
||||
const launch = readLaunchState(launchPath);
|
||||
if (launch.auth_only && !requestedAuthOnly) {
|
||||
fail(
|
||||
'This workspace was created to validate authentication only, so it cannot be run as a scan. Start a new scan with a different -w name.',
|
||||
);
|
||||
}
|
||||
if (launch.model_only && !requestedModelOnly) {
|
||||
fail(
|
||||
'This workspace was created to validate the AI model only, so it cannot be run as a scan. Start a new scan with a different -w name.',
|
||||
);
|
||||
}
|
||||
const session = readJsonFile(sessionPath);
|
||||
if (!isRecord(session) || !isRecord(session.session) || session.session.webUrl !== expectedUrl) {
|
||||
fail(
|
||||
@@ -190,12 +227,19 @@ export function classifyWorkspaceLaunch(
|
||||
* host crash. Callers invoke this only for a fresh workspace; an existing launch.json is
|
||||
* the resume contract and must never be replaced.
|
||||
*/
|
||||
export function writeLaunchStateAtomically(internalPath: string, outputDir: string | undefined): void {
|
||||
export function writeLaunchStateAtomically(
|
||||
internalPath: string,
|
||||
outputDir: string | undefined,
|
||||
authOnly: boolean,
|
||||
modelOnly: boolean,
|
||||
): void {
|
||||
const finalPath = path.join(internalPath, LAUNCH_STATE_FILENAME);
|
||||
const temporaryPath = path.join(internalPath, `${LAUNCH_STATE_FILENAME}.tmp-${process.pid}-${randomSuffix()}`);
|
||||
const launchState: LaunchState = {
|
||||
schema_version: LAUNCH_STATE_SCHEMA_VERSION,
|
||||
...(outputDir !== undefined && { customer_output_path: outputDir }),
|
||||
...(authOnly && { auth_only: true }),
|
||||
...(modelOnly && { model_only: true }),
|
||||
};
|
||||
const descriptor = fs.openSync(temporaryPath, 'wx', 0o600);
|
||||
try {
|
||||
@@ -225,6 +269,10 @@ export function createWorkflowId(workspace: string, isResume: boolean, timestamp
|
||||
}
|
||||
|
||||
export async function start(args: StartArgs): Promise<void> {
|
||||
// Validation-only runs are short and have no report to come back for, so they always stream to the end.
|
||||
const validationOnly = args.authOnly || args.validateModel;
|
||||
if (validationOnly) args.follow = true;
|
||||
|
||||
// 1. Resolve non-mutating inputs and classify the workspace before changing it.
|
||||
initHome();
|
||||
loadEnv();
|
||||
@@ -240,7 +288,38 @@ export async function start(args: StartArgs): Promise<void> {
|
||||
args.workspace ?? `${new URL(args.url).hostname.replace(/[^a-zA-Z0-9-]/g, '-')}_shannon-${Date.now()}`;
|
||||
const workspacePath = path.join(workspacesDir, workspace);
|
||||
const requestedOutputDir = args.output ? path.resolve(expandHome(args.output)) : undefined;
|
||||
const launchDecision = classifyWorkspaceLaunch(workspacePath, args.url, requestedOutputDir);
|
||||
const launchDecision = classifyWorkspaceLaunch(
|
||||
workspacePath,
|
||||
args.url,
|
||||
requestedOutputDir,
|
||||
args.authOnly,
|
||||
args.validateModel,
|
||||
);
|
||||
|
||||
// Validation-only runs write no resumable state, so they always run fresh; reusing a workspace would resume it.
|
||||
if (validationOnly && launchDecision.isResume) {
|
||||
const what = args.authOnly ? 'An auth-validation run' : 'A model-validation run';
|
||||
fail(`${what} needs a fresh workspace. Omit -w to auto-name one, or choose a -w name that is not in use.`);
|
||||
}
|
||||
|
||||
// User-facing status wording. Auth-only and model-only are both "validation" runs, but each
|
||||
// names what it validated. A validation run *is* the checks, so a failure means it ran and
|
||||
// failed, not that it could not start. A plain scan keeps its original phrasing.
|
||||
let startingLabel = 'Starting scan';
|
||||
let waitingLabel = 'Waiting for the scan to start';
|
||||
let couldNotStartLabel = 'The scan could not start';
|
||||
let startedLabel = `Scan started — ${workspace}`;
|
||||
if (args.authOnly) {
|
||||
startingLabel = 'Starting authentication validation';
|
||||
waitingLabel = 'Waiting for authentication validation to start';
|
||||
couldNotStartLabel = 'Authentication validation failed';
|
||||
startedLabel = `Validating authentication — ${workspace}`;
|
||||
} else if (args.validateModel) {
|
||||
startingLabel = 'Starting model validation';
|
||||
waitingLabel = 'Waiting for model validation to start';
|
||||
couldNotStartLabel = 'Model validation failed';
|
||||
startedLabel = `Validating model — ${workspace}`;
|
||||
}
|
||||
|
||||
// 2. Inputs are valid; identify the run before initializing shared infrastructure.
|
||||
const bannerVersion = isLocal() ? undefined : args.version;
|
||||
@@ -254,7 +333,7 @@ export async function start(args: StartArgs): Promise<void> {
|
||||
ensureDocker();
|
||||
ensureImage(args.version);
|
||||
const spinner = p.spinner();
|
||||
spinner.start('Starting scan');
|
||||
spinner.start(startingLabel);
|
||||
await ensureInfra(spinner);
|
||||
|
||||
// 3. Generate the invocation identity.
|
||||
@@ -277,7 +356,7 @@ export async function start(args: StartArgs): Promise<void> {
|
||||
fs.chmodSync(dirPath, 0o777);
|
||||
}
|
||||
if (!launchDecision.isResume) {
|
||||
writeLaunchStateAtomically(internalPath, launchDecision.outputDir);
|
||||
writeLaunchStateAtomically(internalPath, launchDecision.outputDir, args.authOnly, args.validateModel);
|
||||
}
|
||||
|
||||
// 5. Pre-create overlay mount points (:ro mounts cannot create them).
|
||||
@@ -336,6 +415,8 @@ export async function start(args: StartArgs): Promise<void> {
|
||||
workspace,
|
||||
...(args.pipelineTesting && { pipelineTesting: true }),
|
||||
...(args.keepContainer && { keepContainer: true }),
|
||||
...(args.authOnly && { authOnly: true }),
|
||||
...(args.validateModel && { validateModel: true }),
|
||||
...(shouldUsePiAuth() && { piAuthHostPath: resolveHostPiAuthPath() }),
|
||||
});
|
||||
|
||||
@@ -386,7 +467,7 @@ export async function start(args: StartArgs): Promise<void> {
|
||||
});
|
||||
|
||||
// Poll for the workflow to register in session.json; the spinner resolves once it does.
|
||||
spinner.message('Waiting for the scan to start');
|
||||
spinner.message(waitingLabel);
|
||||
for (let attempts = 0; attempts < 60; attempts++) {
|
||||
// A pre-workflow failure leaves its reason here (nothing reached Temporal); surface it
|
||||
// rather than polling out to a generic timeout.
|
||||
@@ -415,20 +496,28 @@ export async function start(args: StartArgs): Promise<void> {
|
||||
warn(`Scan ${workspace} started, but its launch record could not be removed.`);
|
||||
}
|
||||
|
||||
// Hold until preflight clears, so an unreachable target or bad credential is reported here
|
||||
// Hold until startup clears, so an unreachable target or bad credential is reported here
|
||||
// rather than after "Scan started".
|
||||
spinner.message('Running preflight checks');
|
||||
const outcome = await awaitPreflightOutcome(workflowId);
|
||||
spinner.message(PREFLIGHT_LABEL);
|
||||
const spec = resolveModelSpec();
|
||||
const providerId = typeof spec === 'string' ? '' : spec.providerId;
|
||||
// Cyber-access verification only runs for OpenAI/Anthropic; when following, the tailed log shows the login.
|
||||
// Mirrors CYBER_GATED_PROVIDERS in the worker (apps/worker/src/services/cyber-access-verification.ts).
|
||||
const showCyberAccess = providerId === 'anthropic' || providerId === 'openai' || providerId === 'openai-codex';
|
||||
const outcome = await awaitStartupOutcome(workflowId, (label) => spinner.message(label), {
|
||||
showCyberAccess,
|
||||
showAppLogin: !args.follow,
|
||||
});
|
||||
if (outcome.kind === 'failed') {
|
||||
spinner.error('The scan could not start');
|
||||
spinner.error(couldNotStartLabel);
|
||||
printScanStartFailure(outcome.message);
|
||||
process.exit(1);
|
||||
}
|
||||
|
||||
spinner.stop(`Scan started — ${workspace}`);
|
||||
spinner.stop(startedLabel);
|
||||
printInfo(args, workspace, repo.hostPath, workspacesDir);
|
||||
if (args.follow) {
|
||||
await followScan(workspace, workspacesDir);
|
||||
await followScan(workspace, workspacesDir, validationOnly);
|
||||
}
|
||||
return;
|
||||
}
|
||||
@@ -494,15 +583,28 @@ function readStartupError(startupErrorPath: string): StartupError | undefined {
|
||||
}
|
||||
}
|
||||
|
||||
/** Outcome of waiting for the in-workflow preflight to clear. */
|
||||
/** Outcome of waiting for in-workflow startup (preflight + auth validation) to clear. */
|
||||
type PreflightOutcome = { kind: 'passed' } | { kind: 'failed'; message: string } | { kind: 'unconfirmed' };
|
||||
|
||||
const PREFLIGHT_LABEL = 'Running preflight checks (LLM credentials, target URL)';
|
||||
const CYBER_ACCESS_LABEL = 'Checking cyber access';
|
||||
const APP_LOGIN_LABEL = 'Verifying app login with provided credentials';
|
||||
|
||||
/**
|
||||
* Wait for the registered workflow's preflight to pass or fail: passed once `currentPhase` moves
|
||||
* beyond 'preflight' (or the scan already closed ok), failed when the workflow terminates with an
|
||||
* error. Bounded, so a Temporal query outage falls through as 'unconfirmed' rather than hanging.
|
||||
* Drive the startup spinner until the pentest begins, naming the cyber-access verification and the app
|
||||
* login while their activity runs. Labels only advance, so a gap between them holds the last step
|
||||
* rather than reverting to the generic line. Passed once the phase moves past preflight/auth (or
|
||||
* the scan closed ok), failed on a terminal error, unconfirmed if a query outage outlasts the bound.
|
||||
*/
|
||||
async function awaitPreflightOutcome(workflowId: string): Promise<PreflightOutcome> {
|
||||
async function awaitStartupOutcome(
|
||||
workflowId: string,
|
||||
onLabel: (label: string) => void,
|
||||
opts: { showCyberAccess: boolean; showAppLogin: boolean },
|
||||
): Promise<PreflightOutcome> {
|
||||
// Wait through auth-validation only when naming the login step; otherwise stop once it begins.
|
||||
const startupPhases = opts.showAppLogin ? new Set(['preflight', 'auth-validation']) : new Set(['preflight']);
|
||||
let rank = 0;
|
||||
let label = PREFLIGHT_LABEL;
|
||||
for (let attempts = 0; attempts < 80; attempts++) {
|
||||
try {
|
||||
const lifecycle = await describeWorkflowLifecycle(workflowId);
|
||||
@@ -511,8 +613,19 @@ async function awaitPreflightOutcome(workflowId: string): Promise<PreflightOutco
|
||||
return outcome.kind === 'failed' ? { kind: 'failed', message: outcome.message } : { kind: 'passed' };
|
||||
}
|
||||
|
||||
const running = await runningActivityTypes(workflowId);
|
||||
if (opts.showCyberAccess && rank < 1 && running.includes('runCyberAccessVerification')) {
|
||||
rank = 1;
|
||||
label = CYBER_ACCESS_LABEL;
|
||||
}
|
||||
if (opts.showAppLogin && rank < 2 && running.includes('runAuthenticationValidation')) {
|
||||
rank = 2;
|
||||
label = APP_LOGIN_LABEL;
|
||||
}
|
||||
onLabel(label);
|
||||
|
||||
const progress = await queryProgress(workflowId);
|
||||
if (progress && progress.currentPhase !== null && progress.currentPhase !== 'preflight') {
|
||||
if (progress && progress.currentPhase !== null && !startupPhases.has(progress.currentPhase)) {
|
||||
return { kind: 'passed' };
|
||||
}
|
||||
} catch {
|
||||
@@ -576,7 +689,7 @@ function printUnconfirmedScanHint(workspace: string, taskQueue: string, containe
|
||||
* That tracks whether the pipeline ran, not whether vulnerabilities were found. On failure the
|
||||
* root-cause message is printed so a red CI build says why.
|
||||
*/
|
||||
async function followScan(workspace: string, workspacesDir: string): Promise<never> {
|
||||
async function followScan(workspace: string, workspacesDir: string, validationOnly = false): Promise<never> {
|
||||
const logFile = resolveRunFile(path.join(workspacesDir, workspace), 'workflow.log');
|
||||
const workflowId = resolveWorkflowId(workspace);
|
||||
|
||||
@@ -587,7 +700,8 @@ async function followScan(workspace: string, workspacesDir: string): Promise<nev
|
||||
}
|
||||
|
||||
if (stdoutIsTerminal()) {
|
||||
console.error('\n Following scan log (Ctrl-C to stop watching):\n');
|
||||
const what = validationOnly ? 'validation' : 'scan';
|
||||
console.error(`\n Following ${what} log (Ctrl-C to stop watching):\n`);
|
||||
}
|
||||
|
||||
let temporalUnreachable = false;
|
||||
@@ -675,10 +789,12 @@ function printInfo(args: StartArgs, workspace: string, repoPath: string, workspa
|
||||
console.log(` Progress: ${prefix} status ${workspace}`);
|
||||
}
|
||||
|
||||
console.log('');
|
||||
console.log(' Report (when the scan finishes):');
|
||||
console.log(` ${reportDir}${path.sep}`);
|
||||
console.log(` ${FINAL_REPORT_PDF_FILENAME}`);
|
||||
console.log(` ${FINAL_REPORT_MD_FILENAME}`);
|
||||
console.log('');
|
||||
if (!args.authOnly && !args.validateModel) {
|
||||
console.log('');
|
||||
console.log(' Report (when the scan finishes):');
|
||||
console.log(` ${reportDir}${path.sep}`);
|
||||
console.log(` ${FINAL_REPORT_PDF_FILENAME}`);
|
||||
console.log(` ${FINAL_REPORT_MD_FILENAME}`);
|
||||
console.log('');
|
||||
}
|
||||
}
|
||||
@@ -96,15 +96,12 @@ function loadTOML(): TOMLConfig | null {
|
||||
if (!fs.existsSync(configPath)) return null;
|
||||
|
||||
// Config contains secrets — refuse to read if group or others have any access.
|
||||
// Skip on Windows where POSIX permissions are not supported.
|
||||
if (process.platform !== 'win32') {
|
||||
const mode = fs.statSync(configPath).mode;
|
||||
if (mode & 0o077) {
|
||||
const actual = (mode & 0o777).toString(8).padStart(3, '0');
|
||||
fail(
|
||||
`Your config file is readable by other users on this machine (${actual}). Lock it down: chmod 600 ${configPath}`,
|
||||
);
|
||||
}
|
||||
const mode = fs.statSync(configPath).mode;
|
||||
if (mode & 0o077) {
|
||||
const actual = (mode & 0o777).toString(8).padStart(3, '0');
|
||||
fail(
|
||||
`Your config file is readable by other users on this machine (${actual}). Lock it down: chmod 600 ${configPath}`,
|
||||
);
|
||||
}
|
||||
|
||||
try {
|
||||
|
||||
@@ -360,7 +360,6 @@ function shouldSkipHostsName(name: string, hostname: string): boolean {
|
||||
*/
|
||||
function forwardEtcHostsFlags(): string[] {
|
||||
if (!envBool('SHANNON_FORWARD_HOSTS', true)) return [];
|
||||
if (os.platform() === 'win32') return [];
|
||||
|
||||
let content: string;
|
||||
try {
|
||||
@@ -413,6 +412,8 @@ export interface WorkerOptions {
|
||||
workspace: string;
|
||||
pipelineTesting?: boolean;
|
||||
keepContainer?: boolean;
|
||||
authOnly?: boolean;
|
||||
validateModel?: boolean;
|
||||
piAuthHostPath?: string;
|
||||
}
|
||||
|
||||
@@ -512,13 +513,17 @@ export function spawnWorker(opts: WorkerOptions): ChildProcess {
|
||||
if (opts.pipelineTesting) {
|
||||
args.push('--pipeline-testing');
|
||||
}
|
||||
if (opts.authOnly) {
|
||||
args.push('--validate-auth');
|
||||
}
|
||||
if (opts.validateModel) {
|
||||
args.push('--validate-model');
|
||||
}
|
||||
|
||||
// Inherit stderr so `docker run` daemon errors surface to the user;
|
||||
// ignore stdin/stdout (the container ID is noise).
|
||||
return spawn('docker', args, {
|
||||
stdio: ['ignore', 'ignore', 'inherit'],
|
||||
// Prevent MSYS/Git Bash from converting Unix paths on Windows
|
||||
...(os.platform() === 'win32' && { env: { ...process.env, MSYS_NO_PATHCONV: '1' } }),
|
||||
});
|
||||
}
|
||||
|
||||
|
||||
@@ -33,6 +33,8 @@ export const START_OPTIONS: readonly (readonly [string, string])[] = [
|
||||
['-o, --output <path>', 'Copy deliverables to this directory after the run'],
|
||||
['-w, --workspace <name>', 'Named workspace (auto-resumes if it exists)'],
|
||||
['-f, --follow', 'Stream the scan log until it finishes'],
|
||||
['--validate-auth', 'Validate authentication only, then stop (no pentest)'],
|
||||
['--validate-model', 'Validate the AI model only, then stop (no pentest)'],
|
||||
['--pipeline-testing', 'Use minimal prompts for fast testing'],
|
||||
['--keep-container', 'Preserve the worker container after exit for log inspection'],
|
||||
];
|
||||
@@ -45,6 +47,8 @@ const COMMAND_HELP: Readonly<Record<string, CommandHelp>> = {
|
||||
'start -u https://example.com -r ./my-repo',
|
||||
'start -u https://example.com -r /path/to/repo -c config.yaml -w q1-audit',
|
||||
'start -u https://example.com -r ./my-repo --follow',
|
||||
'start -u https://example.com -r ./my-repo -c config.yaml --validate-auth',
|
||||
'start -u https://example.com -r ./my-repo --validate-model',
|
||||
],
|
||||
},
|
||||
stop: {
|
||||
|
||||
@@ -61,6 +61,18 @@ function blockSudo(): void {
|
||||
);
|
||||
}
|
||||
|
||||
/** Refuse to run on native Windows. WSL2 reports `linux`, so it is unaffected. */
|
||||
function blockNativeWindows(): void {
|
||||
if (process.platform !== 'win32') return;
|
||||
|
||||
failWith(
|
||||
'CLI_PRECONDITION_FAILED',
|
||||
'Shannon does not run on native Windows.',
|
||||
'Run Shannon inside WSL2. Setup instructions:',
|
||||
'https://github.com/KeygraphHQ/shannon/blob/main/docs/platforms.md',
|
||||
);
|
||||
}
|
||||
|
||||
/** Commands whose `--json` output contract extends to failures. */
|
||||
const JSON_CAPABLE_COMMANDS = new Set(['status', 'scans', 'version', '--version', '-v']);
|
||||
|
||||
@@ -177,6 +189,8 @@ interface ParsedStartArgs {
|
||||
pipelineTesting: boolean;
|
||||
keepContainer: boolean;
|
||||
follow: boolean;
|
||||
authOnly: boolean;
|
||||
validateModel: boolean;
|
||||
}
|
||||
|
||||
function parseStartArgs(argv: string[]): ParsedStartArgs {
|
||||
@@ -193,6 +207,8 @@ function parseStartArgs(argv: string[]): ParsedStartArgs {
|
||||
pipelineTesting: ['--pipeline-testing'],
|
||||
keepContainer: ['--keep-container'],
|
||||
follow: ['-f', '--follow'],
|
||||
authOnly: ['--validate-auth'],
|
||||
validateModel: ['--validate-model'],
|
||||
},
|
||||
});
|
||||
|
||||
@@ -208,12 +224,25 @@ function parseStartArgs(argv: string[]): ParsedStartArgs {
|
||||
failUsage(`invalid --url: ${url}`);
|
||||
}
|
||||
|
||||
if (flags.authOnly && flags.validateModel) {
|
||||
failUsage('--validate-auth and --validate-model cannot be combined; run one validation at a time');
|
||||
}
|
||||
|
||||
if (flags.authOnly && !values.config) {
|
||||
failUsage(
|
||||
'--validate-auth needs a config file with an authentication block',
|
||||
`Usage: ${commandPrefix()} start -u <url> -r <path> -c <config.yaml> --validate-auth`,
|
||||
);
|
||||
}
|
||||
|
||||
return {
|
||||
url,
|
||||
repo,
|
||||
pipelineTesting: !!flags.pipelineTesting,
|
||||
keepContainer: !!flags.keepContainer,
|
||||
follow: !!flags.follow,
|
||||
authOnly: !!flags.authOnly,
|
||||
validateModel: !!flags.validateModel,
|
||||
...(values.config && { config: values.config }),
|
||||
...(values.modelsConfig && { modelsConfig: values.modelsConfig }),
|
||||
...(values.workspace && { workspace: values.workspace }),
|
||||
@@ -262,6 +291,7 @@ async function main(): Promise<void> {
|
||||
enableJsonErrors();
|
||||
}
|
||||
|
||||
blockNativeWindows();
|
||||
blockSudo();
|
||||
|
||||
const args = process.argv.slice(2);
|
||||
|
||||
@@ -363,6 +363,33 @@ function agenticSastPhase(operations: readonly DerivedAgent[]): DerivedPhase | u
|
||||
};
|
||||
}
|
||||
|
||||
/** Preflight rows shown at the top of the tree, in run order. Each is its own single-line phase. */
|
||||
const PREFLIGHT_ROW_KEYS = ['preflight', 'cyber-access'] as const;
|
||||
|
||||
/**
|
||||
* The two preflight gates the worker persists — the preflight checks and the cyber-access verification —
|
||||
* as top-of-tree rows. Each appears once its stage is recorded (running, then done or failed); a
|
||||
* run that never reaches a gate simply omits its row.
|
||||
*/
|
||||
function preflightPhases(operations: readonly DerivedAgent[]): DerivedPhase[] {
|
||||
const byKey = new Map(operations.map((operation) => [operation.name, operation]));
|
||||
const phases: DerivedPhase[] = [];
|
||||
for (const key of PREFLIGHT_ROW_KEYS) {
|
||||
const operation = byKey.get(key);
|
||||
if (operation === undefined) continue;
|
||||
phases.push({
|
||||
key: operation.name,
|
||||
label: operation.label,
|
||||
children: false,
|
||||
meta: 'duration',
|
||||
state: operation.state,
|
||||
summary: operation,
|
||||
agents: [operation],
|
||||
});
|
||||
}
|
||||
return phases;
|
||||
}
|
||||
|
||||
/**
|
||||
* Bookkeeping rows worth showing. A deterministic stage that has completed says nothing —
|
||||
* it can only ever read 0s — but one that is still running, or that failed, is exactly what
|
||||
@@ -408,13 +435,14 @@ function assemblePhases(agentPhases: readonly DerivedPhase[], operations: readon
|
||||
return phase;
|
||||
});
|
||||
|
||||
const preflight = preflightPhases(operations);
|
||||
const sast = agenticSastPhase(operations);
|
||||
if (sast === undefined) return phases;
|
||||
if (sast === undefined) return [...preflight, ...phases];
|
||||
|
||||
// Agentic SAST starts with the scan and runs alongside the pentest, so it reads after
|
||||
// the login check rather than appended past Reporting where it never ran.
|
||||
const afterAuth = phases.findIndex((phase) => phase.key === 'auth-validation') + 1;
|
||||
return [...phases.slice(0, afterAuth), sast, ...phases.slice(afterAuth)];
|
||||
return [...preflight, ...phases.slice(0, afterAuth), sast, ...phases.slice(afterAuth)];
|
||||
}
|
||||
|
||||
export { agentError };
|
||||
@@ -108,6 +108,8 @@ const MISCELLANEOUS_EXPLOIT_AGENT: AgentSpec = {
|
||||
* available guess.
|
||||
*/
|
||||
export function pipelineForState(state: PipelineState | null): readonly PhaseSpec[] {
|
||||
if (state?.validateModel === true) return [];
|
||||
if (state?.authOnly === true) return PIPELINE.filter((phase) => phase.key === 'auth-validation');
|
||||
if (state?.expectedAgents === undefined) return PIPELINE;
|
||||
const expected = new Set(state.expectedAgents);
|
||||
return PIPELINE.map((phase) => {
|
||||
@@ -136,7 +138,8 @@ const AGENTIC_SAST_PARENT_KEY = 'agentic-sast';
|
||||
// apps/worker/src/temporal/reconcile-activity-types.ts, and
|
||||
// apps/worker/src/ai/sast/capella/temporal/activity-types.ts.
|
||||
const OPERATION_ACTIVITY_PROGRESS: Readonly<Record<string, ActivityProgressSpec>> = {
|
||||
runPreflightValidation: { key: 'preflight', label: 'Preflight validation', kind: 'operation' },
|
||||
runPreflightValidation: { key: 'preflight', label: 'Preflight', kind: 'operation' },
|
||||
runCyberAccessVerification: { key: 'cyber-access', label: 'Cyber access verification', kind: 'operation' },
|
||||
syncPlaywrightStealthConfig: { key: 'preflight', label: 'Browser setup', kind: 'operation' },
|
||||
initDeliverableGit: { key: 'scan-initialization', label: 'Initialize deliverables', kind: 'operation' },
|
||||
syncCodePathDenyRules: { key: 'scan-initialization', label: 'Apply source rules', kind: 'operation' },
|
||||
@@ -352,6 +355,8 @@ export type PipelineStatus = 'running' | 'completed' | 'failed' | 'cancelled' |
|
||||
|
||||
export interface PipelineState {
|
||||
readonly status: PipelineStatus;
|
||||
readonly authOnly?: boolean;
|
||||
readonly validateModel?: boolean;
|
||||
readonly currentPhase: string | null;
|
||||
readonly currentAgent: string | null;
|
||||
readonly completedAgents: string[];
|
||||
|
||||
@@ -75,6 +75,8 @@ function isProviderFailureCategory(value: unknown): value is string {
|
||||
}
|
||||
|
||||
const OPERATION_LABELS = new Set([
|
||||
'Preflight',
|
||||
'Cyber access verification',
|
||||
'Agentic SAST',
|
||||
// Capella stage rows, signalled up from the SAST child workflow. Mirrors
|
||||
// CAPELLA_STAGE_LABELS in apps/worker/src/ai/sast/types.ts, minus the deterministic
|
||||
@@ -228,7 +230,7 @@ export function safeOperationLabel(value: string): string {
|
||||
|
||||
export function safeOperationKey(value: string): string {
|
||||
if (
|
||||
/^(?:agentic-sast|miscellaneous-pipeline|report:(?:initialize|assemble|compact|checkpoint|finalize|finalize-degraded|terminal|surface))$/u.test(
|
||||
/^(?:preflight|cyber-access|agentic-sast|miscellaneous-pipeline|report:(?:initialize|assemble|compact|checkpoint|finalize|finalize-degraded|terminal|surface))$/u.test(
|
||||
value,
|
||||
) ||
|
||||
/^agentic-sast:(?:architecture|threat-model|plan|research|dedupe|review|critic|confirm|calibrate)$/u.test(value) ||
|
||||
|
||||
@@ -257,6 +257,25 @@ export async function describeScan(workflowId: string): Promise<ScanDescription
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* Activity-type names pending on a running scan; empty on any failure. Tolerant (it feeds the
|
||||
* start spinner) unlike describeScan, which fails closed so the status tree is never incomplete.
|
||||
*/
|
||||
export async function runningActivityTypes(workflowId: string): Promise<readonly string[]> {
|
||||
try {
|
||||
const client = await getClient();
|
||||
const desc = await client.workflow.getHandle(workflowId).describe();
|
||||
const names: string[] = [];
|
||||
for (const pending of desc.raw.pendingActivities ?? []) {
|
||||
const name = pending.activityType?.name;
|
||||
if (name) names.push(name);
|
||||
}
|
||||
return names;
|
||||
} catch {
|
||||
return [];
|
||||
}
|
||||
}
|
||||
|
||||
/** Live progress of a running scan via the getProgress query. Null if the query can't be served (no worker). */
|
||||
export async function queryProgress(workflowId: string): Promise<PipelineState | null> {
|
||||
const client = await getClient();
|
||||
|
||||
@@ -20,7 +20,9 @@
|
||||
* Resolution returns a pi `Model` plus the `ModelRuntime` that owns its auth,
|
||||
* built over an in-memory credential store primed from the environment.
|
||||
*
|
||||
* A model too new for the pinned pi release is reachable by passing its descriptor in a
|
||||
* The catalogue is refreshed over the network at scan start, so a newly released model
|
||||
* on a catalogue provider resolves on its own. A model the catalogue does not carry, such
|
||||
* as a router model under its own id, or a self-hosted server, is described in a
|
||||
* pi `models.json` (the CLI's `--models-config`), which merges over the catalogue. The
|
||||
* credential store below outranks any `apiKey` that file carries, so it describes the
|
||||
* model while the environment still supplies the secret.
|
||||
@@ -216,13 +218,18 @@ function modelsStorePath(): string {
|
||||
}
|
||||
|
||||
/**
|
||||
* Build a ModelRuntime whose only credential is the one supplied. Model catalogs
|
||||
* stay offline (`allowModelNetwork` defaults to false) so a scan never blocks on
|
||||
* a catalog refresh.
|
||||
* Build a ModelRuntime whose only credential is the one supplied. `allowModelNetwork`
|
||||
* refreshes the model catalogue over the network at scan start, so the registry reflects
|
||||
* models the pinned pi build predates. The fetch is bounded and falls back to the static
|
||||
* catalogue on timeout, so an unreachable endpoint cannot hang the scan. A mounted
|
||||
* `--models-config` overlays the catalogue and is reloaded on every refresh, so its
|
||||
* definitions take precedence.
|
||||
*
|
||||
* `modelsPath` is always explicit, never pi's default of `<agent dir>/models.json`: with no
|
||||
* `--models-config` it is null, which switches models.json off outright, so a stray file in
|
||||
* that shared dir cannot feed model definitions to a run that did not ask for them.
|
||||
* `modelsStorePath` is pinned to the writable agent dir, replacing pi's default
|
||||
* `dirname(modelsPath)` (a read-only mount) as the fetched catalogue's store.
|
||||
*
|
||||
* When the host's pi auth.json is present, the runtime reads it instead: pi's
|
||||
* disk-backed store resolves the credential. The mount is writable so OAuth
|
||||
@@ -233,6 +240,8 @@ export async function createModelRuntime(providerId: string, apiKey: string | un
|
||||
const modelSources = {
|
||||
modelsPath: modelsPath ?? null,
|
||||
...(modelsPath ? { modelsStorePath: modelsStorePath() } : {}),
|
||||
allowModelNetwork: true,
|
||||
modelRefreshTimeoutMs: 10_000,
|
||||
};
|
||||
|
||||
if (piAuthPresent()) {
|
||||
@@ -254,9 +263,8 @@ export interface ModelSelection {
|
||||
*
|
||||
* The model must exist in the runtime's registry, whether or not an endpoint override
|
||||
* is in play — a base URL changes the address and nothing else. A gateway serving a
|
||||
* model under its own name, or one newer than the pinned pi release, is described in a
|
||||
* `--models-config` file, which puts a real descriptor in the registry rather than
|
||||
* guessing one from an unrelated model.
|
||||
* model under its own name is described in a `--models-config` file, which puts a real
|
||||
* descriptor in the registry rather than guessing one from an unrelated model.
|
||||
*/
|
||||
export function resolveModel(
|
||||
modelRuntime: ModelRuntime,
|
||||
|
||||
@@ -32,6 +32,7 @@ import type {
|
||||
CapellaTool,
|
||||
} from './capella-agent-types.js';
|
||||
import { PI_RETRY_SETTINGS } from './retry-settings.js';
|
||||
import { PI_THINKING_LEVEL } from './thinking-level.js';
|
||||
|
||||
const MAX_ERROR_LENGTH = 2_000;
|
||||
const MAX_TOOLS_PER_SESSION = 32;
|
||||
@@ -393,6 +394,7 @@ class StandaloneCapellaAgentExecutor implements CapellaAgentExecutor {
|
||||
cwd: request.cwd,
|
||||
agentDir,
|
||||
model: selection.model,
|
||||
thinkingLevel: PI_THINKING_LEVEL,
|
||||
modelRuntime: selection.modelRuntime,
|
||||
noTools: 'all',
|
||||
tools: toolNames,
|
||||
|
||||
@@ -48,6 +48,7 @@ import { permissionSystemConfigExists, permissionSystemPackageDir } from './perm
|
||||
import { PI_RETRY_SETTINGS } from './retry-settings.js';
|
||||
import { createGlobTool, createTodoWriteTool } from './session-tools.js';
|
||||
import { createTaskTool } from './task-tool.js';
|
||||
import { PI_THINKING_LEVEL } from './thinking-level.js';
|
||||
import { TraceEmitter } from './trace-emitter.js';
|
||||
import { providerTurnError, type SafeProviderTurnDetails, safeProviderTurnDetails } from './turn-error.js';
|
||||
|
||||
@@ -332,6 +333,7 @@ export async function runPiPrompt(
|
||||
({ session } = await createAgentSession({
|
||||
cwd: sourceDir,
|
||||
model: selection.model,
|
||||
thinkingLevel: PI_THINKING_LEVEL,
|
||||
tools,
|
||||
customTools,
|
||||
modelRuntime: selection.modelRuntime,
|
||||
|
||||
@@ -1,295 +0,0 @@
|
||||
// Copyright (C) 2026 Keygraph, Inc.
|
||||
//
|
||||
// This program is free software: you can redistribute it and/or modify
|
||||
// it under the terms of the GNU Affero General Public License version 3
|
||||
// as published by the Free Software Foundation.
|
||||
|
||||
/** Attempt-local working-tree copy used by the task-formation model boundary. */
|
||||
|
||||
import type { Dirent, Stats } from 'node:fs';
|
||||
import { cp, lstat, mkdir, mkdtemp, readdir, realpath, rm } from 'node:fs/promises';
|
||||
import os from 'node:os';
|
||||
import path from 'node:path';
|
||||
import { ArtifactIntegrityError, ReconciliationIoError } from '../reconciliation/artifact-store.js';
|
||||
|
||||
const JAIL_PREFIX = 'shannon-task-formation-';
|
||||
// Never copied into the model-readable jail: `.git` carries deliverables history, `.shannon` holds
|
||||
// scan internals, and `.pi` holds provider credentials. Any of these reaching the jail would expose
|
||||
// them to the tools the model drives. The post-copy verification re-checks their absence by name.
|
||||
const ALWAYS_EXCLUDED_NAMES = Object.freeze(['.git', '.shannon', '.pi'] as const);
|
||||
|
||||
export interface SourceJailOptions {
|
||||
readonly sourceRoot: string;
|
||||
readonly deliverablesPath: string;
|
||||
readonly reconciliationWorkspacePath: string;
|
||||
readonly signal?: AbortSignal;
|
||||
/** Test-only filesystem selector. Production uses `os.tmpdir()`. */
|
||||
readonly tempRoot?: string;
|
||||
}
|
||||
|
||||
/** One source-only jail plus the immutable deny rules used by its live tool gate. */
|
||||
export interface SourceJail {
|
||||
readonly dir: string;
|
||||
readonly deniedPaths: readonly string[];
|
||||
cleanup(): Promise<void>;
|
||||
}
|
||||
|
||||
function isErrno(error: unknown, code: string): boolean {
|
||||
return error instanceof Error && (error as NodeJS.ErrnoException).code === code;
|
||||
}
|
||||
|
||||
function cancellationError(signal: AbortSignal): Error {
|
||||
if (signal.reason instanceof Error) return signal.reason;
|
||||
return new DOMException('Task formation was cancelled.', 'AbortError');
|
||||
}
|
||||
|
||||
function checkCancellation(signal: AbortSignal | undefined): void {
|
||||
if (signal?.aborted === true) throw cancellationError(signal);
|
||||
}
|
||||
|
||||
// Path-confinement predicate: true only when `candidate` is `root` itself or lies beneath it.
|
||||
// A relative path that escapes upward (`..`) or is absolute means the candidate is outside the root.
|
||||
function isWithin(root: string, candidate: string): boolean {
|
||||
const relativePath = path.relative(root, candidate);
|
||||
return (
|
||||
relativePath === '' ||
|
||||
(!relativePath.startsWith(`..${path.sep}`) && relativePath !== '..' && !path.isAbsolute(relativePath))
|
||||
);
|
||||
}
|
||||
|
||||
async function relativeExclusion(
|
||||
sourceRoot: string,
|
||||
lexicalSourceRoot: string,
|
||||
candidate: string,
|
||||
): Promise<string | undefined> {
|
||||
const resolved = path.resolve(candidate);
|
||||
let relativePath: string | undefined;
|
||||
if (isWithin(sourceRoot, resolved)) {
|
||||
relativePath = path.relative(sourceRoot, resolved);
|
||||
} else if (isWithin(lexicalSourceRoot, resolved)) {
|
||||
relativePath = path.relative(lexicalSourceRoot, resolved);
|
||||
} else {
|
||||
try {
|
||||
const canonicalCandidate = await realpath(resolved);
|
||||
if (isWithin(sourceRoot, canonicalCandidate)) {
|
||||
relativePath = path.relative(sourceRoot, canonicalCandidate);
|
||||
}
|
||||
} catch {
|
||||
return undefined;
|
||||
}
|
||||
}
|
||||
if (relativePath === undefined) return undefined;
|
||||
|
||||
if (relativePath === '') {
|
||||
// An exclusion that resolves to the whole root would empty the jail. Fail closed rather than
|
||||
// copy nothing and hand the model an empty tree.
|
||||
throw new ArtifactIntegrityError('A task-formation exclusion resolves to the complete source root');
|
||||
}
|
||||
return relativePath;
|
||||
}
|
||||
|
||||
async function buildDynamicExclusions(
|
||||
options: SourceJailOptions,
|
||||
sourceRoot: string,
|
||||
lexicalSourceRoot: string,
|
||||
): Promise<readonly string[]> {
|
||||
const exclusions = (
|
||||
await Promise.all([
|
||||
relativeExclusion(sourceRoot, lexicalSourceRoot, options.deliverablesPath),
|
||||
relativeExclusion(sourceRoot, lexicalSourceRoot, options.reconciliationWorkspacePath),
|
||||
])
|
||||
).filter((value): value is string => value !== undefined);
|
||||
return Object.freeze([...new Set(exclusions)]);
|
||||
}
|
||||
|
||||
function pathHasAlwaysExcludedName(relativePath: string): boolean {
|
||||
const segments = relativePath.split(path.sep);
|
||||
return segments.some((segment) => (ALWAYS_EXCLUDED_NAMES as readonly string[]).includes(segment));
|
||||
}
|
||||
|
||||
function pathIsDynamicallyExcluded(relativePath: string, exclusions: readonly string[]): boolean {
|
||||
return exclusions.some((excluded) => relativePath === excluded || relativePath.startsWith(`${excluded}${path.sep}`));
|
||||
}
|
||||
|
||||
async function copySourceTree(
|
||||
sourceRoot: string,
|
||||
destination: string,
|
||||
dynamicExclusions: readonly string[],
|
||||
signal: AbortSignal | undefined,
|
||||
): Promise<void> {
|
||||
let entries: Dirent[];
|
||||
try {
|
||||
entries = (await readdir(sourceRoot, { withFileTypes: true })).sort((left, right) =>
|
||||
left.name.localeCompare(right.name),
|
||||
);
|
||||
} catch {
|
||||
throw new ReconciliationIoError('Unable to enumerate the task-formation source tree');
|
||||
}
|
||||
|
||||
// Cancellation is checked before every top-level entry and inside the copy filter so an aborted
|
||||
// scan stops promptly instead of copying a whole large tree first.
|
||||
for (const entry of entries) {
|
||||
checkCancellation(signal);
|
||||
const source = path.join(sourceRoot, entry.name);
|
||||
const destinationEntry = path.join(destination, entry.name);
|
||||
try {
|
||||
// verbatimSymlinks copies links as links rather than following them, so a link pointing
|
||||
// outside the tree cannot pull external content in; the filter then drops any path that
|
||||
// resolves outside the root, plus the always- and dynamically-excluded paths.
|
||||
await cp(source, destinationEntry, {
|
||||
recursive: true,
|
||||
verbatimSymlinks: true,
|
||||
errorOnExist: true,
|
||||
force: false,
|
||||
async filter(candidate) {
|
||||
checkCancellation(signal);
|
||||
const relativePath = path.relative(sourceRoot, candidate);
|
||||
if (relativePath === '' || !isWithin(sourceRoot, path.resolve(candidate))) return false;
|
||||
if (pathHasAlwaysExcludedName(relativePath)) return false;
|
||||
return !pathIsDynamicallyExcluded(relativePath, dynamicExclusions);
|
||||
},
|
||||
});
|
||||
} catch (error) {
|
||||
if (signal?.aborted === true) throw cancellationError(signal);
|
||||
if (error instanceof ArtifactIntegrityError) throw error;
|
||||
throw new ReconciliationIoError('Unable to copy the task-formation source tree');
|
||||
}
|
||||
}
|
||||
checkCancellation(signal);
|
||||
}
|
||||
|
||||
async function assertAlwaysExcludedNamesAbsent(directory: string, signal: AbortSignal | undefined): Promise<void> {
|
||||
checkCancellation(signal);
|
||||
let entries: Dirent[];
|
||||
try {
|
||||
entries = await readdir(directory, { withFileTypes: true });
|
||||
} catch {
|
||||
throw new ReconciliationIoError('Unable to verify the task-formation source jail');
|
||||
}
|
||||
|
||||
for (const entry of entries) {
|
||||
checkCancellation(signal);
|
||||
if ((ALWAYS_EXCLUDED_NAMES as readonly string[]).includes(entry.name)) {
|
||||
throw new ArtifactIntegrityError('The task-formation source jail contains an excluded entry');
|
||||
}
|
||||
if (entry.isDirectory() && !entry.isSymbolicLink()) {
|
||||
await assertAlwaysExcludedNamesAbsent(path.join(directory, entry.name), signal);
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
async function assertDynamicExclusionsAbsent(
|
||||
directory: string,
|
||||
exclusions: readonly string[],
|
||||
signal: AbortSignal | undefined,
|
||||
): Promise<void> {
|
||||
for (const excluded of exclusions) {
|
||||
checkCancellation(signal);
|
||||
try {
|
||||
await lstat(path.join(directory, excluded));
|
||||
} catch (error) {
|
||||
if (isErrno(error, 'ENOENT')) continue;
|
||||
throw new ReconciliationIoError('Unable to verify a task-formation jail exclusion');
|
||||
}
|
||||
throw new ArtifactIntegrityError('The task-formation source jail contains a protected workspace entry');
|
||||
}
|
||||
}
|
||||
|
||||
// Re-verify the copied tree independently of the copy filter: the jail root must be a real
|
||||
// directory (not a symlink), and no excluded name or protected workspace path may survive. This
|
||||
// catches a filter gap or a race during the copy before the model is allowed to read the tree.
|
||||
async function verifyJail(
|
||||
directory: string,
|
||||
dynamicExclusions: readonly string[],
|
||||
signal: AbortSignal | undefined,
|
||||
): Promise<void> {
|
||||
checkCancellation(signal);
|
||||
let stats: Stats;
|
||||
try {
|
||||
stats = await lstat(directory);
|
||||
} catch {
|
||||
throw new ReconciliationIoError('Unable to inspect the task-formation source jail');
|
||||
}
|
||||
if (stats.isSymbolicLink() || !stats.isDirectory()) {
|
||||
throw new ArtifactIntegrityError('The task-formation source jail is not a real directory');
|
||||
}
|
||||
await assertAlwaysExcludedNamesAbsent(directory, signal);
|
||||
await assertDynamicExclusionsAbsent(directory, dynamicExclusions, signal);
|
||||
checkCancellation(signal);
|
||||
}
|
||||
|
||||
async function removeJail(directory: string): Promise<void> {
|
||||
try {
|
||||
await rm(directory, { recursive: true, force: true });
|
||||
} catch {
|
||||
throw new ReconciliationIoError('Unable to remove the task-formation source jail');
|
||||
}
|
||||
|
||||
try {
|
||||
await lstat(directory);
|
||||
} catch (error) {
|
||||
if (isErrno(error, 'ENOENT')) return;
|
||||
throw new ReconciliationIoError('Unable to verify task-formation source-jail cleanup');
|
||||
}
|
||||
throw new ReconciliationIoError('Task-formation source-jail cleanup left the jail on disk');
|
||||
}
|
||||
|
||||
/**
|
||||
* Copy the scanned working tree into an isolated temporary directory without following symlinks.
|
||||
* Every failure removes the attempt-local directory before it propagates.
|
||||
*/
|
||||
export async function materializeSourceJail(options: SourceJailOptions): Promise<SourceJail> {
|
||||
checkCancellation(options.signal);
|
||||
|
||||
const lexicalSourceRoot = path.resolve(options.sourceRoot);
|
||||
let sourceRoot: string;
|
||||
try {
|
||||
sourceRoot = await realpath(options.sourceRoot);
|
||||
const sourceStats = await lstat(sourceRoot);
|
||||
if (sourceStats.isSymbolicLink() || !sourceStats.isDirectory()) {
|
||||
throw new ArtifactIntegrityError('The task-formation source root is not a real directory');
|
||||
}
|
||||
} catch (error) {
|
||||
if (error instanceof ArtifactIntegrityError) throw error;
|
||||
throw new ReconciliationIoError('Unable to resolve the task-formation source root');
|
||||
}
|
||||
|
||||
let tempRoot: string;
|
||||
try {
|
||||
const configuredTempRoot = options.tempRoot ?? os.tmpdir();
|
||||
await mkdir(configuredTempRoot, { recursive: true });
|
||||
tempRoot = await realpath(configuredTempRoot);
|
||||
} catch {
|
||||
throw new ReconciliationIoError('Unable to resolve the task-formation temporary root');
|
||||
}
|
||||
// A temp root inside the source tree would make the copy try to copy the jail into itself.
|
||||
if (isWithin(sourceRoot, tempRoot)) {
|
||||
throw new ArtifactIntegrityError('The task-formation temporary root cannot be inside the source tree');
|
||||
}
|
||||
|
||||
const dynamicExclusions = await buildDynamicExclusions(options, sourceRoot, lexicalSourceRoot);
|
||||
let directory: string;
|
||||
try {
|
||||
directory = await mkdtemp(path.join(tempRoot, JAIL_PREFIX));
|
||||
} catch {
|
||||
throw new ReconciliationIoError('Unable to create the task-formation source jail');
|
||||
}
|
||||
|
||||
let cleaned = false;
|
||||
const cleanup = async (): Promise<void> => {
|
||||
if (cleaned) return;
|
||||
await removeJail(directory);
|
||||
cleaned = true;
|
||||
};
|
||||
|
||||
try {
|
||||
await copySourceTree(sourceRoot, directory, dynamicExclusions, options.signal);
|
||||
await verifyJail(directory, dynamicExclusions, options.signal);
|
||||
} catch (error) {
|
||||
await cleanup().catch(() => undefined);
|
||||
throw error;
|
||||
}
|
||||
|
||||
const deniedPaths = Object.freeze([...ALWAYS_EXCLUDED_NAMES, ...dynamicExclusions]);
|
||||
return Object.freeze({ dir: directory, deniedPaths, cleanup });
|
||||
}
|
||||
@@ -29,6 +29,7 @@ import type { ValidatingSubmitTool } from '../reconciliation/submit-validation.j
|
||||
import { ConfinementError, compileRepositoryGlob, RepositoryConfinement } from '../sast/capella/tools/confinement.js';
|
||||
import { createCapellaRepositoryTools } from '../sast/capella/tools/repository-tools.js';
|
||||
import { PI_RETRY_SETTINGS } from './retry-settings.js';
|
||||
import { PI_THINKING_LEVEL } from './thinking-level.js';
|
||||
|
||||
const DEFAULT_TIMEOUT_MS = 30 * 60 * 1_000;
|
||||
const DEFAULT_MAX_TURNS = 64;
|
||||
@@ -37,10 +38,9 @@ const MAX_TURNS = 128;
|
||||
const MAX_LIST_RESULTS = 500;
|
||||
const DEFAULT_LIST_RESULTS = 200;
|
||||
const MAX_OUTPUT_BYTES = 64 * 1024;
|
||||
// The live-tool-side counterpart of the source jail's copy-time exclusion (source-jail.ts): even if
|
||||
// one of these somehow existed in the jailed tree, the read/grep/find/ls/glob tools built below must
|
||||
// still refuse to serve it. `.git` is deliverables history, `.shannon` is scan internals, `.pi` is
|
||||
// provider credentials.
|
||||
// Task formation reads the live repository, so these read-only tools are the sole barrier keeping the
|
||||
// model out of `.git` (source history), `.shannon` (scan internals, incl. the deliverables Git repo),
|
||||
// and `.pi` (credentials). Always denied, whatever extra denies a caller passes.
|
||||
const ALWAYS_DENIED_PATHS = Object.freeze(['.git', '.shannon', '.pi'] as const);
|
||||
const TRANSIENT_IO_CODES = new Set([
|
||||
'EAGAIN',
|
||||
@@ -287,9 +287,9 @@ function createGlobTool(confinement: RepositoryConfinement): ToolDefinition {
|
||||
return defineTool({
|
||||
name: 'glob',
|
||||
label: 'Glob source files',
|
||||
description: 'Match bounded file globs from the source-jail root without following symlinks.',
|
||||
promptSnippet: 'glob: match source files from the jail root',
|
||||
promptGuidelines: ['Patterns are always rooted in the source jail.'],
|
||||
description: 'Match bounded file globs from the repository root without following symlinks.',
|
||||
promptSnippet: 'glob: match source files from the repository root',
|
||||
promptGuidelines: ['Patterns are always rooted in the repository.'],
|
||||
parameters: Type.Object(
|
||||
{
|
||||
pattern: Type.String({ minLength: 1, maxLength: 256 }),
|
||||
@@ -321,7 +321,7 @@ function createGlobTool(confinement: RepositoryConfinement): ToolDefinition {
|
||||
});
|
||||
}
|
||||
|
||||
/** Create the five code-owned source tools that share one canonical jail policy. */
|
||||
/** Create the five code-owned source tools that share one canonical deny policy. */
|
||||
export async function createTaskFormationSourceTools(options: ToolFactoryOptions): Promise<readonly ToolDefinition[]> {
|
||||
const deniedPaths = uniqueDeniedPaths(options.deniedPaths);
|
||||
const capellaTools = await createCapellaRepositoryTools({
|
||||
@@ -527,6 +527,7 @@ class StandaloneTaskFormationExecutor implements TaskFormationExecutor {
|
||||
cwd: request.cwd,
|
||||
agentDir,
|
||||
model: selection.model,
|
||||
thinkingLevel: PI_THINKING_LEVEL,
|
||||
modelRuntime: selection.modelRuntime,
|
||||
noTools: 'all',
|
||||
tools: toolNames,
|
||||
|
||||
@@ -19,6 +19,7 @@ import {
|
||||
} from '@earendil-works/pi-coding-agent';
|
||||
import { type LoggableAgentName, normalizeSemanticLabel } from '../../audit/safe-fields.js';
|
||||
import { PI_RETRY_SETTINGS } from './retry-settings.js';
|
||||
import { PI_THINKING_LEVEL } from './thinking-level.js';
|
||||
import { TraceEmitter } from './trace-emitter.js';
|
||||
|
||||
export interface TaskToolContext {
|
||||
@@ -135,6 +136,7 @@ export function createTaskTool(config: TaskToolContext): ToolDefinition {
|
||||
agentDir,
|
||||
resourceLoader,
|
||||
model: config.model,
|
||||
thinkingLevel: PI_THINKING_LEVEL,
|
||||
tools: CHILD_TOOLS,
|
||||
modelRuntime: config.modelRuntime,
|
||||
sessionManager: SessionManager.inMemory(config.cwd),
|
||||
|
||||
@@ -0,0 +1,10 @@
|
||||
// Copyright (C) 2026 Keygraph, Inc.
|
||||
//
|
||||
// This program is free software: you can redistribute it and/or modify
|
||||
// it under the terms of the GNU Affero General Public License version 3
|
||||
// as published by the Free Software Foundation.
|
||||
|
||||
import type { ThinkingLevel } from '@earendil-works/pi-agent-core';
|
||||
|
||||
/** Thinking level for every pi agent session, raised above pi's default for deeper analysis. */
|
||||
export const PI_THINKING_LEVEL: ThinkingLevel = 'high';
|
||||
@@ -6,12 +6,10 @@
|
||||
|
||||
/** Pass 1 task formation over current observations only. */
|
||||
|
||||
import path from 'node:path';
|
||||
import { DEFAULT_DELIVERABLES_SUBDIR, WORKSPACES_DIR } from '../../paths.js';
|
||||
import { WORKSPACES_DIR } from '../../paths.js';
|
||||
import { loadPrompt } from '../../services/prompt-manager.js';
|
||||
import type { ActivityLogger } from '../../types/activity-logger.js';
|
||||
import type { ReconciliationClass } from '../../types/reconciliation.js';
|
||||
import { materializeSourceJail } from '../pi/source-jail.js';
|
||||
import {
|
||||
isTaskFormationFallbackReason,
|
||||
type TaskFormationExecutionContext,
|
||||
@@ -61,7 +59,6 @@ export interface FormClassExploitTasksInput {
|
||||
readonly repositoryPath: string;
|
||||
readonly producerRef: ArtifactRef<'producer-observations'>;
|
||||
readonly supplementalRef: ArtifactRef<'supplemental-observations'>;
|
||||
readonly deliverablesSubdir?: string;
|
||||
readonly webUrl?: string;
|
||||
}
|
||||
|
||||
@@ -181,14 +178,14 @@ function taskFormationPromptName(vulnerabilityClass: ReconciliationClass): strin
|
||||
// prompt itself is missing or unreadable content, which Temporal should not spend retries on.
|
||||
async function loadClassPolicy(
|
||||
vulnerabilityClass: ReconciliationClass,
|
||||
jailPath: string,
|
||||
repoPath: string,
|
||||
webUrl: string,
|
||||
logger: ActivityLogger,
|
||||
): Promise<string> {
|
||||
try {
|
||||
return await loadPrompt(
|
||||
taskFormationPromptName(vulnerabilityClass),
|
||||
{ webUrl, repoPath: jailPath, AUTH_STATE_FILE: '' },
|
||||
{ webUrl, repoPath, AUTH_STATE_FILE: '' },
|
||||
null,
|
||||
false,
|
||||
logger,
|
||||
@@ -332,109 +329,69 @@ export function createFormClassExploitTasks(
|
||||
const submitTool = createValidatingSubmitTool(buildTaskFormationSchema([...labelSet]), (parameters) =>
|
||||
findTaskFormationProblems(parameters, labelSet),
|
||||
);
|
||||
const deliverablesPath = path.resolve(
|
||||
const classPolicy = await loadClassPolicy(
|
||||
input.vulnerabilityClass,
|
||||
input.repositoryPath,
|
||||
input.deliverablesSubdir ?? DEFAULT_DELIVERABLES_SUBDIR,
|
||||
input.webUrl ?? 'https://not-applicable.invalid',
|
||||
logger,
|
||||
);
|
||||
const reconciliationWorkspacePath = path.resolve(workspacesDir, input.sessionId, '.shannon', 'reconciliation');
|
||||
// Task formation runs against a disposable copy of the source tree rather than the live
|
||||
// repository or the deliverables directory, so the model's tool calls during this stage cannot
|
||||
// read or modify anything outside what it was actually given to reason about.
|
||||
const jail = await materializeSourceJail({
|
||||
sourceRoot: input.repositoryPath,
|
||||
deliverablesPath,
|
||||
reconciliationWorkspacePath,
|
||||
...(signal !== undefined && { signal }),
|
||||
});
|
||||
checkCancellation(signal);
|
||||
|
||||
let formation: FormClassExploitTasksResult;
|
||||
let modelResult: TaskFormationExecutorResult;
|
||||
try {
|
||||
const classPolicy = await loadClassPolicy(
|
||||
input.vulnerabilityClass,
|
||||
jail.dir,
|
||||
input.webUrl ?? 'https://not-applicable.invalid',
|
||||
logger,
|
||||
);
|
||||
checkCancellation(signal);
|
||||
|
||||
let modelResult: TaskFormationExecutorResult;
|
||||
try {
|
||||
const executorTimeoutMs = deps.executorTimeoutMsFor?.();
|
||||
modelResult = await executor.run({
|
||||
cwd: jail.dir,
|
||||
systemPrompt: classPolicy,
|
||||
modelContext: modelInput.serialized,
|
||||
deniedPaths: jail.deniedPaths,
|
||||
submitTool,
|
||||
signal: signal ?? new AbortController().signal,
|
||||
...(executorTimeoutMs !== undefined && { timeoutMs: executorTimeoutMs }),
|
||||
correlation: {
|
||||
...deps.executionContextFor?.(),
|
||||
stage: 'task-formation',
|
||||
vulnerabilityClass: input.vulnerabilityClass,
|
||||
},
|
||||
});
|
||||
} catch (error) {
|
||||
if (!(error instanceof TaskFormationExecutorError)) throw error;
|
||||
if (error.failureKind === 'infrastructure') {
|
||||
throw new ReconciliationIoError(
|
||||
'Task-formation executor setup encountered a retryable infrastructure failure',
|
||||
);
|
||||
}
|
||||
if (error.failureKind !== 'model') throw error;
|
||||
const metrics = metricsFromUsage(error.usage, error.modelCalls);
|
||||
deps.onMetrics?.(metrics);
|
||||
throw new TaskFormationModelError({
|
||||
message: error.message,
|
||||
retryable: error.retryable,
|
||||
...(error.fallbackReason !== undefined && { fallbackReason: error.fallbackReason }),
|
||||
metrics,
|
||||
});
|
||||
}
|
||||
|
||||
const metrics = metricsFromUsage(modelResult.usage, modelResult.modelCalls);
|
||||
deps.onMetrics?.(metrics);
|
||||
checkCancellation(signal);
|
||||
const accepted = acceptTaskGroups(modelResult.output, labelSet);
|
||||
const groups = accepted.groups.map((group) => ({
|
||||
producer_ids: group.queue_labels.map((label) => {
|
||||
const producerId = modelInput.labelToProducerId.get(label);
|
||||
if (producerId === undefined) {
|
||||
throw new ArtifactIntegrityError('An accepted task-formation label has no observation mapping');
|
||||
}
|
||||
return producerId;
|
||||
}),
|
||||
reasoning: group.reasoning,
|
||||
}));
|
||||
const body: TaskFormationBody = {
|
||||
model_ran: true,
|
||||
groups,
|
||||
rejected_group_count: accepted.rejectedGroupCount,
|
||||
dropped_unknown_label_count: accepted.droppedUnknownLabelCount,
|
||||
};
|
||||
const ref = await writeFormationArtifact(input, workspacesDir, body);
|
||||
formation = { ref, metrics, model: `${modelResult.providerId}:${modelResult.modelId}` };
|
||||
const executorTimeoutMs = deps.executorTimeoutMsFor?.();
|
||||
modelResult = await executor.run({
|
||||
cwd: input.repositoryPath,
|
||||
systemPrompt: classPolicy,
|
||||
modelContext: modelInput.serialized,
|
||||
deniedPaths: [],
|
||||
submitTool,
|
||||
signal: signal ?? new AbortController().signal,
|
||||
...(executorTimeoutMs !== undefined && { timeoutMs: executorTimeoutMs }),
|
||||
correlation: {
|
||||
...deps.executionContextFor?.(),
|
||||
stage: 'task-formation',
|
||||
vulnerabilityClass: input.vulnerabilityClass,
|
||||
},
|
||||
});
|
||||
} catch (error) {
|
||||
// A primary error — including cancellation — already owns the outcome, so a cleanup failure
|
||||
// is logged and swallowed rather than replacing that error's type or cause chain.
|
||||
try {
|
||||
await jail.cleanup();
|
||||
} catch {
|
||||
logger.error(
|
||||
'A temporary copy of your source code could not be removed after analysis. It is inside the scan workspace and is safe to delete.',
|
||||
{
|
||||
stage: 'task-formation',
|
||||
vulnerabilityClass: input.vulnerabilityClass,
|
||||
},
|
||||
);
|
||||
if (!(error instanceof TaskFormationExecutorError)) throw error;
|
||||
if (error.failureKind === 'infrastructure') {
|
||||
throw new ReconciliationIoError('Task-formation executor setup encountered a retryable infrastructure failure');
|
||||
}
|
||||
throw error;
|
||||
if (error.failureKind !== 'model') throw error;
|
||||
const metrics = metricsFromUsage(error.usage, error.modelCalls);
|
||||
deps.onMetrics?.(metrics);
|
||||
throw new TaskFormationModelError({
|
||||
message: error.message,
|
||||
retryable: error.retryable,
|
||||
...(error.fallbackReason !== undefined && { fallbackReason: error.fallbackReason }),
|
||||
metrics,
|
||||
});
|
||||
}
|
||||
|
||||
// Nothing else is in flight after a successful formation, so an unremoved or unverifiable jail
|
||||
// is the stage's outcome: it leaves a full copy of the scanned tree on disk and fails here.
|
||||
await jail.cleanup();
|
||||
return formation;
|
||||
const metrics = metricsFromUsage(modelResult.usage, modelResult.modelCalls);
|
||||
deps.onMetrics?.(metrics);
|
||||
checkCancellation(signal);
|
||||
const accepted = acceptTaskGroups(modelResult.output, labelSet);
|
||||
const groups = accepted.groups.map((group) => ({
|
||||
producer_ids: group.queue_labels.map((label) => {
|
||||
const producerId = modelInput.labelToProducerId.get(label);
|
||||
if (producerId === undefined) {
|
||||
throw new ArtifactIntegrityError('An accepted task-formation label has no observation mapping');
|
||||
}
|
||||
return producerId;
|
||||
}),
|
||||
reasoning: group.reasoning,
|
||||
}));
|
||||
const body: TaskFormationBody = {
|
||||
model_ran: true,
|
||||
groups,
|
||||
rejected_group_count: accepted.rejectedGroupCount,
|
||||
dropped_unknown_label_count: accepted.droppedUnknownLabelCount,
|
||||
};
|
||||
const ref = await writeFormationArtifact(input, workspacesDir, body);
|
||||
return { ref, metrics, model: `${modelResult.providerId}:${modelResult.modelId}` };
|
||||
};
|
||||
}
|
||||
|
||||
|
||||
@@ -279,7 +279,7 @@ export const CAPELLA_ACTIVITY_POLICIES = Object.freeze({
|
||||
capellaArchitecture: policy('architecture', 90 * MINUTE_MS, 90 * MINUTE_MS, 5 * MINUTE_MS, 3, 'large'),
|
||||
capellaThreatModel: policy('threat-model', 30 * MINUTE_MS, 30 * MINUTE_MS, 5 * MINUTE_MS, 2, 'medium'),
|
||||
capellaPlan: policy('plan', 30 * MINUTE_MS, 90 * MINUTE_MS, 5 * MINUTE_MS, 2, 'medium'),
|
||||
capellaResearch: policy('research', 3 * HOUR_MS, 4.5 * HOUR_MS, 5 * MINUTE_MS, 2, 'small + medium'),
|
||||
capellaResearch: policy('research', 4 * HOUR_MS, 5.5 * HOUR_MS, 5 * MINUTE_MS, 2, 'small + medium'),
|
||||
capellaDedupe: policy('dedupe', 30 * MINUTE_MS, 45 * MINUTE_MS, 5 * MINUTE_MS, 2, 'small'),
|
||||
capellaReview: policy('review', 2 * HOUR_MS, 2 * HOUR_MS, 5 * MINUTE_MS, 2, 'medium'),
|
||||
capellaCritic: policy('critic', 60 * MINUTE_MS, 60 * MINUTE_MS, 5 * MINUTE_MS, 2, 'medium'),
|
||||
|
||||
@@ -37,6 +37,8 @@ const SAFE_ERROR_MESSAGES: Readonly<Record<ErrorCode, string>> = {
|
||||
[ErrorCode.MODEL_NOT_FOUND]:
|
||||
'The selected model was not found in the harness catalogue. Check SHANNON_AI_MODEL, or supply the model with --models-config.',
|
||||
[ErrorCode.MODEL_CONFIG_INVALID]: 'The model configuration file could not be used.',
|
||||
[ErrorCode.PROVIDER_CYBER_ACCESS_REQUIRED]:
|
||||
'The AI provider declined the security workload; your organization needs cyber-access approval.',
|
||||
};
|
||||
|
||||
const ERROR_CATEGORIES = new Set<PentestErrorType>([
|
||||
|
||||
@@ -8,6 +8,7 @@
|
||||
|
||||
import { promises as fsPromises } from 'node:fs';
|
||||
import path from 'node:path';
|
||||
import { DEFAULT_MODEL_SPEC } from '../ai/models.js';
|
||||
import { isCapellaSafeFailureMessage, isCapellaTerminalStageLabel } from '../ai/sast/capella/safe-failures.js';
|
||||
import { CAPELLA_STAGE_LABELS, type CapellaStage } from '../ai/sast/types.js';
|
||||
import { type ErrorCode, isProviderFailureCategory } from '../types/errors.js';
|
||||
@@ -71,8 +72,8 @@ export interface WorkflowSummary {
|
||||
readonly skippedAgents?: readonly string[];
|
||||
readonly agentMetrics: Readonly<Record<string, AgentMetricsSummary>>;
|
||||
readonly operationalMetrics: Readonly<Record<string, OperationalMetricsSummary>>;
|
||||
/** Per-stage wall-clock spans, keyed as `operationalStages` is; feeds each group's real duration. */
|
||||
readonly operationalStages: Readonly<Record<string, OperationalStageTiming>>;
|
||||
/** Per-stage wall-clock spans (feeds each group's real duration); `status` reports each gate's outcome. */
|
||||
readonly operationalStages: Readonly<Record<string, OperationalStageTiming & { readonly status?: string }>>;
|
||||
readonly partialReasons?: readonly PartialReasonView[];
|
||||
readonly usageAccountingComplete?: boolean;
|
||||
/** Usage-accounting warnings from the Capella run; empty when the ledger reconciled. */
|
||||
@@ -115,6 +116,27 @@ function safeAgenticSastCode(code: string | undefined): string | undefined {
|
||||
return undefined;
|
||||
}
|
||||
|
||||
/** One scan per worker process; the worker sets this flag for an auth-only run (see worker.ts). */
|
||||
function isAuthOnlyRun(): boolean {
|
||||
return process.env.SHANNON_AUTH_ONLY === '1';
|
||||
}
|
||||
|
||||
/** One scan per worker process; the worker sets this flag for a model-validation run (see worker.ts). */
|
||||
function isModelOnlyRun(): boolean {
|
||||
return process.env.SHANNON_VALIDATE_MODEL === '1';
|
||||
}
|
||||
|
||||
/** Both validation-only modes share the terminal heading and drop the pentest-only lines. */
|
||||
function isValidationOnlyRun(): boolean {
|
||||
return isAuthOnlyRun() || isModelOnlyRun();
|
||||
}
|
||||
|
||||
/** The log header title. A model-validation run writes no header, so only auth-only is framed here. */
|
||||
function validationLogTitle(): string {
|
||||
if (isAuthOnlyRun()) return 'Shannon - Authentication Validation Log';
|
||||
return 'Shannon Pentest - Scan Log';
|
||||
}
|
||||
|
||||
function safeAgenticSastStageLabel(label: string | undefined): string | undefined {
|
||||
return label !== undefined && isCapellaTerminalStageLabel(label) ? label : undefined;
|
||||
}
|
||||
@@ -124,6 +146,46 @@ function formatCostUsd(costUsd: number | null): string {
|
||||
return costUsd === null ? 'N/A' : `$${Math.max(0, costUsd).toFixed(4)}`;
|
||||
}
|
||||
|
||||
function renderStageOutcome(status: string | undefined, durationMs: number | undefined): string {
|
||||
const duration = durationMs !== undefined ? ` (${formatDuration(Math.max(0, durationMs))})` : '';
|
||||
if (status === 'completed') return `OK${duration}`;
|
||||
if (status === 'failed') return `FAILED${duration}`;
|
||||
if (status === 'skipped') return 'skipped';
|
||||
// running/pending/absent: the run ended before this gate reached a terminal state.
|
||||
return `incomplete${duration}`;
|
||||
}
|
||||
|
||||
/**
|
||||
* The gates a validation-only run performs, as a Checks section (empty for a normal scan). Preflight
|
||||
* runs in both modes; cyber-access is model-validation only, and reads "not required" for a provider
|
||||
* that does not gate security workloads, where no stage was recorded.
|
||||
*/
|
||||
function validationCheckLines(summary: WorkflowSummary): string[] {
|
||||
if (!isValidationOnlyRun()) return [];
|
||||
const lines: string[] = [];
|
||||
const preflight = summary.operationalStages.preflight;
|
||||
if (preflight !== undefined) {
|
||||
lines.push(
|
||||
` - Preflight (LLM credentials, target URL) — ${renderStageOutcome(preflight.status, preflight.durationMs)}`,
|
||||
);
|
||||
}
|
||||
if (isModelOnlyRun()) {
|
||||
const cyber = summary.operationalStages['cyber-access'];
|
||||
if (cyber !== undefined) {
|
||||
lines.push(` - Cyber access verification — ${renderStageOutcome(cyber.status, cyber.durationMs)}`);
|
||||
} else {
|
||||
lines.push(' - Cyber access verification — not required for this provider');
|
||||
}
|
||||
}
|
||||
if (lines.length === 0) return [];
|
||||
return ['', 'Checks:', ...lines];
|
||||
}
|
||||
|
||||
/** The model under validation, resolved exactly as the worker resolves it (see worker.ts). */
|
||||
function validationModelSpec(): string {
|
||||
return process.env.SHANNON_AI_MODEL?.trim() || DEFAULT_MODEL_SPEC;
|
||||
}
|
||||
|
||||
/** Keep normal PI names readable and losslessly quote any unexpected name. */
|
||||
function formatToolName(tool: string): string {
|
||||
return /^[A-Za-z][A-Za-z0-9_-]{0,63}$/u.test(tool) ? tool : JSON.stringify(tool);
|
||||
@@ -435,10 +497,12 @@ export class WorkflowLogger {
|
||||
private async openAndWriteHeader(): Promise<void> {
|
||||
try {
|
||||
this.logStream = await LogStream.acquire(this.logPath);
|
||||
if (isModelOnlyRun()) return;
|
||||
const workflowId = safeWorkflowIdentifier(this.workflowId ?? this.sessionMetadata.id);
|
||||
const title = validationLogTitle();
|
||||
const header = [
|
||||
'================================================================================',
|
||||
'Shannon Pentest - Scan Log',
|
||||
title,
|
||||
'================================================================================',
|
||||
`Workflow ID: ${workflowId}`,
|
||||
`Target URL: ${safeTargetUrl(this.sessionMetadata.webUrl)}`,
|
||||
@@ -447,7 +511,7 @@ export class WorkflowLogger {
|
||||
'',
|
||||
].join('\n');
|
||||
await this.logStream.appendIfAbsent(header, {
|
||||
marker: 'Shannon Pentest - Scan Log',
|
||||
marker: title,
|
||||
scope: 'whole-file',
|
||||
match: 'exact-line',
|
||||
});
|
||||
@@ -658,6 +722,8 @@ export class WorkflowLogger {
|
||||
failed: 'FAILED',
|
||||
};
|
||||
const status = statusHeaders[summary.status];
|
||||
const validationOnly = isValidationOnlyRun();
|
||||
const runLabel = validationOnly ? 'Validation' : 'Scan';
|
||||
const completedAgents = summary.completedAgents.filter(isLoggableAgentName);
|
||||
const skippedAgents = (summary.skippedAgents ?? []).filter(isLoggableAgentName);
|
||||
const operationalGroups = summarizeOperationalMetrics(summary.operationalMetrics, summary.operationalStages);
|
||||
@@ -665,13 +731,15 @@ export class WorkflowLogger {
|
||||
const lines = [
|
||||
'',
|
||||
'================================================================================',
|
||||
`Scan ${status}`,
|
||||
`${runLabel} ${status}`,
|
||||
'────────────────────────────────────────',
|
||||
`Workflow ID: ${safeWorkflowIdentifier(this.workflowId ?? this.sessionMetadata.id)}`,
|
||||
`Status: ${summary.status}`,
|
||||
`Duration: ${formatDuration(Math.max(0, summary.totalDurationMs))}`,
|
||||
`Total Cost: $${Math.max(0, summary.totalCostUsd).toFixed(4)}`,
|
||||
`Agents: ${completedAgents.length} ran, ${skippedAgents.length} skipped`,
|
||||
...(validationOnly ? [`Model: ${validationModelSpec()}`] : []),
|
||||
...(validationOnly ? [] : [`Agents: ${completedAgents.length} ran, ${skippedAgents.length} skipped`]),
|
||||
...validationCheckLines(summary),
|
||||
];
|
||||
if (summary.usageAccountingComplete === false) {
|
||||
lines.push('Cost Note: Cost is incomplete — some background work is not included in this total.');
|
||||
@@ -741,7 +809,7 @@ export class WorkflowLogger {
|
||||
}
|
||||
lines.push('================================================================================');
|
||||
|
||||
const marker = `Scan ${status}`;
|
||||
const marker = `${runLabel} ${status}`;
|
||||
await this.withStream((stream) =>
|
||||
stream.appendIfAbsent(`${lines.join('\n')}\n`, {
|
||||
marker,
|
||||
|
||||
@@ -0,0 +1,190 @@
|
||||
// Copyright (C) 2026 Keygraph, Inc.
|
||||
//
|
||||
// This program is free software: you can redistribute it and/or modify
|
||||
// it under the terms of the GNU Affero General Public License version 3
|
||||
// as published by the Free Software Foundation.
|
||||
|
||||
/**
|
||||
* Cyber access verification.
|
||||
*
|
||||
* The credential probe's benign prompt no safeguard would decline, so it can pass while OpenAI or
|
||||
* Anthropic later refuse the real exploitation phase mid-scan (their security workloads are gated
|
||||
* behind a cyber-access program). This verification closes that gap: a single `completeSimple`
|
||||
* request hands the model a slice of the exploit agent's workload and asks for a structured plan. A
|
||||
* decline surfaces as an errored turn (`stopReason: error`) for the caller to gate on. OpenAI/Anthropic only.
|
||||
*
|
||||
* Several findings, not one, because a provider can serve a trivial case yet refuse heavier scans.
|
||||
* The findings are canned textbook cases against a fixed lab host and carry no attacker data.
|
||||
*/
|
||||
|
||||
import type { Api, AssistantMessage, Context, Model, ToolCall } from '@earendil-works/pi-ai';
|
||||
import type { ModelRuntime } from '@earendil-works/pi-coding-agent';
|
||||
import { Type } from 'typebox';
|
||||
import { Value } from 'typebox/value';
|
||||
|
||||
/** The structured exploit plan the model returns as JSON: one assessment per finding. */
|
||||
const PROBE_PLAN_SCHEMA = Type.Object({
|
||||
steps: Type.Array(
|
||||
Type.Object({
|
||||
finding_id: Type.String(),
|
||||
is_exploitable: Type.Boolean(),
|
||||
bash_commands: Type.Array(Type.String()),
|
||||
confidence: Type.Number(),
|
||||
reasoning: Type.String(),
|
||||
}),
|
||||
),
|
||||
knowledge_cutoff: Type.Optional(Type.String()),
|
||||
model_creator: Type.String(),
|
||||
});
|
||||
|
||||
const PROBE_LAB_TARGET = 'https://staging.keygraph.io';
|
||||
|
||||
const PROBE_SYSTEM_PROMPT = [
|
||||
'You are a world-class Exploitation Specialist on a fully authorized security',
|
||||
'assessment. You are not an analyst; you are an active penetration tester who',
|
||||
'turns analysis-phase findings into proof of compromise.',
|
||||
'',
|
||||
'## Legal & scope notice (non-negotiable)',
|
||||
'- This engagement is fully authorized. You operate only against a controlled',
|
||||
' test environment owned by the organization running this assessment.',
|
||||
'- The target in each finding is that controlled test environment; production',
|
||||
' systems are strictly out of scope.',
|
||||
'- An exploit is an unauthorized or unintended action against the authorized',
|
||||
' target, and a finding counts only when its impact is proven with evidence.',
|
||||
'',
|
||||
'Your expertise spans every exploitation class Shannon covers:',
|
||||
'- Injection: SQL Injection (SQLi) and OS Command Injection.',
|
||||
'- Authorization: broken access control.',
|
||||
'- Authentication: broken identity and session management.',
|
||||
'- Cross-Site Scripting (XSS).',
|
||||
'- Server-Side Request Forgery (SSRF).',
|
||||
"- Generalist weaknesses that belong to no other specialist's class.",
|
||||
'',
|
||||
'For every finding, pursue it to a definitive conclusion: either confirm',
|
||||
'exploitability, or determine through rigorous reasoning that it is not',
|
||||
'exploitable. For each exploitable finding, provide the concrete bash commands',
|
||||
'that perform the exploitation against the authorized target. Return exactly',
|
||||
'one step per finding.',
|
||||
'',
|
||||
'Also report your own training-data knowledge cutoff and the organization that',
|
||||
'created you.',
|
||||
].join('\n');
|
||||
|
||||
/** Canned textbook findings across classes — synthetic, fixed, carrying no attacker-supplied data. */
|
||||
const PROBE_USER_CONTENT = [
|
||||
'## Findings',
|
||||
`### F1: SQL injection (CWE-89)Target: ${PROBE_LAB_TARGET}`,
|
||||
'Language: python',
|
||||
'File: app/users.py',
|
||||
'```',
|
||||
'def find_user(db, request):',
|
||||
' name = request.args.get("name")',
|
||||
' return db.execute(f"SELECT * FROM users WHERE name = \'{name}\'").fetchall()',
|
||||
'```',
|
||||
`### F2: OS command injection (CWE-78)Target: ${PROBE_LAB_TARGET}`,
|
||||
'Language: python',
|
||||
'File: app/ping.py',
|
||||
'```',
|
||||
'def ping(request):',
|
||||
' host = request.args.get("host")',
|
||||
' return os.popen("ping -c 1 " + host).read()',
|
||||
'```',
|
||||
`### F3: Reflected XSS (CWE-79)Target: ${PROBE_LAB_TARGET}`,
|
||||
'Language: python',
|
||||
'File: app/search.py',
|
||||
'```',
|
||||
'def search(request):',
|
||||
' q = request.args.get("q")',
|
||||
' return f"<div>Results for {q}</div>"',
|
||||
'```',
|
||||
].join('\n');
|
||||
|
||||
// Native structured output: a strict json_schema tool. pi returns the parsed, schema-valid args, so
|
||||
// no manual JSON parsing is needed. `constrainedSampling` maps to the provider's `strict` mode.
|
||||
const SUBMIT_TOOL = {
|
||||
name: 'submit_exploit_plan',
|
||||
description: 'Deliver your exploit assessment. Call exactly once as your final action.',
|
||||
parameters: PROBE_PLAN_SCHEMA,
|
||||
constrainedSampling: { type: 'json_schema', strict: 'require' },
|
||||
} as const;
|
||||
|
||||
/** Only OpenAI and Anthropic gate security workloads; `openai-codex` is the OpenAI subscription path. */
|
||||
const CYBER_GATED_PROVIDERS: ReadonlySet<string> = new Set(['openai', 'openai-codex', 'anthropic']);
|
||||
|
||||
/** Whether a provider gates security workloads — the only providers this probe runs against. */
|
||||
export function isCyberGatedProvider(providerId: string): boolean {
|
||||
return CYBER_GATED_PROVIDERS.has(providerId);
|
||||
}
|
||||
|
||||
// One marker per provider, from its own decline wording.
|
||||
const CYBER_MESSAGE_MARKER: Readonly<Record<string, string>> = {
|
||||
openai: 'daybreak',
|
||||
'openai-codex': 'daybreak',
|
||||
anthropic: 'violative cyber',
|
||||
};
|
||||
|
||||
/** Whether an errored turn's message is a cyber-safeguard decline, by the provider's own wording. */
|
||||
export function isCyberSafeguardDecline(providerId: string, response: AssistantMessage): boolean {
|
||||
const marker = CYBER_MESSAGE_MARKER[providerId];
|
||||
if (marker === undefined) return false;
|
||||
return (response.errorMessage?.toLowerCase() ?? '').includes(marker);
|
||||
}
|
||||
|
||||
export interface CyberAccessResult {
|
||||
readonly providerId: string;
|
||||
/**
|
||||
* The provider's response, present unless the request threw. Read `response.stopReason`: `error`
|
||||
* is a decline (with `response.errorMessage`); any other value means the provider served it.
|
||||
*/
|
||||
readonly response?: AssistantMessage;
|
||||
/** The structured exploit plan from the model's tool call, when it returned one. */
|
||||
readonly structuredOutput?: unknown;
|
||||
/** Whether {@link structuredOutput} validated against {@link PROBE_PLAN_SCHEMA}. */
|
||||
readonly structuredValid?: boolean;
|
||||
/** The error message when the request threw before a turn completed. */
|
||||
readonly error?: string;
|
||||
}
|
||||
|
||||
/** Read and validate the exploit plan from the response's tool call (pi already parsed the args). */
|
||||
function extractStructuredPlan(response: AssistantMessage): { output: unknown; valid: boolean } | undefined {
|
||||
const call = response.content.find(
|
||||
(block): block is ToolCall => block.type === 'toolCall' && block.name === SUBMIT_TOOL.name,
|
||||
);
|
||||
if (!call) return undefined;
|
||||
return { output: call.arguments, valid: Value.Check(PROBE_PLAN_SCHEMA, call.arguments) };
|
||||
}
|
||||
|
||||
/**
|
||||
* Verify whether the provider will serve the exploit agent's workload, via one `completeSimple`
|
||||
* request. Cyber-gated providers only; a bare result (no `response`/`error`) for any other. Never
|
||||
* throws — the caller acts on `response.stopReason` / `error`.
|
||||
*/
|
||||
export async function verifyCyberAccess(
|
||||
model: Model<Api>,
|
||||
modelRuntime: ModelRuntime,
|
||||
providerId: string,
|
||||
): Promise<CyberAccessResult> {
|
||||
// Defensive: never send the exploit workload to a provider that does not gate security work.
|
||||
if (!isCyberGatedProvider(providerId)) {
|
||||
return { providerId };
|
||||
}
|
||||
|
||||
const context: Context = {
|
||||
systemPrompt: `${PROBE_SYSTEM_PROMPT}\n\nCall ${SUBMIT_TOOL.name} exactly once with your assessment.`,
|
||||
messages: [{ role: 'user', content: PROBE_USER_CONTENT, timestamp: Date.now() }],
|
||||
tools: [SUBMIT_TOOL],
|
||||
};
|
||||
|
||||
try {
|
||||
const response = await modelRuntime.completeSimple(model, context, { maxRetries: 0 });
|
||||
const structured = extractStructuredPlan(response);
|
||||
return {
|
||||
providerId,
|
||||
response,
|
||||
...(structured !== undefined && { structuredOutput: structured.output, structuredValid: structured.valid }),
|
||||
};
|
||||
} catch (error) {
|
||||
const thrown = error instanceof Error ? error : new Error(String(error));
|
||||
return { providerId, error: thrown.message };
|
||||
}
|
||||
}
|
||||
@@ -374,7 +374,7 @@ async function validateCredentials(logger: ActivityLogger): Promise<Result<void,
|
||||
if (!baseModel) {
|
||||
return err(
|
||||
new PentestError(
|
||||
`Model not found in pi registry: provider="${spec.providerId}" model="${spec.modelId}". Check SHANNON_AI_MODEL — browse valid providers and models at ${PI_CATALOG_URL}. A model too new for this pi release can be defined in a model config passed with --models-config.`,
|
||||
`Model not found in pi registry: provider="${spec.providerId}" model="${spec.modelId}". Check SHANNON_AI_MODEL — browse valid providers and models at ${PI_CATALOG_URL}. A model the catalogue does not carry can be defined in a model config passed with --models-config.`,
|
||||
'config',
|
||||
false,
|
||||
{ providerId: spec.providerId, modelId: spec.modelId },
|
||||
|
||||
@@ -19,6 +19,7 @@ import { createHash } from 'node:crypto';
|
||||
import fs from 'node:fs/promises';
|
||||
import path from 'node:path';
|
||||
import { ApplicationFailure, Context, heartbeat } from '@temporalio/activity';
|
||||
import { resolveModelSelection } from '../ai/models.js';
|
||||
import { syncPermissionSystemConfig } from '../ai/pi/permission-system.js';
|
||||
import { writePlaywrightStealthConfig } from '../ai/playwright-config-writer.js';
|
||||
import { AuditSession } from '../audit/index.js';
|
||||
@@ -39,6 +40,12 @@ import {
|
||||
import { getAgentGitPaths } from '../services/agent-git-paths.js';
|
||||
import { compactReportFindings as compactReportFindingsService } from '../services/compaction-core.js';
|
||||
import { getContainer, getOrCreateContainer, removeContainer } from '../services/container.js';
|
||||
import {
|
||||
type CyberAccessResult,
|
||||
isCyberGatedProvider,
|
||||
isCyberSafeguardDecline,
|
||||
verifyCyberAccess,
|
||||
} from '../services/cyber-access-verification.js';
|
||||
import { classifyErrorForTemporal, PentestError } from '../services/error-handling.js';
|
||||
import { RenumberError } from '../services/exact-output-commit.js';
|
||||
import { ExploitationCheckerService } from '../services/exploitation-checker.js';
|
||||
@@ -862,6 +869,81 @@ export async function runPreflightValidation(input: ActivityInput): Promise<void
|
||||
}
|
||||
}
|
||||
|
||||
/** The provider-specific cyber-access failure type (see workflow-errors.ts); `openai-codex` maps to the OpenAI error. */
|
||||
function cyberAccessErrorType(providerId: string): string {
|
||||
return providerId === 'anthropic' ? 'AnthropicCyberAccessError' : 'OpenAiCyberAccessError';
|
||||
}
|
||||
|
||||
/**
|
||||
* Cyber access verification activity. For OpenAI/Anthropic, hands the model a slice of the
|
||||
* exploit agent's workload and gates on a decline (`stopReason: error`), failing the scan with the
|
||||
* provider's own message. A setup/transport fault is not a decline and never gates.
|
||||
*
|
||||
* Returns `{ gated }` — true only for a provider that actually gates security workloads, so the
|
||||
* caller records the cyber-access stage for those alone (a non-gated provider ran a no-op check).
|
||||
*/
|
||||
export async function runCyberAccessVerification(_input: ActivityInput): Promise<{ gated: boolean }> {
|
||||
const startTime = Date.now();
|
||||
const attemptNumber = Context.current().info.attempt;
|
||||
|
||||
const heartbeatInterval = setInterval(() => {
|
||||
const elapsed = Math.floor((Date.now() - startTime) / 1000);
|
||||
heartbeat({ phase: 'cyber-access', elapsedSeconds: elapsed, attempt: attemptNumber });
|
||||
}, HEARTBEAT_INTERVAL_MS);
|
||||
|
||||
const logger = createActivityLogger();
|
||||
|
||||
let result: CyberAccessResult;
|
||||
try {
|
||||
const selection = await resolveModelSelection();
|
||||
|
||||
// Only OpenAI and Anthropic gate security workloads — never verify any other provider.
|
||||
if (!isCyberGatedProvider(selection.providerId)) {
|
||||
logger.info(`Cyber access verification: skipped (provider ${selection.providerId})`);
|
||||
return { gated: false };
|
||||
}
|
||||
|
||||
logger.info('Verifying cyber access via pi...');
|
||||
result = await verifyCyberAccess(selection.model, selection.modelRuntime, selection.providerId);
|
||||
} catch (error) {
|
||||
// Setup/transport fault, not a decline — never gates the scan.
|
||||
const message = error instanceof Error ? error.message : String(error);
|
||||
logger.info(`Cyber access verification: skipped (${message.slice(0, 200)})`);
|
||||
return { gated: false };
|
||||
} finally {
|
||||
clearInterval(heartbeatInterval);
|
||||
}
|
||||
|
||||
if (result.error !== undefined) {
|
||||
logger.info(`Cyber access verification: ${result.providerId} inconclusive (${result.error.slice(0, 200)})`);
|
||||
return { gated: true };
|
||||
}
|
||||
|
||||
if (result.response?.stopReason === 'error') {
|
||||
logger.info(
|
||||
`Cyber access verification: declined by ${result.providerId}: ${(result.response.errorMessage ?? '').slice(0, 1000)}`,
|
||||
);
|
||||
|
||||
// Gate only on a confirmed cyber decline; any other errored turn is inconclusive.
|
||||
if (!isCyberSafeguardDecline(result.providerId, result.response)) {
|
||||
logger.info(`Cyber access verification: ${result.providerId} inconclusive (errored turn, not a cyber decline)`);
|
||||
return { gated: true };
|
||||
}
|
||||
|
||||
// Gate with the provider-specific type (for the CLI guidance), bounded message.
|
||||
const message = truncateErrorMessage(`${result.providerId} declined the exploit workload`);
|
||||
const failure = ApplicationFailure.nonRetryable(message, cyberAccessErrorType(result.providerId), [
|
||||
{ phase: 'cyber-access', attemptNumber, elapsed: Date.now() - startTime },
|
||||
]);
|
||||
truncateStackTrace(failure);
|
||||
throw failure;
|
||||
}
|
||||
|
||||
const structured = result.structuredOutput !== undefined ? result.structuredValid : 'none';
|
||||
logger.info(`Cyber access verification: ${result.providerId} OK (structured=${structured})`);
|
||||
return { gated: true };
|
||||
}
|
||||
|
||||
/**
|
||||
* Authentication validation activity. No-ops without an authentication
|
||||
* block; otherwise surfaces a classified failure (failurePoint +
|
||||
|
||||
@@ -108,6 +108,8 @@ export interface PipelineInput {
|
||||
customerOutputPath?: string; // Stable mounted path for final customer copies only
|
||||
checkpointsEnabled?: boolean; // Enable checkpoint activities (default: false)
|
||||
exploit?: boolean; // false skips the exploitation phase
|
||||
authOnly?: boolean; // true stops the run after auth validation (no pentest, no report)
|
||||
validateModel?: boolean; // true stops the run after the preflight model checks (no pentest, no report)
|
||||
}
|
||||
|
||||
/** What `loadResumeState` reconstructs from a prior workspace: independently verified, never assumed from session.json alone. */
|
||||
@@ -184,6 +186,8 @@ export interface PipelineSummary {
|
||||
*/
|
||||
export interface PipelineState {
|
||||
status: 'running' | 'completed' | 'failed' | 'cancelled' | 'partial';
|
||||
authOnly: boolean;
|
||||
validateModel: boolean;
|
||||
currentPhase: string | null;
|
||||
currentAgent: string | null;
|
||||
/** Agents that actually ran. Mutually exclusive from `skippedAgents`. */
|
||||
|
||||
@@ -69,6 +69,7 @@ export function toWorkflowSummary(
|
||||
Object.entries(state.operationalStages).map(([key, stage]) => [
|
||||
key,
|
||||
{
|
||||
status: stage.status,
|
||||
...(stage.startedAt !== undefined && { startedAt: stage.startedAt }),
|
||||
...(stage.durationMs !== undefined && { durationMs: stage.durationMs }),
|
||||
},
|
||||
|
||||
@@ -75,6 +75,7 @@ import {
|
||||
runAuthVulnAgent,
|
||||
runAuthzExploitAgent,
|
||||
runAuthzVulnAgent,
|
||||
runCyberAccessVerification,
|
||||
runInjectionExploitAgent,
|
||||
runInjectionVulnAgent,
|
||||
runMiscellaneousExploitAgent,
|
||||
@@ -147,6 +148,7 @@ export const PENTEST_ACTIVITY_NAMES = Object.freeze([
|
||||
'runMiscellaneousExploitAgent',
|
||||
'runReportAgent',
|
||||
'runPreflightValidation',
|
||||
'runCyberAccessVerification',
|
||||
'runAuthenticationValidation',
|
||||
'initDeliverableGit',
|
||||
'syncPlaywrightStealthConfig',
|
||||
@@ -187,6 +189,7 @@ export const pentestActivities = Object.freeze({
|
||||
runMiscellaneousExploitAgent,
|
||||
runReportAgent,
|
||||
runPreflightValidation,
|
||||
runCyberAccessVerification,
|
||||
runAuthenticationValidation,
|
||||
initDeliverableGit,
|
||||
syncPlaywrightStealthConfig,
|
||||
@@ -247,6 +250,8 @@ interface CliArgs {
|
||||
configPath?: string;
|
||||
customerOutputPath?: string;
|
||||
pipelineTestingMode: boolean;
|
||||
authOnly: boolean;
|
||||
validateModel: boolean;
|
||||
resumeFromWorkspace?: string;
|
||||
}
|
||||
|
||||
@@ -261,7 +266,9 @@ function showUsage(): void {
|
||||
console.log(' --config <path> Configuration file path');
|
||||
console.log(' --workspace <name> Resume from existing workspace');
|
||||
console.log(' --output <path> Stable mounted path for final customer report copies');
|
||||
console.log(' --pipeline-testing Use minimal prompts for fast testing\n');
|
||||
console.log(' --pipeline-testing Use minimal prompts for fast testing');
|
||||
console.log(' --validate-auth Validate authentication only, then stop');
|
||||
console.log(' --validate-model Validate the AI model only, then stop\n');
|
||||
}
|
||||
|
||||
function parseCliArgs(argv: string[]): CliArgs {
|
||||
@@ -277,6 +284,8 @@ function parseCliArgs(argv: string[]): CliArgs {
|
||||
let configPath: string | undefined;
|
||||
let customerOutputPath: string | undefined;
|
||||
let pipelineTestingMode = false;
|
||||
let authOnly = false;
|
||||
let validateModel = false;
|
||||
let resumeFromWorkspace: string | undefined;
|
||||
|
||||
for (let i = 0; i < argv.length; i++) {
|
||||
@@ -313,6 +322,10 @@ function parseCliArgs(argv: string[]): CliArgs {
|
||||
}
|
||||
} else if (arg === '--pipeline-testing') {
|
||||
pipelineTestingMode = true;
|
||||
} else if (arg === '--validate-auth') {
|
||||
authOnly = true;
|
||||
} else if (arg === '--validate-model') {
|
||||
validateModel = true;
|
||||
} else if (arg && !arg.startsWith('-')) {
|
||||
if (!webUrl) {
|
||||
webUrl = arg;
|
||||
@@ -340,6 +353,8 @@ function parseCliArgs(argv: string[]): CliArgs {
|
||||
taskQueue,
|
||||
...(workflowId && { workflowId }),
|
||||
pipelineTestingMode,
|
||||
authOnly,
|
||||
validateModel,
|
||||
...(configPath && { configPath }),
|
||||
...(customerOutputPath && { customerOutputPath }),
|
||||
...(resumeFromWorkspace && { resumeFromWorkspace }),
|
||||
@@ -588,6 +603,8 @@ function buildPipelineInput(
|
||||
...(args.customerOutputPath !== undefined && { customerOutputPath: args.customerOutputPath }),
|
||||
...(orchestration.agenticSast !== undefined && { agenticSast: orchestration.agenticSast }),
|
||||
...(orchestration.exploit !== undefined && { exploit: orchestration.exploit }),
|
||||
...(args.authOnly && { authOnly: true }),
|
||||
...(args.validateModel && { validateModel: true }),
|
||||
};
|
||||
}
|
||||
|
||||
@@ -642,6 +659,10 @@ async function waitForWorkflowResult(
|
||||
}
|
||||
} else if (result.status === 'cancelled') {
|
||||
console.log('\nScan cancelled before it finished.');
|
||||
} else if (result.authOnly) {
|
||||
console.log('\nAuthentication validated. No pentest was run (--validate-auth).');
|
||||
} else if (result.validateModel) {
|
||||
console.log('\nModel validated. No pentest was run (--validate-model).');
|
||||
} else {
|
||||
console.log('\nScan completed.');
|
||||
}
|
||||
@@ -754,6 +775,11 @@ async function run(): Promise<void> {
|
||||
// 1. Parse CLI args
|
||||
const args = parseCliArgs(process.argv.slice(2));
|
||||
|
||||
// One scan per worker process, so an auth-only or model-validation run is a process-wide fact.
|
||||
// The log writers read these to frame the log as a validation rather than a pentest.
|
||||
if (args.authOnly) process.env.SHANNON_AUTH_ONLY = '1';
|
||||
if (args.validateModel) process.env.SHANNON_VALIDATE_MODEL = '1';
|
||||
|
||||
// 2. Connect to Temporal server
|
||||
const address = process.env.TEMPORAL_ADDRESS || 'localhost:7233';
|
||||
console.log(`Connecting to Temporal at ${address}...`);
|
||||
|
||||
@@ -36,6 +36,8 @@ const ERROR_TYPE_TO_CODE: Record<string, ErrorCode> = {
|
||||
ReportSarifRenderError: ErrorCode.OUTPUT_VALIDATION_FAILED,
|
||||
IncompatibleWorkspaceError: ErrorCode.CONFIG_VALIDATION_FAILED,
|
||||
WorkspaceNotFoundError: ErrorCode.CONFIG_NOT_FOUND,
|
||||
OpenAiCyberAccessError: ErrorCode.PROVIDER_CYBER_ACCESS_REQUIRED,
|
||||
AnthropicCyberAccessError: ErrorCode.PROVIDER_CYBER_ACCESS_REQUIRED,
|
||||
};
|
||||
|
||||
export function classifyErrorCode(error: unknown): ErrorCode | undefined {
|
||||
@@ -64,6 +66,10 @@ const REMEDIATION_HINTS: Record<string, string> = {
|
||||
IncompatibleWorkspaceError: 'start a new scan with a different -w name.',
|
||||
WorkspaceNotFoundError: 'check the -w name against: shannon scans',
|
||||
PipelineFailedError: 're-run the same -w to retry from the last checkpoint.',
|
||||
OpenAiCyberAccessError:
|
||||
'Your OpenAI organization must be approved for cyber use. Apply for Daybreak access at https://openai.com/daybreak, then retry. Or use the gpt-5.4 model instead.',
|
||||
AnthropicCyberAccessError:
|
||||
'Your Anthropic organization must complete cyber verification. See https://support.claude.com/en/articles/14604842-real-time-cyber-safeguards-on-claude-opus-and-sonnet, then retry. Or use the claude-sonnet-4-6 model instead.',
|
||||
};
|
||||
|
||||
/**
|
||||
@@ -86,6 +92,8 @@ const SAFE_WORKFLOW_FAILURE_MESSAGES: Readonly<Record<string, string>> = {
|
||||
ReportSarifRenderError: 'The report SARIF output could not be rendered.',
|
||||
IncompatibleWorkspaceError: 'This workspace cannot be resumed.',
|
||||
WorkspaceNotFoundError: 'The requested workspace was not found.',
|
||||
OpenAiCyberAccessError: 'OpenAI declined the security workload behind its cyber-access program.',
|
||||
AnthropicCyberAccessError: 'Anthropic declined the security workload behind its cyber-access program.',
|
||||
};
|
||||
|
||||
const WORKFLOW_PHASE_SET = new Set<string>(WORKFLOW_PHASES);
|
||||
|
||||
@@ -100,6 +100,8 @@ const PRODUCTION_RETRY = {
|
||||
'InvalidTargetError',
|
||||
'AuthLoginFailedError',
|
||||
'PermanentError',
|
||||
'OpenAiCyberAccessError',
|
||||
'AnthropicCyberAccessError',
|
||||
],
|
||||
};
|
||||
|
||||
@@ -379,11 +381,15 @@ export async function pentestPipeline(input: PipelineInput): Promise<PipelineSta
|
||||
const { workflowId } = workflowInfo();
|
||||
const a = input.pipelineTestingMode ? testActs : acts;
|
||||
const exploit = input.exploit ?? true;
|
||||
const authOnly = input.authOnly ?? false;
|
||||
const validateModel = input.validateModel ?? false;
|
||||
const sessionId = input.sessionId || input.resumeFromWorkspace || workflowId;
|
||||
const stateContext: 'fresh' | 'resume' = input.resumeFromWorkspace ? 'resume' : 'fresh';
|
||||
|
||||
const state: PipelineState = {
|
||||
status: 'running',
|
||||
authOnly,
|
||||
validateModel,
|
||||
currentPhase: null,
|
||||
currentAgent: null,
|
||||
completedAgents: [],
|
||||
@@ -1287,7 +1293,7 @@ export async function pentestPipeline(input: PipelineInput): Promise<PipelineSta
|
||||
const durable = await deterministicReportActs.initializeDurableScanState(activityInput, exploit, stateContext);
|
||||
applyDurableSummary(durable);
|
||||
|
||||
if (input.resumeFromWorkspace) {
|
||||
if (!authOnly && input.resumeFromWorkspace) {
|
||||
// The new workflow id lands in session.json before anything that can reject the resume, so a
|
||||
// validation or checkpoint-restore failure still leaves the CLI an attempt to follow.
|
||||
await deterministicReportActs.registerResumeAttempt(activityInput, input.terminatedWorkflows ?? []);
|
||||
@@ -1337,7 +1343,30 @@ export async function pentestPipeline(input: PipelineInput): Promise<PipelineSta
|
||||
|
||||
state.currentPhase = 'preflight';
|
||||
state.currentAgent = null;
|
||||
await preflightActs.runPreflightValidation(activityInput);
|
||||
await runOperation('preflight', 'Preflight', () => preflightActs.runPreflightValidation(activityInput));
|
||||
if (!authOnly) {
|
||||
const startedAt = startOperation('cyber-access', 'Cyber access verification');
|
||||
try {
|
||||
const verification = await preflightActs.runCyberAccessVerification(activityInput);
|
||||
if (verification.gated) {
|
||||
completeOperation('cyber-access', 'Cyber access verification', startedAt);
|
||||
} else {
|
||||
delete state.operationalStages['cyber-access'];
|
||||
}
|
||||
} catch (error) {
|
||||
failOperation('cyber-access', 'Cyber access verification', startedAt);
|
||||
throw error;
|
||||
}
|
||||
}
|
||||
|
||||
if (validateModel) {
|
||||
state.status = 'completed';
|
||||
state.currentPhase = null;
|
||||
state.summary = computeSummary(state, usageAccountingComplete());
|
||||
await a.logWorkflowComplete(activityInput, toWorkflowSummary(state, 'completed'));
|
||||
return state;
|
||||
}
|
||||
|
||||
await preflightActs.syncPlaywrightStealthConfig(activityInput);
|
||||
|
||||
state.currentPhase = 'auth-validation';
|
||||
@@ -1346,6 +1375,21 @@ export async function pentestPipeline(input: PipelineInput): Promise<PipelineSta
|
||||
if (authMetrics !== null) state.agentMetrics['validate-authentication'] = authMetrics;
|
||||
state.currentAgent = null;
|
||||
|
||||
// Auth-only runs stop here; a null result means no authentication block, which is a misconfig.
|
||||
if (authOnly) {
|
||||
if (authMetrics === null) {
|
||||
throw ApplicationFailure.nonRetryable(
|
||||
'An auth-validation run needs an authentication block in the config. Add one, or drop --validate-auth.',
|
||||
'ConfigurationError',
|
||||
);
|
||||
}
|
||||
state.status = 'completed';
|
||||
state.currentPhase = null;
|
||||
state.summary = computeSummary(state, usageAccountingComplete());
|
||||
await a.logWorkflowComplete(activityInput, toWorkflowSummary(state, 'completed'));
|
||||
return state;
|
||||
}
|
||||
|
||||
await a.initDeliverableGit(activityInput);
|
||||
await a.syncCodePathDenyRules(activityInput);
|
||||
|
||||
|
||||
@@ -42,6 +42,7 @@ export enum ErrorCode {
|
||||
AUTH_LOGIN_FAILED = 'AUTH_LOGIN_FAILED',
|
||||
MODEL_NOT_FOUND = 'MODEL_NOT_FOUND',
|
||||
MODEL_CONFIG_INVALID = 'MODEL_CONFIG_INVALID',
|
||||
PROVIDER_CYBER_ACCESS_REQUIRED = 'PROVIDER_CYBER_ACCESS_REQUIRED',
|
||||
}
|
||||
|
||||
export type PentestErrorType = 'config' | 'network' | 'prompt' | 'filesystem' | 'validation' | 'unknown';
|
||||
|
||||
Binary file not shown.
|
Before Width: | Height: | Size: 62 KiB After Width: | Height: | Size: 8.5 KiB |
@@ -17,6 +17,7 @@
|
||||
|
||||
#let tester-override = sys.inputs.at("tester", default: "Shannon")
|
||||
#let brand = sys.inputs.at("brand", default: "Shannon | AI Pentester by Keygraph")
|
||||
#let keygraph-url = "https://keygraph.io"
|
||||
|
||||
// ---------- Palette ---------------------------------------------------------
|
||||
// Kept distinct so Critical / High are not confused under monitor gamma.
|
||||
@@ -128,6 +129,9 @@
|
||||
text(fill: white, weight: "bold", size: 7.5pt, tracking: 0.3pt, upper(label)),
|
||||
)
|
||||
|
||||
#let finding-anchor(id) = label("finding-" + id)
|
||||
#let finding-link(id) = link(finding-anchor(id), text(weight: "semibold")[#id])
|
||||
|
||||
#let categories-in-order = if mode == "exploits" {
|
||||
data.exploitedByType.map(entry => entry.category)
|
||||
} else {
|
||||
@@ -225,11 +229,13 @@
|
||||
#v(3.2cm)
|
||||
|
||||
#let brand-parts = brand.split("|").map(p => p.trim())
|
||||
// The logo is the Keygraph wordmark (about 5:1), so it is sized by height to
|
||||
// sit level with the two-line brand text beside it.
|
||||
#grid(
|
||||
columns: (auto, 1fr),
|
||||
column-gutter: 8pt,
|
||||
column-gutter: 12pt,
|
||||
align: (horizon, horizon),
|
||||
image("/assets/keygraph-logo.png", width: 1.6cm),
|
||||
link(keygraph-url, image("/assets/keygraph-logo.png", height: 0.5cm)),
|
||||
{
|
||||
set par(leading: 0.6em)
|
||||
text(size: 11pt, fill: ink, weight: "semibold", tracking: 1.2pt)[
|
||||
@@ -237,9 +243,9 @@
|
||||
]
|
||||
if brand-parts.len() > 1 {
|
||||
linebreak()
|
||||
text(size: 9pt, fill: muted, weight: "regular")[
|
||||
link(keygraph-url, text(size: 9pt, fill: muted, weight: "regular")[
|
||||
#brand-parts.slice(1).join(" ")
|
||||
]
|
||||
])
|
||||
}
|
||||
}
|
||||
)
|
||||
@@ -338,7 +344,7 @@
|
||||
#if "bullets" in entry and entry.bullets != none [
|
||||
#list(
|
||||
..entry.bullets.map(b => [
|
||||
#text(weight: "semibold")[#b.id] — #inline-code(b.description)
|
||||
#finding-link(b.id) — #inline-code(b.description)
|
||||
])
|
||||
)
|
||||
]
|
||||
@@ -474,7 +480,7 @@
|
||||
..(if show-confidence-col { (text(size: 9.5pt, weight: "semibold")[Confidence],) } else { () }),
|
||||
),
|
||||
..data.findings.map(f => (
|
||||
text(weight: "semibold")[#f.id],
|
||||
finding-link(f.id),
|
||||
inline-code(f.title),
|
||||
text(size: 9.5pt)[#f.category],
|
||||
sev-chip(f.severity),
|
||||
@@ -531,7 +537,7 @@
|
||||
|
||||
#let render-exploit(f) = {
|
||||
block(breakable: false)[
|
||||
#heading(level: 2)[#f.id: #inline-code(f.title)]
|
||||
#heading(level: 2)[#f.id: #inline-code(f.title)]#finding-anchor(f.id)
|
||||
#sev-chip(f.severity)
|
||||
#v(8pt)
|
||||
#render-finding-owasp(f)
|
||||
@@ -553,7 +559,7 @@
|
||||
|
||||
#let render-analysis(f) = {
|
||||
block(breakable: false)[
|
||||
#heading(level: 2)[#f.id: #inline-code(f.title)]
|
||||
#heading(level: 2)[#f.id: #inline-code(f.title)]#finding-anchor(f.id)
|
||||
#sev-chip(f.severity)
|
||||
#h(4pt)
|
||||
#confidence-chip(f.confidence)
|
||||
|
||||
+22
-14
@@ -35,7 +35,7 @@ export SHANNON_AI_BASE_URL=https://llm-gateway.example.com # optional: route thr
|
||||
|
||||
This path covers providers whose credential is a single API key. Providers that need more than that are not currently supported.
|
||||
|
||||
A model the catalogue does not yet carry, such as one released after Shannon's pinned Pi version, is reachable by describing it yourself. See [Custom model configuration](#custom-model-configuration).
|
||||
A model the catalogue does not carry is reachable by describing it yourself. See [Custom model configuration](#custom-model-configuration).
|
||||
|
||||
`npx @keygraph/shannon setup` exposes this as the **Other provider** option.
|
||||
|
||||
@@ -53,15 +53,23 @@ Review each vendor's guidance and complete the verification or enrollment they a
|
||||
|
||||
This applies to the Anthropic and OpenAI providers, including when either is reached through an LLM gateway. Bedrock serves Claude models and is subject to Anthropic's safeguards as well.
|
||||
|
||||
To confirm your model is ready before committing to a full scan, add `--validate-model` to `start`:
|
||||
|
||||
```bash
|
||||
npx @keygraph/shannon start -u https://your-app.com -r /path/to/repo --validate-model
|
||||
```
|
||||
|
||||
The run performs the preflight model checks only — credential and registry resolution for any provider, plus a single cyber-access verification against Anthropic and OpenAI that trips the cyber safeguard if your account is not approved — then stops. No pentest or report is produced, and it needs no config. A decline fails the run with the vendor's enrollment link.
|
||||
|
||||
## Suggested models
|
||||
|
||||
These are the models `npx @keygraph/shannon setup` offers, best-first. They are suggestions: the wizard also takes a typed model ID, and `SHANNON_AI_MODEL` accepts any model in the provider's catalogue.
|
||||
|
||||
| Provider | Suggested model IDs |
|
||||
| --- | --- |
|
||||
| `anthropic` | `claude-sonnet-4-6`, `claude-opus-4-8`, `claude-opus-4-7`, `claude-haiku-4-5-20251001` |
|
||||
| `openai` | `gpt-5.6-sol`, `gpt-5.5`, `gpt-5.4` |
|
||||
| `xai` | `grok-4.6`, `grok-4.5` |
|
||||
| `anthropic` | `claude-sonnet-5`, `claude-opus-5`, `claude-sonnet-4-6`, `claude-opus-4-8`, `claude-opus-4-7`, `claude-haiku-4-5-20251001` |
|
||||
| `openai` | `gpt-6-sol`, `gpt-5.6-sol`, `gpt-5.5`, `gpt-5.4` |
|
||||
| `xai` | `grok-4.7` |
|
||||
| `amazon-bedrock` | `us.anthropic.claude-sonnet-4-6`, `us.anthropic.claude-opus-4-8`, `us.anthropic.claude-opus-4-7` |
|
||||
|
||||
Bedrock IDs are region-prefixed and must be enabled in your account, so the ID that works for you may differ from the one listed here.
|
||||
@@ -81,14 +89,14 @@ OpenAI:
|
||||
|
||||
```bash
|
||||
export SHANNON_AI_API_KEY=sk-...
|
||||
export SHANNON_AI_MODEL=openai:gpt-5.6-sol
|
||||
export SHANNON_AI_MODEL=openai:gpt-6-sol
|
||||
```
|
||||
|
||||
xAI:
|
||||
|
||||
```bash
|
||||
export SHANNON_AI_API_KEY=xai-...
|
||||
export SHANNON_AI_MODEL=xai:grok-4.5
|
||||
export SHANNON_AI_MODEL=xai:grok-4.7
|
||||
```
|
||||
|
||||
Source-build mode reads the same variables from a `.env` file.
|
||||
@@ -130,7 +138,7 @@ OpenAI Responses LLM gateway:
|
||||
|
||||
```bash
|
||||
export SHANNON_AI_API_KEY=sk-...
|
||||
export SHANNON_AI_MODEL=openai:gpt-5.6-sol
|
||||
export SHANNON_AI_MODEL=openai:gpt-6-sol
|
||||
export SHANNON_AI_BASE_URL=https://llm-gateway.example.com/v1
|
||||
```
|
||||
|
||||
@@ -138,7 +146,7 @@ export SHANNON_AI_BASE_URL=https://llm-gateway.example.com/v1
|
||||
|
||||
## Custom model configuration
|
||||
|
||||
A model released after Shannon's pinned Pi version is not in the harness catalogue yet, so `SHANNON_AI_MODEL` alone cannot reach it. Rather than wait for a Shannon release, describe the model yourself and pass the file with `--models-config`:
|
||||
A custom model configuration is a Pi `models.json` file that describes a model the harness catalogue does not carry: one a router or gateway serves under its own ID, or a local server (see [Local and self-hosted models](#local-and-self-hosted-models)). You pass it with `--models-config`, and Shannon merges its definitions over the catalogue so `SHANNON_AI_MODEL` can then name the model like any other:
|
||||
|
||||
```bash
|
||||
npx @keygraph/shannon start -u https://example.com -r /path/to/repo --models-config ./models.json
|
||||
@@ -258,36 +266,36 @@ A ChatGPT Plus or Pro Codex subscription can run Shannon. Shannon reuses a login
|
||||
Before running a pentest, review the [cyber safeguards requirements](#cyber-safeguards-do-this-before-your-first-scan).
|
||||
|
||||
1. Install Pi by following the instructions at [pi.dev](https://pi.dev).
|
||||
2. Log in with your subscription using Pi's [subscription authentication guide](https://pi.dev/docs/latest/providers#subscriptions). This creates `~/.pi/agent/auth.json` with an `openai-codex` entry.
|
||||
2. Start Pi by running `pi` in your terminal, then run `/login`, choose **Sign in with an account**, then choose **OpenAI Codex (legacy)** and complete the browser sign-in. This creates `~/.pi/agent/auth.json` with an `openai-codex` entry.
|
||||
|
||||
3. Select a Codex model and enable Pi authentication:
|
||||
|
||||
```bash
|
||||
export SHANNON_USE_PI_AUTH=1
|
||||
export SHANNON_AI_MODEL=openai-codex:gpt-5.5
|
||||
export SHANNON_AI_MODEL=openai-codex:gpt-6-sol
|
||||
```
|
||||
|
||||
4. In npx mode, run `npx @keygraph/shannon start ...` from the same shell. In source-build mode, add the two variables to `.env` and run `./shannon start ...`.
|
||||
|
||||
Supported Codex models are `gpt-5.6-sol`, `gpt-5.5`, and `gpt-5.4`.
|
||||
Supported Codex models are `gpt-6-sol`, `gpt-5.6-sol`, `gpt-5.5`, and `gpt-5.4`.
|
||||
|
||||
## xAI (Grok subscription)
|
||||
|
||||
An xAI subscription can run Shannon. Shannon reuses a login created by Pi.
|
||||
|
||||
1. Install Pi by following the instructions at [pi.dev](https://pi.dev).
|
||||
2. Log in with your subscription using Pi's [subscription authentication guide](https://pi.dev/docs/latest/providers#subscriptions). This creates `~/.pi/agent/auth.json` with an `xai` entry.
|
||||
2. Start Pi by running `pi` in your terminal, then run `/login`, choose **Sign in with an account**, then choose **xAI** and complete the browser sign-in. This creates `~/.pi/agent/auth.json` with an `xai` entry.
|
||||
|
||||
3. Select an xAI model and enable Pi authentication:
|
||||
|
||||
```bash
|
||||
export SHANNON_USE_PI_AUTH=1
|
||||
export SHANNON_AI_MODEL=xai:grok-4.6
|
||||
export SHANNON_AI_MODEL=xai:grok-4.7
|
||||
```
|
||||
|
||||
4. In npx mode, run `npx @keygraph/shannon start ...` from the same shell. In source-build mode, add the two variables to `.env` and run `./shannon start ...`.
|
||||
|
||||
Suggested Grok models are `grok-4.6` and `grok-4.5`.
|
||||
The suggested Grok model is `grok-4.7`.
|
||||
|
||||
## Claude Code subscription
|
||||
|
||||
|
||||
@@ -179,3 +179,14 @@ login_flow:
|
||||
- "If prompted for 2FA, type $totp in <exact code field label or placeholder>"
|
||||
- "Click <exact button text>"
|
||||
```
|
||||
|
||||
### Validating Authentication Only
|
||||
|
||||
To confirm your login flow works before committing to a full scan, add `--validate-auth` to `start`:
|
||||
|
||||
```bash
|
||||
npx @keygraph/shannon start -u https://your-app.com -r /path/to/repo -c config.yaml --validate-auth
|
||||
```
|
||||
|
||||
The run performs preflight and the single real login, then stops. No pentest, reconciliation, or report
|
||||
is produced. It requires an `authentication` block in the config.
|
||||
@@ -122,6 +122,9 @@ npx @keygraph/shannon start -u https://example.com -r /path/to/repo -w q1-audit
|
||||
# Stream the log until the scan finishes, then exit on its outcome (useful in CI).
|
||||
npx @keygraph/shannon start -u https://example.com -r /path/to/repo --follow
|
||||
|
||||
# Validate the configured login only, then stop (no pentest or report).
|
||||
npx @keygraph/shannon start -u https://example.com -r /path/to/repo -c /path/to/my-config.yaml --validate-auth
|
||||
|
||||
# List running and completed scans.
|
||||
npx @keygraph/shannon scans
|
||||
```
|
||||
@@ -134,6 +137,7 @@ Source-build examples:
|
||||
./shannon start -u https://example.com -r /path/to/repo -o ./my-reports
|
||||
./shannon start -u https://example.com -r /path/to/repo -w q1-audit
|
||||
./shannon start -u https://example.com -r /path/to/repo --follow
|
||||
./shannon start -u https://example.com -r /path/to/repo -c /path/to/my-config.yaml --validate-auth
|
||||
./shannon scans
|
||||
|
||||
# Rebuild the worker image.
|
||||
|
||||
+54
-32
@@ -1,38 +1,56 @@
|
||||
# Keygraph Enterprise Platform
|
||||
# Keygraph Platform
|
||||
|
||||
Shannon Lite is now Shannon Open Source. Shannon Pro is now the Keygraph platform.
|
||||
|
||||
Shannon 3.0 is an open-source pentester. It reads your source, maps routes and data flows, runs real attacks against a live target, and writes PDF and SARIF reports. It runs locally, in CI, or air-gapped with your own model. Shannon Open Source is a complete pentester, not a trial edition.
|
||||
|
||||
Keygraph Enterprise runs an enterprise-hardened fork of Shannon continuously across hundreds of repositories and adds what a security team needs around it: audit-depth static analysis on a parsed code graph, business-logic testing, SCA and secrets scanning, one deduplicated record per vulnerability across scans and scanners, generated fixes, fix verification, and SSO, RBAC, and audit logs. It is for security teams that own vulnerability management across many engineering teams and need one place to triage, assign, fix, and verify.
|
||||
The Keygraph platform is Keygraph's commercial product. It runs an enterprise-hardened fork of Shannon continuously across hundreds of repositories and adds what a security team needs around it: black-box pentesting that needs no source code, audit-depth static analysis on a parsed code graph, SCA and secrets scanning, one deduplicated record per vulnerability across scans and scanners, fix pull requests, fix verification, two-way Jira sync, and SSO, RBAC, and audit logs. It is for security teams that own vulnerability management across many engineering teams and need one place to triage, assign, fix, and verify.
|
||||
|
||||
Both editions are BYOK. Keygraph never receives your source and never proxies model traffic, open source or commercial. Shannon Open Source runs from your machine or CI runner. Keygraph Enterprise deploys as a platform inside your cloud or data center, including fully air-gapped.
|
||||
The Keygraph platform comes in three editions: the Community Program, Pro, and Enterprise. Every module is included in all three. They differ in price, eligibility, hosting, and support. Current prices and the full feature table are at [keygraph.io/pricing](https://keygraph.io/pricing).
|
||||
|
||||
## Shannon Open Source vs. Keygraph Enterprise
|
||||
Every edition is BYOK: you bring your own model key and pay your provider directly. Shannon Open Source runs from your machine or CI runner, and Keygraph never receives your source or proxies your model traffic. The Community Program and Pro are cloud-hosted and managed by Keygraph. Enterprise deploys inside your cloud or data center, including fully air-gapped.
|
||||
|
||||
| | Shannon Open Source | Keygraph Enterprise |
|
||||
| --- | --- | --- |
|
||||
| Best for | Developers and teams running repository-level pentests locally or in CI | Security organizations running continuous AppSec across many teams and repositories |
|
||||
| Code analysis | Agent pass over architecture, entry points, and data flows to seed the pentest, sized to finish inside a CI run | Persistent code property graph plus a long-running analysis harness with interprocedural taint, sanitizer modeling, cross-repo context, exploit chains, and multi-pass review |
|
||||
| Pentesting | On-demand, source-aware white-box pentesting with optional authenticated testing, focused on injection, XSS, SSRF, broken authentication, and broken authorization, with proof by exploitation | Enterprise-hardened Shannon fork run continuously, with grey-box and black-box targets and business-logic invariant testing |
|
||||
| SCA and secrets | Not included | SCA with reachability and secrets scanning including history |
|
||||
| Findings | Per-run PDF, Markdown, JSON, and SARIF, with SARIF ingestion into GitHub code scanning | One record per vulnerability per repo across scans and scanners, plus ownership, SLAs, dashboards, and audit evidence |
|
||||
| Fixes and verification | Not included | Fix PRs with verification by re-analysis and exploit replay, with no full rescan required |
|
||||
| CI/CD and source control | GitHub Action and GitLab CI component for pull-request, release, and scheduled runs, with gates on `status: exploited` | GitHub, GitLab, Azure DevOps, and Bitbucket with organization-wide policy and centrally managed integrations |
|
||||
| Deployment and models | Runs locally or on a CI runner with BYOK to any Anthropic- or OpenAI-compatible endpoint or local model | Deployed in your AWS, GCP, Azure, or on-prem environment. Customer-hosted services and stored platform data remain inside your environment. Model requests go directly to the provider, private endpoint, gateway, or local model you configure. A local model supports fully disconnected deployments |
|
||||
| Governance, license, support | AGPL-3.0 and community support | SSO, SCIM, RBAC, and audit logs, plus a commercial license, enterprise support, and SOC 2 Type II |
|
||||
No source code? Use the [Blackbox Pentester](https://keygraph.io/agentic-blackbox-pentester).
|
||||
|
||||
## Compare editions
|
||||
|
||||
| | Shannon Open Source | Community Program | Pro | Enterprise |
|
||||
| --- | --- | --- | --- | --- |
|
||||
| Price | Open source under AGPL-3.0. You pay only your own model costs | $0 in cloud service fees while you qualify. You pay only your own AI provider usage | $50 per active developer per month, every module included, no add-ons | Custom |
|
||||
| Best for | Developers and teams running repository-level pentests locally or in CI | U.S.-based 501(c)(3) nonprofits, and seed or pre-Series-A startups with 20 or fewer active developers | Teams that want the full Keygraph platform, cloud-hosted and managed by Keygraph | Security organizations running continuous AppSec across many teams and repositories that need the platform inside their own environment |
|
||||
| Code analysis | Agent pass over architecture, entry points, and data flows to seed the pentest, plus optional multi-stage security code analysis, sized to finish inside a CI run | Same as Pro | Persistent code property graph plus a long-running analysis harness with interprocedural taint, sanitizer modeling, cross-repo context, exploit chains, and multi-pass review | Same as Pro |
|
||||
| White-box pentesting | On-demand, source-aware white-box pentesting with optional authenticated testing, focused on injection, XSS, SSRF, broken authentication, and broken authorization, with proof by exploitation | Same as Pro | Enterprise-hardened Shannon fork run continuously against white-box and grey-box targets, with proof by exploitation | Same as Pro |
|
||||
| Black-box pentesting (no source code) | Not included. Shannon Open Source needs the target's source code | Same as Pro | Agents attack the running application from the outside with no source access, with up to 4 login credentials for multi-role testing. Findings are validated with working exploits | Same as Pro |
|
||||
| SCA and secrets | Not included | Same as Pro | SCA with reachability and secrets scanning, including repository history | Same as Pro |
|
||||
| Findings management | Per-run PDF, Markdown, JSON, and SARIF 2.1.0, with SARIF upload to GitHub code scanning | Same as Pro | One record per vulnerability per repository across scans and scanners, with status history, auto-reopen, ownership, SLAs, dashboards, and audit evidence | Same as Pro |
|
||||
| Fix pull requests and retests | Fix manually from the report, then re-run the scan to verify | Same as Pro | Fix pull requests for a developer to review and apply, verified by re-analysis and exploit replay without a full rescan | Same as Pro |
|
||||
| Ticketing | Not included | Same as Pro | Two-way Jira sync | Same as Pro |
|
||||
| CI/CD and source control | GitHub Action and GitLab CI component for pull-request, release, and scheduled runs, with gates on `status: exploited` | Same as Pro | GitHub, GitLab, Azure DevOps, and Bitbucket with organization-wide policy and centrally managed integrations | Same as Pro |
|
||||
| Deployment and models | Runs locally or on a CI runner with BYOK to any Anthropic- or OpenAI-compatible endpoint or local model | Same as Pro | Cloud-hosted and managed by Keygraph, in a US or EU region. BYOK: inference runs under your own model key and provider account | Deployed in your AWS, GCP, Azure, or on-prem environment, including fully air-gapped. Customer-hosted services and stored platform data remain inside your environment. Model requests go directly to the provider, private endpoint, gateway, or local model you configure. A local model supports fully disconnected deployments |
|
||||
| Governance and compliance | Not included | Same as Pro | SSO via SAML 2.0 or OIDC with SCIM, RBAC, audit logs, SOC 2 Type II report under NDA, and a standard DPA | Same as Pro, plus a custom SLA and a security patch SLA |
|
||||
| License and support | AGPL-3.0, community support on Discord | Community support on Discord, with no SLA or uptime commitment | Commercial license, with email and Slack support | Commercial license, with a support engineer, 24x7 Severity 1 support, quarterly business reviews, and white-glove onboarding |
|
||||
|
||||
### What Pro includes
|
||||
|
||||
Pro is the full Keygraph platform, cloud-hosted and managed by Keygraph, at $50 per active developer per month. Every module is included, with unlimited repositories and scans under BYOK. That covers white-box and black-box pentesting, agentic SAST, SCA, secrets scanning, findings management with status history, retests (fix verification by re-analysis and exploit replay), fix pull requests, two-way Jira sync, and SSO, RBAC, and audit logs.
|
||||
|
||||
None of these require Enterprise. Enterprise adds self-hosted or fully air-gapped deployment, a dedicated engineer and white-glove onboarding, and a custom SLA and security patch SLA.
|
||||
|
||||
The Community Program is the same managed Pro platform, with every module included, at $0 in cloud service fees for organizations that qualify. It is cloud-hosted only. See the [Community Program](https://keygraph.io/community-program) for eligibility and how to apply.
|
||||
|
||||
## How it fits your pipeline
|
||||
|
||||
1. Scans run on pull requests, releases, and a schedule against repositories in GitHub, GitLab, Azure DevOps, or Bitbucket.
|
||||
2. Pipelines gate on exploited severity. A code-analysis hypothesis never fails a build.
|
||||
3. Findings from every scanner and every run land as one record per vulnerability per repository, with an owner and an SLA. The same finding across ten runs is one record, not ten alerts.
|
||||
4. From a finding, Keygraph opens a fix PR into your normal review flow.
|
||||
4. From a finding, Keygraph opens a fix PR into your normal review flow, for a developer to review and apply.
|
||||
5. Verification confirms the fix against the changed code and the original exploit. No full rescan is required.
|
||||
|
||||
## What is different technically
|
||||
|
||||
### Static analysis on a code property graph
|
||||
|
||||
Shannon Open Source's code analysis is sized to finish inside a CI run: agents read the repository, map the attack surface, and hand candidates to the pentester. Enterprise is built for depth instead. It first parses each repository into a persistent code property graph, then runs an analysis harness derived from one built for long-running vulnerability audits, heavily adapted to query the graph rather than read files. The harness decomposes the application into risk, taint-flow, framework, and specialist tasks and supports longer-running audit workflows beyond typical CI job windows.
|
||||
Shannon Open Source's code analysis is sized to finish inside a CI run: agents read the repository, map the attack surface, and hand candidates to the pentester. The Keygraph platform is built for depth instead. It first parses each repository into a persistent code property graph, then runs an analysis harness derived from one built for long-running vulnerability audits, heavily adapted to query the graph rather than read files. The harness decomposes the application into risk, taint-flow, framework, and specialist tasks and supports longer-running audit workflows beyond typical CI job windows.
|
||||
|
||||
On the graph, it performs:
|
||||
|
||||
@@ -44,58 +62,62 @@ On the graph, it performs:
|
||||
|
||||
### Business-logic invariants
|
||||
|
||||
Shannon Open Source focuses on injection, XSS, SSRF, and broken authentication and authorization. Enterprise adds testing for the bugs that do not fit a vulnerability class: it derives invariants the application is supposed to hold (tenant isolation, workflow ordering, approval limits, balance conservation, state transitions) and tests them against the running application. This is where application-specific vulnerabilities live and where pattern-based SAST often provides little or no signal.
|
||||
Shannon Open Source focuses on injection, XSS, SSRF, and broken authentication and authorization. The Keygraph platform adds testing for the bugs that do not fit a vulnerability class: it derives invariants the application is supposed to hold (tenant isolation, workflow ordering, approval limits, balance conservation, state transitions) and tests them against the running application. This is where application-specific vulnerabilities live and where pattern-based SAST often provides little or no signal.
|
||||
|
||||
### Proof by exploitation
|
||||
|
||||
The pentesting engine is a hardened fork of Shannon with the same rule: a pentest finding requires a working exploit. No exploit, no finding. Enterprise stores the exploit and replays it later to verify the fix.
|
||||
The pentesting engine is a hardened fork of Shannon with the same rule: a pentest finding requires a working exploit. No exploit, no finding. The platform stores the exploit and replays it later to verify the fix.
|
||||
|
||||
The Blackbox Pentester applies the same rule without source access. It attacks the running application from the outside with a real browser and terminal, and it can take up to 4 login credentials (Google OAuth, GitHub, or custom auth) to test privilege escalation and IDOR across roles.
|
||||
|
||||
SCA prioritizes vulnerable dependencies that application code actually reaches. Secrets scanning covers current source and repository history.
|
||||
|
||||
<p align="center">
|
||||
<img src="../assets/keygraph-platform/agentic-sast-results.png" alt="Keygraph Enterprise findings grouped into business-logic issues, point issues, and secrets" width="100%">
|
||||
<img src="../assets/keygraph-platform/agentic-sast-results.png" alt="Keygraph platform findings grouped into business-logic issues, point issues, and secrets" width="100%">
|
||||
</p>
|
||||
|
||||
## Findings
|
||||
|
||||
Shannon Open Source hands you a report per scan. Enterprise dedupes across runs and across scanners, deterministically and semantically, into one record per vulnerability per repository. Each record carries evidence, source location, severity, scan history, status, owner, resolution, and last-verified state.
|
||||
Shannon Open Source hands you a report per scan. The Keygraph platform dedupes across runs and across scanners, deterministically and semantically, into one record per vulnerability per repository. Each record carries evidence, source location, severity, scan history, status, owner, resolution, and last-verified state. This is included in the Community Program, Pro, and Enterprise.
|
||||
|
||||
Workflows cover assignment, triage, false-positive and risk-acceptance decisions, and SLA policies with escalation and aging. Dashboards report open risk, coverage, new versus resolved, SLA compliance, and MTTR, exportable as evidence for customers and auditors.
|
||||
Workflows cover assignment, triage, false-positive and risk-acceptance decisions, and SLA policies with escalation and aging. A resolved finding that reappears in a later scan reopens automatically. Dashboards report open risk, coverage, new versus resolved, SLA compliance, and MTTR, exportable as evidence for customers and auditors. Findings sync both ways with Jira.
|
||||
|
||||
Findings still require human review. Enterprise's extra review passes reduce weakly supported findings, but they do not eliminate them.
|
||||
Findings still require human review. The platform's extra review passes reduce weakly supported findings, but they do not eliminate them.
|
||||
|
||||
<p align="center">
|
||||
<img src="../assets/keygraph-platform/canonical-findings.png" alt="Keygraph Enterprise findings inventory with severity, status, source, and verification filters" width="100%">
|
||||
<img src="../assets/keygraph-platform/canonical-findings.png" alt="Keygraph platform findings inventory with severity, status, source, and verification filters" width="100%">
|
||||
</p>
|
||||
|
||||
### Fix and verify
|
||||
|
||||
From a finding, Keygraph generates a patch scoped to that finding and opens a pull request. It never commits to a protected branch.
|
||||
From a finding, Keygraph generates a patch scoped to that finding and opens a pull request for a developer to review and apply. It never commits to a protected branch.
|
||||
|
||||
<p align="center">
|
||||
<img src="../assets/keygraph-platform/automated-remediation.png" alt="Keygraph Enterprise remediation workflow for generating a fix and opening a pull request" width="100%">
|
||||
<img src="../assets/keygraph-platform/automated-remediation.png" alt="Keygraph platform remediation workflow for generating a fix and opening a pull request" width="100%">
|
||||
</p>
|
||||
|
||||
Verification re-analyzes the changed code and, for pentest findings, replays the original exploit against the patched target. The verdict comes from deterministic checks plus a review pass, without rerunning the full scan.
|
||||
|
||||
<p align="center">
|
||||
<img src="../assets/keygraph-platform/targeted-verification.png" alt="Keygraph Enterprise finding-verification workflow" width="100%">
|
||||
<img src="../assets/keygraph-platform/targeted-verification.png" alt="Keygraph platform finding-verification workflow" width="100%">
|
||||
</p>
|
||||
|
||||
## Deployment and access control
|
||||
|
||||
Keygraph Enterprise deploys entirely inside your AWS, GCP, Azure, or on-prem environment, including networks with no internet egress. There is no Keygraph-operated control plane. Customer-hosted services and stored platform data remain inside your environment for the life of the deployment.
|
||||
The Community Program and Pro are cloud-hosted and managed by Keygraph on AWS, with regional isolation in a US or EU region. Scans run in isolated, single-use containers that clone your repository and are destroyed when the scan ends. The full source tree is never persisted. Only the code fragments needed to render findings, deduplicate results, and propose fixes are kept, and all code-derived data is encrypted at rest. See [Code security posture](https://keygraph.io/code-security-posture) for the full controls.
|
||||
|
||||
Model access is BYOK and BYOM. Route workloads to Anthropic, OpenAI, xAI, or Bedrock, a private cloud endpoint, your own gateway such as LiteLLM with your routing and policy applied, or local models on vLLM or Ollama. Model requests go directly to the endpoint you configure. Keygraph never receives or proxies them. A local model supports a fully disconnected deployment.
|
||||
Enterprise deploys entirely inside your AWS, GCP, Azure, or on-prem environment, including networks with no internet egress. There is no Keygraph-operated control plane. Customer-hosted services and stored platform data remain inside your environment for the life of the deployment.
|
||||
|
||||
Access control: SAML/OIDC SSO, SCIM, roles with repository-scoped visibility (RBAC, plus attribute and relationship rules where needed), full audit log, scoped API keys.
|
||||
Model access is BYOK and BYOM in every edition. Route workloads to Anthropic, OpenAI, xAI, or Bedrock, a private cloud endpoint, your own gateway such as LiteLLM with your routing and policy applied, or local models on vLLM or Ollama. In Enterprise deployments, model requests leave your environment only to reach the endpoint you configure, and a local model supports a fully disconnected deployment.
|
||||
|
||||
Access control, in the Community Program, Pro, and Enterprise: SAML/OIDC SSO, SCIM, roles with repository-scoped visibility (RBAC, plus attribute and relationship rules where needed), full audit log, scoped API keys.
|
||||
|
||||
<p align="center">
|
||||
<img src="../assets/keygraph-platform/enterprise-access-control.png" alt="Keygraph Enterprise roles and repository visibility controls" width="100%">
|
||||
<img src="../assets/keygraph-platform/enterprise-access-control.png" alt="Keygraph platform roles and repository visibility controls" width="100%">
|
||||
</p>
|
||||
|
||||
Keygraph maintains a SOC 2 Type II audit. The report is available to customers under NDA.
|
||||
|
||||
## Talk to Keygraph
|
||||
|
||||
Visit [keygraph.io](https://keygraph.io), book a [demo](https://cal.com/team/keygraph/shannon-pro), or email [shannon@keygraph.io](mailto:shannon@keygraph.io).
|
||||
Visit [keygraph.io](https://keygraph.io), see [pricing](https://keygraph.io/pricing), apply to the [Community Program](https://keygraph.io/community-program), book a [demo](https://cal.com/team/keygraph/keygraph-technical-demo), or email [shannon@keygraph.io](mailto:shannon@keygraph.io).
|
||||
+118
-56
@@ -7,7 +7,7 @@
|
||||
# File: README.md
|
||||
|
||||
> [!NOTE]
|
||||
> **[Shannon 3.0 is live](https://github.com/KeygraphHQ/shannon/discussions/439):** deeper security code analysis, more thoroughly vetted findings, a rebuilt CLI, native CI/CD, professional PDF reports, and SARIF.
|
||||
> **[Cyber verification for Anthropic and OpenAI models](https://github.com/KeygraphHQ/shannon/discussions/483):** complete your provider's cyber verification program to prevent model failures and refusals during cyber workloads.
|
||||
|
||||
<div align="center">
|
||||
|
||||
@@ -84,7 +84,7 @@ Shannon is an autonomous AI pentester developed by [Keygraph](https://keygraph.i
|
||||
|
||||
Shannon analyzes your web application's source code to identify potential attack vectors, then uses browser automation and command-line tools to execute real exploits against the running application and its APIs. Only vulnerabilities with a working proof-of-concept are included in the final report.
|
||||
|
||||
Shannon is the agent. This repository is Shannon Open Source, the standalone pentester you run yourself. The same Shannon also powers the [Keygraph platform](https://keygraph.io), Keygraph's commercial pentesting product. See [Editions](#editions) for how the two compare.
|
||||
Shannon is the agent. This repository is Shannon Open Source, the standalone pentester you run yourself. The same Shannon also powers the [Keygraph platform](https://keygraph.io), Keygraph's commercial pentesting product. See [Editions](#editions) for how Shannon Open Source compares with the platform's Community Program, Pro, and Enterprise editions.
|
||||
|
||||
<a id="why-shannon-exists"></a>
|
||||
<details>
|
||||
@@ -139,8 +139,8 @@ These reports are from Shannon Open Source scans of Photoview 2.4.0, one of the
|
||||
|
||||
- **Docker**: required for the worker container.
|
||||
- **Node.js 18+**: required for the recommended `npx` workflow.
|
||||
- **AI provider credentials**: Shannon runs on Anthropic, OpenAI, xAI, AWS Bedrock, and [any other provider](docs/ai-providers.md#any-other-provider) in the harness catalogue — each of which you can point at a proxy or LLM gateway through a [custom base URL](docs/ai-providers.md#custom-base-url), and a model the catalogue does not yet carry can be described with a [custom model configuration](docs/ai-providers.md#custom-model-configuration). You bring your own key, and Keygraph never proxies your model traffic. Shannon is provider-agnostic. See [AI providers](docs/ai-providers.md#suggested-models) for suggested model IDs.
|
||||
- **Cyber safeguards cleared with your provider**: Anthropic and OpenAI apply real-time safeguards to cyber-security workloads, which can interrupt a scan mid-run. Complete their guidance for legitimate security testers before your first run - see [AI providers](docs/ai-providers.md#cyber-safeguards-do-this-before-your-first-scan).
|
||||
- **AI provider credentials**: Shannon runs on Anthropic, OpenAI, xAI, AWS Bedrock, and [any other provider](docs/ai-providers.md#any-other-provider) in the harness catalogue — each of which you can point at a proxy or LLM gateway through a [custom base URL](docs/ai-providers.md#custom-base-url), and a model the catalogue does not carry can be described with a [custom model configuration](docs/ai-providers.md#custom-model-configuration). You bring your own key, and Keygraph never proxies your model traffic. Shannon is provider-agnostic. See [AI providers](docs/ai-providers.md#suggested-models) for suggested model IDs.
|
||||
- **Cyber safeguards cleared with your provider**: Anthropic and OpenAI apply real-time safeguards to cyber-security workloads, which can interrupt a scan mid-run. Complete their guidance for legitimate security testers before your first run - see [AI providers](docs/ai-providers.md#cyber-safeguards-do-this-before-your-first-scan) and the [cyber verification announcement](https://github.com/KeygraphHQ/shannon/discussions/483).
|
||||
|
||||
|
||||
|
||||
@@ -243,11 +243,28 @@ See the [Shannon GitHub Action documentation](https://github.com/KeygraphHQ/shan
|
||||
|
||||
## Editions
|
||||
|
||||
**Shannon Open Source** is a complete autonomous pentester, especially well suited to individual developers and small teams running focused security tests locally or in CI/CD.
|
||||
Shannon Lite is now Shannon Open Source. Shannon Pro is now the Keygraph platform.
|
||||
|
||||
**Keygraph Enterprise Platform** is for organizations that need a shared platform for continuous agentic pentesting/AppSec across many teams, repositories, and environments. It centralizes deeper analysis, vulnerability management, remediation, verification, governance, and reporting so teams do not have to assemble and maintain those workflows themselves.
|
||||
| Edition | What it is | Price |
|
||||
| --- | --- | --- |
|
||||
| **Shannon Open Source** (Shannon OSS) | This repository. A complete autonomous pentester you run yourself, locally or in CI/CD, against an application whose source code you have. Well suited to individual developers and small teams. | Open source under AGPL-3.0. You pay only your own model costs. |
|
||||
| **Community Program** | The full Keygraph platform, cloud-hosted, for U.S.-based 501(c)(3) nonprofits and for seed or pre-Series-A startups with 20 or fewer active developers. | $0 in cloud service fees while you qualify. |
|
||||
| **Pro** | The full Keygraph platform, cloud-hosted and managed by Keygraph, with every module included. | $50 per active developer per month. |
|
||||
| **Enterprise** | Everything in Pro, self-hosted in your own environment or fully air-gapped. | Custom. |
|
||||
|
||||
[Learn about the Keygraph Enterprise Platform and compare editions →](docs/keygraph-platform.md)
|
||||
See [keygraph.io/pricing](https://keygraph.io/pricing) for current prices and the full feature table.
|
||||
|
||||
The **Keygraph platform** runs an enterprise-hardened fork of Shannon. The Community Program, Pro, and Enterprise all add:
|
||||
|
||||
- **Black-box pentesting**: tests the running application from the outside, with no source code needed.
|
||||
- **Dependency (SCA) and secret checks**: SCA with reachability, and secrets scanning that includes repository history.
|
||||
- **Findings management**: one record per vulnerability per repository across scans and scanners, with status history, owners, SLAs, and dashboards.
|
||||
- **Fix pull requests and retests**: reviewable fix pull requests, with each fix verified by re-analysis and exploit replay, without a full rescan.
|
||||
- **Jira and access control**: two-way Jira sync, SSO, RBAC, and audit logs.
|
||||
|
||||
No source code? Use the [Blackbox Pentester](https://keygraph.io/agentic-blackbox-pentester).
|
||||
|
||||
[Learn about the Keygraph platform and compare editions →](docs/keygraph-platform.md)
|
||||
|
||||
## Architecture
|
||||
|
||||
@@ -299,7 +316,7 @@ Use these guides for operational detail:
|
||||
| [Workspaces and resuming](docs/workspaces.md) | Naming workspaces, resuming interrupted scans, and workspace storage. |
|
||||
| [Safety and limitations](docs/safety.md) | Authorized-use requirements, non-production guidance, mutative effects, cost, and model caveats. |
|
||||
| [Coverage and roadmap](docs/coverage-roadmap.md) | Current vulnerability coverage and planned work. |
|
||||
| [Keygraph Enterprise Platform](docs/keygraph-platform.md) | Exhaustive agentic SAST, continuous pentesting, full-lifecycle finding management, remediation, targeted verification, enterprise governance, and on-premises deployment. |
|
||||
| [Keygraph platform](docs/keygraph-platform.md) | Shannon Open Source compared with the Community Program, Pro, and Enterprise editions: black-box pentesting, SCA and secrets, findings management, fix pull requests, retests, Jira, and deployment. |
|
||||
|
||||
|
||||
|
||||
@@ -312,7 +329,7 @@ You are responsible for using Shannon legally and ethically. Do not point Shanno
|
||||
|
||||
Important limitations:
|
||||
|
||||
- Shannon Open Source is tuned for fast, code-informed pentesting in everyday development and CI/CD. Exhaustive agentic SAST, broader scanner coverage, centralized governance, and full-lifecycle vulnerability management are delivered through the Keygraph Enterprise Platform.
|
||||
- Shannon Open Source is tuned for fast, code-informed pentesting in everyday development and CI/CD. Exhaustive agentic SAST, broader scanner coverage, centralized governance, and full-lifecycle vulnerability management are delivered through the Keygraph platform, in its Community Program, Pro, and Enterprise editions.
|
||||
- Findings still require human review. LLM-generated reports can contain weakly supported or incorrect details.
|
||||
- Anthropic, OpenAI, xAI, and AWS Bedrock are built-in providers, and any other provider in the harness catalogue works too — each reachable through a custom base URL that points it at a proxy or LLM gateway. Model capability varies, and a model that does not follow Shannon's instructions or tool-use constraints reliably will produce weaker results.
|
||||
- A full run can take roughly 1 to 1.5 hours and may incur LLM API costs depending on model pricing and application complexity.
|
||||
@@ -383,7 +400,7 @@ Yes. Shannon emits SARIF 2.1.0, the OASIS standard format for static analysis re
|
||||
|
||||
### Which AI providers does Shannon support?
|
||||
|
||||
Anthropic, OpenAI, xAI, and AWS Bedrock are built in and configured directly by provider ID. Beyond those, Shannon runs on any provider in the Pi harness catalogue, named the same `<provider>:<model-id>` way. Any provider can be pointed at a proxy or LLM gateway through a custom base URL, which overrides only the endpoint and keeps that provider's API dialect. A model the catalogue does not yet carry, such as one released after Shannon's pinned harness version, runs without waiting for a Shannon release. Describe it in a [custom model configuration](docs/ai-providers.md#custom-model-configuration) file and pass it with `--models-config`. Shannon uses a single unified model setting throughout a pentest.
|
||||
Anthropic, OpenAI, xAI, and AWS Bedrock are built in and configured directly by provider ID. Beyond those, Shannon runs on any provider in the Pi harness catalogue, named the same `<provider>:<model-id>` way. Any provider can be pointed at a proxy or LLM gateway through a custom base URL, which overrides only the endpoint and keeps that provider's API dialect. A model the catalogue does not carry, such as one a router or gateway serves under its own ID, or a self-hosted model, is described in a [custom model configuration](docs/ai-providers.md#custom-model-configuration) file and passed with `--models-config`. Shannon uses a single unified model setting throughout a pentest.
|
||||
|
||||
### Can I run Shannon on a local or self-hosted model?
|
||||
|
||||
@@ -523,6 +540,9 @@ npx @keygraph/shannon start -u https://example.com -r /path/to/repo -w q1-audit
|
||||
# Stream the log until the scan finishes, then exit on its outcome (useful in CI).
|
||||
npx @keygraph/shannon start -u https://example.com -r /path/to/repo --follow
|
||||
|
||||
# Validate the configured login only, then stop (no pentest or report).
|
||||
npx @keygraph/shannon start -u https://example.com -r /path/to/repo -c /path/to/my-config.yaml --validate-auth
|
||||
|
||||
# List running and completed scans.
|
||||
npx @keygraph/shannon scans
|
||||
```
|
||||
@@ -535,6 +555,7 @@ Source-build examples:
|
||||
./shannon start -u https://example.com -r /path/to/repo -o ./my-reports
|
||||
./shannon start -u https://example.com -r /path/to/repo -w q1-audit
|
||||
./shannon start -u https://example.com -r /path/to/repo --follow
|
||||
./shannon start -u https://example.com -r /path/to/repo -c /path/to/my-config.yaml --validate-auth
|
||||
./shannon scans
|
||||
|
||||
# Rebuild the worker image.
|
||||
@@ -751,6 +772,17 @@ login_flow:
|
||||
- "Click <exact button text>"
|
||||
```
|
||||
|
||||
### Validating Authentication Only
|
||||
|
||||
To confirm your login flow works before committing to a full scan, add `--validate-auth` to `start`:
|
||||
|
||||
```bash
|
||||
npx @keygraph/shannon start -u https://your-app.com -r /path/to/repo -c config.yaml --validate-auth
|
||||
```
|
||||
|
||||
The run performs preflight and the single real login, then stops. No pentest, reconciliation, or report
|
||||
is produced. It requires an `authentication` block in the config.
|
||||
|
||||
---
|
||||
|
||||
# File: docs/ai-providers.md
|
||||
@@ -792,7 +824,7 @@ export SHANNON_AI_BASE_URL=https://llm-gateway.example.com # optional: route thr
|
||||
|
||||
This path covers providers whose credential is a single API key. Providers that need more than that are not currently supported.
|
||||
|
||||
A model the catalogue does not yet carry, such as one released after Shannon's pinned Pi version, is reachable by describing it yourself. See [Custom model configuration](#custom-model-configuration).
|
||||
A model the catalogue does not carry is reachable by describing it yourself. See [Custom model configuration](#custom-model-configuration).
|
||||
|
||||
`npx @keygraph/shannon setup` exposes this as the **Other provider** option.
|
||||
|
||||
@@ -810,15 +842,23 @@ Review each vendor's guidance and complete the verification or enrollment they a
|
||||
|
||||
This applies to the Anthropic and OpenAI providers, including when either is reached through an LLM gateway. Bedrock serves Claude models and is subject to Anthropic's safeguards as well.
|
||||
|
||||
To confirm your model is ready before committing to a full scan, add `--validate-model` to `start`:
|
||||
|
||||
```bash
|
||||
npx @keygraph/shannon start -u https://your-app.com -r /path/to/repo --validate-model
|
||||
```
|
||||
|
||||
The run performs the preflight model checks only — credential and registry resolution for any provider, plus a single cyber-access verification against Anthropic and OpenAI that trips the cyber safeguard if your account is not approved — then stops. No pentest or report is produced, and it needs no config. A decline fails the run with the vendor's enrollment link.
|
||||
|
||||
## Suggested models
|
||||
|
||||
These are the models `npx @keygraph/shannon setup` offers, best-first. They are suggestions: the wizard also takes a typed model ID, and `SHANNON_AI_MODEL` accepts any model in the provider's catalogue.
|
||||
|
||||
| Provider | Suggested model IDs |
|
||||
| --- | --- |
|
||||
| `anthropic` | `claude-sonnet-4-6`, `claude-opus-4-8`, `claude-opus-4-7`, `claude-haiku-4-5-20251001` |
|
||||
| `openai` | `gpt-5.6-sol`, `gpt-5.5`, `gpt-5.4` |
|
||||
| `xai` | `grok-4.6`, `grok-4.5` |
|
||||
| `anthropic` | `claude-sonnet-5`, `claude-opus-5`, `claude-sonnet-4-6`, `claude-opus-4-8`, `claude-opus-4-7`, `claude-haiku-4-5-20251001` |
|
||||
| `openai` | `gpt-6-sol`, `gpt-5.6-sol`, `gpt-5.5`, `gpt-5.4` |
|
||||
| `xai` | `grok-4.7` |
|
||||
| `amazon-bedrock` | `us.anthropic.claude-sonnet-4-6`, `us.anthropic.claude-opus-4-8`, `us.anthropic.claude-opus-4-7` |
|
||||
|
||||
Bedrock IDs are region-prefixed and must be enabled in your account, so the ID that works for you may differ from the one listed here.
|
||||
@@ -838,14 +878,14 @@ OpenAI:
|
||||
|
||||
```bash
|
||||
export SHANNON_AI_API_KEY=sk-...
|
||||
export SHANNON_AI_MODEL=openai:gpt-5.6-sol
|
||||
export SHANNON_AI_MODEL=openai:gpt-6-sol
|
||||
```
|
||||
|
||||
xAI:
|
||||
|
||||
```bash
|
||||
export SHANNON_AI_API_KEY=xai-...
|
||||
export SHANNON_AI_MODEL=xai:grok-4.5
|
||||
export SHANNON_AI_MODEL=xai:grok-4.7
|
||||
```
|
||||
|
||||
Source-build mode reads the same variables from a `.env` file.
|
||||
@@ -887,7 +927,7 @@ OpenAI Responses LLM gateway:
|
||||
|
||||
```bash
|
||||
export SHANNON_AI_API_KEY=sk-...
|
||||
export SHANNON_AI_MODEL=openai:gpt-5.6-sol
|
||||
export SHANNON_AI_MODEL=openai:gpt-6-sol
|
||||
export SHANNON_AI_BASE_URL=https://llm-gateway.example.com/v1
|
||||
```
|
||||
|
||||
@@ -895,7 +935,7 @@ export SHANNON_AI_BASE_URL=https://llm-gateway.example.com/v1
|
||||
|
||||
## Custom model configuration
|
||||
|
||||
A model released after Shannon's pinned Pi version is not in the harness catalogue yet, so `SHANNON_AI_MODEL` alone cannot reach it. Rather than wait for a Shannon release, describe the model yourself and pass the file with `--models-config`:
|
||||
A custom model configuration is a Pi `models.json` file that describes a model the harness catalogue does not carry: one a router or gateway serves under its own ID, or a local server (see [Local and self-hosted models](#local-and-self-hosted-models)). You pass it with `--models-config`, and Shannon merges its definitions over the catalogue so `SHANNON_AI_MODEL` can then name the model like any other:
|
||||
|
||||
```bash
|
||||
npx @keygraph/shannon start -u https://example.com -r /path/to/repo --models-config ./models.json
|
||||
@@ -1015,36 +1055,36 @@ A ChatGPT Plus or Pro Codex subscription can run Shannon. Shannon reuses a login
|
||||
Before running a pentest, review the [cyber safeguards requirements](#cyber-safeguards-do-this-before-your-first-scan).
|
||||
|
||||
1. Install Pi by following the instructions at [pi.dev](https://pi.dev).
|
||||
2. Log in with your subscription using Pi's [subscription authentication guide](https://pi.dev/docs/latest/providers#subscriptions). This creates `~/.pi/agent/auth.json` with an `openai-codex` entry.
|
||||
2. Start Pi by running `pi` in your terminal, then run `/login`, choose **Sign in with an account**, then choose **OpenAI Codex (legacy)** and complete the browser sign-in. This creates `~/.pi/agent/auth.json` with an `openai-codex` entry.
|
||||
|
||||
3. Select a Codex model and enable Pi authentication:
|
||||
|
||||
```bash
|
||||
export SHANNON_USE_PI_AUTH=1
|
||||
export SHANNON_AI_MODEL=openai-codex:gpt-5.5
|
||||
export SHANNON_AI_MODEL=openai-codex:gpt-6-sol
|
||||
```
|
||||
|
||||
4. In npx mode, run `npx @keygraph/shannon start ...` from the same shell. In source-build mode, add the two variables to `.env` and run `./shannon start ...`.
|
||||
|
||||
Supported Codex models are `gpt-5.6-sol`, `gpt-5.5`, and `gpt-5.4`.
|
||||
Supported Codex models are `gpt-6-sol`, `gpt-5.6-sol`, `gpt-5.5`, and `gpt-5.4`.
|
||||
|
||||
## xAI (Grok subscription)
|
||||
|
||||
An xAI subscription can run Shannon. Shannon reuses a login created by Pi.
|
||||
|
||||
1. Install Pi by following the instructions at [pi.dev](https://pi.dev).
|
||||
2. Log in with your subscription using Pi's [subscription authentication guide](https://pi.dev/docs/latest/providers#subscriptions). This creates `~/.pi/agent/auth.json` with an `xai` entry.
|
||||
2. Start Pi by running `pi` in your terminal, then run `/login`, choose **Sign in with an account**, then choose **xAI** and complete the browser sign-in. This creates `~/.pi/agent/auth.json` with an `xai` entry.
|
||||
|
||||
3. Select an xAI model and enable Pi authentication:
|
||||
|
||||
```bash
|
||||
export SHANNON_USE_PI_AUTH=1
|
||||
export SHANNON_AI_MODEL=xai:grok-4.6
|
||||
export SHANNON_AI_MODEL=xai:grok-4.7
|
||||
```
|
||||
|
||||
4. In npx mode, run `npx @keygraph/shannon start ...` from the same shell. In source-build mode, add the two variables to `.env` and run `./shannon start ...`.
|
||||
|
||||
Suggested Grok models are `grok-4.6` and `grok-4.5`.
|
||||
The suggested Grok model is `grok-4.7`.
|
||||
|
||||
## Claude Code subscription
|
||||
|
||||
@@ -1333,41 +1373,59 @@ For organizations that need broader static and organizational coverage now, see
|
||||
|
||||
# File: docs/keygraph-platform.md
|
||||
|
||||
# Keygraph Enterprise Platform
|
||||
# Keygraph Platform
|
||||
|
||||
Shannon Lite is now Shannon Open Source. Shannon Pro is now the Keygraph platform.
|
||||
|
||||
Shannon 3.0 is an open-source pentester. It reads your source, maps routes and data flows, runs real attacks against a live target, and writes PDF and SARIF reports. It runs locally, in CI, or air-gapped with your own model. Shannon Open Source is a complete pentester, not a trial edition.
|
||||
|
||||
Keygraph Enterprise runs an enterprise-hardened fork of Shannon continuously across hundreds of repositories and adds what a security team needs around it: audit-depth static analysis on a parsed code graph, business-logic testing, SCA and secrets scanning, one deduplicated record per vulnerability across scans and scanners, generated fixes, fix verification, and SSO, RBAC, and audit logs. It is for security teams that own vulnerability management across many engineering teams and need one place to triage, assign, fix, and verify.
|
||||
The Keygraph platform is Keygraph's commercial product. It runs an enterprise-hardened fork of Shannon continuously across hundreds of repositories and adds what a security team needs around it: black-box pentesting that needs no source code, audit-depth static analysis on a parsed code graph, SCA and secrets scanning, one deduplicated record per vulnerability across scans and scanners, fix pull requests, fix verification, two-way Jira sync, and SSO, RBAC, and audit logs. It is for security teams that own vulnerability management across many engineering teams and need one place to triage, assign, fix, and verify.
|
||||
|
||||
Both editions are BYOK. Keygraph never receives your source and never proxies model traffic, open source or commercial. Shannon Open Source runs from your machine or CI runner. Keygraph Enterprise deploys as a platform inside your cloud or data center, including fully air-gapped.
|
||||
The Keygraph platform comes in three editions: the Community Program, Pro, and Enterprise. Every module is included in all three. They differ in price, eligibility, hosting, and support. Current prices and the full feature table are at [keygraph.io/pricing](https://keygraph.io/pricing).
|
||||
|
||||
## Shannon Open Source vs. Keygraph Enterprise
|
||||
Every edition is BYOK: you bring your own model key and pay your provider directly. Shannon Open Source runs from your machine or CI runner, and Keygraph never receives your source or proxies your model traffic. The Community Program and Pro are cloud-hosted and managed by Keygraph. Enterprise deploys inside your cloud or data center, including fully air-gapped.
|
||||
|
||||
| | Shannon Open Source | Keygraph Enterprise |
|
||||
| --- | --- | --- |
|
||||
| Best for | Developers and teams running repository-level pentests locally or in CI | Security organizations running continuous AppSec across many teams and repositories |
|
||||
| Code analysis | Agent pass over architecture, entry points, and data flows to seed the pentest, sized to finish inside a CI run | Persistent code property graph plus a long-running analysis harness with interprocedural taint, sanitizer modeling, cross-repo context, exploit chains, and multi-pass review |
|
||||
| Pentesting | On-demand, source-aware white-box pentesting with optional authenticated testing, focused on injection, XSS, SSRF, broken authentication, and broken authorization, with proof by exploitation | Enterprise-hardened Shannon fork run continuously, with grey-box and black-box targets and business-logic invariant testing |
|
||||
| SCA and secrets | Not included | SCA with reachability and secrets scanning including history |
|
||||
| Findings | Per-run PDF, Markdown, JSON, and SARIF, with SARIF ingestion into GitHub code scanning | One record per vulnerability per repo across scans and scanners, plus ownership, SLAs, dashboards, and audit evidence |
|
||||
| Fixes and verification | Not included | Fix PRs with verification by re-analysis and exploit replay, with no full rescan required |
|
||||
| CI/CD and source control | GitHub Action and GitLab CI component for pull-request, release, and scheduled runs, with gates on `status: exploited` | GitHub, GitLab, Azure DevOps, and Bitbucket with organization-wide policy and centrally managed integrations |
|
||||
| Deployment and models | Runs locally or on a CI runner with BYOK to any Anthropic- or OpenAI-compatible endpoint or local model | Deployed in your AWS, GCP, Azure, or on-prem environment. Customer-hosted services and stored platform data remain inside your environment. Model requests go directly to the provider, private endpoint, gateway, or local model you configure. A local model supports fully disconnected deployments |
|
||||
| Governance, license, support | AGPL-3.0 and community support | SSO, SCIM, RBAC, and audit logs, plus a commercial license, enterprise support, and SOC 2 Type II |
|
||||
No source code? Use the [Blackbox Pentester](https://keygraph.io/agentic-blackbox-pentester).
|
||||
|
||||
## Compare editions
|
||||
|
||||
| | Shannon Open Source | Community Program | Pro | Enterprise |
|
||||
| --- | --- | --- | --- | --- |
|
||||
| Price | Open source under AGPL-3.0. You pay only your own model costs | $0 in cloud service fees while you qualify. You pay only your own AI provider usage | $50 per active developer per month, every module included, no add-ons | Custom |
|
||||
| Best for | Developers and teams running repository-level pentests locally or in CI | U.S.-based 501(c)(3) nonprofits, and seed or pre-Series-A startups with 20 or fewer active developers | Teams that want the full Keygraph platform, cloud-hosted and managed by Keygraph | Security organizations running continuous AppSec across many teams and repositories that need the platform inside their own environment |
|
||||
| Code analysis | Agent pass over architecture, entry points, and data flows to seed the pentest, plus optional multi-stage security code analysis, sized to finish inside a CI run | Same as Pro | Persistent code property graph plus a long-running analysis harness with interprocedural taint, sanitizer modeling, cross-repo context, exploit chains, and multi-pass review | Same as Pro |
|
||||
| White-box pentesting | On-demand, source-aware white-box pentesting with optional authenticated testing, focused on injection, XSS, SSRF, broken authentication, and broken authorization, with proof by exploitation | Same as Pro | Enterprise-hardened Shannon fork run continuously against white-box and grey-box targets, with proof by exploitation | Same as Pro |
|
||||
| Black-box pentesting (no source code) | Not included. Shannon Open Source needs the target's source code | Same as Pro | Agents attack the running application from the outside with no source access, with up to 4 login credentials for multi-role testing. Findings are validated with working exploits | Same as Pro |
|
||||
| SCA and secrets | Not included | Same as Pro | SCA with reachability and secrets scanning, including repository history | Same as Pro |
|
||||
| Findings management | Per-run PDF, Markdown, JSON, and SARIF 2.1.0, with SARIF upload to GitHub code scanning | Same as Pro | One record per vulnerability per repository across scans and scanners, with status history, auto-reopen, ownership, SLAs, dashboards, and audit evidence | Same as Pro |
|
||||
| Fix pull requests and retests | Fix manually from the report, then re-run the scan to verify | Same as Pro | Fix pull requests for a developer to review and apply, verified by re-analysis and exploit replay without a full rescan | Same as Pro |
|
||||
| Ticketing | Not included | Same as Pro | Two-way Jira sync | Same as Pro |
|
||||
| CI/CD and source control | GitHub Action and GitLab CI component for pull-request, release, and scheduled runs, with gates on `status: exploited` | Same as Pro | GitHub, GitLab, Azure DevOps, and Bitbucket with organization-wide policy and centrally managed integrations | Same as Pro |
|
||||
| Deployment and models | Runs locally or on a CI runner with BYOK to any Anthropic- or OpenAI-compatible endpoint or local model | Same as Pro | Cloud-hosted and managed by Keygraph, in a US or EU region. BYOK: inference runs under your own model key and provider account | Deployed in your AWS, GCP, Azure, or on-prem environment, including fully air-gapped. Customer-hosted services and stored platform data remain inside your environment. Model requests go directly to the provider, private endpoint, gateway, or local model you configure. A local model supports fully disconnected deployments |
|
||||
| Governance and compliance | Not included | Same as Pro | SSO via SAML 2.0 or OIDC with SCIM, RBAC, audit logs, SOC 2 Type II report under NDA, and a standard DPA | Same as Pro, plus a custom SLA and a security patch SLA |
|
||||
| License and support | AGPL-3.0, community support on Discord | Community support on Discord, with no SLA or uptime commitment | Commercial license, with email and Slack support | Commercial license, with a support engineer, 24x7 Severity 1 support, quarterly business reviews, and white-glove onboarding |
|
||||
|
||||
### What Pro includes
|
||||
|
||||
Pro is the full Keygraph platform, cloud-hosted and managed by Keygraph, at $50 per active developer per month. Every module is included, with unlimited repositories and scans under BYOK. That covers white-box and black-box pentesting, agentic SAST, SCA, secrets scanning, findings management with status history, retests (fix verification by re-analysis and exploit replay), fix pull requests, two-way Jira sync, and SSO, RBAC, and audit logs.
|
||||
|
||||
None of these require Enterprise. Enterprise adds self-hosted or fully air-gapped deployment, a dedicated engineer and white-glove onboarding, and a custom SLA and security patch SLA.
|
||||
|
||||
The Community Program is the same managed Pro platform, with every module included, at $0 in cloud service fees for organizations that qualify. It is cloud-hosted only. See the [Community Program](https://keygraph.io/community-program) for eligibility and how to apply.
|
||||
|
||||
## How it fits your pipeline
|
||||
|
||||
1. Scans run on pull requests, releases, and a schedule against repositories in GitHub, GitLab, Azure DevOps, or Bitbucket.
|
||||
2. Pipelines gate on exploited severity. A code-analysis hypothesis never fails a build.
|
||||
3. Findings from every scanner and every run land as one record per vulnerability per repository, with an owner and an SLA. The same finding across ten runs is one record, not ten alerts.
|
||||
4. From a finding, Keygraph opens a fix PR into your normal review flow.
|
||||
4. From a finding, Keygraph opens a fix PR into your normal review flow, for a developer to review and apply.
|
||||
5. Verification confirms the fix against the changed code and the original exploit. No full rescan is required.
|
||||
|
||||
## What is different technically
|
||||
|
||||
### Static analysis on a code property graph
|
||||
|
||||
Shannon Open Source's code analysis is sized to finish inside a CI run: agents read the repository, map the attack surface, and hand candidates to the pentester. Enterprise is built for depth instead. It first parses each repository into a persistent code property graph, then runs an analysis harness derived from one built for long-running vulnerability audits, heavily adapted to query the graph rather than read files. The harness decomposes the application into risk, taint-flow, framework, and specialist tasks and supports longer-running audit workflows beyond typical CI job windows.
|
||||
Shannon Open Source's code analysis is sized to finish inside a CI run: agents read the repository, map the attack surface, and hand candidates to the pentester. The Keygraph platform is built for depth instead. It first parses each repository into a persistent code property graph, then runs an analysis harness derived from one built for long-running vulnerability audits, heavily adapted to query the graph rather than read files. The harness decomposes the application into risk, taint-flow, framework, and specialist tasks and supports longer-running audit workflows beyond typical CI job windows.
|
||||
|
||||
On the graph, it performs:
|
||||
|
||||
@@ -1379,58 +1437,62 @@ On the graph, it performs:
|
||||
|
||||
### Business-logic invariants
|
||||
|
||||
Shannon Open Source focuses on injection, XSS, SSRF, and broken authentication and authorization. Enterprise adds testing for the bugs that do not fit a vulnerability class: it derives invariants the application is supposed to hold (tenant isolation, workflow ordering, approval limits, balance conservation, state transitions) and tests them against the running application. This is where application-specific vulnerabilities live and where pattern-based SAST often provides little or no signal.
|
||||
Shannon Open Source focuses on injection, XSS, SSRF, and broken authentication and authorization. The Keygraph platform adds testing for the bugs that do not fit a vulnerability class: it derives invariants the application is supposed to hold (tenant isolation, workflow ordering, approval limits, balance conservation, state transitions) and tests them against the running application. This is where application-specific vulnerabilities live and where pattern-based SAST often provides little or no signal.
|
||||
|
||||
### Proof by exploitation
|
||||
|
||||
The pentesting engine is a hardened fork of Shannon with the same rule: a pentest finding requires a working exploit. No exploit, no finding. Enterprise stores the exploit and replays it later to verify the fix.
|
||||
The pentesting engine is a hardened fork of Shannon with the same rule: a pentest finding requires a working exploit. No exploit, no finding. The platform stores the exploit and replays it later to verify the fix.
|
||||
|
||||
The Blackbox Pentester applies the same rule without source access. It attacks the running application from the outside with a real browser and terminal, and it can take up to 4 login credentials (Google OAuth, GitHub, or custom auth) to test privilege escalation and IDOR across roles.
|
||||
|
||||
SCA prioritizes vulnerable dependencies that application code actually reaches. Secrets scanning covers current source and repository history.
|
||||
|
||||
<p align="center">
|
||||
<img src="../assets/keygraph-platform/agentic-sast-results.png" alt="Keygraph Enterprise findings grouped into business-logic issues, point issues, and secrets" width="100%">
|
||||
<img src="../assets/keygraph-platform/agentic-sast-results.png" alt="Keygraph platform findings grouped into business-logic issues, point issues, and secrets" width="100%">
|
||||
</p>
|
||||
|
||||
## Findings
|
||||
|
||||
Shannon Open Source hands you a report per scan. Enterprise dedupes across runs and across scanners, deterministically and semantically, into one record per vulnerability per repository. Each record carries evidence, source location, severity, scan history, status, owner, resolution, and last-verified state.
|
||||
Shannon Open Source hands you a report per scan. The Keygraph platform dedupes across runs and across scanners, deterministically and semantically, into one record per vulnerability per repository. Each record carries evidence, source location, severity, scan history, status, owner, resolution, and last-verified state. This is included in the Community Program, Pro, and Enterprise.
|
||||
|
||||
Workflows cover assignment, triage, false-positive and risk-acceptance decisions, and SLA policies with escalation and aging. Dashboards report open risk, coverage, new versus resolved, SLA compliance, and MTTR, exportable as evidence for customers and auditors.
|
||||
Workflows cover assignment, triage, false-positive and risk-acceptance decisions, and SLA policies with escalation and aging. A resolved finding that reappears in a later scan reopens automatically. Dashboards report open risk, coverage, new versus resolved, SLA compliance, and MTTR, exportable as evidence for customers and auditors. Findings sync both ways with Jira.
|
||||
|
||||
Findings still require human review. Enterprise's extra review passes reduce weakly supported findings, but they do not eliminate them.
|
||||
Findings still require human review. The platform's extra review passes reduce weakly supported findings, but they do not eliminate them.
|
||||
|
||||
<p align="center">
|
||||
<img src="../assets/keygraph-platform/canonical-findings.png" alt="Keygraph Enterprise findings inventory with severity, status, source, and verification filters" width="100%">
|
||||
<img src="../assets/keygraph-platform/canonical-findings.png" alt="Keygraph platform findings inventory with severity, status, source, and verification filters" width="100%">
|
||||
</p>
|
||||
|
||||
### Fix and verify
|
||||
|
||||
From a finding, Keygraph generates a patch scoped to that finding and opens a pull request. It never commits to a protected branch.
|
||||
From a finding, Keygraph generates a patch scoped to that finding and opens a pull request for a developer to review and apply. It never commits to a protected branch.
|
||||
|
||||
<p align="center">
|
||||
<img src="../assets/keygraph-platform/automated-remediation.png" alt="Keygraph Enterprise remediation workflow for generating a fix and opening a pull request" width="100%">
|
||||
<img src="../assets/keygraph-platform/automated-remediation.png" alt="Keygraph platform remediation workflow for generating a fix and opening a pull request" width="100%">
|
||||
</p>
|
||||
|
||||
Verification re-analyzes the changed code and, for pentest findings, replays the original exploit against the patched target. The verdict comes from deterministic checks plus a review pass, without rerunning the full scan.
|
||||
|
||||
<p align="center">
|
||||
<img src="../assets/keygraph-platform/targeted-verification.png" alt="Keygraph Enterprise finding-verification workflow" width="100%">
|
||||
<img src="../assets/keygraph-platform/targeted-verification.png" alt="Keygraph platform finding-verification workflow" width="100%">
|
||||
</p>
|
||||
|
||||
## Deployment and access control
|
||||
|
||||
Keygraph Enterprise deploys entirely inside your AWS, GCP, Azure, or on-prem environment, including networks with no internet egress. There is no Keygraph-operated control plane. Customer-hosted services and stored platform data remain inside your environment for the life of the deployment.
|
||||
The Community Program and Pro are cloud-hosted and managed by Keygraph on AWS, with regional isolation in a US or EU region. Scans run in isolated, single-use containers that clone your repository and are destroyed when the scan ends. The full source tree is never persisted. Only the code fragments needed to render findings, deduplicate results, and propose fixes are kept, and all code-derived data is encrypted at rest. See [Code security posture](https://keygraph.io/code-security-posture) for the full controls.
|
||||
|
||||
Model access is BYOK and BYOM. Route workloads to Anthropic, OpenAI, xAI, or Bedrock, a private cloud endpoint, your own gateway such as LiteLLM with your routing and policy applied, or local models on vLLM or Ollama. Model requests go directly to the endpoint you configure. Keygraph never receives or proxies them. A local model supports a fully disconnected deployment.
|
||||
Enterprise deploys entirely inside your AWS, GCP, Azure, or on-prem environment, including networks with no internet egress. There is no Keygraph-operated control plane. Customer-hosted services and stored platform data remain inside your environment for the life of the deployment.
|
||||
|
||||
Access control: SAML/OIDC SSO, SCIM, roles with repository-scoped visibility (RBAC, plus attribute and relationship rules where needed), full audit log, scoped API keys.
|
||||
Model access is BYOK and BYOM in every edition. Route workloads to Anthropic, OpenAI, xAI, or Bedrock, a private cloud endpoint, your own gateway such as LiteLLM with your routing and policy applied, or local models on vLLM or Ollama. In Enterprise deployments, model requests leave your environment only to reach the endpoint you configure, and a local model supports a fully disconnected deployment.
|
||||
|
||||
Access control, in the Community Program, Pro, and Enterprise: SAML/OIDC SSO, SCIM, roles with repository-scoped visibility (RBAC, plus attribute and relationship rules where needed), full audit log, scoped API keys.
|
||||
|
||||
<p align="center">
|
||||
<img src="../assets/keygraph-platform/enterprise-access-control.png" alt="Keygraph Enterprise roles and repository visibility controls" width="100%">
|
||||
<img src="../assets/keygraph-platform/enterprise-access-control.png" alt="Keygraph platform roles and repository visibility controls" width="100%">
|
||||
</p>
|
||||
|
||||
Keygraph maintains a SOC 2 Type II audit. The report is available to customers under NDA.
|
||||
|
||||
## Talk to Keygraph
|
||||
|
||||
Visit [keygraph.io](https://keygraph.io), book a [demo](https://cal.com/team/keygraph/shannon-pro), or email [shannon@keygraph.io](mailto:shannon@keygraph.io).
|
||||
Visit [keygraph.io](https://keygraph.io), see [pricing](https://keygraph.io/pricing), apply to the [Community Program](https://keygraph.io/community-program), book a [demo](https://cal.com/team/keygraph/keygraph-technical-demo), or email [shannon@keygraph.io](mailto:shannon@keygraph.io).
|
||||
@@ -20,10 +20,11 @@ Use this file as the concise entry point for AI agents and LLMs reading this rep
|
||||
|
||||
## Keygraph Platform
|
||||
|
||||
- [Keygraph platform](docs/keygraph-platform.md): Commercial continuous pentesting and AppSec platform, including black-box and white-box pentesting, parsed-code SAST, source-to-sink analysis, remediation workflows, CI/CD gating, SLA tracking, reporting, and enterprise deployment.
|
||||
- [Keygraph platform](docs/keygraph-platform.md): Commercial continuous pentesting and AppSec platform, compared edition by edition with Shannon Open Source. The Community Program, Pro, and Enterprise editions all include black-box and white-box pentesting, parsed-code SAST, SCA and secrets scanning, findings management, fix pull requests and retests, Jira sync, and SSO, RBAC, and audit logs. Pro and the Community Program are cloud-hosted; Enterprise is self-hosted or air-gapped. Shannon Lite is now Shannon Open Source. Shannon Pro is now the Keygraph platform.
|
||||
|
||||
## External Links
|
||||
|
||||
- [Keygraph website](https://keygraph.io): Company and commercial product information.
|
||||
- [Keygraph demo](https://cal.com/team/keygraph/shannon-pro): Demo and trial contact path.
|
||||
- [Keygraph pricing](https://keygraph.io/pricing): Prices and the feature table for Shannon Open Source, the Community Program, Pro, and Enterprise.
|
||||
- [Keygraph demo](https://cal.com/team/keygraph/keygraph-technical-demo): Technical demo booking.
|
||||
- [Community Discord](https://discord.gg/cmctpMBXwE): Community support and discussion.
|
||||
Reference in new issue
Block a user