Compare commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
f64a30040e | ||
|
|
b13788d8ef | ||
|
|
53118c6203 | ||
|
|
af1ed2a563 | ||
|
|
12d1c48a78 | ||
|
|
dfb7c69d3b |
No files matched your search
@@ -189,8 +189,6 @@ jobs:
|
||||
|
||||
- name: Publish npm package
|
||||
working-directory: apps/cli
|
||||
env:
|
||||
NODE_AUTH_TOKEN: ${{ secrets.NPM_TOKEN }}
|
||||
run: |
|
||||
if npm view "@keygraph/shannon@${{ needs.preflight.outputs.version }}" version 2>/dev/null; then
|
||||
echo "Version already published, skipping"
|
||||
|
||||
@@ -201,8 +201,6 @@ jobs:
|
||||
|
||||
- name: Publish npm package
|
||||
working-directory: apps/cli
|
||||
env:
|
||||
NODE_AUTH_TOKEN: ${{ secrets.NPM_TOKEN }}
|
||||
run: |
|
||||
if npm view "@keygraph/shannon@${{ needs.preflight.outputs.version }}" version 2>/dev/null; then
|
||||
echo "Version already published, skipping"
|
||||
|
||||
@@ -3,23 +3,26 @@
|
||||
|
||||
<div align="center">
|
||||
|
||||
<picture>
|
||||
<source media="(prefers-color-scheme: dark)" srcset="./assets/github-banner-dark.png">
|
||||
<source media="(prefers-color-scheme: light)" srcset="./assets/github-banner-light.png">
|
||||
<img src="./assets/github-banner.png" alt="Shannon - AI Pentester by Keygraph" width="100%">
|
||||
|
||||
# Shannon - AI Pentester by Keygraph
|
||||
</picture>
|
||||
|
||||
<a href="https://trendshift.io/repositories/15604" target="_blank"><img src="https://trendshift.io/api/badge/repositories/15604" alt="KeygraphHQ%2Fshannon | Trendshift" style="width: 250px; height: 55px;" width="250" height="55"/></a>
|
||||
|
||||
Shannon is an autonomous, AI pentester for web applications and APIs. <br />
|
||||
### Shannon is an autonomous, AI pentester for web applications and APIs.
|
||||
|
||||
It analyzes your source code, identifies attack paths, and executes real exploits to prove vulnerabilities before they reach production.
|
||||
|
||||
**This repository is Shannon Open Source: the full agent, run locally from your command line.**
|
||||
|
||||
---
|
||||
|
||||
<a href="https://discord.gg/9ZqQPuhJB7"><img src="./assets/discord.png" height="40" alt="Join Discord"></a>
|
||||
<a href="https://keygraph.io/"><img src="./assets/Keygraph_Button.png" height="40" alt="Visit Keygraph.io"></a>
|
||||
<a href="https://discord.gg/9ZqQPuhJB7"><picture><source media="(prefers-color-scheme: dark)" srcset="./assets/discord_button_dark.png"><source media="(prefers-color-scheme: light)" srcset="./assets/discord_button_light.png"><img src="./assets/discord_button_light.png" height="40" alt="Join Discord"></picture></a> <a href="https://keygraph.io/"><picture><source media="(prefers-color-scheme: dark)" srcset="./assets/keygraph_button_dark.png"><source media="(prefers-color-scheme: light)" srcset="./assets/keygraph_button_light.png"><img src="./assets/keygraph_button_light.png" height="40" alt="Visit Keygraph.io"></picture></a>
|
||||
|
||||
---
|
||||
|
||||
</div>
|
||||
|
||||
> [!TIP]
|
||||
@@ -38,6 +41,7 @@ It analyzes your source code, identifies attack paths, and executes real exploit
|
||||
- [License](#license)
|
||||
- [About Keygraph](#about-keygraph)
|
||||
- [Community and Support](#community-and-support)
|
||||
- [Common Questions](#common-questions)
|
||||
|
||||
## What is Shannon?
|
||||
|
||||
@@ -73,7 +77,7 @@ Sample penetration test reports from intentionally vulnerable applications, prod
|
||||
|
||||
- **Docker**: required for the worker container.
|
||||
- **Node.js 18+**: required for the recommended `npx` workflow.
|
||||
- **AI provider credentials**: Anthropic, OpenAI, xAI, or AWS Bedrock - or [any other provider](docs/ai-providers.md#any-other-provider). Claude models are recommended. For suggested model IDs per provider, plus gateways and custom base URLs, see [AI providers](docs/ai-providers.md#suggested-models).
|
||||
- **AI provider credentials**: Shannon runs on Anthropic, OpenAI, xAI, AWS Bedrock, [any other provider](docs/ai-providers.md#any-other-provider) in the harness catalogue, and any endpoint that speaks the Anthropic Messages API or the OpenAI Chat Completions or Responses API through a [custom base URL](docs/ai-providers.md#custom-base-url). You bring your own key, and Keygraph never proxies your model traffic. Shannon is provider-agnostic. See [AI providers](docs/ai-providers.md#suggested-models) for suggested model IDs.
|
||||
- **Cyber safeguards cleared with your provider**: Anthropic and OpenAI apply real-time safeguards to cyber-security workloads, which can interrupt a scan mid-run. Complete their guidance for legitimate security testers before your first run - see [AI providers](docs/ai-providers.md#cyber-safeguards-do-this-before-your-first-scan).
|
||||
|
||||
### Run Shannon
|
||||
@@ -107,6 +111,8 @@ For source builds, authenticated scans, provider-specific setup, and platform no
|
||||
- **Authenticated testing**: configuration files can describe login flows, test credentials, TOTP, email-based login flows, focus areas, and rules of engagement.
|
||||
- **OWASP-focused coverage**: Shannon targets exploitable Injection, XSS, SSRF, Broken Authentication, and Broken Authorization issues.
|
||||
- **Resumable workspaces**: Shannon can resume interrupted runs without re-running completed agents.
|
||||
- **Machine-readable output**: Shannon emits findings as structured JSON, and as SARIF 2.1.0 when you enable it in configuration. SARIF is the OASIS standard for static analysis results, so findings flow into any code scanning service, vulnerability management platform, security dashboard, or CI/CD pipeline that reads it.
|
||||
- **Bring your own key, provider-agnostic**: Shannon runs on Anthropic, OpenAI, xAI, AWS Bedrock, and any endpoint speaking the Anthropic Messages API or the OpenAI Chat Completions or Responses API, including self-hosted models served through Ollama, vLLM, or LM Studio and gateways such as OpenRouter and LiteLLM. You supply the credentials, so source code and model traffic stay inside your infrastructure. Local and self-hosted models are technically supported but not recommended: they may not follow Shannon's instructions or tool-use constraints as reliably as frontier models, so take that path only if you know how your chosen model behaves.
|
||||
|
||||
## Editions
|
||||
|
||||
@@ -210,7 +216,7 @@ Important limitations:
|
||||
|
||||
- Shannon Open Source focuses on actively exploitable issues such as Injection, XSS, SSRF, Broken Authentication, and Broken Authorization. Broader static-analysis coverage, including vulnerable dependencies and insecure configurations, is delivered through the Keygraph platform.
|
||||
- Findings still require human review. LLM-generated reports can contain weakly supported or incorrect details.
|
||||
- Shannon is officially supported with Claude models. Smaller, alternative, or proxied non-Claude models may be incomplete or unstable.
|
||||
- Anthropic, OpenAI, xAI, and AWS Bedrock are built-in providers, and any Anthropic Messages API or OpenAI Chat Completions or Responses API endpoint works through a custom base URL. Model capability varies, and a model that does not follow Shannon's instructions or tool-use constraints reliably will produce weaker results.
|
||||
- A full run can take roughly 1 to 1.5 hours and may incur LLM API costs depending on model pricing and application complexity.
|
||||
- Do not scan untrusted or adversarial codebases. AI-powered tools that read source code can be exposed to prompt injection.
|
||||
|
||||
@@ -249,6 +255,32 @@ Stay connected:
|
||||
- [Twitter/X: @KeygraphHQ](https://twitter.com/KeygraphHQ)
|
||||
- [LinkedIn: Keygraph](https://linkedin.com/company/keygraph)
|
||||
|
||||
## Common Questions
|
||||
|
||||
### Can I self-host Shannon?
|
||||
|
||||
Yes. Shannon Open Source runs entirely on your own infrastructure in an ephemeral Docker container. Your source code is mounted read-only and never leaves your environment.
|
||||
|
||||
### Does Shannon support bring your own key (BYOK)?
|
||||
|
||||
Yes, always. You provide the LLM credentials Shannon uses to run a pentest, in every deployment, open source and commercial. Keygraph never proxies your model traffic.
|
||||
|
||||
### Does Shannon output SARIF?
|
||||
|
||||
Yes. Shannon emits SARIF 2.1.0, the OASIS standard format for static analysis results, alongside structured JSON. Any SARIF consumer reads it: code scanning services, vulnerability management platforms, security dashboards, and CI/CD pipelines. Set `report.sarif` to `"true"` in your configuration file to enable the SARIF log.
|
||||
|
||||
### Which AI providers does Shannon support?
|
||||
|
||||
Anthropic, OpenAI, xAI, and AWS Bedrock are built in and configured directly by provider ID. Beyond those, Shannon runs on any endpoint that implements the Anthropic Messages API or the OpenAI Chat Completions or Responses API, reached through a custom base URL. The rule is the API format, not the vendor. Shannon uses a single unified model setting throughout a pentest.
|
||||
|
||||
### Can I run Shannon on a local or self-hosted model?
|
||||
|
||||
Technically yes, but it is not recommended. Shannon works with local models served through Ollama, vLLM, or LM Studio, which expose an OpenAI-compatible endpoint, as well as routers such as OpenRouter and gateways such as LiteLLM. Point Shannon at the endpoint with a custom base URL. Capability varies, and a model that does not follow Shannon's instructions or tool-use constraints reliably will produce weaker pentests than a frontier model, so take this path only if you know how your chosen model behaves. See [AI providers](docs/ai-providers.md#custom-base-url).
|
||||
|
||||
### Does Shannon actually exploit vulnerabilities, or just scan?
|
||||
|
||||
Shannon executes real exploits. It reports a finding only when it has produced a working proof-of-concept, and discards hypotheses it cannot prove. It is a pentester, not a scanner.
|
||||
|
||||
<p align="center">
|
||||
<b>Built by <a href="https://keygraph.io">Keygraph</a></b>
|
||||
</p>
|
||||
@@ -1,21 +1,23 @@
|
||||
/**
|
||||
* `shannon logs` command — tail a scan's live log.
|
||||
*
|
||||
* Uses chokidar for reliable cross-platform file watching and
|
||||
* bounded synchronous reads to prevent duplicate output.
|
||||
* The log file is streamed for its content; completion is decided by Temporal (the
|
||||
* workflow's status), so a worker that dies mid-run can't leave the tail hanging. Uses
|
||||
* chokidar for reliable cross-platform file watching and bounded synchronous reads to
|
||||
* prevent duplicate output.
|
||||
*/
|
||||
|
||||
import fs from 'node:fs';
|
||||
import path from 'node:path';
|
||||
import { setTimeout as sleep } from 'node:timers/promises';
|
||||
import { watch } from 'chokidar';
|
||||
import { fail } from '../errors.js';
|
||||
import { getWorkspacesDir } from '../home.js';
|
||||
import { resolveRunFile } from '../paths.js';
|
||||
import { resolveWorkflowId } from '../session.js';
|
||||
import { waitForWorkflowClose } from '../temporal-client.js';
|
||||
import { stdoutIsTerminal } from '../tty.js';
|
||||
|
||||
// Match the exact line the worker writes — anchored to prevent false positives from agent output
|
||||
const COMPLETION_PATTERN = /^Scan (COMPLETED|FAILED)$/m;
|
||||
|
||||
/** Read a byte range from a file and return it as a UTF-8 string. */
|
||||
function readRange(filePath: string, start: number, end: number): string {
|
||||
const length = end - start;
|
||||
@@ -62,60 +64,114 @@ export function resolveLogFile(workspaceId: string): string {
|
||||
);
|
||||
}
|
||||
|
||||
export interface TailOptions {
|
||||
/** Workflow whose Temporal status decides when the tail stops. Without it, only Ctrl-C ends the tail. */
|
||||
readonly workflowId?: string;
|
||||
/** Called if the tail ends because Temporal became unreachable, with the captured error. */
|
||||
readonly onUnreachable?: (lastError: string) => void;
|
||||
}
|
||||
|
||||
/** Outcome of a tail: whether the streamed log already contained the worker's `Scan FAILED` block. */
|
||||
export interface TailResult {
|
||||
readonly sawFailure: boolean;
|
||||
}
|
||||
|
||||
// The worker writes this exact line at the head of its terminal failure summary.
|
||||
const FAILURE_MARKER = /^Scan FAILED$/m;
|
||||
|
||||
/**
|
||||
* Tail a scan's log until it reports completion, resolving when the completion marker appears
|
||||
* (or the file is gone, or Ctrl-C stops it). Never exits the process, so the caller decides what
|
||||
* happens next: plain `logs` exits 0; `start --follow` reads the workflow outcome first.
|
||||
* Stream a scan's log to the terminal until the workflow closes (completion comes from Temporal,
|
||||
* or Ctrl-C). A Temporal outage is warned about and, if sustained, ends the tail with a diagnostic.
|
||||
* Never exits the process: plain `logs` exits; `start --follow` reads the workflow outcome first.
|
||||
* Reports whether the log already showed the failure, so a caller need not print it a second time.
|
||||
*/
|
||||
export function tailUntilComplete(logFile: string): Promise<void> {
|
||||
export function tailUntilComplete(logFile: string, opts: TailOptions = {}): Promise<TailResult> {
|
||||
return new Promise((resolve) => {
|
||||
let position = 0;
|
||||
let done = false;
|
||||
let sawFailure = false;
|
||||
const controller = new AbortController();
|
||||
let watcher: ReturnType<typeof watch> | undefined;
|
||||
|
||||
/**
|
||||
* Output any new content appended since the last read.
|
||||
* Returns true when the workflow completion marker is detected.
|
||||
*/
|
||||
function flush(): boolean {
|
||||
/** Output any new content appended since the last read. */
|
||||
function flush(): void {
|
||||
try {
|
||||
const { size } = fs.statSync(logFile);
|
||||
if (size <= position) return false;
|
||||
|
||||
if (size <= position) return;
|
||||
const data = readRange(logFile, position, size);
|
||||
process.stdout.write(data);
|
||||
position = size;
|
||||
|
||||
return COMPLETION_PATTERN.test(data);
|
||||
if (!sawFailure && FAILURE_MARKER.test(data)) {
|
||||
sawFailure = true;
|
||||
}
|
||||
} catch {
|
||||
// File deleted or unreadable — treat as done
|
||||
return true;
|
||||
// File not present yet or transiently unreadable — nothing to flush this round.
|
||||
}
|
||||
}
|
||||
|
||||
// 1. Output existing content
|
||||
if (flush()) {
|
||||
resolve();
|
||||
return;
|
||||
function finish(): void {
|
||||
if (done) return;
|
||||
done = true;
|
||||
controller.abort();
|
||||
if (watcher) {
|
||||
watcher.close().finally(() => resolve({ sawFailure }));
|
||||
// Safety net — resolve anyway if watcher.close() stalls.
|
||||
setTimeout(() => resolve({ sawFailure }), 1000).unref();
|
||||
} else {
|
||||
resolve({ sawFailure });
|
||||
}
|
||||
}
|
||||
|
||||
// 2. Watch for appended content via chokidar
|
||||
const watcher = watch(logFile, { persistent: true });
|
||||
// 1. Output existing content, then stream anything appended.
|
||||
flush();
|
||||
watcher = watch(logFile, { persistent: true });
|
||||
watcher.on('change', () => flush());
|
||||
|
||||
const stop = (): void => {
|
||||
watcher.close().finally(() => resolve());
|
||||
// Safety net — resolve anyway if watcher.close() stalls
|
||||
setTimeout(() => resolve(), 1000).unref();
|
||||
};
|
||||
// 2. Ctrl-C stops watching.
|
||||
process.on('SIGINT', finish);
|
||||
|
||||
watcher.on('change', () => {
|
||||
if (flush()) stop();
|
||||
});
|
||||
|
||||
process.on('SIGINT', stop);
|
||||
// 3. Temporal decides completion. Without a workflow id, the tail relies on Ctrl-C alone.
|
||||
if (opts.workflowId) {
|
||||
waitForWorkflowClose(opts.workflowId, {
|
||||
signal: controller.signal,
|
||||
onConnectionTrouble: (lastError) => {
|
||||
if (!done) console.error(`\n⚠ Lost contact with Temporal, retrying… (${lastError})`);
|
||||
},
|
||||
onReconnected: () => {
|
||||
if (!done) console.error(' Reconnected to Temporal.');
|
||||
},
|
||||
})
|
||||
.then(async (end) => {
|
||||
if (done) return;
|
||||
// Flush, let a just-written final summary land, then flush the tail once more.
|
||||
flush();
|
||||
await sleep(750).catch(() => {});
|
||||
flush();
|
||||
if (end.reason === 'unreachable') {
|
||||
console.error('\nScan watch aborted: lost contact with Temporal.');
|
||||
console.error(` Last error: ${end.lastError}`);
|
||||
console.error(' Temporal may have crashed — check `docker compose logs temporal`.');
|
||||
opts.onUnreachable?.(end.lastError);
|
||||
}
|
||||
finish();
|
||||
})
|
||||
.catch(() => {
|
||||
// waitForWorkflowClose never rejects; guard only against an aborted race.
|
||||
});
|
||||
}
|
||||
});
|
||||
}
|
||||
|
||||
export function logs(workspaceId: string): void {
|
||||
const logFile = resolveLogFile(workspaceId);
|
||||
const workflowId = resolveWorkflowId(workspaceId);
|
||||
console.error(stdoutIsTerminal() ? `Tailing scan log: ${logFile}` : 'Tailing scan log');
|
||||
tailUntilComplete(logFile).finally(() => process.exit(0));
|
||||
|
||||
let unreachable = false;
|
||||
tailUntilComplete(logFile, {
|
||||
...(workflowId ? { workflowId } : {}),
|
||||
onUnreachable: () => {
|
||||
unreachable = true;
|
||||
},
|
||||
}).finally(() => process.exit(unreachable ? 1 : 0));
|
||||
}
|
||||
@@ -11,7 +11,9 @@ import path from 'node:path';
|
||||
import * as p from '@clack/prompts';
|
||||
import { type ShannonConfig, saveConfig } from '../config/writer.js';
|
||||
import { CURATED_PROVIDERS, type CuratedProviderId, isCuratedProvider, type OpenAiFormat } from '../model-spec.js';
|
||||
import { displaySplash } from '../splash.js';
|
||||
import { requireInteractive } from '../tty.js';
|
||||
import { getVersion } from '../version.js';
|
||||
|
||||
const SHANNON_HOME = path.join(os.homedir(), '.shannon');
|
||||
|
||||
@@ -63,7 +65,8 @@ function modelIdPlaceholder(provider: string): string | undefined {
|
||||
|
||||
export async function setup(): Promise<void> {
|
||||
requireInteractive('setup', 'For non-interactive use, export credentials as env vars (e.g. ANTHROPIC_API_KEY).');
|
||||
p.intro('Shannon Setup');
|
||||
displaySplash(getVersion());
|
||||
p.intro('Setup');
|
||||
|
||||
// 1. Select provider. "Custom Base URL" is a route, not a provider — it asks
|
||||
// which API dialect the gateway speaks and configures that provider. "Other
|
||||
|
||||
@@ -24,6 +24,7 @@ import {
|
||||
resolveRepo,
|
||||
resolveRunFile,
|
||||
} from '../paths.js';
|
||||
import { indentFailureSegments } from '../scan/failure.js';
|
||||
import { resolveWorkflowId } from '../session.js';
|
||||
import { displaySplash } from '../splash.js';
|
||||
import { getTerminalOutcome } from '../temporal-client.js';
|
||||
@@ -81,7 +82,10 @@ export async function start(args: StartArgs): Promise<void> {
|
||||
const config = args.config ? resolveConfig(args.config) : undefined;
|
||||
|
||||
// Inputs are valid — show the splash before the Docker/Temporal setup work.
|
||||
displaySplash(isLocal() ? undefined : args.version);
|
||||
// Skip it off a real terminal (e.g. CI) so piped/logged output stays clean.
|
||||
if (stdoutIsTerminal()) {
|
||||
displaySplash(isLocal() ? undefined : args.version);
|
||||
}
|
||||
|
||||
// 4. Ensure workspaces dir is writable by container user (UID 1001)
|
||||
const workspacesDir = getWorkspacesDir();
|
||||
@@ -237,12 +241,14 @@ export async function start(args: StartArgs): Promise<void> {
|
||||
}
|
||||
|
||||
/**
|
||||
* Follow a just-started scan (for `--follow`, aimed at CI): stream its log to completion, then
|
||||
* exit on the workflow outcome — 0 if the assessment ran, 1 if the scan failed. That tracks
|
||||
* whether the pipeline ran, not whether vulnerabilities were found.
|
||||
* Follow a just-started scan (for `--follow`, aimed at CI): stream its log while Temporal drives
|
||||
* completion, then exit on the workflow outcome — 0 if the assessment ran, 1 if the scan failed.
|
||||
* That tracks whether the pipeline ran, not whether vulnerabilities were found. On failure the
|
||||
* root-cause message is printed so a red CI build says why.
|
||||
*/
|
||||
async function followScan(workspace: string, workspacesDir: string): Promise<never> {
|
||||
const logFile = resolveRunFile(path.join(workspacesDir, workspace), 'workflow.log');
|
||||
const workflowId = resolveWorkflowId(workspace);
|
||||
|
||||
// The worker creates workflow.log as it starts; wait briefly so the first read doesn't
|
||||
// mistake a not-yet-created file for an already-finished scan.
|
||||
@@ -253,18 +259,38 @@ async function followScan(workspace: string, workspacesDir: string): Promise<nev
|
||||
if (stdoutIsTerminal()) {
|
||||
console.error('\n Following scan log (Ctrl-C to stop watching):\n');
|
||||
}
|
||||
await tailUntilComplete(logFile);
|
||||
|
||||
const workflowId = resolveWorkflowId(workspace);
|
||||
let temporalUnreachable = false;
|
||||
const { sawFailure } = await tailUntilComplete(logFile, {
|
||||
...(workflowId && { workflowId }),
|
||||
onUnreachable: () => {
|
||||
temporalUnreachable = true;
|
||||
},
|
||||
});
|
||||
|
||||
// The tail already printed the diagnostic; reading the outcome would only fail the same way.
|
||||
if (temporalUnreachable) {
|
||||
process.exit(1);
|
||||
}
|
||||
|
||||
if (!workflowId) {
|
||||
fail('Scan finished but its workflow id could not be resolved from session.json.');
|
||||
}
|
||||
|
||||
try {
|
||||
const outcome = await getTerminalOutcome(workflowId);
|
||||
process.exit(outcome.kind === 'success' ? 0 : 1);
|
||||
} catch {
|
||||
fail('Could not reach Temporal at 127.0.0.1:7233 to read the scan outcome.');
|
||||
if (outcome.kind === 'failed') {
|
||||
// Print the reason only when the streamed log didn't already show the worker's failure
|
||||
// summary — otherwise the worker crashed before writing it, and this is the only report.
|
||||
if (!sawFailure) {
|
||||
console.error(`\nScan failed:\n${indentFailureSegments(outcome.message)}`);
|
||||
}
|
||||
process.exit(1);
|
||||
}
|
||||
process.exit(0);
|
||||
} catch (err) {
|
||||
const detail = err instanceof Error ? err.message : String(err);
|
||||
fail('Could not read the scan outcome from Temporal at 127.0.0.1:7233.', ` ${detail}`);
|
||||
}
|
||||
}
|
||||
|
||||
|
||||
@@ -20,8 +20,10 @@ import { status } from './commands/status.js';
|
||||
import { stop } from './commands/stop.js';
|
||||
import { crash, fail, failUsage } from './errors.js';
|
||||
import { availableCommands, isHelpableCommand, printCommandHelp, START_OPTIONS } from './help.js';
|
||||
import { commandPrefix, getMode } from './mode.js';
|
||||
import { commandPrefix, getMode, isLocal, type Mode } from './mode.js';
|
||||
import { displaySplash } from './splash.js';
|
||||
import { closestMatch } from './suggest.js';
|
||||
import { stdoutIsTerminal } from './tty.js';
|
||||
import { getVersion, getVersionLine } from './version.js';
|
||||
|
||||
function blockSudo(): void {
|
||||
@@ -50,33 +52,38 @@ function renderStartOptions(): string {
|
||||
return START_OPTIONS.map(([flag, desc]) => ` ${flag.padEnd(flagWidth)} ${desc}`).join('\n');
|
||||
}
|
||||
|
||||
function showHelp(): void {
|
||||
/**
|
||||
* Render the command list with the description column aligned. Padding is computed from the
|
||||
* widest command, so it lines up regardless of the prefix (`npx @keygraph/shannon` vs `./shannon`).
|
||||
*/
|
||||
function renderUsage(prefix: string, mode: Mode): string {
|
||||
const rows: ReadonlyArray<readonly [string, string]> = [
|
||||
...(mode === 'local' ? [] : [[`${prefix} setup`, 'Configure credentials'] as const]),
|
||||
[`${prefix} start --url <url> --repo <path> [options]`, 'Start a pentest scan'],
|
||||
[`${prefix} stop <workspace> [--yes]`, 'Stop one scan'],
|
||||
[`${prefix} stop --all [--yes]`, 'Stop all scans (Temporal stays up)'],
|
||||
[`${prefix} reset`, 'Stop everything and wipe all Temporal data'],
|
||||
[`${prefix} logs <workspace>`, "Show a scan's live log"],
|
||||
[`${prefix} status <workspace> [--json]`, 'Live phase/agent progress of one scan'],
|
||||
[`${prefix} scans [--json]`, 'List completed scans and their reports'],
|
||||
...(mode === 'local' ? [[`${prefix} build [--no-cache]`, 'Build worker image'] as const] : []),
|
||||
[`${prefix} version [--json]`, 'Show version'],
|
||||
[`${prefix} help`, 'Show this help'],
|
||||
];
|
||||
|
||||
const commandWidth = Math.max(...rows.map(([command]) => command.length));
|
||||
return rows.map(([command, desc]) => ` ${command.padEnd(commandWidth)} ${desc}`).join('\n');
|
||||
}
|
||||
|
||||
function showHelp(withSplash: boolean): void {
|
||||
const mode = getMode();
|
||||
const prefix = commandPrefix();
|
||||
|
||||
console.log(`
|
||||
Shannon - AI Penetration Testing Framework
|
||||
const header = withSplash ? '' : '\nShannon - AI Penetration Testing Framework\n';
|
||||
|
||||
Usage:${
|
||||
mode === 'local'
|
||||
? ''
|
||||
: `
|
||||
${prefix} setup Configure credentials`
|
||||
}
|
||||
${prefix} start --url <url> --repo <path> [options] Start a pentest scan
|
||||
${prefix} stop <workspace> [--yes] Stop one scan
|
||||
${prefix} stop --all [--yes] Stop all scans (Temporal stays up)
|
||||
${prefix} reset Stop everything and wipe all Temporal data
|
||||
${prefix} logs <workspace> Show a scan's live log
|
||||
${prefix} status <workspace> [--json] Live phase/agent progress of one scan
|
||||
${prefix} scans [--json] List completed scans and their reports${
|
||||
mode === 'local'
|
||||
? `
|
||||
${prefix} build [--no-cache] Build worker image`
|
||||
: ''
|
||||
}
|
||||
${prefix} version [--json] Show version
|
||||
${prefix} help Show this help
|
||||
console.log(`${header}
|
||||
Usage:
|
||||
${renderUsage(prefix, mode)}
|
||||
|
||||
Options for 'start':
|
||||
${renderStartOptions()}
|
||||
@@ -89,14 +96,7 @@ Examples:
|
||||
${prefix} reset
|
||||
|
||||
Run '${prefix} <command> --help' for help on a specific command.
|
||||
${
|
||||
mode === 'local'
|
||||
? `
|
||||
State directory: ./workspaces/`
|
||||
: `
|
||||
State directory: ~/.shannon/`
|
||||
}
|
||||
Monitor scans at http://localhost:8233
|
||||
|
||||
Docs & source: https://github.com/KeygraphHQ/shannon
|
||||
`);
|
||||
}
|
||||
@@ -174,7 +174,9 @@ async function main(): Promise<void> {
|
||||
if (topic && isHelpableCommand(topic)) {
|
||||
printCommandHelp(topic);
|
||||
} else {
|
||||
showHelp();
|
||||
const bare = command === undefined;
|
||||
if (bare && stdoutIsTerminal()) displaySplash(isLocal() ? undefined : getVersion());
|
||||
showHelp(bare);
|
||||
}
|
||||
return;
|
||||
}
|
||||
|
||||
@@ -0,0 +1,31 @@
|
||||
/**
|
||||
* Rendering for the worker's '|'-delimited failure string.
|
||||
*
|
||||
* `formatWorkflowError` in the worker joins error segments — phase context, error type,
|
||||
* message, and remediation hint — with '|' as a delimiter. These helpers turn that raw
|
||||
* string into readable output for the CLI's own surfaces.
|
||||
*/
|
||||
|
||||
/**
|
||||
* Split the failure string into trimmed, non-empty lines. Segments are delimited by '|', and a
|
||||
* segment's own embedded newlines (e.g. a multi-line validation message) become their own lines so
|
||||
* each aligns with the rest of the block.
|
||||
*/
|
||||
export function parseFailureSegments(message: string): string[] {
|
||||
return message
|
||||
.split(/[|\n]/)
|
||||
.map((segment) => segment.trim())
|
||||
.filter((segment) => segment.length > 0);
|
||||
}
|
||||
|
||||
/** Multi-line block: one segment per indented line (the caller prints the header). */
|
||||
export function indentFailureSegments(message: string, indent = ' '): string {
|
||||
return parseFailureSegments(message)
|
||||
.map((segment) => `${indent}${segment}`)
|
||||
.join('\n');
|
||||
}
|
||||
|
||||
/** Single-line summary for compact contexts like the status footer. */
|
||||
export function inlineFailureReason(message: string): string {
|
||||
return parseFailureSegments(message).join(' — ');
|
||||
}
|
||||
@@ -11,6 +11,7 @@ import { BOLD, DIM, GOLD, paint, RED, YELLOW } from '../colors.js';
|
||||
import { commandPrefix } from '../mode.js';
|
||||
import type { RunningAgent } from '../temporal-client.js';
|
||||
import { agentError, deriveAgentStates, isTerminal, phaseGlyphState, type RunState, scanElapsedMs } from './derive.js';
|
||||
import { inlineFailureReason } from './failure.js';
|
||||
import { PIPELINE, type PipelineState } from './pipeline.js';
|
||||
|
||||
export interface RenderInput {
|
||||
@@ -229,7 +230,8 @@ function footerLines(input: RenderInput, opts: RenderOptions): string[] {
|
||||
const temporalValue = temporalDashboardUrl(input.workflowId);
|
||||
|
||||
if (isTerminal(input.temporalStatus)) {
|
||||
const reason = input.failureMessage ?? input.state?.error ?? 'no result recorded';
|
||||
const rawReason = input.failureMessage ?? input.state?.error;
|
||||
const reason = rawReason ? inlineFailureReason(rawReason) : 'no result recorded';
|
||||
return [
|
||||
footerDivider(opts),
|
||||
paint(
|
||||
|
||||
@@ -7,12 +7,17 @@
|
||||
* publishes — so this needs Temporal up, but no worker of its own.
|
||||
*/
|
||||
|
||||
import { setTimeout as sleep } from 'node:timers/promises';
|
||||
import { Client, Connection, WorkflowFailedError, WorkflowNotFoundError } from '@temporalio/client';
|
||||
import { ACTIVITY_TO_AGENT, type PipelineState } from './scan/pipeline.js';
|
||||
|
||||
const ADDRESS = '127.0.0.1:7233';
|
||||
const NAMESPACE = 'default';
|
||||
|
||||
// WorkflowExecutionStatusName values that mean the scan has closed. RUNNING (and the unused
|
||||
// CONTINUED_AS_NEW) are the only non-terminal states.
|
||||
const TERMINAL_STATUSES: ReadonlySet<string> = new Set(['COMPLETED', 'FAILED', 'CANCELLED', 'TERMINATED', 'TIMED_OUT']);
|
||||
|
||||
export interface RunningAgent {
|
||||
readonly agent: string;
|
||||
readonly attempt: number;
|
||||
@@ -111,6 +116,75 @@ function rootFailureMessage(err: WorkflowFailedError): string {
|
||||
return message;
|
||||
}
|
||||
|
||||
/** How a {@link waitForWorkflowClose} watch ended. */
|
||||
export type WatchEnd = { readonly reason: 'closed' } | { readonly reason: 'unreachable'; readonly lastError: string };
|
||||
|
||||
export interface WatchOptions {
|
||||
/** Poll interval in ms (default 3000). */
|
||||
readonly pollMs?: number;
|
||||
/** Consecutive connection failures before giving up (default 10 → ~30s at the default interval). */
|
||||
readonly maxConnectFailures?: number;
|
||||
/** Consecutive connection failures before {@link onConnectionTrouble} fires once (default 3). */
|
||||
readonly warnAfterFailures?: number;
|
||||
/** Abort the watch (the caller stopped for another reason, e.g. Ctrl-C). */
|
||||
readonly signal?: AbortSignal;
|
||||
/** Called once when contact is first lost, so a live follower's log isn't silent during the outage. */
|
||||
readonly onConnectionTrouble?: (lastError: string) => void;
|
||||
/** Called once when contact is regained after {@link onConnectionTrouble} fired. */
|
||||
readonly onReconnected?: () => void;
|
||||
}
|
||||
|
||||
/**
|
||||
* Resolve once the scan is no longer running, using the workflow's Temporal status as the
|
||||
* completion signal. Ends on a terminal status, a not-found workflow (closed past retention), or
|
||||
* maxConnectFailures consecutive unreachable polls (a scan can't progress while its Temporal is
|
||||
* down, so sustained no-contact is a safe stop). Never rejects; connection errors surface via the
|
||||
* callbacks and the returned {@link WatchEnd}.
|
||||
*/
|
||||
export async function waitForWorkflowClose(workflowId: string, opts: WatchOptions = {}): Promise<WatchEnd> {
|
||||
const pollMs = opts.pollMs ?? 3000;
|
||||
const maxConnectFailures = opts.maxConnectFailures ?? 10;
|
||||
const warnAfterFailures = opts.warnAfterFailures ?? 3;
|
||||
const signal = opts.signal;
|
||||
|
||||
let connectFailures = 0;
|
||||
let lastError = '';
|
||||
let warned = false;
|
||||
|
||||
while (!signal?.aborted) {
|
||||
try {
|
||||
const desc = await describeScan(workflowId);
|
||||
if (desc === null || TERMINAL_STATUSES.has(desc.status)) {
|
||||
return { reason: 'closed' };
|
||||
}
|
||||
// Reachable and still RUNNING — reset the failure streak and note any recovery.
|
||||
if (warned) {
|
||||
warned = false;
|
||||
opts.onReconnected?.();
|
||||
}
|
||||
connectFailures = 0;
|
||||
} catch (err) {
|
||||
connectFailures++;
|
||||
lastError = err instanceof Error ? err.message : String(err);
|
||||
if (!warned && connectFailures >= warnAfterFailures) {
|
||||
warned = true;
|
||||
opts.onConnectionTrouble?.(lastError);
|
||||
}
|
||||
if (connectFailures >= maxConnectFailures) {
|
||||
return { reason: 'unreachable', lastError };
|
||||
}
|
||||
}
|
||||
|
||||
try {
|
||||
await sleep(pollMs, undefined, { signal });
|
||||
} catch {
|
||||
break; // Aborted mid-wait by the caller.
|
||||
}
|
||||
}
|
||||
|
||||
return { reason: 'closed' };
|
||||
}
|
||||
|
||||
/** Final state of a closed scan: success carries the full PipelineState, failure carries the message. */
|
||||
export async function getTerminalOutcome(workflowId: string): Promise<TerminalOutcome> {
|
||||
const client = await getClient();
|
||||
|
||||
@@ -303,13 +303,17 @@ export class WorkflowLogger {
|
||||
* Output: "Error: phase context\n ErrorType\n ..."
|
||||
*/
|
||||
private formatErrorBlock(errorString: string): string {
|
||||
const segments = errorString.split('|');
|
||||
const label = 'Error: ';
|
||||
const indent = ' '.repeat(label.length);
|
||||
|
||||
const lines = segments.map((segment, i) => (i === 0 ? `${label}${segment.trim()}` : `${indent}${segment.trim()}`));
|
||||
// Segments are delimited by '|'; a segment's own embedded newlines (e.g. a multi-line
|
||||
// validation message) become their own lines so each aligns under the label.
|
||||
const lines = errorString
|
||||
.split(/[|\n]/)
|
||||
.map((segment) => segment.trim())
|
||||
.filter((segment) => segment.length > 0);
|
||||
|
||||
return `${lines.join('\n')}\n`;
|
||||
return `${lines.map((line, i) => (i === 0 ? `${label}${line}` : `${indent}${line}`)).join('\n')}\n`;
|
||||
}
|
||||
|
||||
/**
|
||||
@@ -336,17 +340,19 @@ export class WorkflowLogger {
|
||||
lines.push(this.formatErrorBlock(summary.error).trimEnd());
|
||||
}
|
||||
|
||||
lines.push('');
|
||||
lines.push('Agent Breakdown:');
|
||||
if (summary.completedAgents.length > 0) {
|
||||
lines.push('');
|
||||
lines.push('Agent Breakdown:');
|
||||
|
||||
for (const agentName of summary.completedAgents) {
|
||||
const metrics = summary.agentMetrics[agentName];
|
||||
if (metrics) {
|
||||
const duration = formatDuration(metrics.durationMs);
|
||||
const cost = metrics.costUsd !== null ? `$${metrics.costUsd.toFixed(4)}` : 'N/A';
|
||||
lines.push(` - ${agentName} (${duration}, ${cost})`);
|
||||
} else {
|
||||
lines.push(` - ${agentName}`);
|
||||
for (const agentName of summary.completedAgents) {
|
||||
const metrics = summary.agentMetrics[agentName];
|
||||
if (metrics) {
|
||||
const duration = formatDuration(metrics.durationMs);
|
||||
const cost = metrics.costUsd !== null ? `$${metrics.costUsd.toFixed(4)}` : 'N/A';
|
||||
lines.push(` - ${agentName} (${duration}, ${cost})`);
|
||||
} else {
|
||||
lines.push(` - ${agentName}`);
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
|
||||
@@ -452,9 +452,11 @@ export async function runAuthzExploitAgent(input: ActivityInput): Promise<AgentM
|
||||
/**
|
||||
* Write report.sarif when the run is exploitative and the operator asked for it.
|
||||
*
|
||||
* Skipped entirely for analysis-only runs: those findings carry no severity, so every
|
||||
* `result.level` would be invented. Failures are logged and swallowed — the SARIF log is a
|
||||
* secondary artifact and must not fail a run whose report is already written.
|
||||
* Skipped entirely for analysis-only runs. The original reason was that those findings carried
|
||||
* no severity, so every `result.level` would have been invented; since severity is recorded in
|
||||
* both modes an analysis run could now populate `level`, but it would report an assessed
|
||||
* severity as a measured one, so the gate stays. Failures are logged and swallowed — the SARIF
|
||||
* log is a secondary artifact and must not fail a run whose report is already written.
|
||||
*/
|
||||
async function writeSarifIfEnabled(
|
||||
input: ActivityInput,
|
||||
|
||||
@@ -14,5 +14,4 @@ export type {
|
||||
ResumeState,
|
||||
VulnExploitPipelineResult,
|
||||
} from './shared.js';
|
||||
export { PipelineExecutionError } from './shared.js';
|
||||
export { pentestPipeline } from './workflows.js';
|
||||
@@ -60,21 +60,6 @@ export interface PipelineState {
|
||||
summary: PipelineSummary | null;
|
||||
}
|
||||
|
||||
/**
|
||||
* Thrown by pentestPipeline() when the run fails, carrying the fully-populated
|
||||
* PipelineState (real agentMetrics, completedAgents, summary) so a consumer can
|
||||
* report actual spend instead of synthesizing a zeroed failed state. `cause`
|
||||
* preserves the original error for classification and Temporal failure reporting.
|
||||
*/
|
||||
export class PipelineExecutionError extends Error {
|
||||
override name = 'PipelineExecutionError' as const;
|
||||
readonly state: PipelineState;
|
||||
constructor(message: string, state: PipelineState, options?: { cause?: unknown }) {
|
||||
super(message, options);
|
||||
this.state = state;
|
||||
}
|
||||
}
|
||||
|
||||
// Extended state returned by getProgress query (includes computed fields)
|
||||
export interface PipelineProgress extends PipelineState {
|
||||
workflowId: string;
|
||||
|
||||
@@ -41,7 +41,6 @@ import type { ActivityInput } from './activities.js';
|
||||
import {
|
||||
type AgentMetrics,
|
||||
getProgress,
|
||||
PipelineExecutionError,
|
||||
type PipelineInput,
|
||||
type PipelineProgress,
|
||||
type PipelineState,
|
||||
@@ -718,9 +717,10 @@ export async function pentestPipeline(input: PipelineInput): Promise<PipelineSta
|
||||
});
|
||||
}
|
||||
|
||||
// Carry the populated state so a consumer can report real spend instead of a zeroed
|
||||
// failed state. The original error rides as `cause` for classification/reporting.
|
||||
throw new PipelineExecutionError(state.error ?? 'Pipeline failed', state, { cause: error });
|
||||
// Terminate the workflow in Temporal's FAILED state. WARNING: this must be an
|
||||
// ApplicationFailure — any other thrown type becomes an unhandled workflow-task failure
|
||||
// that Temporal retries indefinitely, leaving the run stuck in RUNNING.
|
||||
throw ApplicationFailure.nonRetryable(state.error ?? 'Pipeline failed', 'PipelineExecutionError');
|
||||
}
|
||||
}
|
||||
|
||||
|
||||
|
Before Width: | Height: | Size: 79 KiB |
|
Before Width: | Height: | Size: 50 KiB |
|
After Width: | Height: | Size: 49 KiB |
|
After Width: | Height: | Size: 50 KiB |
|
After Width: | Height: | Size: 75 KiB |
|
After Width: | Height: | Size: 80 KiB |
|
Before Width: | Height: | Size: 91 KiB |
|
After Width: | Height: | Size: 39 KiB |
|
After Width: | Height: | Size: 40 KiB |
@@ -25,7 +25,7 @@ Shannon forwards only the selected provider's credential into the scan container
|
||||
|
||||
### Any other provider
|
||||
|
||||
Shannon accepts any provider and model present in the Pi harness catalogue. Browse them at [pi.dev/models](https://pi.dev/models). These are technically supported but not recommended. Claude models are best-supported (see the note below).
|
||||
Shannon accepts any provider and model present in the Pi harness catalogue. Browse them at [pi.dev/models](https://pi.dev/models).
|
||||
|
||||
```bash
|
||||
export SHANNON_AI_API_KEY=your-api-key # the provider's API key
|
||||
@@ -37,7 +37,7 @@ This path covers providers whose credential is a single API key. Providers that
|
||||
`npx @keygraph/shannon setup` exposes this as the **Other provider** option.
|
||||
|
||||
> [!IMPORTANT]
|
||||
> Claude models are the best-supported option. Shannon's evaluations, internal testing, and agent harness are tuned for Claude. Other models are permitted and validated against the harness catalogue, but may not follow Shannon's instructions or tool-use constraints as reliably. Use them at your own risk.
|
||||
> Models are validated against the harness catalogue, but capability varies. A model that does not follow Shannon's instructions or tool-use constraints reliably will produce weaker pentests. Evaluate the model you choose against your own targets before depending on its results.
|
||||
|
||||
## Cyber safeguards (do this before your first scan)
|
||||
|
||||
|
||||
@@ -106,12 +106,12 @@ rules:
|
||||
|
||||
| Key | Effect |
|
||||
| --- | --- |
|
||||
| `min_severity` | Drops findings rated below this severity. Applies only when `exploit` is `"true"`. |
|
||||
| `min_severity` | Drops findings rated below this severity. Applies in both exploitative and analysis-only runs. |
|
||||
| `min_confidence` | Drops findings rated below this confidence. Applies only when `exploit` is `"false"`. |
|
||||
| `guidance` | Free-text instruction to the report agent, such as which topics to exclude. |
|
||||
| `sarif` | Emits a SARIF 2.1.0 log alongside the Markdown report. Requires `exploit: "true"`. |
|
||||
|
||||
A finding carries one rating or the other, never both: an exploited finding is rated by severity, an analysis-only finding by confidence. Setting the threshold that does not apply to the run is ignored, and Shannon logs a warning naming the one to use instead.
|
||||
Every finding carries a severity, but it does not mean the same thing in each mode: an exploitative run measures severity from what the exploit demonstrated, while an analysis-only run assesses it from the class of flaw and the impact it would have. An analysis-only finding carries a confidence rating alongside its severity, since nothing was proven. Setting `min_confidence` on an exploitative run is ignored, and Shannon logs a warning naming the threshold to use instead.
|
||||
|
||||
### SARIF Output
|
||||
|
||||
@@ -125,7 +125,7 @@ report:
|
||||
|
||||
Each finding becomes one SARIF result, filed under a rule per vulnerability class (`shannon/injection`, `shannon/xss`, `shannon/auth`, `shannon/authz`, `shannon/ssrf`) and tagged with its OWASP Top Ten 2025 category. Results are anchored to the code location the analysis phase recorded, falling back to the HTTP entry point when the finding names no file. Severity maps onto SARIF's three levels: `critical` and `high` become `error`, `medium` becomes `warning`, everything else becomes `note`.
|
||||
|
||||
The log is written only for exploitative runs. An analysis-only run rates findings by confidence and produces no severity, so there is nothing to populate `level` with; `sarif` is ignored when `exploit` is `"false"`.
|
||||
The log is written only for exploitative runs. `sarif` is ignored when `exploit` is `"false"`.
|
||||
|
||||
Supported rule types include `url_path`, `subdomain`, `domain`, `method`, `header`, `parameter`, and `code_path`.
|
||||
|
||||
|
||||
@@ -28,7 +28,7 @@ For maximum isolation, run Shannon inside a disposable virtual machine.
|
||||
## LLM and Automation Caveats
|
||||
|
||||
- **Verification is required**: Shannon uses a proof-by-exploitation methodology, but final reports can still contain weakly supported or incorrect details. Human review is essential.
|
||||
- **Model support**: Shannon is officially supported only with Claude models. Alternative models may be incomplete, inaccurate, or unstable.
|
||||
- **Model support**: results vary by model. A model that does not follow Shannon's instructions or tool-use constraints reliably may produce incomplete, inaccurate, or unstable runs.
|
||||
- **Prompt injection risk**: Do not point Shannon at untrusted or adversarial codebases. AI-powered tools that read source code can be influenced by malicious repository content.
|
||||
|
||||
## Scope of Analysis
|
||||
|
||||
@@ -8,27 +8,30 @@
|
||||
# File: README.md
|
||||
|
||||
> [!NOTE]
|
||||
> **[Shannon 2.0 now runs on the Pi harness](https://github.com/KeygraphHQ/shannon/discussions/393)**
|
||||
> **[Shannon 2.0 is officially here](https://github.com/KeygraphHQ/shannon/discussions/405)**
|
||||
|
||||
<div align="center">
|
||||
|
||||
<picture>
|
||||
<source media="(prefers-color-scheme: dark)" srcset="./assets/github-banner-dark.png">
|
||||
<source media="(prefers-color-scheme: light)" srcset="./assets/github-banner-light.png">
|
||||
<img src="./assets/github-banner.png" alt="Shannon - AI Pentester by Keygraph" width="100%">
|
||||
|
||||
# Shannon - AI Pentester by Keygraph
|
||||
</picture>
|
||||
|
||||
<a href="https://trendshift.io/repositories/15604" target="_blank"><img src="https://trendshift.io/api/badge/repositories/15604" alt="KeygraphHQ%2Fshannon | Trendshift" style="width: 250px; height: 55px;" width="250" height="55"/></a>
|
||||
|
||||
Shannon is an autonomous, white-box AI pentester for web applications and APIs. <br />
|
||||
### Shannon is an autonomous, AI pentester for web applications and APIs.
|
||||
|
||||
It analyzes your source code, identifies attack paths, and executes real exploits to prove vulnerabilities before they reach production.
|
||||
|
||||
**This repository is Shannon Open Source: the full agent, run locally from your command line.**
|
||||
|
||||
---
|
||||
|
||||
<a href="https://discord.gg/9ZqQPuhJB7"><img src="./assets/discord.png" height="40" alt="Join Discord"></a>
|
||||
<a href="https://keygraph.io/"><img src="./assets/Keygraph_Button.png" height="40" alt="Visit Keygraph.io"></a>
|
||||
<a href="https://discord.gg/9ZqQPuhJB7"><picture><source media="(prefers-color-scheme: dark)" srcset="./assets/discord_button_dark.png"><source media="(prefers-color-scheme: light)" srcset="./assets/discord_button_light.png"><img src="./assets/discord_button_light.png" height="40" alt="Join Discord"></picture></a> <a href="https://keygraph.io/"><picture><source media="(prefers-color-scheme: dark)" srcset="./assets/keygraph_button_dark.png"><source media="(prefers-color-scheme: light)" srcset="./assets/keygraph_button_light.png"><img src="./assets/keygraph_button_light.png" height="40" alt="Visit Keygraph.io"></picture></a>
|
||||
|
||||
---
|
||||
|
||||
</div>
|
||||
|
||||
> [!TIP]
|
||||
@@ -44,13 +47,14 @@ It analyzes your source code, identifies attack paths, and executes real exploit
|
||||
- [Architecture](#architecture)
|
||||
- [Documentation](#documentation)
|
||||
- [Safety, Scope, and Limitations](#safety-scope-and-limitations)
|
||||
- [License and Enterprise Licensing](#license-and-enterprise-licensing)
|
||||
- [License](#license)
|
||||
- [About Keygraph](#about-keygraph)
|
||||
- [Community and Support](#community-and-support)
|
||||
- [Common Questions](#common-questions)
|
||||
|
||||
## What is Shannon?
|
||||
|
||||
Shannon is an autonomous AI pentester developed by [Keygraph](https://keygraph.io). It performs white-box security testing of web applications and their underlying APIs by combining source-code analysis with live exploitation.
|
||||
Shannon is an autonomous AI pentester developed by [Keygraph](https://keygraph.io). It performs security testing of web applications and their underlying APIs by combining source-code analysis with live exploitation.
|
||||
|
||||
Shannon analyzes your web application's source code to identify potential attack vectors, then uses browser automation and command-line tools to execute real exploits against the running application and its APIs. Only vulnerabilities with a working proof-of-concept are included in the final report.
|
||||
|
||||
@@ -82,7 +86,7 @@ Sample penetration test reports from intentionally vulnerable applications, prod
|
||||
|
||||
- **Docker**: required for the worker container.
|
||||
- **Node.js 18+**: required for the recommended `npx` workflow.
|
||||
- **AI provider credentials**: Anthropic, OpenAI, xAI, or AWS Bedrock - or [any other provider](docs/ai-providers.md#any-other-provider). Claude models are recommended. Gateway and proxy setups are documented separately.
|
||||
- **AI provider credentials**: Shannon runs on Anthropic, OpenAI, xAI, AWS Bedrock, [any other provider](docs/ai-providers.md#any-other-provider) in the harness catalogue, and any endpoint that speaks the Anthropic Messages API or the OpenAI Chat Completions or Responses API through a [custom base URL](docs/ai-providers.md#custom-base-url). You bring your own key, and Keygraph never proxies your model traffic. Shannon is provider-agnostic. See [AI providers](docs/ai-providers.md#suggested-models) for suggested model IDs.
|
||||
- **Cyber safeguards cleared with your provider**: Anthropic and OpenAI apply real-time safeguards to cyber-security workloads, which can interrupt a scan mid-run. Complete their guidance for legitimate security testers before your first run - see [AI providers](docs/ai-providers.md#cyber-safeguards-do-this-before-your-first-scan).
|
||||
|
||||
### Run Shannon
|
||||
@@ -103,7 +107,10 @@ Shannon pulls the worker image from Docker Hub, starts the required local infras
|
||||
For source builds, authenticated scans, provider-specific setup, and platform notes, see [Documentation](#documentation).
|
||||
|
||||
> [!TIP]
|
||||
> **Prefer to run on your Claude Code subscription instead of API credits?** The [`shannon-v1`](https://github.com/KeygraphHQ/shannon/tree/shannon-v1) branch is the last release built on the Claude Agent SDK, so it accepts a Claude Code OAuth token. Generate one with `claude setup-token`, then run `npx @keygraph/shannon@1.9.0 setup` and pick **OAuth Token**. Pentests then cost nothing beyond your existing subscription.
|
||||
> **Prefer to use a subscription instead of API credits?**
|
||||
>
|
||||
> - **OpenAI Codex:** The latest version of Shannon supports ChatGPT Plus and Pro subscriptions. Follow the [OpenAI Codex subscription setup guide](docs/ai-providers.md#openai-codex-chatgpt-pluspro-subscription) to get started.
|
||||
> - **Claude Code:** The latest version of Shannon does not support Claude Code subscriptions. Follow the [Claude Code subscription setup guide](docs/ai-providers.md#claude-code-subscription) to use version `1.9.0`, which is the final release built on the Claude Agent SDK.
|
||||
|
||||
## Key Capabilities
|
||||
|
||||
@@ -113,6 +120,8 @@ For source builds, authenticated scans, provider-specific setup, and platform no
|
||||
- **Authenticated testing**: configuration files can describe login flows, test credentials, TOTP, email-based login flows, focus areas, and rules of engagement.
|
||||
- **OWASP-focused coverage**: Shannon targets exploitable Injection, XSS, SSRF, Broken Authentication, and Broken Authorization issues.
|
||||
- **Resumable workspaces**: Shannon can resume interrupted runs without re-running completed agents.
|
||||
- **Machine-readable output**: Shannon emits findings as structured JSON, and as SARIF 2.1.0 when you enable it in configuration. SARIF is the OASIS standard for static analysis results, so findings flow into any code scanning service, vulnerability management platform, security dashboard, or CI/CD pipeline that reads it.
|
||||
- **Bring your own key, provider-agnostic**: Shannon runs on Anthropic, OpenAI, xAI, AWS Bedrock, and any endpoint speaking the Anthropic Messages API or the OpenAI Chat Completions or Responses API, including self-hosted models served through Ollama, vLLM, or LM Studio and gateways such as OpenRouter and LiteLLM. You supply the credentials, so source code and model traffic stay inside your infrastructure. Local and self-hosted models are technically supported but not recommended: they may not follow Shannon's instructions or tool-use constraints as reliably as frontier models, so take that path only if you know how your chosen model behaves.
|
||||
|
||||
## Editions
|
||||
|
||||
@@ -199,7 +208,7 @@ Use these guides for operational detail:
|
||||
| --- | --- |
|
||||
| [Source build and CLI commands](docs/development.md) | Cloning, building, common commands, output paths, and local development. |
|
||||
| [Configuration](docs/configuration.md) | Authenticated testing, login flows, rules of engagement, and report filters. |
|
||||
| [AI providers](docs/ai-providers.md) | Selecting the model, the supported providers (Anthropic, OpenAI, xAI, AWS Bedrock), and custom gateways. |
|
||||
| [AI providers](docs/ai-providers.md) | Selecting the model, the supported providers (Anthropic, OpenAI, xAI, AWS Bedrock, and any other Pi-supported provider), and custom gateways. |
|
||||
| [Platforms and networking](docs/platforms.md) | Windows/WSL2, Linux, macOS, Docker networking, local apps, and custom hostnames. |
|
||||
| [Workspaces and resuming](docs/workspaces.md) | Naming workspaces, resuming interrupted scans, and workspace storage. |
|
||||
| [Safety and limitations](docs/safety.md) | Authorized-use requirements, non-production guidance, mutative effects, cost, and model caveats. |
|
||||
@@ -216,13 +225,13 @@ Important limitations:
|
||||
|
||||
- Shannon Open Source focuses on actively exploitable issues such as Injection, XSS, SSRF, Broken Authentication, and Broken Authorization. Broader static-analysis coverage, including vulnerable dependencies and insecure configurations, is delivered through the Keygraph platform.
|
||||
- Findings still require human review. LLM-generated reports can contain weakly supported or incorrect details.
|
||||
- Shannon is officially supported with Claude models. Smaller, alternative, or proxied non-Claude models may be incomplete or unstable.
|
||||
- Anthropic, OpenAI, xAI, and AWS Bedrock are built-in providers, and any Anthropic Messages API or OpenAI Chat Completions or Responses API endpoint works through a custom base URL. Model capability varies, and a model that does not follow Shannon's instructions or tool-use constraints reliably will produce weaker results.
|
||||
- A full run can take roughly 1 to 1.5 hours and may incur LLM API costs depending on model pricing and application complexity.
|
||||
- Do not scan untrusted or adversarial codebases. AI-powered tools that read source code can be exposed to prompt injection.
|
||||
|
||||
Read the full [Safety and limitations](docs/safety.md) guide before running Shannon in a new environment.
|
||||
|
||||
## License and Enterprise Licensing
|
||||
## License
|
||||
|
||||
Shannon Open Source is licensed under the [GNU Affero General Public License v3.0](LICENSE).
|
||||
|
||||
@@ -255,6 +264,40 @@ Stay connected:
|
||||
- [Twitter/X: @KeygraphHQ](https://twitter.com/KeygraphHQ)
|
||||
- [LinkedIn: Keygraph](https://linkedin.com/company/keygraph)
|
||||
|
||||
## Common Questions
|
||||
|
||||
### Is Shannon free?
|
||||
|
||||
Yes. Shannon Open Source is free and licensed under AGPL-3.0. You run it yourself from the command line. Your only cost is the AI provider credits you supply.
|
||||
|
||||
### Can I self-host Shannon?
|
||||
|
||||
Yes. Shannon Open Source runs entirely on your own infrastructure in an ephemeral Docker container. Your source code is mounted read-only and never leaves your environment.
|
||||
|
||||
### Does Shannon support bring your own key (BYOK)?
|
||||
|
||||
Yes, always. You provide the LLM credentials Shannon uses to run a pentest, in every deployment, open source and commercial. Keygraph never proxies your model traffic.
|
||||
|
||||
### Does Shannon output SARIF?
|
||||
|
||||
Yes. Shannon emits SARIF 2.1.0, the OASIS standard format for static analysis results, alongside structured JSON. Any SARIF consumer reads it: code scanning services, vulnerability management platforms, security dashboards, and CI/CD pipelines. Set `report.sarif` to `"true"` in your configuration file to enable the SARIF log.
|
||||
|
||||
### Which AI providers does Shannon support?
|
||||
|
||||
Anthropic, OpenAI, xAI, and AWS Bedrock are built in and configured directly by provider ID. Beyond those, Shannon runs on any endpoint that implements the Anthropic Messages API or the OpenAI Chat Completions or Responses API, reached through a custom base URL. The rule is the API format, not the vendor. Shannon uses a single unified model setting throughout a pentest.
|
||||
|
||||
### Can I run Shannon on a local or self-hosted model?
|
||||
|
||||
Technically yes, but it is not recommended. Shannon works with local models served through Ollama, vLLM, or LM Studio, which expose an OpenAI-compatible endpoint, as well as routers such as OpenRouter and gateways such as LiteLLM. Point Shannon at the endpoint with a custom base URL. Capability varies, and a model that does not follow Shannon's instructions or tool-use constraints reliably will produce weaker pentests than a frontier model, so take this path only if you know how your chosen model behaves. See [AI providers](docs/ai-providers.md#custom-base-url).
|
||||
|
||||
### Does Shannon actually exploit vulnerabilities, or just scan?
|
||||
|
||||
Shannon executes real exploits. It reports a finding only when it has produced a working proof-of-concept, and discards hypotheses it cannot prove. It is a pentester, not a scanner.
|
||||
|
||||
### Is Shannon free for startups and nonprofits?
|
||||
|
||||
Shannon Open Source is free for everyone. In addition, the Keygraph Community Program gives eligible nonprofits and early-stage startups free access to the commercial Keygraph platform. See [keygraph.io](https://keygraph.io).
|
||||
|
||||
<p align="center">
|
||||
<b>Built by <a href="https://keygraph.io">Keygraph</a></b>
|
||||
</p>
|
||||
@@ -531,12 +574,12 @@ rules:
|
||||
|
||||
| Key | Effect |
|
||||
| --- | --- |
|
||||
| `min_severity` | Drops findings rated below this severity. Applies only when `exploit` is `"true"`. |
|
||||
| `min_severity` | Drops findings rated below this severity. Applies in both exploitative and analysis-only runs. |
|
||||
| `min_confidence` | Drops findings rated below this confidence. Applies only when `exploit` is `"false"`. |
|
||||
| `guidance` | Free-text instruction to the report agent, such as which topics to exclude. |
|
||||
| `sarif` | Emits a SARIF 2.1.0 log alongside the Markdown report. Requires `exploit: "true"`. |
|
||||
|
||||
A finding carries one rating or the other, never both: an exploited finding is rated by severity, an analysis-only finding by confidence. Setting the threshold that does not apply to the run is ignored, and Shannon logs a warning naming the one to use instead.
|
||||
Every finding carries a severity, but it does not mean the same thing in each mode: an exploitative run measures severity from what the exploit demonstrated, while an analysis-only run assesses it from the class of flaw and the impact it would have. An analysis-only finding carries a confidence rating alongside its severity, since nothing was proven. Setting `min_confidence` on an exploitative run is ignored, and Shannon logs a warning naming the threshold to use instead.
|
||||
|
||||
### SARIF Output
|
||||
|
||||
@@ -550,7 +593,7 @@ report:
|
||||
|
||||
Each finding becomes one SARIF result, filed under a rule per vulnerability class (`shannon/injection`, `shannon/xss`, `shannon/auth`, `shannon/authz`, `shannon/ssrf`) and tagged with its OWASP Top Ten 2025 category. Results are anchored to the code location the analysis phase recorded, falling back to the HTTP entry point when the finding names no file. Severity maps onto SARIF's three levels: `critical` and `high` become `error`, `medium` becomes `warning`, everything else becomes `note`.
|
||||
|
||||
The log is written only for exploitative runs. An analysis-only run rates findings by confidence and produces no severity, so there is nothing to populate `level` with; `sarif` is ignored when `exploit` is `"false"`.
|
||||
The log is written only for exploitative runs. `sarif` is ignored when `exploit` is `"false"`.
|
||||
|
||||
Supported rule types include `url_path`, `subdomain`, `domain`, `method`, `header`, `parameter`, and `code_path`.
|
||||
|
||||
@@ -613,7 +656,7 @@ Shannon forwards only the selected provider's credential into the scan container
|
||||
|
||||
### Any other provider
|
||||
|
||||
Shannon accepts any provider and model present in the Pi harness catalogue. Browse them at [pi.dev/models](https://pi.dev/models). These are technically supported but not recommended. Claude models are best-supported (see the note below).
|
||||
Shannon accepts any provider and model present in the Pi harness catalogue. Browse them at [pi.dev/models](https://pi.dev/models).
|
||||
|
||||
```bash
|
||||
export SHANNON_AI_API_KEY=your-api-key # the provider's API key
|
||||
@@ -625,7 +668,7 @@ This path covers providers whose credential is a single API key. Providers that
|
||||
`npx @keygraph/shannon setup` exposes this as the **Other provider** option.
|
||||
|
||||
> [!IMPORTANT]
|
||||
> Claude models are the best-supported option. Shannon's evaluations, internal testing, and agent harness are tuned for Claude. Other models are permitted and validated against the harness catalogue, but may not follow Shannon's instructions or tool-use constraints as reliably. Use them at your own risk.
|
||||
> Models are validated against the harness catalogue, but capability varies. A model that does not follow Shannon's instructions or tool-use constraints reliably will produce weaker pentests. Evaluate the model you choose against your own targets before depending on its results.
|
||||
|
||||
## Cyber safeguards (do this before your first scan)
|
||||
|
||||
@@ -732,12 +775,55 @@ The variable is rejected in preflight where it cannot take effect: with a non-`o
|
||||
|
||||
`npx @keygraph/shannon setup` covers this under **Custom Base URL**, which asks which API your gateway serves and configures the matching provider for you.
|
||||
|
||||
## OpenAI Codex (ChatGPT Plus/Pro subscription)
|
||||
|
||||
A ChatGPT Plus or Pro Codex subscription can run Shannon. Shannon reuses a login created by Pi.
|
||||
|
||||
Before running a pentest, review the [cyber safeguards requirements](#cyber-safeguards-do-this-before-your-first-scan).
|
||||
|
||||
1. Install Pi by following the instructions at [pi.dev](https://pi.dev).
|
||||
2. Log in with your subscription using Pi's [subscription authentication guide](https://pi.dev/docs/latest/providers#subscriptions). This creates `~/.pi/agent/auth.json` with an `openai-codex` entry.
|
||||
|
||||
3. Select a Codex model and enable Pi authentication:
|
||||
|
||||
```bash
|
||||
export SHANNON_USE_PI_AUTH=1
|
||||
export SHANNON_AI_MODEL=openai-codex:gpt-5.5
|
||||
```
|
||||
|
||||
4. In npx mode, run `npx @keygraph/shannon start ...` from the same shell. In source-build mode, add the two variables to `.env` and run `./shannon start ...`.
|
||||
|
||||
Supported Codex models are `gpt-5.6-sol`, `gpt-5.5`, and `gpt-5.4`.
|
||||
|
||||
## Claude Code subscription
|
||||
|
||||
The latest version of Shannon does not support Claude Code subscriptions. The [`shannon-v1`](https://github.com/KeygraphHQ/shannon/tree/shannon-v1) branch is the final release built on the Claude Agent SDK and supports Claude Code OAuth.
|
||||
|
||||
Before running a pentest, review the [cyber safeguards requirements](#cyber-safeguards-do-this-before-your-first-scan).
|
||||
|
||||
1. Generate a Claude Code OAuth token:
|
||||
|
||||
```bash
|
||||
claude setup-token
|
||||
```
|
||||
|
||||
2. Run the setup flow for the final `shannon-v1` release:
|
||||
|
||||
```bash
|
||||
npx @keygraph/shannon@1.9.0 setup
|
||||
```
|
||||
|
||||
3. Select **OAuth Token** and enter the token generated by Claude Code.
|
||||
4. Start the pentest with `npx @keygraph/shannon@1.9.0 start ...`.
|
||||
|
||||
These instructions apply only to `shannon-v1`.
|
||||
|
||||
## Validation
|
||||
|
||||
Checks run before a scan starts, so mistakes fail immediately rather than partway through a run:
|
||||
|
||||
- **Provider and model ID** — validated against the Pi harness catalogue. An unknown provider or model ID fails preflight with a pointer to [pi.dev/models](https://pi.dev/models). A custom base URL exempts the model ID, since a gateway may serve its own names.
|
||||
- **Credential presence** — always validated for the selected provider.
|
||||
- **Credential presence** — validated for the selected provider, or read from Pi when `SHANNON_USE_PI_AUTH=1`.
|
||||
- **Credential validity** — one minimal request against the model the scan will use, so a rejected key, an exhausted quota, or a model the account cannot reach fails before any agent runs. Bedrock included: its bearer token and region go through the same probe.
|
||||
|
||||
## Migrating from the three-tier configuration
|
||||
@@ -942,7 +1028,7 @@ For maximum isolation, run Shannon inside a disposable virtual machine.
|
||||
## LLM and Automation Caveats
|
||||
|
||||
- **Verification is required**: Shannon uses a proof-by-exploitation methodology, but final reports can still contain weakly supported or incorrect details. Human review is essential.
|
||||
- **Model support**: Shannon is officially supported only with Claude models. Alternative models may be incomplete, inaccurate, or unstable.
|
||||
- **Model support**: results vary by model. A model that does not follow Shannon's instructions or tool-use constraints reliably may produce incomplete, inaccurate, or unstable runs.
|
||||
- **Prompt injection risk**: Do not point Shannon at untrusted or adversarial codebases. AI-powered tools that read source code can be influenced by malicious repository content.
|
||||
|
||||
## Scope of Analysis
|
||||
@@ -963,7 +1049,6 @@ For broader coverage, the Keygraph platform adds black-box and white-box agentic
|
||||
|
||||
A full test run typically takes roughly 1 to 1.5 hours. LLM API costs vary by model pricing, target complexity, selected provider, and concurrency.
|
||||
|
||||
|
||||
---
|
||||
|
||||
# File: docs/coverage-roadmap.md
|
||||
|
||||
@@ -6,7 +6,7 @@ Use this file as the concise entry point for AI agents and LLMs reading this rep
|
||||
|
||||
## Start Here
|
||||
|
||||
- [README](README.md): Main project overview, editions, quick start, Shannon capabilities, Keygraph platform positioning, safety notes, licensing, and support links.
|
||||
- [README](README.md): Main project overview, editions, quick start, Shannon capabilities, Keygraph platform positioning, common questions, safety notes, licensing, and support links.
|
||||
- [Full Combined Context](llms-full.txt): README and documentation combined into one file for agents that need maximum local context.
|
||||
|
||||
## Shannon
|
||||
|
||||