mirror of
https://github.com/KeygraphHQ/shannon.git
synced 2026-10-02 22:36:49 +02:00
Compare commits
8
Commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
e90cdb5424 | ||
|
|
8d732f9fa2 | ||
|
|
a46adf1c59 | ||
|
|
f926c5ae69 | ||
|
|
c81553271f | ||
|
|
ad2d069563 | ||
|
|
a14c7944d8 | ||
|
|
57c511ff8e |
No files matched your search
+3
-3
@@ -13,7 +13,7 @@ SHANNON_AI_MODEL=anthropic:claude-sonnet-4-6
|
||||
|
||||
# --- xAI ---------------------------------------------------------------------
|
||||
# SHANNON_AI_API_KEY=your-api-key-here
|
||||
# SHANNON_AI_MODEL=xai:grok-4.5
|
||||
# SHANNON_AI_MODEL=xai:grok-4.7
|
||||
|
||||
# --- AWS Bedrock -------------------------------------------------------------
|
||||
# Bearer token only; model must be enabled in your region.
|
||||
@@ -48,9 +48,9 @@ SHANNON_AI_MODEL=anthropic:claude-sonnet-4-6
|
||||
# See the guide below to use an OpenAI subscription
|
||||
# https://github.com/KeygraphHQ/shannon/blob/main/docs/ai-providers.md#openai-codex-chatgpt-pluspro-subscription
|
||||
# SHANNON_USE_PI_AUTH=1
|
||||
# SHANNON_AI_MODEL=openai-codex:gpt-5.5
|
||||
# SHANNON_AI_MODEL=openai-codex:gpt-6-sol
|
||||
|
||||
# Or the guide below to use an xAI subscription
|
||||
# https://github.com/KeygraphHQ/shannon/blob/main/docs/ai-providers.md#xai-grok-subscription
|
||||
# SHANNON_USE_PI_AUTH=1
|
||||
# SHANNON_AI_MODEL=xai:grok-4.6
|
||||
# SHANNON_AI_MODEL=xai:grok-4.7
|
||||
@@ -87,7 +87,7 @@ pnpm biome:fix # Auto-fix lint, format, and import sorting
|
||||
|
||||
**Monorepo tooling:** pnpm workspaces, Turborepo for task orchestration, Biome for linting/formatting. TypeScript compiler options shared via `tsconfig.base.json` at the root. All packages extend it, overriding only `rootDir` and `outDir`. Shared devDependencies (`typescript`, `@types/node`, `turbo`, `@biomejs/biome`) are hoisted to the root workspace.
|
||||
|
||||
**Options:** `-c <file>` (YAML config), `--models-config <file>` (pi `models.json` defining models pi's catalogue lacks), `-o <path>` (output directory), `-w <name>` (named workspace; auto-resumes if exists), `--pipeline-testing` (minimal prompts, 10s retries), `--keep-container` (preserve worker container after exit for log inspection), `--yes`/`-y` (skip the confirmation prompt on `stop`; required for non-interactive use; `reset` requires a typed `confirm` and cannot be skipped)
|
||||
**Options:** `-c <file>` (YAML config), `--models-config <file>` (pi `models.json` defining models pi's catalogue lacks), `-o <path>` (output directory), `-w <name>` (named workspace; auto-resumes if exists), `--validate-auth` (run preflight and auth validation only, then stop; no pentest or report; requires a fresh workspace and an `authentication` block in the config), `--pipeline-testing` (minimal prompts, 10s retries), `--keep-container` (preserve worker container after exit for log inspection), `--yes`/`-y` (skip the confirmation prompt on `stop`; required for non-interactive use; `reset` requires a typed `confirm` and cannot be skipped)
|
||||
|
||||
## Architecture
|
||||
|
||||
@@ -165,7 +165,7 @@ Around those phases:
|
||||
- **Configuration** — YAML configs in `apps/worker/configs/` use the closed JSON Schema in `config-schema.json`. Every fresh scan runs the fixed five analysis classes; there is no public class selector. `agentic_sast.enabled` is the only public agentic-SAST setting. Finding reconciliation runs on every scan and has no public setting of its own. Config also supports authentication (MFA/TOTP), URL/code rule scoping (`rules.avoid`/`rules.focus`), `exploit`, free-form `rules_of_engagement`, and post-hoc `report` options (`min_severity`, `min_confidence`, `guidance`, and exploit-only `sarif` output via `apps/worker/src/services/sarif-renderer.ts`, on by default for exploit runs and opt out with `report.sarif: "false"`). `code_path` avoid rules are enforced via the `@gotgenes/pi-permission-system` extension: `apps/worker/src/temporal/activities.ts:syncCodePathDenyRules` writes a global `path` deny config once per workflow (`apps/worker/src/ai/pi/permission-system.ts:syncPermissionSystemConfig`), and the executor loads the extension when that config is present (`apps/worker/src/ai/pi/pi-executor.ts`), so denies fire across every tool and child `task` session. Credential resolution — local mode: env vars → `./.env`; npx mode: env vars → `~/.shannon/config.toml` (via `npx @keygraph/shannon setup`)
|
||||
- **Agentic SAST progress** — Capella runs as a child workflow, so its activities are absent from the parent's `pendingActivities` and invisible to the CLI. The child signals each stage boundary up via `capellaStageProgress` (`apps/worker/src/temporal/shared.ts`); the parent's handler validates the payload and writes the child-supplied `startedAt` and `durationMs` directly to `operationalStages['agentic-sast:<stage>']`, so both the live `getProgress` query and the terminal result carry per-stage rows. Signalling is best-effort and every failure is swallowed — a closed or unreachable parent must never fail a SAST run. `CAPELLA_STAGE_LABELS` in `apps/worker/src/ai/sast/types.ts` is the one label table, shared by the scan log and the status tree; `CAPELLA_PROGRESS_STAGES` omits `export`, which runs no model and so never becomes a row. Scans predating the signal keep the aggregate `agentic-sast` span and render as a bare phase line
|
||||
- **Prompts** — Per-phase templates in `apps/worker/prompts/` with variable substitution (`{{TARGET_URL}}`, `{{CONFIG_CONTEXT}}`). Shared partials in `apps/worker/prompts/shared/` via `apps/worker/src/services/prompt-manager.ts`, including `_code-path-rules.txt` (focus/avoid `[FILE]`/`[GLOB]` routing) and `_rules-of-engagement.txt` (free-text engagement rules). When `exploit: false`, `apps/worker/src/services/findings-renderer.ts` deterministically converts each `*_exploitation_queue.json` into a `*_findings.md` for report assembly — no LLM in the loop
|
||||
- **Agent Harness (pi)** — Uses the **pi harness** (`@earendil-works/pi-coding-agent`, requires Node ≥ 22.19) via `apps/worker/src/ai/pi/pi-executor.ts` (`runPiPrompt` → `createAgentSession`). Retry is split in `apps/worker/src/ai/pi/retry-settings.ts`: pi's agent-level loop is off so Temporal owns agent restarts, while `provider.maxRetries` stays on — pi reads the `provider` block independently of the `enabled` flag — so transport faults are absorbed in-session rather than costing a full agent re-run. `maxRetryDelayMs` is left at pi's 60s default. One model runs every phase, named by `SHANNON_AI_MODEL=<provider>:<model-id>` (default `anthropic:claude-sonnet-4-6`). `apps/worker/src/ai/models.ts` parses the spec — splitting on the **first** colon only, so Bedrock IDs keep theirs — and resolves it through pi's `ModelRuntime`. pi ships the `CredentialStore` interface but no in-memory implementation (its own reads `auth.json` from disk), so `RuntimeCredentialStore` in that file supplies one: credentials arrive as env vars in an ephemeral container and must never touch disk. `createModelRuntime(providerId, apiKey)` builds the runtime with `allowModelNetwork: true`, so `ModelRuntime.create()` refreshes the model catalogue over the network at scan start and a freshly released model resolves without a `--models-config` file. The fetch is bounded (10s) and falls back to the static catalogue on timeout, so an unreachable catalogue endpoint cannot hang the scan. The refresh does not override a `--models-config`: pi reloads and re-applies that file as a config overlay on every refresh (it reloads `this.config` at the top of `refresh()`), so custom definitions still win over the fetched catalogue; the merge semantics below are unchanged, just layered over a fresher base. `resolveModelSelection()` is **async** because `ModelRuntime.create()` is. Any pi-ai provider id is accepted — `parseModelSpec` no longer rejects against a hardcoded list, so pi's registry is the authority (an unknown provider/model surfaces as a clear "not found in pi registry" error at preflight, which points to the browsable catalogue at `pi.dev/models` — `PI_CATALOG_URL` in `apps/worker/src/ai/models.ts`, appended to the not-found errors and shown in the setup wizard's "Other provider" hint). Four providers are **curated** (`CURATED_PROVIDERS`: `anthropic`, `openai`, `xai`, `amazon-bedrock`) with their own credential variables, config sections, and setup flows; each provider's API key env var is declared once in `PROVIDER_API_KEY_ENV` — Shannon uses each vendor's own variable name (`OPENAI_API_KEY`, `XAI_API_KEY`, …), never an invented one; Bedrock's entry is `AWS_BEARER_TOKEN_BEDROCK`, paired with `AWS_REGION`, which preflight requires separately as provider config rather than a credential. Any other provider uses the **generic** credential path: `SHANNON_AI_API_KEY` (`GENERIC_API_KEY_ENV`) supplies the key for any provider whose credential is a plain API key. Curated providers' own variables take precedence over it, and it also works as a fallback for them — Bedrock is the sole exception (it authenticates through its AWS_ variables, so the generic key never stands in for it). The CLI forwards `SHANNON_AI_API_KEY` in `COMMON_FORWARD_VARS` (it is provider-neutral, binding to whatever `SHANNON_AI_MODEL` names, so the "only one provider configured" guard counts only named credentials), and stores it under a generic `[provider]` config.toml section (`provider.api_key`). `npx @keygraph/shannon setup` exposes this as the "Other provider" option: free-text provider id + model id + key (a curated provider id is rejected there, since it has its own option). A model pi's catalogue does not carry, such as a self-hosted model, is reachable without an SDK bump: `--models-config <file>` mounts a pi `models.json` read-only at `/app/models.json`. The mount is the entire CLI→worker protocol: nothing is forwarded through the environment, and `modelsConfigPath()` detects the file at that fixed path, exactly as `piAuthPresent()` detects the pi auth mount whose flag is likewise not forwarded (`MODELS_CONFIG_CONTAINER_PATH` in the CLI and `MODELS_CONFIG_PATH` in `apps/worker/src/paths.ts` must stay in sync). `createModelRuntime` always names `modelsPath` explicitly — the mounted path, or **`null` when no config was supplied**, which switches models.json off outright. It is never left to pi's default of `<agent dir>/models.json`, because that dir is shared with the pi auth mount, so a file landing there must not silently contribute model definitions to a scan that did not ask for one. `modelsStorePath` is pinned to the agent dir alongside it, since pi otherwise derives it from `dirname(modelsPath)` and would try to write beside a read-only mount. Custom definitions merge over the built-in catalogue: a matching model id replaces the built-in entry, a new id is added alongside, and `modelOverrides` adjusts a built-in without replacing the proviLine truncated
|
||||
- **Agent Harness (pi)** — Uses the **pi harness** (`@earendil-works/pi-coding-agent`, requires Node ≥ 22.19) via `apps/worker/src/ai/pi/pi-executor.ts` (`runPiPrompt` → `createAgentSession`). Retry is split in `apps/worker/src/ai/pi/retry-settings.ts`: pi's agent-level loop is off so Temporal owns agent restarts, while `provider.maxRetries` stays on — pi reads the `provider` block independently of the `enabled` flag — so transport faults are absorbed in-session rather than costing a full agent re-run. `maxRetryDelayMs` is left at pi's 60s default. One model runs every phase, named by `SHANNON_AI_MODEL=<provider>:<model-id>` (default `anthropic:claude-sonnet-4-6`). `apps/worker/src/ai/models.ts` parses the spec — splitting on the **first** colon only, so Bedrock IDs keep theirs — and resolves it through pi's `ModelRuntime`. pi ships the `CredentialStore` interface but no in-memory implementation (its own reads `auth.json` from disk), so `RuntimeCredentialStore` in that file supplies one: credentials arrive as env vars in an ephemeral container and must never touch disk. `createModelRuntime(providerId, apiKey)` builds the runtime with `allowModelNetwork: true`, so `ModelRuntime.create()` refreshes the model catalogue over the network at scan start and a freshly released model resolves without a `--models-config` file. The fetch is bounded (10s) and falls back to the static catalogue on timeout, so an unreachable catalogue endpoint cannot hang the scan. The refresh does not override a `--models-config`: pi reloads and re-applies that file as a config overlay on every refresh (it reloads `this.config` at the top of `refresh()`), so custom definitions still win over the fetched catalogue; the merge semantics below are unchanged, just layered over a fresher base. `resolveModelSelection()` is **async** because `ModelRuntime.create()` is. Any pi-ai provider id is accepted — `parseModelSpec` no longer rejects against a hardcoded list, so pi's registry is the authority (an unknown provider/model surfaces as a clear "not found in pi registry" error at preflight, which points to the browsable catalogue at `pi.dev/models` — `PI_CATALOG_URL` in `apps/worker/src/ai/models.ts`, appended to the not-found errors and shown in the setup wizard's "Other provider" hint). Four providers are **curated** (`CURATED_PROVIDERS`: `anthropic`, `openai`, `xai`, `amazon-bedrock`) with their own credential variables, config sections, and setup flows; each provider's API key env var is declared once in `PROVIDER_API_KEY_ENV` — Shannon uses each vendor's own variable name (`OPENAI_API_KEY`, `XAI_API_KEY`, …), never an invented one; Bedrock's entry is `AWS_BEARER_TOKEN_BEDROCK`, paired with `AWS_REGION`, which preflight requires separately as provider config rather than a credential. Any other provider uses the **generic** credential path: `SHANNON_AI_API_KEY` (`GENERIC_API_KEY_ENV`) supplies the key for any provider whose credential is a plain API key. Curated providers' own variables take precedence over it, and it also works as a fallback for them — Bedrock is the sole exception (it authenticates through its AWS_ variables, so the generic key never stands in for it). The CLI forwards `SHANNON_AI_API_KEY` in `COMMON_FORWARD_VARS` (it is provider-neutral, binding to whatever `SHANNON_AI_MODEL` names, so the "only one provider configured" guard counts only named credentials), and stores it under a generic `[provider]` config.toml section (`provider.api_key`). `npx @keygraph/shannon setup` exposes this as the "Other provider" option: free-text provider id + model id + key (a curated provider id is rejected there, since it has its own option). A model pi's catalogue does not carry, such as a self-hosted model, is reachable without an SDK bump: `--models-config <file>` mounts a pi `models.json` read-only at `/app/models.json`. The mount is the entire CLI→worker protocol: nothing is forwarded through the environment, and `modelsConfigPath()` detects the file at that fixed path, exactly as `piAuthPresent()` detects the pi auth mount whose flag is likewise not forwarded (`MODELS_CONFIG_CONTAINER_PATH` in the CLI and `MODELS_CONFIG_PATH` in `apps/worker/src/paths.ts` must stay in sync). `createModelRuntime` always names `modelsPath` explicitly — the mounted path, or **`null` when no config was supplied**, which switches models.json off outright. It is never left to pi's default of `<agent dir>/models.json`, because that dir is shared with the pi auth mount, so a file landing there must not silently contribute model definitions to a scan that did not ask for one. `modelsStorePath` is pinned to the agent dir alongside it, since pi otherwise derives it from `dirname(modelsPath)` and would try to write beside a read-only mount. Custom definitions merge over the built-in catalogue: a matching model id replaces the built-in entry, a new id is added alongside, and `modelOverrides` adjusts a built-in without replacing the proviLine truncated
|
||||
- **Pi Credential Reuse** — `SHANNON_USE_PI_AUTH=1` opts into reusing the host's Pi login, including an `openai-codex` ChatGPT Plus/Pro subscription (`SHANNON_AI_MODEL=openai-codex:<model-id>`) or an `xai` Grok subscription (`SHANNON_AI_MODEL=xai:<model-id>`); the mechanism is provider-agnostic and works for any Pi login. `apps/cli/src/env.ts` requires `~/.pi/agent/auth.json`; `start.ts` passes its path to `spawnWorker`, which mounts only that file read-write at `/tmp/.pi/agent/auth.json`. The flag itself is not forwarded: the worker detects the file with `piAuthPresent()` and passes its path to `ModelRuntime.create`. CLI and worker API-key presence checks are skipped on this path, but the normal preflight model probe still validates the credential. The image and UID-remapping entrypoint keep `/tmp/.pi/agent` owned by `pentest` so adjacent Pi/Shannon configuration remains writable. Refreshed OAuth state is persisted to the host for subsequent scans.
|
||||
- **Audit System** — Crash-safe append-only logging in `workspaces/{hostname}_{sessionId}/`. The run directory's top level holds the human-facing report in both formats (`Security-Assessment-Report.pdf` and `Security-Assessment-Report.md`, `FINAL_REPORT_PDF_FILENAME`/`FINAL_REPORT_MD_FILENAME` in `apps/worker/src/paths.ts`); everything else — deliverables, per-agent logs, prompts, `session.json`, `workflow.log`, and browser artifacts — is nested under a hidden `.shannon/` internals dir (`INTERNAL_DIR`) so a customer sees only the report. Audit path helpers route through `generateInternalPath` (`apps/worker/src/audit/utils.ts`); the CLI nests the overlay backing dirs under the same `.shannon/` (`apps/cli/src/docker.ts`, `start.ts`). `session.json`/`workflow.log` reads use dual-read resolvers (`resolveSessionJsonPath`, `resolveRunFile`) that prefer `.shannon/` and fall back to the legacy run-root layout, so pre-restructure workspaces stay listable (`scans`/`logs`) without migration. A pre-restructure workspace cannot be resumed: `classifyWorkspaceLaunch` (`apps/cli/src/commands/start.ts`) requires `.shannon/launch.json`, and its absence fails the launch as "created by an earlier version of Shannon" before anything on disk is touched. There is no in-place migration — the workspace's files and report are left untouched, and the operator starts a new scan under a different `-w` name. The report agent writes structured findings to `report.json`, from which `report-renderer.ts` renders the assembled markdown and `report-json-adapter.ts` produces the Typst-shaped JSON that `pdf-renderer.ts` compiles into `comprehensive_security_assessment_report.pdf` using the bundled `apps/worker/templates/typst/report.typ` template (the `typst` binary is installed in the worker image). `copyReportToRunRoot` (`apps/worker/src/services/reporting.ts`) surfaces both the PDF and the markdown to the run root as `Security-Assessment-Report.pdf` and `Security-Assessment-Report.md`; the deliverables-dir copies remain as the git-checkpointed sources. PDF compilation is best-effort — a failure is logged and the run still completes. WorkflowLogger (`apps/worker/src/audit/workflow-logger.ts`) provides unified human-readable per-workflow logs, backed by LogStream (`apps/worker/src/audit/log-stream.ts`) shared stream primitive. Every combined-log line is also projected into a per-agent file under `.shannon/agents/<slug>.log` (one per pipeline agent, one per Capella stage; subagents fold into the parent's file, and a stage's concurrent sessions share its file with an inline session label). The projection boundary is `apps/worker/src/audit/actor-projection.ts` (`projectActor` maps a `TraceActor` to its combined prefix and owning file slug — slugs come only from closed fields); fan-out is best-effort and never blocks the canonical combined log. A lifecycle owner holds a `LogStream` lease per agent file (the pipeline agent's `logAgent` span, or a Capella stage activity's `try/finally`) so per-line writes ride the reference count; `CapellaStageTrace.drain()` flushes a stage's trace queue before its activity returns. The CLI tails one file with `shannon logs --agent <name>` (`--list-agents` to enumerate); the default `shannon logs` path is unchanged
|
||||
- **Deliverables** — Saved to `.shannon/deliverables/` in the target repo via the `save-deliverable` CLI script (`apps/worker/src/scripts/save-deliverable.ts`)
|
||||
|
||||
@@ -21,7 +21,15 @@ import { resolveWorkflowId } from '../session.js';
|
||||
import { waitForWorkflowClose } from '../temporal-client.js';
|
||||
import { stdoutIsTerminal } from '../tty.js';
|
||||
|
||||
const TERMINAL_HEADINGS = new Set(['Scan COMPLETED', 'Scan PARTIAL', 'Scan FAILED', 'Scan CANCELLED']);
|
||||
const TERMINAL_HEADINGS = new Set([
|
||||
'Scan COMPLETED',
|
||||
'Scan PARTIAL',
|
||||
'Scan FAILED',
|
||||
'Scan CANCELLED',
|
||||
'Validation COMPLETED',
|
||||
'Validation FAILED',
|
||||
'Validation CANCELLED',
|
||||
]);
|
||||
|
||||
// The combined log resets completion on the bare `RESUMED` heading; a per-agent file carries the
|
||||
// distinct `--- RESUMED (<workflow id>) ---` boundary that WorkflowLogger.logResumeBoundary writes
|
||||
@@ -48,7 +56,7 @@ export class LogCompletionState {
|
||||
this.failureIsLastMarker = false;
|
||||
} else if (TERMINAL_HEADINGS.has(line)) {
|
||||
this.terminalIsLastMarker = true;
|
||||
this.failureIsLastMarker = line === 'Scan FAILED';
|
||||
this.failureIsLastMarker = line.endsWith('FAILED');
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
@@ -36,17 +36,24 @@ const GATEWAY_DIALECTS: readonly {
|
||||
|
||||
/** Suggested models per curated provider, best-first. Free-text entry accepts any model in the provider's catalogue. */
|
||||
const MODEL_SUGGESTIONS: Readonly<Record<CuratedProviderId, readonly string[]>> = {
|
||||
anthropic: ['claude-sonnet-4-6', 'claude-opus-4-8', 'claude-opus-4-7', 'claude-haiku-4-5-20251001'],
|
||||
openai: ['gpt-5.6-sol', 'gpt-5.5', 'gpt-5.4'],
|
||||
xai: ['grok-4.5'],
|
||||
anthropic: [
|
||||
'claude-sonnet-5',
|
||||
'claude-opus-5',
|
||||
'claude-sonnet-4-6',
|
||||
'claude-opus-4-8',
|
||||
'claude-opus-4-7',
|
||||
'claude-haiku-4-5-20251001',
|
||||
],
|
||||
openai: ['gpt-6-sol', 'gpt-5.6-sol', 'gpt-5.5', 'gpt-5.4'],
|
||||
xai: ['grok-4.7'],
|
||||
'amazon-bedrock': ['us.anthropic.claude-sonnet-4-6', 'us.anthropic.claude-opus-4-8', 'us.anthropic.claude-opus-4-7'],
|
||||
};
|
||||
|
||||
/** Placeholder shown in the free-text model ID prompt, per curated provider. */
|
||||
const MODEL_ID_PLACEHOLDER: Readonly<Record<CuratedProviderId, string>> = {
|
||||
anthropic: 'claude-sonnet-4-6',
|
||||
openai: 'gpt-5.6-sol',
|
||||
xai: 'grok-4.5',
|
||||
openai: 'gpt-6-sol',
|
||||
xai: 'grok-4.7',
|
||||
'amazon-bedrock': 'us.anthropic.claude-opus-4-8',
|
||||
};
|
||||
|
||||
|
||||
@@ -45,6 +45,7 @@ export interface StartArgs {
|
||||
pipelineTesting: boolean;
|
||||
keepContainer: boolean;
|
||||
follow: boolean;
|
||||
authOnly: boolean;
|
||||
version: string;
|
||||
}
|
||||
|
||||
@@ -60,6 +61,8 @@ const FIXED_CLASSES = ['injection', 'xss', 'auth', 'authz', 'ssrf'] as const;
|
||||
interface LaunchState {
|
||||
readonly schema_version: typeof LAUNCH_STATE_SCHEMA_VERSION;
|
||||
readonly customer_output_path?: string;
|
||||
/** True when the workspace was created by an auth-validation run; such a workspace is not a scan. */
|
||||
readonly auth_only?: boolean;
|
||||
}
|
||||
|
||||
export interface WorkspaceLaunchDecision {
|
||||
@@ -124,17 +127,22 @@ function readLaunchState(filePath: string): LaunchState {
|
||||
if (!isRecord(value)) fail(NEWER_RELEASE_MESSAGE);
|
||||
// Unknown keys mean a newer release wrote this workspace; refuse rather than half-read it.
|
||||
const keys = Object.keys(value).sort();
|
||||
const keysAreValid = keys.every((key) => key === 'customer_output_path' || key === 'schema_version');
|
||||
const keysAreValid = keys.every(
|
||||
(key) => key === 'auth_only' || key === 'customer_output_path' || key === 'schema_version',
|
||||
);
|
||||
const customerPath = value.customer_output_path;
|
||||
const pathIsValid =
|
||||
customerPath === undefined ||
|
||||
(typeof customerPath === 'string' && path.isAbsolute(customerPath) && path.resolve(customerPath) === customerPath);
|
||||
if (value.schema_version !== LAUNCH_STATE_SCHEMA_VERSION || !keysAreValid || !pathIsValid) {
|
||||
const authOnly = value.auth_only;
|
||||
const authOnlyIsValid = authOnly === undefined || typeof authOnly === 'boolean';
|
||||
if (value.schema_version !== LAUNCH_STATE_SCHEMA_VERSION || !keysAreValid || !pathIsValid || !authOnlyIsValid) {
|
||||
fail(NEWER_RELEASE_MESSAGE);
|
||||
}
|
||||
return {
|
||||
schema_version: LAUNCH_STATE_SCHEMA_VERSION,
|
||||
...(typeof customerPath === 'string' && { customer_output_path: customerPath }),
|
||||
...(authOnly === true && { auth_only: true }),
|
||||
};
|
||||
}
|
||||
|
||||
@@ -149,6 +157,7 @@ export function classifyWorkspaceLaunch(
|
||||
workspacePath: string,
|
||||
expectedUrl: string,
|
||||
requestedOutputDir: string | undefined,
|
||||
requestedAuthOnly: boolean,
|
||||
): WorkspaceLaunchDecision {
|
||||
const sessionPath = resolveRunFile(workspacePath, 'session.json');
|
||||
const sessionExists = fs.existsSync(sessionPath);
|
||||
@@ -163,6 +172,11 @@ export function classifyWorkspaceLaunch(
|
||||
|
||||
const launchPath = path.join(workspacePath, INTERNAL_DIR, LAUNCH_STATE_FILENAME);
|
||||
const launch = readLaunchState(launchPath);
|
||||
if (launch.auth_only && !requestedAuthOnly) {
|
||||
fail(
|
||||
'This workspace was created to validate authentication only, so it cannot be run as a scan. Start a new scan with a different -w name.',
|
||||
);
|
||||
}
|
||||
const session = readJsonFile(sessionPath);
|
||||
if (!isRecord(session) || !isRecord(session.session) || session.session.webUrl !== expectedUrl) {
|
||||
fail(
|
||||
@@ -190,12 +204,17 @@ export function classifyWorkspaceLaunch(
|
||||
* host crash. Callers invoke this only for a fresh workspace; an existing launch.json is
|
||||
* the resume contract and must never be replaced.
|
||||
*/
|
||||
export function writeLaunchStateAtomically(internalPath: string, outputDir: string | undefined): void {
|
||||
export function writeLaunchStateAtomically(
|
||||
internalPath: string,
|
||||
outputDir: string | undefined,
|
||||
authOnly: boolean,
|
||||
): void {
|
||||
const finalPath = path.join(internalPath, LAUNCH_STATE_FILENAME);
|
||||
const temporaryPath = path.join(internalPath, `${LAUNCH_STATE_FILENAME}.tmp-${process.pid}-${randomSuffix()}`);
|
||||
const launchState: LaunchState = {
|
||||
schema_version: LAUNCH_STATE_SCHEMA_VERSION,
|
||||
...(outputDir !== undefined && { customer_output_path: outputDir }),
|
||||
...(authOnly && { auth_only: true }),
|
||||
};
|
||||
const descriptor = fs.openSync(temporaryPath, 'wx', 0o600);
|
||||
try {
|
||||
@@ -225,6 +244,9 @@ export function createWorkflowId(workspace: string, isResume: boolean, timestamp
|
||||
}
|
||||
|
||||
export async function start(args: StartArgs): Promise<void> {
|
||||
// Auth-only runs are short and have no report to come back for, so they always stream to the end.
|
||||
if (args.authOnly) args.follow = true;
|
||||
|
||||
// 1. Resolve non-mutating inputs and classify the workspace before changing it.
|
||||
initHome();
|
||||
loadEnv();
|
||||
@@ -240,7 +262,14 @@ export async function start(args: StartArgs): Promise<void> {
|
||||
args.workspace ?? `${new URL(args.url).hostname.replace(/[^a-zA-Z0-9-]/g, '-')}_shannon-${Date.now()}`;
|
||||
const workspacePath = path.join(workspacesDir, workspace);
|
||||
const requestedOutputDir = args.output ? path.resolve(expandHome(args.output)) : undefined;
|
||||
const launchDecision = classifyWorkspaceLaunch(workspacePath, args.url, requestedOutputDir);
|
||||
const launchDecision = classifyWorkspaceLaunch(workspacePath, args.url, requestedOutputDir, args.authOnly);
|
||||
|
||||
// Auth-only runs write no resumable state, so they always run fresh; reusing a workspace would resume it.
|
||||
if (args.authOnly && launchDecision.isResume) {
|
||||
fail(
|
||||
'An auth-validation run needs a fresh workspace. Omit -w to auto-name one, or choose a -w name that is not in use.',
|
||||
);
|
||||
}
|
||||
|
||||
// 2. Inputs are valid; identify the run before initializing shared infrastructure.
|
||||
const bannerVersion = isLocal() ? undefined : args.version;
|
||||
@@ -254,7 +283,7 @@ export async function start(args: StartArgs): Promise<void> {
|
||||
ensureDocker();
|
||||
ensureImage(args.version);
|
||||
const spinner = p.spinner();
|
||||
spinner.start('Starting scan');
|
||||
spinner.start(args.authOnly ? 'Starting authentication validation' : 'Starting scan');
|
||||
await ensureInfra(spinner);
|
||||
|
||||
// 3. Generate the invocation identity.
|
||||
@@ -277,7 +306,7 @@ export async function start(args: StartArgs): Promise<void> {
|
||||
fs.chmodSync(dirPath, 0o777);
|
||||
}
|
||||
if (!launchDecision.isResume) {
|
||||
writeLaunchStateAtomically(internalPath, launchDecision.outputDir);
|
||||
writeLaunchStateAtomically(internalPath, launchDecision.outputDir, args.authOnly);
|
||||
}
|
||||
|
||||
// 5. Pre-create overlay mount points (:ro mounts cannot create them).
|
||||
@@ -336,6 +365,7 @@ export async function start(args: StartArgs): Promise<void> {
|
||||
workspace,
|
||||
...(args.pipelineTesting && { pipelineTesting: true }),
|
||||
...(args.keepContainer && { keepContainer: true }),
|
||||
...(args.authOnly && { authOnly: true }),
|
||||
...(shouldUsePiAuth() && { piAuthHostPath: resolveHostPiAuthPath() }),
|
||||
});
|
||||
|
||||
@@ -386,7 +416,7 @@ export async function start(args: StartArgs): Promise<void> {
|
||||
});
|
||||
|
||||
// Poll for the workflow to register in session.json; the spinner resolves once it does.
|
||||
spinner.message('Waiting for the scan to start');
|
||||
spinner.message(args.authOnly ? 'Waiting for authentication validation to start' : 'Waiting for the scan to start');
|
||||
for (let attempts = 0; attempts < 60; attempts++) {
|
||||
// A pre-workflow failure leaves its reason here (nothing reached Temporal); surface it
|
||||
// rather than polling out to a generic timeout.
|
||||
@@ -420,15 +450,15 @@ export async function start(args: StartArgs): Promise<void> {
|
||||
spinner.message('Running preflight checks');
|
||||
const outcome = await awaitPreflightOutcome(workflowId);
|
||||
if (outcome.kind === 'failed') {
|
||||
spinner.error('The scan could not start');
|
||||
spinner.error(args.authOnly ? 'Authentication validation could not start' : 'The scan could not start');
|
||||
printScanStartFailure(outcome.message);
|
||||
process.exit(1);
|
||||
}
|
||||
|
||||
spinner.stop(`Scan started — ${workspace}`);
|
||||
spinner.stop(args.authOnly ? `Validating authentication — ${workspace}` : `Scan started — ${workspace}`);
|
||||
printInfo(args, workspace, repo.hostPath, workspacesDir);
|
||||
if (args.follow) {
|
||||
await followScan(workspace, workspacesDir);
|
||||
await followScan(workspace, workspacesDir, args.authOnly);
|
||||
}
|
||||
return;
|
||||
}
|
||||
@@ -576,7 +606,7 @@ function printUnconfirmedScanHint(workspace: string, taskQueue: string, containe
|
||||
* That tracks whether the pipeline ran, not whether vulnerabilities were found. On failure the
|
||||
* root-cause message is printed so a red CI build says why.
|
||||
*/
|
||||
async function followScan(workspace: string, workspacesDir: string): Promise<never> {
|
||||
async function followScan(workspace: string, workspacesDir: string, authOnly = false): Promise<never> {
|
||||
const logFile = resolveRunFile(path.join(workspacesDir, workspace), 'workflow.log');
|
||||
const workflowId = resolveWorkflowId(workspace);
|
||||
|
||||
@@ -587,7 +617,8 @@ async function followScan(workspace: string, workspacesDir: string): Promise<nev
|
||||
}
|
||||
|
||||
if (stdoutIsTerminal()) {
|
||||
console.error('\n Following scan log (Ctrl-C to stop watching):\n');
|
||||
const what = authOnly ? 'validation' : 'scan';
|
||||
console.error(`\n Following ${what} log (Ctrl-C to stop watching):\n`);
|
||||
}
|
||||
|
||||
let temporalUnreachable = false;
|
||||
@@ -675,10 +706,12 @@ function printInfo(args: StartArgs, workspace: string, repoPath: string, workspa
|
||||
console.log(` Progress: ${prefix} status ${workspace}`);
|
||||
}
|
||||
|
||||
console.log('');
|
||||
console.log(' Report (when the scan finishes):');
|
||||
console.log(` ${reportDir}${path.sep}`);
|
||||
console.log(` ${FINAL_REPORT_PDF_FILENAME}`);
|
||||
console.log(` ${FINAL_REPORT_MD_FILENAME}`);
|
||||
console.log('');
|
||||
if (!args.authOnly) {
|
||||
console.log('');
|
||||
console.log(' Report (when the scan finishes):');
|
||||
console.log(` ${reportDir}${path.sep}`);
|
||||
console.log(` ${FINAL_REPORT_PDF_FILENAME}`);
|
||||
console.log(` ${FINAL_REPORT_MD_FILENAME}`);
|
||||
console.log('');
|
||||
}
|
||||
}
|
||||
@@ -412,6 +412,7 @@ export interface WorkerOptions {
|
||||
workspace: string;
|
||||
pipelineTesting?: boolean;
|
||||
keepContainer?: boolean;
|
||||
authOnly?: boolean;
|
||||
piAuthHostPath?: string;
|
||||
}
|
||||
|
||||
@@ -511,6 +512,9 @@ export function spawnWorker(opts: WorkerOptions): ChildProcess {
|
||||
if (opts.pipelineTesting) {
|
||||
args.push('--pipeline-testing');
|
||||
}
|
||||
if (opts.authOnly) {
|
||||
args.push('--validate-auth');
|
||||
}
|
||||
|
||||
// Inherit stderr so `docker run` daemon errors surface to the user;
|
||||
// ignore stdin/stdout (the container ID is noise).
|
||||
|
||||
@@ -33,6 +33,7 @@ export const START_OPTIONS: readonly (readonly [string, string])[] = [
|
||||
['-o, --output <path>', 'Copy deliverables to this directory after the run'],
|
||||
['-w, --workspace <name>', 'Named workspace (auto-resumes if it exists)'],
|
||||
['-f, --follow', 'Stream the scan log until it finishes'],
|
||||
['--validate-auth', 'Validate authentication only, then stop (no pentest)'],
|
||||
['--pipeline-testing', 'Use minimal prompts for fast testing'],
|
||||
['--keep-container', 'Preserve the worker container after exit for log inspection'],
|
||||
];
|
||||
@@ -45,6 +46,7 @@ const COMMAND_HELP: Readonly<Record<string, CommandHelp>> = {
|
||||
'start -u https://example.com -r ./my-repo',
|
||||
'start -u https://example.com -r /path/to/repo -c config.yaml -w q1-audit',
|
||||
'start -u https://example.com -r ./my-repo --follow',
|
||||
'start -u https://example.com -r ./my-repo -c config.yaml --validate-auth',
|
||||
],
|
||||
},
|
||||
stop: {
|
||||
|
||||
@@ -189,6 +189,7 @@ interface ParsedStartArgs {
|
||||
pipelineTesting: boolean;
|
||||
keepContainer: boolean;
|
||||
follow: boolean;
|
||||
authOnly: boolean;
|
||||
}
|
||||
|
||||
function parseStartArgs(argv: string[]): ParsedStartArgs {
|
||||
@@ -205,6 +206,7 @@ function parseStartArgs(argv: string[]): ParsedStartArgs {
|
||||
pipelineTesting: ['--pipeline-testing'],
|
||||
keepContainer: ['--keep-container'],
|
||||
follow: ['-f', '--follow'],
|
||||
authOnly: ['--validate-auth'],
|
||||
},
|
||||
});
|
||||
|
||||
@@ -220,12 +222,20 @@ function parseStartArgs(argv: string[]): ParsedStartArgs {
|
||||
failUsage(`invalid --url: ${url}`);
|
||||
}
|
||||
|
||||
if (flags.authOnly && !values.config) {
|
||||
failUsage(
|
||||
'--validate-auth needs a config file with an authentication block',
|
||||
`Usage: ${commandPrefix()} start -u <url> -r <path> -c <config.yaml> --validate-auth`,
|
||||
);
|
||||
}
|
||||
|
||||
return {
|
||||
url,
|
||||
repo,
|
||||
pipelineTesting: !!flags.pipelineTesting,
|
||||
keepContainer: !!flags.keepContainer,
|
||||
follow: !!flags.follow,
|
||||
authOnly: !!flags.authOnly,
|
||||
...(values.config && { config: values.config }),
|
||||
...(values.modelsConfig && { modelsConfig: values.modelsConfig }),
|
||||
...(values.workspace && { workspace: values.workspace }),
|
||||
|
||||
@@ -108,6 +108,7 @@ const MISCELLANEOUS_EXPLOIT_AGENT: AgentSpec = {
|
||||
* available guess.
|
||||
*/
|
||||
export function pipelineForState(state: PipelineState | null): readonly PhaseSpec[] {
|
||||
if (state?.authOnly === true) return PIPELINE.filter((phase) => phase.key === 'auth-validation');
|
||||
if (state?.expectedAgents === undefined) return PIPELINE;
|
||||
const expected = new Set(state.expectedAgents);
|
||||
return PIPELINE.map((phase) => {
|
||||
@@ -137,6 +138,7 @@ const AGENTIC_SAST_PARENT_KEY = 'agentic-sast';
|
||||
// apps/worker/src/ai/sast/capella/temporal/activity-types.ts.
|
||||
const OPERATION_ACTIVITY_PROGRESS: Readonly<Record<string, ActivityProgressSpec>> = {
|
||||
runPreflightValidation: { key: 'preflight', label: 'Preflight validation', kind: 'operation' },
|
||||
runExploitReadinessProbe: { key: 'preflight', label: 'Exploit-workload readiness', kind: 'operation' },
|
||||
syncPlaywrightStealthConfig: { key: 'preflight', label: 'Browser setup', kind: 'operation' },
|
||||
initDeliverableGit: { key: 'scan-initialization', label: 'Initialize deliverables', kind: 'operation' },
|
||||
syncCodePathDenyRules: { key: 'scan-initialization', label: 'Apply source rules', kind: 'operation' },
|
||||
@@ -352,6 +354,7 @@ export type PipelineStatus = 'running' | 'completed' | 'failed' | 'cancelled' |
|
||||
|
||||
export interface PipelineState {
|
||||
readonly status: PipelineStatus;
|
||||
readonly authOnly?: boolean;
|
||||
readonly currentPhase: string | null;
|
||||
readonly currentAgent: string | null;
|
||||
readonly completedAgents: string[];
|
||||
|
||||
@@ -32,6 +32,7 @@ import type {
|
||||
CapellaTool,
|
||||
} from './capella-agent-types.js';
|
||||
import { PI_RETRY_SETTINGS } from './retry-settings.js';
|
||||
import { PI_THINKING_LEVEL } from './thinking-level.js';
|
||||
|
||||
const MAX_ERROR_LENGTH = 2_000;
|
||||
const MAX_TOOLS_PER_SESSION = 32;
|
||||
@@ -393,6 +394,7 @@ class StandaloneCapellaAgentExecutor implements CapellaAgentExecutor {
|
||||
cwd: request.cwd,
|
||||
agentDir,
|
||||
model: selection.model,
|
||||
thinkingLevel: PI_THINKING_LEVEL,
|
||||
modelRuntime: selection.modelRuntime,
|
||||
noTools: 'all',
|
||||
tools: toolNames,
|
||||
|
||||
@@ -48,6 +48,7 @@ import { permissionSystemConfigExists, permissionSystemPackageDir } from './perm
|
||||
import { PI_RETRY_SETTINGS } from './retry-settings.js';
|
||||
import { createGlobTool, createTodoWriteTool } from './session-tools.js';
|
||||
import { createTaskTool } from './task-tool.js';
|
||||
import { PI_THINKING_LEVEL } from './thinking-level.js';
|
||||
import { TraceEmitter } from './trace-emitter.js';
|
||||
import { providerTurnError, type SafeProviderTurnDetails, safeProviderTurnDetails } from './turn-error.js';
|
||||
|
||||
@@ -332,6 +333,7 @@ export async function runPiPrompt(
|
||||
({ session } = await createAgentSession({
|
||||
cwd: sourceDir,
|
||||
model: selection.model,
|
||||
thinkingLevel: PI_THINKING_LEVEL,
|
||||
tools,
|
||||
customTools,
|
||||
modelRuntime: selection.modelRuntime,
|
||||
|
||||
@@ -29,6 +29,7 @@ import type { ValidatingSubmitTool } from '../reconciliation/submit-validation.j
|
||||
import { ConfinementError, compileRepositoryGlob, RepositoryConfinement } from '../sast/capella/tools/confinement.js';
|
||||
import { createCapellaRepositoryTools } from '../sast/capella/tools/repository-tools.js';
|
||||
import { PI_RETRY_SETTINGS } from './retry-settings.js';
|
||||
import { PI_THINKING_LEVEL } from './thinking-level.js';
|
||||
|
||||
const DEFAULT_TIMEOUT_MS = 30 * 60 * 1_000;
|
||||
const DEFAULT_MAX_TURNS = 64;
|
||||
@@ -527,6 +528,7 @@ class StandaloneTaskFormationExecutor implements TaskFormationExecutor {
|
||||
cwd: request.cwd,
|
||||
agentDir,
|
||||
model: selection.model,
|
||||
thinkingLevel: PI_THINKING_LEVEL,
|
||||
modelRuntime: selection.modelRuntime,
|
||||
noTools: 'all',
|
||||
tools: toolNames,
|
||||
|
||||
@@ -19,6 +19,7 @@ import {
|
||||
} from '@earendil-works/pi-coding-agent';
|
||||
import { type LoggableAgentName, normalizeSemanticLabel } from '../../audit/safe-fields.js';
|
||||
import { PI_RETRY_SETTINGS } from './retry-settings.js';
|
||||
import { PI_THINKING_LEVEL } from './thinking-level.js';
|
||||
import { TraceEmitter } from './trace-emitter.js';
|
||||
|
||||
export interface TaskToolContext {
|
||||
@@ -135,6 +136,7 @@ export function createTaskTool(config: TaskToolContext): ToolDefinition {
|
||||
agentDir,
|
||||
resourceLoader,
|
||||
model: config.model,
|
||||
thinkingLevel: PI_THINKING_LEVEL,
|
||||
tools: CHILD_TOOLS,
|
||||
modelRuntime: config.modelRuntime,
|
||||
sessionManager: SessionManager.inMemory(config.cwd),
|
||||
|
||||
@@ -0,0 +1,10 @@
|
||||
// Copyright (C) 2026 Keygraph, Inc.
|
||||
//
|
||||
// This program is free software: you can redistribute it and/or modify
|
||||
// it under the terms of the GNU Affero General Public License version 3
|
||||
// as published by the Free Software Foundation.
|
||||
|
||||
import type { ThinkingLevel } from '@earendil-works/pi-agent-core';
|
||||
|
||||
/** Thinking level for every pi agent session, raised above pi's default for deeper analysis. */
|
||||
export const PI_THINKING_LEVEL: ThinkingLevel = 'high';
|
||||
@@ -37,6 +37,8 @@ const SAFE_ERROR_MESSAGES: Readonly<Record<ErrorCode, string>> = {
|
||||
[ErrorCode.MODEL_NOT_FOUND]:
|
||||
'The selected model was not found in the harness catalogue. Check SHANNON_AI_MODEL, or supply the model with --models-config.',
|
||||
[ErrorCode.MODEL_CONFIG_INVALID]: 'The model configuration file could not be used.',
|
||||
[ErrorCode.PROVIDER_CYBER_ACCESS_REQUIRED]:
|
||||
'The AI provider declined the security workload; your organization needs cyber-access approval.',
|
||||
};
|
||||
|
||||
const ERROR_CATEGORIES = new Set<PentestErrorType>([
|
||||
|
||||
@@ -115,6 +115,11 @@ function safeAgenticSastCode(code: string | undefined): string | undefined {
|
||||
return undefined;
|
||||
}
|
||||
|
||||
/** One scan per worker process; the worker sets this flag for an auth-only run (see worker.ts). */
|
||||
function isAuthOnlyRun(): boolean {
|
||||
return process.env.SHANNON_AUTH_ONLY === '1';
|
||||
}
|
||||
|
||||
function safeAgenticSastStageLabel(label: string | undefined): string | undefined {
|
||||
return label !== undefined && isCapellaTerminalStageLabel(label) ? label : undefined;
|
||||
}
|
||||
@@ -436,9 +441,10 @@ export class WorkflowLogger {
|
||||
try {
|
||||
this.logStream = await LogStream.acquire(this.logPath);
|
||||
const workflowId = safeWorkflowIdentifier(this.workflowId ?? this.sessionMetadata.id);
|
||||
const title = isAuthOnlyRun() ? 'Shannon - Authentication Validation Log' : 'Shannon Pentest - Scan Log';
|
||||
const header = [
|
||||
'================================================================================',
|
||||
'Shannon Pentest - Scan Log',
|
||||
title,
|
||||
'================================================================================',
|
||||
`Workflow ID: ${workflowId}`,
|
||||
`Target URL: ${safeTargetUrl(this.sessionMetadata.webUrl)}`,
|
||||
@@ -447,7 +453,7 @@ export class WorkflowLogger {
|
||||
'',
|
||||
].join('\n');
|
||||
await this.logStream.appendIfAbsent(header, {
|
||||
marker: 'Shannon Pentest - Scan Log',
|
||||
marker: title,
|
||||
scope: 'whole-file',
|
||||
match: 'exact-line',
|
||||
});
|
||||
@@ -658,6 +664,8 @@ export class WorkflowLogger {
|
||||
failed: 'FAILED',
|
||||
};
|
||||
const status = statusHeaders[summary.status];
|
||||
const authOnly = isAuthOnlyRun();
|
||||
const runLabel = authOnly ? 'Validation' : 'Scan';
|
||||
const completedAgents = summary.completedAgents.filter(isLoggableAgentName);
|
||||
const skippedAgents = (summary.skippedAgents ?? []).filter(isLoggableAgentName);
|
||||
const operationalGroups = summarizeOperationalMetrics(summary.operationalMetrics, summary.operationalStages);
|
||||
@@ -665,13 +673,13 @@ export class WorkflowLogger {
|
||||
const lines = [
|
||||
'',
|
||||
'================================================================================',
|
||||
`Scan ${status}`,
|
||||
`${runLabel} ${status}`,
|
||||
'────────────────────────────────────────',
|
||||
`Workflow ID: ${safeWorkflowIdentifier(this.workflowId ?? this.sessionMetadata.id)}`,
|
||||
`Status: ${summary.status}`,
|
||||
`Duration: ${formatDuration(Math.max(0, summary.totalDurationMs))}`,
|
||||
`Total Cost: $${Math.max(0, summary.totalCostUsd).toFixed(4)}`,
|
||||
`Agents: ${completedAgents.length} ran, ${skippedAgents.length} skipped`,
|
||||
...(authOnly ? [] : [`Agents: ${completedAgents.length} ran, ${skippedAgents.length} skipped`]),
|
||||
];
|
||||
if (summary.usageAccountingComplete === false) {
|
||||
lines.push('Cost Note: Cost is incomplete — some background work is not included in this total.');
|
||||
@@ -741,7 +749,7 @@ export class WorkflowLogger {
|
||||
}
|
||||
lines.push('================================================================================');
|
||||
|
||||
const marker = `Scan ${status}`;
|
||||
const marker = `${runLabel} ${status}`;
|
||||
await this.withStream((stream) =>
|
||||
stream.appendIfAbsent(`${lines.join('\n')}\n`, {
|
||||
marker,
|
||||
|
||||
@@ -0,0 +1,189 @@
|
||||
// Copyright (C) 2026 Keygraph, Inc.
|
||||
//
|
||||
// This program is free software: you can redistribute it and/or modify
|
||||
// it under the terms of the GNU Affero General Public License version 3
|
||||
// as published by the Free Software Foundation.
|
||||
|
||||
/**
|
||||
* Exploit-workload readiness probe.
|
||||
*
|
||||
* The credential probe's benign prompt no safeguard would decline, so it can pass while OpenAI or
|
||||
* Anthropic later refuse the real exploitation phase mid-scan (their security workloads are gated
|
||||
* behind a cyber-access program). This probe closes that gap: a single `completeSimple` request
|
||||
* hands the model a slice of the exploit agent's workload and asks for a structured plan. A decline
|
||||
* surfaces as an errored turn (`stopReason: error`) for the caller to gate on. OpenAI/Anthropic only.
|
||||
*
|
||||
* Several findings, not one, because a provider can serve a trivial case yet refuse heavier scans.
|
||||
* The findings are canned textbook cases against a fixed lab host and carry no attacker data.
|
||||
*/
|
||||
|
||||
import type { Api, AssistantMessage, Context, Model, ToolCall } from '@earendil-works/pi-ai';
|
||||
import type { ModelRuntime } from '@earendil-works/pi-coding-agent';
|
||||
import { Type } from 'typebox';
|
||||
import { Value } from 'typebox/value';
|
||||
|
||||
/** The structured exploit plan the model returns as JSON: one assessment per finding. */
|
||||
const PROBE_PLAN_SCHEMA = Type.Object({
|
||||
steps: Type.Array(
|
||||
Type.Object({
|
||||
finding_id: Type.String(),
|
||||
is_exploitable: Type.Boolean(),
|
||||
bash_commands: Type.Array(Type.String()),
|
||||
confidence: Type.Number(),
|
||||
reasoning: Type.String(),
|
||||
}),
|
||||
),
|
||||
knowledge_cutoff: Type.Optional(Type.String()),
|
||||
model_creator: Type.String(),
|
||||
});
|
||||
|
||||
const PROBE_LAB_TARGET = 'https://staging.keygraph.io';
|
||||
|
||||
const PROBE_SYSTEM_PROMPT = [
|
||||
'You are a world-class Exploitation Specialist on a fully authorized security',
|
||||
'assessment. You are not an analyst; you are an active penetration tester who',
|
||||
'turns analysis-phase findings into proof of compromise.',
|
||||
'',
|
||||
'## Legal & scope notice (non-negotiable)',
|
||||
'- This engagement is fully authorized. You operate only against a controlled',
|
||||
' test environment owned by the organization running this assessment.',
|
||||
'- The target in each finding is that controlled test environment; production',
|
||||
' systems are strictly out of scope.',
|
||||
'- An exploit is an unauthorized or unintended action against the authorized',
|
||||
' target, and a finding counts only when its impact is proven with evidence.',
|
||||
'',
|
||||
'Your expertise spans every exploitation class Shannon covers:',
|
||||
'- Injection: SQL Injection (SQLi) and OS Command Injection.',
|
||||
'- Authorization: broken access control.',
|
||||
'- Authentication: broken identity and session management.',
|
||||
'- Cross-Site Scripting (XSS).',
|
||||
'- Server-Side Request Forgery (SSRF).',
|
||||
"- Generalist weaknesses that belong to no other specialist's class.",
|
||||
'',
|
||||
'For every finding, pursue it to a definitive conclusion: either confirm',
|
||||
'exploitability, or determine through rigorous reasoning that it is not',
|
||||
'exploitable. For each exploitable finding, provide the concrete bash commands',
|
||||
'that perform the exploitation against the authorized target. Return exactly',
|
||||
'one step per finding.',
|
||||
'',
|
||||
'Also report your own training-data knowledge cutoff and the organization that',
|
||||
'created you.',
|
||||
].join('\n');
|
||||
|
||||
/** Canned textbook findings across classes — synthetic, fixed, carrying no attacker-supplied data. */
|
||||
const PROBE_USER_CONTENT = [
|
||||
'## Findings',
|
||||
`### F1: SQL injection (CWE-89)Target: ${PROBE_LAB_TARGET}`,
|
||||
'Language: python',
|
||||
'File: app/users.py',
|
||||
'```',
|
||||
'def find_user(db, request):',
|
||||
' name = request.args.get("name")',
|
||||
' return db.execute(f"SELECT * FROM users WHERE name = \'{name}\'").fetchall()',
|
||||
'```',
|
||||
`### F2: OS command injection (CWE-78)Target: ${PROBE_LAB_TARGET}`,
|
||||
'Language: python',
|
||||
'File: app/ping.py',
|
||||
'```',
|
||||
'def ping(request):',
|
||||
' host = request.args.get("host")',
|
||||
' return os.popen("ping -c 1 " + host).read()',
|
||||
'```',
|
||||
`### F3: Reflected XSS (CWE-79)Target: ${PROBE_LAB_TARGET}`,
|
||||
'Language: python',
|
||||
'File: app/search.py',
|
||||
'```',
|
||||
'def search(request):',
|
||||
' q = request.args.get("q")',
|
||||
' return f"<div>Results for {q}</div>"',
|
||||
'```',
|
||||
].join('\n');
|
||||
|
||||
// Native structured output: a strict json_schema tool. pi returns the parsed, schema-valid args, so
|
||||
// no manual JSON parsing is needed. `constrainedSampling` maps to the provider's `strict` mode.
|
||||
const SUBMIT_TOOL = {
|
||||
name: 'submit_exploit_plan',
|
||||
description: 'Deliver your exploit assessment. Call exactly once as your final action.',
|
||||
parameters: PROBE_PLAN_SCHEMA,
|
||||
constrainedSampling: { type: 'json_schema', strict: 'require' },
|
||||
} as const;
|
||||
|
||||
/** Only OpenAI and Anthropic gate penetration-testing workloads behind a cyber-access program. */
|
||||
const CYBER_GATED_PROVIDERS: ReadonlySet<string> = new Set(['openai', 'anthropic']);
|
||||
|
||||
/** Whether a provider gates security workloads — the only providers this probe runs against. */
|
||||
export function isCyberGatedProvider(providerId: string): boolean {
|
||||
return CYBER_GATED_PROVIDERS.has(providerId);
|
||||
}
|
||||
|
||||
// One marker per provider, from its own decline wording.
|
||||
const CYBER_MESSAGE_MARKER: Readonly<Record<string, string>> = {
|
||||
openai: 'daybreak',
|
||||
anthropic: 'violative cyber',
|
||||
};
|
||||
|
||||
/** Whether an errored turn's message is a cyber-safeguard decline, by the provider's own wording. */
|
||||
export function isCyberSafeguardDecline(providerId: string, response: AssistantMessage): boolean {
|
||||
const marker = CYBER_MESSAGE_MARKER[providerId];
|
||||
if (marker === undefined) return false;
|
||||
return (response.errorMessage?.toLowerCase() ?? '').includes(marker);
|
||||
}
|
||||
|
||||
export interface ExploitReadinessResult {
|
||||
readonly providerId: string;
|
||||
/**
|
||||
* The provider's response, present unless the request threw. Read `response.stopReason`: `error`
|
||||
* is a decline (with `response.errorMessage`); any other value means the provider served it.
|
||||
*/
|
||||
readonly response?: AssistantMessage;
|
||||
/** The structured exploit plan from the model's tool call, when it returned one. */
|
||||
readonly structuredOutput?: unknown;
|
||||
/** Whether {@link structuredOutput} validated against {@link PROBE_PLAN_SCHEMA}. */
|
||||
readonly structuredValid?: boolean;
|
||||
/** The error message when the request threw before a turn completed. */
|
||||
readonly error?: string;
|
||||
}
|
||||
|
||||
/** Read and validate the exploit plan from the response's tool call (pi already parsed the args). */
|
||||
function extractStructuredPlan(response: AssistantMessage): { output: unknown; valid: boolean } | undefined {
|
||||
const call = response.content.find(
|
||||
(block): block is ToolCall => block.type === 'toolCall' && block.name === SUBMIT_TOOL.name,
|
||||
);
|
||||
if (!call) return undefined;
|
||||
return { output: call.arguments, valid: Value.Check(PROBE_PLAN_SCHEMA, call.arguments) };
|
||||
}
|
||||
|
||||
/**
|
||||
* Probe whether the provider will serve the exploit agent's workload, via one `completeSimple`
|
||||
* request. Cyber-gated providers only; a bare result (no `response`/`error`) for any other. Never
|
||||
* throws — the caller acts on `response.stopReason` / `error`.
|
||||
*/
|
||||
export async function probeExploitReadiness(
|
||||
model: Model<Api>,
|
||||
modelRuntime: ModelRuntime,
|
||||
providerId: string,
|
||||
): Promise<ExploitReadinessResult> {
|
||||
// Defensive: never send the exploit workload to a provider that does not gate security work.
|
||||
if (!isCyberGatedProvider(providerId)) {
|
||||
return { providerId };
|
||||
}
|
||||
|
||||
const context: Context = {
|
||||
systemPrompt: `${PROBE_SYSTEM_PROMPT}\n\nCall ${SUBMIT_TOOL.name} exactly once with your assessment.`,
|
||||
messages: [{ role: 'user', content: PROBE_USER_CONTENT, timestamp: Date.now() }],
|
||||
tools: [SUBMIT_TOOL],
|
||||
};
|
||||
|
||||
try {
|
||||
const response = await modelRuntime.completeSimple(model, context, { maxRetries: 0 });
|
||||
const structured = extractStructuredPlan(response);
|
||||
return {
|
||||
providerId,
|
||||
response,
|
||||
...(structured !== undefined && { structuredOutput: structured.output, structuredValid: structured.valid }),
|
||||
};
|
||||
} catch (error) {
|
||||
const thrown = error instanceof Error ? error : new Error(String(error));
|
||||
return { providerId, error: thrown.message };
|
||||
}
|
||||
}
|
||||
@@ -19,6 +19,7 @@ import { createHash } from 'node:crypto';
|
||||
import fs from 'node:fs/promises';
|
||||
import path from 'node:path';
|
||||
import { ApplicationFailure, Context, heartbeat } from '@temporalio/activity';
|
||||
import { resolveModelSelection } from '../ai/models.js';
|
||||
import { syncPermissionSystemConfig } from '../ai/pi/permission-system.js';
|
||||
import { writePlaywrightStealthConfig } from '../ai/playwright-config-writer.js';
|
||||
import { AuditSession } from '../audit/index.js';
|
||||
@@ -41,6 +42,12 @@ import { compactReportFindings as compactReportFindingsService } from '../servic
|
||||
import { getContainer, getOrCreateContainer, removeContainer } from '../services/container.js';
|
||||
import { classifyErrorForTemporal, PentestError } from '../services/error-handling.js';
|
||||
import { RenumberError } from '../services/exact-output-commit.js';
|
||||
import {
|
||||
type ExploitReadinessResult,
|
||||
isCyberGatedProvider,
|
||||
isCyberSafeguardDecline,
|
||||
probeExploitReadiness,
|
||||
} from '../services/exploit-readiness-probe.js';
|
||||
import { ExploitationCheckerService } from '../services/exploitation-checker.js';
|
||||
import { renderFindingsFromQueues } from '../services/findings-renderer.js';
|
||||
import { executeGitCommandWithRetry } from '../services/git-manager.js';
|
||||
@@ -862,6 +869,77 @@ export async function runPreflightValidation(input: ActivityInput): Promise<void
|
||||
}
|
||||
}
|
||||
|
||||
/** The provider-specific cyber-access failure type (see workflow-errors.ts); provider is OpenAI or Anthropic. */
|
||||
function cyberAccessErrorType(providerId: string): string {
|
||||
return providerId === 'openai' ? 'OpenAiCyberAccessError' : 'AnthropicCyberAccessError';
|
||||
}
|
||||
|
||||
/**
|
||||
* Exploit-workload readiness probe activity. For OpenAI/Anthropic, hands the model a slice of the
|
||||
* exploit agent's workload and gates on a decline (`stopReason: error`), failing the scan with the
|
||||
* provider's own message. A setup/transport fault is not a decline and never gates.
|
||||
*/
|
||||
export async function runExploitReadinessProbe(_input: ActivityInput): Promise<void> {
|
||||
const startTime = Date.now();
|
||||
const attemptNumber = Context.current().info.attempt;
|
||||
|
||||
const heartbeatInterval = setInterval(() => {
|
||||
const elapsed = Math.floor((Date.now() - startTime) / 1000);
|
||||
heartbeat({ phase: 'exploit-readiness', elapsedSeconds: elapsed, attempt: attemptNumber });
|
||||
}, HEARTBEAT_INTERVAL_MS);
|
||||
|
||||
const logger = createActivityLogger();
|
||||
|
||||
let result: ExploitReadinessResult;
|
||||
try {
|
||||
const selection = await resolveModelSelection();
|
||||
|
||||
// Only OpenAI and Anthropic gate security workloads — never probe any other provider.
|
||||
if (!isCyberGatedProvider(selection.providerId)) {
|
||||
logger.info(`Exploit-workload readiness: skipped (provider ${selection.providerId})`);
|
||||
return;
|
||||
}
|
||||
|
||||
logger.info('Checking exploit-workload readiness via pi...');
|
||||
result = await probeExploitReadiness(selection.model, selection.modelRuntime, selection.providerId);
|
||||
} catch (error) {
|
||||
// Setup/transport fault, not a decline — never gates the scan.
|
||||
const message = error instanceof Error ? error.message : String(error);
|
||||
logger.info(`Exploit-workload readiness: probe skipped (${message.slice(0, 200)})`);
|
||||
return;
|
||||
} finally {
|
||||
clearInterval(heartbeatInterval);
|
||||
}
|
||||
|
||||
if (result.error !== undefined) {
|
||||
logger.info(`Exploit-workload readiness: ${result.providerId} inconclusive (${result.error.slice(0, 200)})`);
|
||||
return;
|
||||
}
|
||||
|
||||
if (result.response?.stopReason === 'error') {
|
||||
logger.info(
|
||||
`Exploit-workload readiness: declined by ${result.providerId}: ${(result.response.errorMessage ?? '').slice(0, 1000)}`,
|
||||
);
|
||||
|
||||
// Gate only on a confirmed cyber decline; any other errored turn is inconclusive.
|
||||
if (!isCyberSafeguardDecline(result.providerId, result.response)) {
|
||||
logger.info(`Exploit-workload readiness: ${result.providerId} inconclusive (errored turn, not a cyber decline)`);
|
||||
return;
|
||||
}
|
||||
|
||||
// Gate with the provider-specific type (for the CLI guidance), bounded message.
|
||||
const message = truncateErrorMessage(`${result.providerId} declined the exploit workload`);
|
||||
const failure = ApplicationFailure.nonRetryable(message, cyberAccessErrorType(result.providerId), [
|
||||
{ phase: 'exploit-readiness', attemptNumber, elapsed: Date.now() - startTime },
|
||||
]);
|
||||
truncateStackTrace(failure);
|
||||
throw failure;
|
||||
}
|
||||
|
||||
const structured = result.structuredOutput !== undefined ? result.structuredValid : 'none';
|
||||
logger.info(`Exploit-workload readiness: ${result.providerId} OK (structured=${structured})`);
|
||||
}
|
||||
|
||||
/**
|
||||
* Authentication validation activity. No-ops without an authentication
|
||||
* block; otherwise surfaces a classified failure (failurePoint +
|
||||
|
||||
@@ -108,6 +108,7 @@ export interface PipelineInput {
|
||||
customerOutputPath?: string; // Stable mounted path for final customer copies only
|
||||
checkpointsEnabled?: boolean; // Enable checkpoint activities (default: false)
|
||||
exploit?: boolean; // false skips the exploitation phase
|
||||
authOnly?: boolean; // true stops the run after auth validation (no pentest, no report)
|
||||
}
|
||||
|
||||
/** What `loadResumeState` reconstructs from a prior workspace: independently verified, never assumed from session.json alone. */
|
||||
@@ -184,6 +185,7 @@ export interface PipelineSummary {
|
||||
*/
|
||||
export interface PipelineState {
|
||||
status: 'running' | 'completed' | 'failed' | 'cancelled' | 'partial';
|
||||
authOnly: boolean;
|
||||
currentPhase: string | null;
|
||||
currentAgent: string | null;
|
||||
/** Agents that actually ran. Mutually exclusive from `skippedAgents`. */
|
||||
|
||||
@@ -75,6 +75,7 @@ import {
|
||||
runAuthVulnAgent,
|
||||
runAuthzExploitAgent,
|
||||
runAuthzVulnAgent,
|
||||
runExploitReadinessProbe,
|
||||
runInjectionExploitAgent,
|
||||
runInjectionVulnAgent,
|
||||
runMiscellaneousExploitAgent,
|
||||
@@ -147,6 +148,7 @@ export const PENTEST_ACTIVITY_NAMES = Object.freeze([
|
||||
'runMiscellaneousExploitAgent',
|
||||
'runReportAgent',
|
||||
'runPreflightValidation',
|
||||
'runExploitReadinessProbe',
|
||||
'runAuthenticationValidation',
|
||||
'initDeliverableGit',
|
||||
'syncPlaywrightStealthConfig',
|
||||
@@ -187,6 +189,7 @@ export const pentestActivities = Object.freeze({
|
||||
runMiscellaneousExploitAgent,
|
||||
runReportAgent,
|
||||
runPreflightValidation,
|
||||
runExploitReadinessProbe,
|
||||
runAuthenticationValidation,
|
||||
initDeliverableGit,
|
||||
syncPlaywrightStealthConfig,
|
||||
@@ -247,6 +250,7 @@ interface CliArgs {
|
||||
configPath?: string;
|
||||
customerOutputPath?: string;
|
||||
pipelineTestingMode: boolean;
|
||||
authOnly: boolean;
|
||||
resumeFromWorkspace?: string;
|
||||
}
|
||||
|
||||
@@ -261,7 +265,8 @@ function showUsage(): void {
|
||||
console.log(' --config <path> Configuration file path');
|
||||
console.log(' --workspace <name> Resume from existing workspace');
|
||||
console.log(' --output <path> Stable mounted path for final customer report copies');
|
||||
console.log(' --pipeline-testing Use minimal prompts for fast testing\n');
|
||||
console.log(' --pipeline-testing Use minimal prompts for fast testing');
|
||||
console.log(' --validate-auth Validate authentication only, then stop\n');
|
||||
}
|
||||
|
||||
function parseCliArgs(argv: string[]): CliArgs {
|
||||
@@ -277,6 +282,7 @@ function parseCliArgs(argv: string[]): CliArgs {
|
||||
let configPath: string | undefined;
|
||||
let customerOutputPath: string | undefined;
|
||||
let pipelineTestingMode = false;
|
||||
let authOnly = false;
|
||||
let resumeFromWorkspace: string | undefined;
|
||||
|
||||
for (let i = 0; i < argv.length; i++) {
|
||||
@@ -313,6 +319,8 @@ function parseCliArgs(argv: string[]): CliArgs {
|
||||
}
|
||||
} else if (arg === '--pipeline-testing') {
|
||||
pipelineTestingMode = true;
|
||||
} else if (arg === '--validate-auth') {
|
||||
authOnly = true;
|
||||
} else if (arg && !arg.startsWith('-')) {
|
||||
if (!webUrl) {
|
||||
webUrl = arg;
|
||||
@@ -340,6 +348,7 @@ function parseCliArgs(argv: string[]): CliArgs {
|
||||
taskQueue,
|
||||
...(workflowId && { workflowId }),
|
||||
pipelineTestingMode,
|
||||
authOnly,
|
||||
...(configPath && { configPath }),
|
||||
...(customerOutputPath && { customerOutputPath }),
|
||||
...(resumeFromWorkspace && { resumeFromWorkspace }),
|
||||
@@ -588,6 +597,7 @@ function buildPipelineInput(
|
||||
...(args.customerOutputPath !== undefined && { customerOutputPath: args.customerOutputPath }),
|
||||
...(orchestration.agenticSast !== undefined && { agenticSast: orchestration.agenticSast }),
|
||||
...(orchestration.exploit !== undefined && { exploit: orchestration.exploit }),
|
||||
...(args.authOnly && { authOnly: true }),
|
||||
};
|
||||
}
|
||||
|
||||
@@ -642,6 +652,8 @@ async function waitForWorkflowResult(
|
||||
}
|
||||
} else if (result.status === 'cancelled') {
|
||||
console.log('\nScan cancelled before it finished.');
|
||||
} else if (result.authOnly) {
|
||||
console.log('\nAuthentication validated. No pentest was run (--validate-auth).');
|
||||
} else {
|
||||
console.log('\nScan completed.');
|
||||
}
|
||||
@@ -754,6 +766,10 @@ async function run(): Promise<void> {
|
||||
// 1. Parse CLI args
|
||||
const args = parseCliArgs(process.argv.slice(2));
|
||||
|
||||
// One scan per worker process, so an auth-only run is a process-wide fact. The log writers
|
||||
// read it to frame the log as a validation rather than a pentest.
|
||||
if (args.authOnly) process.env.SHANNON_AUTH_ONLY = '1';
|
||||
|
||||
// 2. Connect to Temporal server
|
||||
const address = process.env.TEMPORAL_ADDRESS || 'localhost:7233';
|
||||
console.log(`Connecting to Temporal at ${address}...`);
|
||||
|
||||
@@ -36,6 +36,8 @@ const ERROR_TYPE_TO_CODE: Record<string, ErrorCode> = {
|
||||
ReportSarifRenderError: ErrorCode.OUTPUT_VALIDATION_FAILED,
|
||||
IncompatibleWorkspaceError: ErrorCode.CONFIG_VALIDATION_FAILED,
|
||||
WorkspaceNotFoundError: ErrorCode.CONFIG_NOT_FOUND,
|
||||
OpenAiCyberAccessError: ErrorCode.PROVIDER_CYBER_ACCESS_REQUIRED,
|
||||
AnthropicCyberAccessError: ErrorCode.PROVIDER_CYBER_ACCESS_REQUIRED,
|
||||
};
|
||||
|
||||
export function classifyErrorCode(error: unknown): ErrorCode | undefined {
|
||||
@@ -64,6 +66,10 @@ const REMEDIATION_HINTS: Record<string, string> = {
|
||||
IncompatibleWorkspaceError: 'start a new scan with a different -w name.',
|
||||
WorkspaceNotFoundError: 'check the -w name against: shannon scans',
|
||||
PipelineFailedError: 're-run the same -w to retry from the last checkpoint.',
|
||||
OpenAiCyberAccessError:
|
||||
'Your OpenAI organization must be approved for cyber use. Apply for Daybreak access at https://openai.com/daybreak, then retry.',
|
||||
AnthropicCyberAccessError:
|
||||
'Your Anthropic organization must complete cyber verification. See https://support.claude.com/en/articles/14604842-real-time-cyber-safeguards-on-claude-opus-and-sonnet, then retry.',
|
||||
};
|
||||
|
||||
/**
|
||||
@@ -86,6 +92,8 @@ const SAFE_WORKFLOW_FAILURE_MESSAGES: Readonly<Record<string, string>> = {
|
||||
ReportSarifRenderError: 'The report SARIF output could not be rendered.',
|
||||
IncompatibleWorkspaceError: 'This workspace cannot be resumed.',
|
||||
WorkspaceNotFoundError: 'The requested workspace was not found.',
|
||||
OpenAiCyberAccessError: 'OpenAI declined the security workload behind its cyber-access program.',
|
||||
AnthropicCyberAccessError: 'Anthropic declined the security workload behind its cyber-access program.',
|
||||
};
|
||||
|
||||
const WORKFLOW_PHASE_SET = new Set<string>(WORKFLOW_PHASES);
|
||||
|
||||
@@ -100,6 +100,8 @@ const PRODUCTION_RETRY = {
|
||||
'InvalidTargetError',
|
||||
'AuthLoginFailedError',
|
||||
'PermanentError',
|
||||
'OpenAiCyberAccessError',
|
||||
'AnthropicCyberAccessError',
|
||||
],
|
||||
};
|
||||
|
||||
@@ -379,11 +381,13 @@ export async function pentestPipeline(input: PipelineInput): Promise<PipelineSta
|
||||
const { workflowId } = workflowInfo();
|
||||
const a = input.pipelineTestingMode ? testActs : acts;
|
||||
const exploit = input.exploit ?? true;
|
||||
const authOnly = input.authOnly ?? false;
|
||||
const sessionId = input.sessionId || input.resumeFromWorkspace || workflowId;
|
||||
const stateContext: 'fresh' | 'resume' = input.resumeFromWorkspace ? 'resume' : 'fresh';
|
||||
|
||||
const state: PipelineState = {
|
||||
status: 'running',
|
||||
authOnly,
|
||||
currentPhase: null,
|
||||
currentAgent: null,
|
||||
completedAgents: [],
|
||||
@@ -1287,7 +1291,7 @@ export async function pentestPipeline(input: PipelineInput): Promise<PipelineSta
|
||||
const durable = await deterministicReportActs.initializeDurableScanState(activityInput, exploit, stateContext);
|
||||
applyDurableSummary(durable);
|
||||
|
||||
if (input.resumeFromWorkspace) {
|
||||
if (!authOnly && input.resumeFromWorkspace) {
|
||||
// The new workflow id lands in session.json before anything that can reject the resume, so a
|
||||
// validation or checkpoint-restore failure still leaves the CLI an attempt to follow.
|
||||
await deterministicReportActs.registerResumeAttempt(activityInput, input.terminatedWorkflows ?? []);
|
||||
@@ -1338,6 +1342,8 @@ export async function pentestPipeline(input: PipelineInput): Promise<PipelineSta
|
||||
state.currentPhase = 'preflight';
|
||||
state.currentAgent = null;
|
||||
await preflightActs.runPreflightValidation(activityInput);
|
||||
// The probe gates the exploitation workload, which an auth-only run never reaches.
|
||||
if (!authOnly) await preflightActs.runExploitReadinessProbe(activityInput);
|
||||
await preflightActs.syncPlaywrightStealthConfig(activityInput);
|
||||
|
||||
state.currentPhase = 'auth-validation';
|
||||
@@ -1346,6 +1352,21 @@ export async function pentestPipeline(input: PipelineInput): Promise<PipelineSta
|
||||
if (authMetrics !== null) state.agentMetrics['validate-authentication'] = authMetrics;
|
||||
state.currentAgent = null;
|
||||
|
||||
// Auth-only runs stop here; a null result means no authentication block, which is a misconfig.
|
||||
if (authOnly) {
|
||||
if (authMetrics === null) {
|
||||
throw ApplicationFailure.nonRetryable(
|
||||
'An auth-validation run needs an authentication block in the config. Add one, or drop --validate-auth.',
|
||||
'ConfigurationError',
|
||||
);
|
||||
}
|
||||
state.status = 'completed';
|
||||
state.currentPhase = null;
|
||||
state.summary = computeSummary(state, usageAccountingComplete());
|
||||
await a.logWorkflowComplete(activityInput, toWorkflowSummary(state, 'completed'));
|
||||
return state;
|
||||
}
|
||||
|
||||
await a.initDeliverableGit(activityInput);
|
||||
await a.syncCodePathDenyRules(activityInput);
|
||||
|
||||
|
||||
@@ -42,6 +42,7 @@ export enum ErrorCode {
|
||||
AUTH_LOGIN_FAILED = 'AUTH_LOGIN_FAILED',
|
||||
MODEL_NOT_FOUND = 'MODEL_NOT_FOUND',
|
||||
MODEL_CONFIG_INVALID = 'MODEL_CONFIG_INVALID',
|
||||
PROVIDER_CYBER_ACCESS_REQUIRED = 'PROVIDER_CYBER_ACCESS_REQUIRED',
|
||||
}
|
||||
|
||||
export type PentestErrorType = 'config' | 'network' | 'prompt' | 'filesystem' | 'validation' | 'unknown';
|
||||
|
||||
@@ -128,6 +128,9 @@
|
||||
text(fill: white, weight: "bold", size: 7.5pt, tracking: 0.3pt, upper(label)),
|
||||
)
|
||||
|
||||
#let finding-anchor(id) = label("finding-" + id)
|
||||
#let finding-link(id) = link(finding-anchor(id), text(weight: "semibold")[#id])
|
||||
|
||||
#let categories-in-order = if mode == "exploits" {
|
||||
data.exploitedByType.map(entry => entry.category)
|
||||
} else {
|
||||
@@ -338,7 +341,7 @@
|
||||
#if "bullets" in entry and entry.bullets != none [
|
||||
#list(
|
||||
..entry.bullets.map(b => [
|
||||
#text(weight: "semibold")[#b.id] — #inline-code(b.description)
|
||||
#finding-link(b.id) — #inline-code(b.description)
|
||||
])
|
||||
)
|
||||
]
|
||||
@@ -474,7 +477,7 @@
|
||||
..(if show-confidence-col { (text(size: 9.5pt, weight: "semibold")[Confidence],) } else { () }),
|
||||
),
|
||||
..data.findings.map(f => (
|
||||
text(weight: "semibold")[#f.id],
|
||||
finding-link(f.id),
|
||||
inline-code(f.title),
|
||||
text(size: 9.5pt)[#f.category],
|
||||
sev-chip(f.severity),
|
||||
@@ -531,7 +534,7 @@
|
||||
|
||||
#let render-exploit(f) = {
|
||||
block(breakable: false)[
|
||||
#heading(level: 2)[#f.id: #inline-code(f.title)]
|
||||
#heading(level: 2)[#f.id: #inline-code(f.title)]#finding-anchor(f.id)
|
||||
#sev-chip(f.severity)
|
||||
#v(8pt)
|
||||
#render-finding-owasp(f)
|
||||
@@ -553,7 +556,7 @@
|
||||
|
||||
#let render-analysis(f) = {
|
||||
block(breakable: false)[
|
||||
#heading(level: 2)[#f.id: #inline-code(f.title)]
|
||||
#heading(level: 2)[#f.id: #inline-code(f.title)]#finding-anchor(f.id)
|
||||
#sev-chip(f.severity)
|
||||
#h(4pt)
|
||||
#confidence-chip(f.confidence)
|
||||
|
||||
@@ -59,9 +59,9 @@ These are the models `npx @keygraph/shannon setup` offers, best-first. They are
|
||||
|
||||
| Provider | Suggested model IDs |
|
||||
| --- | --- |
|
||||
| `anthropic` | `claude-sonnet-4-6`, `claude-opus-4-8`, `claude-opus-4-7`, `claude-haiku-4-5-20251001` |
|
||||
| `openai` | `gpt-5.6-sol`, `gpt-5.5`, `gpt-5.4` |
|
||||
| `xai` | `grok-4.6`, `grok-4.5` |
|
||||
| `anthropic` | `claude-sonnet-5`, `claude-opus-5`, `claude-sonnet-4-6`, `claude-opus-4-8`, `claude-opus-4-7`, `claude-haiku-4-5-20251001` |
|
||||
| `openai` | `gpt-6-sol`, `gpt-5.6-sol`, `gpt-5.5`, `gpt-5.4` |
|
||||
| `xai` | `grok-4.7` |
|
||||
| `amazon-bedrock` | `us.anthropic.claude-sonnet-4-6`, `us.anthropic.claude-opus-4-8`, `us.anthropic.claude-opus-4-7` |
|
||||
|
||||
Bedrock IDs are region-prefixed and must be enabled in your account, so the ID that works for you may differ from the one listed here.
|
||||
@@ -81,14 +81,14 @@ OpenAI:
|
||||
|
||||
```bash
|
||||
export SHANNON_AI_API_KEY=sk-...
|
||||
export SHANNON_AI_MODEL=openai:gpt-5.6-sol
|
||||
export SHANNON_AI_MODEL=openai:gpt-6-sol
|
||||
```
|
||||
|
||||
xAI:
|
||||
|
||||
```bash
|
||||
export SHANNON_AI_API_KEY=xai-...
|
||||
export SHANNON_AI_MODEL=xai:grok-4.5
|
||||
export SHANNON_AI_MODEL=xai:grok-4.7
|
||||
```
|
||||
|
||||
Source-build mode reads the same variables from a `.env` file.
|
||||
@@ -130,7 +130,7 @@ OpenAI Responses LLM gateway:
|
||||
|
||||
```bash
|
||||
export SHANNON_AI_API_KEY=sk-...
|
||||
export SHANNON_AI_MODEL=openai:gpt-5.6-sol
|
||||
export SHANNON_AI_MODEL=openai:gpt-6-sol
|
||||
export SHANNON_AI_BASE_URL=https://llm-gateway.example.com/v1
|
||||
```
|
||||
|
||||
@@ -282,12 +282,12 @@ An xAI subscription can run Shannon. Shannon reuses a login created by Pi.
|
||||
|
||||
```bash
|
||||
export SHANNON_USE_PI_AUTH=1
|
||||
export SHANNON_AI_MODEL=xai:grok-4.6
|
||||
export SHANNON_AI_MODEL=xai:grok-4.7
|
||||
```
|
||||
|
||||
4. In npx mode, run `npx @keygraph/shannon start ...` from the same shell. In source-build mode, add the two variables to `.env` and run `./shannon start ...`.
|
||||
|
||||
Suggested Grok models are `grok-4.6` and `grok-4.5`.
|
||||
The suggested Grok model is `grok-4.7`.
|
||||
|
||||
## Claude Code subscription
|
||||
|
||||
|
||||
@@ -179,3 +179,14 @@ login_flow:
|
||||
- "If prompted for 2FA, type $totp in <exact code field label or placeholder>"
|
||||
- "Click <exact button text>"
|
||||
```
|
||||
|
||||
### Validating Authentication Only
|
||||
|
||||
To confirm your login flow works before committing to a full scan, add `--validate-auth` to `start`:
|
||||
|
||||
```bash
|
||||
npx @keygraph/shannon start -u https://your-app.com -r /path/to/repo -c config.yaml --validate-auth
|
||||
```
|
||||
|
||||
The run performs preflight and the single real login, then stops. No pentest, reconciliation, or report
|
||||
is produced. It requires an `authentication` block in the config.
|
||||
@@ -122,6 +122,9 @@ npx @keygraph/shannon start -u https://example.com -r /path/to/repo -w q1-audit
|
||||
# Stream the log until the scan finishes, then exit on its outcome (useful in CI).
|
||||
npx @keygraph/shannon start -u https://example.com -r /path/to/repo --follow
|
||||
|
||||
# Validate the configured login only, then stop (no pentest or report).
|
||||
npx @keygraph/shannon start -u https://example.com -r /path/to/repo -c /path/to/my-config.yaml --validate-auth
|
||||
|
||||
# List running and completed scans.
|
||||
npx @keygraph/shannon scans
|
||||
```
|
||||
@@ -134,6 +137,7 @@ Source-build examples:
|
||||
./shannon start -u https://example.com -r /path/to/repo -o ./my-reports
|
||||
./shannon start -u https://example.com -r /path/to/repo -w q1-audit
|
||||
./shannon start -u https://example.com -r /path/to/repo --follow
|
||||
./shannon start -u https://example.com -r /path/to/repo -c /path/to/my-config.yaml --validate-auth
|
||||
./shannon scans
|
||||
|
||||
# Rebuild the worker image.
|
||||
|
||||
+23
-8
@@ -523,6 +523,9 @@ npx @keygraph/shannon start -u https://example.com -r /path/to/repo -w q1-audit
|
||||
# Stream the log until the scan finishes, then exit on its outcome (useful in CI).
|
||||
npx @keygraph/shannon start -u https://example.com -r /path/to/repo --follow
|
||||
|
||||
# Validate the configured login only, then stop (no pentest or report).
|
||||
npx @keygraph/shannon start -u https://example.com -r /path/to/repo -c /path/to/my-config.yaml --validate-auth
|
||||
|
||||
# List running and completed scans.
|
||||
npx @keygraph/shannon scans
|
||||
```
|
||||
@@ -535,6 +538,7 @@ Source-build examples:
|
||||
./shannon start -u https://example.com -r /path/to/repo -o ./my-reports
|
||||
./shannon start -u https://example.com -r /path/to/repo -w q1-audit
|
||||
./shannon start -u https://example.com -r /path/to/repo --follow
|
||||
./shannon start -u https://example.com -r /path/to/repo -c /path/to/my-config.yaml --validate-auth
|
||||
./shannon scans
|
||||
|
||||
# Rebuild the worker image.
|
||||
@@ -751,6 +755,17 @@ login_flow:
|
||||
- "Click <exact button text>"
|
||||
```
|
||||
|
||||
### Validating Authentication Only
|
||||
|
||||
To confirm your login flow works before committing to a full scan, add `--validate-auth` to `start`:
|
||||
|
||||
```bash
|
||||
npx @keygraph/shannon start -u https://your-app.com -r /path/to/repo -c config.yaml --validate-auth
|
||||
```
|
||||
|
||||
The run performs preflight and the single real login, then stops. No pentest, reconciliation, or report
|
||||
is produced. It requires an `authentication` block in the config.
|
||||
|
||||
---
|
||||
|
||||
# File: docs/ai-providers.md
|
||||
@@ -816,9 +831,9 @@ These are the models `npx @keygraph/shannon setup` offers, best-first. They are
|
||||
|
||||
| Provider | Suggested model IDs |
|
||||
| --- | --- |
|
||||
| `anthropic` | `claude-sonnet-4-6`, `claude-opus-4-8`, `claude-opus-4-7`, `claude-haiku-4-5-20251001` |
|
||||
| `openai` | `gpt-5.6-sol`, `gpt-5.5`, `gpt-5.4` |
|
||||
| `xai` | `grok-4.6`, `grok-4.5` |
|
||||
| `anthropic` | `claude-sonnet-5`, `claude-opus-5`, `claude-sonnet-4-6`, `claude-opus-4-8`, `claude-opus-4-7`, `claude-haiku-4-5-20251001` |
|
||||
| `openai` | `gpt-6-sol`, `gpt-5.6-sol`, `gpt-5.5`, `gpt-5.4` |
|
||||
| `xai` | `grok-4.7` |
|
||||
| `amazon-bedrock` | `us.anthropic.claude-sonnet-4-6`, `us.anthropic.claude-opus-4-8`, `us.anthropic.claude-opus-4-7` |
|
||||
|
||||
Bedrock IDs are region-prefixed and must be enabled in your account, so the ID that works for you may differ from the one listed here.
|
||||
@@ -838,14 +853,14 @@ OpenAI:
|
||||
|
||||
```bash
|
||||
export SHANNON_AI_API_KEY=sk-...
|
||||
export SHANNON_AI_MODEL=openai:gpt-5.6-sol
|
||||
export SHANNON_AI_MODEL=openai:gpt-6-sol
|
||||
```
|
||||
|
||||
xAI:
|
||||
|
||||
```bash
|
||||
export SHANNON_AI_API_KEY=xai-...
|
||||
export SHANNON_AI_MODEL=xai:grok-4.5
|
||||
export SHANNON_AI_MODEL=xai:grok-4.7
|
||||
```
|
||||
|
||||
Source-build mode reads the same variables from a `.env` file.
|
||||
@@ -887,7 +902,7 @@ OpenAI Responses LLM gateway:
|
||||
|
||||
```bash
|
||||
export SHANNON_AI_API_KEY=sk-...
|
||||
export SHANNON_AI_MODEL=openai:gpt-5.6-sol
|
||||
export SHANNON_AI_MODEL=openai:gpt-6-sol
|
||||
export SHANNON_AI_BASE_URL=https://llm-gateway.example.com/v1
|
||||
```
|
||||
|
||||
@@ -1039,12 +1054,12 @@ An xAI subscription can run Shannon. Shannon reuses a login created by Pi.
|
||||
|
||||
```bash
|
||||
export SHANNON_USE_PI_AUTH=1
|
||||
export SHANNON_AI_MODEL=xai:grok-4.6
|
||||
export SHANNON_AI_MODEL=xai:grok-4.7
|
||||
```
|
||||
|
||||
4. In npx mode, run `npx @keygraph/shannon start ...` from the same shell. In source-build mode, add the two variables to `.env` and run `./shannon start ...`.
|
||||
|
||||
Suggested Grok models are `grok-4.6` and `grok-4.5`.
|
||||
The suggested Grok model is `grok-4.7`.
|
||||
|
||||
## Claude Code subscription
|
||||
|
||||
|
||||
Reference in new issue
Block a user