mirror of
https://github.com/KeygraphHQ/shannon.git
synced 2026-09-17 23:42:23 +02:00
* feat(worker): add agentic static analysis Add the ten-stage Agentic SAST pipeline, confined repository tools, model runtime, prompt templates, and SARIF export. Make retries, repair sessions, reduced coverage, usage accounting, and model-output drift durable across Temporal replay and resume. Keep retry diagnostics in their actionable closed vocabulary. Package the Mantis-derived license material with the prompts that require it. * feat(worker): deduplicate static and runtime findings before exploitation Parse Agentic SAST SARIF into typed observations, enrich and route those observations, and reconcile them with pentest findings before exploitation. Publish deterministic exploitation queues with stable lineage, exact-path Git commits, retry-safe manifests, named drop reasons, and confined task formation. Reject duplicate producer IDs before commit and adopt either legal provenance shape after a lost acknowledgement. * feat(config)!: replace vuln_classes with agentic_sast Wire Agentic SAST and reconciliation into the main pipeline, persist their durable state, and add the Miscellaneous finding and exploitation lane. Make scan completion, cancellation, partial outcomes, resume identity, and report recovery use the integrated final workflow contract. Introduce the atomic finalization, ordering, renumbering, compaction, and output services that workflow calls. Keep completed Miscellaneous work and report drafts idempotent across resume, preserve public main's default-on exploit SARIF behavior, and describe stage-fallback candidates without claiming they were exported. BREAKING CHANGE: `vuln_classes` has been removed. Configs containing it now fail validation, and all five core pentest classes run on every scan. Workspaces created by Shannon 2.x cannot be resumed. Finish or discard in-flight scans before upgrading, then start a new workspace name. * perf: overlap static analysis and the Miscellaneous lane with the pentest Run Agentic SAST alongside vulnerability analysis and run Miscellaneous exploitation alongside the specialist exploitation lanes. Keep reconciliation dependent on the completed static-analysis result while preserving parallel work everywhere that has no data dependency. * feat(cli)!: default the scan target and add a JSON error contract List local scans, resolve the active or most recent workspace automatically, and make logs, status, and stop use one canonical scan identity. Add stable machine-readable failures, richer status output, explicit help errors, and seven-day Temporal retention. Treat absent Temporal pending-activity failures as absent whether the decoder represents them as `null` or missing. BREAKING CHANGE: `status --json` now returns a fixed `failureMessage`. Read `partialReasons`, `agenticSast`, and `workflow.log` for diagnostic detail. * feat(logging): trace tool calls and write a log per agent Record complete tool-call arguments in the workflow log and project each agent's events into its own durable log. Add agent listing and agent-specific log tailing while preserving byte-exact output and draining log handles before activities return. * feat(worker): standardize severity and reporting guidance in exploit prompts Give every exploit agent the same status, confidence, severity-reasoning, report-writing, credential-handling, and scope contract. Apply the same task-formation and SAST-enrichment procedure to the Miscellaneous lane. * feat(worker): disclose scan coverage and make reporting auditable Build on the retry-safe finalization foundation to preserve correct identities, source locations, scan dates, partial-coverage limitations, and consistent report JSON, Markdown, SARIF, and PDF output. Report Agentic SAST, reconciliation wall-clock time, stage usage, retry spend, and background work without duplicate or hardcoded totals. Keep report findings canonical, drop cross-class restatements, name enrichment losses, and render the executive-summary narrative in the PDF. * chore(license): attribute Mantis and Pi and refresh the docs Add the final Mantis and Pi notices, license copies, acknowledgements, and residual copyright updates. Update the README, maintained documentation, contributor guidance, and hand-maintained mirrors to describe Agentic SAST, reconciliation, the Miscellaneous lane, current CLI behavior, and the final release contract. Correct stale workspace and container guidance and annotate long-standing internals for maintainers. * fix(logging): treat a slash as a word separator in agent labels * feat(cli)!: rebuild scan status around model work - show Capella stages beneath the concurrent Agentic SAST phase - attach reconciliation time to the class row it feeds - hide completed bookkeeping and the duplicate miscellaneous wrapper - carry validated child-workflow progress into durable parent state - derive the terminal tree and status JSON from the same phase shape BREAKING CHANGE: `status --json` replaces phase `parallel` with `children` and `meta`, adds phase summaries and notes plus agent attachment fields, and removes the `analysis-engines` and `operational-work` phases. * fix(report): drop the empty Critical Findings section from the PDF summary * fix(sast): align Capella export with the submit-time code-path contract The export gate required every code_paths entry to be file:line, but submit only requires the primary sink to be file:line and accepts bare trace steps. A single malformed trace step therefore dropped an otherwise-valid finding at export. - add isValidPrimaryCodePath as the one shared primary-sink contract - validate only the primary at export; buildResult already drops unusable steps - route the submit-time validator through the same helper so the two cannot drift * feat(sast): tolerate hygiene-only Capella reductions instead of going partial A reduction only makes a run partial when it loses real coverage or a whole finding. Malformed model output, salvaged turn-limit work, and rejected duplicate verdicts are recorded as evidence but no longer flip the run to partial. - add reductionIsTolerable: partial only when genuine-loss counts are nonzero - drive runCapella's partial reasons and display coverage off non-tolerable ones - keep every reduction in agenticSast.reductions so nothing is lost as evidence * feat(logging): record the provider reason for a failed agent turn A failed provider turn collapsed to AGENT_EXECUTION_FAILED/unknown with the underlying reason discarded, so a model-side rejection or safeguard was indistinguishable from a transport fault in the error log. - add safeProviderTurnDetails: write bounded, non-sensitive fields (provider, model, responseId, stop reason, tool-in-flight, category, retryable) to error.log - gate a sanitized errorMessage snippet behind SHANNON_DEBUG_PROVIDER_ERRORS, off by default - forward SHANNON_DEBUG_PROVIDER_ERRORS from the CLI into the worker container * fix(cli): keep shannon logs tailing through a Temporal blip - End the interactive tail on the log's own terminal marker or Ctrl-C, so a transient Temporal outage no longer aborts the command with exit 1. - Rebuild the memoized Temporal client after a failed poll: a wedged gRPC channel was cached forever, so "retrying…" could never reconnect. - Keep start --follow (CI) bounded — a genuinely dead Temporal still fails the run instead of hanging. * fix(worker): correct PDF finding reporting - Render OWASP category, authentication state, and remediation - Omit the redundant per-finding exploited status - Preserve canonical category and field ordering across report modes - Continue Proof of Impact numbering across embedded code blocks - Wrap long PDF code lines without changing canonical report content * fix: attribute a reconciliation failure to exploitation only - Stop marking a class's vulnerability-analysis agent failed when that agent succeeded and only reconciliation failed; the status tree now renders the analysis row completed and the exploitation row failed - Consume the worker's failedReconciliations signal in the CLI, which the mirrored PipelineState already declared but never read - Correct the class_reconciliation_failed message, which claimed the class's analysis results were still in the report when the class is excluded from it * fix(pi): give each task sub-session its own resource loader to prevent stale extension ctx * fix(prompts): scope exploit agents to in-band proof, mark OOB-only findings blocked * fix(cli): reject a shell credential that shadows a gateway config.toml key * fix(cli): make scan shutdown verifiable - preselect and persist workflow identity before worker launch - cancel first, then verify bounded Temporal termination - reconcile Docker workers with Temporal open workflows - fail closed on stale images and unavailable lifecycle state - mark cancellation only after confirmed shutdown * feat(cli): prompt for setup on a bare npx invocation with no credentials * fix(cli): don't blame anthropic when no credentials are configured at all * chore(release): bump beta base version to 3.0.0 * feat(cli): show a 'start your first scan' box in help on a TTY * docs: refresh README and platform overview for Shannon 3.0 - lead with the 3.0 launch note and rewrite key capabilities around security code analysis, the rebuilt terminal experience, native CI/CD, and PDF/SARIF - recast the editions table as Shannon Open Source against the Keygraph Enterprise Platform, stating open source is not a trial edition - rewrite the platform overview around exhaustive agentic SAST, canonical findings, automated remediation, targeted verification, and governance - add five product screenshots under assets/keygraph-platform/, referenced relative to docs/ * docs: add the Shannon naming section and swap in the 3.0 demo GIF - explain the Claude Shannon information-theory origin under "What is Shannon?" - point "Shannon in Action" at the 3.0 recording in assets/Shannon3GIF.gif Both taken from the README half of #438. * docs: document CI/CD integrations and the reconciled analysis pipeline - add a CI/CD Integrations section covering the official GitHub Action and GitLab component, pipeline artifacts, and exploit-only severity gates - redraw the architecture section as a Mermaid flow: agentic code analysis and recon feed finding reconciliation, then exploitation and reporting - describe open-source code analysis as a multi-stage agentic workflow and reserve parsed-code CPGs and exhaustive verification for Enterprise - sharpen the privacy wording: results stay local, but model requests carry source context to whichever endpoint you configure - drop the "not recommended" framing on local models and add a section on why Shannon complements rather than replaces human pentesters - regenerate llms-full.txt from the updated README and docs * docs: add the Photoview benchmark across three models - Add a "Shannon in Action" table for Photoview 2.4.0 runs on DeepSeek v4 Flash, Grok 4.6, and Claude Opus 5, each linking its PDF report and SARIF output - Store the per-model reports under benchmark/ - Link the (forthcoming) benchmark writeup from the section intro * docs: add the Shannon vs XBOW/Aikido Photoview benchmark writeup - Add docs/shannon-xbow-aikido-benchmark.md with methodology, per-model cost/coverage tables, and links to each model's report and SARIF - Link the writeup from the README "Shannon in Action" section * docs: link the benchmark announcement discussion from the README * fix(readme): restore theme-aware banner, badge, and buttons * feat!: trigger the Shannon 3.0 major release --------- Co-authored-by: ezl-keygraph <ezhil@keygraph.io>
744 lines
24 KiB
TypeScript
744 lines
24 KiB
TypeScript
// Copyright (C) 2026 Keygraph, Inc.
|
|
//
|
|
// This program is free software: you can redistribute it and/or modify
|
|
// it under the terms of the GNU Affero General Public License version 3
|
|
// as published by the Free Software Foundation.
|
|
|
|
import { createRequire } from 'node:module';
|
|
import { Ajv, type ErrorObject, type ValidateFunction } from 'ajv';
|
|
import type { FormatsPlugin } from 'ajv-formats';
|
|
import yaml from 'js-yaml';
|
|
import { fs } from 'zx';
|
|
import { PentestError } from './services/error-handling.js';
|
|
import type { Authentication, Config, DistributedConfig, Rule } from './types/config.js';
|
|
import { ErrorCode } from './types/errors.js';
|
|
|
|
/**
|
|
* Parses and validates scan configuration YAML against config-schema.json, then
|
|
* distributes it into the plain values consumed by prompts and services.
|
|
*
|
|
* The schema is closed: every object in config-schema.json sets `additionalProperties:
|
|
* false`, so an unrecognized field anywhere in the config is a hard validation failure
|
|
* rather than a silently ignored typo. There is no public way to select which analysis
|
|
* classes run; the schema only exposes steering knobs (rules, authentication,
|
|
* agentic_sast.enabled, exploit, report, rules_of_engagement) on top of the fixed
|
|
* five-class pipeline.
|
|
*/
|
|
|
|
// Handle ESM/CJS interop for ajv-formats using require
|
|
const require = createRequire(import.meta.url);
|
|
const addFormats: FormatsPlugin = require('ajv-formats');
|
|
|
|
const ajv = new Ajv({ allErrors: true, verbose: true });
|
|
addFormats(ajv);
|
|
|
|
let configSchema: object;
|
|
let validateSchema: ValidateFunction;
|
|
|
|
try {
|
|
const schemaPath = new URL('../configs/config-schema.json', import.meta.url);
|
|
const schemaContent = await fs.readFile(schemaPath, 'utf8');
|
|
configSchema = JSON.parse(schemaContent) as object;
|
|
validateSchema = ajv.compile(configSchema);
|
|
} catch (error) {
|
|
const errMsg = error instanceof Error ? error.message : String(error);
|
|
throw new PentestError(`Failed to load configuration schema: ${errMsg}`, 'config', false, {
|
|
schemaPath: '../configs/config-schema.json',
|
|
originalError: errMsg,
|
|
});
|
|
}
|
|
|
|
// Free-text config fields (description, rules_of_engagement, rule values, login fields,
|
|
// report.guidance) get interpolated verbatim into agent prompts via prompt-manager.ts.
|
|
// These patterns block the more obvious ways a scan config could smuggle markup, script
|
|
// URLs, or path traversal into that prompt text or into a rendered value.
|
|
const DANGEROUS_PATTERNS: RegExp[] = [
|
|
/\.\.\//, // Path traversal
|
|
/[<>]/, // HTML/XML injection
|
|
/javascript:/i, // JavaScript URLs
|
|
/data:/i, // Data URLs
|
|
/file:/i, // File URLs
|
|
];
|
|
|
|
/**
|
|
* Format a single AJV error into a human-readable message.
|
|
* Translates AJV error keywords into plain English descriptions.
|
|
*/
|
|
function formatAjvError(error: ErrorObject): string {
|
|
const path = error.instancePath || 'root';
|
|
const params = error.params as Record<string, unknown>;
|
|
|
|
switch (error.keyword) {
|
|
case 'required': {
|
|
const missingProperty = params.missingProperty as string;
|
|
return `Missing required field: "${missingProperty}" at ${path || 'root'}`;
|
|
}
|
|
|
|
case 'type': {
|
|
const expectedType = params.type as string;
|
|
return `Invalid type at ${path}: expected ${expectedType}`;
|
|
}
|
|
|
|
case 'enum': {
|
|
const allowedValues = params.allowedValues as unknown[];
|
|
const formattedValues = allowedValues.map((v) => `"${v}"`).join(', ');
|
|
return `Invalid value at ${path}: must be one of [${formattedValues}]`;
|
|
}
|
|
|
|
case 'additionalProperties': {
|
|
const additionalProperty = params.additionalProperty as string;
|
|
return `Unknown field at ${path}: "${additionalProperty}" is not allowed`;
|
|
}
|
|
|
|
case 'minLength': {
|
|
const limit = params.limit as number;
|
|
return `Value at ${path} is too short: must have at least ${limit} character(s)`;
|
|
}
|
|
|
|
case 'maxLength': {
|
|
const limit = params.limit as number;
|
|
return `Value at ${path} is too long: must have at most ${limit} character(s)`;
|
|
}
|
|
|
|
case 'minimum': {
|
|
const limit = params.limit as number;
|
|
return `Value at ${path} is too small: must be >= ${limit}`;
|
|
}
|
|
|
|
case 'maximum': {
|
|
const limit = params.limit as number;
|
|
return `Value at ${path} is too large: must be <= ${limit}`;
|
|
}
|
|
|
|
case 'minItems': {
|
|
const limit = params.limit as number;
|
|
return `Array at ${path} has too few items: must have at least ${limit} item(s)`;
|
|
}
|
|
|
|
case 'maxItems': {
|
|
const limit = params.limit as number;
|
|
return `Array at ${path} has too many items: must have at most ${limit} item(s)`;
|
|
}
|
|
|
|
case 'pattern': {
|
|
const pattern = params.pattern as string;
|
|
return `Value at ${path} does not match required pattern: ${pattern}`;
|
|
}
|
|
|
|
case 'format': {
|
|
const format = params.format as string;
|
|
return `Value at ${path} must be a valid ${format}`;
|
|
}
|
|
|
|
case 'const': {
|
|
const allowedValue = params.allowedValue as unknown;
|
|
return `Value at ${path} must be exactly "${allowedValue}"`;
|
|
}
|
|
|
|
case 'oneOf': {
|
|
return `Value at ${path} must match exactly one schema (matched ${params.passingSchemas ?? 0})`;
|
|
}
|
|
|
|
case 'anyOf': {
|
|
return `Value at ${path} must match at least one of the allowed schemas`;
|
|
}
|
|
|
|
case 'not': {
|
|
return `Value at ${path} matches a schema it should not match`;
|
|
}
|
|
|
|
case 'if': {
|
|
return `Value at ${path} does not satisfy conditional schema requirements`;
|
|
}
|
|
|
|
case 'uniqueItems': {
|
|
const i = params.i as number;
|
|
const j = params.j as number;
|
|
return `Array at ${path} contains duplicate items at positions ${j} and ${i}`;
|
|
}
|
|
|
|
case 'propertyNames': {
|
|
const propertyName = params.propertyName as string;
|
|
return `Invalid property name at ${path}: "${propertyName}" does not match naming requirements`;
|
|
}
|
|
|
|
case 'dependencies':
|
|
case 'dependentRequired': {
|
|
const property = params.property as string;
|
|
const missingProperty = params.missingProperty as string;
|
|
return `Missing dependent field at ${path}: "${missingProperty}" is required when "${property}" is present`;
|
|
}
|
|
|
|
default: {
|
|
// Fallback for any unhandled keywords - use AJV's message if available
|
|
const message = error.message || `validation failed for keyword "${error.keyword}"`;
|
|
return `${path}: ${message}`;
|
|
}
|
|
}
|
|
}
|
|
|
|
/**
|
|
* Format all AJV errors into a list of human-readable messages.
|
|
* Returns an array of formatted error strings.
|
|
*/
|
|
function formatAjvErrors(errors: ErrorObject[]): string[] {
|
|
return errors.map(formatAjvError);
|
|
}
|
|
|
|
export const parseConfig = async (configPath: string): Promise<Config> => {
|
|
try {
|
|
// 1. Verify file exists
|
|
if (!(await fs.pathExists(configPath))) {
|
|
throw new PentestError(
|
|
`Configuration file not found: ${configPath}`,
|
|
'config',
|
|
false,
|
|
{ configPath },
|
|
ErrorCode.CONFIG_NOT_FOUND,
|
|
);
|
|
}
|
|
|
|
// 2. Check file size
|
|
const stats = await fs.stat(configPath);
|
|
const maxFileSize = 1024 * 1024; // 1MB
|
|
if (stats.size > maxFileSize) {
|
|
throw new PentestError(
|
|
`Configuration file too large: ${stats.size} bytes (maximum: ${maxFileSize} bytes)`,
|
|
'config',
|
|
false,
|
|
{ configPath, fileSize: stats.size, maxFileSize },
|
|
ErrorCode.CONFIG_VALIDATION_FAILED,
|
|
);
|
|
}
|
|
|
|
// 3. Read and check for empty content
|
|
const configContent = await fs.readFile(configPath, 'utf8');
|
|
|
|
if (!configContent.trim()) {
|
|
throw new PentestError(
|
|
'Configuration file is empty',
|
|
'config',
|
|
false,
|
|
{ configPath },
|
|
ErrorCode.CONFIG_VALIDATION_FAILED,
|
|
);
|
|
}
|
|
|
|
// 4. Parse YAML with safe schema
|
|
let config: unknown;
|
|
try {
|
|
config = yaml.load(configContent, {
|
|
schema: yaml.FAILSAFE_SCHEMA, // Only basic YAML types, no JS evaluation
|
|
json: false, // Don't allow JSON-specific syntax
|
|
filename: configPath,
|
|
});
|
|
} catch (yamlError) {
|
|
const errMsg = yamlError instanceof Error ? yamlError.message : String(yamlError);
|
|
throw new PentestError(
|
|
`YAML parsing failed: ${errMsg}`,
|
|
'config',
|
|
false,
|
|
{ configPath, originalError: errMsg },
|
|
ErrorCode.CONFIG_PARSE_ERROR,
|
|
);
|
|
}
|
|
|
|
// 5. Guard against null/undefined parse result
|
|
if (config === null || config === undefined) {
|
|
throw new PentestError(
|
|
'Configuration file resulted in null/undefined after parsing',
|
|
'config',
|
|
false,
|
|
{ configPath },
|
|
ErrorCode.CONFIG_PARSE_ERROR,
|
|
);
|
|
}
|
|
|
|
// 6. Validate schema, security rules, and return
|
|
validateConfig(config as Config);
|
|
|
|
return config as Config;
|
|
} catch (error) {
|
|
// PentestError instances are already well-formatted, re-throw as-is
|
|
if (error instanceof PentestError) {
|
|
throw error;
|
|
}
|
|
const errMsg = error instanceof Error ? error.message : String(error);
|
|
throw new PentestError(
|
|
`Failed to parse configuration file '${configPath}': ${errMsg}`,
|
|
'config',
|
|
false,
|
|
{ configPath, originalError: errMsg },
|
|
ErrorCode.CONFIG_PARSE_ERROR,
|
|
);
|
|
}
|
|
};
|
|
|
|
/**
|
|
* Parse a raw YAML string into a validated Config object.
|
|
*
|
|
* Same validation as parseConfig but accepts a string instead of a file path.
|
|
* Used when config YAML is passed inline (e.g., from a parent workflow).
|
|
*/
|
|
export const parseConfigYAML = (yamlContent: string): Config => {
|
|
if (!yamlContent.trim()) {
|
|
throw new PentestError(
|
|
'Configuration YAML string is empty',
|
|
'config',
|
|
false,
|
|
{},
|
|
ErrorCode.CONFIG_VALIDATION_FAILED,
|
|
);
|
|
}
|
|
|
|
let config: unknown;
|
|
try {
|
|
config = yaml.load(yamlContent, {
|
|
schema: yaml.FAILSAFE_SCHEMA,
|
|
json: false,
|
|
});
|
|
} catch (yamlError) {
|
|
const errMsg = yamlError instanceof Error ? yamlError.message : String(yamlError);
|
|
throw new PentestError(
|
|
`YAML parsing failed: ${errMsg}`,
|
|
'config',
|
|
false,
|
|
{ originalError: errMsg },
|
|
ErrorCode.CONFIG_PARSE_ERROR,
|
|
);
|
|
}
|
|
|
|
if (config === null || config === undefined) {
|
|
throw new PentestError(
|
|
'Configuration YAML resulted in null/undefined after parsing',
|
|
'config',
|
|
false,
|
|
{},
|
|
ErrorCode.CONFIG_PARSE_ERROR,
|
|
);
|
|
}
|
|
|
|
validateConfig(config as Config);
|
|
return config as Config;
|
|
};
|
|
|
|
// Runs before schema validation so a renamed field fails with a specific "renamed to X"
|
|
// message instead of the generic "additionalProperties" rejection the closed schema
|
|
// would otherwise produce for the old field name.
|
|
function checkDeprecatedFields(config: Config): void {
|
|
const messages: string[] = [];
|
|
|
|
const checkRules = (rules: unknown, where: string): void => {
|
|
if (!Array.isArray(rules)) return;
|
|
rules.forEach((rule, idx) => {
|
|
if (typeof rule !== 'object' || rule === null) return;
|
|
const r = rule as Record<string, unknown>;
|
|
if (r.type === 'path') {
|
|
messages.push(`rules.${where}[${idx}].type: 'path' has been renamed to 'url_path'.`);
|
|
}
|
|
if ('url_path' in r && !('value' in r)) {
|
|
messages.push(`rules.${where}[${idx}]: the rule field 'url_path' has been renamed to 'value'.`);
|
|
}
|
|
});
|
|
};
|
|
|
|
const raw = config as Record<string, unknown>;
|
|
const rules = raw.rules as { avoid?: unknown; focus?: unknown } | undefined;
|
|
checkRules(rules?.avoid, 'avoid');
|
|
checkRules(rules?.focus, 'focus');
|
|
|
|
if (messages.length > 0) {
|
|
throw new PentestError(
|
|
`Configuration uses deprecated fields. Please update:\n - ${messages.join('\n - ')}`,
|
|
'config',
|
|
false,
|
|
{ deprecatedFields: messages },
|
|
ErrorCode.CONFIG_VALIDATION_FAILED,
|
|
);
|
|
}
|
|
}
|
|
|
|
const validateConfig = (config: Config): void => {
|
|
if (!config || typeof config !== 'object') {
|
|
throw new PentestError(
|
|
'Configuration must be a valid object',
|
|
'config',
|
|
false,
|
|
{},
|
|
ErrorCode.CONFIG_VALIDATION_FAILED,
|
|
);
|
|
}
|
|
|
|
if (Array.isArray(config)) {
|
|
throw new PentestError(
|
|
'Configuration must be an object, not an array',
|
|
'config',
|
|
false,
|
|
{},
|
|
ErrorCode.CONFIG_VALIDATION_FAILED,
|
|
);
|
|
}
|
|
|
|
checkDeprecatedFields(config);
|
|
|
|
const isValid = validateSchema(config);
|
|
if (!isValid) {
|
|
const errors = validateSchema.errors || [];
|
|
const errorMessages = formatAjvErrors(errors);
|
|
throw new PentestError(
|
|
`Configuration validation failed:\n - ${errorMessages.join('\n - ')}`,
|
|
'config',
|
|
false,
|
|
{ validationErrors: errorMessages },
|
|
ErrorCode.CONFIG_VALIDATION_FAILED,
|
|
);
|
|
}
|
|
|
|
performSecurityValidation(config);
|
|
|
|
const hasAnySteering =
|
|
!!config.rules ||
|
|
!!config.authentication ||
|
|
!!config.description ||
|
|
!!config.agentic_sast ||
|
|
config.exploit !== undefined ||
|
|
!!config.report ||
|
|
!!config.rules_of_engagement;
|
|
if (!hasAnySteering) {
|
|
console.warn('⚠️ Configuration file contains no steering fields. The pentest will run with all defaults.');
|
|
} else if (config.rules && !config.rules.avoid && !config.rules.focus) {
|
|
console.warn('⚠️ Configuration file contains no rules. The pentest will run without any scoping restrictions.');
|
|
}
|
|
};
|
|
|
|
const performSecurityValidation = (config: Config): void => {
|
|
if (config.authentication) {
|
|
const auth = config.authentication;
|
|
|
|
// Check login_url for dangerous patterns (AJV's "uri" format allows javascript: per RFC 3986)
|
|
if (auth.login_url) {
|
|
for (const pattern of DANGEROUS_PATTERNS) {
|
|
if (pattern.test(auth.login_url)) {
|
|
throw new PentestError(
|
|
`authentication.login_url contains potentially dangerous pattern: ${pattern.source}`,
|
|
'config',
|
|
false,
|
|
{ field: 'login_url', pattern: pattern.source },
|
|
ErrorCode.CONFIG_VALIDATION_FAILED,
|
|
);
|
|
}
|
|
}
|
|
}
|
|
|
|
if (auth.credentials) {
|
|
for (const pattern of DANGEROUS_PATTERNS) {
|
|
if (pattern.test(auth.credentials.username)) {
|
|
throw new PentestError(
|
|
`authentication.credentials.username contains potentially dangerous pattern: ${pattern.source}`,
|
|
'config',
|
|
false,
|
|
{ field: 'credentials.username', pattern: pattern.source },
|
|
ErrorCode.CONFIG_VALIDATION_FAILED,
|
|
);
|
|
}
|
|
}
|
|
}
|
|
|
|
if (auth.login_flow) {
|
|
auth.login_flow.forEach((step, index) => {
|
|
for (const pattern of DANGEROUS_PATTERNS) {
|
|
if (pattern.test(step)) {
|
|
throw new PentestError(
|
|
`authentication.login_flow[${index}] contains potentially dangerous pattern: ${pattern.source}`,
|
|
'config',
|
|
false,
|
|
{ field: `login_flow[${index}]`, pattern: pattern.source },
|
|
ErrorCode.CONFIG_VALIDATION_FAILED,
|
|
);
|
|
}
|
|
}
|
|
});
|
|
}
|
|
}
|
|
|
|
if (config.rules) {
|
|
validateRulesSecurity(config.rules.avoid, 'avoid');
|
|
validateRulesSecurity(config.rules.focus, 'focus');
|
|
|
|
checkForDuplicates(config.rules.avoid || [], 'avoid');
|
|
checkForDuplicates(config.rules.focus || [], 'focus');
|
|
checkForConflicts(config.rules.avoid, config.rules.focus);
|
|
}
|
|
|
|
if (config.description) {
|
|
for (const pattern of DANGEROUS_PATTERNS) {
|
|
if (pattern.test(config.description)) {
|
|
throw new PentestError(
|
|
`description contains potentially dangerous pattern: ${pattern.source}`,
|
|
'config',
|
|
false,
|
|
{ field: 'description', pattern: pattern.source },
|
|
ErrorCode.CONFIG_VALIDATION_FAILED,
|
|
);
|
|
}
|
|
}
|
|
}
|
|
|
|
if (config.rules_of_engagement) {
|
|
for (const pattern of DANGEROUS_PATTERNS) {
|
|
if (pattern.test(config.rules_of_engagement)) {
|
|
throw new PentestError(
|
|
`rules_of_engagement contains potentially dangerous pattern: ${pattern.source}`,
|
|
'config',
|
|
false,
|
|
{ field: 'rules_of_engagement', pattern: pattern.source },
|
|
ErrorCode.CONFIG_VALIDATION_FAILED,
|
|
);
|
|
}
|
|
}
|
|
}
|
|
|
|
if (config.report?.guidance) {
|
|
for (const pattern of DANGEROUS_PATTERNS) {
|
|
if (pattern.test(config.report.guidance)) {
|
|
throw new PentestError(
|
|
`report.guidance contains potentially dangerous pattern: ${pattern.source}`,
|
|
'config',
|
|
false,
|
|
{ field: 'report.guidance', pattern: pattern.source },
|
|
ErrorCode.CONFIG_VALIDATION_FAILED,
|
|
);
|
|
}
|
|
}
|
|
}
|
|
};
|
|
|
|
const validateRulesSecurity = (rules: Rule[] | undefined, ruleType: string): void => {
|
|
if (!rules) return;
|
|
|
|
rules.forEach((rule, index) => {
|
|
for (const pattern of DANGEROUS_PATTERNS) {
|
|
if (pattern.test(rule.value)) {
|
|
throw new PentestError(
|
|
`rules.${ruleType}[${index}].value contains potentially dangerous pattern: ${pattern.source}`,
|
|
'config',
|
|
false,
|
|
{ field: `rules.${ruleType}[${index}].value`, pattern: pattern.source },
|
|
ErrorCode.CONFIG_VALIDATION_FAILED,
|
|
);
|
|
}
|
|
if (rule.description !== undefined && pattern.test(rule.description)) {
|
|
throw new PentestError(
|
|
`rules.${ruleType}[${index}].description contains potentially dangerous pattern: ${pattern.source}`,
|
|
'config',
|
|
false,
|
|
{ field: `rules.${ruleType}[${index}].description`, pattern: pattern.source },
|
|
ErrorCode.CONFIG_VALIDATION_FAILED,
|
|
);
|
|
}
|
|
}
|
|
|
|
validateRuleTypeSpecific(rule, ruleType, index);
|
|
});
|
|
};
|
|
|
|
const validateRuleTypeSpecific = (rule: Rule, ruleType: string, index: number): void => {
|
|
const field = `rules.${ruleType}[${index}].value`;
|
|
|
|
switch (rule.type) {
|
|
case 'url_path':
|
|
if (!rule.value.startsWith('/')) {
|
|
throw new PentestError(
|
|
`${field} for type 'url_path' must start with '/'`,
|
|
'config',
|
|
false,
|
|
{ field, ruleType: rule.type },
|
|
ErrorCode.CONFIG_VALIDATION_FAILED,
|
|
);
|
|
}
|
|
break;
|
|
|
|
case 'code_path':
|
|
if (rule.value.includes('://')) {
|
|
throw new PentestError(
|
|
`${field} for type 'code_path' must not contain a URL protocol (got '${rule.value}')`,
|
|
'config',
|
|
false,
|
|
{ field, ruleType: rule.type },
|
|
ErrorCode.CONFIG_VALIDATION_FAILED,
|
|
);
|
|
}
|
|
break;
|
|
|
|
case 'subdomain':
|
|
case 'domain':
|
|
// Basic domain validation - no slashes allowed
|
|
if (rule.value.includes('/')) {
|
|
throw new PentestError(
|
|
`${field} for type '${rule.type}' cannot contain '/' characters`,
|
|
'config',
|
|
false,
|
|
{ field, ruleType: rule.type },
|
|
ErrorCode.CONFIG_VALIDATION_FAILED,
|
|
);
|
|
}
|
|
// Must contain at least one dot for domains
|
|
if (rule.type === 'domain' && !rule.value.includes('.')) {
|
|
throw new PentestError(
|
|
`${field} for type 'domain' must be a valid domain name`,
|
|
'config',
|
|
false,
|
|
{ field, ruleType: rule.type },
|
|
ErrorCode.CONFIG_VALIDATION_FAILED,
|
|
);
|
|
}
|
|
break;
|
|
|
|
case 'method': {
|
|
const allowedMethods = ['GET', 'POST', 'PUT', 'DELETE', 'PATCH', 'HEAD', 'OPTIONS'];
|
|
if (!allowedMethods.includes(rule.value.toUpperCase())) {
|
|
throw new PentestError(
|
|
`${field} for type 'method' must be one of: ${allowedMethods.join(', ')}`,
|
|
'config',
|
|
false,
|
|
{ field, ruleType: rule.type, allowedMethods },
|
|
ErrorCode.CONFIG_VALIDATION_FAILED,
|
|
);
|
|
}
|
|
break;
|
|
}
|
|
|
|
case 'header':
|
|
if (!rule.value.match(/^[a-zA-Z0-9\-_]+$/)) {
|
|
throw new PentestError(
|
|
`${field} for type 'header' must be a valid header name (alphanumeric, hyphens, underscores only)`,
|
|
'config',
|
|
false,
|
|
{ field, ruleType: rule.type },
|
|
ErrorCode.CONFIG_VALIDATION_FAILED,
|
|
);
|
|
}
|
|
break;
|
|
|
|
case 'parameter':
|
|
if (!rule.value.match(/^[a-zA-Z0-9\-_]+$/)) {
|
|
throw new PentestError(
|
|
`${field} for type 'parameter' must be a valid parameter name (alphanumeric, hyphens, underscores only)`,
|
|
'config',
|
|
false,
|
|
{ field, ruleType: rule.type },
|
|
ErrorCode.CONFIG_VALIDATION_FAILED,
|
|
);
|
|
}
|
|
break;
|
|
}
|
|
};
|
|
|
|
const checkForDuplicates = (rules: Rule[], ruleType: string): void => {
|
|
const seen = new Set<string>();
|
|
rules.forEach((rule, index) => {
|
|
const key = `${rule.type}:${rule.value}`;
|
|
if (seen.has(key)) {
|
|
throw new PentestError(
|
|
`Duplicate rule found in rules.${ruleType}[${index}]: ${rule.type} '${rule.value}'`,
|
|
'config',
|
|
false,
|
|
{ field: `rules.${ruleType}[${index}]`, ruleType: rule.type, value: rule.value },
|
|
ErrorCode.CONFIG_VALIDATION_FAILED,
|
|
);
|
|
}
|
|
seen.add(key);
|
|
});
|
|
};
|
|
|
|
const checkForConflicts = (avoidRules: Rule[] = [], focusRules: Rule[] = []): void => {
|
|
const avoidSet = new Set(avoidRules.map((rule) => `${rule.type}:${rule.value}`));
|
|
|
|
focusRules.forEach((rule, index) => {
|
|
const key = `${rule.type}:${rule.value}`;
|
|
if (avoidSet.has(key)) {
|
|
throw new PentestError(
|
|
`Conflicting rule found: rules.focus[${index}] '${rule.value}' also exists in rules.avoid`,
|
|
'config',
|
|
false,
|
|
{ field: `rules.focus[${index}]`, value: rule.value },
|
|
ErrorCode.CONFIG_VALIDATION_FAILED,
|
|
);
|
|
}
|
|
});
|
|
};
|
|
|
|
const sanitizeRule = (rule: Rule): Rule => {
|
|
const sanitized: Rule = {
|
|
type: rule.type.toLowerCase().trim() as Rule['type'],
|
|
value: rule.value.trim(),
|
|
};
|
|
const description = rule.description?.trim();
|
|
if (description) {
|
|
sanitized.description = description;
|
|
}
|
|
return sanitized;
|
|
};
|
|
|
|
export const distributeConfig = (config: Config | null): DistributedConfig => {
|
|
const avoid = config?.rules?.avoid || [];
|
|
const focus = config?.rules?.focus || [];
|
|
const authentication = config?.authentication || null;
|
|
const description = config?.description?.trim() || '';
|
|
|
|
// The schema types boolean-shaped fields (exploit, report.sarif, agentic_sast.enabled)
|
|
// as a string enum ("true"/"false") rather than JSON boolean, since YAML's FAILSAFE_SCHEMA
|
|
// parses bareword true/false as strings. The string comparison here is intentional, not
|
|
// a leftover from a looser type.
|
|
const exploit = config?.exploit !== undefined ? config.exploit === 'true' : true;
|
|
|
|
const report = {
|
|
// Default on; only an explicit "false" opts out.
|
|
sarif: config?.report?.sarif !== 'false',
|
|
...(config?.report?.min_severity && { min_severity: config.report.min_severity }),
|
|
...(config?.report?.min_confidence && { min_confidence: config.report.min_confidence }),
|
|
...(config?.report?.guidance && { guidance: config.report.guidance.trim() }),
|
|
};
|
|
|
|
const rules_of_engagement = config?.rules_of_engagement?.trim() ?? '';
|
|
|
|
return {
|
|
avoid: avoid.map(sanitizeRule),
|
|
focus: focus.map(sanitizeRule),
|
|
authentication: authentication ? sanitizeAuthentication(authentication) : null,
|
|
description,
|
|
...(config?.agentic_sast?.enabled === 'true' && { agenticSast: true as const }),
|
|
exploit,
|
|
report,
|
|
rules_of_engagement,
|
|
};
|
|
};
|
|
|
|
const sanitizeAuthentication = (auth: Authentication): Authentication => {
|
|
return {
|
|
login_type: auth.login_type.toLowerCase().trim() as Authentication['login_type'],
|
|
login_url: auth.login_url.trim(),
|
|
credentials: {
|
|
username: auth.credentials.username.trim(),
|
|
...(auth.credentials.password && { password: auth.credentials.password }),
|
|
...(auth.credentials.totp_secret && {
|
|
totp_secret: auth.credentials.totp_secret.replace(/\s/g, ''),
|
|
}),
|
|
...(auth.credentials.email_login && {
|
|
email_login: {
|
|
address: auth.credentials.email_login.address.trim(),
|
|
password: auth.credentials.email_login.password,
|
|
...(auth.credentials.email_login.totp_secret && {
|
|
totp_secret: auth.credentials.email_login.totp_secret.replace(/\s/g, ''),
|
|
}),
|
|
},
|
|
}),
|
|
},
|
|
...(auth.login_flow && { login_flow: auth.login_flow.map((step) => step.trim()) }),
|
|
success_condition: {
|
|
type: auth.success_condition.type.toLowerCase().trim() as Authentication['success_condition']['type'],
|
|
value: auth.success_condition.value.trim(),
|
|
},
|
|
};
|
|
};
|