feat: Shannon 3.0 Agentic SAST (#433)

* feat(worker): add agentic static analysis

Add the ten-stage Agentic SAST pipeline, confined repository tools, model runtime, prompt templates, and SARIF export.

Make retries, repair sessions, reduced coverage, usage accounting, and model-output drift durable across Temporal replay and resume. Keep retry diagnostics in their actionable closed vocabulary. Package the Mantis-derived license material with the prompts that require it.

* feat(worker): deduplicate static and runtime findings before exploitation

Parse Agentic SAST SARIF into typed observations, enrich and route those observations, and reconcile them with pentest findings before exploitation.

Publish deterministic exploitation queues with stable lineage, exact-path Git commits, retry-safe manifests, named drop reasons, and confined task formation. Reject duplicate producer IDs before commit and adopt either legal provenance shape after a lost acknowledgement.

* feat(config)!: replace vuln_classes with agentic_sast

Wire Agentic SAST and reconciliation into the main pipeline, persist their durable state, and add the Miscellaneous finding and exploitation lane.

Make scan completion, cancellation, partial outcomes, resume identity, and report recovery use the integrated final workflow contract. Introduce the atomic finalization, ordering, renumbering, compaction, and output services that workflow calls. Keep completed Miscellaneous work and report drafts idempotent across resume, preserve public main's default-on exploit SARIF behavior, and describe stage-fallback candidates without claiming they were exported.

BREAKING CHANGE: `vuln_classes` has been removed. Configs containing it now fail validation, and all five core pentest classes run on every scan.

Workspaces created by Shannon 2.x cannot be resumed. Finish or discard in-flight scans before upgrading, then start a new workspace name.

* perf: overlap static analysis and the Miscellaneous lane with the pentest

Run Agentic SAST alongside vulnerability analysis and run Miscellaneous exploitation alongside the specialist exploitation lanes.

Keep reconciliation dependent on the completed static-analysis result while preserving parallel work everywhere that has no data dependency.

* feat(cli)!: default the scan target and add a JSON error contract

List local scans, resolve the active or most recent workspace automatically, and make logs, status, and stop use one canonical scan identity.

Add stable machine-readable failures, richer status output, explicit help errors, and seven-day Temporal retention. Treat absent Temporal pending-activity failures as absent whether the decoder represents them as `null` or missing.

BREAKING CHANGE: `status --json` now returns a fixed `failureMessage`. Read `partialReasons`, `agenticSast`, and `workflow.log` for diagnostic detail.

* feat(logging): trace tool calls and write a log per agent

Record complete tool-call arguments in the workflow log and project each agent's events into its own durable log.

Add agent listing and agent-specific log tailing while preserving byte-exact output and draining log handles before
activities return.

* feat(worker): standardize severity and reporting guidance in exploit prompts

Give every exploit agent the same status, confidence, severity-reasoning, report-writing, credential-handling, and
scope contract.

Apply the same task-formation and SAST-enrichment procedure to the Miscellaneous lane.

* feat(worker): disclose scan coverage and make reporting auditable

Build on the retry-safe finalization foundation to preserve correct identities, source locations, scan dates,
partial-coverage limitations, and consistent report JSON, Markdown, SARIF, and PDF output.

Report Agentic SAST, reconciliation wall-clock time, stage usage, retry spend, and background work without duplicate
or hardcoded totals. Keep report findings canonical, drop cross-class restatements, name enrichment losses, and render
the executive-summary narrative in the PDF.

* chore(license): attribute Mantis and Pi and refresh the docs

Add the final Mantis and Pi notices, license copies, acknowledgements, and residual copyright updates.

Update the README, maintained documentation, contributor guidance, and hand-maintained mirrors to describe Agentic
SAST, reconciliation, the Miscellaneous lane, current CLI behavior, and the final release contract. Correct stale
workspace and container guidance and annotate long-standing internals for maintainers.

* fix(logging): treat a slash as a word separator in agent labels

* feat(cli)!: rebuild scan status around model work

- show Capella stages beneath the concurrent Agentic SAST phase
- attach reconciliation time to the class row it feeds
- hide completed bookkeeping and the duplicate miscellaneous wrapper
- carry validated child-workflow progress into durable parent state
- derive the terminal tree and status JSON from the same phase shape

BREAKING CHANGE: `status --json` replaces phase `parallel` with `children` and `meta`, adds phase summaries and notes plus agent attachment fields, and removes the `analysis-engines` and `operational-work` phases.

* fix(report): drop the empty Critical Findings section from the PDF summary

* fix(sast): align Capella export with the submit-time code-path contract

The export gate required every code_paths entry to be file:line, but submit only
requires the primary sink to be file:line and accepts bare trace steps. A single
malformed trace step therefore dropped an otherwise-valid finding at export.

- add isValidPrimaryCodePath as the one shared primary-sink contract
- validate only the primary at export; buildResult already drops unusable steps
- route the submit-time validator through the same helper so the two cannot drift

* feat(sast): tolerate hygiene-only Capella reductions instead of going partial

A reduction only makes a run partial when it loses real coverage or a whole
finding. Malformed model output, salvaged turn-limit work, and rejected duplicate
verdicts are recorded as evidence but no longer flip the run to partial.

- add reductionIsTolerable: partial only when genuine-loss counts are nonzero
- drive runCapella's partial reasons and display coverage off non-tolerable ones
- keep every reduction in agenticSast.reductions so nothing is lost as evidence

* feat(logging): record the provider reason for a failed agent turn

A failed provider turn collapsed to AGENT_EXECUTION_FAILED/unknown with the
underlying reason discarded, so a model-side rejection or safeguard was
indistinguishable from a transport fault in the error log.

- add safeProviderTurnDetails: write bounded, non-sensitive fields (provider,
  model, responseId, stop reason, tool-in-flight, category, retryable) to error.log
- gate a sanitized errorMessage snippet behind SHANNON_DEBUG_PROVIDER_ERRORS, off by default
- forward SHANNON_DEBUG_PROVIDER_ERRORS from the CLI into the worker container

* fix(cli): keep shannon logs tailing through a Temporal blip

- End the interactive tail on the log's own terminal marker or Ctrl-C, so a
  transient Temporal outage no longer aborts the command with exit 1.
- Rebuild the memoized Temporal client after a failed poll: a wedged gRPC
  channel was cached forever, so "retrying…" could never reconnect.
- Keep start --follow (CI) bounded — a genuinely dead Temporal still fails
  the run instead of hanging.

* fix(worker): correct PDF finding reporting

- Render OWASP category, authentication state, and remediation
- Omit the redundant per-finding exploited status
- Preserve canonical category and field ordering across report modes
- Continue Proof of Impact numbering across embedded code blocks
- Wrap long PDF code lines without changing canonical report content

* fix: attribute a reconciliation failure to exploitation only

- Stop marking a class's vulnerability-analysis agent failed when that agent
  succeeded and only reconciliation failed; the status tree now renders the
  analysis row completed and the exploitation row failed
- Consume the worker's failedReconciliations signal in the CLI, which the
  mirrored PipelineState already declared but never read
- Correct the class_reconciliation_failed message, which claimed the class's
  analysis results were still in the report when the class is excluded from it

* fix(pi): give each task sub-session its own resource loader to prevent stale extension ctx

* fix(prompts): scope exploit agents to in-band proof, mark OOB-only findings blocked

* fix(cli): reject a shell credential that shadows a gateway config.toml key

* fix(cli): make scan shutdown verifiable

- preselect and persist workflow identity before worker launch
- cancel first, then verify bounded Temporal termination
- reconcile Docker workers with Temporal open workflows
- fail closed on stale images and unavailable lifecycle state
- mark cancellation only after confirmed shutdown

* feat(cli): prompt for setup on a bare npx invocation with no credentials

* fix(cli): don't blame anthropic when no credentials are configured at all

* chore(release): bump beta base version to 3.0.0

* feat(cli): show a 'start your first scan' box in help on a TTY

* docs: refresh README and platform overview for Shannon 3.0

- lead with the 3.0 launch note and rewrite key capabilities around security
  code analysis, the rebuilt terminal experience, native CI/CD, and PDF/SARIF
- recast the editions table as Shannon Open Source against the Keygraph
  Enterprise Platform, stating open source is not a trial edition
- rewrite the platform overview around exhaustive agentic SAST, canonical
  findings, automated remediation, targeted verification, and governance
- add five product screenshots under assets/keygraph-platform/, referenced
  relative to docs/

* docs: add the Shannon naming section and swap in the 3.0 demo GIF

- explain the Claude Shannon information-theory origin under "What is Shannon?"
- point "Shannon in Action" at the 3.0 recording in assets/Shannon3GIF.gif

Both taken from the README half of #438.

* docs: document CI/CD integrations and the reconciled analysis pipeline

- add a CI/CD Integrations section covering the official GitHub Action and
  GitLab component, pipeline artifacts, and exploit-only severity gates
- redraw the architecture section as a Mermaid flow: agentic code analysis
  and recon feed finding reconciliation, then exploitation and reporting
- describe open-source code analysis as a multi-stage agentic workflow and
  reserve parsed-code CPGs and exhaustive verification for Enterprise
- sharpen the privacy wording: results stay local, but model requests carry
  source context to whichever endpoint you configure
- drop the "not recommended" framing on local models and add a section on
  why Shannon complements rather than replaces human pentesters
- regenerate llms-full.txt from the updated README and docs

* docs: add the Photoview benchmark across three models

- Add a "Shannon in Action" table for Photoview 2.4.0 runs on
  DeepSeek v4 Flash, Grok 4.6, and Claude Opus 5, each linking its
  PDF report and SARIF output
- Store the per-model reports under benchmark/
- Link the (forthcoming) benchmark writeup from the section intro

* docs: add the Shannon vs XBOW/Aikido Photoview benchmark writeup

- Add docs/shannon-xbow-aikido-benchmark.md with methodology, per-model
  cost/coverage tables, and links to each model's report and SARIF
- Link the writeup from the README "Shannon in Action" section

* docs: link the benchmark announcement discussion from the README

* fix(readme): restore theme-aware banner, badge, and buttons

* feat!: trigger the Shannon 3.0 major release

---------

Co-authored-by: ezl-keygraph <ezhil@keygraph.io>
This commit is contained in:
Arjun Malleswaranandezl-keygraph authored and GitHub committed 2026-09-02 14:34:26 +05:30
1 parent 6108de3cfc
commit 9767ebe633
267 files changed
+127389 -3268

No files matched your search

+285 -20
View File
@@ -8,8 +8,17 @@
*/
import type { RunningAgent } from '../temporal-client.js';
import { agentClass, PIPELINE, type PipelineState } from './pipeline.js';
import {
AGENTIC_SAST_STAGE_ORDER,
agentClass,
isModelBackedOperation,
type OperationalStageState,
operationFamilyKey,
type PipelineState,
pipelineForState,
} from './pipeline.js';
import type { RenderInput } from './render.js';
import { safeFailureDetail, safeOperationKey, safeOperationLabel } from './safe-fields.js';
export type RunState = 'pending' | 'running' | 'completed' | 'failed' | 'skipped';
@@ -22,14 +31,33 @@ export interface DerivedAgent {
readonly durationMs: number | null;
readonly runningElapsedMs: number | null;
readonly attempt: number | null;
/** The step a running operation row is currently on, merged in from its child activity. */
readonly detail?: string;
/** Reconciliation time for this agent's class, rendered as a trailing `+ duration`.
* Reconciliation is model work that produces this agent's inputs, so it is shown
* attached to the agent it feeds rather than as free-floating background work. */
readonly attachedMs?: number;
/** This class's findings could not be grouped, so each one became its own task. */
readonly ungrouped?: boolean;
readonly error?: string;
}
/** How a phase line summarizes itself: its own wall time, or a k/N tally over its children. */
export type PhaseMetaKind = 'duration' | 'count';
export interface DerivedPhase {
readonly key: string;
readonly label: string;
readonly parallel: boolean;
/** Whether the phase renders its agents as sub-rows. Independent of {@link meta}:
* Agentic SAST lists its stages under a duration, exploitation lists its classes under a tally. */
readonly children: boolean;
readonly meta: PhaseMetaKind;
readonly state: RunState;
/** The phase's own span, when the worker records one for the phase rather than for a single
* agent inside it (Agentic SAST). The phase line presents this exactly like an agent row. */
readonly summary?: DerivedAgent;
/** Rendered after the phase's summary, e.g. to mark work that overlaps other phases. */
readonly note?: string;
readonly agents: readonly DerivedAgent[];
}
@@ -38,8 +66,25 @@ export function isTerminal(status: string): boolean {
return status !== 'RUNNING' && status !== 'UNSPECIFIED';
}
/**
* Whether the class-level failure recorded for this agent's class applies to this agent.
*
* A class failure is recorded against the class as a whole, so it matches both of that class's
* agents. A reconciliation failure, though, happens only after the analysis agent has already
* succeeded, so it belongs to the exploitation lane: attributing it to the analysis row as well
* would report an agent that completed as failed.
*/
function classFailureApplies(name: string, state: PipelineState | null): boolean {
if (!state) return false;
const vulnClass = agentClass(name);
if (!state.failedPipelines.some((f) => f.vulnType === vulnClass)) return false;
const reconciliationFailed = (state.failedReconciliations ?? []).some((r) => r.vulnerabilityClass === vulnClass);
const isAnalysisAgent = name.endsWith('-vuln');
return !(reconciliationFailed && isAnalysisAgent);
}
function isFailedAgent(name: string, state: PipelineState | null): boolean {
return !!state && (state.failedAgent === name || state.failedPipelines.some((f) => f.vulnType === agentClass(name)));
return !!state && (state.failedAgent === name || classFailureApplies(name, state));
}
/** An agent has entered play once it is running, has metrics, or has failed. */
@@ -48,12 +93,12 @@ function isAgentActive(name: string, state: PipelineState | null, running: Set<s
}
/**
* Resolve one agent's state. "Ran" is signalled by a metrics entry, not by
* completedAgents — the workflow lists conditionally-skipped agents (e.g. exploit
* agents when there is nothing to exploit) as completed but records no metrics for
* them. `resolved` is true once we've moved past this agent's phase (the scan is
* terminal, or a later phase is already active), at which point a metric-less,
* non-running agent is skipped rather than still pending.
* Resolve one agent's state. "Ran" is signalled by a metrics entry: a
* conditionally-skipped agent (e.g. an exploit agent when there is nothing to
* exploit) records no metrics, and the workflow tracks it in skippedAgents rather
* than completedAgents. `resolved` is true once we've moved past this agent's phase
* (the scan is terminal, or a later phase is already active), at which point a
* metric-less, non-running agent is skipped rather than still pending.
*/
function agentState(name: string, state: PipelineState | null, running: Set<string>, resolved: boolean): RunState {
if (running.has(name)) return 'running';
@@ -62,13 +107,15 @@ function agentState(name: string, state: PipelineState | null, running: Set<stri
return resolved ? 'skipped' : 'pending';
}
/**
* Only the presence of a class failure is used here, never its `.error` text: that string is
* the worker's raw error for the failed class, not vetted for display, so it is reduced to
* a boolean before reaching safeFailureDetail's fixed sentence.
*/
function agentError(name: string, state: PipelineState | null, byAgent: Map<string, RunningAgent>): string | undefined {
const failed = state?.failedPipelines.find((f) => f.vulnType === agentClass(name));
return (
failed?.error ??
byAgent.get(name)?.lastFailure ??
(state?.failedAgent === name ? (state.error ?? undefined) : undefined)
);
const hasFailure =
classFailureApplies(name, state) || byAgent.get(name)?.lastFailure !== undefined || state?.failedAgent === name;
return safeFailureDetail(hasFailure);
}
/** Scan wall-clock elapsed ms: recorded duration for a closed scan, live elapsed for a running one. */
@@ -99,16 +146,17 @@ export function phaseGlyphState(states: readonly RunState[]): RunState {
* class had anything to exploit), not still pending.
*/
export function deriveAgentStates(input: RenderInput): Map<string, RunState> {
const runningSet = new Set(input.running.map((r) => r.agent));
const pipeline = pipelineForState(input.state);
const runningSet = new Set(input.running.filter((runner) => runner.kind === 'agent').map((runner) => runner.agent));
const terminal = isTerminal(input.temporalStatus);
let frontier = -1;
PIPELINE.forEach((phase, idx) => {
pipeline.forEach((phase, idx) => {
if (phase.agents.some((a) => isAgentActive(a.name, input.state, runningSet))) frontier = idx;
});
const states = new Map<string, RunState>();
for (const [phaseIdx, phase] of PIPELINE.entries()) {
for (const [phaseIdx, phase] of pipeline.entries()) {
const resolved = terminal || phaseIdx < frontier;
for (const agent of phase.agents) {
states.set(agent.name, agentState(agent.name, input.state, runningSet, resolved));
@@ -117,6 +165,52 @@ export function deriveAgentStates(input: RenderInput): Map<string, RunState> {
return states;
}
/** Which operation families have a running parent stage, and the step to show on it. */
interface OperationFamilyView {
/** Families whose parent stage row already represents their child activities. */
readonly runningFamilies: ReadonlySet<string>;
/** Family to current step, present only where the child activities agree on one. */
readonly stepByFamily: ReadonlyMap<string, string>;
}
/**
* Resolve the parent stage rows that own their family's child activities. A family only
* resolves to a step when its running children agree: several classes reconcile at once and
* their pending activities carry no class, so a family caught mid-stride shows its parent
* rows without a step rather than attributing one to the wrong class.
*/
function operationFamilyView(
running: readonly RunningAgent[],
persistedOperations: readonly OperationalStageState[],
): OperationFamilyView {
const runningFamilies = new Set(
persistedOperations
.filter((operation) => operation.status === 'running')
.map((operation) => operationFamilyKey(operation.key)),
);
const labelsByFamily = new Map<string, Set<string>>();
for (const runner of running) {
if (runner.kind !== 'operation' || runner.parentKey === undefined) continue;
if (!runningFamilies.has(runner.parentKey)) continue;
const labels = labelsByFamily.get(runner.parentKey) ?? new Set<string>();
labels.add(runner.label);
labelsByFamily.set(runner.parentKey, labels);
}
const stepByFamily = new Map<string, string>();
for (const [family, labels] of labelsByFamily) {
const [onlyLabel] = labels;
if (labels.size === 1 && onlyLabel !== undefined) stepByFamily.set(family, lowercaseFirst(onlyLabel));
}
return { runningFamilies, stepByFamily };
}
/** Progress labels are written to start a row; as a detail they continue a sentence. */
function lowercaseFirst(label: string): string {
return label.charAt(0).toLowerCase() + label.slice(1);
}
/**
* Full structured view of the pipeline: every agent's state plus the raw
* metrics/timing needed to present it, and each phase's collapsed state.
@@ -124,8 +218,9 @@ export function deriveAgentStates(input: RenderInput): Map<string, RunState> {
export function derivePipeline(input: RenderInput, now: number): DerivedPhase[] {
const states = deriveAgentStates(input);
const byAgent = new Map(input.running.map((r) => [r.agent, r]));
const pipeline = pipelineForState(input.state);
return PIPELINE.map((phase) => {
const agentPhases = pipeline.map((phase) => {
const agents = phase.agents.map((a): DerivedAgent => {
const state = states.get(a.name) ?? 'pending';
const metrics = input.state?.agentMetrics[a.name];
@@ -145,11 +240,181 @@ export function derivePipeline(input: RenderInput, now: number): DerivedPhase[]
return {
key: phase.key,
label: phase.label,
parallel: phase.parallel,
children: phase.parallel,
meta: phase.parallel ? ('count' as const) : ('duration' as const),
state: phaseGlyphState(agents.map((ag) => ag.state)),
agents,
};
});
// Operational rows merge two sources: stages the worker has persisted (durable truth,
// including terminal outcomes) and pending activities whose stage record has not landed
// yet. Persisted keys win, so a stage is never listed twice while the two views overlap.
const persistedOperations = Object.values(input.state?.operationalStages ?? {});
const persistedKeys = new Set(persistedOperations.map((operation) => operation.key));
const { runningFamilies, stepByFamily } = operationFamilyView(input.running, persistedOperations);
const unpersistedRunning = input.running
.filter((runner) => runner.kind === 'operation' && !persistedKeys.has(runner.agent))
// A child activity whose family already has a running parent stage is that stage's current
// step, not separate work: the parent row below represents it, with the step as its detail
// where the family's children agree on one. Without such a parent it keeps its own row.
.filter((runner) => runner.parentKey === undefined || !runningFamilies.has(runner.parentKey))
.map((runner) => ({
key: runner.agent,
label: runner.label,
status: 'running' as const,
...(runner.startedAt !== undefined && { startedAt: runner.startedAt }),
...(runner.lastFailure !== undefined && { error: safeFailureDetail(true) }),
}));
const operationalAgents: DerivedAgent[] = [...persistedOperations, ...unpersistedRunning].map((operation) => {
const runner = byAgent.get(operation.key);
const operationState = operation.status as RunState;
const persistedDurationMs = 'durationMs' in operation ? (operation.durationMs ?? null) : null;
const detail = operationState === 'running' ? stepByFamily.get(operationFamilyKey(operation.key)) : undefined;
return {
name: safeOperationKey(operation.key),
label: safeOperationLabel(operation.label),
state: operationState,
durationMs: operationState === 'completed' ? persistedDurationMs : null,
runningElapsedMs:
operationState === 'running' && operation.startedAt !== undefined ? now - operation.startedAt : null,
attempt: operationState === 'running' ? (runner?.attempt ?? null) : null,
...(detail !== undefined && { detail }),
...(operation.error !== undefined && { error: safeFailureDetail(true) }),
};
});
// Operational rows are not peers of the agents. Each one is either model work that
// belongs to an agent (reconciliation), model work that belongs to the SAST engine
// (its stages), or bookkeeping that only earns a row when it is stuck or broken.
return assemblePhases(agentPhases, operationalAgents);
}
/** Reconciliation wall time per vulnerability class, plus the classes whose grouping degraded. */
interface ReconciliationView {
readonly durationByClass: ReadonlyMap<string, number>;
readonly ungroupedClasses: ReadonlySet<string>;
}
function reconciliationView(operations: readonly DerivedAgent[]): ReconciliationView {
const durationByClass = new Map<string, number>();
const ungroupedClasses = new Set<string>();
for (const operation of operations) {
if (operationFamilyKey(operation.name) !== 'reconciliation') continue;
const [, vulnerabilityClass] = operation.name.split(':');
if (vulnerabilityClass === undefined) continue;
if (operation.name.endsWith(':fallback')) {
ungroupedClasses.add(vulnerabilityClass);
continue;
}
if (operation.durationMs !== null) durationByClass.set(vulnerabilityClass, operation.durationMs);
}
return { durationByClass, ungroupedClasses };
}
/** Attach each class's reconciliation time to the agent row it feeds. */
function withReconciliation(phase: DerivedPhase, view: ReconciliationView): DerivedPhase {
const agents = phase.agents.map((agent): DerivedAgent => {
const vulnerabilityClass = agentClass(agent.name);
const attachedMs = view.durationByClass.get(vulnerabilityClass);
const ungrouped = view.ungroupedClasses.has(vulnerabilityClass);
return {
...agent,
...(attachedMs !== undefined && { attachedMs }),
...(ungrouped && { ungrouped }),
};
});
return { ...phase, agents };
}
/**
* Build the Agentic SAST phase from the aggregate span the parent workflow records and the
* per-stage rows the SAST child signals up. Scans that predate stage signalling have the
* aggregate but no stages, and render as a bare phase line rather than an error.
*/
function agenticSastPhase(operations: readonly DerivedAgent[]): DerivedPhase | undefined {
const aggregate = operations.find((operation) => operation.name === 'agentic-sast');
if (aggregate === undefined) return undefined;
const byStage = new Map<string, DerivedAgent>();
for (const operation of operations) {
const [family, stage] = operation.name.split(':');
if (family !== 'agentic-sast' || stage === undefined) continue;
// The worker's label is the scan log's Title Case form. These rows sit beside the
// lowercase class rows below them, so they read in the same register here.
byStage.set(stage, { ...operation, label: lowercaseFirst(operation.label) });
}
// Run order, not insertion order: a resumed or replayed run can persist stages out of order.
const stages = AGENTIC_SAST_STAGE_ORDER.map((stage) => byStage.get(stage)).filter(
(stage): stage is DerivedAgent => stage !== undefined,
);
return {
key: 'agentic-sast',
label: 'Agentic SAST',
children: stages.length > 0,
meta: 'duration',
state: aggregate.state,
summary: aggregate,
// It shares wall time with the pentest phases below it, so the times do not add up
// in sequence. Saying so is cheaper than a layout that pretends to be two columns.
note: 'concurrent',
agents: stages,
};
}
/**
* Bookkeeping rows worth showing. A deterministic stage that has completed says nothing —
* it can only ever read 0s — but one that is still running, or that failed, is exactly what
* an operator needs to see, so those keep a row under the phase they belong to.
*/
function troubledReportSteps(operations: readonly DerivedAgent[]): readonly DerivedAgent[] {
return operations.filter((operation) => {
if (isModelBackedOperation(operation.name)) return false;
if (operationFamilyKey(operation.name) !== 'report') return false;
return operation.state === 'running' || operation.state === 'failed';
});
}
/**
* Fold operational rows into the agent phases. Nothing here becomes a bucket of its own:
* every surviving row is either a SAST stage, time attached to an agent, or a report step
* that is currently in trouble.
*/
function assemblePhases(agentPhases: readonly DerivedPhase[], operations: readonly DerivedAgent[]): DerivedPhase[] {
const view = reconciliationView(operations);
// Reconciliation produces the exploitation queue, so its time belongs on the exploitation
// row it feeds. With exploitation off there is no such row, and it falls back to the
// analysis row for the same class so the time is never silently dropped.
const attachTo = agentPhases.some((phase) => phase.key === 'exploitation')
? 'exploitation'
: 'vulnerability-analysis';
const reportSteps = troubledReportSteps(operations);
const phases = agentPhases.map((phase) => {
if (phase.key === attachTo) return withReconciliation(phase, view);
if (phase.key === 'reporting' && reportSteps.length > 0) {
// The report agent stays on the phase line it already titles; the steps in trouble
// become its children, so nothing is listed twice.
const summary = phase.agents[0];
return {
...phase,
children: true,
...(summary !== undefined && { summary }),
state: phaseGlyphState([...phase.agents, ...reportSteps].map((row) => row.state)),
agents: reportSteps,
};
}
return phase;
});
const sast = agenticSastPhase(operations);
if (sast === undefined) return phases;
// Agentic SAST starts with the scan and runs alongside the pentest, so it reads after
// the login check rather than appended past Reporting where it never ran.
const afterAuth = phases.findIndex((phase) => phase.key === 'auth-validation') + 1;
return [...phases.slice(0, afterAuth), sast, ...phases.slice(afterAuth)];
}
export { agentError };
+262 -2
View File
@@ -8,6 +8,7 @@
* - apps/worker/src/temporal/activities.ts (the run*Agent activity names → `activityType`)
* - apps/worker/src/temporal/shared.ts (PipelineState / PipelineSummary)
* - apps/worker/src/types/metrics.ts (AgentMetrics)
* - apps/worker/src/types/run-state.ts (PartialReasonView)
*/
export interface AgentSpec {
@@ -26,6 +27,18 @@ export interface PhaseSpec {
readonly agents: readonly AgentSpec[];
}
export interface ActivityProgressSpec {
readonly key: string;
readonly label: string;
readonly kind: 'agent' | 'operation';
/**
* Operation rows whose work is already represented by a persisted parent stage. The parent
* owns the row; this activity supplies the step shown as its detail. Parent stage keys are
* the family key itself or the family key followed by ':' and a class or stage suffix.
*/
readonly parentKey?: string;
}
/** The pipeline phases in execution order, each with its agents. */
export const PIPELINE: readonly PhaseSpec[] = [
{
@@ -80,9 +93,183 @@ export const PIPELINE: readonly PhaseSpec[] = [
},
];
/** Temporal activity type name → canonical agent name, for mapping pendingActivities. */
const MISCELLANEOUS_EXPLOIT_AGENT: AgentSpec = {
name: 'miscellaneous-exploit',
label: 'miscellaneous',
activityType: 'runMiscellaneousExploitAgent',
};
/**
* Shape the static PIPELINE to one scan's durable truth. expectedAgents, persisted by the
* worker at scan start, names every exploit agent the scan can ever run: exploit rows it
* excludes are dropped, 'miscellaneous-exploit' is appended only once the miscellaneous pipeline has
* admitted findings, and a phase left with no agents disappears entirely. Without state
* (the scan has not initialized durable state yet) the full static pipeline is the best
* available guess.
*/
export function pipelineForState(state: PipelineState | null): readonly PhaseSpec[] {
if (state?.expectedAgents === undefined) return PIPELINE;
const expected = new Set(state.expectedAgents);
return PIPELINE.map((phase) => {
if (phase.key !== 'exploitation') return phase;
const agents = phase.agents.filter((agent) => expected.has(agent.name));
if (expected.has(MISCELLANEOUS_EXPLOIT_AGENT.name)) agents.push(MISCELLANEOUS_EXPLOIT_AGENT);
return { ...phase, agents };
}).filter((phase) => phase.agents.length > 0);
}
const AGENT_ACTIVITY_PROGRESS: Readonly<Record<string, ActivityProgressSpec>> = Object.fromEntries(
[...PIPELINE.flatMap((phase) => phase.agents), MISCELLANEOUS_EXPLOIT_AGENT].map((agent) => [
agent.activityType,
{ key: agent.name, label: agent.label, kind: 'agent' },
]),
);
/** Families whose per-class or per-stage work is already carried by one persisted stage row. */
const RECONCILIATION_PARENT_KEY = 'reconciliation';
const AGENTIC_SAST_PARENT_KEY = 'agentic-sast';
// Every production activity that is not an agent run must have a row here. describeScan
// throws on an unmapped activity type, so adding a worker activity without updating this
// table breaks `shannon status` loudly instead of hiding the new work. The authoritative
// name lists live in apps/worker/src/temporal/worker.ts,
// apps/worker/src/temporal/reconcile-activity-types.ts, and
// apps/worker/src/ai/sast/capella/temporal/activity-types.ts.
const OPERATION_ACTIVITY_PROGRESS: Readonly<Record<string, ActivityProgressSpec>> = {
runPreflightValidation: { key: 'preflight', label: 'Preflight validation', kind: 'operation' },
syncPlaywrightStealthConfig: { key: 'preflight', label: 'Browser setup', kind: 'operation' },
initDeliverableGit: { key: 'scan-initialization', label: 'Initialize deliverables', kind: 'operation' },
syncCodePathDenyRules: { key: 'scan-initialization', label: 'Apply source rules', kind: 'operation' },
initializeDurableScanState: { key: 'durable-state', label: 'Saving scan state', kind: 'operation' },
persistMiscellaneousOutcome: {
key: 'miscellaneous-pipeline',
label: 'Including miscellaneous findings',
kind: 'operation',
},
initializeReportProgress: { key: 'report:initialize', label: 'Initialize report state', kind: 'operation' },
renumberClassFindings: { key: 'report:renumber', label: 'Renumber findings', kind: 'operation' },
assembleReportActivity: { key: 'report:assemble', label: 'Assemble report inputs', kind: 'operation' },
compactReportFindings: { key: 'report:compact', label: 'Compact report findings', kind: 'operation' },
persistCanonicalReportProgress: { key: 'report:checkpoint', label: 'Saving report progress', kind: 'operation' },
finalizeReportOutputs: { key: 'report:finalize', label: 'Finalize report outputs', kind: 'operation' },
persistFinalizedReportProgress: { key: 'report:terminal', label: 'Saving final report state', kind: 'operation' },
surfaceReportOutputs: { key: 'report:surface', label: 'Surface customer report', kind: 'operation' },
checkExploitationQueue: { key: 'queue-check', label: 'Check exploitation queue', kind: 'operation' },
loadResumeState: { key: 'resume-validation', label: 'Validate resume state', kind: 'operation' },
restoreGitCheckpoint: { key: 'resume-restore', label: 'Restore checkpoint', kind: 'operation' },
registerResumeAttempt: { key: 'resume-registration', label: 'Register resume', kind: 'operation' },
recordResumeAttempt: { key: 'resume-registration', label: 'Record resume', kind: 'operation' },
logPhaseTransition: { key: 'audit-log', label: 'Update audit log', kind: 'operation' },
logWorkflowComplete: { key: 'audit-log', label: 'Finalize audit log', kind: 'operation' },
saveCheckpoint: { key: 'checkpoint', label: 'Save checkpoint', kind: 'operation' },
seedEmptyProducerQueue: {
key: 'miscellaneous-pipeline',
label: 'Preparing miscellaneous findings',
kind: 'operation',
},
prepareClassReconciliation: {
key: 'reconciliation',
label: 'Preparing findings',
kind: 'operation',
parentKey: RECONCILIATION_PARENT_KEY,
},
enrichClassSastObservations: {
key: 'reconciliation',
label: 'Adding code context',
kind: 'operation',
parentKey: RECONCILIATION_PARENT_KEY,
},
formClassExploitTasks: {
key: 'reconciliation',
label: 'Grouping into test cases',
kind: 'operation',
parentKey: RECONCILIATION_PARENT_KEY,
},
materializeClassExploitTasks: {
key: 'reconciliation',
label: 'Writing test cases',
kind: 'operation',
parentKey: RECONCILIATION_PARENT_KEY,
},
publishClassReconciliationOss: {
key: 'reconciliation',
label: 'Saving results',
kind: 'operation',
parentKey: RECONCILIATION_PARENT_KEY,
},
capellaArchitecture: {
key: 'agentic-sast:architecture',
label: 'Mapping architecture',
kind: 'operation',
parentKey: AGENTIC_SAST_PARENT_KEY,
},
capellaThreatModel: {
key: 'agentic-sast:threat-model',
label: 'Modelling threats',
kind: 'operation',
parentKey: AGENTIC_SAST_PARENT_KEY,
},
capellaPlan: {
key: 'agentic-sast:plan',
label: 'Planning the review',
kind: 'operation',
parentKey: AGENTIC_SAST_PARENT_KEY,
},
capellaResearch: {
key: 'agentic-sast:research',
label: 'Researching code',
kind: 'operation',
parentKey: AGENTIC_SAST_PARENT_KEY,
},
capellaDedupe: {
key: 'agentic-sast:dedupe',
label: 'Merging duplicates',
kind: 'operation',
parentKey: AGENTIC_SAST_PARENT_KEY,
},
capellaReview: {
key: 'agentic-sast:review',
label: 'Reviewing findings',
kind: 'operation',
parentKey: AGENTIC_SAST_PARENT_KEY,
},
capellaCritic: {
key: 'agentic-sast:critic',
label: 'Critiquing findings',
kind: 'operation',
parentKey: AGENTIC_SAST_PARENT_KEY,
},
capellaConfirm: {
key: 'agentic-sast:confirm',
label: 'Confirming findings',
kind: 'operation',
parentKey: AGENTIC_SAST_PARENT_KEY,
},
capellaCalibrate: {
key: 'agentic-sast:calibrate',
label: 'Calibrating risk',
kind: 'operation',
parentKey: AGENTIC_SAST_PARENT_KEY,
},
capellaExport: {
key: 'agentic-sast:export',
label: 'Exporting findings',
kind: 'operation',
parentKey: AGENTIC_SAST_PARENT_KEY,
},
};
/** Complete production activity mirror. Unknown names are errors, never hidden progress. */
export const ACTIVITY_TO_PROGRESS: Readonly<Record<string, ActivityProgressSpec>> = Object.freeze({
...AGENT_ACTIVITY_PROGRESS,
...OPERATION_ACTIVITY_PROGRESS,
});
/** Agent-only projection of ACTIVITY_TO_PROGRESS: activity type name to canonical agent name. */
export const ACTIVITY_TO_AGENT: Readonly<Record<string, string>> = Object.fromEntries(
PIPELINE.flatMap((phase) => phase.agents.map((agent) => [agent.activityType, agent.name])),
Object.entries(ACTIVITY_TO_PROGRESS)
.filter(([, progress]) => progress.kind === 'agent')
.map(([activityType, progress]) => [activityType, progress.key]),
);
/** The vuln/exploit class of an agent (e.g. "authz-vuln" → "authz"), for failedPipelines matching. */
@@ -100,11 +287,65 @@ export interface AgentMetrics {
readonly skipped?: boolean;
}
export interface OperationalStageState {
readonly key: string;
readonly label: string;
readonly status: 'pending' | 'running' | 'completed' | 'failed' | 'skipped';
readonly startedAt?: number;
readonly durationMs?: number;
readonly error?: string;
}
/** Family key a persisted operational stage belongs to, e.g. `reconciliation:xss` to `reconciliation`. */
export function operationFamilyKey(stageKey: string): string {
const separator = stageKey.indexOf(':');
return separator === -1 ? stageKey : stageKey.slice(0, separator);
}
/** The Capella stages that get a progress row, in run order. Mirrors CAPELLA_PROGRESS_STAGES
* in apps/worker/src/ai/sast/types.ts — the deterministic `export` stage is not among them. */
export const AGENTIC_SAST_STAGE_ORDER: readonly string[] = [
'architecture',
'threat-model',
'plan',
'research',
'dedupe',
'review',
'critic',
'confirm',
'calibrate',
];
/**
* Whether an operational stage represents model work rather than bookkeeping.
*
* Only the agentic-SAST stages and per-class reconciliation run a model; every other
* operational stage is a git commit or a durable-state write that can only ever record
* sub-second wall time. The progress tree shows model work, so this is what decides
* whether a stage is worth a row at all.
*/
export function isModelBackedOperation(stageKey: string): boolean {
const family = operationFamilyKey(stageKey);
if (family === 'agentic-sast') return true;
// A `reconciliation:<class>:fallback` marker records a degradation, not a model span.
return family === 'reconciliation' && !stageKey.endsWith(':fallback');
}
export interface PipelineSummary {
readonly totalCostUsd: number;
readonly totalDurationMs: number; // Wall-clock (end - start)
readonly totalTurns: number;
readonly agentCount: number;
/** False when operational (Capella/reconciliation) spend is known to be incomplete. */
readonly usageAccountingComplete?: boolean;
}
/** One durable degradation reason with its derived safe message (mirror of PartialReasonView). */
export interface PartialReasonView {
readonly code: string;
readonly vulnerabilityClass?: string;
readonly stage?: string;
readonly message: string;
}
export type PipelineStatus = 'running' | 'completed' | 'failed' | 'cancelled' | 'partial';
@@ -114,10 +355,29 @@ export interface PipelineState {
readonly currentPhase: string | null;
readonly currentAgent: string | null;
readonly completedAgents: string[];
readonly expectedAgents?: string[];
readonly participatingClasses?: string[];
readonly failedPipelines: { vulnType: string; error: string }[];
readonly failedReconciliations?: { vulnerabilityClass: string; error: string }[];
readonly failedAgent: string | null;
readonly error: string | null;
readonly startTime: number;
readonly agentMetrics: Record<string, AgentMetrics>;
readonly operationalMetrics?: Record<string, AgentMetrics>;
readonly operationalStages?: Record<string, OperationalStageState>;
/** `error` is the worker's sanitized failure sentence, safe to print verbatim. */
readonly agenticSast?: {
readonly status: string;
readonly durationMs?: number;
/** Reader-facing name of the failed stage, already projected by the worker. */
readonly failedStageLabel?: string;
readonly error?: string;
readonly errorCode?: string;
/** Usage-accounting warnings projected by the worker; empty when the ledger reconciled. */
readonly warnings?: readonly string[];
};
readonly nonFatalFailures?: { readonly phase: string; readonly error: string }[];
/** Ordered durable degradation reasons with safe messages; empty or absent for full success. */
readonly partialReasons?: readonly PartialReasonView[];
readonly summary: PipelineSummary | null;
}
+99 -28
View File
@@ -10,9 +10,9 @@
import { BOLD, DIM, GOLD, paint, RED, YELLOW } from '../colors.js';
import { commandPrefix } from '../mode.js';
import type { RunningAgent } from '../temporal-client.js';
import { agentError, deriveAgentStates, isTerminal, phaseGlyphState, type RunState, scanElapsedMs } from './derive.js';
import { inlineFailureReason } from './failure.js';
import { PIPELINE, type PipelineState } from './pipeline.js';
import { derivePipeline, isTerminal, type RunState, scanElapsedMs } from './derive.js';
import type { PipelineState } from './pipeline.js';
import { safeAgenticSast, safeCliIdentifier, safePartialReasons, safeTerminalFailure } from './safe-fields.js';
export interface RenderInput {
readonly workspace: string;
@@ -68,7 +68,7 @@ function truncate(text: string, max: number): string {
/** Temporal Web UI, published by compose on 8233; deep-links to the workflow when its id is known. */
function temporalDashboardUrl(workflowId: string | undefined): string {
const base = 'http://localhost:8233';
return workflowId ? `${base}/namespaces/default/workflows/${workflowId}` : base;
return workflowId ? `${base}/namespaces/default/workflows/${safeCliIdentifier(workflowId)}` : base;
}
// === Glyphs & status ===
@@ -95,6 +95,12 @@ const STATE_COLOR: Record<RunState, string> = {
skipped: COLORS.dim,
};
/** Column width for an agent or background-work label inside a phase. */
const AGENT_LABEL_WIDTH = 18;
/** Inline budget for a failure sentence, wide enough to carry a whole first sentence. */
const FAILURE_DETAIL_WIDTH = 120;
/** Braille spinner frames for running agents — the clack loader style. */
const SPINNER_FRAMES = ['⠋', '⠙', '⠹', '⠸', '⠼', '⠴', '⠦', '⠧', '⠇', '⠏'] as const;
@@ -112,42 +118,68 @@ function statusBadge(input: RenderInput, opts: RenderOptions): string {
const workflowStatus = input.state?.status;
if (!isTerminal(input.temporalStatus)) return paint('running', COLORS.gold, opts.color);
if (workflowStatus === 'partial') return paint('partial', COLORS.yellow, opts.color);
if (workflowStatus === 'cancelled') return paint('cancelled', COLORS.yellow, opts.color);
if (input.temporalStatus === 'COMPLETED') return paint('completed', COLORS.gold, opts.color);
if (input.temporalStatus === 'TERMINATED') return paint('stopped', COLORS.yellow, opts.color);
if (input.temporalStatus === 'CANCELLED' || input.temporalStatus === 'CANCELED') {
return paint('cancelled', COLORS.yellow, opts.color);
}
if (input.temporalStatus === 'TIMED_OUT') return paint('timed out', COLORS.red, opts.color);
return paint('FAILED', COLORS.red, opts.color);
return paint('failed', COLORS.red, opts.color);
}
// === Line builders ===
/** The parts of a derived row agentMeta reads beyond its state and metrics. */
interface RowExtras {
readonly runningElapsedMs?: number | null;
readonly attachedMs?: number;
readonly ungrouped?: boolean;
}
function agentMeta(
state: RunState,
metrics: { durationMs: number } | undefined,
runner: RunningAgent | undefined,
error: string | undefined,
opts: RenderOptions,
step?: string,
extras?: RowExtras,
): string {
if (state === 'completed') {
const duration = metrics?.durationMs != null ? formatDuration(metrics.durationMs) : 'done';
return paint(duration, COLORS.dim, opts.color);
return paint(`${duration}${attachedSuffix(extras)}`, COLORS.dim, opts.color);
}
if (state === 'running') {
const parts = ['running'];
if (runner?.startedAt !== undefined) parts.push(formatDuration(opts.now - runner.startedAt));
if (step !== undefined) parts.push(step);
// An operational row carries its own elapsed time: it is derived from the persisted stage
// span, and has no pending activity on the parent workflow to read a start time from.
const elapsedMs =
runner?.startedAt !== undefined ? opts.now - runner.startedAt : (extras?.runningElapsedMs ?? null);
if (elapsedMs !== null) parts.push(formatDuration(elapsedMs));
if (runner && runner.attempt > 1) parts.push(`retry ${runner.attempt}`);
return paint(parts.join(' · '), COLORS.gold, opts.color);
}
if (state === 'failed') {
const detail = error ? ` · ${truncate(error, 46)}` : '';
const detail = error ? ` · ${truncate(error, FAILURE_DETAIL_WIDTH)}` : '';
return paint(`failed${detail}`, COLORS.red, opts.color);
}
if (state === 'skipped') return paint('skipped', COLORS.dim, opts.color);
return paint('queued', COLORS.dim, opts.color);
}
/**
* Time a reconciliation lane contributed to this agent's class, shown as `+ duration` on the
* row it feeds. `ungrouped` marks a class whose findings could not be grouped, so each one
* was tested separately and duplicates are expected.
*/
function attachedSuffix(extras: RowExtras | undefined): string {
if (extras === undefined) return '';
const time = extras.attachedMs === undefined ? '' : ` + ${formatDuration(extras.attachedMs)}`;
return extras.ungrouped ? `${time} · ungrouped` : time;
}
function phaseMeta(states: readonly RunState[], inPlay: number, parallel: boolean, opts: RenderOptions): string {
if (states.every((s) => s === 'pending')) return paint('pending', COLORS.dim, opts.color);
if (states.every((s) => s === 'skipped')) return paint('skipped', COLORS.dim, opts.color);
@@ -163,35 +195,43 @@ function phaseMeta(states: readonly RunState[], inPlay: number, parallel: boolea
/** Render the full progress frame as one string (no trailing newline). */
export function renderScan(input: RenderInput, opts: RenderOptions): string {
const byAgent = new Map(input.running.map((r) => [r.agent, r]));
const stateMap = deriveAgentStates(input);
const phases = derivePipeline(input, opts.now);
const lines: string[] = ['', ...headerLines(input, opts), ''];
const metaFor = (name: string, state: RunState): string =>
agentMeta(state, input.state?.agentMetrics[name], byAgent.get(name), agentError(name, input.state, byAgent), opts);
// Only agents that have actually entered play are shown; pending/skipped ones stay hidden.
const inPlay = (s: RunState): boolean => s === 'running' || s === 'completed' || s === 'failed';
for (const phase of PIPELINE) {
const states = phase.agents.map((a) => stateMap.get(a.name) ?? 'pending');
for (const phase of phases) {
const states = phase.agents.map((agent) => agent.state);
const playing = states.filter(inPlay).length;
const phaseRunState: RunState = phaseGlyphState(states);
const phaseRunState = phase.state;
const metaFor = (agent: (typeof phase.agents)[number]): string => {
const metrics = agent.durationMs === null ? undefined : { durationMs: agent.durationMs };
return agentMeta(agent.state, metrics, byAgent.get(agent.name), agent.error, opts, agent.detail, agent);
};
// A single-agent phase carries that agent's own duration/cost on the phase line once it
// starts; a parallel phase gets a "k/N done" summary over the agents in play.
// A phase summarizes itself by wall time or by a "k/N done" tally. A phase with its own
// recorded span (Agentic SAST) presents it like any agent row; otherwise a single-agent
// phase borrows its one agent's duration once that agent starts.
const first = phase.agents[0];
const firstState = states[0];
const phaseMetaStr =
!phase.parallel && first && firstState && inPlay(firstState)
? metaFor(first.name, firstState)
: phaseMeta(states, playing, phase.parallel, opts);
lines.push(` ${glyph(phaseRunState, opts)} ${phase.label.padEnd(26)}${phaseMetaStr}`);
const borrowed = first && firstState && inPlay(firstState) ? metaFor(first) : undefined;
const durationMeta = phase.summary === undefined ? borrowed : metaFor(phase.summary);
const summaryMeta =
phase.meta === 'duration' && durationMeta !== undefined
? durationMeta
: phaseMeta(states, playing, phase.meta === 'count', opts);
const note = phase.note === undefined ? '' : paint(` · ${phase.note}`, COLORS.dim, opts.color);
lines.push(` ${glyph(phaseRunState, opts)} ${phase.label.padEnd(26)}${summaryMeta}${note}`);
if (!phase.parallel) continue;
if (!phase.children) continue;
for (let i = 0; i < phase.agents.length; i++) {
const agent = phase.agents[i];
const state = states[i];
if (!agent || !state || !inPlay(state)) continue;
lines.push(` ${glyph(state, opts)} ${agent.label.padEnd(18)}${metaFor(agent.name, state)}`);
// Two trailing spaces before padding, so a label wider than the column still separates
// from its meta text; a label inside the column pads to the same width as before.
lines.push(` ${glyph(state, opts)} ${`${agent.label} `.padEnd(AGENT_LABEL_WIDTH)}${metaFor(agent)}`);
}
}
@@ -202,7 +242,7 @@ export function renderScan(input: RenderInput, opts: RenderOptions): string {
function headerLines(input: RenderInput, opts: RenderOptions): string[] {
const elapsedMs = scanElapsedMs(input, opts.now);
const meta = [statusBadge(input, opts), elapsedMs !== undefined ? formatDuration(elapsedMs) : '—'].join(' · ');
return [` ${paint('Scan:', COLORS.bold, opts.color)} ${input.workspace.padEnd(22)} ${meta}`];
return [` ${paint('Scan:', COLORS.bold, opts.color)} ${safeCliIdentifier(input.workspace).padEnd(22)} ${meta}`];
}
/** Aligned label column for the footer's Logs / Temporal rows. */
@@ -223,15 +263,46 @@ function footerLines(input: RenderInput, opts: RenderOptions): string[] {
if (isTerminal(input.temporalStatus) && input.state?.summary) {
const wall = formatDuration(input.state.summary.totalDurationMs);
return ['', ` Time Taken ${wall}`];
const lines = ['', ` Time Taken ${wall}`];
// A partial scan names each durable degradation reason through its safe message,
// so the operator never has to guess why the badge is not "completed".
const reasons = safePartialReasons(input.state.partialReasons ?? []);
if (reasons.length > 0) {
lines.push('', ` ${paint('Why this scan is partial:', COLORS.yellow, opts.color)}`);
for (const reason of reasons) {
lines.push(paint(` - ${reason.message}`, COLORS.dim, opts.color));
}
// The safe message names what degraded; these three name the agentic-SAST failure
// behind it, under the same labels the scan log and worker output use.
const agenticSast = safeAgenticSast(input.state.agenticSast);
if (agenticSast?.status === 'failed') {
if (agenticSast.failedStageLabel !== undefined) {
lines.push(paint(` Agentic SAST stopped at: ${agenticSast.failedStageLabel}`, COLORS.dim, opts.color));
}
if (agenticSast.error !== undefined) {
lines.push(paint(` What happened: ${agenticSast.error}`, COLORS.dim, opts.color));
}
if (agenticSast.errorCode !== undefined) {
lines.push(paint(` Reference code (for a bug report): ${agenticSast.errorCode}`, COLORS.dim, opts.color));
}
}
}
if (input.state.summary.usageAccountingComplete === false) {
lines.push(
paint(' Cost is incomplete — some background work is not included in this total.', COLORS.dim, opts.color),
);
}
return lines;
}
const logsValue = `${prefix} logs ${input.workspace}`;
const logsValue = `${prefix} logs ${safeCliIdentifier(input.workspace)}`;
const temporalValue = temporalDashboardUrl(input.workflowId);
if (isTerminal(input.temporalStatus)) {
const rawReason = input.failureMessage ?? input.state?.error;
const reason = rawReason ? inlineFailureReason(rawReason) : 'no result recorded';
const hasRecordedFailure =
input.failureMessage !== undefined || (input.state !== null && input.state.error !== null);
const reason = safeTerminalFailure(hasRecordedFailure) ?? 'no result recorded';
return [
footerDivider(opts),
paint(
+278
View File
@@ -0,0 +1,278 @@
/**
* Closed-field projection for Temporal values displayed by the CLI.
*
* PipelineState travels through Temporal from a worker container this process does not
* control, so free-text fields are treated as unvetted: this module either matches a
* value against a known closed set (safe to print as-is) or collapses it to a fixed,
* bounded message. A value with no case here should fail closed to something generic,
* never pass through untouched.
*/
import type { PartialReasonView, PipelineState } from './pipeline.js';
const CLASS_NAMES: Readonly<Record<string, string>> = Object.freeze({
injection: 'Injection',
xss: 'Cross-Site Scripting',
auth: 'Authentication',
authz: 'Authorization',
ssrf: 'Server-Side Request Forgery',
miscellaneous: 'Miscellaneous',
});
const STAGE_NAMES: Readonly<Record<string, string>> = Object.freeze({
architecture: 'architecture mapping',
'threat-model': 'threat modelling',
plan: 'review planning',
research: 'deep code research',
dedupe: 'duplicate merging',
review: 'independent review',
critic: 'viability critique',
confirm: 'static confirmation',
calibrate: 'risk calibration',
export: 'findings export',
workflow: 'orchestration',
});
const TERMINAL_STAGE_NAMES = new Set([
'architecture',
'threat model',
'planning',
'audit wave',
'deduplication',
'review',
'critic',
'confirmation',
'calibration',
'export',
'orchestration',
]);
const CAPELLA_FAILURE_MESSAGES = new Set([
'Provider authentication failed. Verify the configured credential.',
'Agentic SAST configuration is invalid.',
'Agentic SAST received invalid input.',
'An agentic SAST step returned an unusable result.',
'An agentic SAST step failed.',
'Agentic SAST infrastructure failed before producing a usable result.',
'Agentic SAST had not finished when the scan stopped.',
]);
// Mirrors apps/worker/src/types/errors.ts. The CLI cannot import from the worker package,
// so keep this exact closed set in sync with ProviderFailureCategory.
const PROVIDER_FAILURE_CATEGORIES = new Set([
'rate_limit',
'overloaded',
'transport',
'context_limit',
'quota',
'authentication',
'configuration',
'unknown',
]);
function isProviderFailureCategory(value: unknown): value is string {
return typeof value === 'string' && PROVIDER_FAILURE_CATEGORIES.has(value);
}
const OPERATION_LABELS = new Set([
'Agentic SAST',
// Capella stage rows, signalled up from the SAST child workflow. Mirrors
// CAPELLA_STAGE_LABELS in apps/worker/src/ai/sast/types.ts, minus the deterministic
// export stage, which never becomes a row.
'Architecture',
'Threat model',
'Plan',
'Research',
'Dedupe',
'Review',
'Critique',
'Confirm',
'Calibrate',
'Reconcile injection',
'Reconcile xss',
'Reconcile auth',
'Reconcile authz',
'Reconcile ssrf',
'Reconcile miscellaneous',
'Prepare reconciliation',
'Enrich observations',
'Form exploit tasks',
'Materialize exploit tasks',
'Publish reconciliation',
'Renumber injection',
'Renumber xss',
'Renumber auth',
'Renumber authz',
'Renumber ssrf',
'Renumber miscellaneous',
'Initialize report state',
'Assemble report inputs',
'Compact report findings',
'Saving report progress',
'Finalize report outputs',
'Finalize report without SARIF',
'Saving final report state',
'Surface customer report',
]);
function safeClassName(value: string | undefined): string | undefined {
return value === undefined ? undefined : CLASS_NAMES[value];
}
function safeStageName(value: string | undefined): string | undefined {
return value === undefined ? undefined : STAGE_NAMES[value];
}
function reasonMessage(reason: PartialReasonView): string | undefined {
const className = safeClassName(reason.vulnerabilityClass);
switch (reason.code) {
case 'agentic_sast_failed': {
const stageName = safeStageName(reason.stage);
return stageName === undefined
? 'Agentic SAST failed, so the pentest continued without its findings.'
: `Agentic SAST failed during ${stageName}, so the pentest continued without its findings.`;
}
case 'agentic_sast_reduced':
return 'Agentic SAST completed with reduced coverage.';
case 'class_pipeline_failed':
return className === undefined
? undefined
: `${className} could not be fully assessed. The other classes completed. Re-running this workspace retries only the part that failed.`;
case 'class_reconciliation_failed':
return className === undefined
? undefined
: `${className} findings could not be grouped into test cases, so that class was not exploited and its findings are not in the report.`;
case 'report_renumber_failed':
return className === undefined
? undefined
: `${className} findings kept their working reference numbers, so numbering in the report may have gaps. The findings themselves are complete.`;
case 'report_compaction_failed':
return 'Finding reference numbers in the report may have gaps. Every finding is present; only the numbering is affected.';
case 'report_class_omitted':
return className === undefined
? undefined
: `${className} was assessed but could not be included in the final report.`;
case 'report_sarif_failed':
return 'Report SARIF could not be generated. JSON and Markdown remain available.';
default:
return undefined;
}
}
export function safePartialReasons(reasons: readonly PartialReasonView[]): readonly PartialReasonView[] {
return reasons.flatMap((reason) => {
const message = reasonMessage(reason);
if (message === undefined) return [];
const vulnerabilityClass =
safeClassName(reason.vulnerabilityClass) === undefined ? undefined : reason.vulnerabilityClass;
const stage = safeStageName(reason.stage) === undefined ? undefined : reason.stage;
return [
{
code: reason.code,
message,
...(vulnerabilityClass !== undefined && { vulnerabilityClass }),
...(stage !== undefined && { stage }),
},
];
});
}
/** Upper bounds on the warning array crossing into cli.status.json, so a malformed state cannot bloat it. */
const MAX_AGENTIC_SAST_WARNINGS = 20;
const MAX_AGENTIC_SAST_WARNING_LENGTH = 2_000;
/** Sanitize the worker's usage-accounting warnings: strings only, bounded count and length. */
function safeAgenticSastWarnings(value: PipelineState['agenticSast']): readonly string[] {
const warnings = value?.warnings;
if (!Array.isArray(warnings)) return [];
return warnings
.filter((warning): warning is string => typeof warning === 'string')
.slice(0, MAX_AGENTIC_SAST_WARNINGS)
.map((warning) => warning.slice(0, MAX_AGENTIC_SAST_WARNING_LENGTH));
}
export function safeAgenticSast(value: PipelineState['agenticSast']):
| {
readonly status: string;
readonly failedStageLabel?: string;
readonly error?: string;
readonly errorCode?: string;
readonly warnings: readonly string[];
}
| undefined {
if (value === undefined || !['disabled', 'running', 'succeeded', 'failed'].includes(value.status)) return undefined;
const failedStageLabel = TERMINAL_STAGE_NAMES.has(value.failedStageLabel ?? '') ? value.failedStageLabel : undefined;
let error: string | undefined;
if (value.error !== undefined && CAPELLA_FAILURE_MESSAGES.has(value.error)) {
error = value.error;
} else if (value.status === 'failed') {
error = 'An agentic SAST step failed.';
}
const errorCode =
value.errorCode !== undefined &&
(/^[A-Z][A-Z0-9_]{0,63}$/u.test(value.errorCode) || isProviderFailureCategory(value.errorCode))
? value.errorCode
: undefined;
return {
status: value.status,
...(failedStageLabel !== undefined && { failedStageLabel }),
...(error !== undefined && { error }),
...(errorCode !== undefined && { errorCode }),
warnings: safeAgenticSastWarnings(value),
};
}
export function safeOperationLabel(value: string): string {
return OPERATION_LABELS.has(value) ? value : 'Background task';
}
export function safeOperationKey(value: string): string {
if (
/^(?:agentic-sast|miscellaneous-pipeline|report:(?:initialize|assemble|compact|checkpoint|finalize|finalize-degraded|terminal|surface))$/u.test(
value,
) ||
/^agentic-sast:(?:architecture|threat-model|plan|research|dedupe|review|critic|confirm|calibrate)$/u.test(value) ||
/^(?:reconciliation|report:renumber):(?:injection|xss|auth|authz|ssrf|miscellaneous)$/u.test(value) ||
/^reconciliation:(?:injection|xss|auth|authz|ssrf|miscellaneous):fallback$/u.test(value)
) {
return value;
}
return 'background-task';
}
/**
* A workspace or workflow id is printed straight into the progress display, so this
* confines it to a plain identifier charset before that happens: no control or escape
* characters survive to reach the terminal.
*/
export function safeCliIdentifier(value: string): string {
return /^[A-Za-z0-9][A-Za-z0-9._-]{0,127}$/u.test(value) ? value : 'unknown';
}
export function safeTemporalStatus(value: string): string {
return [
'RUNNING',
'UNSPECIFIED',
'COMPLETED',
'FAILED',
'CANCELLED',
'CANCELED',
'TERMINATED',
'TIMED_OUT',
'CONTINUED_AS_NEW',
].includes(value)
? value
: 'UNKNOWN';
}
export function safeFailureDetail(hasFailure: true): string;
export function safeFailureDetail(hasFailure: false): undefined;
export function safeFailureDetail(hasFailure: boolean): string | undefined;
export function safeFailureDetail(hasFailure: boolean): string | undefined {
return hasFailure ? 'This scan step could not be completed.' : undefined;
}
/** Same closed-set trade-off as safeFailureDetail, for the scan-level (not per-agent) failure. */
export function safeTerminalFailure(hasFailure: boolean): string | undefined {
return hasFailure ? 'The scan could not be completed.' : undefined;
}
+41 -4
View File
@@ -8,7 +8,15 @@
import type { DerivedPhase } from './derive.js';
import { derivePipeline, isTerminal, scanElapsedMs } from './derive.js';
import type { PartialReasonView } from './pipeline.js';
import type { RenderInput } from './render.js';
import {
safeAgenticSast,
safeCliIdentifier,
safePartialReasons,
safeTemporalStatus,
safeTerminalFailure,
} from './safe-fields.js';
/** Coarse scan status token, mirroring the human status badge in machine-friendly form. */
export type ScanStatus = 'running' | 'completed' | 'partial' | 'failed' | 'stopped' | 'cancelled' | 'timed_out';
@@ -27,6 +35,18 @@ export interface StatusJson {
readonly endedAt?: string;
/** Failure text when a failed scan left no readable state. */
readonly failureMessage?: string;
/** Ordered durable degradation reasons with safe messages; present only when non-empty. */
readonly partialReasons?: readonly PartialReasonView[];
/** Agentic SAST outcome, with the worker's sanitized failure sentence and bounded code. */
readonly agenticSast?: {
readonly status: string;
readonly error?: string;
readonly errorCode?: string;
/** Usage-accounting warnings; always present (empty when the ledger reconciled) so it is never null. */
readonly warnings: readonly string[];
};
/** False when operational (Capella/reconciliation) spend is known to be incomplete. */
readonly usageAccountingComplete?: boolean;
readonly phases: readonly DerivedPhase[];
}
@@ -34,6 +54,7 @@ export interface StatusJson {
function deriveStatus(input: RenderInput): ScanStatus {
if (!isTerminal(input.temporalStatus)) return 'running';
if (input.state?.status === 'partial') return 'partial';
if (input.state?.status === 'cancelled') return 'cancelled';
switch (input.temporalStatus) {
case 'COMPLETED':
@@ -53,16 +74,32 @@ function deriveStatus(input: RenderInput): ScanStatus {
/** Build the JSON snapshot for a scan at instant `now`. */
export function toStatusJson(input: RenderInput, now: number): StatusJson {
const elapsedMs = scanElapsedMs(input, now);
const partialReasons = safePartialReasons(input.state?.partialReasons ?? []);
const agenticSast = safeAgenticSast(input.state?.agenticSast);
const usageAccountingComplete = input.state?.summary?.usageAccountingComplete;
const failureMessage = safeTerminalFailure(input.failureMessage !== undefined);
return {
workspace: input.workspace,
...(input.workflowId !== undefined && { workflowId: input.workflowId }),
workspace: safeCliIdentifier(input.workspace),
...(input.workflowId !== undefined && { workflowId: safeCliIdentifier(input.workflowId) }),
status: deriveStatus(input),
temporalStatus: input.temporalStatus,
temporalStatus: safeTemporalStatus(input.temporalStatus),
elapsedMs: elapsedMs ?? null,
...(input.startedAt !== undefined && { startedAt: new Date(input.startedAt).toISOString() }),
...(input.endedAt !== undefined && { endedAt: new Date(input.endedAt).toISOString() }),
...(input.failureMessage !== undefined && { failureMessage: input.failureMessage }),
...(failureMessage !== undefined && { failureMessage }),
...(partialReasons.length > 0 && { partialReasons }),
// Present only when agentic SAST actually ran; a disabled scan omits the key entirely.
...(agenticSast !== undefined &&
agenticSast.status !== 'disabled' && {
agenticSast: {
status: agenticSast.status,
...(agenticSast.error !== undefined && { error: agenticSast.error }),
...(agenticSast.errorCode !== undefined && { errorCode: agenticSast.errorCode }),
warnings: [...agenticSast.warnings],
},
}),
...(usageAccountingComplete !== undefined && { usageAccountingComplete }),
phases: derivePipeline(input, now),
};
}