shannon

mirror of https://github.com/KeygraphHQ/shannon.git synced 2026-02-12 17:22:50 +00:00

Author	SHA1	Message	Date
Arjun Malleswaran	accb9562ba	Merge pull request #19 from KeygraphHQ/additional-flags chore: added flag additions for minimizing logs	2025-12-09 10:33:36 -08:00
Khaushik-keygraph	38e49eb1eb	chore: added flag additions for minimizing logs	2025-12-09 23:59:12 +05:30
Arjun Malleswaran	c664000458	Merge pull request #18 from KeygraphHQ/16-windows-defender-flags-benchmark-deliverables-as-backdoorphpperhetshell-during-local-use docs: add Windows Defender false positive guidance	2025-12-08 10:20:51 -08:00
ajmallesh	af41570ae9	docs: add Windows Defender false positive guidance Closes #16	2025-12-02 19:07:37 -08:00
ajmallesh	2c410d90b3	docs: update Discord invite links	2025-12-01 09:24:19 -08:00
ajmallesh	534b18e303	chore: change license to AGPL-3.0	2025-11-26 18:45:36 -08:00
ajmallesh	8f2825b32f	docs: clarify Shannon is a white-box pentesting tool - Add prominent callout that Shannon Lite is designed for white-box (source-available) application security testing - Update XBOW benchmark description to "hint-free, source-aware" - Clarify benchmark comparison context (white-box vs black-box results) - Update benchmark performance comparison image 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-Authored-By: Claude <noreply@anthropic.com>	2025-11-24 12:37:55 -08:00
Khaushik-keygraph	369e3a34cf	chore: added licensing to dockerfile	2025-11-22 20:46:15 +05:30
keygraphVarun	7f7285702e	fix link	2025-11-22 20:43:09 +05:30
keygraphVarun	deb4e51f98	cleanup	2025-11-22 20:43:09 +05:30
keygraphVarun	2b14282ff6	consistency on score	2025-11-22 20:43:09 +05:30
ajmallesh	5bbd757b45	fix: resolve Docker build failure and clarify env var configuration - Remove .env file with incorrect CLAUDE_CODE_MAX_TOKENS variable - Remove .env copy from Dockerfile that was causing build to fail - Update README to distinguish local (export) vs Docker (-e) env var usage - Add CLAUDE_CODE_MAX_OUTPUT_TOKENS to all Docker run examples The correct variable is CLAUDE_CODE_MAX_OUTPUT_TOKENS (not CLAUDE_CODE_MAX_TOKENS) and should be passed at runtime via -e flag for Docker or export for local runs. 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-Authored-By: Claude <noreply@anthropic.com>	2025-11-19 10:28:44 -08:00
Khaushik-keygraph	d2519322d2	fix: removed comments	2025-11-13 20:33:58 +05:30
keygraphVarun	456d852b87	style changes	2025-11-13 20:28:15 +05:30
keygraphVarun	341448c8a3	Link to benchmark	2025-11-13 20:27:26 +05:30
ajmallesh	b32e71a9b4	chore: add licensing comments to prompts	2025-11-13 17:53:41 +05:30
ajmallesh	30f324be5e	Update license references from BSL to MPL in documentation 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-Authored-By: Claude <noreply@anthropic.com>	2025-11-13 17:48:05 +05:30
Arjun Malleswaran	378585a4a3	Merge pull request #14 from KeygraphHQ/license-change License change	2025-11-13 16:57:18 +05:30
Arjun Malleswaran	c040efc6b5	Update LICENSE	2025-11-13 16:56:19 +05:30
ajmallesh	1051d40527	chore: add MPL license comments	2025-11-13 16:55:13 +05:30
Arjun Malleswaran	fbf24c5d10	Update README.md	2025-11-04 08:47:18 -08:00
Arjun Malleswaran	13231ae016	Update README.md	2025-11-04 08:46:15 -08:00
ajmallesh	76591ae5e2	Update README.md	2025-11-03 20:23:16 -08:00
ajmallesh	23e072a236	Merge branch 'main' of github.com:KeygraphHQ/shannon	2025-11-03 20:22:27 -08:00
ajmallesh	9d9b81d8a9	Update README.md	2025-11-03 20:22:18 -08:00
Arjun Malleswaran	3ff1609151	Merge pull request #9 from KeygraphHQ/adding-xben-results Update README.md	2025-11-03 20:19:55 -08:00
ajmallesh	aa045b65da	Update README.md	2025-11-03 20:16:08 -08:00
Arjun Malleswaran	b25cbaa643	Merge pull request #7 from KeygraphHQ/adding-xben-results Adding xben results	2025-11-03 20:04:45 -08:00
ajmallesh	40f80fc4cb	Update README.md	2025-11-03 20:04:21 -08:00
ajmallesh	c8d7ec1e29	docs: add benchmarks README	2025-11-03 20:03:06 -08:00
ajmallesh	cb54ad46a0	Rename SQLi/Command Injection to Injection throughout README Consolidates SQL Injection and Command Injection references to the unified "Injection" terminology for consistency with agent naming and OWASP categorization. Changes: - Updated feature descriptions and vulnerability lists - Modified architecture diagrams - Simplified targeted vulnerability scope 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-Authored-By: Claude <noreply@anthropic.com>	2025-11-03 16:56:40 -08:00
ajmallesh	7454b1a581	Add audit logs and update gitignore for xben results Updates .gitignore to only ignore top-level audit-logs/ directory, allowing xben-benchmark-results audit logs to be tracked. This enables full reproducibility of benchmark runs with complete session data, prompts, and agent execution logs. 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-Authored-By: Claude <noreply@anthropic.com>	2025-11-03 16:29:56 -08:00
ajmallesh	aa49f9fc66	Add X-Bow benchmark performance visualization This commit adds a professional performance comparison chart showing Shannon's 96% success rate against other autonomous pentesting systems on the X-Bow benchmark. Chart features: - Y-axis properly starts at 0% (honest data visualization) - Shannon bar highlighted in brand orange - Descriptive title with sample size (104 challenges) - SVG format for scalability 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-Authored-By: Claude <noreply@anthropic.com>	2025-11-03 12:34:55 -08:00
ajmallesh	c686d7f800	Add X-Bow benchmark results (104 test cases) This commit adds comprehensive X-Bow (XBEN) benchmark results demonstrating Shannon's performance across 104 CTF security challenges. Each test case includes detailed penetration testing reports and exploitation evidence for reproducible research. Contents: - 104 XBEN test case directories (XBEN-001-24 through XBEN-104-24) - Deliverables including analysis reports and exploitation evidence - Individual test case results with vulnerability assessments 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-Authored-By: Claude <noreply@anthropic.com>	2025-11-03 12:34:41 -08:00
ajmallesh	33f17dd570	docs: add ctf-mode branch documentation to README Add a TIP callout in the Overview section documenting the ctf-mode branch for users who want to run Shannon against Capture-The-Flag challenges with optimized flag extraction prompts. 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-Authored-By: Claude <noreply@anthropic.com>	2025-11-03 10:35:45 -08:00
ajmallesh	939398074f	refactor: update injection display name and add max tokens docs - Change agent prefix from [SQLi/Cmd] to [Injection] to reflect expanded scope - Add README documentation for CLAUDE_CODE_MAX_OUTPUT_TOKENS environment variable This update aligns the display naming with the expanded injection analysis scope that now covers SQLi, Command Injection, LFI/RFI, SSTI, Path Traversal, and Insecure Deserialization vulnerabilities. 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-Authored-By: Claude <noreply@anthropic.com>	2025-11-03 10:21:17 -08:00
ajmallesh	4224d1c4f4	feat: expand injection analysis scope to cover LFI/RFI/SSTI/Path Traversal/Deserialization Fixes responsibility gap where agents found vulnerabilities but rejected them as "out of scope" Changes: - vuln-injection.txt: Added LFI/RFI, SSTI, Path Traversal, Deserialization to scope - Updated role definition and objective - Added new vulnerability_type and slot_type enums - Added sink definitions and defense rules for new injection classes - Added witness payload examples - pre-recon-code.txt: Expanded sink hunter agent to find file/template/deserialize sinks - recon.txt: Updated Section 9 with clear injection source definitions for all types - exploit-injection.txt: Updated evidence template to handle all injection types Token-optimized: Condensed verbose sections while preserving critical guidance Addresses XBEN benchmark failures where LFI/SSTI/Path Traversal were detected but excluded from exploitation queues 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-Authored-By: Claude <noreply@anthropic.com>	2025-11-03 10:20:15 -08:00
ajmallesh	52d7cc46a6	feat: add environment variable support for Claude Code token limits Introduces .env file configuration to manage CLAUDE_CODE_MAX_TOKENS, allowing flexible control of the context window size for AI analysis sessions. This enables users to tune token limits based on their specific penetration testing needs without modifying code. 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-Authored-By: Claude <noreply@anthropic.com>	2025-10-30 10:53:42 -07:00
ajmallesh	9d338ec948	fix: err handling for claude code session limit	2025-10-30 10:28:35 -07:00
ajmallesh	905156bbc1	chore: print audit logs folder location	2025-10-28 10:31:00 -07:00
ajmallesh	4be7f969a9	Merge pull request #3 from KeygraphHQ/feature/improve-audit-log-naming Feature/improve audit log naming	2025-10-27 14:56:57 -07:00
ajmallesh	7f3bff9b36	Revert "feat: improve audit log naming with timestamp and app context" This reverts the timestamp-based naming scheme that was causing audit log fragmentation. Each agent execution was creating a new folder because the timestamp kept changing. Reverting back to simple, stable naming: {hostname}_{sessionId} This ensures ONE folder per session, preventing the bug where multiple folders were created for the same session. 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-Authored-By: Claude <noreply@anthropic.com>	2025-10-27 13:30:25 -07:00
ajmallesh	95b8e876eb	fix: use session's original createdAt instead of current time Fixed bug where audit system would create duplicate folders for the same session because it was using current time instead of the session's original createdAt timestamp. Bug behavior: - Session created at T1 → folder: {T1}_app_host_id/ - Audit re-initialized at T2 → NEW folder: {T2}_app_host_id/ - Result: 2 folders per session with same ID but different timestamps Root cause: - metrics-tracker.js:65 was calling formatTimestamp() (current time) - Should use sessionMetadata.createdAt (original creation time) Impact: Each running benchmark was creating 2 audit log folders instead of 1 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-Authored-By: Claude <noreply@anthropic.com>	2025-10-27 10:55:53 -07:00
ajmallesh	6e89f26474	feat: improve audit log naming with timestamp and app context Enhances audit log directory naming from `{hostname}_{uuid}` to `{timestamp}_{appName}_{hostname}_{shortId}` for better discoverability and benchmarking analysis. Changes: - Add extractAppName() helper to extract app name from config files - Add smart fallback: use port number for localhost without config - Update generateSessionIdentifier() to include timestamp prefix - Shorten session ID to first 8 characters for readability Examples: - With config: 20251025T193847Z_myapp_localhost_efc60ee0/ - Without config: 20251025T193913Z_8080_localhost_d47e3bfd/ - Remote: 20251024T004401Z_noconfig_example-com_d47e3bfd/ Benefits: - Chronologically sortable audit logs - Instant app identification in directory listings - Efficient filtering for benchmarking queries - Non-breaking: existing logs keep their names 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-Authored-By: Claude <noreply@anthropic.com>	2025-10-27 10:14:19 -07:00
ajmallesh	b4cd1066d6	Merge pull request #2 from KeygraphHQ/fixing-bugs Fixing bugs	2025-10-23 18:18:21 -07:00
ajmallesh	dcae34af81	fix: enable Playwright MCP browser automation in Docker containers Resolves Playwright browser installation failures in Docker by using Wolfi's system Chromium instead of downloading Playwright's bundled browsers at runtime. ## Problem When running in Docker, agents attempted to install browsers via `browser_install` tool, which failed due to: - Permission issues (non-root user couldn't install system dependencies) - npx @playwright/mcp spawns with its own Playwright dependency separate from global installations - Playwright's bundled browsers require runtime download (~280MB) and glibc deps - Environment variables alone (PLAYWRIGHT_BROWSERS_PATH) weren't sufficient ## Solution Dockerfile changes: - Use Wolfi's native `chromium` package (guaranteed compatible, already installed) - Remove Playwright browser installation step (saves ~280MB and build time) - Add explicit `SHANNON_DOCKER=true` environment variable for reliable detection - Set PLAYWRIGHT_CHROMIUM_EXECUTABLE_PATH to point to system Chromium Code changes (claude-executor.js): - Detect Docker via `process.env.SHANNON_DOCKER` (more reliable than /.dockerenv) - Conditionally add `--executable-path /usr/bin/chromium-browser` CLI arg for Docker - Local: Use Playwright's bundled browsers (downloaded to ~/Library/Caches/) - Docker: Use system Chromium with no runtime downloads ## Research Findings - @playwright/mcp has separate playwright-core dependency (v1.56.0-alpha) - MCP server spawned via npx doesn't inherit browser binaries from global install - --executable-path CLI argument is required (env vars insufficient) - /.dockerenv file is unreliable (missing in BuildKit, K8s, can be spoofed) ## Testing ✅ Docker: All 5 parallel agents successfully navigate, screenshot, create deliverables ✅ Local: All 5 parallel agents successfully navigate, screenshot, create deliverables ✅ No browser_install calls, no permission errors ✅ Image size reduced by ~280MB Fixes #docker-playwright-browser-issues 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-Authored-By: Claude <noreply@anthropic.com>	2025-10-23 17:56:19 -07:00
ajmallesh	3094862310	refactor: simplify pipeline testing report prompt by 78% Reduce prompts/pipeline-testing/report-executive.txt from 137 to 30 lines by: - Removing hardcoded detailed vulnerability content - Testing actual workflow (read → modify → save) instead of creating from scratch - Removing meta-commentary, keeping only direct instructions - Making it consistent with other pipeline testing prompts (30 lines like exploit agents) The prompt now properly mimics the real reporting agent behavior where the orchestration code stitches files first, then the agent modifies the result. 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-Authored-By: Claude <noreply@anthropic.com>	2025-10-23 17:13:25 -07:00
ajmallesh	d372f87297	refactor: remove ~500 lines of dead code and consolidate duplicates Comprehensive codebase cleanup based on parallel agent analysis and automated dead code detection (knip, depcheck). Reduces codebase by ~10% with zero functional changes. ## Phase 1: Obsolete MCP Setup Removal (~82 lines) - Delete setupMCP() and cleanupMCP() functions from environment.js - Remove all calls to cleanupMCP() (8 instances across 3 files) - Migrate from claude CLI to SDK's mcpServers option - Remove --log flag (obsolete logging system) ## Phase 2: Dead Code Removal (~317 lines) - Delete src/utils/logger.js entirely (127 lines, superseded by audit system) - Remove handleConfigError() and handleError() from error-handling.js - Remove isToolAvailable() from tool-checker.js - Remove 5 dead methods from audit-session.js (logSessionFailure, logMessage, markRolledBack, updateValidation, getValidation) - Remove 6 wrapper methods from audit/logger.js (all callers use logEvent directly) - Remove formatCost(), updateMessage(), compose() utilities (unused) ## Phase 3: Consolidation (~195 lines) - Extract SessionMutex to src/utils/concurrency.js (was duplicated in 2 files) - Consolidate formatDuration to src/audit/utils.js (was in 3 files) - Extract readline prompts to src/cli/prompts.js (was duplicated in 2 files) - Create validator factories in constants.js (reduce 72 lines to 30) ## Impact - Total reduction: 488 lines (20 files modified, 2 created, 1 deleted) - Codebase: ~4,900 → ~4,400 LOC (10% reduction) - Zero functional changes, all tests pass - Improved maintainability and DRY compliance 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-Authored-By: Claude <noreply@anthropic.com>	2025-10-23 17:01:17 -07:00
ajmallesh	369bf29588	refactor: deduplicate prompt templates with shared content system Implemented @include() directive system to eliminate ~800 lines of duplicated content across 10 specialist prompt files. All prompt-related content now consolidated under prompts/ directory for better maintainability. Changes: - Added processIncludes() to prompt-manager.js for generic @include() support - Created prompts/shared/ with 5 reusable template files - Refactored all 10 specialist prompts to use @include() for common sections - Moved login_instructions.txt to prompts/shared/ (deleted login_resources/) - Updated CLAUDE.md to reflect new structure Impact: -137 net lines, zero breaking changes, infinitely scalable for future shared content. 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-Authored-By: Claude <noreply@anthropic.com>	2025-10-23 16:19:25 -07:00
ajmallesh	dafd9148f6	chore: remove ~500 lines of dead code identified by knip Remove unused files and exports to improve codebase maintainability: Phase 1 - Deleted files (5): - login_resources/generate-totp-standalone.mjs (replaced by MCP tool) - mcp-server/src/tools/index.js (unused barrel export) - mcp-server/src/utils/index.js (unused barrel export) - mcp-server/src/validation/index.js (unused barrel export) - src/agent-status.js (deprecated 309-line status manager) Phase 2 - Removed unused exports (3): - mcp-server/src/index.js: shannonHelperServer constant - mcp-server/src/utils/error-formatter.js: createFileSystemError function - src/utils/git-manager.js: cleanWorkspace (now internal-only) Phase 3 - Unexported internal functions (4): - src/checkpoint-manager.js: runSingleAgent, runAgentRange, runParallelVuln, runParallelExploit (internal use only) All Shannon CLI commands tested and verified working. 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-Authored-By: Claude <noreply@anthropic.com>	2025-10-23 12:46:51 -07:00

1 2

79 Commits