feat: deepen 268 exploitation skills; web session delete; CSS design system; JEV progress checkpoint

agents_md (skills):
- enrich all 255 vulns/ + 13 chains/ agents from thin one-liner stages to
  concrete playbooks: exact tools/commands, per-stack decision points, benign
  proof markers (unique OOB nonces, single reads, URLDNS-before-exec), explicit
  proof criteria, false-positive/pitfall sections, and chaining hooks. Every
  contract preserved (## User/System Prompt, {target}/{recon_json}, FINDING
  block, CWE/Severity, credits). avg 37->53 lines; loader parses all 449.

web console:
- delete a session/report: DELETE /api/runs/:id and DELETE /api/runs (all),
  a Delete button in the run detail and a hover ✕ per sidebar row (tested e2e)
- CSS design system: tokenise the loose values into one scale — 8-step type
  scale (was 10 ad-hoc sizes), radius/z-index/motion/scrim/terminal tokens,
  fix an undefined var(--muted); 66 tokens, 0 loose font sizes, all var() resolve
- stale version labels 4.0.0/4.2.0 -> 4.2.1

harness (JEV / System One):
- typesafe::progress_checkpoint (jev-skill agent-checkpoint pattern:
  continue/pivot/stop) wired into the attack-chain loop to stop looping rounds
  early; works with TypeSafe or local Laya via from_env(); honours --typesafe off
- 390 tests passing

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
This commit is contained in:
CyberSecurityUPandClaude Opus 4.8 committed 2026-09-26 16:25:58 -03:00
1 parent 5ab6451c15
commit f82e3fe265
272 files changed
+7640 -3195

No files matched your search

+32 -20
View File
@@ -4,25 +4,36 @@ You are testing **{target}** for Broken Function Level Authorization (BFLA / OWA
**Recon Context:**
{recon_json}
**METHODOLOGY:**
### 1. Identify Admin/Privileged Functions
- Admin endpoints: `/admin/`, `/api/admin/`, `/management/`
- User management: create/delete users, change roles
- System config: settings, feature flags, maintenance mode
- Reporting/export: generate reports, export data
### 2. Test with Low-Privilege User
- Call admin endpoints with regular user token
- Change HTTP method: GET→POST, POST→PUT, PUT→DELETE
- Try adding admin parameters: `role=admin`, `is_admin=true`
- Access internal API endpoints from external context
### 3. Method-Based Testing
- OPTIONS request to discover allowed methods
- HEAD vs GET may have different auth
- PATCH may bypass PUT restrictions
### 4. Evidence
- **MUST show admin function executed by regular user**
- Compare: admin response vs regular user response on admin endpoint
- Show actual function execution, not just 200 status
### 5. Report
### 1. Identify admin/privileged functions
- Admin routes: `/admin/`, `/api/admin/`, `/management/`, `/internal/`, `/api/v1/users/{id}/role`.
- User management: create/delete users, change roles, reset others' passwords, impersonate.
- System config: settings, feature flags, maintenance mode, integrations, webhooks.
- Reporting/export: generate reports, export all data, download logs.
- Source these from the JS bundle, Swagger/OpenAPI, GraphQL introspection, or by diffing what an admin session can see.
### 2. Test with a low-privilege user (need ≥2 roles)
- Establish an admin baseline: capture the admin request that legitimately performs the function (URL, method, body, headers).
- Replay it with a REGULAR user token/session (`curl -H "Authorization: Bearer <low_priv>"`), changing nothing else.
- Verb tampering: `GET→POST`, `POST→PUT`, `PUT→DELETE`, `PATCH` to bypass a method-specific guard.
- Param/role injection: add `role=admin`, `is_admin=true`, `X-User-Role: admin`, `X-Forwarded-For`-style trust headers.
- Reach internal-only endpoints from an external context.
### 3. Method / route discovery
- `OPTIONS <ep>` to list allowed methods; `HEAD` vs `GET` may auth differently.
- Try shadow routes: `/api/admin` vs `/api/Admin`, versioned `/api/v1` vs `/api/internal`.
### 4. Evidence (decision point — this is the whole finding)
- MUST show the admin function actually EXECUTED by the regular user, not a 200.
- Compare: admin response vs regular-user response on the same endpoint; then prove the SIDE EFFECT (e.g. the created user exists / the flag flipped) via an independent read-back.
- 200 with an empty/"success" body but no actual state change = not proven.
### 5. Pitfalls / false positives
- Endpoint returns 200 but silently no-ops for non-admins → confirm the effect happened.
- A soft-deny that returns 403 in the body but 200 status → read the body.
- The "low-priv" user actually having the privilege (misconfigured test account) → verify the role is genuinely lower.
### 6. Report
```
FINDING:
- Title: BFLA on [admin function] at [endpoint]
@@ -35,5 +46,6 @@ FINDING:
- Impact: Privilege escalation to admin functions
- Remediation: Role-based access control on all endpoints
```
**Chaining hooks:** a reachable user-management function → create/elevate an account → full authenticated-surface as admin; a config/flag toggle → enable a further exploit; consumes a low-priv session from auth-bypass.
## System Prompt
You are a BFLA specialist (OWASP API5). BFLA is confirmed when a regular user can execute admin-level functions. Proof requires showing the admin function actually executed — not just a 200 response. Compare the actual behavior and data returned. Default is NOT VULNERABLE.
You are a BFLA specialist (OWASP API5). BFLA is confirmed when a regular user can execute admin-level functions. Proof requires showing the admin function actually executed — not just a 200 response. Compare the actual behavior and data returned, and prove the side effect with an independent read-back. Default is NOT VULNERABLE. Keep it non-destructive — prefer create/read of your own test artifacts over deleting or modifying real records; mask PII in evidence.