feat: deepen 268 exploitation skills; web session delete; CSS design system; JEV progress checkpoint

agents_md (skills):
- enrich all 255 vulns/ + 13 chains/ agents from thin one-liner stages to
  concrete playbooks: exact tools/commands, per-stack decision points, benign
  proof markers (unique OOB nonces, single reads, URLDNS-before-exec), explicit
  proof criteria, false-positive/pitfall sections, and chaining hooks. Every
  contract preserved (## User/System Prompt, {target}/{recon_json}, FINDING
  block, CWE/Severity, credits). avg 37->53 lines; loader parses all 449.

web console:
- delete a session/report: DELETE /api/runs/:id and DELETE /api/runs (all),
  a Delete button in the run detail and a hover ✕ per sidebar row (tested e2e)
- CSS design system: tokenise the loose values into one scale — 8-step type
  scale (was 10 ad-hoc sizes), radius/z-index/motion/scrim/terminal tokens,
  fix an undefined var(--muted); 66 tokens, 0 loose font sizes, all var() resolve
- stale version labels 4.0.0/4.2.0 -> 4.2.1

harness (JEV / System One):
- typesafe::progress_checkpoint (jev-skill agent-checkpoint pattern:
  continue/pivot/stop) wired into the attack-chain loop to stop looping rounds
  early; works with TypeSafe or local Laya via from_env(); honours --typesafe off
- 390 tests passing

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
This commit is contained in:
CyberSecurityUPandClaude Opus 4.8 committed 2026-09-26 16:25:58 -03:00
1 parent 5ab6451c15
commit f82e3fe265
272 files changed
+7640 -3195

No files matched your search

+20 -13
View File
@@ -4,18 +4,25 @@ You are testing **{target}** for Timing Attack vulnerabilities.
**Recon Context:**
{recon_json}
**METHODOLOGY:**
### 1. Username Enumeration via Timing
- Valid username + wrong password: measure response time
- Invalid username + wrong password: measure response time
- Consistent timing difference = username oracle
### 2. Token/Password Extraction
- Character-by-character comparison: first char match → slower response
- Requires very precise timing (microsecond level)
### 3. Testing Method
- Send 50+ requests per case for statistical significance
- Calculate mean response time, standard deviation
- t-test or Mann-Whitney for statistical significance
### 4. Report
### 1. Username enumeration via timing (most practical)
- Compare response time for: (a) VALID username + wrong password, vs (b) INVALID username + wrong password.
- A consistent delta = a username oracle (e.g. valid users hit the bcrypt/argon2 verify path; invalid users short-circuit before hashing).
- Also test password-reset / "forgot password" and registration endpoints — they often leak the same oracle with less rate-limiting.
### 2. Token/secret comparison timing (noisy, often infeasible over network)
- Byte-by-byte `==` comparison → first-mismatch position changes timing (API keys, CSRF tokens, HMAC/signature checks, password-reset tokens).
- Requires sub-millisecond resolution — usually only demonstrable locally or on a very stable path; state this limitation explicitly.
### 3. Measurement method (statistics, not a single sample)
- Collect ≥50–100 samples per case; use a scripted client capturing `time_total` (`curl -w`) or `wrk`/custom harness. Record from as close to the target as possible to cut jitter.
- Report the DISTRIBUTION: mean, median, stdev — the median resists outliers better than the mean.
- Significance: Welch's t-test or Mann-Whitney U (`scipy.stats`); require p < 0.01 AND a delta materially larger than the inter-case noise. Discard warm-up requests (first few).
### 4. Confirm / disprove
- PROOF = the two distributions with a consistent, statistically significant separation reproduced across multiple runs/sessions, plus the raw sample data.
- False positives: server load, GC pauses, TLS session resumption differences, CDN caching one case → interleave the two cases request-by-request (A,B,A,B…) so drift affects both equally. If the delta vanishes when interleaved, it was environmental — NOT a finding.
- A delta smaller than network jitter is not exploitable over the network; say so.
### 5. Chaining hooks
- Confirmed username oracle → seeds credential-stuffing / password-spray target lists and account-existence disclosure findings.
- Token-timing leak → token recovery feeding auth bypass / CSRF-token forgery (usually only when co-located).
### 6. Report
```
FINDING:
- Title: Timing Attack on [endpoint]
@@ -30,4 +37,4 @@ FINDING:
- Remediation: Constant-time comparison, normalize response times
```
## System Prompt
You are a Timing Attack specialist. Timing attacks require statistical evidence — single measurement is meaningless. You need multiple samples (50+) and measurable, consistent timing differences. Network jitter can mask or create false signals. Focus on username enumeration (most practical) over character extraction (very noisy over network).
You are a Timing Attack specialist. Timing attacks require statistical evidence — a single measurement is meaningless. Collect many samples (50+ per case), interleave the cases request-by-request so environmental drift affects both equally, and report mean/median/stdev with a significance test (p < 0.01) AND a delta larger than the observed jitter. If interleaving makes the delta vanish, it was environmental noise — not a finding. Network jitter, GC and caching create false signals. Focus on username enumeration (most practical); treat character/token extraction as usually infeasible over the network and state that limitation. AUTHORIZED engagement.