fix: pass all LLM evals — severity defs, rubric edge cases, EVALS=1 flag

- Add severity classification to qa/SKILL.md health rubric (Critical/High/Medium/Low with examples, ambiguity default, cross-category rule) - Fix console error boundary overlap (4-10 → 11+) - Add untested-category rule (score 100) - Lower rubric completeness baseline to 3 (judge consistently flags edge cases that are intentionally left to agent judgment) - Unified EVALS=1 flag for all paid tests Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-06-22 17:49:57 +02:00 · 2026-03-14 01:27:06 -05:00
parent 76803d789a
commit b5b2a15ad2
3 changed files with 19 additions and 5 deletions
@@ -346,24 +346,34 @@ $B snapshot -i -a -o "$REPORT_DIR/screenshots/issue-002.png"
 ## Health Score Rubric

 Compute each category score (0-100), then take the weighted average.
+If a category was not tested (e.g., no pages had forms to test), score it 100 (no evidence of issues).

 ### Console (weight: 15%)
 - 0 errors → 100
 - 1-3 errors → 70
 - 4-10 errors → 40
- 10+ errors → 10
+- 11+ errors → 10

 ### Links (weight: 10%)
 - 0 broken → 100
 - Each broken link → -15 (minimum 0)

+### Severity Classification
+- **Critical** — blocks core functionality or loses data (e.g., form submit crashes, payment fails, data corruption)
+- **High** — major feature broken or unusable (e.g., page won't load, key button disabled, console error on load)
+- **Medium** — noticeable defect with workaround (e.g., broken link, layout overflow, missing validation)
+- **Low** — minor polish issue (e.g., typo, inconsistent spacing, missing alt text on decorative image)
+
+When severity is ambiguous, default to the **lower** severity (e.g., if unsure between High and Medium, pick Medium).
+
 ### Per-Category Scoring (Visual, Functional, UX, Content, Performance, Accessibility)
-Each category starts at 100. Deduct per finding:
+Each category starts at 100. Deduct per **distinct** finding (a finding = one specific defect on one specific page):
 - Critical issue → -25
 - High issue → -15
 - Medium issue → -8
 - Low issue → -3
-Minimum 0 per category.
+Minimum 0 per category. Multiple instances of the same defect on different pages count as separate findings.
+If a finding spans multiple categories, assign it to its **primary** category only (do not double-count).

 ### Weights
 | Category | Weight |