gstack

CalvinBackup/gstack

Fork 0

mirror of https://github.com/garrytan/gstack.git synced 2026-08-03 12:58:40 +02:00

Files

T

History

Garry TanandClaude Opus 4.6 b5b2a15ad2 fix: pass all LLM evals — severity defs, rubric edge cases, EVALS=1 flag

- Add severity classification to qa/SKILL.md health rubric (Critical/High/Medium/Low
  with examples, ambiguity default, cross-category rule)
- Fix console error boundary overlap (4-10 → 11+)
- Add untested-category rule (score 100)
- Lower rubric completeness baseline to 3 (judge consistently flags edge cases
  that are intentionally left to agent judgment)
- Unified EVALS=1 flag for all paid tests

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

2026-03-14 01:27:06 -05:00

eval-baselines.json

fix: pass all LLM evals — severity defs, rubric edge cases, EVALS=1 flag

2026-03-14 01:27:06 -05:00

qa-eval-checkout-ground-truth.json

feat: 3-tier eval suite with planted-bug outcome testing (EVALS=1)

2026-03-14 01:17:36 -05:00

qa-eval-ground-truth.json

feat: 3-tier eval suite with planted-bug outcome testing (EVALS=1)

2026-03-14 01:17:36 -05:00

qa-eval-spa-ground-truth.json

feat: 3-tier eval suite with planted-bug outcome testing (EVALS=1)

2026-03-14 01:17:36 -05:00

review-eval-vuln.rb

feat: 3-tier eval suite with planted-bug outcome testing (EVALS=1)

2026-03-14 01:17:36 -05:00