gstack

mirror of https://github.com/garrytan/gstack.git synced 2026-05-02 11:45:20 +02:00

Files

T

History

Garry Tan 59752fc510 feat: wire eval-cache + eval-tier into LLM judge, pin E2E model

callJudge/judge now return {result, meta} with SHA-based caching
(~$0.18/run savings when SKILL.md unchanged) and dynamic model
selection via EVAL_JUDGE_TIER env var. E2E tests pass --model from
EVAL_TIER to claude -p. outcomeJudge retains simple return type.
All 8 LLM eval test sites updated with real costs and costs[].

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

2026-03-15 16:47:35 -05:00

eval-store.test.ts

feat: wire costs[] from modelUsage into eval results

2026-03-15 16:47:27 -05:00

eval-store.ts

feat: wire costs[] from modelUsage into eval results

2026-03-15 16:47:27 -05:00

llm-judge.test.ts

feat: wire eval-cache + eval-tier into LLM judge, pin E2E model

2026-03-15 16:47:35 -05:00

llm-judge.ts

feat: wire eval-cache + eval-tier into LLM judge, pin E2E model