mirror of
https://github.com/garrytan/gstack.git
synced 2026-09-11 07:29:00 +02:00
docs: sync docs for v1.66.0.0 (test/evals/CI speedup)
CONTRIBUTING.md, AGENTS.md, and ARCHITECTURE.md still taught bare `bun test` for the suite; the shipped runner deprecates it (walks the whole repo, loads paid eval files, misses the strict classifier). All suite-level references now say `bun run test`, the Tier 1 section describes the strict shard runner (~90-100s, --verbose, --wall-timeout), the sharded paid-runner paragraph documents diff-based shard skipping and the EVALS_JOBS / EVALS_CONCURRENCY split, the Tier 3 row points at the actual judge-only invocation, and GSTACK_EVAL_MODEL_JUDGE is documented at the judge it overrides. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
This commit is contained in:
co-authored by
Claude Fable 5
parent
f627c676fc
commit
9b9bc84bdc
+2
-2
@@ -321,7 +321,7 @@ Three reasons:
|
||||
| 2 — E2E via `claude -p` | Spawn real Claude session, run each skill, check for errors | ~$3.85 | ~20min |
|
||||
| 3 — LLM-as-judge | Sonnet scores docs on clarity/completeness/actionability | ~$0.15 | ~30s |
|
||||
|
||||
Tier 1 runs on every `bun test`. Tiers 2+3 are gated behind `EVALS=1`. The idea is: catch 95% of issues for free, use LLMs only for judgment calls.
|
||||
Tier 1 runs on every `bun run test`. Tiers 2+3 are gated behind `EVALS=1`. The idea is: catch 95% of issues for free, use LLMs only for judgment calls.
|
||||
|
||||
## Command dispatch
|
||||
|
||||
@@ -435,7 +435,7 @@ The `EvalCollector` accumulates test results and writes them in two ways:
|
||||
| 2 — E2E via `claude -p` | Spawn real Claude session, run each skill, scan for errors | ~$3.85 | ~20min |
|
||||
| 3 — LLM-as-judge | Sonnet scores docs on clarity/completeness/actionability | ~$0.15 | ~30s |
|
||||
|
||||
Tier 1 runs on every `bun test`. Tiers 2+3 are gated behind `EVALS=1`. The idea: catch 95% of issues for free, use LLMs only for judgment calls and integration testing.
|
||||
Tier 1 runs on every `bun run test`. Tiers 2+3 are gated behind `EVALS=1`. The idea: catch 95% of issues for free, use LLMs only for judgment calls and integration testing.
|
||||
|
||||
## What's intentionally not here
|
||||
|
||||
|
||||
Reference in New Issue
Block a user