mirror of
https://github.com/garrytan/gstack.git
synced 2026-10-08 04:11:18 +02:00
test: clean up the paid eval lane (B1-B4, B6, B7)
- B1: delete paid files that assert nothing or cannot pass meaningfully: skill-llm-eval-spec and skill-e2e-spec-execute (test.todo), gemini-e2e (+ gemini-session-runner; no gemini CLI in CI), ship-idempotency (red since v1.63), the two opus-4-7 *-sonnet overlay wrappers, conductor-prose (+ its source-evaluation replay), codex-e2e-plan-format; drop their keys, scripts and census rows. - B2: skill-llm-eval grades browse/sections/command-list.md with one union judge that also carries the baseline score pin; regression-vs-baseline deleted (paid run: pass, c4/c4/a4). - B3: memory-pipeline, ios-qa, ios-qa-swift-build and plan-tune-cathedral make no model calls; renamed out of the paid glob so they run on every PR. Swift builds need GSTACK_TEST_SWIFT=1; device stub deleted. - B4: codex-e2e*, outside-voice, aside and ios-device cannot run in the CI image; excluded from the weekly lane with a tracked re-entry condition. - B6: fold opus-47's negative routing controls into skill-routing-e2e journey-negatives (paid run: 3/3 unrouted) and delete the file. - B7: delete the never-green brain-privacy-gate eval; a free gstack-skill-start test now proves consent precedes artifacts egress.
This commit is contained in:
1 parent
5d032ef299
commit
53e7f3212f
58 files changed
+273
-2594
No files matched your search
@@ -60,7 +60,7 @@ describe('paid test enumeration', () => {
|
||||
const files = collectPaidTestFiles();
|
||||
expect(files.length).toBeGreaterThan(0);
|
||||
expect(files.every(isPaidTestFile)).toBe(true);
|
||||
expect(PAID_TEST_GLOBS.length).toBe(7);
|
||||
expect(PAID_TEST_GLOBS.length).toBe(6);
|
||||
|
||||
const shards = planPaidShards(files);
|
||||
expect(shards.flat().sort()).toEqual([...files].sort());
|
||||
|
||||
Reference in new issue
Block a user