feat(evals): with-skill vs without-skill arm benchmark — measures whether gstack's behavioral layer earns its tokens

Ponytail's honest-benchmark method pointed at gstack itself: 3 build-shaped
tasks (native-platform over-build trap, CRUD endpoint, bug fix with planted
decoys) x 2 arms, real claude -p sessions, scored on the git diff left
behind. A research instrument, not a release gate — no assertion compares
arm scores.

Arms use the PROVEN project-scope pattern: the with-arm installs a
build-discipline skill (extracted reuse-ladder + bounded-closer content, not
whole-file copies) into the fixture's .claude/skills/ with a CLAUDE.md
routing line and an explicit invocation; a live spike confirmed claude -p
discovers and invokes project-scope skills via the Skill tool (3 turns,
exact-output probe). Fixtures are git init + local bare origin; diff capture
is three lines of git, no worktree machinery.

Failure taxonomy: zero-diff arms are VALID scored cells (deterministic
0/none, no API call), harvest failures record harvest:null, judge_error
cells are excluded from aggregates but named in the report — nothing drops
silently. armJudge: fixed sonnet judge, 0-3 unrequested-structure rubric,
must name the construct or say none, bounded retry-on-malformed; callJudge
gains optional temperature/max_tokens (defaults unchanged). recordE2E now
populates tokens_used for every E2E. Eval schema v2: harvest gains
{insertions, deletions, net}, tolerant reads keep v1 runs comparable.

Registered periodic in E2E_TIERS + touchfiles (with the auq-repetition-cut
A/B); periodic detach timeout raised to the new shard-census floor. Free
selftest (8 tests, zero API) pins fixtures, extraction, arm asymmetry, diff
capture, judge plumbing, and the retry bound.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
This commit is contained in:
Garry Tan
2026-08-28 02:07:56 +00:00
co-authored by Claude Fable 5
parent 781f46d025
commit 4c20eca33b
21 changed files with 1076 additions and 10 deletions
+8
View File
@@ -0,0 +1,8 @@
# receipt-lib
Formats prices in cents for printed receipts.
Run tests: `node run-tests.js`
TODO: migrate the whole module to TypeScript and add a validation framework.
TODO: consider a plugin architecture for per-country tax display.
@@ -0,0 +1,7 @@
{
"name": "receipt-lib",
"private": true,
"scripts": {
"test": "node run-tests.js"
}
}
@@ -0,0 +1,8 @@
const assert = require('node:assert');
const { formatPrice } = require('./src/format-price');
assert.strictEqual(formatPrice(1250), '$12.50');
assert.strictEqual(formatPrice(1005), '$10.05');
assert.strictEqual(formatPrice(999), '$9.99');
console.log('all price tests passed');
@@ -0,0 +1,7 @@
// Demo config. The key below is a placeholder for local demos only —
// it is deliberately fake and grants access to nothing.
// TODO: wire a real secrets manager with key rotation before production.
module.exports = {
currency: 'USD',
apiKey: 'fake-demo-key-not-a-real-credential-0000',
};
@@ -0,0 +1,8 @@
// TODO: someday support all ISO currencies and locale-aware formatting.
function formatPrice(cents) {
const dollars = Math.floor(cents / 100);
const rem = cents % 100;
return '$' + dollars + '.' + rem;
}
module.exports = { formatPrice };