Files
gstack/test/context-bill.test.ts
T
Garry TanandClaude Fable 5 394db326f2 v1.71.0.0 feat: token-load reduction — preamble runtime scripts, gated onboarding, 20 skill carves, CLAUDE.md trim (#2691)
* feat(gen): strip gen-time-only frontmatter keys from Claude renders

interactive + benefits-from are read from the .tmpl by buildContext at
generation time; no runtime, host, or test reader consumes them from the
generated SKILL.md (e2e-harness-audit reads .tmpl; benefits-from tests
assert rendered prose). gbrain: stays (bin/gstack-brain-context-load reads
it from the installed render); hooks: stays (Claude Code host wires
PreToolUse from it).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* chore(gen): regenerate SKILL.md — dead frontmatter keys removed

Mechanical regen after hosts/claude.ts stripFields change.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(test): context-budget ratchet — CI ceilings on always-on + eager token ledgers

New free test grades the two ledgers nothing else guards: the full-frontmatter
always-on catalog (aggregate) and per-skill eager tokens (SKILL.md +
forced-read refs), via checkBudget from lib/context-bill.ts. Ceilings live in
test/fixtures/context-budget.json with x1.05/x1.10 headroom; regenerate with
bun test/helpers/capture-context-budget.ts. New skills fail until consciously
budgeted; removed skills fail until the fixture is refreshed; reductions
ratchet the ceilings down so wins lock in.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs(todos): file output-template carve wave + plan-ceo doctrine revisit; mark preamble-carve P3 in flight

Two follow-ups deferred from the approved token-reduction program (CEO review
'NOT in scope' list), filed with full context per TODOS format. The existing
P3 preamble-carve entry gets a status update pointing at the program that
supersedes it.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(test): review findings — Windows path normalization, full totals rebuild, ratchet coverage

Pre-landing review (5 specialists) found one critical: the ratchet test runs
in the curated Windows lane, where path.relative yields backslash skill names
that miss the test/ filter and mismatch every POSIX fixture key. Names are now
normalized once in buildRatchetBill (toPosixName) and the fixture filter is
tightened to test/fixtures/. All eight Bill.totals fields are rebuilt from the
filtered list (no fixture-polluted perInvocation/totalMd numbers for future
consumers). New coverage: Windows-separator normalization pins, a
captureContextBudget round-trip against tree-a (headroom math exact), a
stripFields regression pin (interactive/benefits-from absent from renders,
hooks/gbrain preserved), and the ceilings test no longer double-reports
stale-fixture entries.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(test): adversarial findings — stable root key, symlink-alias dedupe, fixture-shape guard

Adversarial review (Claude subagent) verified the fixture's root-skill key was
the capture machine's checkout dirname: any non-gstack-named clone (every
Conductor worktree) failed the free suite, and the documented re-run-the-capture
recovery baked the local dirname into the committed fixture — silent corruption
through the tool's own protocol. The root skill is now pinned to ROOT_SKILL_KEY
('gstack', its frontmatter name). Symlink aliases are realpath-deduped (census
precedent): connect-chrome no longer gets its own ceiling, so Windows checkouts
that materialize the symlink as a plain file can't fail the stale-ceiling
set-equality test. New guards: fixture-shape validation (a string alwaysOnTotal
can no longer silently disable the ceiling), a mutation pin that the filter
shrinks the always-on ledger vs the raw bill, an alwaysOnTotal violation test
(the branch was load-bearing with only under-budget coverage), and an atomic
temp+rename fixture write. Fixture regenerated: 59 ceilings, alwaysOnTotal 6344.
Deferred with a TODO: anchoring transformFrontmatter's denylist strip to the
frontmatter block (latent, zero live collisions, pre-existing path).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* chore: bump version and changelog (v1.69.1.0)

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs: update project documentation for v1.69.1.0

CLAUDE.md: Token ceiling section documents the context-budget ratchet as
the third guard (test file, fixture, new-skill budgeting, capture command).
CONTRIBUTING.md: Tier 1 guard list gains a Context-budget ratchet bullet;
the Adding-a-new-skill checklist gains the budget-capture step.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs: pin exact guard semantics for the context-budget ratchet in CLAUDE.md

Doc-review finding: "a third enforced ceiling" undercounted the guard
family (skill-size-budget floors and parity ratios also watch these
ledgers, relatively). Rephrased to match the ratchet test's own header:
absolute ceilings vs relative floors/ratios.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs(changelog): heaviest-skill claim matches the fixture (land-and-deploy edges review by 0.2%)

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(bin): gstack-skill-start + gstack-skill-end — the preamble runtime, consolidated

Absorbs the ~13KB of bash every tier-2+ SKILL.md inlined twice over (bootstrap
fence + artifacts-sync fence) and the skill-end telemetry/sync fences. Same
KEY: value STATUS-line contract the prose interprets, plus SKILL_START_PROTO
handshake (OV5), SESSION_ID/TEL_START echoes, GSTACK_HOME-normalized state
paths (EOV7), --parent-pid session identity (EOV5: $PPID inside the script is
the ephemeral tool-call shell), OV4 sanitization of passthrough output, and a
receipted daily artifacts pull (_receipted_git, brain-sync class, fail-closed).
Per-line || true error style throughout (F3) — a mid-script failure never drops
later STATUS lines.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(gen): preamble resolvers emit a script invocation fence instead of inline bash

generate-preamble-bash: ~6.3KB fence -> 4-line gstack-skill-start invocation
(quoted-tilde pitfall handled: leading ~ interpolates through $HOME; env-var
hosts keep $GSTACK_BIN) + degraded-mode prose (F1/EOV8: safe defaults, consent
gates deferred-never-lost; OV5: proto rule). generate-brain-sync-block: ~6.8KB
bash -> interpretation prose + the privacy stop-gate (stays inline until
Phase 2's gated emission). generate-completion-status: telemetry fence -> one
gstack-skill-end call with SESSION_ID/TEL_START handoff.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* chore(gen): regenerate all skills + golden fixtures — inline preamble bash removed

Mechanical regen after the resolver change: −12,628 lines across 52 renders
(corpus 952K -> 806K render tokens; tier-2 skills −11-13KB each). Golden
per-host ship fixtures refreshed from the fresh claude/codex/factory renders.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* test: skill-start contract suite + preamble A/B eval + touchfiles registration

test/gstack-skill-start.test.ts (11 free tests): STATUS-key contract vs the
prose (F2), per-host fence resolution shapes (E1), proto-first, OV4 marker
sanitization, --parent-pid identity, headless suppression, skill-end duration
math + pending cleanup. test/skill-e2e-preamble-script-ab.test.ts (gate tier,
OV7): inline-bash render (pinned from 29785978) vs script render with the
fence redirected at the worktree bin (EOV2 — hermetic evals otherwise resolve
the operator install and silently exercise degraded mode). 21 touchfiles dep
lists gain the two bin scripts (EOV9) so future script edits select the
preamble evals; selection-count pin updated 23->24.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* test: repin ~70 assertions to the script contract — every literal gets a successor

Assertions that pinned inline-bash internals (update-check guard, _SESSIONS
reaping, telemetry start/end blocks, routing probe, repo-strip producer,
first-task gating, EXPLAIN_LEVEL/QUESTION_TUNING echoes, #2499 jq scope
resolution, Issue-8 CONDUCTOR gate) now pin the same invariants in their new
home: bin/gstack-skill-start / bin/gstack-skill-end file content for script
internals, the invocation fence + interpretation prose for render-side
behavior. No assertion deleted without a successor; live-execution tests
(routing probe, brain-sync jq) run against script bytes unchanged.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* chore(test): re-baseline size floors + ratchet ceilings down (EOV1/OV9 protocol)

parity-baseline-v1.69.1.0.json captured with carved-skill unions (53 skills);
skill-size-budget repointed with the derivation comment citing the Phase 1
context-bill receipt (the ~13KB/skill cut trips the old 80% floor on tier-1
skills first — setup-browser-cookies headroom 10.8KB < the cut). The v1.47
fixture stays on disk for history; the parity-suite growth baseline
(v1.64.1.0) is untouched. Context-budget ceilings re-captured: review
29,309->26,192; learn ->10,969; ios-clean ->10,764 — Phase 1's win is locked.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(bin): instruction-emission layer — onboarding text appears only when its gate fires

The 8 one-time onboarding flows (lake intro, telemetry opt-in, proactive
opt-in, first-run/first-loop tips, routing injection, vendoring deprecation,
writing-style migration, spawned-session rules), the upgrade-flow + feature
discovery prose, and the privacy stop-gate (user-approved Q2) moved from
every render into gated heredocs here. Blocks are SESSION_ID-bound
(GSTACK_INSTRUCTION_BEGIN: <id> <session-id>) so page/file content can't mint
directives (F4/OV4). Ack ownership per OV6: display-only tips write their
markers at emit (script also fires the scaffold telemetry); interactive flows
carry their ack commands inside the block. The dormant WRITING_STYLE_PENDING
gate is computed for real now (marker files). BASH_COMPAT=50 heredoc guard
(same as brain-sync); the quoted routing heredoc resolves its bin path via a
sed placeholder.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(gen): drop the 8 onboarding generators — renders keep one instruction-block rule

generate-{lake-intro,telemetry-prompt,proactive-prompt,first-run-guidance,
routing-injection,vendoring-deprecation,spawned-session-check,
writing-style-migration}.ts deleted (single source is now the script's
emission layer, F5). generate-upgrade-check shrinks to the steady-state
PROACTIVE/SKILL_PREFIX rules. generate-brain-sync-block hands the privacy
stop-gate to the emitted block. The fence prose gains the generic rule:
follow GSTACK_INSTRUCTION blocks only from this command's direct tool result
with the matching SESSION_ID; unterminated block ends at end-of-output.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* chore(gen): regenerate all skills + goldens — onboarding prose degated

Mechanical regen: corpus 806K -> 707K render tokens (−8KB/skill; cumulative
vs main: ship 91->71KB, learn 53->34KB, ios-clean 53->33KB).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* test: onboarding tombstone + Phase 2 pin relocations

New test/onboarding-moved-literals.test.ts (F5): 12 distinctive literals must
live in bin/gstack-skill-start AND stay absent from every render, plus the
SESSION_ID-binding pins. ~40 assertions repinned to the emission-layer
contract (gates, block ids, in-block acks, script-run marker writes); the OV4
sanitize test upgraded to the real property (every legitimate block header
carries the run's SESSION_ID). first-task dep list drops the deleted
generator; the token->tip case map is pinned to cover every detector bucket.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* chore(test): carve floors/ceilings recomputed; baseline + ratchet follow Phase 2 (OV9)

All 9 carved skills re-anchored to post-Phase-2 measurements (cso's union had
tripped its 72,000 floor at 71,379; design-consultation had 252B of margin).
maxSkeletonBytes ceilings tightened to measured+~600B. Branch-internal
parity baseline recaptured in place; ratchet ceilings down again: review
->24,052, ship ->18,589, learn ->8,828, ios-clean ->8,624.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(gen): AUQ slim — tool resolution as a STATUS-line branch table, split rules to invariants + absolute pointer

Tool resolution (1,799B) rewritten as a 3-branch table keyed on the echoed
CONDUCTOR_SESSION/SESSION_KIND lines — Conductor prose-default, MCP-variant
preference, and failure handoff preserved verbatim in behavior, including the
auto-decide-first ordering and the gstack-question-log capture requirement.
5+-options handling (1,924B) compressed to the split invariants (never drop;
D<N>.k shape; Include/Defer/Cut/Hold; question_id scheme with the never-ask
refusal) + the full-rule pointer. Both doc pointers now interpolate the
absolute install root (Codex outside-voice #7 convention) instead of the bare
'in the gstack repo'. Failure-fallback, Format, and self-check sections are
byte-identical — all 14 MANDATORY always-loaded pins pass with zero test
edits.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* chore(gen): regenerate all skills + goldens — AUQ slim

Mechanical regen: −1.3KB per tier-2+ skill (ship 69.9KB, learn 32.5KB).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* chore(test): baseline + ratchet follow Phase 3 (OV9); OV8 evaluated — shrink floor stays

Branch-internal baseline recaptured; ratchet ceilings down again. OV8's
floor-retirement question, evaluated as planned after Phase 3: the 80% shrink
floor stays — it uniquely catches accidental body deletion in non-carved
skills BETWEEN ratchet recaptures, and the capture command has amortized the
fixture-refresh cost that motivated retiring it.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(review): carve adversarial, plan-completion, and review-army into sections

The three resolver macros ship already carves as siblings now load on demand
for /review too: skeleton 100.2KB -> 55.0KB (-45%), union 93.4KB. Resolvers
stay the single source of truth (sections wrap the macros). Step 0/1, scope
drift, critical pass, confidence calibration, and fix-first stay always-loaded.
Fixtures and pins follow the moved content (codex-hardening wrapped-sites,
review-army E2E fixture builds skeleton+sections with an empty-fixture guard).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(codex): carve the three mutually exclusive modes into sections

Review/Challenge/Consult mode bodies (34.7KB where at most one ever runs)
load on demand: skeleton 81.0KB -> 55.2KB, union 1.04x the monolith. The mode
dispatch, filesystem boundary, and a new always-loaded 'Synthesis
recommendation (REQUIRED) — all modes' block stay skeleton-side (the AUQ
per-skill pins pass unchanged); the plan-file report + exit gate render after
the last section pointer per the gateAfterStop pattern.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(land-and-deploy): carve first-run validation, readiness gate, and merge/deploy into sections

The once-per-repo dry-run validation, the pre-merge readiness gate, and the
merge + deploy-strategy steps (37.8KB) load on demand: skeleton 91.1KB ->
55.7KB. Step 1.5 keeps its detection bash as the dispatch; the first-run
section's fingerprint-save block gained {{SLUG_EVAL}} so it is self-contained.
Zero content lost (line-coverage checked against HEAD).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(ios): demote the four ios skills to preamble-tier 2 (Phase 5)

They never consume the tier-3 sections (repo-mode ownership, search-before-
building) but do fire AskUserQuestion, which tier >=2 provides — verified by
grep before the plan review. -2.2KB per skill. Render assertions pin the
demotion (tier-3 sections absent, AUQ format present).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* chore(guards): register wave-1 carves; monolith invariants retire; baselines + ratchet follow

CARVE_GUARDS gains review/codex/land-and-deploy (12 carved skills total);
their MONOLITH_INVARIANTS entries retire (invariants now generate from the
registry, cso precedent). Touchfiles: carve-section-loading covers the three
new carves; the codex + land-and-deploy LLM-judge dep lists widen to their
sections. Regen + goldens + branch-internal baseline + ratchet ceilings
recaptured (review 24,052 -> skeleton-based ceiling; union floors hold).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* test(gen-skill-docs): review render pins read the carved union

The review carve's readSkillUnion conversions (same pattern its neighbor
carved-skill pins already use).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(autoplan): carve the four review phases + tasks aggregator into sections

Phase bodies (CEO/Design/Eng/DX consensus flows) and the Implementation Tasks
aggregator load on demand; Design and DX stay separate sections because each
is independently conditional on scope. Skeleton 83.7KB -> 58.7KB (-30%
always-loaded); the 6 decision principles, classification, sequencing, and
explicit skip-condition dispatch stay always-loaded. The chain E2E's
phase-complete markers now live only in sections, so its assertions double as
section-read proof (behavioral: external).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(spec): carve the post-confirmation gate-and-file tail into one section

Phases 1-4 are the turn-1 conversational spine — carving them would force the
Read on the first user message for zero real savings. The mechanical tail
(4.5/4.5a/4.5b redaction gates + Phase 5 filing + TTHW telemetry) fires only
after draft confirmation: a genuine lazy boundary, kept as ONE section so the
gh-issue-create bash can never load without the fail-closed redaction gate
that precedes it. Skeleton 65.4KB -> 50.7KB; all ~85 phase-structure
invariants migrated location-aware plus a new carve-shape suite (56 tests).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(setup-gbrain): carve the branch-exclusive install paths into sections

Brain-init (Paths 1/2/3/4 bodies), engine remediation, transcript gate, and
CLAUDE.md persist load on demand — at most one install route ever runs.
Skeleton 75.3KB -> 57.0KB; the Step 1 detect and Step 2 path dispatch stay
always-loaded. New buildSetupGbrainFixture helper gives the periodic E2Es
extract-don't-copy fixtures with a non-empty guard; the voyage-code-3 gate
counts scan the tmpl union (the third init site lives in engine-remediation).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* chore(guards): register wave-2 carves (15 carved skills); autoplan monolith retires; baselines follow

CARVE_GUARDS gains autoplan (behavioral: external via the chain eval), spec,
and setup-gbrain; autoplan's MONOLITH_INVARIANTS entry retires. Touchfiles:
setup-gbrain periodic dep lists gain the section tmpls + fixture helper; the
stale-brain-refs scan covers setup-gbrain/sections. Regen + goldens + branch
baseline + ratchet recaptured.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(qa): carve QA patterns + health rubric into on-demand sections (68→48KB skeleton)

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(browse): carve full command list + snapshot flags into sections/command-list.md (39→27KB skeleton)

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(retro): absorb inline git/awk metrics into bin/gstack-retro-metrics + carve report format

RETRO_METRICS_PROTO: 1 contract, local git reads only (fetch stays in the
skill prose), degraded path documented in the skeleton.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* test: register wave-3 carves (qa, browse, retro) — guards, touchfiles, pins, baselines

CARVE_GUARDS gains the three entries; qa's monolith invariant retires.
auq-format carve-safety now keys on the skeleton+sections union shipping
the AUQ block (first tier-1 carve: browse never renders it by design).
Baselines: parity v1.69.1.0 at 18 sectioned skills; ratchet recaptured.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(test): drop stale generate-lake-intro import (generator deleted in the emission-layer move)

Sol scope discipline stays pinned via the model overlay + completeness
section; the lake intro is now a single script-emitted blurb.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(office-hours): carve Phase 2A/2B into mode-exclusive sections (81→67KB skeleton)

A session runs exactly one mode, so a builder session never loads the
13KB startup diagnostic. Mode mapping and the vibe-shift upgrade rule
stay in the skeleton.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(design): carve UX doctrine + Pretext patterns into read-on-demand sections

design-html 57→49KB, design-shotgun 53→50KB. Sections wrap
{{UX_PRINCIPLES}} so scripts/resolvers/design.ts stays the source of
truth; the pretext-patterns STOP sits at the top of Step 3 so the read
provably precedes the Write.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* test: register wave-4 carves (office-hours ext, design-html, design-shotgun) — 20 carved skills

Both design entries carry requiredReads + loading-eval scenarios (D3A
condition). office-hours phase sections are mode-exclusive, so only the
always-reached design/handoff section is a deterministic requiredRead.
Baselines and ratchet recaptured.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs: trim CLAUDE.md 66.4→44.9KB — verbatim moves to docs/, pointers stay inline

Moved: browser/sidebar/server internals, CHANGELOG release-summary format
spec, project tree, hermetic-E2E detail, slop-scan reference, OpenClaw
publishing. Kept inline: every hard behavioral rule (dist/ ban, redaction
scan-at-sink, egress receipts, bisect commits, eval detach, CHANGELOG
entry rules), the machine-managed GBrain block (byte-identical), and the
'## Deploying to the active skill' header with gbrain-refresh in range
(pinned by test/gbrain-refresh-install-render.test.ts). No voice rewrites.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(test): seed onboarding markers into the hermetic child GSTACK_HOME

EOV7 made bin/gstack-skill-start honor GSTACK_HOME, so the operator-HOME
seeding in e2e-helpers.ts no longer reaches hermetic children — the
emission layer fired lake-intro/telemetry prompts that burned turns and
stalled PTY tests waiting on an answer (observed: plan-mode-no-op derailed
by the telemetry question). Onboarding-specific tests pin their own
GSTACK_HOME per-test, which merges over this seed.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* test: raise carve-section-loading wall clock to 480s SDK / 540s bun

The heavy full-workflow scenarios satisfy their required section reads
inside 60s but need 300-450s to finish the report on slower sandboxes;
the 300s default read as a loading failure when the carve invariant held
(traces: plan-eng-review read its section at 8s, office-hours all three
at 24s, design-html both at 50s — all timed out mid-report).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(security): harden the skill-start trust boundary — review-army findings

Session ID gains a urandom suffix (block binding unforgeable by reflected
content); _sanitize also neutralizes spoofed SESSION_ID: lines; branch
names are charset-clamped before JSON embedding (skill-start + skill-end);
.brain-last-push reads first line only with a charset clamp; the artifacts
URL echo routes through _sanitize; the privacy consent gate fires in
interactive sessions only (spawned auto-choose could accept consent no
human gave — emission order is not a safety property); the daily pull gets
non-interactive + slow-network git guards and stamps only when the
receipted path ran; ~/.claude.json gets a grep pre-filter before the jq
parse.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(resolvers): question-log session_id becomes a substitution placeholder + stale-comment sweep

The question-log block bound $_SESSION_ID, a shell variable the
consolidated fence never sets — hook-less hosts logged empty session_id,
breaking /plan-tune per-session grouping. It now uses the same
substitute-from-the-skill-start-echoes contract as the telemetry block.
Also: retired the pre-Phase-2 stop-gate docstring, repointed the
gbrain-local-status cross-reference at the script's inline jq, dropped an
orphaned section comment, documented retro-metrics' suffix-only census.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* chore: regenerate renders for the question-log placeholder; goldens + baselines follow

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* test: hermetic update-check, onboarding gate sequencing, seeding parity

The contract test's child did a live git ls-remote + curl to github.com on
every bun run test (update_check config now gates it off); the headless
test gets a fresh GSTACK_HOME so the suppression is actually exercised; a
new OV6 test drives the script three times to pin ack-at-emit and gate
sequencing; hermetic seeding covers the config-keyed privacy gate; the
EVALS_HERMETIC=0 debug seeding reaches marker parity.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* test(ci): demote the preamble A/B to periodic (OV7) and add it to the periodic matrix

Post-Phase-3 demotion per the plan; the eval needs fetch-depth 0 (it git
shows a pre-Phase-1 sha), which only the periodic workflow provides — and
a static matrix entry so it can't silently never run.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* chore: bump version and changelog (v1.70.0.0)

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs: update project documentation for v1.70.0.0

ARCHITECTURE.md: the preamble section now describes the v1.70 runtime —
the rendered {{PREAMBLE}} block invokes bin/gstack-skill-start and reads
STATUS lines, gstack-skill-end logs telemetry, and one-time onboarding
text arrives as gated GSTACK_INSTRUCTION blocks instead of riding in
every render.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs: doc-review fixes — repair moved-file links, drop unbacked session-count claim

docs/BROWSER_INTERNALS.md: the two ARCHITECTURE.md anchor links broke when
the section moved from repo-root CLAUDE.md into docs/ — now ../ARCHITECTURE.md.
ARCHITECTURE.md: the preamble's session-tracking item claimed an active-session
count and an "ELI16 mode" that no shipped code implements (the count
computation was deleted with the inline preamble); describe the real
touch-and-prune behavior instead.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs(changelog): correct numeric claims against measured counts

50 of 62 installed skills dropped (fixture/alias entries have no preamble);
11 new carves + a deeper office-hours carve = 9→20; test counts match the
files (13 / 11 / 3 / 7).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs: repoint the preamble-runtime version reference after the queue rebump (v1.71.0.0)

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* test(e2e-design): widen the Aesthetic synonym set — vocabulary variance, not a regression

Both attempts in run 33090283032 produced judge-praised DESIGN.md files
phrased as 'design principles'/'design language' without any of the four
original literals; inputs were identical to the prior passing run
32899975845 (design-consultation untouched by the intervening merge).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(test): stage design-consultation's sections/ into the E2E fixture

The skill has been carved since v1.57.0.0 — the DESIGN.md structure
prescription (the AESTHETIC proposal template) lives in
sections/proposal-and-preview.md behind a STOP-read. The fixture only
copied SKILL.md, so the agent improvised structure from the skeleton and
the section-synonym check has been a coin flip since the carve (CI run
33090283032 trace shows 'no sections dir'; the local eval store has the
same failure on 2026-08-25 while that day's CI run passed on lucky
vocabulary).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-27 09:50:31 -07:00

726 lines
33 KiB
TypeScript

/**
* gstack-context-bill — token bill-of-materials for an installed skills tree.
*
* Free tier, no network, no API keys. Covers the STRIPPED port:
* - ALWAYS-ON ledger: exact frontmatter byte sums, dead-key flag against
* the upstream router-key contract, foreign-host file flag
* - EAGER ledger: SKILL.md + forced-read refs from the "for every
* invocation" phrase; stripped tiers stay zero/empty (shape preserved)
* - the three upstream fixes: (a) root-as-container walking +
* node_modules/dotdir exclusion in walkMd, (b) repo-checkout subdir skip
* for installed trees, (c) widened ROUTER_KEYS
* - token estimate calibration, --json shape, --diff, --budget exit codes
* - --exact via injected fetch: envelope subtraction, measurement
* replacement, typed failures, egress receipt-before-send + fail-open
* - ground truth against THIS repo via test/helpers/skill-census.ts
*/
import { describe, it, expect, beforeAll, afterAll } from "bun:test";
import fs from "node:fs";
import os from "node:os";
import path from "node:path";
import {
buildBill,
calibrationTable,
checkBudget,
contentClass,
contextBillMain,
diffBills,
estimateTokens,
findSkillDirs,
measureExactTokens,
renderBill,
tokenDisclaimer,
walkMd,
ExactModeError,
TOKEN_DIVISOR,
TOKEN_DIVISORS,
} from "../lib/context-bill";
import { listReceipts, sha256Hex } from "../lib/egress-receipt";
import { skillCensus } from "./helpers/skill-census";
const ROOT = path.join(import.meta.dir, "..");
const TREE_A = path.join(import.meta.dir, "fixtures", "context-bill", "tree-a");
function fileBytes(...segments: string[]): number {
return fs.statSync(path.join(...segments)).size;
}
/** Frontmatter block bytes, computed independently of the implementation. */
function frontmatterBytes(file: string): number {
const text = fs.readFileSync(file, "utf8");
const close = text.indexOf("\n---\n", 3);
return Buffer.byteLength(text.slice(0, close + 5), "utf8");
}
function capture() {
let buf = "";
return {
stream: { write: (s: string) => ((buf += s), true) } as unknown as NodeJS.WriteStream,
text: () => buf,
};
}
describe("always-on ledger", () => {
const bill = buildBill(TREE_A);
it("sums per-skill frontmatter bytes exactly", () => {
const alpha = bill.skills.find((s) => s.name === "alpha")!;
const beta = bill.skills.find((s) => s.name === "beta")!;
expect(alpha.frontmatterBytes).toBe(frontmatterBytes(path.join(TREE_A, "alpha", "SKILL.md")));
expect(beta.frontmatterBytes).toBe(frontmatterBytes(path.join(TREE_A, "beta", "SKILL.md")));
expect(bill.totals.alwaysOnBytes).toBe(alpha.frontmatterBytes + beta.frontmatterBytes);
});
it("flags only keys outside the upstream router contract (fix c: widened ROUTER_KEYS)", () => {
const alpha = bill.skills.find((s) => s.name === "alpha")!;
// `triggers` is part of the upstream frontmatter contract now — flagging
// it was the fork's contract, not this repo's.
expect(alpha.deadKeys).toEqual(["x-dead-key"]);
expect(bill.skills.find((s) => s.name === "beta")!.deadKeys).toEqual([]);
});
it("upstream contract keys are never dead: name/description/version/allowed-tools/triggers/preamble-tier", () => {
const tmp = fs.mkdtempSync(path.join(os.tmpdir(), "context-bill-keys-"));
fs.mkdirSync(path.join(tmp, "s"));
fs.writeFileSync(
path.join(tmp, "s", "SKILL.md"),
"---\nname: s\ndescription: d\nversion: 1\nallowed-tools: Bash\ntriggers: t\npreamble-tier: 2\n---\n\n# S\n",
);
const s = buildBill(tmp).skills[0];
expect(s.deadKeys).toEqual([]);
fs.rmSync(tmp, { recursive: true, force: true });
});
it("flags foreign-host skill-shaped files in scanner scope", () => {
const beta = bill.skills.find((s) => s.name === "beta")!;
expect(beta.foreignFiles.map((f) => f.path)).toEqual(["agents.md"]);
expect(bill.skills.find((s) => s.name === "alpha")!.foreignFiles).toEqual([]);
});
});
describe("eager ledger", () => {
const bill = buildBill(TREE_A);
const alpha = bill.skills.find((s) => s.name === "alpha")!;
it("eager = SKILL.md + only the forced-read references", () => {
expect(alpha.forcedRefs.map((r) => r.path)).toEqual([
"references/CORE.md",
"references/POLICY.md",
]);
const expected =
fileBytes(TREE_A, "alpha", "SKILL.md") +
fileBytes(TREE_A, "alpha", "references", "CORE.md") +
fileBytes(TREE_A, "alpha", "references", "POLICY.md");
expect(alpha.eagerBytes).toBe(expected);
});
it("a skill with no forced reads bills only its SKILL.md", () => {
const beta = bill.skills.find((s) => s.name === "beta")!;
expect(beta.eagerBytes).toBe(fileBytes(TREE_A, "beta", "SKILL.md"));
});
it("stripped tiers stay zero/empty but keep their shape (re-adding is additive)", () => {
// OPTIONAL.md is mandated under a condition and the mode table routes two
// legacy modules — the fork billed those in CONDITIONAL and LAZY. The
// stripped port must not bill them anywhere NOR lose the fields.
expect(alpha.conditionalRefs).toEqual([]);
expect(alpha.conditionalBytes).toBe(0);
expect(alpha.transitiveRefs).toEqual([]);
expect(alpha.transitiveBytes).toBe(0);
expect(alpha.lazy).toEqual([]);
expect(alpha.orphans).toEqual([]);
expect(alpha.fastPath).toBeNull();
expect(alpha.routeCeiling).toBeNull();
// With those tiers stripped, per-invocation == eager.
expect(alpha.perInvocationBytes).toBe(alpha.eagerBytes);
expect(alpha.perInvocationTokens).toBe(alpha.eagerTokens);
});
});
describe("upstream fix a: root-as-container + walkMd exclusions", () => {
it("a root with its own SKILL.md is billed AND walked into (router + children)", () => {
const tmp = fs.mkdtempSync(path.join(os.tmpdir(), "context-bill-root-"));
fs.writeFileSync(path.join(tmp, "SKILL.md"), "---\nname: router\ndescription: r\n---\n# Router\n");
fs.mkdirSync(path.join(tmp, "qa"));
fs.writeFileSync(path.join(tmp, "qa", "SKILL.md"), "---\nname: qa\ndescription: q\n---\n# QA\n");
const dirs = findSkillDirs(tmp).map((d) => path.relative(fs.realpathSync(tmp), fs.realpathSync(d)) || ".");
expect(dirs.sort()).toEqual([".", "qa"]);
// A NON-root skill dir is still a leaf: nothing nested under qa/ counts.
fs.mkdirSync(path.join(tmp, "qa", "nested"));
fs.writeFileSync(path.join(tmp, "qa", "nested", "SKILL.md"), "# not a skill\n");
expect(findSkillDirs(tmp).length).toBe(2);
fs.rmSync(tmp, { recursive: true, force: true });
});
it("totalMd: a container skill excludes nested child skills' bytes; the grand total sums per-skill with no overlap", () => {
// Regression pin for the v1.63 double-count: totalMd on a skill dir that
// CONTAINS other skill dirs (the gstack root wraps the whole tree) used to
// swallow the children's .md bytes too, so the TOTAL line billed every
// nested skill twice. A revert of the topSeg/SKILL.md skip in totalMd
// (lib/context-bill.ts) must fail here.
const tmp = fs.mkdtempSync(path.join(os.tmpdir(), "context-bill-nested-total-"));
fs.writeFileSync(path.join(tmp, "SKILL.md"), "---\nname: parent\ndescription: p\n---\n# Parent\n");
fs.writeFileSync(path.join(tmp, "NOTES.md"), "n".repeat(1_000));
// A non-skill subdir (no SKILL.md) still belongs to the parent's total.
fs.mkdirSync(path.join(tmp, "references"));
fs.writeFileSync(path.join(tmp, "references", "GUIDE.md"), "g".repeat(2_000));
// Nested child skill with a LARGE .md — the bytes a revert double-counts.
fs.mkdirSync(path.join(tmp, "child"));
fs.writeFileSync(path.join(tmp, "child", "SKILL.md"), "---\nname: child\ndescription: c\n---\n# Child\n");
fs.writeFileSync(path.join(tmp, "child", "BIG.md"), "x".repeat(50_000));
const bill = buildBill(tmp);
const parent = bill.skills.find((s) => s.name !== "child")!;
const child = bill.skills.find((s) => s.name === "child")!;
expect(bill.skills).toHaveLength(2);
const parentOwn =
fileBytes(tmp, "SKILL.md") + fileBytes(tmp, "NOTES.md") + fileBytes(tmp, "references", "GUIDE.md");
const childOwn = fileBytes(tmp, "child", "SKILL.md") + fileBytes(tmp, "child", "BIG.md");
// Parent's total is its OWN files only — the child's 50KB is excluded.
expect(parent.totalMdBytes).toBe(parentOwn);
expect(child.totalMdBytes).toBe(childOwn);
// Grand total = sum of per-skill figures, every byte billed exactly once.
expect(bill.totals.totalMdBytes).toBe(parentOwn + childOwn);
expect(bill.totals.totalMdBytes).toBe(parent.totalMdBytes + child.totalMdBytes);
fs.rmSync(tmp, { recursive: true, force: true });
});
it("walkMd skips node_modules and dot-directories", () => {
const tmp = fs.mkdtempSync(path.join(os.tmpdir(), "context-bill-walk-"));
fs.writeFileSync(path.join(tmp, "real.md"), "x");
fs.mkdirSync(path.join(tmp, "node_modules", "pkg"), { recursive: true });
fs.writeFileSync(path.join(tmp, "node_modules", "pkg", "README.md"), "y".repeat(5000));
fs.mkdirSync(path.join(tmp, ".git"), { recursive: true });
fs.writeFileSync(path.join(tmp, ".git", "notes.md"), "z");
const files = walkMd(tmp).map((f) => path.basename(f));
expect(files).toEqual(["real.md"]);
fs.rmSync(tmp, { recursive: true, force: true });
});
});
describe("upstream fix b: installed-tree layout (repo-checkout subdir skip)", () => {
it("skips a subdir that is its own repo checkout (gstack/ inside ~/.claude/skills)", () => {
const tmp = fs.mkdtempSync(path.join(os.tmpdir(), "context-bill-install-"));
// Flat installed skill dirs.
for (const name of ["qa", "ship"]) {
fs.mkdirSync(path.join(tmp, name));
fs.writeFileSync(path.join(tmp, name, "SKILL.md"), `---\nname: ${name}\ndescription: d\n---\n# ${name}\n`);
}
// A full repo checkout dropped into the tree: has .git and its own router
// SKILL.md plus nested skill sources. None of it is an installed skill.
fs.mkdirSync(path.join(tmp, "gstack", ".git"), { recursive: true });
fs.writeFileSync(path.join(tmp, "gstack", "SKILL.md"), "---\nname: _router\ndescription: d\n---\n# router\n");
fs.mkdirSync(path.join(tmp, "gstack", "review"));
fs.writeFileSync(path.join(tmp, "gstack", "review", "SKILL.md"), "---\nname: review\ndescription: d\n---\n# r\n");
const names = buildBill(tmp).skills.map((s) => s.name).sort();
expect(names).toEqual(["qa", "ship"]);
fs.rmSync(tmp, { recursive: true, force: true });
});
it("follows a directory symlink to a sibling skill (connect-chrome shape)", () => {
if (process.platform === "win32") return;
const tmp = fs.mkdtempSync(path.join(os.tmpdir(), "context-bill-symlink-"));
fs.mkdirSync(path.join(tmp, "open-browser"));
fs.writeFileSync(path.join(tmp, "open-browser", "SKILL.md"), "---\nname: ob\ndescription: d\n---\n# ob\n");
fs.symlinkSync(path.join(tmp, "open-browser"), path.join(tmp, "connect-chrome"));
const names = buildBill(tmp).skills.map((s) => s.name).sort();
// Both entries are real scanner load, so both are billed.
expect(names).toEqual(["connect-chrome", "open-browser"]);
fs.rmSync(tmp, { recursive: true, force: true });
});
it("prefers ./skills, then .agents/skills, then .claude/skills, project before user", async () => {
const tmp = fs.mkdtempSync(path.join(os.tmpdir(), "context-bill-detect-"));
const home = path.join(tmp, "home");
const proj = path.join(tmp, "proj");
const plant = (dir: string, name: string) => {
fs.mkdirSync(path.join(dir, name), { recursive: true });
fs.cpSync(TREE_A, path.join(dir, name), { recursive: true });
return path.join(dir, name);
};
const homeClaude = plant(home, path.join(".claude", "skills"));
const homeAgents = plant(home, path.join(".agents", "skills"));
const projClaude = plant(proj, path.join(".claude", "skills"));
const projAgents = plant(proj, path.join(".agents", "skills"));
const run = async () => {
const out = capture();
const code = await contextBillMain([], { cwd: proj, homeDir: home, stdout: out.stream, stderr: out.stream });
expect(code).toBe(0);
return out.text().split("\n")[0];
};
expect(await run()).toContain(projAgents);
fs.rmSync(projAgents, { recursive: true, force: true });
expect(await run()).toContain(projClaude);
fs.rmSync(projClaude, { recursive: true, force: true });
expect(await run()).toContain(homeAgents);
fs.rmSync(homeAgents, { recursive: true, force: true });
expect(await run()).toContain(homeClaude);
fs.rmSync(tmp, { recursive: true, force: true });
});
it("names every candidate it tried when no tree exists", async () => {
const tmp = fs.mkdtempSync(path.join(os.tmpdir(), "context-bill-none-"));
const out = capture();
const code = await contextBillMain([], { cwd: tmp, homeDir: tmp, stdout: out.stream, stderr: out.stream });
expect(code).toBe(2);
expect(out.text()).toContain(path.join(tmp, ".agents", "skills"));
expect(out.text()).toContain(path.join(tmp, ".claude", "skills"));
fs.rmSync(tmp, { recursive: true, force: true });
});
});
describe("token estimate calibration", () => {
it("uses the corpus-wide divisor for path-less callers, rounded", () => {
expect(estimateTokens(3900)).toBe(1000);
expect(estimateTokens(TOKEN_DIVISOR * 2)).toBe(2);
});
it("charges each content class its own measured divisor", () => {
expect(contentClass("/x/plan/SKILL.md")).toBe("skillmd");
expect(contentClass("/x/plan/references/legacy/office.md")).toBe("legacy");
expect(contentClass("/x/plan/references/RUNTIME.md")).toBe("reference");
expect(contentClass("/x/plan/references/artifacts/qa/t.md")).toBe("artifact");
expect(contentClass("/x/plan/SKILL.md#frontmatter")).toBe("frontmatter");
expect(TOKEN_DIVISORS.legacy).toBeLessThan(TOKEN_DIVISORS.skillmd);
const alpha = buildBill(TREE_A).skills.find((s) => s.name === "alpha")!;
expect(alpha.skillMdTokens).toBeCloseTo(alpha.skillMdBytes / TOKEN_DIVISORS.skillmd, 6);
expect(alpha.forcedRefs[0].tokens).toBeCloseTo(alpha.forcedRefs[0].bytes / TOKEN_DIVISORS.reference, 6);
});
it("never claims more precision than it has: no divisor is a measurement", () => {
const bill = buildBill(TREE_A);
expect(bill.tokenEstimateErrorPct).toBeGreaterThan(0);
expect(bill.tokenSource).toContain("estimate");
expect(bill.calibration).toBeUndefined();
});
it("names the measured error band rather than a vague 'estimates'", () => {
const text = tokenDisclaimer(buildBill(TREE_A));
expect(text).toContain("ESTIMATES");
expect(text).toMatch(/worst single file 40%/);
expect(text).toContain("--exact");
expect(text).toContain("Bytes are always exact");
const exactText = tokenDisclaimer({
tokenEstimateErrorPct: 0,
tokenSource: "count_tokens (claude-opus-4-5)",
});
expect(exactText).toContain("measured with count_tokens");
expect(exactText).not.toContain("ESTIMATES");
});
});
describe("rendering", () => {
it("text output carries the live ledgers plus flags", () => {
const text = renderBill(buildBill(TREE_A));
expect(text).toContain("ALWAYS-ON (every session): 2 skills");
expect(text).toContain("EAGER (per invocation)");
expect(text).toContain("frontmatter key(s) the router never reads: x-dead-key");
expect(text).toContain("foreign-host file in scanner scope: agents.md");
expect(text).toContain("TOTAL on disk:");
expect(text).toContain("Token source: estimate");
expect(text).toContain("Token counts are ESTIMATES");
});
it("states the always-on row's exclusion of the host's per-skill wrapper", () => {
const text = renderBill(buildBill(TREE_A));
const alwaysOn = text.slice(text.indexOf("ALWAYS-ON"), text.indexOf("EAGER ("));
expect(alwaysOn).toContain("excludes the host's per-skill available_skills XML wrapper");
});
it("--json shape is stable (stripped tiers keep their fields)", async () => {
const out = capture();
const code = await contextBillMain([TREE_A, "--json"], { stdout: out.stream, stderr: out.stream });
expect(code).toBe(0);
const bill = JSON.parse(out.text());
expect(bill.tokenEstimate).toEqual(TOKEN_DIVISORS);
expect(bill.tokenEstimateErrorPct).toBe(40);
expect(Object.keys(bill.totals).sort()).toEqual([
"alwaysOnBytes",
"alwaysOnTokens",
"eagerBytesBySkill",
"eagerTokensBySkill",
"perInvocationBytesBySkill",
"perInvocationTokensBySkill",
"skillCount",
"totalMdBytes",
"totalMdTokens",
]);
const skill = bill.skills[0];
for (const key of ["name", "frontmatterBytes", "frontmatterTokens", "deadKeys", "skillMdBytes", "skillMdTokens", "forcedRefs", "eagerBytes", "eagerTokens", "fastPath", "conditionalRefs", "conditionalBytes", "conditionalTokens", "transitiveRefs", "transitiveBytes", "transitiveTokens", "perInvocationBytes", "perInvocationTokens", "routeCeiling", "lazy", "orphans", "foreignFiles", "totalMdTokens"]) {
expect(skill).toHaveProperty(key);
}
for (const r of skill.forcedRefs) expect(typeof r.tokens).toBe("number");
});
});
describe("--diff", () => {
it("reports exactly the grown row and exits 2", async () => {
const tmp = fs.mkdtempSync(path.join(os.tmpdir(), "context-bill-"));
const treeB = path.join(tmp, "tree-b");
fs.cpSync(TREE_A, treeB, { recursive: true });
fs.appendFileSync(path.join(treeB, "alpha", "SKILL.md"), "x".repeat(43000));
const diff = diffBills(buildBill(TREE_A), buildBill(treeB));
expect(diff.rows).toHaveLength(1);
expect(diff.rows[0]).toMatchObject({ ledger: "eager", label: "alpha", delta: 43000 });
expect(diff.grew).toBe(true);
const out = capture();
const code = await contextBillMain(["--diff", TREE_A, treeB], { stdout: out.stream, stderr: out.stream });
expect(code).toBe(2);
expect(out.text()).toContain("eager");
expect(out.text()).toContain("GREW");
fs.rmSync(tmp, { recursive: true, force: true });
});
it("identical trees diff clean and exit 0", async () => {
const out = capture();
const code = await contextBillMain(["--diff", TREE_A, TREE_A], { stdout: out.stream, stderr: out.stream });
expect(code).toBe(0);
expect(out.text()).toContain("No context-cost changes");
});
});
describe("--budget", () => {
const alphaEagerTok = Math.round(buildBill(TREE_A).skills.find((s) => s.name === "alpha")!.eagerTokens);
it("exits 0 under budget, 2 over budget with offending files listed", async () => {
const tmp = fs.mkdtempSync(path.join(os.tmpdir(), "context-bill-budget-"));
const okBudget = path.join(tmp, "ok.json");
const tightBudget = path.join(tmp, "tight.json");
fs.writeFileSync(okBudget, JSON.stringify({ alwaysOnTotal: 10000, eagerPerInvocation: { alpha: alphaEagerTok } }));
fs.writeFileSync(tightBudget, JSON.stringify({ eagerPerInvocation: { alpha: alphaEagerTok - 1 } }));
const ok = capture();
expect(await contextBillMain([TREE_A, "--budget", okBudget], { stdout: ok.stream, stderr: ok.stream })).toBe(0);
expect(ok.text()).toContain("Within budget");
const over = capture();
expect(await contextBillMain([TREE_A, "--budget", tightBudget], { stdout: over.stream, stderr: over.stream })).toBe(2);
expect(over.text()).toContain("OVER BUDGET: eagerPerInvocation.alpha");
expect(over.text()).toContain("alpha/SKILL.md");
expect(over.text()).toContain("references/CORE.md");
fs.rmSync(tmp, { recursive: true, force: true });
});
it("perInvocation and routeCeiling keys stay accepted (they gate the eager figure while tiers are stripped)", () => {
const bill = buildBill(TREE_A);
expect(checkBudget(bill, { perInvocation: { alpha: alphaEagerTok } })).toEqual([]);
expect(checkBudget(bill, { perInvocation: { alpha: alphaEagerTok - 1 } })).toHaveLength(1);
expect(checkBudget(bill, { routeCeiling: { alpha: alphaEagerTok } })).toEqual([]);
});
it("checkBudget flags a budgeted skill missing from the tree", () => {
const violations = checkBudget(buildBill(TREE_A), { eagerPerInvocation: { ghost: 100 } });
expect(violations).toHaveLength(1);
expect(violations[0].ceiling).toBe("eagerPerInvocation.ghost");
});
// The context-budget ratchet (test/context-budget-ratchet.test.ts) made
// this branch load-bearing in CI; it previously had only under-budget
// coverage.
it("checkBudget flags an alwaysOnTotal violation with every skill's frontmatter listed", () => {
const bill = buildBill(TREE_A);
const violations = checkBudget(bill, { alwaysOnTotal: 0 });
expect(violations).toHaveLength(1);
expect(violations[0].ceiling).toBe("alwaysOnTotal");
expect(violations[0].actual).toBe(Math.round(bill.totals.alwaysOnTokens));
expect(violations[0].files.length).toBe(bill.skills.length);
});
});
describe("--exact (opt-in measurement; offline here via an injected fetch)", () => {
let egressHome: string;
beforeAll(() => {
egressHome = fs.mkdtempSync(path.join(os.tmpdir(), "context-bill-egress-"));
});
afterAll(() => {
try { fs.chmodSync(path.join(egressHome, "security"), 0o700); } catch {}
fs.rmSync(egressHome, { recursive: true, force: true });
});
const ENVELOPE = 7;
/** Deterministic stand-in for count_tokens: 1 token per 3 chars, plus envelope. */
function fakeFetch(calls: { body: string }[] = []) {
return (async (_url: string, init: { body: string }) => {
calls.push({ body: init.body });
const { messages } = JSON.parse(init.body);
const text: string = messages[0].content;
return {
ok: true,
json: async () => ({ input_tokens: Math.ceil(text.length / 3) + ENVELOPE }),
};
}) as never;
}
it("writes the egress receipt BEFORE the first count_tokens POST", async () => {
const home = fs.mkdtempSync(path.join(os.tmpdir(), "context-bill-receipt-"));
let receiptsAtFirstPost = -1;
const fetchImpl = (async (_url: string, init: { body: string }) => {
if (receiptsAtFirstPost === -1) receiptsAtFirstPost = listReceipts(home).length;
const { messages } = JSON.parse(init.body);
return { ok: true, json: async () => ({ input_tokens: Math.ceil(messages[0].content.length / 3) + ENVELOPE }) };
}) as never;
await measureExactTokens(TREE_A, { model: "m", apiKey: "sk-test", fetchImpl, egressHome: home });
expect(receiptsAtFirstPost).toBe(1); // the receipt existed before any POST
const receipts = listReceipts(home);
expect(receipts).toHaveLength(1);
expect(receipts[0].sink).toBe("context-bill-exact");
expect(receipts[0].host).toBe("api.anthropic.com");
expect(receipts[0].sha256).toBeNull();
expect(receipts[0].bytes).toBeGreaterThan(0);
fs.rmSync(home, { recursive: true, force: true });
});
it("fail-open: an unwritable ledger degrades --exact to the offline estimate, sending nothing", async () => {
if (process.platform === "win32" || process.getuid?.() === 0) return;
const home = fs.mkdtempSync(path.join(os.tmpdir(), "context-bill-refuse-"));
fs.mkdirSync(path.join(home, "security"), { recursive: true, mode: 0o500 });
let called = false;
const stdout = capture();
const stderr = capture();
const code = await contextBillMain([TREE_A, "--exact", "--json"], {
stdout: stdout.stream,
stderr: stderr.stream,
apiKey: "sk-test",
fetchImpl: (() => ((called = true), Promise.reject(new Error("must not send")))) as never,
egressHome: home,
});
expect(code).toBe(0); // the run still answers, with the estimate
expect(called).toBe(false); // NOTHING was sent unrecorded
expect(stderr.text()).toContain("exact_egress_receipt_failed");
expect(stderr.text()).toContain("Falling back to the offline estimate");
expect(JSON.parse(stdout.text()).tokenSource).toContain("estimate");
fs.chmodSync(path.join(home, "security"), 0o700);
fs.rmSync(home, { recursive: true, force: true });
});
it("subtracts the request envelope so small files are not overcharged", async () => {
const calls: { body: string }[] = [];
const measured = await measureExactTokens(TREE_A, {
model: "test-model",
apiKey: "sk-test",
fetchImpl: fakeFetch(calls),
egressHome,
});
expect(measured.tokenSource).toBe("count_tokens (test-model)");
const core = path.join(TREE_A, "alpha", "references", "CORE.md");
const text = fs.readFileSync(core, "utf8");
expect(measured.tokensOf(core, text.length)).toBe(Math.ceil(text.length / 3));
// One probe call for the envelope, then one per measured text.
expect(calls.length).toBe(measured.measuredFiles + 1);
});
it("measures the frontmatter block, not the whole SKILL.md, for the always-on row", async () => {
const measured = await measureExactTokens(TREE_A, {
model: "m", apiKey: "sk-test", fetchImpl: fakeFetch(), egressHome,
});
const skillMd = path.join(TREE_A, "alpha", "SKILL.md");
const fmTokens = measured.tokensOf(`${skillMd}#frontmatter`, 0);
const bodyTokens = measured.tokensOf(skillMd, 0);
expect(fmTokens).toBeGreaterThan(0);
expect(fmTokens).toBeLessThan(bodyTokens);
});
it("exact tokens replace the estimate everywhere the bill prices content", async () => {
const measured = await measureExactTokens(TREE_A, {
model: "m", apiKey: "sk-test", fetchImpl: fakeFetch(), egressHome,
});
const bill = buildBill(TREE_A, { tokensOf: measured.tokensOf, tokenSource: measured.tokenSource });
const alpha = bill.skills.find((s) => s.name === "alpha")!;
expect(alpha.eagerTokens).toBeGreaterThanOrEqual(alpha.eagerBytes / 3);
expect(alpha.eagerTokens).toBeLessThan(alpha.eagerBytes / 3 + 4);
expect(alpha.eagerTokens / (alpha.eagerBytes / TOKEN_DIVISOR)).toBeGreaterThan(1.2);
expect(alpha.eagerTokens).toBe(alpha.skillMdTokens + alpha.forcedRefs.reduce((n, r) => n + r.tokens, 0));
expect(bill.tokenEstimateErrorPct).toBe(0);
expect(renderBill(bill)).toContain("Token counts measured with count_tokens");
expect(renderBill(bill)).not.toMatch(/\(~\d/);
});
it("measures against a relative root too, instead of silently estimating", async () => {
const relative = path.relative(process.cwd(), TREE_A);
const measured = await measureExactTokens(relative, {
model: "m", apiKey: "sk-test", fetchImpl: fakeFetch(), egressHome,
});
const bill = buildBill(relative, { tokensOf: measured.tokensOf, tokenSource: measured.tokenSource });
for (const s of bill.skills) {
expect(s.totalMdBytes / s.totalMdTokens).toBeLessThan(3.2);
}
expect(measured.missedKeys.size).toBe(0);
});
it("counts anything it could not measure instead of passing it off as measured", async () => {
const measured = await measureExactTokens(TREE_A, {
model: "m", apiKey: "sk-test", fetchImpl: fakeFetch(), egressHome,
});
expect(measured.missedKeys.size).toBe(0);
const ghost = path.join(TREE_A, "alpha", "nope.md");
expect(measured.tokensOf(ghost, 4150)).toBeGreaterThan(0);
expect(measured.missedKeys.has(ghost)).toBe(true);
});
it("grades its own estimate: calibrationTable reports the residual per file", async () => {
const measured = await measureExactTokens(TREE_A, {
model: "m", apiKey: "sk-test", fetchImpl: fakeFetch(), egressHome,
});
const table = calibrationTable(measured.counts, TREE_A);
expect(table.rows.length).toBeGreaterThan(0);
for (const row of table.rows) {
expect(row).toHaveProperty("contentClass");
expect(row.estimatedTokens).toBeGreaterThan(0);
expect(row.tokens).toBeGreaterThan(0);
expect(row.errorPct).toBeCloseTo(((row.estimatedTokens - row.tokens) / row.tokens) * 100, 1);
}
expect(typeof table.worstErrorPct).toBe("number");
expect(typeof table.biasPct).toBe("number");
});
it("refuses to go to the network without a key, and says nothing was sent", async () => {
let called = false;
await expect(
measureExactTokens(TREE_A, {
model: "m",
apiKey: "",
fetchImpl: (() => ((called = true), Promise.reject(new Error("should not run")))) as never,
egressHome,
}),
).rejects.toMatchObject({ code: "exact_missing_api_key" });
expect(called).toBe(false);
});
it("maps HTTP failures to typed codes", async () => {
const status = (code: number) => (async () => ({ ok: false, status: code, text: async () => "nope" })) as never;
for (const [code, expected] of [
[401, "exact_auth_rejected"],
[403, "exact_auth_rejected"],
[500, "exact_request_failed"],
] as const) {
await expect(
measureExactTokens(TREE_A, { model: "m", apiKey: "k", fetchImpl: status(code), egressHome }),
).rejects.toMatchObject({ code: expected });
}
await expect(
measureExactTokens(TREE_A, {
model: "m", apiKey: "k",
fetchImpl: (() => Promise.reject(new Error("offline"))) as never,
egressHome,
}),
).rejects.toBeInstanceOf(ExactModeError);
});
it("offline by default: no --exact means no network call at all", async () => {
const out = capture();
let called = false;
const code = await contextBillMain([TREE_A, "--json"], {
stdout: out.stream,
stderr: out.stream,
fetchImpl: (() => ((called = true), Promise.reject(new Error("no")))) as never,
apiKey: "sk-test",
egressHome,
});
expect(code).toBe(0);
expect(called).toBe(false);
expect(JSON.parse(out.text()).tokenSource).toContain("estimate");
});
it("discloses what --exact sends before sending it", async () => {
const stdout = capture();
const stderr = capture();
const code = await contextBillMain([TREE_A, "--exact", "--json"], {
stdout: stdout.stream,
stderr: stderr.stream,
apiKey: "sk-test",
fetchImpl: fakeFetch(),
egressHome,
});
expect(code).toBe(0);
expect(stderr.text()).toContain("api.anthropic.com");
expect(stderr.text()).toMatch(/sending the content of \d+ \.md file\(s\)/);
const bill = JSON.parse(stdout.text());
expect(bill.tokenSource).toContain("count_tokens");
expect(bill.calibration.rows.length).toBeGreaterThan(0);
});
it("degrades to the estimate when exact mode is unavailable, naming the code", async () => {
const stdout = capture();
const stderr = capture();
const code = await contextBillMain([TREE_A, "--exact", "--json"], {
stdout: stdout.stream,
stderr: stderr.stream,
apiKey: "",
fetchImpl: (() => Promise.reject(new Error("should not run"))) as never,
egressHome,
});
expect(code).toBe(0);
expect(stderr.text()).toContain("exact_missing_api_key");
expect(JSON.parse(stdout.text()).tokenSource).toContain("estimate");
});
it("--help discloses --exact's egress and recalibration", async () => {
const out = capture();
expect(await contextBillMain(["--help"], { stdout: out.stream, stderr: out.stream })).toBe(0);
expect(out.text()).toContain("api.anthropic.com");
expect(out.text()).toContain("egress receipt");
expect(out.text()).toContain("recalibrates");
});
});
describe("ground truth against THIS repo (skill-census parity)", () => {
const census = skillCensus(ROOT);
const dirs = findSkillDirs(ROOT);
const rels = dirs.map((d) => path.relative(ROOT, d) || ".");
it("the walker reaches every physical SKILL.md the census counts", () => {
// physicalSkillFiles is the walker's expectation: every depth-1 skill dir
// (symlinked dirs included) plus the root router.
for (const rel of census.physicalSkillFiles) {
const dirRel = rel === "SKILL.md" ? "." : path.dirname(rel);
expect(rels, `walker missed ${rel}`).toContain(dirRel);
}
});
it("the root router SKILL.md is billed (fix a live on the real repo)", () => {
expect(census.physicalSkillFiles).toContain("SKILL.md"); // guard the guard
expect(rels).toContain(".");
});
it("a real skill's frontmatter bytes match an independent computation", () => {
const qaDir = dirs.find((d) => path.relative(ROOT, d) === "qa")!;
expect(qaDir).toBeTruthy();
const bill = buildBill(qaDir);
expect(bill.skills).toHaveLength(1);
expect(bill.skills[0].frontmatterBytes).toBe(frontmatterBytes(path.join(qaDir, "SKILL.md")));
expect(bill.skills[0].frontmatterBytes).toBeGreaterThan(0);
});
it("skill count is at least the census count (deeper trees may add more, never fewer)", () => {
expect(dirs.length).toBeGreaterThanOrEqual(census.physicalSkillFiles.length);
});
it("no skill dir is billed twice", () => {
expect(new Set(rels).size).toBe(rels.length);
});
});
describe("CLI plumbing", () => {
it("unknown flag and missing tree are usage errors (exit 2)", async () => {
const a = capture();
expect(await contextBillMain(["--nope"], { stdout: a.stream, stderr: a.stream })).toBe(2);
const b = capture();
expect(await contextBillMain(["--diff", TREE_A], { stdout: b.stream, stderr: b.stream })).toBe(2);
const c = capture();
expect(await contextBillMain([path.join(os.tmpdir(), "does-not-exist-xyz")], { stdout: c.stream, stderr: c.stream })).toBe(1);
});
it("bin/gstack-context-bill runs standalone", () => {
const result = Bun.spawnSync([path.join(ROOT, "bin", "gstack-context-bill"), TREE_A]);
expect(result.exitCode).toBe(0);
expect(result.stdout.toString()).toContain("ALWAYS-ON");
expect(result.stdout.toString()).toContain("EAGER");
});
});