mirror of
https://github.com/garrytan/gstack.git
synced 2026-09-09 14:38:59 +02:00
* feat(aside): browser-driver contract, cookbook, research and fallback resolvers
{{ASIDE_SETUP}} (readiness probe + ten rules for driving the user's real browser), {{ASIDE_COOKBOOK}} (script shapes verified live against Aside CLI 1.26: one flow per aside repl script, CDP console hook before navigation, evidence lines, session-directory artifact handoff, GSTACK_STEP_OK sentinel), {{ASIDE_RESEARCH}} (research through aside exec, WebSearch when Aside is absent, knowledge otherwise) and {{BROWSE_FALLBACK}} (the fifteen-row Aside-step to $B-command table plus the rules that differ, so every browsing skill keeps working on gstack's own headless browser). test/aside-driver.test.ts pins the sentences and asserts every browsing skill carries the Aside block followed by the fallback; test/helpers/aside-available.ts is the shared live-Aside probe.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* feat(render): Aside-first local-HTML renderer with the bundled browser as fallback
lib/aside-render.ts serves the HTML's directory on loopback (Aside refuses file:// URLs), opens it with waitUntil load, prints through CDP Page.printToPDF so tagged output, outlines, header/footer templates and page numbers survive, emulates device metrics for sized screenshots, and writes in-page evaluations to files; when Aside is absent it runs the same spec through the browse daemon (newtab, load, js, pdf, screenshot, closetab) and reports ENGINE=aside|browse. bin/gstack-render.ts is the CLI skill templates call. lib/claude-bin.ts and lib/error-handling.ts become the canonical copies (browse/src re-exports them).
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* refactor(browse): /browse drives Aside first, with the $B reference behind the fallback
Contract, cookbook, mode choice (aside repl by default, aside exec for reading), report format, the fallback section, and the full command reference carved on demand.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* refactor(qa): /qa and /qa-only drive Aside, fall back to $B
QA_METHODOLOGY runs every phase as Aside scripts (orient, explore, document, re-test, mobile viewport via CDP emulation, links via HEAD fetch); the authenticate phase is 'you are already signed in'; a 13th rule requires consent before mutating actions on non-local targets; the fallback section translates each step onto $B. The qa E2E tests run on whichever engine is present.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* refactor(design): design-review, design-consultation, design-shotgun, plan-design-review, design-html drive Aside
Design-system extraction is one script printing FONTS/COLORS/HEADINGS/TOUCH_TARGETS/NAV; competitor research confirms the exact URLs before opening them in the real browser and runs on the bundled browser when Aside is absent; design-html's viewport screenshots, sketches and comparison boards render through gstack-render.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* refactor(deploy): benchmark, canary, land-and-deploy Step 7, devex-review drive Aside
One aside repl script per page prints NAV/PAINT/LCP/RESOURCES/SCRIPTS/CSS/SUMMARY (benchmark), CONSOLE_ERRORS/NAV/TEXT + screenshot (canary, re-run every 60s), and the post-deploy check reads responseStatus from the navigation entry; each carries the $B fallback.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* feat(third-party-actions): Aside is the recommended driver; gstack's visible browser stays the fallback
The readiness probe is lifted from {{ASIDE_SETUP}} at gen time (byte-identity pinned) and rule 3 points at browse/SKILL.md for how to drive; the consent question offers Aside first and gstack's own visible browser (handoff/resume for sign-in) as the fallback, as v1.72 framed it.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* refactor(scrape): /scrape reads pages through Aside; the browser-skills runtime rides the fallback
Look-then-extract scripts build the JSON inside the page and print it between JSON_START/JSON_END; aside exec for fuzzy intents; on the $B fallback the browser-skills match/prototype flow and /skillify apply as before.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* refactor(make-pdf): print through Aside first, the bundled browser otherwise
asideClient.ts replaces the direct $B client with one render() call per PDF (the exact option mapping the browse pdf command had: paper, margins, header/footer/page numbers, tagged, outline, printBackground, preferCSSPageSize, Paged.js wait); the diagram pre-pass, oversized-image downscale and DOCX rasters each run as one render script with per-fence try/catch; exit 4 now means no browser is available and names both remedies; $P setup reports which engine it found. The e2e gates run on whichever engine is present, so the Linux lane exercises the fallback.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* refactor(diagram): the triplet is one gstack-render call
SVG, PNG and excalidraw from one invocation over the content-addressed bundle staged under /tmp/gstack-render; every diagram type gets an excalidraw export; gstack-render picks the engine and prints ENGINE=; the diagram E2E gates on either engine.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* feat(research): web research runs in Aside first, WebSearch second
The planning, review, design, security and investigate skills research through {{ASIDE_RESEARCH}}; WebSearch stays in allowed-tools as the fallback; testing.ts's bootstrap step follows; skeleton ceilings ratcheted for the research block.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* feat(setup,gen-skill-docs): prune renders of skills that no longer exist
setup gains _prune_stale_generated for every host tree and the doc generator removes gstack-* output dirs it did not write, so a skill removed from the source tree can never linger in an install.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* test: registries, budgets and suite reconciled for Aside-first with the $B fallback
Touchfiles + E2E tiers gain the Aside keys, coverage matrix and eval baselines updated, size budget re-baselined to parity-baseline-v1.80.0.0.json (the contract plus fallback ride in every browsing skill), parity ceilings ratcheted with measured values, LLM-judge prompts and the E2E fixtures speak Aside-first, browse-fallback.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* docs: Aside first, gstack browser fallback
README, BROWSER.md, docs/, CONTRIBUTING, CLAUDE.md, ARCHITECTURE, AGENTS.md, TODOS and the root router describe the one product story: Aside is the browser gstack drives first; the bundled headless browser is the automatic fallback (Linux, Windows, app closed) where cookie import, GStack Browser, pair-agent and browser-skills still apply.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* chore: regenerate SKILL.md docs, llms.txt, agents digest, ship goldens, context-budget fixture
bun run gen:skill-docs over the templates; goldens re-rendered; context-budget ceilings recaptured.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* v1.80.0.0: Aside is the browser gstack drives first; the bundled browser is the fallback
MINOR: new capability across ten skills, the renderer and research; nothing removed. CHANGELOG release summary + itemized changes; VERSION 1.80.0.0; package.json 1.80.0.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* docs(todos): file non-Claude host ownership-gate and version-heading pin follow-ups
Two follow-ups from the /plan-ceo-review + /plan-eng-review pass on merging
PR #2804 with main's v1.80.0.0 ownership gate: bring the Codex/Factory/
OpenCode/Cursor/Kiro copy loops and the stale-render prune under the
.gstack-owned marker rule, and a free test pinning that the CHANGELOG top
heading equals VERSION (the collision that git cannot see).
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* fix: pre-landing review fixes for the Aside-first branch
Review army + adversarial passes (Claude and Codex) on the merged branch:
setup
- _prune_stale_generated scans the host dirs too (the generator already
removed the render before setup ran, so the host branch was dead), skips
symlinks in the render tree (rm -rf on a slash-terminated link empties its
target), removes a host symlink only when it resolves into gstack, cleans a
bannered real dir through _cleanup_weak_dir, recognizes frontmatter-renamed
skills, and logs through log. The always-run codex render passes every host
dir that may link to it.
- NEEDS_BUILD checks all three binaries (with $_EXE) and lib/ sources; the
browser hint and the bootstrap summary honor GSTACK_SKIP_ASIDE, treat a
requested skip as a request, and derive one skill list.
lib/aside-render.ts + bin/gstack-render.ts
- The loopback server carries a per-render secret path, checks containment on
the real path (symlink escapes are 403), and rejects malformed encoding.
- Inline eval results are one base64 line, so page text cannot forge
ASIDE_DIR= or the sentinel; the last ASIDE_DIR wins.
- runProc escalates SIGTERM to SIGKILL, bounds every wait, and clears every
timer (an uncleared one kept gstack-render alive after printing OK).
- renderTmpDir refuses a shared /tmp name owned by someone else; the work dir
and server are created inside try; goto's budget follows the render budget.
- probeAside classifies a present-but-failing CLI as ASIDE_NOT_RUNNING like
the skills' bash probe; render() retries on gstack's own browser when Aside
could not start or its private CDP bridge is gone (never on a page error
or a timeout of a running script); the CLI reports the engine that actually
rendered, exits 0 on --help, rejects non-numeric flags, documents
--wait-timeout, fences EVAL/PAGE_ERRORS as untrusted content, and names the
daemon's cookie-import JS lock remedy.
- The browse path passes --scale only when asked (a scale change rebuilds
the daemon context) and restores the viewport after a sized screenshot.
resolvers / templates
- The bash probe honors GSTACK_SKIP_ASIDE and has a perl deadline on stock
macOS; .local is no longer LOCAL (mDNS); same-origin filters compare parsed
origins; link status is HEAD-checked only on LOCAL targets; every
aside exec goes through the receipted _aside_exec prelude
({{ASIDE_EXEC_PRELUDE}}), including nine template blocks that called it
bare; the design sketch and diagram staging use private directories.
- The generator prunes only bannered renders and never a host whose
generation failed.
Docs, stale comments and dead code cleaned; goldens re-rendered; tests
updated and added for every behavior above.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* test: coverage for the render CLI, setup rebuild check, make-pdf exit codes, and prose $B spans
New free tests from the ship coverage audit: test/gstack-render-cli.test.ts
(argv guards, --help, output contract with a fake daemon, failure and
serve-root paths, no-browser case, prompt exit), test/setup-needs-build.test.ts
(every binary and source set flips NEEDS_BUILD, Windows suffixes),
make-pdf/test/cli-exit-codes.test.ts and setup-smoke.test.ts (error to exit
code mapping, runSetup stages, renderPdf's engine), and prose-span cases for
extractBrowseCommands in test/skill-parser.test.ts.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* docs: CHANGELOG and TODOS cover the review fixes (v1.81.0.0)
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* docs: sync project docs with the v1.81.0.0 review fixes
BROWSER.md, ARCHITECTURE.md, CONTRIBUTING.md, README.md, CLAUDE.md,
docs/TESTING_INTERNALS.md and docs/PROJECT_STRUCTURE.md now describe the
shipped renderer and setup: the loopback render server's per-render secret
path and real-path containment, ENGINE= naming the engine that actually
rendered (mid-run retry on gstack's own browser), EVAL/PAGE_ERRORS fenced as
untrusted content, --wait-timeout and the CLI's argv guards, the receipted
_aside_exec prelude ({{ASIDE_EXEC_PRELUDE}} in the placeholder table), the
LOCAL host rule without .local, LOCAL-only HEAD checks in the links script,
GSTACK_SKIP_ASIDE across probe/renderer/setup, the ownership-gated
retired-skill prune, the widened NEEDS_BUILD check, and the new free tests
(gstack-render-cli, setup-prune-stale-generated, setup-browser-hint,
setup-needs-build, make-pdf cli-exit-codes and setup-smoke).
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* docs: CHANGELOG states the precise mid-run retry rule
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* fix(test): skill-e2e-bws slices the $B setup block from the Browser fallback section
browse/SKILL.md no longer has '## SETUP' / '## Core QA Patterns' (Aside is the
primary driver; the $B block moved under 'Browser fallback'), so the gate test
sliced an empty block and handed the agent nothing to run. Anchor on
'### Find the `$B` binary' up to the next heading. 7/7 pass.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* fix(test): gate POSIX-only fixtures off Windows
windows-free-tests: the gstack-render CLI tests drive a shebang fake browse
that CreateProcess cannot exec, and two NEEDS_BUILD cases assert an execute
bit and a bare-name miss that MSYS bash does not have (test -x ignores mode
bits and resolves design -> design.exe). Those describes and cases now
self-skip on win32; argument guards, --help, the no-browser case, and every
other rebuild-check case still run there.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* fix(render): runProc waits for the exit code until the kill deadline; newtab retries once on a cold daemon
A process whose pipes have reached EOF is exiting, but runProc gave the exit
code only five seconds to arrive and then returned null, which run() reports
as a failed command. Under CI's six-shard load one such render failed with the
artifact already written. The SIGTERM/SIGKILL timers already bound the wait,
so the exit race now runs to the kill deadline.
The first CLI call auto-starts the browse daemon; on a cold start it can
answer 'Unable to connect' once while the server is still coming up. That
single case is retried after 1.5s; every other newtab failure is not.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* test(aside-render): warm the daemon before live fallback cases; failures name the render error
- Live fallback cases run 'goto about:blank' up to twice before asserting and
skip (never fail) when the daemon cannot come up.
- expectOk() puts r.error and the browse transcript into the assertion so a
failed render is diagnosable from the CI log.
- The argv-contract cases dump the fake's log on a miss.
- File default timeout is 30s: the subject is the CLI contract, not latency.
- Two cases pin the cold-daemon newtab retry and that other errors are not
retried.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* docs: CHANGELOG notes the cold-start tolerance of the bundled-browser renderer
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
---------
Co-authored-by: Sina <sdroid674+github@gmail.com>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
290 lines
14 KiB
TypeScript
290 lines
14 KiB
TypeScript
/**
|
||
* Per-skill SKILL.md size budget regression (v1.46.0.0 T5).
|
||
*
|
||
* Asserts that no skill's generated SKILL.md grew beyond the v1.47.0.0
|
||
* baseline. Catches preamble/resolver changes that bloat skills back to
|
||
* the pre-compression size. Free — pure file IO + JSON diff.
|
||
*
|
||
* Baseline rebased v1.44.1 → v1.47.0.0 in the AskUserQuestion split-rule
|
||
* PR after main merged GSTACK_PLAN_MODE + /spec, pushing the v1.44.1
|
||
* anchor past the 5% ratchet. Historical v1.44.1.json and v1.46.0.0.json
|
||
* are retained in test/fixtures/ for reference.
|
||
*
|
||
* Why a separate test from skill-budget-regression.test.ts: that one
|
||
* compares LIVE eval runs (tool calls, turns, cost); this one compares
|
||
* static SKILL.md sizes. Both gate-tier.
|
||
*
|
||
* Baseline rebased v1.69.1.0 → v1.81.0.0: the Aside-first browser contract
|
||
* ({{ASIDE_SETUP}}) plus the gstack-browser fallback block now ride in every
|
||
* browsing skill (~9KB), which pushed benchmark and scrape past 1.5× of the
|
||
* v1.69.1.0 anchor. Deliberate, corpus-wide, receipted in the v1.81.0.0
|
||
* CHANGELOG; the v1.69.1.0 fixture stays on disk for history.
|
||
*
|
||
* The previous baseline lived at test/fixtures/parity-baseline-v1.69.1.0.json,
|
||
* re-captured 2026-08-25 during token-reduction Phase 1 (bash consolidation
|
||
* moved ~11-13KB of inline preamble bash per skill into bin/gstack-skill-start
|
||
* and bin/gstack-skill-end — a deliberate corpus-wide shrink; receipt:
|
||
* gstack-context-bill --diff in PR #2691). The prior v1.47.0.0 fixture stays
|
||
* on disk for history. Live pins at capture time: this test (shrink floor)
|
||
* and test/parity-suite.test.ts vs parity-baseline-v1.64.1.0.json (growth).
|
||
*
|
||
* Override:
|
||
* - GSTACK_SIZE_BUDGET_RATIO=<n> changes the per-skill regression ratio.
|
||
* Default 1.0 (no growth allowed). Set to 1.10 to permit 10% growth
|
||
* (e.g., during deliberate feature additions that the catalog trim
|
||
* doesn't offset).
|
||
* - GSTACK_SIZE_BUDGET_OVERRIDE_REASON="text" allows a regression to
|
||
* pass and logs the reason to ~/.gstack/analytics/spend-overrides.jsonl
|
||
* for audit. Use sparingly; the next baseline should bake in the new
|
||
* size.
|
||
*/
|
||
|
||
import { describe, test, expect } from 'bun:test';
|
||
import * as fs from 'fs';
|
||
import * as path from 'path';
|
||
import { execSync } from 'child_process';
|
||
import { captureBaseline, extractDescription, type ParityBaseline } from './helpers/capture-parity-baseline';
|
||
import { logBudgetOverride } from './helpers/budget-override';
|
||
import { CARVED_SKILLS } from './helpers/carve-guards';
|
||
|
||
const REPO_ROOT = path.resolve(import.meta.dir, '..');
|
||
const BASELINE_PATH = path.join(REPO_ROOT, 'test', 'fixtures', 'parity-baseline-v1.81.0.0.json');
|
||
|
||
// Default per-skill ratio is 1.50 (50% growth tolerance). Adjusted v1.52.0.0
|
||
// (cathedral cap audit) from 1.05 → 1.50: a 5% ratio tripped on legitimate
|
||
// feature additions (e.g., plan-tune cathedral T13 grew SKILL.md ×1.24
|
||
// adding load-bearing Dream cycle + Audit unmarked + Recent auto-decisions
|
||
// surfaces). Real bloat is 2-3×; this catches that while not tripping on
|
||
// normal feature scope. The always-loaded catalog cost is enforced
|
||
// separately with a hard ceiling.
|
||
const DEFAULT_RATIO = 1.50;
|
||
const RATIO = Number(process.env.GSTACK_SIZE_BUDGET_RATIO) || DEFAULT_RATIO;
|
||
|
||
interface Regression {
|
||
skill: string;
|
||
beforeBytes: number;
|
||
afterBytes: number;
|
||
growth: number;
|
||
}
|
||
|
||
describe('SKILL.md size budget regression (gate, free)', () => {
|
||
test('parity-baseline-v1.81.0.0.json exists', () => {
|
||
expect(fs.existsSync(BASELINE_PATH)).toBe(true);
|
||
});
|
||
|
||
test('no skill exceeds v1.81.0.0 baseline size × ratio', () => {
|
||
const baseline: ParityBaseline = JSON.parse(fs.readFileSync(BASELINE_PATH, 'utf-8'));
|
||
const current = captureBaseline({ repoRoot: REPO_ROOT });
|
||
|
||
const regressions: Regression[] = [];
|
||
for (const [skill, before] of Object.entries(baseline.skills)) {
|
||
const after = current.skills[skill];
|
||
if (!after) continue; // skill removed since v1.44 — not a regression
|
||
if (after.skillMdBytes <= before.skillMdBytes * RATIO) continue;
|
||
regressions.push({
|
||
skill,
|
||
beforeBytes: before.skillMdBytes,
|
||
afterBytes: after.skillMdBytes,
|
||
growth: after.skillMdBytes / before.skillMdBytes,
|
||
});
|
||
}
|
||
|
||
if (regressions.length === 0) return;
|
||
|
||
const overrideReason = process.env.GSTACK_SIZE_BUDGET_OVERRIDE_REASON?.trim();
|
||
if (overrideReason) {
|
||
logBudgetOverride({
|
||
scope: 'skill-size-budget',
|
||
reason: overrideReason,
|
||
details: { ratio: RATIO, regressions },
|
||
});
|
||
// eslint-disable-next-line no-console
|
||
console.warn(
|
||
`[skill-size-budget] OVERRIDE APPLIED (${overrideReason}) — ${regressions.length} regression(s) allowed:`,
|
||
);
|
||
for (const r of regressions) {
|
||
// eslint-disable-next-line no-console
|
||
console.warn(` ${r.skill}: ${r.beforeBytes} → ${r.afterBytes} bytes (×${r.growth.toFixed(2)})`);
|
||
}
|
||
return;
|
||
}
|
||
|
||
const msg = regressions.map(r =>
|
||
` ${r.skill}: ${r.beforeBytes} → ${r.afterBytes} bytes (×${r.growth.toFixed(2)})`,
|
||
).join('\n');
|
||
throw new Error(
|
||
`${regressions.length} skill(s) regressed past v1.47.0.0 baseline × ${RATIO}:\n${msg}\n` +
|
||
`Override: set GSTACK_SIZE_BUDGET_OVERRIDE_REASON="why this is OK" to allow and audit-log.`,
|
||
);
|
||
});
|
||
|
||
test('total corpus byte count does not regress past baseline × ratio', () => {
|
||
const baseline: ParityBaseline = JSON.parse(fs.readFileSync(BASELINE_PATH, 'utf-8'));
|
||
const current = captureBaseline({ repoRoot: REPO_ROOT });
|
||
const ratio = current.totalCorpusBytes / baseline.totalCorpusBytes;
|
||
if (current.totalCorpusBytes <= baseline.totalCorpusBytes * RATIO) {
|
||
// eslint-disable-next-line no-console
|
||
console.log(
|
||
`[skill-size-budget] corpus OK: ${baseline.totalCorpusBytes} → ${current.totalCorpusBytes} bytes (×${ratio.toFixed(3)})`,
|
||
);
|
||
return;
|
||
}
|
||
const overrideReason = process.env.GSTACK_SIZE_BUDGET_OVERRIDE_REASON?.trim();
|
||
if (overrideReason) {
|
||
logBudgetOverride({
|
||
scope: 'skill-size-budget-corpus',
|
||
reason: overrideReason,
|
||
details: { ratio: RATIO, observed: ratio, before: baseline.totalCorpusBytes, after: current.totalCorpusBytes },
|
||
});
|
||
return;
|
||
}
|
||
throw new Error(
|
||
`Total corpus regressed past v1.47.0.0 baseline × ${RATIO}: ` +
|
||
`${baseline.totalCorpusBytes} → ${current.totalCorpusBytes} bytes (×${ratio.toFixed(3)}). ` +
|
||
`Override: set GSTACK_SIZE_BUDGET_OVERRIDE_REASON to allow.`,
|
||
);
|
||
});
|
||
|
||
/**
|
||
* Gap E (v1.46.0.0): per-skill min-size floor.
|
||
*
|
||
* The existing skill-coverage-floor enforces body ≥ 200 bytes, which is
|
||
* a tiny noise floor. A skill that was 100 KB at v1.47.0.0 and shrinks to
|
||
* 250 bytes passes that check despite losing 99.75% of content. The
|
||
* parity-suite content invariants cover this for 10 hand-picked skills
|
||
* (cso, ship, plan-ceo, etc.); the remaining 41 skills had no per-skill
|
||
* shrinkage floor.
|
||
*
|
||
* Floor: 80% of the v1.47.0.0 baseline. v1.46 actual shrinkage is <1% per
|
||
* skill, so this is a comfortable ceiling that still catches accidental
|
||
* mass deletion (e.g., a refactor that strips the body of a skill).
|
||
*
|
||
* v2.0.0.0 introduces the sections/ pattern for 5 heavyweights
|
||
* (ship, plan-ceo-review, office-hours, plan-eng-review,
|
||
* plan-design-review). Carved so far: ship (skeleton ~83 KB) and
|
||
* plan-ceo-review (skeleton ~81 KB, down from the 138 KB monolith). Those
|
||
* skeletons legitimately fall below the 80% body-strip floor, so each carved
|
||
* skill is added to SECTIONS_EXTRACTED; its union is guarded instead by the
|
||
* sectioned invariant in parity-harness.ts (minBytes on skeleton+sections).
|
||
* Add the remaining three here as they carve.
|
||
*/
|
||
test('no skill shrinks past 80% of v1.69.1.0 baseline (catches accidental body strip)', () => {
|
||
const baseline: ParityBaseline = JSON.parse(fs.readFileSync(BASELINE_PATH, 'utf-8'));
|
||
const current = captureBaseline({ repoRoot: REPO_ROOT });
|
||
const MIN_RATIO = 0.80; // a skill at <80% of its v1.44 size signals mass-deletion
|
||
// Carved skills (v2 plan T9): the skeleton SKILL.md intentionally shrinks
|
||
// because prose moved into sections/*.md. The union size is guarded instead
|
||
// by the sectioned invariant in parity-harness.ts (minBytes on the
|
||
// skeleton+sections union), so exempt the skeleton from the body-strip floor.
|
||
// EQ1: derived from the canonical CARVE_GUARDS registry — no parallel list.
|
||
const SECTIONS_EXTRACTED = new Set<string>(CARVED_SKILLS);
|
||
// Intentional one-off shrinks vs the frozen baseline (each needs a reason):
|
||
// - spec: the baseline measured a template bug — prose at Phase 5 mentioned
|
||
// {{PREAMBLE}} literally, so the generator expanded the ENTIRE preamble a
|
||
// second time mid-sentence (~47 KB of duplication). Fixed by rewording the
|
||
// prose; spec/SKILL.md now carries exactly one preamble (~80.9 KB, ×0.79).
|
||
// - scrape/diagram/open-gstack-browser/landing-report/pair-agent/skillify:
|
||
// the baseline measured these at the silent tier-4 default (a missing
|
||
// preamble-tier frontmatter fell through `?? 4`). Their tiers are now
|
||
// declared correctly (1-2), shedding the tier-2..4 onboarding prose they
|
||
// never should have carried (-271 lines each for tier 1).
|
||
// - browse: the baseline measured the headless-browse skill with its ~17 KB
|
||
// $B command reference + snapshot-flag tables inline. /browse drives the
|
||
// Aside browser first and carries the Aside contract in the skeleton; the
|
||
// command tables live in the carved browse/sections/command-list.md.
|
||
const INTENTIONAL_SHRINKS = new Set<string>([
|
||
'spec',
|
||
'scrape', 'diagram', 'open-gstack-browser',
|
||
'landing-report', 'pair-agent', 'skillify',
|
||
'browse',
|
||
]);
|
||
|
||
const undershoots: Array<{
|
||
skill: string; beforeBytes: number; afterBytes: number; ratio: number;
|
||
}> = [];
|
||
for (const [skill, before] of Object.entries(baseline.skills)) {
|
||
if (SECTIONS_EXTRACTED.has(skill)) continue;
|
||
if (INTENTIONAL_SHRINKS.has(skill)) continue;
|
||
const after = current.skills[skill];
|
||
if (!after) continue; // skill removed since baseline — separate concern
|
||
const ratio = after.skillMdBytes / before.skillMdBytes;
|
||
if (ratio < MIN_RATIO) {
|
||
undershoots.push({
|
||
skill, beforeBytes: before.skillMdBytes, afterBytes: after.skillMdBytes, ratio,
|
||
});
|
||
}
|
||
}
|
||
|
||
if (undershoots.length === 0) return;
|
||
|
||
const overrideReason = process.env.GSTACK_SIZE_BUDGET_OVERRIDE_REASON?.trim();
|
||
if (overrideReason) {
|
||
logBudgetOverride({
|
||
scope: 'skill-size-budget-floor',
|
||
reason: overrideReason,
|
||
details: { min_ratio: MIN_RATIO, undershoots },
|
||
});
|
||
// eslint-disable-next-line no-console
|
||
console.warn(
|
||
`[skill-size-budget-floor] OVERRIDE APPLIED (${overrideReason}) — ${undershoots.length} undershoot(s) allowed`,
|
||
);
|
||
return;
|
||
}
|
||
|
||
const msg = undershoots.map(u =>
|
||
` ${u.skill}: ${u.beforeBytes} → ${u.afterBytes} bytes (×${u.ratio.toFixed(2)} — below ${MIN_RATIO} floor)`,
|
||
).join('\n');
|
||
throw new Error(
|
||
`${undershoots.length} skill(s) shrunk past v1.47.0.0 × ${MIN_RATIO} floor:\n${msg}\n` +
|
||
`This usually signals accidental body strip (e.g., a resolver returning empty, a ` +
|
||
`template losing a section). If the shrinkage is intentional (e.g., the skill moved ` +
|
||
`to the sections/ pattern), add it to SECTIONS_EXTRACTED in this test. Override: ` +
|
||
`GSTACK_SIZE_BUDGET_OVERRIDE_REASON="why" allows + audit-logs.`,
|
||
);
|
||
});
|
||
|
||
test('catalog token estimate stays compressed (v1.45 target ≤ 7000)', () => {
|
||
// Measure COMMITTED content (git show HEAD:), not the live tree. Under
|
||
// the parallel free-suite runner, sibling workers regenerate real
|
||
// SKILL.md files mid-run (gen-skill-docs regen tests), so the live-tree
|
||
// estimate was a moving target: 4177 solo, 8356 and 8041 in two parallel
|
||
// runs. A repo-budget ratchet measures the catalog that ships; CI always
|
||
// checks the PR's committed tree anyway.
|
||
const trackedPaths = execSync('git ls-files -- "*/SKILL.md"', { cwd: REPO_ROOT, encoding: 'utf-8', timeout: 30_000 })
|
||
.split('\n')
|
||
.filter(Boolean)
|
||
.filter((p) => p.split('/').length === 2);
|
||
let descriptionBytes = 0;
|
||
for (const rel of trackedPaths) {
|
||
const committed = execSync(`git show HEAD:${JSON.stringify(rel)}`, {
|
||
cwd: REPO_ROOT,
|
||
encoding: 'utf-8',
|
||
maxBuffer: 8 * 1024 * 1024,
|
||
timeout: 30_000,
|
||
});
|
||
descriptionBytes += Buffer.byteLength(extractDescription(committed), 'utf-8');
|
||
}
|
||
const catalogTokens = Math.round(descriptionBytes / 4);
|
||
const trackedCount = trackedPaths.length;
|
||
const v145Target = 7000;
|
||
if (catalogTokens <= v145Target) {
|
||
// eslint-disable-next-line no-console
|
||
console.log(`[skill-size-budget] catalog OK: ~${catalogTokens} tokens (target ≤${v145Target}, ${trackedCount} tracked skills)`);
|
||
return;
|
||
}
|
||
const overrideReason = process.env.GSTACK_SIZE_BUDGET_OVERRIDE_REASON?.trim();
|
||
if (overrideReason) {
|
||
logBudgetOverride({
|
||
scope: 'skill-size-budget-catalog',
|
||
reason: overrideReason,
|
||
details: { target: v145Target, observed: catalogTokens },
|
||
});
|
||
return;
|
||
}
|
||
throw new Error(
|
||
`Catalog token estimate regressed past v1.45 target: ${catalogTokens} tokens > ${v145Target}. ` +
|
||
`T4 catalog trim should keep this under control. Override: set GSTACK_SIZE_BUDGET_OVERRIDE_REASON to allow.`,
|
||
);
|
||
});
|
||
});
|