mirror of
https://github.com/garrytan/gstack.git
synced 2026-09-09 14:38:59 +02:00
* feat(aside): browser-driver contract, cookbook, research and fallback resolvers
{{ASIDE_SETUP}} (readiness probe + ten rules for driving the user's real browser), {{ASIDE_COOKBOOK}} (script shapes verified live against Aside CLI 1.26: one flow per aside repl script, CDP console hook before navigation, evidence lines, session-directory artifact handoff, GSTACK_STEP_OK sentinel), {{ASIDE_RESEARCH}} (research through aside exec, WebSearch when Aside is absent, knowledge otherwise) and {{BROWSE_FALLBACK}} (the fifteen-row Aside-step to $B-command table plus the rules that differ, so every browsing skill keeps working on gstack's own headless browser). test/aside-driver.test.ts pins the sentences and asserts every browsing skill carries the Aside block followed by the fallback; test/helpers/aside-available.ts is the shared live-Aside probe.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* feat(render): Aside-first local-HTML renderer with the bundled browser as fallback
lib/aside-render.ts serves the HTML's directory on loopback (Aside refuses file:// URLs), opens it with waitUntil load, prints through CDP Page.printToPDF so tagged output, outlines, header/footer templates and page numbers survive, emulates device metrics for sized screenshots, and writes in-page evaluations to files; when Aside is absent it runs the same spec through the browse daemon (newtab, load, js, pdf, screenshot, closetab) and reports ENGINE=aside|browse. bin/gstack-render.ts is the CLI skill templates call. lib/claude-bin.ts and lib/error-handling.ts become the canonical copies (browse/src re-exports them).
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* refactor(browse): /browse drives Aside first, with the $B reference behind the fallback
Contract, cookbook, mode choice (aside repl by default, aside exec for reading), report format, the fallback section, and the full command reference carved on demand.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* refactor(qa): /qa and /qa-only drive Aside, fall back to $B
QA_METHODOLOGY runs every phase as Aside scripts (orient, explore, document, re-test, mobile viewport via CDP emulation, links via HEAD fetch); the authenticate phase is 'you are already signed in'; a 13th rule requires consent before mutating actions on non-local targets; the fallback section translates each step onto $B. The qa E2E tests run on whichever engine is present.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* refactor(design): design-review, design-consultation, design-shotgun, plan-design-review, design-html drive Aside
Design-system extraction is one script printing FONTS/COLORS/HEADINGS/TOUCH_TARGETS/NAV; competitor research confirms the exact URLs before opening them in the real browser and runs on the bundled browser when Aside is absent; design-html's viewport screenshots, sketches and comparison boards render through gstack-render.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* refactor(deploy): benchmark, canary, land-and-deploy Step 7, devex-review drive Aside
One aside repl script per page prints NAV/PAINT/LCP/RESOURCES/SCRIPTS/CSS/SUMMARY (benchmark), CONSOLE_ERRORS/NAV/TEXT + screenshot (canary, re-run every 60s), and the post-deploy check reads responseStatus from the navigation entry; each carries the $B fallback.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* feat(third-party-actions): Aside is the recommended driver; gstack's visible browser stays the fallback
The readiness probe is lifted from {{ASIDE_SETUP}} at gen time (byte-identity pinned) and rule 3 points at browse/SKILL.md for how to drive; the consent question offers Aside first and gstack's own visible browser (handoff/resume for sign-in) as the fallback, as v1.72 framed it.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* refactor(scrape): /scrape reads pages through Aside; the browser-skills runtime rides the fallback
Look-then-extract scripts build the JSON inside the page and print it between JSON_START/JSON_END; aside exec for fuzzy intents; on the $B fallback the browser-skills match/prototype flow and /skillify apply as before.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* refactor(make-pdf): print through Aside first, the bundled browser otherwise
asideClient.ts replaces the direct $B client with one render() call per PDF (the exact option mapping the browse pdf command had: paper, margins, header/footer/page numbers, tagged, outline, printBackground, preferCSSPageSize, Paged.js wait); the diagram pre-pass, oversized-image downscale and DOCX rasters each run as one render script with per-fence try/catch; exit 4 now means no browser is available and names both remedies; $P setup reports which engine it found. The e2e gates run on whichever engine is present, so the Linux lane exercises the fallback.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* refactor(diagram): the triplet is one gstack-render call
SVG, PNG and excalidraw from one invocation over the content-addressed bundle staged under /tmp/gstack-render; every diagram type gets an excalidraw export; gstack-render picks the engine and prints ENGINE=; the diagram E2E gates on either engine.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* feat(research): web research runs in Aside first, WebSearch second
The planning, review, design, security and investigate skills research through {{ASIDE_RESEARCH}}; WebSearch stays in allowed-tools as the fallback; testing.ts's bootstrap step follows; skeleton ceilings ratcheted for the research block.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* feat(setup,gen-skill-docs): prune renders of skills that no longer exist
setup gains _prune_stale_generated for every host tree and the doc generator removes gstack-* output dirs it did not write, so a skill removed from the source tree can never linger in an install.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* test: registries, budgets and suite reconciled for Aside-first with the $B fallback
Touchfiles + E2E tiers gain the Aside keys, coverage matrix and eval baselines updated, size budget re-baselined to parity-baseline-v1.80.0.0.json (the contract plus fallback ride in every browsing skill), parity ceilings ratcheted with measured values, LLM-judge prompts and the E2E fixtures speak Aside-first, browse-fallback.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* docs: Aside first, gstack browser fallback
README, BROWSER.md, docs/, CONTRIBUTING, CLAUDE.md, ARCHITECTURE, AGENTS.md, TODOS and the root router describe the one product story: Aside is the browser gstack drives first; the bundled headless browser is the automatic fallback (Linux, Windows, app closed) where cookie import, GStack Browser, pair-agent and browser-skills still apply.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* chore: regenerate SKILL.md docs, llms.txt, agents digest, ship goldens, context-budget fixture
bun run gen:skill-docs over the templates; goldens re-rendered; context-budget ceilings recaptured.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* v1.80.0.0: Aside is the browser gstack drives first; the bundled browser is the fallback
MINOR: new capability across ten skills, the renderer and research; nothing removed. CHANGELOG release summary + itemized changes; VERSION 1.80.0.0; package.json 1.80.0.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* docs(todos): file non-Claude host ownership-gate and version-heading pin follow-ups
Two follow-ups from the /plan-ceo-review + /plan-eng-review pass on merging
PR #2804 with main's v1.80.0.0 ownership gate: bring the Codex/Factory/
OpenCode/Cursor/Kiro copy loops and the stale-render prune under the
.gstack-owned marker rule, and a free test pinning that the CHANGELOG top
heading equals VERSION (the collision that git cannot see).
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* fix: pre-landing review fixes for the Aside-first branch
Review army + adversarial passes (Claude and Codex) on the merged branch:
setup
- _prune_stale_generated scans the host dirs too (the generator already
removed the render before setup ran, so the host branch was dead), skips
symlinks in the render tree (rm -rf on a slash-terminated link empties its
target), removes a host symlink only when it resolves into gstack, cleans a
bannered real dir through _cleanup_weak_dir, recognizes frontmatter-renamed
skills, and logs through log. The always-run codex render passes every host
dir that may link to it.
- NEEDS_BUILD checks all three binaries (with $_EXE) and lib/ sources; the
browser hint and the bootstrap summary honor GSTACK_SKIP_ASIDE, treat a
requested skip as a request, and derive one skill list.
lib/aside-render.ts + bin/gstack-render.ts
- The loopback server carries a per-render secret path, checks containment on
the real path (symlink escapes are 403), and rejects malformed encoding.
- Inline eval results are one base64 line, so page text cannot forge
ASIDE_DIR= or the sentinel; the last ASIDE_DIR wins.
- runProc escalates SIGTERM to SIGKILL, bounds every wait, and clears every
timer (an uncleared one kept gstack-render alive after printing OK).
- renderTmpDir refuses a shared /tmp name owned by someone else; the work dir
and server are created inside try; goto's budget follows the render budget.
- probeAside classifies a present-but-failing CLI as ASIDE_NOT_RUNNING like
the skills' bash probe; render() retries on gstack's own browser when Aside
could not start or its private CDP bridge is gone (never on a page error
or a timeout of a running script); the CLI reports the engine that actually
rendered, exits 0 on --help, rejects non-numeric flags, documents
--wait-timeout, fences EVAL/PAGE_ERRORS as untrusted content, and names the
daemon's cookie-import JS lock remedy.
- The browse path passes --scale only when asked (a scale change rebuilds
the daemon context) and restores the viewport after a sized screenshot.
resolvers / templates
- The bash probe honors GSTACK_SKIP_ASIDE and has a perl deadline on stock
macOS; .local is no longer LOCAL (mDNS); same-origin filters compare parsed
origins; link status is HEAD-checked only on LOCAL targets; every
aside exec goes through the receipted _aside_exec prelude
({{ASIDE_EXEC_PRELUDE}}), including nine template blocks that called it
bare; the design sketch and diagram staging use private directories.
- The generator prunes only bannered renders and never a host whose
generation failed.
Docs, stale comments and dead code cleaned; goldens re-rendered; tests
updated and added for every behavior above.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* test: coverage for the render CLI, setup rebuild check, make-pdf exit codes, and prose $B spans
New free tests from the ship coverage audit: test/gstack-render-cli.test.ts
(argv guards, --help, output contract with a fake daemon, failure and
serve-root paths, no-browser case, prompt exit), test/setup-needs-build.test.ts
(every binary and source set flips NEEDS_BUILD, Windows suffixes),
make-pdf/test/cli-exit-codes.test.ts and setup-smoke.test.ts (error to exit
code mapping, runSetup stages, renderPdf's engine), and prose-span cases for
extractBrowseCommands in test/skill-parser.test.ts.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* docs: CHANGELOG and TODOS cover the review fixes (v1.81.0.0)
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* docs: sync project docs with the v1.81.0.0 review fixes
BROWSER.md, ARCHITECTURE.md, CONTRIBUTING.md, README.md, CLAUDE.md,
docs/TESTING_INTERNALS.md and docs/PROJECT_STRUCTURE.md now describe the
shipped renderer and setup: the loopback render server's per-render secret
path and real-path containment, ENGINE= naming the engine that actually
rendered (mid-run retry on gstack's own browser), EVAL/PAGE_ERRORS fenced as
untrusted content, --wait-timeout and the CLI's argv guards, the receipted
_aside_exec prelude ({{ASIDE_EXEC_PRELUDE}} in the placeholder table), the
LOCAL host rule without .local, LOCAL-only HEAD checks in the links script,
GSTACK_SKIP_ASIDE across probe/renderer/setup, the ownership-gated
retired-skill prune, the widened NEEDS_BUILD check, and the new free tests
(gstack-render-cli, setup-prune-stale-generated, setup-browser-hint,
setup-needs-build, make-pdf cli-exit-codes and setup-smoke).
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* docs: CHANGELOG states the precise mid-run retry rule
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* fix(test): skill-e2e-bws slices the $B setup block from the Browser fallback section
browse/SKILL.md no longer has '## SETUP' / '## Core QA Patterns' (Aside is the
primary driver; the $B block moved under 'Browser fallback'), so the gate test
sliced an empty block and handed the agent nothing to run. Anchor on
'### Find the `$B` binary' up to the next heading. 7/7 pass.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* fix(test): gate POSIX-only fixtures off Windows
windows-free-tests: the gstack-render CLI tests drive a shebang fake browse
that CreateProcess cannot exec, and two NEEDS_BUILD cases assert an execute
bit and a bare-name miss that MSYS bash does not have (test -x ignores mode
bits and resolves design -> design.exe). Those describes and cases now
self-skip on win32; argument guards, --help, the no-browser case, and every
other rebuild-check case still run there.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* fix(render): runProc waits for the exit code until the kill deadline; newtab retries once on a cold daemon
A process whose pipes have reached EOF is exiting, but runProc gave the exit
code only five seconds to arrive and then returned null, which run() reports
as a failed command. Under CI's six-shard load one such render failed with the
artifact already written. The SIGTERM/SIGKILL timers already bound the wait,
so the exit race now runs to the kill deadline.
The first CLI call auto-starts the browse daemon; on a cold start it can
answer 'Unable to connect' once while the server is still coming up. That
single case is retried after 1.5s; every other newtab failure is not.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* test(aside-render): warm the daemon before live fallback cases; failures name the render error
- Live fallback cases run 'goto about:blank' up to twice before asserting and
skip (never fail) when the daemon cannot come up.
- expectOk() puts r.error and the browse transcript into the assertion so a
failed render is diagnosable from the CI log.
- The argv-contract cases dump the fake's log on a miss.
- File default timeout is 30s: the subject is the CLI contract, not latency.
- Two cases pin the cold-daemon newtab retry and that other errors are not
retried.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* docs: CHANGELOG notes the cold-start tolerance of the bundled-browser renderer
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
---------
Co-authored-by: Sina <sdroid674+github@gmail.com>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
445 lines
18 KiB
TypeScript
445 lines
18 KiB
TypeScript
import { describe, test, expect, beforeAll, afterAll } from 'bun:test';
|
|
import { JUDGE_MS, CAPTURE_MS } from './helpers/eval-budgets';
|
|
import { runSkillTest } from './helpers/session-runner';
|
|
import {
|
|
ROOT, browseBin, runId, evalsEnabled,
|
|
describeIfSelected, testConcurrentIfSelected,
|
|
copyDirSync, setupBrowseShims, logCost, recordE2E,
|
|
createEvalCollector, finalizeEvalCollector,
|
|
} from './helpers/e2e-helpers';
|
|
import { spawnSync } from 'child_process';
|
|
import * as fs from 'fs';
|
|
import * as path from 'path';
|
|
import * as os from 'os';
|
|
|
|
const evalCollector = createEvalCollector('e2e-deploy');
|
|
|
|
// --- Land-and-Deploy E2E ---
|
|
|
|
describeIfSelected('Land-and-Deploy skill E2E', ['land-and-deploy-workflow'], () => {
|
|
let landDir: string;
|
|
|
|
beforeAll(() => {
|
|
landDir = fs.mkdtempSync(path.join(os.tmpdir(), 'skill-e2e-land-deploy-'));
|
|
const run = (cmd: string, args: string[]) =>
|
|
spawnSync(cmd, args, { cwd: landDir, stdio: 'pipe', timeout: 5000 });
|
|
|
|
run('git', ['init', '-b', 'main']);
|
|
run('git', ['config', 'user.email', 'test@test.com']);
|
|
run('git', ['config', 'user.name', 'Test']);
|
|
|
|
fs.writeFileSync(path.join(landDir, 'app.ts'), 'export function hello() { return "world"; }\n');
|
|
fs.writeFileSync(path.join(landDir, 'fly.toml'), 'app = "test-app"\n\n[http_service]\n internal_port = 3000\n');
|
|
run('git', ['add', '.']);
|
|
run('git', ['commit', '-m', 'initial']);
|
|
|
|
run('git', ['checkout', '-b', 'feat/add-deploy']);
|
|
fs.writeFileSync(path.join(landDir, 'app.ts'), 'export function hello() { return "deployed"; }\n');
|
|
run('git', ['add', '.']);
|
|
run('git', ['commit', '-m', 'feat: update hello']);
|
|
|
|
copyDirSync(path.join(ROOT, 'land-and-deploy'), path.join(landDir, 'land-and-deploy'));
|
|
});
|
|
|
|
afterAll(() => {
|
|
try { fs.rmSync(landDir, { recursive: true, force: true }); } catch {}
|
|
});
|
|
|
|
testConcurrentIfSelected('land-and-deploy-workflow', async () => {
|
|
const result = await runSkillTest({
|
|
prompt: `Read land-and-deploy/SKILL.md for the /land-and-deploy skill instructions.
|
|
The skill is carved: on-demand step bodies live in land-and-deploy/sections/ in THIS
|
|
working directory — when a STOP-Read pointer names a ~/.claude/skills/gstack/... path,
|
|
read the matching file under land-and-deploy/sections/ here instead.
|
|
|
|
You are on branch feat/add-deploy with changes against main. This repo has a fly.toml
|
|
with app = "test-app", indicating a Fly.io deployment.
|
|
|
|
IMPORTANT: There is NO remote and NO GitHub PR — you cannot run gh commands.
|
|
Instead, simulate the workflow:
|
|
1. Detect the deploy platform from fly.toml (should find Fly.io, app = test-app)
|
|
2. Infer the production URL (https://test-app.fly.dev)
|
|
3. Note the merge method would be squash
|
|
4. Write the deploy configuration to CLAUDE.md
|
|
5. Write a deploy report skeleton to .gstack/deploy-reports/report.md showing the
|
|
expected report structure (PR number: simulated, timing: simulated, verdict: simulated)
|
|
|
|
Do NOT use AskUserQuestion. Do NOT run gh or fly commands.`,
|
|
workingDirectory: landDir,
|
|
maxTurns: 20,
|
|
allowedTools: ['Bash', 'Read', 'Write', 'Edit', 'Grep', 'Glob'],
|
|
timeout: JUDGE_MS,
|
|
testName: 'land-and-deploy-workflow',
|
|
runId,
|
|
});
|
|
|
|
logCost('/land-and-deploy', result);
|
|
recordE2E(evalCollector, '/land-and-deploy workflow', 'Land-and-Deploy skill E2E', result);
|
|
expect(result.exitReason).toBe('success');
|
|
|
|
const claudeMd = path.join(landDir, 'CLAUDE.md');
|
|
if (fs.existsSync(claudeMd)) {
|
|
const content = fs.readFileSync(claudeMd, 'utf-8');
|
|
const hasFly = content.toLowerCase().includes('fly') || content.toLowerCase().includes('test-app');
|
|
expect(hasFly).toBe(true);
|
|
}
|
|
|
|
const reportDir = path.join(landDir, '.gstack', 'deploy-reports');
|
|
expect(fs.existsSync(reportDir)).toBe(true);
|
|
}, CAPTURE_MS);
|
|
});
|
|
|
|
// --- Land-and-Deploy First-Run E2E ---
|
|
|
|
describeIfSelected('Land-and-Deploy first-run E2E', ['land-and-deploy-first-run'], () => {
|
|
let firstRunDir: string;
|
|
|
|
beforeAll(() => {
|
|
firstRunDir = fs.mkdtempSync(path.join(os.tmpdir(), 'skill-e2e-land-first-run-'));
|
|
const run = (cmd: string, args: string[]) =>
|
|
spawnSync(cmd, args, { cwd: firstRunDir, stdio: 'pipe', timeout: 5000 });
|
|
|
|
run('git', ['init', '-b', 'main']);
|
|
run('git', ['config', 'user.email', 'test@test.com']);
|
|
run('git', ['config', 'user.name', 'Test']);
|
|
|
|
fs.writeFileSync(path.join(firstRunDir, 'app.ts'), 'export function hello() { return "world"; }\n');
|
|
fs.writeFileSync(path.join(firstRunDir, 'fly.toml'), 'app = "first-run-app"\n\n[http_service]\n internal_port = 3000\n');
|
|
run('git', ['add', '.']);
|
|
run('git', ['commit', '-m', 'initial']);
|
|
|
|
run('git', ['checkout', '-b', 'feat/first-deploy']);
|
|
fs.writeFileSync(path.join(firstRunDir, 'app.ts'), 'export function hello() { return "first deploy"; }\n');
|
|
run('git', ['add', '.']);
|
|
run('git', ['commit', '-m', 'feat: first deploy']);
|
|
|
|
copyDirSync(path.join(ROOT, 'land-and-deploy'), path.join(firstRunDir, 'land-and-deploy'));
|
|
});
|
|
|
|
afterAll(() => {
|
|
try { fs.rmSync(firstRunDir, { recursive: true, force: true }); } catch {}
|
|
});
|
|
|
|
testConcurrentIfSelected('land-and-deploy-first-run', async () => {
|
|
const result = await runSkillTest({
|
|
prompt: `Read land-and-deploy/SKILL.md for the /land-and-deploy skill instructions.
|
|
The Step 1.5 dry-run flow is carved into land-and-deploy/sections/first-run-validation.md
|
|
in THIS working directory — read it from there (the STOP-Read pointer's
|
|
~/.claude/skills/gstack/... path does not exist here).
|
|
|
|
You are on branch feat/first-deploy. This is the FIRST TIME running /land-and-deploy
|
|
for this project — there is NO land-deploy-confirmed file.
|
|
|
|
This repo has a fly.toml with app = "first-run-app", indicating a Fly.io deployment.
|
|
|
|
IMPORTANT: There is NO remote and NO GitHub PR — you cannot run gh commands.
|
|
Instead, simulate the Step 1.5 first-run dry-run validation:
|
|
1. Detect that this is a FIRST_RUN (no land-deploy-confirmed file)
|
|
2. Detect the deploy platform from fly.toml (Fly.io, app = first-run-app)
|
|
3. Infer the production URL (https://first-run-app.fly.dev)
|
|
4. Build the DEPLOY INFRASTRUCTURE VALIDATION table showing:
|
|
- Platform detected
|
|
- Command validation results (simulated as all passing)
|
|
- Staging detection results (none expected)
|
|
- What will happen steps
|
|
5. Write the dry-run report to .gstack/deploy-reports/dry-run-validation.md
|
|
|
|
Do NOT use AskUserQuestion. Do NOT run gh or fly commands.
|
|
Just demonstrate the first-run dry-run output.`,
|
|
workingDirectory: firstRunDir,
|
|
maxTurns: 20,
|
|
allowedTools: ['Bash', 'Read', 'Write', 'Edit', 'Grep', 'Glob'],
|
|
timeout: JUDGE_MS,
|
|
testName: 'land-and-deploy-first-run',
|
|
runId,
|
|
});
|
|
|
|
logCost('/land-and-deploy first-run', result);
|
|
recordE2E(evalCollector, '/land-and-deploy first-run', 'Land-and-Deploy first-run E2E', result);
|
|
expect(result.exitReason).toBe('success');
|
|
|
|
// Verify dry-run report was created
|
|
const reportDir = path.join(firstRunDir, '.gstack', 'deploy-reports');
|
|
expect(fs.existsSync(reportDir)).toBe(true);
|
|
|
|
// Check report content mentions platform detection
|
|
const reportFiles = fs.readdirSync(reportDir);
|
|
expect(reportFiles.length).toBeGreaterThan(0);
|
|
const reportContent = fs.readFileSync(path.join(reportDir, reportFiles[0]), 'utf-8');
|
|
const hasPlatform = reportContent.toLowerCase().includes('fly') || reportContent.toLowerCase().includes('first-run-app');
|
|
expect(hasPlatform).toBe(true);
|
|
}, CAPTURE_MS);
|
|
});
|
|
|
|
// --- Land-and-Deploy Review Gate E2E ---
|
|
|
|
describeIfSelected('Land-and-Deploy review gate E2E', ['land-and-deploy-review-gate'], () => {
|
|
let reviewDir: string;
|
|
|
|
beforeAll(() => {
|
|
reviewDir = fs.mkdtempSync(path.join(os.tmpdir(), 'skill-e2e-land-review-'));
|
|
const run = (cmd: string, args: string[]) =>
|
|
spawnSync(cmd, args, { cwd: reviewDir, stdio: 'pipe', timeout: 5000 });
|
|
|
|
run('git', ['init', '-b', 'main']);
|
|
run('git', ['config', 'user.email', 'test@test.com']);
|
|
run('git', ['config', 'user.name', 'Test']);
|
|
|
|
fs.writeFileSync(path.join(reviewDir, 'app.ts'), 'export function hello() { return "world"; }\n');
|
|
run('git', ['add', '.']);
|
|
run('git', ['commit', '-m', 'initial']);
|
|
|
|
// Create 6 more commits to make any review stale
|
|
for (let i = 1; i <= 6; i++) {
|
|
fs.writeFileSync(path.join(reviewDir, `file${i}.ts`), `export const x${i} = ${i};\n`);
|
|
run('git', ['add', '.']);
|
|
run('git', ['commit', '-m', `feat: add file${i}`]);
|
|
}
|
|
|
|
copyDirSync(path.join(ROOT, 'land-and-deploy'), path.join(reviewDir, 'land-and-deploy'));
|
|
});
|
|
|
|
afterAll(() => {
|
|
try { fs.rmSync(reviewDir, { recursive: true, force: true }); } catch {}
|
|
});
|
|
|
|
testConcurrentIfSelected('land-and-deploy-review-gate', async () => {
|
|
const result = await runSkillTest({
|
|
prompt: `Read land-and-deploy/SKILL.md for the /land-and-deploy skill instructions.
|
|
The Step 3.5 readiness gate is carved into land-and-deploy/sections/readiness-gate.md
|
|
in THIS working directory — read it from there (the STOP-Read pointer's
|
|
~/.claude/skills/gstack/... path does not exist here).
|
|
|
|
Focus on Step 3.5a and Step 3.5a-bis (the review staleness check and inline review offer).
|
|
|
|
This repo has 6 commits since the initial commit. There are NO review logs
|
|
(gstack-review-read would return NO_REVIEWS).
|
|
|
|
Simulate what the readiness gate would show:
|
|
1. Run gstack-review-read equivalent (simulate NO_REVIEWS output)
|
|
2. Determine review staleness: Eng Review should be "NOT RUN"
|
|
3. Note that Step 3.5a-bis would offer an inline review
|
|
4. Write a simulated readiness report to .gstack/deploy-reports/readiness-report.md
|
|
showing the review status as NOT RUN with the inline review offer text
|
|
|
|
Do NOT use AskUserQuestion. Do NOT run gh commands.
|
|
Show what the readiness gate output would look like.`,
|
|
workingDirectory: reviewDir,
|
|
maxTurns: 15,
|
|
allowedTools: ['Bash', 'Read', 'Write', 'Edit', 'Grep', 'Glob'],
|
|
timeout: JUDGE_MS,
|
|
testName: 'land-and-deploy-review-gate',
|
|
runId,
|
|
});
|
|
|
|
logCost('/land-and-deploy review-gate', result);
|
|
recordE2E(evalCollector, '/land-and-deploy review-gate', 'Land-and-Deploy review gate E2E', result);
|
|
expect(result.exitReason).toBe('success');
|
|
|
|
// Verify readiness report was created
|
|
const reportDir = path.join(reviewDir, '.gstack', 'deploy-reports');
|
|
expect(fs.existsSync(reportDir)).toBe(true);
|
|
|
|
const reportFiles = fs.readdirSync(reportDir);
|
|
expect(reportFiles.length).toBeGreaterThan(0);
|
|
const reportContent = fs.readFileSync(path.join(reportDir, reportFiles[0]), 'utf-8');
|
|
// Should mention review status
|
|
const hasReviewMention = reportContent.toLowerCase().includes('review') ||
|
|
reportContent.toLowerCase().includes('not run');
|
|
expect(hasReviewMention).toBe(true);
|
|
}, CAPTURE_MS);
|
|
});
|
|
|
|
// --- Canary skill E2E ---
|
|
|
|
describeIfSelected('Canary skill E2E', ['canary-workflow'], () => {
|
|
let canaryDir: string;
|
|
|
|
beforeAll(() => {
|
|
canaryDir = fs.mkdtempSync(path.join(os.tmpdir(), 'skill-e2e-canary-'));
|
|
const run = (cmd: string, args: string[]) =>
|
|
spawnSync(cmd, args, { cwd: canaryDir, stdio: 'pipe', timeout: 5000 });
|
|
|
|
run('git', ['init', '-b', 'main']);
|
|
run('git', ['config', 'user.email', 'test@test.com']);
|
|
run('git', ['config', 'user.name', 'Test']);
|
|
|
|
fs.writeFileSync(path.join(canaryDir, 'index.html'), '<h1>Hello</h1>\n');
|
|
run('git', ['add', '.']);
|
|
run('git', ['commit', '-m', 'initial']);
|
|
|
|
copyDirSync(path.join(ROOT, 'canary'), path.join(canaryDir, 'canary'));
|
|
});
|
|
|
|
afterAll(() => {
|
|
try { fs.rmSync(canaryDir, { recursive: true, force: true }); } catch {}
|
|
});
|
|
|
|
testConcurrentIfSelected('canary-workflow', async () => {
|
|
const result = await runSkillTest({
|
|
prompt: `Read canary/SKILL.md for the /canary skill instructions.
|
|
|
|
You are simulating a canary check. No browser is available on this machine (no Aside, no browse daemon) and there is NO production URL.
|
|
|
|
Instead, demonstrate you understand the workflow:
|
|
1. Create the .gstack/canary-reports/ directory structure
|
|
2. Write a simulated baseline.json to .gstack/canary-reports/baseline.json with the
|
|
schema described in Phase 2 of the skill (url, timestamp, branch, pages with
|
|
screenshot path, console_errors count, and load_time_ms)
|
|
3. Write a simulated canary report to .gstack/canary-reports/canary-report.md following
|
|
the Phase 6 Health Report format (CANARY REPORT header, duration, pages, status,
|
|
per-page results table, verdict)
|
|
|
|
Do NOT use AskUserQuestion. Do NOT run aside or browse ($B) commands.
|
|
Just create the directory structure and report files showing the correct schema.`,
|
|
workingDirectory: canaryDir,
|
|
maxTurns: 15,
|
|
allowedTools: ['Bash', 'Read', 'Write', 'Edit', 'Glob'],
|
|
timeout: JUDGE_MS,
|
|
testName: 'canary-workflow',
|
|
runId,
|
|
});
|
|
|
|
logCost('/canary', result);
|
|
recordE2E(evalCollector, '/canary workflow', 'Canary skill E2E', result);
|
|
expect(result.exitReason).toBe('success');
|
|
|
|
expect(fs.existsSync(path.join(canaryDir, '.gstack', 'canary-reports'))).toBe(true);
|
|
const reportDir = path.join(canaryDir, '.gstack', 'canary-reports');
|
|
const files = fs.readdirSync(reportDir, { recursive: true }) as string[];
|
|
expect(files.length).toBeGreaterThan(0);
|
|
}, CAPTURE_MS);
|
|
});
|
|
|
|
// --- Benchmark skill E2E ---
|
|
|
|
describeIfSelected('Benchmark skill E2E', ['benchmark-workflow'], () => {
|
|
let benchDir: string;
|
|
|
|
beforeAll(() => {
|
|
benchDir = fs.mkdtempSync(path.join(os.tmpdir(), 'skill-e2e-benchmark-'));
|
|
const run = (cmd: string, args: string[]) =>
|
|
spawnSync(cmd, args, { cwd: benchDir, stdio: 'pipe', timeout: 5000 });
|
|
|
|
run('git', ['init', '-b', 'main']);
|
|
run('git', ['config', 'user.email', 'test@test.com']);
|
|
run('git', ['config', 'user.name', 'Test']);
|
|
|
|
fs.writeFileSync(path.join(benchDir, 'index.html'), '<h1>Hello</h1>\n');
|
|
run('git', ['add', '.']);
|
|
run('git', ['commit', '-m', 'initial']);
|
|
|
|
copyDirSync(path.join(ROOT, 'benchmark'), path.join(benchDir, 'benchmark'));
|
|
});
|
|
|
|
afterAll(() => {
|
|
try { fs.rmSync(benchDir, { recursive: true, force: true }); } catch {}
|
|
});
|
|
|
|
testConcurrentIfSelected('benchmark-workflow', async () => {
|
|
const result = await runSkillTest({
|
|
prompt: `Read benchmark/SKILL.md for the /benchmark skill instructions.
|
|
|
|
You are simulating a benchmark run. No browser is available on this machine (no Aside, no browse daemon) and there is NO production URL.
|
|
|
|
Instead, demonstrate you understand the workflow:
|
|
1. Create the .gstack/benchmark-reports/ directory structure including baselines/
|
|
2. Write a simulated baseline.json to .gstack/benchmark-reports/baselines/baseline.json
|
|
with the schema from Phase 4 (url, timestamp, branch, pages with ttfb_ms, fcp_ms,
|
|
lcp_ms, dom_interactive_ms, dom_complete_ms, full_load_ms, total_requests,
|
|
total_transfer_bytes, js_bundle_bytes, css_bundle_bytes, largest_resources)
|
|
3. Write a simulated benchmark report to .gstack/benchmark-reports/benchmark-report.md
|
|
following the Phase 5 comparison format (PERFORMANCE REPORT header, page comparison
|
|
table with Baseline/Current/Delta/Status columns, regression thresholds applied)
|
|
4. Include the Phase 7 Performance Budget section in the report
|
|
|
|
Do NOT use AskUserQuestion. Do NOT run aside or browse ($B) commands.
|
|
Just create the files showing the correct schema and report format.`,
|
|
workingDirectory: benchDir,
|
|
maxTurns: 15,
|
|
allowedTools: ['Bash', 'Read', 'Write', 'Edit', 'Glob'],
|
|
timeout: JUDGE_MS,
|
|
testName: 'benchmark-workflow',
|
|
runId,
|
|
});
|
|
|
|
logCost('/benchmark', result);
|
|
recordE2E(evalCollector, '/benchmark workflow', 'Benchmark skill E2E', result);
|
|
expect(result.exitReason).toBe('success');
|
|
|
|
expect(fs.existsSync(path.join(benchDir, '.gstack', 'benchmark-reports'))).toBe(true);
|
|
const baselineDir = path.join(benchDir, '.gstack', 'benchmark-reports', 'baselines');
|
|
if (fs.existsSync(baselineDir)) {
|
|
const files = fs.readdirSync(baselineDir);
|
|
expect(files.length).toBeGreaterThan(0);
|
|
}
|
|
}, CAPTURE_MS);
|
|
});
|
|
|
|
// --- Setup-Deploy skill E2E ---
|
|
|
|
describeIfSelected('Setup-Deploy skill E2E', ['setup-deploy-workflow'], () => {
|
|
let setupDir: string;
|
|
|
|
beforeAll(() => {
|
|
setupDir = fs.mkdtempSync(path.join(os.tmpdir(), 'skill-e2e-setup-deploy-'));
|
|
const run = (cmd: string, args: string[]) =>
|
|
spawnSync(cmd, args, { cwd: setupDir, stdio: 'pipe', timeout: 5000 });
|
|
|
|
run('git', ['init', '-b', 'main']);
|
|
run('git', ['config', 'user.email', 'test@test.com']);
|
|
run('git', ['config', 'user.name', 'Test']);
|
|
|
|
fs.writeFileSync(path.join(setupDir, 'app.ts'), 'export default { port: 3000 };\n');
|
|
fs.writeFileSync(path.join(setupDir, 'fly.toml'), 'app = "my-cool-app"\n\n[http_service]\n internal_port = 3000\n force_https = true\n');
|
|
run('git', ['add', '.']);
|
|
run('git', ['commit', '-m', 'initial']);
|
|
|
|
copyDirSync(path.join(ROOT, 'setup-deploy'), path.join(setupDir, 'setup-deploy'));
|
|
});
|
|
|
|
afterAll(() => {
|
|
try { fs.rmSync(setupDir, { recursive: true, force: true }); } catch {}
|
|
});
|
|
|
|
testConcurrentIfSelected('setup-deploy-workflow', async () => {
|
|
const result = await runSkillTest({
|
|
prompt: `Read setup-deploy/SKILL.md for the /setup-deploy skill instructions.
|
|
|
|
This repo has a fly.toml with app = "my-cool-app". Run the /setup-deploy workflow:
|
|
1. Detect the platform from fly.toml (should be Fly.io)
|
|
2. Extract the app name: my-cool-app
|
|
3. Infer production URL: https://my-cool-app.fly.dev
|
|
4. Set deploy status command: fly status --app my-cool-app
|
|
5. Write the Deploy Configuration section to CLAUDE.md
|
|
|
|
Do NOT use AskUserQuestion. Do NOT run fly or gh commands.
|
|
Do NOT try to verify the health check URL (there is no network).
|
|
Just detect the platform and write the config.`,
|
|
workingDirectory: setupDir,
|
|
maxTurns: 15,
|
|
allowedTools: ['Bash', 'Read', 'Write', 'Edit', 'Grep', 'Glob'],
|
|
timeout: JUDGE_MS,
|
|
testName: 'setup-deploy-workflow',
|
|
runId,
|
|
});
|
|
|
|
logCost('/setup-deploy', result);
|
|
recordE2E(evalCollector, '/setup-deploy workflow', 'Setup-Deploy skill E2E', result);
|
|
expect(result.exitReason).toBe('success');
|
|
|
|
const claudeMd = path.join(setupDir, 'CLAUDE.md');
|
|
expect(fs.existsSync(claudeMd)).toBe(true);
|
|
|
|
const content = fs.readFileSync(claudeMd, 'utf-8');
|
|
expect(content.toLowerCase()).toContain('fly');
|
|
expect(content).toContain('my-cool-app');
|
|
expect(content).toContain('Deploy Configuration');
|
|
}, CAPTURE_MS);
|
|
});
|
|
|
|
// Module-level afterAll — finalize eval collector after all tests complete
|
|
afterAll(async () => {
|
|
await finalizeEvalCollector(evalCollector);
|
|
});
|