mirror of
https://github.com/garrytan/gstack.git
synced 2026-08-11 16:50:20 +02:00
* fix(careful): warn on chained rm even when the last target is safe The safe-exception block whitelisted rm -rf of build artifacts by extracting targets with a single greedy match (.*rm ...), which only ever inspects the LAST rm in the command. A chain like 'rm -rf /; rm -rf node_modules' was therefore judged solely by its trailing safe target and allowed without warning, waving through the destructive 'rm -rf /'. Gate the shortcut to single rm invocations: when any shell separator (; | & newline, incl. JSON-escaped \n/\r from the grep extraction path) is present, fall through to the destructive-pattern check, which warns on any recursive rm. Single-command artifact cleanups still allow. Adds 3 regression tests covering semicolon and && chains in both orders. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * harden(careful): substitution separators + capital -R recursive flag (#2039) Two residual fail-opens in the same guard PR #2040 hardened, both verified by executing the script pre-fix: - rm -rf $(./wipe-all)/node_modules silently allowed: the substitution token ends in a whitelisted suffix and the safe-exception early exit skipped ALL downstream checks. $( and backtick now count as chain separators; plain $VAR expansion stays allowed. - rm -R / silently allowed: both greps required a lowercase r in the flag cluster; capital -R is the documented BSD/macOS recursive flag. Both greps now match -[a-zA-Z]*[rR]. Six new tests: substitution x2 -> ask, capital-R x2 -> ask, rm -Rf node_modules single-command -> still allowed, escaped-newline branch (existing code, previously untested), and a pinned deliberate FP (cd app && rm -rf node_modules -> ask) documenting the fail-closed direction on chains. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(context-restore): prefer the current branch's own checkpoint (#2052) All worktrees of a repo share one origin-derived slug, so they share one `~/.gstack/projects/<slug>/checkpoints/` dir. `/context-restore` loaded the newest checkpoint across the whole dir, so in one worktree it could silently restore a *sibling worktree's* newer checkpoint. Step 1 now orders candidates current-branch-first (read from each file's `branch:` frontmatter), keeping other branches as a fallback. A branch is checked out in at most one worktree, so this stops cross-worktree contamination while preserving Conductor cross-branch handoff: when the current branch has no checkpoint of its own, the full newest-first set is still used. - scan the 200 newest before partitioning so a current-branch checkpoint sitting below a burst of sibling saves is still found; output still capped at 20 - non-git / detached HEAD / branchless legacy saves fall back to the old newest-first behavior (back-compat) - +5 regression tests in context-save-hardening.test.ts (the #2052 bug case fails on the old pipeline); regenerated SKILL.md + proactive-suggestions.json Fixes #2052 Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(gbrain): pass --confirm-destructive on drift re-register (#1985) ensureSourceRegistered() handles match-but-different-path by removing the old source then re-adding it at the new path. The remove was issued as `gbrain sources remove <id> --yes`, but gbrain >= 0.42 gates `sources remove` behind `--confirm-destructive` (`--yes` alone no longer suppresses the data-loss prompt). The remove therefore fails with "To proceed, pass --confirm-destructive", which ensureSourceRegistered surfaces as "source registration failed" — aborting the entire /sync-gbrain code stage for any already-registered source whose path has drifted. The memory and brain-sync stages still pass, so the code index silently stops refreshing. The orchestrator's own safeSourcesRemove() already passes --confirm-destructive; this brings the lib helper in line with that convention. Keeps --yes for older gbrain. Tests: extend the fake gbrain shim in gbrain-sources.test.ts to simulate the gbrain >= 0.42 guard (remove without --confirm-destructive exits 1), update the drift re-register assertion, and add a regression test that proves the drift path no longer throws. Both fail on main with the exact "To proceed, pass --confirm-destructive" error and pass with the fix. Fixes #1985 Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * harden(gbrain-sources): route drift remove through #1734 guards + realpath drift check Absorbing #2031 un-blocked a destructive remove that bypassed the #1734 data-loss guards: ensureSourceRegistered's drift path issued `gbrain sources remove` directly, without the detectAutopilot + decideSourceRemove checks every other remove routes through via safeSourcesRemove. gbrain >= 0.42's own prompt was accidentally blocking that path; with --confirm-destructive passed it is live again. - Drift remove now refuses LOUDLY (throws, actionable message) while an autopilot is active or when decideSourceRemove disallows; a silent changed=false would hide the drifted registration. - decideSourceRemove's extraArgs (--keep-storage when supported) propagate to the remove call, matching safeSourcesRemove. - Drift is realpath-normalized before being declared: a symlink alias of the same directory (macOS /tmp -> /private/tmp) is a match, not drift — the probable cause of #1985's reporter hitting the remove on an unmoved repo. - Drift fires a loud stderr line (old -> new path); perpetual drift in logs is the trigger for promoting #1985's reindex-in-place design. Tests: autopilot-active refusal (no remove in call log), fail-closed refusal on unreadable sources list, --keep-storage propagation, symlink-alias no-drift; existing drift tests pin the guard probes so a live autopilot on the dev machine can't flip them. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(developer-profile): exclude mode:resources rows from SESSION_COUNT, TIER, NUDGE_ELIGIBLE (#2067) Every /office-hours run appends a mode:"resources" bookkeeping row alongside the real session row, so --read double-counted sessions (~2x): tiers promoted early and the builder-to-founder nudge armed prematurely. The file already filtered resources rows for LAST_*/CROSS_PROJECT; the same realSessions filter now feeds SESSION_COUNT/TIER, and the nudge predicate is the faithful allowlist (mode === 'builder') so a future mode #4 fails closed instead of re-opening this bug. 8 regression tests: count vs resources noise, tier boundaries both sides, nudge false-with-noise / true-at-3-builders, cross-project trailing row. Absorbed from PR #1991 by @mvann (fix + tests commits; the PR's version-bump commit is superseded by this wave's consolidated release commit). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(hooks): passThrough() two-branch contract — never emit permissionDecision:'defer' (#2035, #2006) Every AskUserQuestion died with "Tool result missing due to internal error" on current Claude Code builds (Desktop 1.14271.0, CC 2.1.177). Root cause: the question-preference-hook emitted permissionDecision:'defer' on every pass-through path. 'defer' is a real PreToolUse value, but since CC v2.1.89 its semantics are "pause this tool call for external resumption" (headless resume) — never "abstain". Interactive sessions have nothing to resume the paused call, so the tool orphaned. Pre-2.1.89 builds ignored the unknown value, which is why the hook worked when it shipped and broke later. The fix is the two-branch pass-through contract: - no context -> exit 0 with EXACTLY empty stdout - memory nuggets present -> hookSpecificOutput with hookEventName + additionalContext ONLY (the documented shape; plan-tune Layer 8 memory injection ships through this branch and keeps working) defer() is renamed passThrough() so the function says what it does, and docs/spikes/claude-code-hook-mutation.md's protocol contract (cited by the hook header) is corrected in the same commit — it taught '"defer" — let permission flow continue' and was the reintroduction vector. Test contract rewritten in the same commit (13 assertions across 3 files, verified fail-first against the unfixed hook): pass-through paths assert exact-empty stdout (a garbage/partial write cannot slip past an optional-chained parse), the nugget path asserts permissionDecision is ABSENT while additionalContext survives, and a new tripwire asserts no non-deny path ever puts the string "permissionDecision" on stdout. The deny (auto-decide) and Conductor prose-redirect paths are unchanged. Deployment: no migration needed — settings.json points at the absolute bash shim which execs the .ts live; /gstack-upgrade delivers the fix. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(one-way-doors): unify credential noun net + wire it into the runtime (#2024) Library fix: revoke/reset/rotate now share ONE noun alternation (api key, token, secret, credential, access key, password) with optional plural s?. Pre-fix leaks: "reset my secret", "reset my access key", "revoke my secret" (mismatched per-verb lists) and every plural form ("rotate the credentials", "revoke all tokens" — \b(...)\b cannot match a trailing s). Runtime wiring — the regexes could never fire in production before: - gstack-question-preference --check gains --summary-stdin: the question text pipes via stdin (never argv — summaries carry quotes/newlines/shell metacharacters) and feeds isOneWayDoor alongside the id, so an ad-hoc destructive question with a stored never-ask preference now forces ASK_NORMALLY. Empty/absent stdin keeps exact id-only semantics. - question-preference-hook falls back to classifyQuestion(question text) when the registry lookup misses, so unregistered destructive questions pass through to a human instead of auto-deciding. - question-tuning resolver prose shows the piped form (SKILL.md regen lands in the wave's release commit). Tripwires (verified fail-first): full verbs x nouns x singular/plural matrix with the #2024 repro rows, benign-summary no-over-match rows, stdin transport survival (quotes/newlines), empty-stdin fail-safe, and hook fallback both directions (destructive -> pass-through, benign -> deny). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(design): loud integer-flag contract for --count/--retry/--timeout (#2032) design variants --count abc silently generated ZERO variants and exited 0: parseInt(NaN) flowed through Math.min into the generation loop bound. The same NaN class was live on the two sibling flags in the same file: --retry abc made generate() a silent no-op (attempt <= NaN never true, null output, exit 0) and --timeout abc killed the serve board ~immediately (setTimeout(NaN)). New design/src/flag-utils.ts: parseIntFlag (pure, unit-testable) + normalizeIntFlag (CLI wrapper). Contract matches the --viewports precedent (error loudly on nonsense — these commands spend real image-API money, a silent fixup hides typos from calling agents): undefined -> default; bare flag/empty/non-integer ("3.7" rejected, not truncated)/below-min -> exit 1 with usage hint; above-max -> clamp with stderr warning. --count normalizes at the variants() consumption site so programmatic callers are covered, with the ceiling derived from STYLE_VARIATIONS.length instead of a magic 7; the CLI passes the raw flag through (a pre-parseInt would truncate "3.7"). Tripwires live in test/design-flag-utils.test.ts — deliberately under test/, not design/test/, which is invisible to the bun test glob, TEST_ROOTS, and every workflow (wiring design/test/ into CI is a captured TODO). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(gbrain): thin-client state — remote-MCP brains no longer classify as broken-config (#2051) A thin client (remote-HTTP MCP brain, no local engine by design) probed `gbrain sources list`, which gbrain's dispatch guard REFUSES on thin clients (exit 1, no recognized error string), so the classifier fell to its defensive broken-config default and every suppression gate silently hid brain-aware blocks from exactly the users on a shared team brain. New 'thin-client' state, detected PRE-probe from gbrain's own remote_mcp config marker via the existing gbrainConfigPath() helper (mirrors gbrain's isThinClient(); honors GBRAIN_HOME; zero network, immune to error-string drift), with a /thin[- ]client/ stderr backstop in the probe catch. Remote reachability is deliberately NOT probed by the classifier — that is the #1964 pathology; gbrain calls degrade gracefully at use time, and the detect JSON says so honestly (gbrain_thin_client: {probed: false}). The state is admitted at every suppression gate — gstack-gbrain-detect --is-ok (drives setup + gbrain-refresh), gen-skill-docs' detection override, gstack-config gbrain-refresh — while the sync stages (code/memory/dream) SKIP with an accurate reason: code indexing runs on the brain server, memory syncs via the remote brain's artifacts pull. The two consumer classes need opposite answers, which is why this is a distinct state and not a skip-the-probe special case. sync-gbrain Step 1.5 and setup-gbrain prose route thin-client to proceed, never into broken-config remediation. detectMcpMode secondary generalization: url-match against the config's remote_mcp.mcp_url (deterministic — gbrain mounts at the generic /mcp path) -> name pattern gbrain[-_]* -> stdio command token; gbrain_mcp_mode stays a 3-value enum. Tripwires: end-to-end --is-ok exits 0 on a thin-client fixture AND still exits 1 on broken-config (the gate didn't widen); pre-probe + stderr-fallback classifier paths; 4 detectMcpMode identification cases incl. a non-matching url that must NOT false-positive. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * release: v1.60.0.0 — regen SKILL.md, VERSION, CHANGELOG, TODOS follow-ups - Regenerate all SKILL.md from templates (question-tuning --summary-stdin prose from #2024, context-restore branch preference from PR #2054, sync-gbrain/setup-gbrain thin-client prose from #2051) + llms.txt. - VERSION + package.json -> 1.60.0.0 (bin/gstack-next-version, queue-aware: #1815 claims 1.59.0.0, #2213 claims 1.59.1.0). - CHANGELOG release summary + itemized entry crediting @jbetala7 (x3) and @mvann. - TODOS.md: three eng-review follow-ups (design/test CI wiring + documented pre-existing retry-after flake, /context-save worktree identity, gbrain reindex-in-place conditional on the new drift log). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(resolvers): compress --summary-stdin preamble prose to fit parity budget; re-bless ship goldens The v1.57.7.0 parity suite caps investigate's generated size at 1.09x baseline; the #2024 question-tuning prose (duplicated into every tier->=2 skill) tipped it to 1.092. Compressed to a single inline command + short pointer (the full rationale lives in bin/gstack-question-preference's header and the one-way-doors module docs). Ship goldens re-blessed against the final resolver text (conscious template-change acknowledgment, per the golden-file regression contract). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(e2e): office-hours-spec-review turn budget fits the carved skill layout (#2473) The test failed deterministically with error_max_turns at 9 turns on main and this branch alike (CI attempt logs + local main repro). Root cause from the failing transcript: the Spec Review Loop content is carved out of office-hours/SKILL.md into office-hours/sections/, so the agent needs discovery hops (grep SKILL.md -> ls sections/ -> read the section) before it can write — 8 tool turns + the closing text turn = 9 > the 8-turn budget, which predates the carve. Observed failures wrote a CORRECT summary on tool turn 8 and died on the closing turn. maxTurns 8 -> 12. Verified: PASS locally post-fix (7 turns this run — the extra headroom absorbs discovery-path nondeterminism). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(e2e): review-dashboard-via session budget survives runner contention (#2473) The test failed on CI (and its baseline run) with the timeout signature: 0 turns, $0.00, exactly 183s, 3/3 attempts — the spawned claude -p session never emitted a single stream event before the 180s inner timeout. The file's tests run concurrently on one runner; session startup queues behind sibling sessions, and this test had the tightest budget in the file (the 240s-budget tests in the same job passed). A clean local run takes 270s wall for 4 turns, confirming 180s was too tight even without contention. Inner timeout 180s -> 300s; outer bun timeout 240s -> 360s to keep headroom over the inner budget. Verified: PASS locally post-fix (4 turns, 270s). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(e2e): retro-base-branch session budget survives runner contention (#2473) Same class as review-dashboard-via, one test over in the same file: /retro is a long multi-step flow whose clean pass measures 225-239s — a coin flip against the 240s inner budget. First CI run passed at 225s; the rerun timed out at the 240s line on all 3 attempts (exitReason "timeout"); the local verification run passed at 239s, ONE second under the old cap. Inner timeout 240s -> 360s; outer bun timeout 300s -> 480s for headroom. Verified: PASS locally post-fix (17 turns, 239s). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: Jayesh Betala <jayesh.betala7@gmail.com> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Co-authored-by: Michael Vann <9221873+mvann@users.noreply.github.com>
435 lines
17 KiB
TypeScript
435 lines
17 KiB
TypeScript
/**
|
|
* Tier-2 hardening tests for context-save + context-restore.
|
|
*
|
|
* These exercise the exact bash snippets from the SKILL.md templates,
|
|
* without spawning claude -p. Free tier, runs in milliseconds.
|
|
*
|
|
* Covers the hardening work from commit 3df8ea86:
|
|
* - Bash-side title sanitizer (allowlist a-z0-9.-, cap 60, default "untitled")
|
|
* - Collision-safe filenames (random suffix on same-second double-save)
|
|
* - head -20 cap on the restore-flow directory listing
|
|
* - Migration HOME unset guard
|
|
* - Empty-set "NO_CHECKPOINTS" fallback
|
|
*/
|
|
|
|
import { describe, test, expect, beforeEach, afterEach } from 'bun:test';
|
|
import { spawnSync } from 'child_process';
|
|
import * as fs from 'fs';
|
|
import * as path from 'path';
|
|
import * as os from 'os';
|
|
|
|
const ROOT = path.resolve(import.meta.dir, '..');
|
|
|
|
// The exact sanitize+collision bash used by context-save/SKILL.md Step 4.
|
|
// Kept in sync with context-save/SKILL.md.tmpl. If the template changes
|
|
// this helper out of alignment, the title-sanitize tests fail — intended.
|
|
const TITLE_BASH = `
|
|
RAW="\${TITLE_RAW:-untitled}"
|
|
TITLE_SLUG=$(printf '%s' "$RAW" | tr '[:upper:]' '[:lower:]' | tr -s ' \\t' '-' | tr -cd 'a-z0-9.-' | cut -c1-60)
|
|
TITLE_SLUG="\${TITLE_SLUG:-untitled}"
|
|
FILE="\${CHECKPOINT_DIR}/\${TIMESTAMP}-\${TITLE_SLUG}.md"
|
|
if [ -e "$FILE" ]; then
|
|
SUFFIX=$(LC_ALL=C tr -dc 'a-z0-9' < /dev/urandom 2>/dev/null | head -c 4 || printf '%04x' "$$")
|
|
FILE="\${CHECKPOINT_DIR}/\${TIMESTAMP}-\${TITLE_SLUG}-\${SUFFIX}.md"
|
|
fi
|
|
echo "TITLE_SLUG=$TITLE_SLUG"
|
|
echo "FILE=$FILE"
|
|
`;
|
|
|
|
// The exact selection used by context-restore/SKILL.md Step 1: scan newest 200,
|
|
// order current-branch checkpoints first (fallback: all branches), cap at 20.
|
|
// CURRENT_BRANCH is injected via env in tests; the skill resolves it from git.
|
|
const RESTORE_FIND_BASH = `
|
|
if [ ! -d "$CHECKPOINT_DIR" ]; then
|
|
echo "NO_CHECKPOINTS"
|
|
else
|
|
ALL=$(find "$CHECKPOINT_DIR" -maxdepth 1 -name "*.md" -type f 2>/dev/null | sort -r | head -200)
|
|
if [ -z "$ALL" ]; then
|
|
echo "NO_CHECKPOINTS"
|
|
else
|
|
: "\${CURRENT_BRANCH:=$(git rev-parse --abbrev-ref HEAD 2>/dev/null)}"
|
|
SAME=""; OTHER=""
|
|
while IFS= read -r f; do
|
|
[ -n "$f" ] || continue
|
|
b=$(grep -m1 '^branch:' "$f" 2>/dev/null | sed 's/^branch:[[:space:]]*//')
|
|
if [ -n "$CURRENT_BRANCH" ] && [ "$b" = "$CURRENT_BRANCH" ]; then
|
|
SAME="\${SAME}\${f}
|
|
"
|
|
else
|
|
OTHER="\${OTHER}\${f}
|
|
"
|
|
fi
|
|
done <<EOF
|
|
$ALL
|
|
EOF
|
|
FILES=$(printf '%s%s' "$SAME" "$OTHER" | grep -v '^[[:space:]]*$' | head -20)
|
|
echo "$FILES"
|
|
fi
|
|
fi
|
|
`;
|
|
|
|
function runBash(script: string, env: Record<string, string>): { stdout: string; stderr: string; exitCode: number } {
|
|
const result = spawnSync('bash', ['-c', script], {
|
|
env: { ...process.env, ...env },
|
|
stdio: ['ignore', 'pipe', 'pipe'],
|
|
timeout: 5000,
|
|
});
|
|
return {
|
|
stdout: result.stdout.toString(),
|
|
stderr: result.stderr.toString(),
|
|
exitCode: result.status ?? 1,
|
|
};
|
|
}
|
|
|
|
function parseKV(stdout: string): Record<string, string> {
|
|
const out: Record<string, string> = {};
|
|
for (const line of stdout.split('\n')) {
|
|
const eq = line.indexOf('=');
|
|
if (eq > 0) out[line.slice(0, eq)] = line.slice(eq + 1);
|
|
}
|
|
return out;
|
|
}
|
|
|
|
// ─── Title sanitizer ───────────────────────────────────────────────────────
|
|
|
|
describe('context-save: title sanitizer', () => {
|
|
let tmp: string;
|
|
beforeEach(() => { tmp = fs.mkdtempSync(path.join(os.tmpdir(), 'ctx-san-')); });
|
|
afterEach(() => { try { fs.rmSync(tmp, { recursive: true, force: true }); } catch {} });
|
|
|
|
test('shell metachars stripped to allowlist', () => {
|
|
const kv = parseKV(runBash(TITLE_BASH, {
|
|
TITLE_RAW: '$(rm -rf /) `whoami` ; echo pwned',
|
|
CHECKPOINT_DIR: tmp,
|
|
TIMESTAMP: '20260419-120000',
|
|
}).stdout);
|
|
expect(kv.TITLE_SLUG).toMatch(/^[a-z0-9.-]*$/);
|
|
expect(kv.TITLE_SLUG).not.toContain('$');
|
|
expect(kv.TITLE_SLUG).not.toContain('(');
|
|
expect(kv.TITLE_SLUG).not.toContain(';');
|
|
expect(kv.TITLE_SLUG).not.toContain('`');
|
|
});
|
|
|
|
test('path traversal attempt stripped', () => {
|
|
const kv = parseKV(runBash(TITLE_BASH, {
|
|
TITLE_RAW: '../../../etc/passwd',
|
|
CHECKPOINT_DIR: tmp,
|
|
TIMESTAMP: '20260419-120000',
|
|
}).stdout);
|
|
expect(kv.TITLE_SLUG).not.toContain('/');
|
|
// Slashes stripped, dots retained — result is contained within the
|
|
// checkpoint directory (no path escape possible). The exact number of dots
|
|
// depends on the input; what matters is the file stays inside $CHECKPOINT_DIR.
|
|
expect(kv.FILE.startsWith(`${tmp}/`)).toBe(true);
|
|
expect(path.dirname(kv.FILE)).toBe(tmp);
|
|
});
|
|
|
|
test('uppercase lowercased', () => {
|
|
const kv = parseKV(runBash(TITLE_BASH, {
|
|
TITLE_RAW: 'Wintermute Progress',
|
|
CHECKPOINT_DIR: tmp,
|
|
TIMESTAMP: '20260419-120000',
|
|
}).stdout);
|
|
expect(kv.TITLE_SLUG).toBe('wintermute-progress');
|
|
});
|
|
|
|
test('whitespace collapsed to single hyphen', () => {
|
|
const kv = parseKV(runBash(TITLE_BASH, {
|
|
TITLE_RAW: 'foo bar\t\tbaz',
|
|
CHECKPOINT_DIR: tmp,
|
|
TIMESTAMP: '20260419-120000',
|
|
}).stdout);
|
|
expect(kv.TITLE_SLUG).toBe('foo-bar-baz');
|
|
});
|
|
|
|
test('length capped at 60 chars', () => {
|
|
const kv = parseKV(runBash(TITLE_BASH, {
|
|
TITLE_RAW: 'a'.repeat(200),
|
|
CHECKPOINT_DIR: tmp,
|
|
TIMESTAMP: '20260419-120000',
|
|
}).stdout);
|
|
expect(kv.TITLE_SLUG.length).toBe(60);
|
|
});
|
|
|
|
test('empty title falls back to "untitled"', () => {
|
|
const kv = parseKV(runBash(TITLE_BASH, {
|
|
TITLE_RAW: '',
|
|
CHECKPOINT_DIR: tmp,
|
|
TIMESTAMP: '20260419-120000',
|
|
}).stdout);
|
|
expect(kv.TITLE_SLUG).toBe('untitled');
|
|
});
|
|
|
|
test('only-special-chars title falls back to "untitled"', () => {
|
|
const kv = parseKV(runBash(TITLE_BASH, {
|
|
TITLE_RAW: '!@#$%^&*()+=<>?',
|
|
CHECKPOINT_DIR: tmp,
|
|
TIMESTAMP: '20260419-120000',
|
|
}).stdout);
|
|
expect(kv.TITLE_SLUG).toBe('untitled');
|
|
});
|
|
|
|
test('unicode stripped to ASCII allowlist', () => {
|
|
const kv = parseKV(runBash(TITLE_BASH, {
|
|
TITLE_RAW: '日本語 emoji 🚀 test',
|
|
CHECKPOINT_DIR: tmp,
|
|
TIMESTAMP: '20260419-120000',
|
|
}).stdout);
|
|
expect(kv.TITLE_SLUG).toMatch(/^[a-z0-9.-]*$/);
|
|
// Must contain the ASCII words that survived
|
|
expect(kv.TITLE_SLUG).toContain('emoji');
|
|
expect(kv.TITLE_SLUG).toContain('test');
|
|
});
|
|
|
|
test('numbers + dots + hyphens preserved', () => {
|
|
const kv = parseKV(runBash(TITLE_BASH, {
|
|
TITLE_RAW: 'v1.0.1-release-notes',
|
|
CHECKPOINT_DIR: tmp,
|
|
TIMESTAMP: '20260419-120000',
|
|
}).stdout);
|
|
expect(kv.TITLE_SLUG).toBe('v1.0.1-release-notes');
|
|
});
|
|
});
|
|
|
|
// ─── Filename collision handling ───────────────────────────────────────────
|
|
|
|
describe('context-save: filename collision', () => {
|
|
let tmp: string;
|
|
beforeEach(() => { tmp = fs.mkdtempSync(path.join(os.tmpdir(), 'ctx-col-')); });
|
|
afterEach(() => { try { fs.rmSync(tmp, { recursive: true, force: true }); } catch {} });
|
|
|
|
test('first save with title uses predictable path', () => {
|
|
const kv = parseKV(runBash(TITLE_BASH, {
|
|
TITLE_RAW: 'foo',
|
|
CHECKPOINT_DIR: tmp,
|
|
TIMESTAMP: '20260419-120000',
|
|
}).stdout);
|
|
expect(kv.FILE).toBe(`${tmp}/20260419-120000-foo.md`);
|
|
});
|
|
|
|
test('second save same-second same-title gets random suffix', () => {
|
|
// Pre-seed: file already exists at the predictable path.
|
|
fs.writeFileSync(`${tmp}/20260419-120000-foo.md`, 'prior save');
|
|
const kv = parseKV(runBash(TITLE_BASH, {
|
|
TITLE_RAW: 'foo',
|
|
CHECKPOINT_DIR: tmp,
|
|
TIMESTAMP: '20260419-120000',
|
|
}).stdout);
|
|
// Path must differ (append-only contract).
|
|
expect(kv.FILE).not.toBe(`${tmp}/20260419-120000-foo.md`);
|
|
// Suffix format: base-XXXX.md where XXXX matches the suffix allowlist.
|
|
expect(kv.FILE).toMatch(new RegExp(`^${tmp.replace(/[/.]/g, '\\$&')}/20260419-120000-foo-[a-z0-9]+\\.md$`));
|
|
});
|
|
|
|
test('collision suffix preserves append-only — prior file intact', () => {
|
|
const priorPath = `${tmp}/20260419-120000-foo.md`;
|
|
fs.writeFileSync(priorPath, 'critical prior save');
|
|
const kv = parseKV(runBash(TITLE_BASH, {
|
|
TITLE_RAW: 'foo',
|
|
CHECKPOINT_DIR: tmp,
|
|
TIMESTAMP: '20260419-120000',
|
|
}).stdout);
|
|
// Write a new file at the collision-safe path.
|
|
fs.writeFileSync(kv.FILE, 'new save');
|
|
// Prior file must still exist and be untouched.
|
|
expect(fs.readFileSync(priorPath, 'utf-8')).toBe('critical prior save');
|
|
expect(fs.readFileSync(kv.FILE, 'utf-8')).toBe('new save');
|
|
// Directory should have exactly 2 files.
|
|
expect(fs.readdirSync(tmp).length).toBe(2);
|
|
});
|
|
|
|
test('different titles same second — no collision, no suffix', () => {
|
|
fs.writeFileSync(`${tmp}/20260419-120000-foo.md`, 'first save');
|
|
const kv = parseKV(runBash(TITLE_BASH, {
|
|
TITLE_RAW: 'bar',
|
|
CHECKPOINT_DIR: tmp,
|
|
TIMESTAMP: '20260419-120000',
|
|
}).stdout);
|
|
// Different title → predictable path, no suffix.
|
|
expect(kv.FILE).toBe(`${tmp}/20260419-120000-bar.md`);
|
|
});
|
|
});
|
|
|
|
// ─── Restore flow: head-20 cap + empty-set ─────────────────────────────────
|
|
|
|
describe('context-restore: find + sort + head cap', () => {
|
|
let tmp: string;
|
|
beforeEach(() => { tmp = fs.mkdtempSync(path.join(os.tmpdir(), 'ctx-rest-')); });
|
|
afterEach(() => { try { fs.rmSync(tmp, { recursive: true, force: true }); } catch {} });
|
|
|
|
test('missing directory → NO_CHECKPOINTS', () => {
|
|
const out = runBash(RESTORE_FIND_BASH, {
|
|
CHECKPOINT_DIR: `${tmp}/nonexistent`,
|
|
}).stdout;
|
|
expect(out.trim()).toBe('NO_CHECKPOINTS');
|
|
});
|
|
|
|
test('empty directory → NO_CHECKPOINTS', () => {
|
|
const out = runBash(RESTORE_FIND_BASH, {
|
|
CHECKPOINT_DIR: tmp,
|
|
}).stdout;
|
|
expect(out.trim()).toBe('NO_CHECKPOINTS');
|
|
});
|
|
|
|
test('directory with non-.md files → NO_CHECKPOINTS', () => {
|
|
fs.writeFileSync(`${tmp}/not-a-save.txt`, 'noise');
|
|
fs.writeFileSync(`${tmp}/.DS_Store`, 'macos');
|
|
const out = runBash(RESTORE_FIND_BASH, {
|
|
CHECKPOINT_DIR: tmp,
|
|
}).stdout;
|
|
expect(out.trim()).toBe('NO_CHECKPOINTS');
|
|
});
|
|
|
|
test('50 .md files → only 20 returned, newest first by filename', () => {
|
|
// Seed 50 files with monotonically increasing timestamps.
|
|
for (let i = 0; i < 50; i++) {
|
|
const ts = `20260419-${String(120000 + i).padStart(6, '0')}`;
|
|
fs.writeFileSync(`${tmp}/${ts}-file${i}.md`, `content ${i}`);
|
|
}
|
|
const out = runBash(RESTORE_FIND_BASH, {
|
|
CHECKPOINT_DIR: tmp,
|
|
}).stdout;
|
|
const lines = out.trim().split('\n').filter(Boolean);
|
|
expect(lines.length).toBe(20);
|
|
// sort -r → newest first by filename. Highest timestamps (files 30-49).
|
|
expect(lines[0]).toContain('file49');
|
|
expect(lines[19]).toContain('file30');
|
|
});
|
|
|
|
test('sort is by filename prefix, NOT mtime', () => {
|
|
// Older filename, newer mtime. Sort -r must still put newer filename first.
|
|
const olderByFilename = `${tmp}/20260101-120000-old.md`;
|
|
const newerByFilename = `${tmp}/20260419-120000-new.md`;
|
|
fs.writeFileSync(olderByFilename, 'old content');
|
|
fs.writeFileSync(newerByFilename, 'new content');
|
|
// Scramble mtimes: older filename gets newer mtime.
|
|
const now = Math.floor(Date.now() / 1000);
|
|
fs.utimesSync(olderByFilename, now, now);
|
|
fs.utimesSync(newerByFilename, now - 86400 * 30, now - 86400 * 30);
|
|
|
|
const out = runBash(RESTORE_FIND_BASH, {
|
|
CHECKPOINT_DIR: tmp,
|
|
}).stdout;
|
|
const lines = out.trim().split('\n').filter(Boolean);
|
|
expect(lines[0]).toBe(newerByFilename);
|
|
expect(lines[1]).toBe(olderByFilename);
|
|
});
|
|
|
|
test('no listing-cwd fallback when empty (macOS xargs ls gotcha)', () => {
|
|
// On macOS, `find ... | xargs ls -1t` with zero results falls back to
|
|
// listing the current working directory. Our find|sort|head pattern must
|
|
// NOT have that behavior. Running from a dir with many .md files.
|
|
const out = runBash(RESTORE_FIND_BASH, {
|
|
CHECKPOINT_DIR: tmp,
|
|
// Intentionally: working directory is the gstack repo which has many .md files.
|
|
}).stdout;
|
|
expect(out.trim()).toBe('NO_CHECKPOINTS');
|
|
// Must NOT contain any .md filename from cwd.
|
|
expect(out).not.toContain('SKILL.md');
|
|
expect(out).not.toContain('README.md');
|
|
});
|
|
});
|
|
|
|
// ─── Current-branch preference (#2052) ──────────────────────────────────────
|
|
//
|
|
// All worktrees of a repo share one origin-derived slug → one checkpoints dir.
|
|
// Restore must prefer the CURRENT branch's own checkpoint so a sibling
|
|
// worktree's newer save can't shadow it, while still falling back across
|
|
// branches (Conductor handoff) when the current branch has none.
|
|
|
|
describe('context-restore: current-branch preference (#2052)', () => {
|
|
let tmp: string;
|
|
beforeEach(() => { tmp = fs.mkdtempSync(path.join(os.tmpdir(), 'ctx-branch-')); });
|
|
afterEach(() => { try { fs.rmSync(tmp, { recursive: true, force: true }); } catch {} });
|
|
|
|
function writeCheckpoint(ts: string, branch: string | null): string {
|
|
const file = `${tmp}/${ts}-work.md`;
|
|
const fm = branch === null
|
|
? `---\nstatus: in-progress\n---\n`
|
|
: `---\nstatus: in-progress\nbranch: ${branch}\n---\n`;
|
|
fs.writeFileSync(file, fm);
|
|
return file;
|
|
}
|
|
|
|
function firstCandidate(currentBranch?: string): string {
|
|
const env: Record<string, string> = { CHECKPOINT_DIR: tmp };
|
|
if (currentBranch !== undefined) env.CURRENT_BRANCH = currentBranch;
|
|
const out = runBash(RESTORE_FIND_BASH, env).stdout;
|
|
return out.trim().split('\n').filter(Boolean)[0] ?? '';
|
|
}
|
|
|
|
test('the bug: current-branch save is NOT shadowed by a newer sibling-worktree save', () => {
|
|
const mine = writeCheckpoint('20260101-120000', 'feature-a'); // older, my branch
|
|
writeCheckpoint('20260619-120000', 'feature-b'); // newer, sibling worktree
|
|
// On feature-a, restore must load feature-a's own (older) checkpoint.
|
|
expect(firstCandidate('feature-a')).toBe(mine);
|
|
});
|
|
|
|
test('fallback: current branch has no checkpoint → newest across all branches (Conductor handoff)', () => {
|
|
writeCheckpoint('20260101-120000', 'feature-a');
|
|
const newest = writeCheckpoint('20260619-120000', 'feature-b');
|
|
// On feature-c (no own checkpoint), cross-branch resume still works.
|
|
expect(firstCandidate('feature-c')).toBe(newest);
|
|
});
|
|
|
|
test('back-compat: empty current branch (non-git) → newest across all', () => {
|
|
writeCheckpoint('20260101-120000', 'feature-a');
|
|
const newest = writeCheckpoint('20260619-120000', 'feature-b');
|
|
expect(firstCandidate('')).toBe(newest);
|
|
});
|
|
|
|
test('checkpoints without a branch frontmatter still rank as fallback, never lost', () => {
|
|
const mine = writeCheckpoint('20260101-120000', 'feature-a');
|
|
writeCheckpoint('20260301-120000', null); // legacy save, no branch field
|
|
const out = runBash(RESTORE_FIND_BASH, { CHECKPOINT_DIR: tmp, CURRENT_BRANCH: 'feature-a' }).stdout;
|
|
const lines = out.trim().split('\n').filter(Boolean);
|
|
expect(lines[0]).toBe(mine); // current branch first
|
|
expect(lines.length).toBe(2); // legacy file is still present
|
|
});
|
|
|
|
test('within the current branch, ordering stays newest-first', () => {
|
|
const older = writeCheckpoint('20260101-120000', 'feature-a');
|
|
const newer = writeCheckpoint('20260619-120000', 'feature-a');
|
|
const out = runBash(RESTORE_FIND_BASH, { CHECKPOINT_DIR: tmp, CURRENT_BRANCH: 'feature-a' }).stdout;
|
|
const lines = out.trim().split('\n').filter(Boolean);
|
|
expect(lines[0]).toBe(newer);
|
|
expect(lines[1]).toBe(older);
|
|
});
|
|
});
|
|
|
|
// ─── Migration HOME guard ──────────────────────────────────────────────────
|
|
|
|
describe('migration v1.1.3.0: HOME guard', () => {
|
|
let tmp: string;
|
|
const MIGRATION = path.join(ROOT, 'gstack-upgrade', 'migrations', 'v1.1.3.0.sh');
|
|
|
|
beforeEach(() => { tmp = fs.mkdtempSync(path.join(os.tmpdir(), 'ctx-home-')); });
|
|
afterEach(() => { try { fs.rmSync(tmp, { recursive: true, force: true }); } catch {} });
|
|
|
|
test('HOME unset → exits 0 with diagnostic, no filesystem changes', () => {
|
|
// Create a file that would be wiped by an HOME="" bug: /.claude/skills/gstack/checkpoint
|
|
// (not actually writable by the test, but we verify the script doesn't TRY).
|
|
// Spawn without HOME in env.
|
|
const env = { PATH: process.env.PATH || '/usr/bin:/bin' } as Record<string, string>;
|
|
const result = spawnSync('bash', [MIGRATION], {
|
|
env,
|
|
stdio: ['ignore', 'pipe', 'pipe'],
|
|
timeout: 5000,
|
|
});
|
|
expect(result.status).toBe(0);
|
|
expect(result.stderr.toString()).toContain('HOME is unset');
|
|
});
|
|
|
|
test('HOME="" → exits 0 with diagnostic', () => {
|
|
const result = spawnSync('bash', [MIGRATION], {
|
|
env: { HOME: '', PATH: process.env.PATH || '/usr/bin:/bin' },
|
|
stdio: ['ignore', 'pipe', 'pipe'],
|
|
timeout: 5000,
|
|
});
|
|
expect(result.status).toBe(0);
|
|
expect(result.stderr.toString()).toContain('HOME is unset or empty');
|
|
// Critical: no stdout (no "Removed stale" messages — nothing touched).
|
|
expect(result.stdout.toString().trim()).toBe('');
|
|
});
|
|
});
|