mirror of
https://github.com/garrytan/gstack.git
synced 2026-08-29 01:10:50 +02:00
* test(wireup): make gbrain-missing PATH fixture hermetic The gbrain-missing test appended the host PATH (and a hardcoded /opt/homebrew/bin) to the fixture PATH, so on any machine with a real gbrain installed the 'missing' case saw it, exited 0 instead of 2, and could never fail where the bug exists — a false green for a whole machine class. The fixture now keeps only root-owned OS dirs on the child PATH, and a new determinism check plants a host-like gbrain to prove it is unreachable. Absorbed from PR #2615 with authorship preserved; the PR-thread liveness screenshot (docs/images/gstack-pr-liveness-2255.png) is dropped — referenced by nothing in the tree. Fixes #2255 Co-authored-by: CommandCodeBot <noreply@commandcode.ai> * fix(evidence): stop bun's dotenv autoload from reaching the spawned command `bin/gstack-evidence` has a `#!/usr/bin/env bun` shebang, and bun AUTO-LOADS `.env`, `.env.<NODE_ENV>` and `.env.local` from the cwd into `process.env`. The wrapper then spawned the command with no `env` override, so every command run through it inherited those variables — and a repo `.env.local` routinely holds production credentials. Two things go wrong, and the second is worse than the leak: 1. Secrets reach a child that would not otherwise have them. `npm test` run by hand in the same shell sees none of them; the same command through the wrapper sees all of them. 2. THE COMMAND UNDER TEST BEHAVES DIFFERENTLY, so the ledger certifies a run that is not the run CI performs. Observed in a Next.js repo on 2026-08-20: four tests failed 4/4 through the wrapper and passed 5/5 without it, because app code branched on env vars only the wrapper supplied. Nearly an hour went into chasing a "flake" that was the measuring instrument. The wrapper exists to record trustworthy evidence, so silently altering the environment defeats its purpose. The fix builds the child env from `process.env` minus the keys bun injected, and detection is exact rather than heuristic: verified on bun 1.3.11, a dotenv file does NOT override a variable the shell already exported (the shell's value wins). So a key whose live value equals the dotenv file's value was injected by bun, and dropping it restores the environment the user's own shell would have given the command. A key whose live value differs is genuinely the caller's and survives. `BUN_DOTENV_FILES()` mirrors bun's precedence, including that `.env.local` is skipped when NODE_ENV is "test" — scrubbing a key bun never loaded would strip a variable the caller legitimately provided. Escape hatch: GSTACK_EVIDENCE_KEEP_DOTENV=1 keeps the old behaviour. When keys are scrubbed the wrapper warns with the KEY NAMES ONLY, so the diagnostic cannot become the leak it prevents. Tests: 6 cases, mutation-verified — removing `env: spawnEnv` reddens exactly the two leak tests and restoring it gives 30/30. Every leak test asserts the scrub warning fired, because `bun test` sets NODE_ENV=test and the first version of these tests passed vacuously against a `.env.local` bun had never loaded. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Absorbed from PR #2652 with authorship preserved. Wave additions: a doc-comment on the ${VAR}-expansion limitation (bun expands refs, the reader compares raw text — those keys are left in the child env, failing open) and a regression pin for the unreadable-.env fail-open path with a functional DAC-override skip guard. Fixes #2624 * fix(setup): reap dangling skill dirs when the payload is gone cleanup_old_claude_symlinks derived its work list from the payload directory, so when the payload was gone — precisely when orphans exist — the glob matched nothing and the loop never ran; the -f guard also followed symlinks, hiding dangling SKILL.md links even with a payload present. The cleanup now scans the DESTINATION skills dir (-e/-L, so dangling symlinks are visible) and anchors SKILL.md provenance to path segments (gstack/*, */gstack/*, */.gstack/render/claude/*) instead of a bare *gstack* substring that would eat a user skill under ~/tools/gstack-fork/. The Windows real-file arm stays payload-gated: a real file has no provable owner. Absorbed from PR #2634 (2 commits squashed) with authorship preserved. The symmetric cleanup_prefixed_claude_symlinks hole is filed as a TODOS.md residual in this wave. Fixes #2204 * fix(redact): tolerate EEXIST from recursive mkdir in install-prepush-hook on bun/Windows (#2635) fs.mkdirSync(dir, { recursive: true }) is a no-op on an existing directory in Node, but bun on Windows throws EEXIST - crashing hook install on any repo whose .git/hooks already existed, leaving the repo unprotected. Add lib/fs-utils.ts mkdirpSync: swallow EEXIST only when statSync confirms the path is an existing directory; a regular file occupying the path, a stat failure, or any other errno still rethrows. Use it in installPrepushHook(). The regression test emulates the Windows bun fs semantics via a bun --preload fixture, so the exact crash path runs (and fails on the old code) on any platform, including CI Linux. Absorbed from PR #2641 with authorship preserved. Fixes #2635 * fix(bin): route remaining Windows-reachable mkdirSync sites through mkdirpSync Sweep follow-up to #2641's lib/fs-utils.ts helper: bun on Windows throws EEXIST from a recursive mkdir on an existing dir, so every unguarded recursive mkdirSync on a Windows-reachable path is a latent crash. Converted: bin/gstack-decision-log (unguarded, runs on every decision log — the second call on any machine hits the pre-existing projects dir), bin/gstack-evidence logsDir + ledger dir sites, and bin/gstack-redact-prepush's skip-log site (already try-wrapped, so its failure mode was a silent skip-log loss rather than a crash — the fix makes the log survive). The ~15 remaining gbrain/mac-lane sites are deliberately left alone. Regression: fs-utils.test.ts drives gstack-decision-log twice, the second run under the bun-Windows EEXIST preload fixture — the pre-sweep code exits 1 with EEXIST there; verified red against v1.68.3.0. * fix(setup-gbrain): warn about the ZeroEntropy sunset before Sept 4 ZeroEntropy was acquired by Notion and sunsets its hosted API on September 4, 2026. A gbrain configured with the zeroentropyai embedding recipe keeps importing pages after that date but embedding silently fails — pages land structurally with no semantic search, this repo's tracker P1 (TODOS.md NEXT PRIORITY). Nothing in gstack ever recommended ZeroEntropy (the dependency is gbrain-internal), so the gstack side is detection + advisory: the wireup helper warns when ~/.gbrain/config.json names the recipe (fail-open grep — a missing, unreadable, or other-provider config stays silent and never blocks a working setup), the setup-gbrain provider-default comments say never to select the legacy recipe for a new brain, and USING_GBRAIN_WITH_GSTACK.md gains a troubleshooting entry. The gbrain-side provider migration stays open upstream. Refs #2365 * fix(gbrain-source-wireup): first sync targets the registered source, not --repo The wireup registered a federated source by id, then ran 'gbrain sync --repo $WORKTREE' — which resolves against the brain's DEFAULT source and (on gbrain 0.46.x) rewrites that source's local_path anchor to our worktree. Net effect: the user's primary knowledge source silently repointed at the gstack brain worktree while the just-registered source got zero pages, and pages_synced still reported success. The sync now targets the registered id ('gbrain sync --source $id', the same form the repo's own troubleshooting documents). Because the script's stated floor is gbrain >= 0.18.0 and nothing proves --source exists there, support is probed via 'gbrain sync --help' first: an older gbrain keeps the wrong-but-working --repo call with an upgrade warning instead of converting it into a hard failure. The probe sits after the GSTACK_BRAIN_NO_SYNC early-exit and is unreachable in --probe mode. Regression tests (fail on v1.68.3.0): a no-skip sync case asserting the call log shows 'sync --source gstack-brain-<id>' and never 'sync --repo', and an old-gbrain fallback case (fake sync --help without --source) asserting --repo plus the upgrade warning. Fixes #2662 * fix(setup): --host slate exits informatively instead of silently installing nothing slate passed --host validation (added to the accept-list in v1.64.1.0) but never got a dispatch arm, and the all-INSTALL_*-zero fallback lives inside the auto branch — so './setup --host slate' configured nothing and exited 0, a silent no-op strictly worse than the original hard rejection. slate is now an informational arm (per docs/designs/SLATE_HOST.md it is blocked on the host-config refactor; Slate reads .claude/skills as a compatibility fallback, so the arm points at './setup --host claude'), and a defensive guard after the dispatch chain errors loudly (naming the host, the missing arm, and the valid targets, exit 1) if a future host is ever accepted without being wired. Regression tests (fail on v1.68.3.0): a dispatch-arm ratchet asserting every accept-listed install target has a matching dispatch branch — the exact drift class; a registry cross-check deriving both sides from hosts/index.ts and setup's case arms; a behavioral slate probe (exit 0, points at --host claude, never reaches the installer — on unfixed code it fell through into the installer); and a static pin on the guard's shape. Fixes #2361 * fix(make-pdf): resolve the sibling browse binary from execPath, not argv[0] In a bun-compiled binary process.argv[0] is the raw invocation string — often relative ('./pdf', 'pdf') — so dirname(argv[0]) yielded '.' and the sibling candidates (../browse/dist/browse etc.) resolved against the CWD instead of the install dir. Resolution was cwd-dependent: correct-by-luck when the fallbacks rescued it, wrong when a cwd-relative path matched. process.execPath is always the absolute binary path. The resolution step takes an injectable selfPath (defaulted) because under bun test the process path is the bun runtime and the compiled-binary shapes are otherwise unreachable. The issue's other half — pdf setup failing on newtab('about:blank') — was already fixed on main in v1.64.0.0 (browse/src/url-validation.ts exact-match allows about:blank; its comment names this exact smoke). This commit closes what remains. Regression tests (the sibling-via-selfPath case fails on v1.68.3.0 — pre-fix code ignores the seam and either resolves the global install or throws): sibling resolution from an install-shaped tree, and a decoy-browse-DIRECTORY case pinning that a directory never wins resolution. Fixes #2156 * fix(memory-ingest): store the normalized git_remote so unattributed pages hit the policy filter buildTranscriptPage wrote the normalized '_unattributed' sentinel into the page FRONTMATTER but stored the raw resolved remote ('' when unresolvable) on the page object. The policy filter fast-paths !p.git_remote, so under --include-unattributed an explicit '_unattributed → deny' (or read-only) policy never applied to exactly the pages it names — they ingested unpoliced. The stored value now matches the frontmatter. Regression test (fails on v1.68.3.0): seeds the REAL bin/gstack-gbrain-repo-policy store with '_unattributed → deny' through its own set verb, ingests an unresolvable-remote session with --include-unattributed, and asserts nothing reaches gbrain — pre-fix the '' remote bypassed the filter and the import ran. A fake echoing tiers would pass on both sides of the fix; the real helper prints 'none' for unknown keys, so only a genuinely applied deny distinguishes the two. Fixes #2353 * fix(land-and-deploy): MERGED recovery reconciles and reports remote-branch cleanup Step 4's merge commands carry --delete-branch, and the success path tells the user 'The branch has been cleaned up.' When gh exits non-zero AFTER GitHub already merged (routine in worktree layouts: gh's local cleanup runs git checkout <base> and fails), the §4a-postfail MERGED recovery re-established everything EXCEPT the branch deletion — and said nothing about it, so the discrepancy was invisible. The MERGED path now reconciles: git ls-remote --heads distinguishes branch-already-gone (exit 0, empty → 'already cleaned up', idempotent on re-runs) from branch-survived (offer confirm-first deletion, matching the section's worktree posture; -d not -D for any local branch) from check-itself-failed (non-zero exit → 'couldn't verify', skip the offer — never read a failed check as a clean branch). Template + regenerated SKILL.md + test extensions land in one commit (the md-sync assertion goes red otherwise). Regression assertions (fail on v1.68.3.0: no delete-branch reconciliation existed in test/ at all) pin the ls-remote check, the confirm-first delete, and the absent-vs-failed distinction. Fixes #2656 * fix(scripts): stop heredoc bodies deadlocking under Homebrew bash `./setup --help` can hang forever on macOS, printing nothing, with no way to tell it apart from a slow install. Eleven scripts carry the same latent hang, `setup` itself being the one every user hits first. bash 5.2+ delivers a heredoc body of 64KiB or less through a pipe: the forked child writes the entire body before exec, and nothing reads the other end until the command starts. Under macOS pipe-KVA pressure the kernel hands a fresh pipe a 512-byte buffer instead of the usual 16-64KiB, so any body of 512 bytes or more blocks write() permanently. The capacity check bash would need to notice (F_GETPIPE_SZ) is Linux-only, so it never fires here. It is pressure-dependent, which is why it reads as "worked on my machine" — the same script runs fine all day and then wedges. Homebrew bash is what `#!/usr/bin/env bash` resolves to on a Mac with brew on PATH, which is most of them. Apple's /bin/bash 3.2 predates the pipe path and is unaffected, so the bug is invisible to anyone testing with the system shell. The fix is `BASH_COMPAT=50` in each affected script, which restores the pre-5.2 tempfile path: $ bash -c 'probe() { [ -p /dev/stdin ] && echo PIPE || echo TEMPFILE; } probe <<EOF $(printf "x%.0s" $(seq 1 1000)) EOF' PIPE $ BASH_COMPAT=50 bash -c '...same...' TEMPFILE - Not a `#!/bin/bash` shebang swap: that pins the script to whatever bash lives at /bin (3.2 on macOS, absent on some Linux distributions) and is bypassed entirely by `bash script.sh` call sites. The variable survives both. - Not exported, so child processes keep their own compat level. - Placed below any `--help` sed range that reads $0, so usage output is unchanged (verified on all eleven). - Every guarded script is bash-3.2-clean — no associative arrays, case conversion, or mapfile — so compat level 50 costs them nothing. test/heredoc-pipe-deadlock.test.ts scans every tracked shell script for a heredoc body in the 512B-64KiB window and fails without the guard, and proves the mechanism at runtime on bash 5.2+ by asserting the body moves from PIPE to TEMPFILE. On older bash the runtime half is skipped, since the pipe path does not exist there. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Absorbed from PR #2640 with authorship preserved. Wave adaptations: the pipe-probe test skips on minimal-/dev environments without /dev/stdin (it would report OTHER for an unobservable fd), and one caveat verified during review: on bash 4.3/4.4 (e.g. Git Bash), assigning BASH_COMPAT=50 prints a non-fatal 'invalid value' warning to stderr — those bashes are already on tempfiles, so the guard is a no-op there; windows-setup-e2e exercises this empirically. * docs: TODOS.md v1.69 wave close-out Move the slate P4 entry and the ZeroEntropy P1's gstack-side half to Completed (v1.69.0.0); reframe the ZeroEntropy NEXT PRIORITY entry around the remaining gbrain-side work; file the wave's four residuals with rationale — the prefixed-cleanup symmetric conversion, the #2163 legacy-slug checkpoint heal, the invited #2657 --reconcile contribution, and the table-driven setup host dispatch behind the new cross-check ratchet. * chore: bump version and changelog (v1.69.0.0) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: Som Samantray <som.samantray@gmail.com> Co-authored-by: CommandCodeBot <noreply@commandcode.ai> Co-authored-by: Connex Client Access <paul@paulkortman.com> Co-authored-by: y$un_ <forrest.sun527@gmail.com> Co-authored-by: Lockyer <135391289+Lockyer228@users.noreply.github.com> Co-authored-by: Benjamin D. Smith <benjamin.smith@binarysword.com> Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
964 lines
41 KiB
Bash
Executable File
964 lines
41 KiB
Bash
Executable File
#!/usr/bin/env bash
|
||
# gstack-brain-sync — drain queue, commit allowlisted paths, push to remote.
|
||
#
|
||
# Usage:
|
||
# gstack-brain-sync --once drain queue, commit, push (default)
|
||
# gstack-brain-sync --status print sync health as JSON
|
||
# gstack-brain-sync --skip-file <p> add <p> to ~/.gstack/.brain-skip.txt
|
||
# gstack-brain-sync --drop-queue --yes clear queue without committing
|
||
# gstack-brain-sync --discover-new scan allowlist dirs, enqueue changed files
|
||
#
|
||
# Invoked by the preamble at skill START and END boundaries. No persistent
|
||
# daemon. Typical run <1s when queue empty; ~200-800ms with network push.
|
||
#
|
||
# Singleton enforcement: flock on ~/.gstack/.brain-sync.lock. Concurrent
|
||
# invocations queue and serialize.
|
||
#
|
||
# Env:
|
||
# GSTACK_HOME — override ~/.gstack (aligns with writers).
|
||
|
||
# Heredoc delivery guard. bash 5.2+ writes a heredoc body <=64KiB through a
|
||
# pipe in the forked child before exec, with no reader on the other end. On
|
||
# macOS under pipe-KVA pressure a fresh pipe gets a 512-byte buffer, so any
|
||
# body >=512B blocks write() forever and the script hangs at startup with no
|
||
# output. Compat level 50 restores the tempfile path. These scripts are
|
||
# bash-3.2-clean, so the compat level costs them nothing. Not exported: the
|
||
# guard is per-script, and it survives `bash script.sh` call sites that
|
||
# bypass the shebang.
|
||
BASH_COMPAT=50
|
||
|
||
set -uo pipefail
|
||
|
||
GSTACK_HOME="${GSTACK_HOME:-$HOME/.gstack}"
|
||
# Maildir-style spool: one FILE per record, <epoch>-<pid>-<uniq>.json.
|
||
# Writers (gstack-brain-enqueue, --discover-new) create records via tmp-file
|
||
# + atomic rename; the drain deletes exactly the files it snapshotted. No
|
||
# shared inode between writer and drainer → no append/rewrite race.
|
||
QUEUE_DIR="$GSTACK_HOME/.brain-queue.d"
|
||
# Legacy single-file queue: kept ONLY for migration. Pre-spool writers
|
||
# appended lines here; migrate_legacy_queue converts them to spool files.
|
||
QUEUE="$GSTACK_HOME/.brain-queue.jsonl"
|
||
ALLOWLIST="$GSTACK_HOME/.brain-allowlist"
|
||
PRIVACY_MAP="$GSTACK_HOME/.brain-privacy-map.json"
|
||
SKIP_FILE="$GSTACK_HOME/.brain-skip.txt"
|
||
STATUS_FILE="$GSTACK_HOME/.brain-sync-status.json"
|
||
LAST_PUSH_FILE="$GSTACK_HOME/.brain-last-push"
|
||
LOCK_FILE="$GSTACK_HOME/.brain-sync.lock"
|
||
DISCOVER_CURSOR="$GSTACK_HOME/.brain-discover-cursor"
|
||
|
||
SCRIPT_DIR="$(cd "$(dirname "$0")" && pwd)"
|
||
CONFIG_BIN="$SCRIPT_DIR/gstack-config"
|
||
|
||
# Egress receipt helpers (_receipted_git): receipt-before-send, fail-closed.
|
||
. "$SCRIPT_DIR/gstack-egress-lib.sh"
|
||
|
||
# origin host for receipt records (github.com etc).
|
||
remote_host() {
|
||
local url host
|
||
url=$(git -C "$GSTACK_HOME" remote get-url origin 2>/dev/null || echo "")
|
||
host="${url#*://}"; host="${host#*@}"; host="${host%%[/:]*}"
|
||
echo "${host:-unknown}"
|
||
}
|
||
|
||
# Remote-specific hint for auth errors (branch on origin URL).
|
||
remote_auth_hint() {
|
||
local url
|
||
url=$(git -C "$GSTACK_HOME" remote get-url origin 2>/dev/null || echo "")
|
||
case "$url" in
|
||
*github.com*|*@github.*) echo "run: gh auth status (and gh auth refresh if needed)" ;;
|
||
*gitlab*) echo "run: glab auth status" ;;
|
||
*) echo "check 'git remote -v' and your credentials" ;;
|
||
esac
|
||
}
|
||
|
||
write_status() {
|
||
# args: status_code message [extra_json_blob]
|
||
local code="$1"
|
||
local msg="$2"
|
||
local extra="${3:-{\}}"
|
||
local ts
|
||
ts=$(date -u +%Y-%m-%dT%H:%M:%SZ 2>/dev/null || echo "")
|
||
python3 - "$STATUS_FILE" "$code" "$msg" "$ts" "$extra" <<'PYEOF' 2>/dev/null || true
|
||
import json, sys
|
||
path, code, msg, ts, extra = sys.argv[1:6]
|
||
try:
|
||
extra_obj = json.loads(extra) if extra else {}
|
||
except Exception:
|
||
extra_obj = {}
|
||
data = {"status": code, "message": msg, "ts": ts, **extra_obj}
|
||
with open(path, "w") as f:
|
||
json.dump(data, f)
|
||
f.write("\n")
|
||
PYEOF
|
||
}
|
||
|
||
# Read config; return 0 if sync active, 1 otherwise.
|
||
sync_active() {
|
||
if [ ! -d "$GSTACK_HOME/.git" ]; then
|
||
return 1
|
||
fi
|
||
local mode
|
||
mode=$("$CONFIG_BIN" get artifacts_sync_mode 2>/dev/null || echo off)
|
||
[ "$mode" = "off" ] && return 1
|
||
return 0
|
||
}
|
||
|
||
# Secret regex families — stdin scan. Exits 0 clean, 1 if hit.
|
||
# Echoes the matching pattern family name on hit. Uses python3 -c (not
|
||
# heredoc) so sys.stdin stays available for the diff content.
|
||
secret_scan_stdin() {
|
||
python3 -c "
|
||
import sys, re
|
||
patterns = [
|
||
('aws-access-key', re.compile(r'AKIA[0-9A-Z]{16}')),
|
||
('github-token', re.compile(r'\\b(gh[pousr]_[A-Za-z0-9]{20,}|github_pat_[A-Za-z0-9_]{20,})')),
|
||
('openai-key', re.compile(r'\\bsk-[A-Za-z0-9_-]{20,}')),
|
||
('pem-block', re.compile(r'-----BEGIN [A-Z ]{3,}-----')),
|
||
('jwt', re.compile(r'\\beyJ[A-Za-z0-9_-]{10,}\\.[A-Za-z0-9_-]{10,}\\.[A-Za-z0-9_-]{10,}\\b')),
|
||
('bearer-token-json',
|
||
# JSON-embedded auth headers. The optional Bearer/Basic/Token prefix
|
||
# matters: real auth values include a literal space after the scheme
|
||
# name, but the value charset below does not include spaces, so
|
||
# without the optional prefix every Bearer token in a JSON blob slips
|
||
# past the scanner.
|
||
re.compile(r'\"(authorization|api[_-]?key|apikey|token|secret|password)\"\\s*:\\s*\"(Bearer |Basic |Token )?[A-Za-z0-9_./+=-]{16,}\"',
|
||
re.IGNORECASE)),
|
||
]
|
||
text = sys.stdin.read()
|
||
for name, rx in patterns:
|
||
m = rx.search(text)
|
||
if m:
|
||
snippet = m.group(0)
|
||
if len(snippet) > 30:
|
||
snippet = snippet[:30] + '...'
|
||
print(name + ':' + snippet)
|
||
sys.exit(1)
|
||
sys.exit(0)
|
||
"
|
||
}
|
||
|
||
# True (0) if the spool holds at least one record file.
|
||
spool_has_records() {
|
||
local f
|
||
for f in "$QUEUE_DIR"/*.json; do
|
||
[ -e "$f" ] && return 0
|
||
done
|
||
return 1
|
||
}
|
||
|
||
# Convert one legacy queue file's lines into spool record files (tmp +
|
||
# os.replace, one file per line). Reads the file TWICE before unlinking: a
|
||
# pre-rename writer can still append through its already-open fd after our
|
||
# rename, and those appends land in the renamed file — the second pass
|
||
# NARROWS the tail-race window (transition-only: it applies to pre-spool
|
||
# writers, and a writer that appends after the second read but before the
|
||
# unlink can still lose that line; spool-native writers are immune).
|
||
# Unparseable lines migrate as-is; finalize_queue quarantines + warns on them.
|
||
convert_legacy_file() {
|
||
local legacy="$1"
|
||
python3 - "$legacy" "$QUEUE_DIR" <<'PYEOF' 2>/dev/null || true
|
||
import os, sys, time
|
||
|
||
legacy, spool = sys.argv[1:3]
|
||
|
||
def read_lines(path):
|
||
try:
|
||
with open(path) as f:
|
||
return [l.rstrip("\r\n") for l in f if l.strip()]
|
||
except (FileNotFoundError, OSError):
|
||
return []
|
||
|
||
seq = 0
|
||
def write_spool(line):
|
||
global seq
|
||
seq += 1
|
||
tmp = os.path.join(spool, f".tmp-{os.getpid()}-m{seq}")
|
||
with open(tmp, "w") as f:
|
||
f.write(line + "\n")
|
||
os.replace(tmp, os.path.join(spool, f"{int(time.time())}-{os.getpid()}-m{seq}.json"))
|
||
|
||
written = set()
|
||
for _pass in (1, 2): # second read closes the pre-rename-fd tail race
|
||
for line in read_lines(legacy):
|
||
if line not in written: # identical duplicates collapse, as the old rewrite did
|
||
write_spool(line)
|
||
written.add(line)
|
||
os.unlink(legacy)
|
||
PYEOF
|
||
}
|
||
|
||
# Legacy migration (transition window only). If the single-file queue holds
|
||
# records, atomically rename it aside and convert each line to a spool file.
|
||
# A concurrent OLD writer that recreates a fresh legacy file after the rename
|
||
# simply gets migrated on the NEXT drain — nothing is lost, only deferred one
|
||
# boundary. Runs inside the run lock, before the drain reads the spool.
|
||
migrate_legacy_queue() {
|
||
local migrating="$QUEUE.migrating"
|
||
# Crash leftover: a prior migration renamed but died before unlink. Some of
|
||
# its lines may already exist as spool files — re-converting duplicates is
|
||
# safe (at-least-once; the drain dedups paths per snapshot and downstream
|
||
# content-hash dedup absorbs re-syncs). Losing the file would not be. If
|
||
# the conversion itself fails, the file stays for the next run (never rm a
|
||
# non-empty .migrating file outside convert_legacy_file's own unlink).
|
||
if [ -f "$migrating" ]; then
|
||
if [ -s "$migrating" ]; then
|
||
mkdir -p "$QUEUE_DIR" 2>/dev/null || return 0
|
||
convert_legacy_file "$migrating"
|
||
else
|
||
rm -f "$migrating" 2>/dev/null || true
|
||
fi
|
||
fi
|
||
# If the leftover STILL holds records, the conversion failed (e.g. python3
|
||
# unavailable). The mv below would overwrite it and destroy those records —
|
||
# exactly the never-destroy invariant above. Defer this run's migration;
|
||
# the next run retries both files.
|
||
[ -s "$migrating" ] && return 0
|
||
if [ -s "$QUEUE" ]; then
|
||
mkdir -p "$QUEUE_DIR" 2>/dev/null || return 0
|
||
mv -f "$QUEUE" "$migrating" 2>/dev/null || return 0
|
||
convert_legacy_file "$migrating"
|
||
fi
|
||
return 0
|
||
}
|
||
|
||
# Compute matched allowlisted, privacy-filtered path set from the spool.
|
||
# Output: newline-delimited relative paths that should be staged.
|
||
#
|
||
# #2549: every non-staged queue entry is CLASSIFIED, never silently discarded.
|
||
# When $2 is given, a JSON classification lands there:
|
||
# {"retained": [privacy/mode-held paths that stay queued],
|
||
# "dropped": {"skipped": [...], "invalid": [...], "unmatched": [...], "missing": [...]}}
|
||
# retained entries would sync if the user raises artifacts_sync_mode, so they
|
||
# stay in the queue; dropped classes can never sync (explicit skip, escape
|
||
# attempt, no allowlist glob, not on disk) and are removed WITH a counted
|
||
# status — the old behavior truncated the whole queue and reported every one
|
||
# of these, including privacy holds, as "no allowlisted changes".
|
||
#
|
||
# Spool snapshot ($3): the sorted list of spool record filenames read here is
|
||
# written to the snapshot manifest, one filename per line. finalize_queue
|
||
# deletes exactly the manifest's files and never touches records created
|
||
# after this listing — a concurrent enqueue is a separate file by
|
||
# construction, so it simply rides to the next drain.
|
||
compute_paths_to_stage() {
|
||
local mode="$1"
|
||
local class_file="${2:-}"
|
||
local snapshot_file="${3:-}"
|
||
python3 - "$GSTACK_HOME" "$QUEUE_DIR" "$ALLOWLIST" "$PRIVACY_MAP" "$SKIP_FILE" "$mode" "$class_file" "$snapshot_file" <<'PYEOF'
|
||
import sys, json, os, fnmatch, glob
|
||
|
||
gstack_home, spool_dir, allowlist_path, privacy_path, skip_path, mode, class_file, snapshot_file = sys.argv[1:9]
|
||
|
||
def load_lines(path):
|
||
try:
|
||
with open(path) as f:
|
||
return [l.strip() for l in f if l.strip() and not l.lstrip().startswith("#")]
|
||
except FileNotFoundError:
|
||
return []
|
||
|
||
def load_privacy_map(path):
|
||
# Returns (entries, corrupt). Non-dict entries are filtered out
|
||
# defensively — the map may be PULLED from the artifacts remote, so a
|
||
# malformed entry like ["bad"] is remotely triggerable and used to raise
|
||
# mid-classification (after the snapshot manifest was written), which the
|
||
# old finalize turned into a full queue wipe. Any malformed shape also
|
||
# marks the map CORRUPT: privacy classification cannot be trusted, so the
|
||
# caller holds every queued record instead of guessing (a corrupt privacy
|
||
# map silently treated as empty would over-share behavioral data).
|
||
try:
|
||
with open(path) as f:
|
||
data = json.load(f)
|
||
except FileNotFoundError:
|
||
return [], False
|
||
except json.JSONDecodeError:
|
||
return [], True
|
||
if not isinstance(data, list):
|
||
return [], True
|
||
# Expected: [{"pattern": "glob", "class": "artifact" | "behavioral"}]
|
||
entries = [e for e in data if isinstance(e, dict)]
|
||
return entries, len(entries) != len(data)
|
||
|
||
allowlist_globs = load_lines(allowlist_path)
|
||
privacy_map, privacy_corrupt = load_privacy_map(privacy_path)
|
||
# Normalize skip entries to the POSIX form queued paths use, so a backslash
|
||
# entry in .brain-skip.txt still matches on Windows. The drain is the safety
|
||
# boundary that actually stages files, so it must normalize identically to
|
||
# discover_new — otherwise an explicitly-skipped file gets committed.
|
||
skip_lines = {s.replace(os.sep, "/") for s in load_lines(skip_path)}
|
||
|
||
# Snapshot the spool: sorted (= chronological, filenames are epoch-first)
|
||
# list of record files at read time. Records that appear after this listing
|
||
# belong to the NEXT drain. Files we cannot read stay OUT of the manifest so
|
||
# finalize never deletes a record this drain didn't actually consume.
|
||
try:
|
||
snapshot = sorted(n for n in os.listdir(spool_dir) if n.endswith(".json"))
|
||
except (FileNotFoundError, NotADirectoryError):
|
||
snapshot = []
|
||
|
||
queue_paths = set()
|
||
consumed = []
|
||
for name in snapshot:
|
||
try:
|
||
with open(os.path.join(spool_dir, name)) as f:
|
||
line = f.readline().strip()
|
||
except OSError:
|
||
continue
|
||
consumed.append(name)
|
||
if not line:
|
||
continue
|
||
try:
|
||
obj = json.loads(line)
|
||
p = obj.get("file")
|
||
if isinstance(p, str):
|
||
queue_paths.add(p)
|
||
except json.JSONDecodeError:
|
||
continue # unparseable record: finalize keeps + warns
|
||
|
||
if snapshot_file:
|
||
with open(snapshot_file, "w") as f:
|
||
for name in consumed:
|
||
f.write(name + "\n")
|
||
|
||
def path_matches_any(path, globs):
|
||
for pattern in globs:
|
||
if fnmatch.fnmatchcase(path, pattern):
|
||
return True
|
||
return False
|
||
|
||
def privacy_class(path, mapping):
|
||
for entry in mapping:
|
||
pat = entry.get("pattern")
|
||
if pat and fnmatch.fnmatchcase(path, pat):
|
||
return entry.get("class", "artifact")
|
||
# Default class when no pattern matches: artifact (safe default).
|
||
return "artifact"
|
||
|
||
# mode filter: 'off' → nothing; 'artifacts-only' → only artifact class;
|
||
# 'full' → both classes.
|
||
def mode_allows(cls, mode):
|
||
if mode == "off":
|
||
return False
|
||
if mode == "artifacts-only":
|
||
return cls == "artifact"
|
||
return True # full
|
||
|
||
final = []
|
||
classified = {"retained": [], "dropped": {"skipped": [], "invalid": [], "unmatched": [], "missing": []}}
|
||
if privacy_corrupt:
|
||
# Fail-safe: with an untrustworthy privacy map, stage NOTHING and drop
|
||
# NOTHING — retain every queued record until the map is fixed. The next
|
||
# drain re-classifies from scratch.
|
||
print("BRAIN_SYNC: warning: privacy map at " + privacy_path +
|
||
" is malformed — holding all queued records until it is fixed", file=sys.stderr)
|
||
classified["retained"] = sorted(queue_paths)
|
||
queue_paths = set()
|
||
for p in sorted(queue_paths):
|
||
if p in skip_lines:
|
||
classified["dropped"]["skipped"].append(p)
|
||
continue
|
||
# Must be under GSTACK_HOME root. Reject absolute + reject ../ escape.
|
||
if p.startswith("/") or ".." in p.split("/"):
|
||
classified["dropped"]["invalid"].append(p)
|
||
continue
|
||
# Must match at least one allowlist glob.
|
||
if not path_matches_any(p, allowlist_globs):
|
||
classified["dropped"]["unmatched"].append(p)
|
||
continue
|
||
# Must survive privacy mode filter — held entries STAY QUEUED (retained):
|
||
# they would sync under a higher artifacts_sync_mode, and reporting them
|
||
# as "no allowlisted changes" was #2549's misattribution.
|
||
cls = privacy_class(p, privacy_map)
|
||
if not mode_allows(cls, mode):
|
||
classified["retained"].append(p)
|
||
continue
|
||
# Must exist on disk — can't stage what isn't there.
|
||
if not os.path.exists(os.path.join(gstack_home, p)):
|
||
classified["dropped"]["missing"].append(p)
|
||
continue
|
||
final.append(p)
|
||
|
||
if class_file:
|
||
with open(class_file, "w") as f:
|
||
json.dump(classified, f)
|
||
|
||
for p in final:
|
||
print(p)
|
||
PYEOF
|
||
}
|
||
|
||
# Finalize the drain: delete exactly the spool record files this drain
|
||
# consumed (per the snapshot manifest) AND positively classified. Deletion is
|
||
# EXPLICIT-DELETE-ONLY: a record is unlinked only when its path appears in
|
||
# (staged paths ∪ classified dropped). The old polarity ("delete unless
|
||
# retained") turned a missing/unparseable classification into retained=∅ and
|
||
# wiped every snapshotted record — remotely triggerable via a malformed
|
||
# pulled privacy map that raised AFTER the manifest write. Now a
|
||
# missing/unparseable class_file or paths_file deletes NOTHING (warn +
|
||
# return), and a path the classification never mentions stays queued.
|
||
# The predecessor (a shared-file queue rewrite) had a lockless-append race
|
||
# between its live re-read and the os.replace; with one file per record that
|
||
# race class is structurally gone — a concurrent enqueue is a separate file
|
||
# the snapshot never listed, so finalize cannot touch it. Crash semantics are
|
||
# at-least-once: a drain that dies before finalize leaves its spool files in
|
||
# place and the next run re-drains them; downstream content-hash dedup
|
||
# absorbs the duplicates. Unparseable records move to $QUEUE_DIR/quarantine/
|
||
# (never deleted) so they stop re-warning at every boundary. Dropped-path
|
||
# detail goes to a 0600 sidecar so the status line can stay content-free
|
||
# (counts only).
|
||
finalize_queue() {
|
||
local snapshot_file="$1" # spool filenames this drain consumed, one per line
|
||
local class_file="$2" # classification JSON from compute_paths_to_stage
|
||
local paths_file="$3" # staged paths (compute_paths_to_stage stdout), one per line
|
||
# Fail-open by design (a failed finalize self-corrects next run: re-stage →
|
||
# nothing-to-commit), but say so — a silent failure here would let the
|
||
# subsequent "ok/idle" status claim a drain that did not happen.
|
||
python3 - "$QUEUE_DIR" "$snapshot_file" "$class_file" "$paths_file" "$GSTACK_HOME/.brain-sync-drops.json" <<'PYEOF' || echo "BRAIN_SYNC: warning: queue finalize failed — entries retained; next run re-drains" >&2
|
||
import json, os, sys, time
|
||
spool_dir, snapshot_file, class_file, paths_file, drops_file = sys.argv[1:6]
|
||
|
||
def lines(path):
|
||
try:
|
||
with open(path) as f:
|
||
return [l.rstrip("\r\n") for l in f if l.strip()]
|
||
except FileNotFoundError:
|
||
return []
|
||
|
||
# Explicit-delete-only inputs. Either input unreadable → delete NOTHING.
|
||
try:
|
||
with open(class_file) as f:
|
||
classified = json.load(f)
|
||
if not isinstance(classified, dict):
|
||
raise ValueError("classification is not an object")
|
||
except Exception:
|
||
print("BRAIN_SYNC: warning: classification unreadable — no queue records deleted; next run re-drains", file=sys.stderr)
|
||
sys.exit(0)
|
||
try:
|
||
with open(paths_file) as f:
|
||
staged = {l.strip() for l in f if l.strip()}
|
||
except Exception:
|
||
print("BRAIN_SYNC: warning: staged-paths file unreadable — no queue records deleted; next run re-drains", file=sys.stderr)
|
||
sys.exit(0)
|
||
|
||
dropped = set()
|
||
for group in (classified.get("dropped", {}) or {}).values():
|
||
dropped.update(group)
|
||
deletable = staged | dropped
|
||
|
||
unparseable = 0
|
||
for name in lines(snapshot_file):
|
||
full = os.path.join(spool_dir, name)
|
||
try:
|
||
with open(full) as f:
|
||
rec = f.readline().strip()
|
||
except OSError:
|
||
continue # unreadable now: leave it for the next drain
|
||
p = None
|
||
try:
|
||
p = json.loads(rec).get("file")
|
||
except Exception:
|
||
pass
|
||
if not isinstance(p, str):
|
||
# Never destroy what we can't read — but don't leave it re-warning at
|
||
# every boundary either: move it aside for inspection.
|
||
unparseable += 1
|
||
try:
|
||
qdir = os.path.join(spool_dir, "quarantine")
|
||
os.makedirs(qdir, exist_ok=True)
|
||
os.replace(full, os.path.join(qdir, name))
|
||
except OSError:
|
||
pass # quarantine move failed — leave in place; next run retries
|
||
continue
|
||
if p not in deletable:
|
||
continue # retained / unclassified: stays queued (explicit-delete-only)
|
||
try:
|
||
os.unlink(full) # staged or dropped: fully processed
|
||
except FileNotFoundError:
|
||
pass
|
||
if unparseable:
|
||
print(f"BRAIN_SYNC: {unparseable} unparseable spool record(s) moved to quarantine (inspect {os.path.join(spool_dir, 'quarantine')})", file=sys.stderr)
|
||
|
||
if dropped:
|
||
fd = os.open(drops_file, os.O_WRONLY | os.O_CREAT | os.O_TRUNC, 0o600)
|
||
with os.fdopen(fd, "w") as f:
|
||
json.dump({"ts": time.strftime("%Y-%m-%dT%H:%M:%SZ", time.gmtime()),
|
||
"dropped": classified.get("dropped", {})}, f)
|
||
PYEOF
|
||
}
|
||
|
||
# Human-readable classification counts for status messages.
|
||
queue_summary() {
|
||
local class_file="$1"
|
||
python3 - "$class_file" <<'PYEOF' 2>/dev/null || echo ""
|
||
import json, sys
|
||
try:
|
||
with open(sys.argv[1]) as f:
|
||
c = json.load(f)
|
||
except Exception:
|
||
print(""); sys.exit(0)
|
||
d = c.get("dropped", {}) or {}
|
||
parts = []
|
||
r = len(c.get("retained", []))
|
||
if r: parts.append(f"{r} privacy-held retained")
|
||
for k in ("skipped", "unmatched", "missing", "invalid"):
|
||
n = len(d.get(k, []))
|
||
if n: parts.append(f"{n} {k} dropped")
|
||
print("; ".join(parts))
|
||
PYEOF
|
||
}
|
||
|
||
subcmd_once() {
|
||
if ! sync_active; then
|
||
# Silent no-op when feature not initialized / disabled.
|
||
exit 0
|
||
fi
|
||
|
||
# Singleton lock via atomic mkdir. `flock(1)` isn't on macOS by default;
|
||
# `mkdir` is atomic on every POSIX filesystem. If another --once is already
|
||
# running, skip (don't wait) — the next skill boundary will catch up.
|
||
local lock_dir="${LOCK_FILE}.d"
|
||
if ! mkdir "$lock_dir" 2>/dev/null; then
|
||
# Is the lock stale? Check the pidfile inside. If process is dead, clear it.
|
||
if [ -f "$lock_dir/pid" ]; then
|
||
local lock_pid
|
||
lock_pid=$(cat "$lock_dir/pid" 2>/dev/null || echo "")
|
||
if [ -n "$lock_pid" ] && ! kill -0 "$lock_pid" 2>/dev/null; then
|
||
# Stale lock — clear and retry once.
|
||
rm -rf "$lock_dir" 2>/dev/null || true
|
||
if ! mkdir "$lock_dir" 2>/dev/null; then
|
||
exit 0
|
||
fi
|
||
else
|
||
# Lock is held by a live process.
|
||
exit 0
|
||
fi
|
||
else
|
||
# Lock dir without pidfile — treat as held; don't touch.
|
||
exit 0
|
||
fi
|
||
fi
|
||
echo "$$" > "$lock_dir/pid" 2>/dev/null || true
|
||
# Release the lock on EVERY exit from here on — including the empty-queue
|
||
# fast path and an INT during the detector's network push. Leaking it would
|
||
# rely on next-run stale-pid detection, which PID reuse can defeat (kill -0
|
||
# matching an unrelated live process wedges sync at every boundary). The
|
||
# mktemp block below re-traps with tempfile cleanup added; both traps keep
|
||
# the lock removal.
|
||
trap 'rm -rf "$lock_dir" 2>/dev/null || true' EXIT INT TERM
|
||
|
||
# Convert any legacy single-file queue lines into spool records before the
|
||
# drain reads the spool (transition window for pre-spool writers).
|
||
migrate_legacy_queue
|
||
|
||
# Janitor: reap orphaned enqueue temp files. A writer killed between its
|
||
# tmp write and the atomic rename leaves `.tmp-*` behind forever — it never
|
||
# becomes a record and nothing else touches it. One hour is far beyond any
|
||
# live writer's write→rename window, so a fresh tmp (an in-flight enqueue)
|
||
# is never touched. Runs inside the run lock, so it can't race the drain.
|
||
find "$QUEUE_DIR" -maxdepth 1 -type f -name '.tmp-*' -mmin +60 -delete 2>/dev/null || true
|
||
|
||
local mode
|
||
mode=$("$CONFIG_BIN" get artifacts_sync_mode 2>/dev/null || echo off)
|
||
|
||
# #2516: advance the brain worktree gbrain indexes to the artifacts repo's
|
||
# HEAD once a day — previously it only moved when setup-gbrain / sync-gbrain
|
||
# / brain-restore ran, so brains silently served stale code forever. Runs
|
||
# inside THIS run lock (never concurrent with the ingest steps below) and
|
||
# before they touch the worktree. Attempt-throttled: the stamp is written on
|
||
# ATTEMPT, so a persistently-failing advance warns once per 24h, not at
|
||
# every skill boundary. The advance itself refuses dirty or unmanaged
|
||
# worktrees and never force-removes (see gstack-gbrain-source-wireup).
|
||
if [ -e "${GSTACK_BRAIN_WORKTREE:-$HOME/.gstack-brain-worktree}" ]; then
|
||
local adv_stamp adv_now adv_last adv_age
|
||
adv_stamp="$GSTACK_HOME/.brain-worktree-last-advance"
|
||
adv_now=$(date +%s)
|
||
adv_last=$(cat "$adv_stamp" 2>/dev/null || echo 0)
|
||
case "$adv_last" in ''|*[!0-9]*) adv_last=0 ;; esac
|
||
adv_age=$(( adv_now - adv_last ))
|
||
if [ "$adv_age" -ge 86400 ]; then
|
||
echo "$adv_now" > "$adv_stamp" 2>/dev/null || true
|
||
if ! "$SCRIPT_DIR/gstack-gbrain-source-wireup" --advance-only 1>&2; then
|
||
echo "BRAIN_SYNC: warning: brain worktree advance failed — gbrain may be indexing stale code (run gstack-gbrain-source-wireup to repair)" >&2
|
||
fi
|
||
fi
|
||
fi
|
||
|
||
# #2549 unpushed-commit detector: a prior drain may have COMMITTED but
|
||
# failed to push (auth blip, offline). The data was never lost — it sits in
|
||
# a local commit — but nothing re-pushed it until NEW changes arrived.
|
||
# Retry the push up front, inside the lock. Receipted fail-closed like
|
||
# every other push; a receipt REFUSAL skips the retry without blocking the
|
||
# rest of the drain (local staging must not wedge on receipt problems).
|
||
# Guards: origin/<branch> may not exist yet (first sync, deleted remote).
|
||
#
|
||
# Throttled: the preamble runs --once at EVERY skill boundary, so an
|
||
# unthrottled retry would pay a full network push attempt per boundary in
|
||
# exactly the steady states this targets (offline, broken auth) — and a
|
||
# captive-portal push can block 30-75s against the header's "<1s when
|
||
# idle" promise. Attempts are recorded (success or fail) and retried at
|
||
# most every 10 minutes; the push itself never prompts for credentials and
|
||
# bounds stalled transfers via git's own low-speed limits (portable — stock
|
||
# macOS ships no `timeout` binary).
|
||
#
|
||
# Author-scoped — EXCLUSIVELY: `git push origin HEAD` publishes every
|
||
# unpushed commit, so the retry fires only when ALL unpushed commits are
|
||
# gstack-brain-sync's own. One interleaved user commit disables the
|
||
# auto-retry entirely (adversarial review: an existential check would
|
||
# silently auto-publish a user's manual ~/.gstack commit the moment a bot
|
||
# commit sat in front of it). User commits ride along when a REAL drain
|
||
# pushes, as before — the detector never publishes work it didn't create.
|
||
local det_branch det_unpushed det_total det_now det_last
|
||
det_branch=$(git -C "$GSTACK_HOME" rev-parse --abbrev-ref HEAD 2>/dev/null || echo "")
|
||
# Detached HEAD reads as the literal "HEAD" — origin/HEAD usually resolves,
|
||
# so without this exclusion the detector would retry a doomed push forever.
|
||
[ "$det_branch" = "HEAD" ] && det_branch=""
|
||
if [ -n "$det_branch" ] && git -C "$GSTACK_HOME" rev-parse --verify --quiet "origin/$det_branch" >/dev/null 2>&1; then
|
||
det_unpushed=$(git -C "$GSTACK_HOME" rev-list --count --author="gstack-brain-sync" "origin/$det_branch..HEAD" 2>/dev/null || echo 0)
|
||
det_total=$(git -C "$GSTACK_HOME" rev-list --count "origin/$det_branch..HEAD" 2>/dev/null || echo 0)
|
||
case "$det_unpushed" in ''|*[!0-9]*) det_unpushed=0 ;; esac
|
||
case "$det_total" in ''|*[!0-9]*) det_total=0 ;; esac
|
||
det_now=$(date +%s)
|
||
det_last=$(cat "$GSTACK_HOME/.brain-last-push-attempt" 2>/dev/null || echo 0)
|
||
case "$det_last" in ''|*[!0-9]*) det_last=0 ;; esac
|
||
if [ "$det_unpushed" -gt 0 ] && [ "$det_unpushed" -eq "$det_total" ] && [ $(( det_now - det_last )) -ge 600 ]; then
|
||
echo "$det_now" > "$GSTACK_HOME/.brain-last-push-attempt" 2>/dev/null || true
|
||
local det_host
|
||
det_host=$(remote_host)
|
||
if GSTACK_HOME="$GSTACK_HOME" _receipted_git closed brain-sync "$det_host" curated-memory-git-push "artifacts_sync_mode!=off" \
|
||
bash -c 'GIT_TERMINAL_PROMPT=0 git -c http.lowSpeedLimit=1024 -c http.lowSpeedTime=30 -C "$1" push origin HEAD 2>/dev/null' _ "$GSTACK_HOME"; then
|
||
date -u +%Y-%m-%dT%H:%M:%SZ > "$LAST_PUSH_FILE"
|
||
fi
|
||
fi
|
||
fi
|
||
|
||
# Empty-queue fast path: this is the steady state at every skill boundary.
|
||
# Skipping compute/finalize here is safe — with zero spool records there is
|
||
# nothing to classify, retain, or drop, and a record created after this
|
||
# check simply waits for the next boundary. The legacy file is checked too:
|
||
# an OLD writer may have recreated it after the migration above (it gets
|
||
# migrated next run, but the depth is honest now) — and so is a leftover
|
||
# .migrating file: if its conversion failed above (e.g. python3 missing),
|
||
# records are still pending, so "idle" would be dishonest. (The detector
|
||
# above already ran: its whole point is re-pushing stranded commits when
|
||
# the queue is empty.) The lock-release trap installed at acquisition
|
||
# covers this exit.
|
||
if ! spool_has_records && [ ! -s "$QUEUE" ] && [ ! -s "$QUEUE.migrating" ]; then
|
||
write_status "idle" "queue empty"
|
||
exit 0
|
||
fi
|
||
|
||
local paths_file class_file snapshot_file
|
||
paths_file=$(mktemp /tmp/brain-sync-paths.XXXXXX) || { rm -rf "$lock_dir" 2>/dev/null; write_status "error" "mktemp failed"; exit 1; }
|
||
class_file=$(mktemp /tmp/brain-sync-class.XXXXXX) || { rm -f "$paths_file"; rm -rf "$lock_dir" 2>/dev/null; write_status "error" "mktemp failed"; exit 1; }
|
||
snapshot_file=$(mktemp /tmp/brain-sync-snapshot.XXXXXX) || { rm -f "$paths_file" "$class_file"; rm -rf "$lock_dir" 2>/dev/null; write_status "error" "mktemp failed"; exit 1; }
|
||
# Single trap covers all: lock cleanup AND tempfile cleanup.
|
||
trap 'rm -f "$paths_file" "$class_file" "$snapshot_file" 2>/dev/null; rm -rf "$lock_dir" 2>/dev/null || true' EXIT INT TERM
|
||
|
||
# Fail-safe (G1): a classifier that dies mid-run (ENOSPC/OOM/SIGKILL, or a
|
||
# shape the defensive filters don't cover) may have already written the
|
||
# snapshot manifest but no classification. Finalizing on that state is what
|
||
# used to wipe the queue — so on a nonzero exit, warn loudly, do NOT call
|
||
# finalize_queue, and leave everything queued for the next drain.
|
||
if ! compute_paths_to_stage "$mode" "$class_file" "$snapshot_file" > "$paths_file"; then
|
||
echo "BRAIN_SYNC: warning: queue classification failed — no records consumed; next run re-drains" >&2
|
||
write_status "error" "classification failed; queue preserved (next run retries)"
|
||
exit 0
|
||
fi
|
||
if [ ! -s "$paths_file" ]; then
|
||
# Nothing stageable. Finalize the snapshot (retained entries survive;
|
||
# classified drops removed; records created after the snapshot untouched).
|
||
finalize_queue "$snapshot_file" "$class_file" "$paths_file"
|
||
local summary
|
||
summary=$(queue_summary "$class_file")
|
||
write_status "idle" "no stageable changes${summary:+ ($summary)}"
|
||
exit 0
|
||
fi
|
||
|
||
# Stage with git add -f (forces past .gitignore=*) explicit paths only.
|
||
while IFS= read -r p; do
|
||
p="${p%$'\r'}" # Windows: compute_paths_to_stage's python print() emits CRLF;
|
||
# a trailing CR makes the pathspec match nothing (silent no-stage).
|
||
[ -z "$p" ] && continue
|
||
git -C "$GSTACK_HOME" add -f -- "$p" 2>/dev/null || true
|
||
done < "$paths_file"
|
||
|
||
# Secret-scan staged diff.
|
||
local scan_out
|
||
scan_out=$(git -C "$GSTACK_HOME" diff --cached 2>/dev/null | secret_scan_stdin || true)
|
||
if [ -n "$scan_out" ]; then
|
||
# Hit — unstage, preserve queue, write loud status.
|
||
git -C "$GSTACK_HOME" reset HEAD -- . >/dev/null 2>&1 || true
|
||
local hint
|
||
hint="secret pattern detected ($scan_out). Remediation: review the staged file, then run: gstack-brain-sync --skip-file <path> OR edit the content."
|
||
write_status "blocked" "$hint"
|
||
echo "BRAIN_SYNC: blocked: $scan_out" >&2
|
||
exit 0
|
||
fi
|
||
|
||
# Egress receipt for the push, written BEFORE the commit consumes the
|
||
# queue (amendment C7 ordering): a refused receipt exits HERE, before any
|
||
# queue mutation or local commit, so the queue stays intact and the next
|
||
# run retries the whole drain. Content-free: git owns the bytes
|
||
# (sha256:null). Fail-closed.
|
||
local push_host receipt_err
|
||
push_host=$(remote_host)
|
||
if ! receipt_err=$(GSTACK_HOME="$GSTACK_HOME" "$SCRIPT_DIR/gstack-egress-receipt" write \
|
||
--sink brain-sync --host "$push_host" --class curated-memory-git-push \
|
||
--no-payload --consent "artifacts_sync_mode!=off" 2>&1 >/dev/null); then
|
||
write_status "push_failed" "EGRESS_RECEIPT_FAILED: receipt not writable; push refused (queue preserved)"
|
||
_gstack_egress_refusal "brain-sync push" "$(printf '%s' "$receipt_err" | head -c 300)"
|
||
exit 1
|
||
fi
|
||
|
||
# Commit with template message.
|
||
local n ts
|
||
n=$(wc -l < "$paths_file" | tr -d ' ')
|
||
ts=$(date -u +%Y-%m-%dT%H:%M:%SZ)
|
||
local msg="sync: $n file(s) | $ts"
|
||
git -C "$GSTACK_HOME" -c user.email="gstack@localhost" -c user.name="gstack-brain-sync" \
|
||
commit -q -m "$msg" 2>/dev/null || {
|
||
# Nothing to commit (e.g. all files already committed). The drained
|
||
# records leave the spool; retained + post-snapshot records survive.
|
||
finalize_queue "$snapshot_file" "$class_file" "$paths_file"
|
||
write_status "idle" "queue drained but no new changes to commit"
|
||
exit 0
|
||
}
|
||
|
||
# Push. On reject, fetch + merge (merge driver handles JSONL) + retry once.
|
||
local push_err
|
||
push_err=$(git -C "$GSTACK_HOME" push origin HEAD 2>&1 >/dev/null) || {
|
||
# Check if this is an auth error first — no point retrying.
|
||
if echo "$push_err" | grep -qiE "auth|permission|403|401|forbidden"; then
|
||
local hint
|
||
hint=$(remote_auth_hint)
|
||
write_status "push_failed" "push failed: auth error; commit retained locally, will retry next run. fix: $hint"
|
||
echo "BRAIN_SYNC: push failed: auth. fix: $hint" >&2
|
||
# Drained records leave the spool — they live in the local commit, which
|
||
# the run-start detector re-pushes next time (#2549). Retained +
|
||
# post-snapshot records survive the finalize.
|
||
finalize_queue "$snapshot_file" "$class_file" "$paths_file"
|
||
exit 0
|
||
fi
|
||
|
||
# Try a fetch-and-merge + retry. The fetch and the retry push are their
|
||
# own attempted-egress ops, each receipted fail-closed (a refusal falls
|
||
# through to the push_failed path below).
|
||
if GSTACK_HOME="$GSTACK_HOME" _receipted_git closed brain-sync "$push_host" curated-memory-git-fetch "artifacts_sync_mode!=off" \
|
||
bash -c 'git -C "$1" fetch origin 2>/dev/null' _ "$GSTACK_HOME"; then
|
||
local branch
|
||
branch=$(git -C "$GSTACK_HOME" rev-parse --abbrev-ref HEAD 2>/dev/null || echo main)
|
||
if git -C "$GSTACK_HOME" merge --no-edit "origin/$branch" >/dev/null 2>&1; then
|
||
if GSTACK_HOME="$GSTACK_HOME" _receipted_git closed brain-sync "$push_host" curated-memory-git-push "artifacts_sync_mode!=off" \
|
||
bash -c 'git -C "$1" push origin HEAD 2>/dev/null' _ "$GSTACK_HOME"; then
|
||
finalize_queue "$snapshot_file" "$class_file" "$paths_file"
|
||
date -u +%Y-%m-%dT%H:%M:%SZ > "$LAST_PUSH_FILE"
|
||
write_status "ok" "pushed $n file(s) after rebase"
|
||
exit 0
|
||
fi
|
||
fi
|
||
fi
|
||
# Commit exists locally; the run-start detector re-pushes it next time.
|
||
write_status "push_failed" "push failed: $(printf '%s' "$push_err" | head -1); commit retained locally, will retry next run"
|
||
finalize_queue "$snapshot_file" "$class_file" "$paths_file"
|
||
exit 0
|
||
}
|
||
|
||
# Success: drained records leave the spool (retained + post-snapshot survive).
|
||
finalize_queue "$snapshot_file" "$class_file" "$paths_file"
|
||
date -u +%Y-%m-%dT%H:%M:%SZ > "$LAST_PUSH_FILE"
|
||
write_status "ok" "pushed $n file(s)"
|
||
exit 0
|
||
}
|
||
|
||
subcmd_status() {
|
||
if [ -f "$STATUS_FILE" ]; then
|
||
cat "$STATUS_FILE"
|
||
else
|
||
echo '{"status":"unknown","message":"no status file yet"}'
|
||
fi
|
||
# Supplemental info (not in status file). Depth = spool record files plus
|
||
# any not-yet-migrated legacy queue lines (transition window), including a
|
||
# crash-leftover .migrating file — its records are still pending too.
|
||
local queue_depth spool_depth legacy_depth
|
||
spool_depth=$(ls "$QUEUE_DIR"/*.json 2>/dev/null | wc -l | tr -d ' ')
|
||
legacy_depth=0
|
||
[ -f "$QUEUE" ] && legacy_depth=$(wc -l < "$QUEUE" | tr -d ' ')
|
||
[ -f "$QUEUE.migrating" ] && legacy_depth=$(( legacy_depth + $(wc -l < "$QUEUE.migrating" | tr -d ' ') ))
|
||
queue_depth=$(( spool_depth + legacy_depth ))
|
||
local last_push="never"
|
||
[ -f "$LAST_PUSH_FILE" ] && last_push=$(cat "$LAST_PUSH_FILE" 2>/dev/null || echo never)
|
||
local mode
|
||
mode=$("$CONFIG_BIN" get artifacts_sync_mode 2>/dev/null || echo off)
|
||
printf '{"queue_depth":%s,"last_push":"%s","mode":"%s"}\n' "$queue_depth" "$last_push" "$mode"
|
||
}
|
||
|
||
subcmd_skip_file() {
|
||
local path="${1:-}"
|
||
if [ -z "$path" ]; then
|
||
echo "Usage: gstack-brain-sync --skip-file <path>" >&2
|
||
exit 1
|
||
fi
|
||
mkdir -p "$GSTACK_HOME"
|
||
# Avoid duplicate entries.
|
||
if [ -f "$SKIP_FILE" ] && grep -Fxq "$path" "$SKIP_FILE"; then
|
||
echo "already in skip list: $path"
|
||
exit 0
|
||
fi
|
||
echo "$path" >> "$SKIP_FILE"
|
||
echo "added to skip list: $path"
|
||
echo "(future writers will not enqueue this path; existing queue entries ignored on next --once)"
|
||
}
|
||
|
||
subcmd_drop_queue() {
|
||
local force="${1:-}"
|
||
if [ "$force" != "--yes" ]; then
|
||
echo "Refusing: --drop-queue discards pending syncs. Pass --yes to confirm." >&2
|
||
exit 1
|
||
fi
|
||
# Remove spool record files, then truncate any legacy queue remnant —
|
||
# including a crash-leftover .migrating file, whose records would otherwise
|
||
# resurrect on the next drain via migrate_legacy_queue after the user
|
||
# explicitly discarded the queue.
|
||
local n=0 f
|
||
for f in "$QUEUE_DIR"/*.json; do
|
||
[ -e "$f" ] || continue
|
||
rm -f "$f" 2>/dev/null && n=$(( n + 1 ))
|
||
done
|
||
if [ -f "$QUEUE" ]; then
|
||
local legacy_n
|
||
legacy_n=$(wc -l < "$QUEUE" | tr -d ' ')
|
||
n=$(( n + legacy_n ))
|
||
: > "$QUEUE"
|
||
fi
|
||
if [ -f "$QUEUE.migrating" ]; then
|
||
local mig_n
|
||
mig_n=$(wc -l < "$QUEUE.migrating" | tr -d ' ')
|
||
n=$(( n + mig_n ))
|
||
rm -f "$QUEUE.migrating" 2>/dev/null || true
|
||
fi
|
||
if [ "$n" -eq 0 ]; then
|
||
echo "queue already empty"
|
||
exit 0
|
||
fi
|
||
echo "dropped $n queue entries"
|
||
}
|
||
|
||
subcmd_discover_new() {
|
||
if ! sync_active; then
|
||
exit 0
|
||
fi
|
||
# Walk allowlist globs; enqueue any file where mtime+size differs from cursor.
|
||
python3 - "$GSTACK_HOME" "$ALLOWLIST" "$DISCOVER_CURSOR" <<'PYEOF' 2>/dev/null || true
|
||
import sys, os, json, fnmatch, time
|
||
from datetime import datetime, timezone
|
||
|
||
gstack_home, allowlist_path, cursor_path = sys.argv[1:4]
|
||
spool_dir = os.path.join(gstack_home, ".brain-queue.d")
|
||
skip_path = os.path.join(gstack_home, ".brain-skip.txt")
|
||
|
||
def load_lines(path):
|
||
try:
|
||
with open(path) as f:
|
||
return [l.strip() for l in f if l.strip() and not l.lstrip().startswith("#")]
|
||
except FileNotFoundError:
|
||
return []
|
||
|
||
def load_cursor(path):
|
||
try:
|
||
with open(path) as f:
|
||
return json.load(f)
|
||
except (FileNotFoundError, json.JSONDecodeError):
|
||
return {}
|
||
|
||
def save_cursor(path, data):
|
||
try:
|
||
with open(path, "w") as f:
|
||
json.dump(data, f)
|
||
except OSError:
|
||
pass
|
||
|
||
allowlist = load_lines(allowlist_path)
|
||
# Normalize skip entries to the same POSIX form as `rel` below, so a
|
||
# backslash entry in .brain-skip.txt still matches a normalized path on Windows.
|
||
skip = {s.replace(os.sep, "/") for s in load_lines(skip_path)}
|
||
cursor = load_cursor(cursor_path)
|
||
new_cursor = dict(cursor)
|
||
to_enqueue = []
|
||
|
||
# Walk all files under gstack_home, match against allowlist.
|
||
for root, dirs, files in os.walk(gstack_home):
|
||
# Skip .git and .brain-* state files.
|
||
if ".git" in root.split(os.sep):
|
||
continue
|
||
for name in files:
|
||
full = os.path.join(root, name)
|
||
# Repo paths are POSIX-relative. os.path.relpath yields backslash
|
||
# separators on Windows, which never match the forward-slash allowlist
|
||
# globs (e.g. "projects/*/learnings.jsonl"), so discovery silently
|
||
# enqueued nothing under projects/ on Windows. Normalize to "/".
|
||
rel = os.path.relpath(full, gstack_home).replace(os.sep, "/")
|
||
if rel.startswith(".brain-"):
|
||
continue
|
||
if not any(fnmatch.fnmatchcase(rel, pat) for pat in allowlist):
|
||
continue
|
||
if rel in skip:
|
||
continue
|
||
try:
|
||
st = os.stat(full)
|
||
key = f"{int(st.st_mtime)}:{st.st_size}"
|
||
except OSError:
|
||
continue
|
||
if cursor.get(rel) != key:
|
||
to_enqueue.append((rel, key))
|
||
|
||
# Write spool records directly. The previous implementation shelled out to
|
||
# gstack-brain-enqueue once per file, but Windows Python cannot exec a
|
||
# bash-shebang script (the spawn fails with a fork error), so discovery
|
||
# enqueued nothing on Windows even after the path-match fix above.
|
||
# Writing the record here is platform-agnostic; the drain step
|
||
# (compute_paths_to_stage) still re-applies the skip-list + privacy filters.
|
||
if to_enqueue:
|
||
ts = datetime.now(timezone.utc).strftime("%Y-%m-%dT%H:%M:%SZ")
|
||
written = []
|
||
try:
|
||
# One spool FILE per record (tmp write + atomic os.replace), matching
|
||
# gstack-brain-enqueue's maildir contract: writers and the drain never
|
||
# share an inode, so a parallel writer or drain can't race this.
|
||
# Compact separators match the shim's JSON shape.
|
||
os.makedirs(spool_dir, exist_ok=True)
|
||
for i, (rel, key) in enumerate(to_enqueue):
|
||
rec = json.dumps({"file": rel, "ts": ts}, separators=(",", ":"))
|
||
tmp = os.path.join(spool_dir, f".tmp-{os.getpid()}-d{i}")
|
||
with open(tmp, "w") as f:
|
||
f.write(rec + "\n")
|
||
os.replace(tmp, os.path.join(spool_dir, f"{int(time.time())}-{os.getpid()}-d{i}.json"))
|
||
written.append((rel, key))
|
||
except OSError:
|
||
# Spool write failed (disk full, AV file lock). Leave the cursor
|
||
# unadvanced for unwritten records so they are retried on the next
|
||
# discover instead of being silently recorded as synced (which loses
|
||
# the change until the file next changes).
|
||
pass
|
||
# Advance the cursor only for records actually written.
|
||
for rel, key in written:
|
||
new_cursor[rel] = key
|
||
|
||
save_cursor(cursor_path, new_cursor)
|
||
PYEOF
|
||
}
|
||
|
||
# -------- dispatch --------
|
||
case "${1:-}" in
|
||
--once|"") subcmd_once ;;
|
||
--status) subcmd_status ;;
|
||
--skip-file) shift; subcmd_skip_file "${1:-}" ;;
|
||
--drop-queue) shift; subcmd_drop_queue "${1:-}" ;;
|
||
--discover-new) subcmd_discover_new ;;
|
||
--help|-h)
|
||
sed -n '2,18p' "$0" | sed 's/^# \{0,1\}//'
|
||
;;
|
||
*)
|
||
echo "Unknown subcommand: $1" >&2
|
||
echo "Run: gstack-brain-sync --help" >&2
|
||
exit 1
|
||
;;
|
||
esac
|