mirror of
https://github.com/garrytan/gstack.git
synced 2026-10-04 02:16:56 +02:00
* refactor(resolvers): split review.ts into MECE resolver modules (pure move) Move every function from scripts/resolvers/review.ts, unchanged, into: - review-dashboard.ts: review dashboard, plan-file review report - plan-gates.ts: approval check, exit-plan-mode gate, plan-file discovery, plan-completion audit/gate (ship + review), plan verification exec - spec-review.ts: both spec review loops, benefits-from, anti-shortcut clause - outside-voice-steps.ts: Codex second opinion, adversarial step, Codex plan review, Codex doc review, disabled-outside record - review-scope.ts: scope drift, cross-review dedup, shared-code reuse review.ts is deleted; index.ts imports the new modules. gen-skill-docs output is byte-identical for every host (--host all). Test imports and source-path references are re-pointed; the two source-text report/gate tests in gen-skill-docs.test.ts become behavioral renders across every consuming skill and host. All 46 touchfile entries that named review.ts now name all five modules, guarded by a recorded selection golden. * test(browse): black-box auth matrix for every server route and both surfaces Drives buildFetchHandler fetchLocal/fetchTunnel with no token, wrong token, root token, scoped token and the SSE cookie for all 33 routes, plus unmatched paths and wrong methods. Denials assert today's exact status, body and content type; allowed credentials assert the handler was reached. Written against the unchanged if-chain server so the W3 route-table refactor must keep it green. * refactor(shard-engine): move scripts/test-strict-output.ts to scripts/lib/shard-engine.ts The shared shard engine grows from the existing strict-output module (runShardChild, killProcessGroup, signal forwarding, strict classifier). scripts/test-strict-output.ts stays as a re-export so existing importers, mock.module paths and the strict-output/run-shard-child tests are unchanged. The engine inherits the global touchfile entry; the free runner's CLI-routing fixture copies the new module. * refactor(resolvers): decompose the three >150-line review resolvers (output-neutral) Split generateAdversarialStep, generateCodexPlanReview and generatePlanCompletionAuditInner into per-section helpers whose template literals are copied verbatim, so every function in the new modules is at or under 150 lines. gen-skill-docs output is byte-identical for every host (--host all, compared against96764e80with a fixed --link-root). * refactor(resolvers): one outside-voice failure policy (deliberate prose unification) outsideVoiceFailurePolicy(ctx, opts) in outside-voice.ts now renders the auth / timeout / empty-response bullets for all four call sites that hand-typed them (Codex second opinion, adversarial step, Codex plan review, design outside voices). Options are explicit per site (timeoutMinutes, onTimeout, stderrOnEmpty, fallback, escape) with no defaults. Deliberate generated-prose changes (every host): - office-hours: 'Fall back to <native> subagent.' becomes 'Fall back to the <native> subagent below.' - plan-devex-review: the plain 'Auth failure (stderr contains ...)' bullets become the canonical bold bullets; auth also triggers on 'API key'; 'auth failed' becomes 'authentication failed'. - review/ship adversarial: 'exceeded 9 minutes and was terminated' becomes 'timed out after 9 minutes and was terminated'; the timeout is still MISSING COVERAGE. - design outside voices: unchanged. Adds ratchet (d) (test/outside-voice-failure-policy.test.ts) with a reasoned allowlist for /codex's own CLI errors, the MISSING COVERAGE retention test, refreshed codex/factory ship goldens, and outside-voice.ts in every touchfile entry of review.ts and design.ts (selection golden extended). * test(pty): fake PTY session driver with an injectable clock through the runner launch seam The three plan-skill runners take an optional PtyDriver (launch, now, monotonic, sleep); omitted, they use the real launcher and clocks exactly as before. test/helpers/pty/fake-session.ts feeds scripted frames through that seam, and claude-pty-runner.runners.unit.test.ts runs observation, counting and floor for success, deadline timeout, permission prompt and plan-ready outcomes with no CLI or real timers. These cases must stay green unchanged through the W4 split and the runPtySession extraction. Touchfiles: every entry that lists claude-pty-runner.ts or pty-screen.ts now also lists test/helpers/pty/**. * refactor(shard-engine): run both lanes on the shared engine; lane policy injected Engine (scripts/lib/shard-engine.ts) gains the W2 primitives: per-shard tmp/Chromium sandbox + async cleanup backstop, log-path allocation and full-stream log capture, one duration-seed reader/writer with a lane predicate, LanePolicy (seed predicate + zero-execution verdict), strictShardStatus, and the shared CLI flag loop. runShardChild takes an optional companion (signal/settle) and waits a bounded 250ms to reap a wall-killed child. Free lane stops spawning shards itself: runFreeShard uses runShardChild with trackShardBrowser as the companion (win32 path unchanged: no process group, no negative-pid kill). Its sync state-dir removal stays lane policy. Paid lane uses the sandbox, log, seed, verdict and flag primitives; the hollow-shard guard applies PAID_LANE_POLICY. Lane outcomes are unchanged (free keeps >= 0 seeds and file-count zero-exec rule; paid keeps > 0 seeds, warning under selection and passed-empty under EVALS_ALL). paid-free-boundary's closure assertion now names the engine module, where the strict classifier lives. * test(shard-engine): engine unit tests, fixture-corpus equivalence, per-lane CLI parity - test/shard-engine.test.ts: failing/unhandled/module-load output fails both lanes, per-lane zero-execution and seed rules, whole-group kill on a wall timeout (both lanes), mocked-win32 path with no negative-pid kill, companion settle order, log capture, sandbox isolation, flag loop. - test/shard-engine-equivalence.test.ts + test/fixtures/shard-equivalence: seven outcome fixtures plus one real shard, run through both lanes and compared with classifications recorded from the base runners (96764e80). - test/shard-cli-parity.test.ts + test/fixtures/shard-cli-parity: flag set, defaults, validation errors and the Unknown argument error per lane match the base runners. * refactor(shard-engine): decompose runFreeShard and runPaidShard to <= 150 lines Output-neutral extraction under the fixture-corpus equivalence and runner tests: captureFreeStream, explainFreeVerdict and logFreeRecovery (free); paidShardCommand, settleShardSpool, settleBootstrapRetention and printLogTail (paid). The bootstrap scope-creation block that bootstrap-retention.test.ts evaluates stays verbatim. * refactor(pty): split claude-pty-runner.ts into test/helpers/pty/* behind a barrel Pure move: every line of the former 5,047-line runner lands verbatim in one module (four private helpers gain `export` for cross-module use): binary, screen (absorbs test/helpers/pty-screen.ts, which now re-exports it), launch, session (PtyDriver), judge, classify, auq, plan-native, boundaries, runners/{observation,counting,floor}. claude-pty-runner.ts re-exports the original public surface by name; pty/ modules import siblings directly. Tests that read the runner's source text: - rewritten as behavioral: the unit test's model-pin tripwire (fake CLI argv: fallback chain, --model before extraArgs, hermetic --strict-mcp-config), pty-skill-seeding-wiring (runners through the fake driver; launcher through a fake CLI reporting CLAUDE_CONFIG_DIR). The "three wrappers forward model" grep is replaced by the runners' fake-driver launch assertions. - pty-screen-session / pty-screen-supervision: stop copying runner source; they mock.module the real pty/screen.ts (and the fixture cleanup) instead. - re-pointed to the owning module (they execute a sliced runner body with injected boundaries; no seam exists for those boundaries yet): eng-seeded-completion-ai, plan-floor-permission, plan-create-prepublication, plan-count-completion; hermetic-wiring's source guard now reads pty/launch.ts and scans every pty/ module for raw process.env spreads. - plan-count-timeout and pty-output-wake mock the viewport at pty/screen.ts. * test(ratchet-c): enforcing module/function size ratchet and moved-code touchfile coverage Ratchet (c) ships enforcing: test/helpers/module-size.ts counts file and top-level function lengths by brace matching over masked source (strings, comments, regex literals and template text masked; ${} expressions kept), covering function declarations, arrow functions assigned to consts and route-table handler properties, with no parser dependency. Its self-test uses template literals and code-fence braces copied from scripts/resolvers/review.ts and design.ts. test/fixtures/module-size-ratchet.json binds scripts/lib/shard-engine.ts (<= 800 lines, <= 150 per function) and records the residual runner sizes (free 2352, paid 1921) as non-growth caps; allowlist entries are keyed on file plus matched text and need a reason. Failure output lists file:line, the rule, Fix: and the allowlist path. touchfiles.test.ts gains the moved-code superset check over test/fixtures/touchfile-move-goldens/ (W2 golden recorded at96764e80: test-strict-output.ts and test-paid-shards.ts global, test-free-shards.ts none). * refactor(browse): declared route table replaces the buildFetchHandler if-chain The ~1,300-line if-chain in buildFetchHandler becomes a route table: each entry declares method, path, auth kind and surfaces, and one auth gate in browse/src/routes/table.ts returns the per-kind denial (root-bearer, scoped, root-or-sse-cookie: 401 Unauthorized; root-token: 403 Root token required; extension-origin: 403 Forbidden). Unmatched requests take the declared fallthrough (root-bearer check, then plain-text 404). Handlers move to browse/src/routes/{core,pairing,pty,tokens,tunnel,activity,commands,files, inspector}.ts and receive a RouteContext with auth checks as functions instead of closing over factory locals. Dispatch order is unchanged: tunnel filter, beforeRoute overlay, gate, handler. TUNNEL_PATHS stays a literal in server.ts. Behavior-preserving: the black-box auth matrix from the previous commit passes unchanged. /memory and /inspector/events are declared root-bearer because the blanket check always ran before their SSE-cookie branch. Source-text route tests are rewritten as behavioral tests through buildFetchHandler or a route's real handler with a stub RouteContext (browse/test/route-test-harness.ts). Checks with no runtime seam are re-pointed to the route modules: Surface type, /inspector/events SSE helper, sanitizeReplacer imports, /pty-inject-scan sidecar-client import, and the ngrok config lookup and startTunnel wiring that stay in server.ts. * test(browse): stubbed-handler auth matrix and route inventory for the route table Every ROUTES entry runs through the real dispatcher and gate with stub handlers on each declared surface and six credentials; denials assert the exact status and body each auth kind returned at96764e8, admitted credentials assert the handler ran (with the gate's TokenInfo for scoped routes). Also pins the reviewed route inventory (method, path, auth kind, surfaces), that every entry declares auth and surfaces, that the table's tunnel paths equal the TUNNEL_PATHS literal with GET /connect admitted, the unmatched fallthrough, and that the root token is rejected on every tunnel route through buildFetchHandler. * test(browse): ratchet (b) keeps route dispatch inside the route table Scans browse/src/server.ts and browse/src/routes/*.ts for pathname comparisons; only the table matcher and the tunnel-surface filter are allowed, listed with reasons in browse/test/fixtures/route-dispatch-allowlist.json (keyed on file plus line text). Also checks every entry declares auth and surfaces and that gstack registers no beforeRoute overlay itself. Self-tests plant a violation and assert the file:line, Fix: and allowlist path in the message, that a shifted line stays allowlisted, and that a reasonless entry is rejected. * test: touchfile superset check for modules moved out of browse/src/server.ts Records the paid evals selected by touching browse/src/server.ts at96764e80(17 E2E, 1 LLM judge) and asserts every browse/src/routes/*.ts module selects a superset. The test reads every golden in test/fixtures/moved-module-selection/ so other moved-code goldens can sit beside it. * test(shard-engine): give non-timeout corpus fixtures CI headroom; keep the POSIX golden off the Windows lane Only the wall-timeout fixture keeps a 3s wall; the rest get 60s so a loaded host cannot turn a pass into a timeout. Base and branch runners still agree on every classification under the new walls. The Windows exclusion entry moves the free runner's ratchet (c) residual cap to 2356 lines. * refactor(pty): one runPtySession loop drives observation, counting and floor test/helpers/pty/session.ts owns launch -> start -> (poll -> tick)* -> timeout and the failure contract the three runners each hand-rolled: the run's own error wins over capture and close errors, close always runs, owned fixture cleanup runs last (also when launch fails). Each runner now supplies a PtySessionPlan: its boot/command step, poll cadence (2s observation/floor sleep; counting's output wake + 250ms coalesce), tick policy (permission handling, native identity, terminal rules stay per runner because they differ) and capture hooks. The runner bodies are decomposed into top-level steps so no function exceeds 150 lines; behavior is unchanged and the fake-driver cases from the first W4 commit pass unmodified. The counting capture step and the native completion-summary predicate are now named functions (countingCapture, isNativeCompletionSummary), so plan-create-prepublication and plan-count-completion call them directly instead of executing sliced source. The two harnesses that still execute a sliced runner body with injected boundaries (eng-seeded-completion-ai, plan-floor-permission) pass the PtyDriver seam instead of overriding Date/Bun.sleep. * test(ratchet-c): register route modules, review resolver modules and server.ts residual cap * refactor(pty): decompose launchClaudePty and engNumberedFindingAUQ under 150 lines launchClaudePty (349 lines) becomes launch preparation (args, hermetic child env, owned state roots), recorder creation, spawn, the trust-dialog watcher, close, and the session handle over one PtyProcess state object. The failure order is unchanged: abort the viewport, dispose any recorders created so far, dispose the viewport, rethrow. The --model / --strict-mcp-config ordering and seedSkills wiring stay pinned by the behavioral fake-CLI tests. engNumberedFindingAUQ (345 lines) keeps its guards and dispatch; each self-contained issue family (declared cache, library retry hooks, cache owner, injected singleton, shared writers, injected export) moves verbatim into its own function. Every pty/ module is now <= 800 lines and every top-level function <= 150 lines. * test(pty): split claude-pty-runner.unit.test.ts along the pty/ module seams The 188 unit tests move verbatim into claude-pty-runner.{screen,classify, auq,launch,plan-native,boundaries}.unit.test.ts (test names unchanged; each file imports only what it uses from the barrel). The five files that no longer read a SKILL.md template join the test-of-test ratchet baseline with a reason. * test(touchfiles): moved PTY modules keep their paid-eval selection test/fixtures/touchfile-selection/w4-pty.json records, at96764e8, the paid evals selected by touching test/helpers/claude-pty-runner.ts (20) and test/helpers/pty-screen.ts (20). touchfiles.test.ts now asserts every .ts file under test/helpers/pty/ (and pty/screen.ts for both sources) selects a superset, reading every golden in that directory so later moves can add one; a planted-violation case pins the report and its Fix line. * fix(browse): unexchanged pair setup keys no longer authenticate bearer requests validateToken accepted a gsk_setup_ key as a bearer on /command, /batch and /file (found while building the W3 auth matrix). A setup key now only authenticates the /connect exchange. * W1: one state-root owner (lib/state-root.ts + bin/gstack-state-root.sh), gstack-paths --explain and fail-stop, parity tests * W1: guarded migration of every executable state-root site; uninstall deletes only ~/.gstack Bins, careful/freeze hooks, setup, upgrade migrations, browse/src, design, ios-qa daemon, lib and scripts resolve the state root through bin/gstack-state-root.sh (bash) or lib/state-root.ts (TS). Bins source the twin and stop with a reinstall message when it is missing; hooks source it and never spawn gstack-paths. browse/src/config.ts and lib/cso/state.ts delegate to resolveStateRoot. Analytics writers and readers move together so the usage log stays one file. gstack-uninstall deletes state only at ~/.gstack, refuses (exit 2) when it resolves to /, $HOME or an ancestor, the checkout or the git root, and leaves any other resolved root in place with the removal command. Fixtures that copy single bins now copy the twin. * W1: privacy keys and trust-policy deny tiers merge across state roots; gstack-config reporting; test hermeticity readConfigKey / gstack_read_config_key return the most restrictive telemetry, memorable_recall, codex_reviews and update_check across the resolved root and ~/.gstack; other keys read the resolved root only. gstack-config set reports an overriding root with the exact override command, list shows the winning root and a root-variable disagreement line. gstack-gbrain-repo-policy get merges deny/read-only tiers. gstack-egress reads through readConfigKey. test-setup.ts strips inherited GSTACK_STATE_ROOT/GSTACK_STATE_DIR and redirects the legacy root. * W1: shared hook logging helper (hosts/claude/hooks/hook-log.ts) One hook-errors.log writer: root from resolveStateRoot, 0600 on every append, opt-in rate limit used only by memorable-user-prompt. The five hooks route through it. * W1: docs/state-root.md and README troubleshooting pointer Precedence table, a real --explain example, the move-your-state recipe, merged privacy keys, the uninstall rule, the resolver-failure fix, and the plugin-mode note (evidence gate: no official plugin distribution). * W1b: template and resolver prose resolve state through guarded gstack-paths; ratchet (a) Every gstack-paths eval in templates and resolvers carries the fail-stop guard; executable ~/.gstack paths in bash blocks (context recovery preamble, eureka log, analytics, project artifacts, upgrade snooze, setup-gbrain lock, retro snapshots, ship consent marker) use $GSTACK_STATE_ROOT, and the writer prose that pairs with them points at the printed PROJECT_DIR / RETRO_FILE. ship drops export GSTACK_STATE_ROOT. SKILL.md regenerated (claude + codex), ship goldens re-pinned, parity and context-budget caps raised to the measured sizes with notes. test/state-root-ratchet.test.ts enforces the rule with a reasoned allowlist; W1 touchfile entries plus a superset golden. * refactor: apply W1 state-root edits in W2/W3/W5-owned files; one moved-code touchfile golden for all workstreams * test: fold the moved-code touchfile golden into touchfiles.test.ts; fix integration fixture closure and caps * v1.91.11.0: CHANGELOG, TODOS, docs and conventions for the refactor wave * test: re-measure plan-ceo/design-consultation caps and ship goldens after the guarded plan-discovery and spec-review blocks; add the state-root twin to the workflow-boundaries fixture * fix(windows): migrations resolve their directory with either path separator; state-root parity compares under the HOME Git Bash actually sees * fix(review,ship): state plan-check timing after smoke expiry and test_stub Skip semantics (review workflow judge clarity) * test(qa-eval): webhook fix eval asks for the fix loop's post-repair probes; eight-scenario coverage stays in the report-only case and the harness recheck * test(qa-eval): re-pin the webhook prompt contract to the fix-loop stage; R29 coverage omissions stay bound by the report-only case * fix(review,ship): plan checks publish a checkpoint before each probe; only the smoke expiry stop is skipped * fix(qa): carry #2999's checkpoint receipt link, report-template line and full-revision placeholder (identical hunks) * test(qa-callers): disable git auto maintenance in the caller fixture Git 2.47+ runs auto maintenance detached after commit; on the CI runner's git 2.55 it rewrote .git/objects fan-out directories while the write observer was running, which surfaced as unauthorized mutations. Same gc.auto=0 / maintenance.auto=false guard the shared-libs fixture already uses. * test(plan-mode-no-op): require prose evidence for the prose-fallback members so a spinner-frame judge verdict cannot end the run as asked * test(ship-docsync): carry #2999's seeded-attempt docsync harness (identical files) The doc-sync fault cases replayed attempt 1 before reaching their gate and ran out of their 285s budget. The fixture now seeds attempt 1 and the parent starts at the gate under test. Taken byte-identical from origin/capy/audit-fix-wave (fb526898,e6ac813d,6ce10ff7,d0c53577,77cce3be). Local: stale-before, recovery and late-result 6/6 PASS (97-164s); the whole file 12/12 PASS.
765 lines
38 KiB
Cheetah
765 lines
38 KiB
Cheetah
---
|
||
name: retro
|
||
preamble-tier: 2
|
||
version: 2.0.0
|
||
description: |
|
||
Weekly engineering retrospective. Analyzes commit history, work patterns,
|
||
and code quality metrics with persistent history and trend tracking.
|
||
Team-aware: breaks down per-person contributions with praise and growth areas.
|
||
Use when asked to "weekly retro", "what did we ship", or "engineering retrospective".
|
||
Proactively suggest at the end of a work week or sprint. (gstack)
|
||
allowed-tools:
|
||
- Bash
|
||
- Read
|
||
- Write
|
||
- Glob
|
||
- AskUserQuestion
|
||
triggers:
|
||
- weekly retro
|
||
- what did we ship
|
||
- engineering retrospective
|
||
gbrain:
|
||
schema: 1
|
||
context_queries:
|
||
- id: prior-retros
|
||
kind: filesystem
|
||
# #2552: /retro writes .context/retros/*.json (repo-local; see the save
|
||
# step below) — the old ~/.gstack/.../retros/*.md glob matched a
|
||
# directory and extension nothing ever writes, so this query was dead.
|
||
glob: ".context/retros/*.json"
|
||
sort: mtime_desc
|
||
limit: 5
|
||
render_as: "## Prior retros for this project"
|
||
- id: recent-timeline
|
||
kind: filesystem
|
||
glob: "~/.gstack/projects/{repo_slug}/timeline.jsonl"
|
||
tail: 30
|
||
render_as: "## Recent timeline events"
|
||
- id: recent-learnings
|
||
kind: filesystem
|
||
glob: "~/.gstack/projects/{repo_slug}/learnings.jsonl"
|
||
tail: 10
|
||
render_as: "## Recent learnings"
|
||
---
|
||
|
||
{{PREAMBLE}}
|
||
|
||
{{BASE_BRANCH_DETECT}}
|
||
|
||
# /retro — Weekly Engineering Retrospective
|
||
|
||
Analyze commit history, work patterns, and code quality for the current user and every contributor, with evidence-backed praise and growth opportunities.
|
||
|
||
## User-invocable
|
||
When the user types `/retro`, run this skill.
|
||
|
||
## Arguments
|
||
- `/retro` — default: last 7 days
|
||
- `/retro 24h` — last 24 hours
|
||
- `/retro 14d` — last 14 days
|
||
- `/retro 30d` — last 30 days
|
||
- `/retro compare` — compare current window vs prior same-length window
|
||
- `/retro compare 14d` — compare with explicit window
|
||
- `/retro global` — cross-project retro across all AI coding tools (7d default)
|
||
- `/retro global 14d` — cross-project retro with explicit window
|
||
|
||
{{GBRAIN_CONTEXT_LOAD}}
|
||
|
||
{{SECTION_INDEX:retro}}
|
||
|
||
## Instructions
|
||
|
||
Parse the argument to determine the time window. Default to 7 days if no argument given. All times should be reported in the user's **local timezone** (use the system default — do NOT set `TZ`).
|
||
|
||
**Midnight-aligned windows:** For day (`d`) and week (`w`) units, compute an absolute start date at local midnight, not a relative string. For example, if today is 2026-03-18 and the window is 7 days: the start date is 2026-03-11. Use `--since "2026-03-11T00:00:00"` — the explicit `T00:00:00` suffix ensures git starts from midnight. Without it, git uses the current wall-clock time (e.g., `--since "2026-03-11"` at 11pm means 11pm, not midnight). For week units, multiply by 7 to get days (e.g., `2w` = 14 days back). For hour (`h`) units, use `--since "N hours ago"` since midnight alignment does not apply to sub-day windows. Compute "today" from the user-visible `## currentDate` tag in the session reminder — NEVER from `date` (the system clock can be hours off in containerized harnesses). If you cannot reliably compute "today", stop and ask the user via AskUserQuestion rather than proceeding.
|
||
|
||
**Argument validation:** If the argument doesn't match a number followed by `d`, `h`, or `w`, the word `compare` (optionally followed by a window), or the word `global` (optionally followed by a window), show this usage and stop:
|
||
```
|
||
Usage: /retro [window | compare | global]
|
||
/retro — last 7 days (default)
|
||
/retro 24h — last 24 hours
|
||
/retro 14d — last 14 days
|
||
/retro 30d — last 30 days
|
||
/retro compare — compare this period vs prior period
|
||
/retro compare 14d — compare with explicit window
|
||
/retro global — cross-project retro across all AI tools (7d default)
|
||
/retro global 14d — cross-project retro with explicit window
|
||
```
|
||
|
||
**Routing:** `global` skips all repo-scoped steps, including Prior Learnings, Step 0.5, and post-report capture; follow **Global Retrospective Mode** (no git repo required). `compare` follows **Compare Mode**. Both accept an optional window (default 7d). Otherwise run the repo-scoped flow below.
|
||
|
||
`<default>` is the base branch from the preceding **Step 0: Detect platform and base branch**. `<today>` is the session-reminder date; reuse it in all snapshot filenames, never re-read the clock.
|
||
|
||
{{LEARNINGS_SEARCH}}
|
||
|
||
### Step 0.5: Freshness pre-flight (fetch)
|
||
|
||
Refresh `origin/<default>` so the retro doesn't misreport against a stale local ref. If the repo has no `origin` remote this fails harmlessly — the metrics script (Step 1) falls back to the local branch and its guard lines disclose it:
|
||
|
||
```bash
|
||
git fetch origin <default> --quiet 2>/dev/null \
|
||
|| echo "RETRO_FETCH: failed (offline or no remote) — proceeding against last-known refs"
|
||
```
|
||
|
||
Remember whether the fetch succeeded — the stale-base guard in Step 1 only BLOCKs when it did.
|
||
|
||
### Step 1: Gather Metrics (one command)
|
||
|
||
Run `gstack-retro-metrics` with the detected base branch and computed start:
|
||
|
||
```bash
|
||
_RM="$HOME/.claude/skills/gstack/bin/gstack-retro-metrics"
|
||
[ -x "$_RM" ] || _RM=".claude/skills/gstack/bin/gstack-retro-metrics"
|
||
"$_RM" --base "<default>" --since "<since>" \
|
||
|| echo "RETRO_METRICS: unavailable — stale install (read the helper source for manual computation)"
|
||
```
|
||
|
||
Read the `METRIC_NAME: value` lines. **Degraded mode:** without `RETRO_METRICS_PROTO: 1`, reproduce computations from the installed `bin/gstack-retro-metrics` source, not the presentation steps below. If source or metrics are unavailable, say so; never invent values. Suggest `/gstack-upgrade` to restore the helper.
|
||
|
||
**Identity:** `USER_NAME` is **"you"** — the person reading this retro. All other authors are teammates. Orient the narrative around this: "your" commits vs teammate contributions.
|
||
|
||
**Stale-base + bad-today-anchor guard.** `GUARD_LATEST_COMMIT: <DATE>` is the newest commit on the analyzed ref. A wrong "today" or stale ref can produce an empty window. Evaluate in order:
|
||
|
||
1. If `GUARD_REMOTE: none` or `GUARD_HEAD: detached` or the Step 0.5 fetch failed: proceed, but carry the disclosure into the narrative ("offline run, window not freshness-verified") rather than silently misreporting.
|
||
2. If the Step 0.5 fetch succeeded AND the `GUARD_LATEST_COMMIT` date is **older than (today − window-days)**: BLOCK with: "Retro window is stale. Latest commit on `origin/<default>` was `<DATE>`, but the window covers `<since>` to `<today>`. This usually means either (a) today's date is wrong in this session or (b) `origin/<default>` is materially behind the remote. Confirm today's date via the session reminder; if today is correct, run `git fetch origin <default>` manually and re-run /retro." Stop the skill until the user resolves.
|
||
3. Otherwise, write: "RETRO_GUARD: latest commit `<DATE>` within window — proceeding."
|
||
|
||
Also check `RETRO_REF`: if it is not `origin/<default>` (local-only repo, missing remote branch), disclose which ref the retro analyzed.
|
||
|
||
**Metric line reference** (what the script emits):
|
||
|
||
| Line | Meaning |
|
||
|------|---------|
|
||
| `COMMIT: hash\|author\|datetime\|+ins/-del\|subject` | One per commit, newest first (capped at 300) — the raw material for narrative anchoring |
|
||
| `COMMITS` / `MERGE_COMMITS` / `CONTRIBUTORS` | Window totals on the analyzed ref |
|
||
| `INSERTIONS` / `DELETIONS` / `NET_LOC` | Raw LOC |
|
||
| `LOGICAL_SLOC_ADDED` | Non-blank, non-comment added lines — the primary code-volume metric |
|
||
| `TEST_INSERTIONS` / `TEST_RATIO` | Test LOC (test/spec paths + .test./.spec. suffixes) and its share of insertions |
|
||
| `WEIGHTED_COMMITS` | Commits × files-touched, capped at 20 per commit |
|
||
| `ACTIVE_DAYS` | Distinct local dates with commits |
|
||
| `SESSIONS` / `DEEP_SESSIONS` / `MEDIUM_SESSIONS` / `MICRO_SESSIONS` | 45-minute-gap session detection: deep 50+ min, medium 20-50, micro <20 |
|
||
| `TOTAL_ACTIVE_MINUTES` / `AVG_SESSION_MINUTES` / `LOC_PER_SESSION_HOUR` | Session time aggregates (LOC/hour pre-rounded to nearest 50) |
|
||
| `COMMIT_TYPES` / `FIX_RATIO` | Conventional-commit prefix mix |
|
||
| `COMMIT_SIZE_BUCKETS` | small <100 / medium 100-500 / large 500-1500 / xl 1500+ LOC per commit |
|
||
| `HOURS` / `PEAK_HOUR` | Hourly commit histogram (local time), nonzero hours only |
|
||
| `FOCUS_SCORE` | % of file changes in the single busiest top-level directory |
|
||
| `BIGGEST_COMMIT` | Highest-LOC commit in the window (ship-of-the-week candidate) |
|
||
| `HOTSPOT: count file` | Top 10 most-changed files |
|
||
| `AUTHOR: name\|commits\|ins\|del\|test_ratio\|top_areas\|types\|peak_hour` | Per-contributor rollup, sorted by commits desc |
|
||
| `AUTHOR_BIGGEST: name\|hash\|loc\|subject` | Each contributor's biggest ship |
|
||
| `COAUTHOR: hash\|name` / `AI_ASSISTED_COMMITS` | Human co-author credit lines; count of commits with AI trailers |
|
||
| `WEEK: wN\|commits\|ins\|del\|test_ratio` | Weekly buckets, w0 = newest (for Step 10 trends) |
|
||
| `PR_REFS` / `PRS_REFERENCED` | PR/MR numbers from commit subjects (GitHub #NNN, GitLab !NNN) |
|
||
| `TEST_FILES_TOTAL` / `TEST_FILES_CHANGED` / `REGRESSION_TEST_COMMITS` / `REGRESSION_COMMIT` | Test health: repo-wide test file count, test files changed in window, `test(qa):` / `test(design):` / `test: coverage` commits |
|
||
| `VERSION_RANGE` | First → last VERSION file value in the window (when tracked) |
|
||
| `TEAM_STREAK` / `USER_STREAK` | Consecutive commit days with anchor date (Step 11) |
|
||
| `RETRO_CONTEXT` / `GREPTILE_HISTORY` / `TODOS_FILE` / `SKILL_USAGE_LOG` / `EUREKA_LOG` | Presence of optional inputs — Read the ones marked present |
|
||
|
||
**Optional inputs** (Read each file the script marks `present`):
|
||
|
||
- `RETRO_CONTEXT: present` → Read `~/.gstack/retro-context.md`. It is user-authored and may contain meeting notes, calendar events, decisions, and other context that doesn't appear in git history. Incorporate it into the retro narrative where relevant.
|
||
- `GREPTILE_HISTORY: present` → Read `~/.gstack/greptile-history.md`. Filter entries to the retro window by date. Count by type: `fix`, `fp`, `already-fixed`. Signal ratio = `(fix + already-fixed) / (fix + already-fixed + fp)`. Skip unparseable lines silently; if no entries fall in the window, skip the Greptile metric row.
|
||
- `TODOS_FILE: present` → Read `TODOS.md`. Compute: total open TODOs (exclude the `## Completed` section), P0/P1 count, P2 count, items completed this period (Completed entries dated within the window), items added this period (cross-reference `COMMIT:` lines that touched TODOS.md).
|
||
- `SKILL_USAGE_LOG: present` → Read `~/.gstack/analytics/skill-usage.jsonl`. Filter to the window by `ts`. Separate skill activations (no `event` field) from hook fires (`event: "hook_fire"`). Aggregate by skill name.
|
||
- `EUREKA_LOG: present` → Read `~/.gstack/analytics/eureka.jsonl`. Filter to the window by `ts`. For each eureka moment note the skill that flagged it, the branch, and a one-line summary of the insight.
|
||
|
||
### Step 2: Compute Metrics
|
||
|
||
Most rows come directly from the metric lines. Gather the two shipping outcomes separately before building the table:
|
||
|
||
- **Merged PRs:** On GitHub, run `gh pr list --state merged --base "<default>" --search "merged:>=<start-date>" --limit 1000 --json number,title,mergedAt`. Filter `mergedAt` to the exact requested window, including its upper bound in compare mode. If the result hits the limit, paginate via the hosting API or label the count partial. On GitLab use the equivalent merged-MR listing. If hosting data is unavailable, show **PRs referenced** = `PRS_REFERENCED` instead; these are not verified merges. Save `prs_merged: null` in that case.
|
||
- **Features shipped:** Read CHANGELOG changes on `RETRO_REF` in the same window (`git log <ref> --since "<since>" -p -- CHANGELOG.md`, adding `--until` for the prior window). Combine newly added user-visible capabilities with verified merged PR titles. Deduplicate entries referring to the same capability, excluding fixes, chores, and reverted work. Keep a short list of feature names with their source commit/PR beside the count. If neither source is available, show unavailable, not zero. This is an evidence-backed classification, not a metric-script line.
|
||
|
||
Use the analyzed ref in the commit-count label (not always `main`). Test health counts **files changed**, not tests added or test cases; use `TEST_FILES_TOTAL`, `TEST_FILES_CHANGED`, and `REGRESSION_TEST_COMMITS` respectively.
|
||
|
||
| Metric | Value |
|
||
|--------|-------|
|
||
| **Features shipped** (from CHANGELOG + merged PR titles) | N |
|
||
| Commits to analyzed ref | N |
|
||
| Weighted commits (`WEIGHTED_COMMITS`) | N |
|
||
| Contributors | N |
|
||
| PRs merged | N |
|
||
| **Logical SLOC added** (`LOGICAL_SLOC_ADDED` — primary code-volume metric) | N |
|
||
| Raw LOC: insertions | N |
|
||
| Raw LOC: deletions | N |
|
||
| Raw LOC: net | N |
|
||
| Test LOC (insertions) | N |
|
||
| Test LOC ratio | N% |
|
||
| Version range | vX.Y.Z.W → vX.Y.Z.W |
|
||
| Active days | N |
|
||
| Detected sessions | N |
|
||
| Avg raw LOC/session-hour | N |
|
||
| Greptile signal | N% (Y catches, Z FPs) |
|
||
| Test Health | N test files · M changed this period · K regression test commits |
|
||
|
||
Lead with user-visible features, then commit and logical-SLOC metrics; raw LOC
|
||
is only context, not impact (PLAN_TUNING_V1.md, Workstream C).
|
||
|
||
Then show a **per-author leaderboard** immediately below, from the `AUTHOR:` lines:
|
||
|
||
```
|
||
Contributor Commits +/- Top area
|
||
You (garry) 32 +2400/-300 browse/
|
||
alice 12 +800/-150 app/services/
|
||
bob 3 +120/-40 tests/
|
||
```
|
||
|
||
Sort by commits descending. The current user (`USER_NAME`) always appears first, labeled "You (name)".
|
||
|
||
Conditional rows (skip each when its input is absent or empty in the window):
|
||
|
||
```
|
||
| Backlog Health | N open (X P0/P1, Y P2) · Z completed this period |
|
||
| Skill Usage | /ship(12) /qa(8) /review(5) · 3 safety hook fires |
|
||
| Eureka Moments | 2 this period |
|
||
```
|
||
|
||
If eureka moments exist, list them:
|
||
```
|
||
EUREKA /office-hours (branch: garrytan/auth-rethink): "Session tokens don't need server storage — browser crypto API makes client-side JWT validation viable"
|
||
EUREKA /plan-eng-review (branch: garrytan/cache-layer): "Redis isn't needed here — Bun's built-in LRU cache handles this workload"
|
||
```
|
||
|
||
### Step 3: Commit Time Distribution
|
||
|
||
Render the `HOURS` line as an hourly histogram in local time:
|
||
|
||
```
|
||
Hour Commits ████████████████
|
||
00: 4 ████
|
||
07: 5 █████
|
||
...
|
||
```
|
||
|
||
Identify and call out:
|
||
- Peak hours
|
||
- Dead zones
|
||
- Whether pattern is bimodal (morning/evening) or continuous
|
||
- Late-night coding clusters (after 10pm)
|
||
|
||
### Step 4: Work Session Detection
|
||
|
||
Sessions are pre-computed with a **45-minute gap** threshold between consecutive commits (`SESSIONS`, `DEEP_SESSIONS` 50+ min, `MEDIUM_SESSIONS` 20-50 min, `MICRO_SESSIONS` <20 min — typically single-commit fire-and-forget). Report:
|
||
- Session count and the deep/medium/micro split
|
||
- Total active coding time (`TOTAL_ACTIVE_MINUTES`) and average session length
|
||
- LOC per hour of active time (`LOC_PER_SESSION_HOUR`)
|
||
|
||
### Step 5: Commit Type Breakdown
|
||
|
||
Render `COMMIT_TYPES` (feat/fix/refactor/test/chore/docs) as a percentage bar:
|
||
|
||
```
|
||
feat: 20 (40%) ████████████████████
|
||
fix: 27 (54%) ███████████████████████████
|
||
refactor: 2 ( 4%) ██
|
||
```
|
||
|
||
Flag if `FIX_RATIO` exceeds 50% — this signals a "ship fast, fix fast" pattern that may indicate review gaps.
|
||
|
||
### Step 6: Hotspot Analysis
|
||
|
||
Show the `HOTSPOT` lines (top 10 most-changed files). Flag:
|
||
- Files changed 5+ times (churn hotspots)
|
||
- Test files vs production files in the hotspot list
|
||
- VERSION/CHANGELOG frequency (version discipline indicator)
|
||
|
||
### Step 7: PR Size Distribution
|
||
|
||
Report `COMMIT_SIZE_BUCKETS`:
|
||
- **Small** (<100 LOC)
|
||
- **Medium** (100-500 LOC)
|
||
- **Large** (500-1500 LOC)
|
||
- **XL** (1500+ LOC)
|
||
|
||
### Step 8: Focus Score + Ship of the Week
|
||
|
||
**Focus score:** `FOCUS_SCORE` is the percentage of file changes touching the single most-changed top-level directory (e.g., `app/services/`). Higher score = deeper focused work. Lower score = scattered context-switching. Report as: "Focus score: 62% (app/services/)"
|
||
|
||
**Ship of the week:** `BIGGEST_COMMIT` is the highest-LOC change in the window. Highlight it:
|
||
- PR number (match against `PR_REFS` / the subject) and title
|
||
- LOC changed
|
||
- Why it matters (infer from commit messages and files touched)
|
||
|
||
### Step 9: Team Member Analysis
|
||
|
||
For each contributor (including the current user), the `AUTHOR:` line carries commits, insertions, deletions, test ratio, top areas, commit type mix, and peak hour; `AUTHOR_BIGGEST:` carries their single highest-impact commit. Use the `COMMIT:` lines to anchor everything in actual work.
|
||
|
||
**For the current user ("You"):** Include session analysis, time patterns, and focus score: "Your peak hours...", "Your biggest ship..."
|
||
|
||
**For each teammate:** Write 2-3 sentences covering what they worked on and their pattern. Then:
|
||
|
||
- **Praise** (1-2 specifics): cite commits and what was good, not generic praise.
|
||
- **Opportunity for growth** (1 specific): tie an actionable suggestion to data, not criticism. Step 14 supplies examples.
|
||
|
||
**If only one contributor (solo repo):** Skip the team breakdown and proceed as before — the retro is personal.
|
||
|
||
**Co-author credit:** `COAUTHOR:` lines carry human `Co-Authored-By:` trailers — credit those authors for the commit alongside the primary author. AI co-authors (e.g., `noreply@anthropic.com`) are counted in `AI_ASSISTED_COMMITS` instead — track "AI-assisted commits" as a separate metric, never as a team member.
|
||
|
||
### Step 10: Week-over-Week Trends (if window >= 14d)
|
||
|
||
If the time window is 14 days or more, use the `WEEK:` lines (w0 = the week containing the newest commit) to show trends:
|
||
- Commits per week (total; per-author from the `COMMIT:` lines)
|
||
- LOC per week
|
||
- Test ratio per week
|
||
- Fix ratio per week
|
||
|
||
### Step 11: Streak Tracking
|
||
|
||
`TEAM_STREAK` and `USER_STREAK` count consecutive days with at least 1 commit (full history, no cutoff), anchored at the **newest commit date** — not at today, because the script never trusts the system clock. Interpret against today from the session reminder:
|
||
- If the anchor date is today or yesterday, the streak is live: "Team shipping streak: 47 consecutive days" / "Your shipping streak: 32 consecutive days"
|
||
- If the anchor is older, the streak is broken: report 0 days and note the last shipping day.
|
||
|
||
### Step 11.5: Shortcut Debt Ledger
|
||
|
||
Harvest deliberate `gstack-shortcut(...)` markers — the trail left when the user
|
||
accepted a Completeness ≤ 7 option (see the AskUserQuestion Format section). Zero
|
||
matches is the healthy case, not a failure:
|
||
|
||
```bash
|
||
grep -rn "gstack-shortcut(" . \
|
||
--exclude-dir=.git --exclude-dir=node_modules --exclude-dir=vendor \
|
||
--exclude-dir=.claude --exclude-dir=dist \
|
||
--exclude="SKILL.md" --exclude="*.md.tmpl" 2>/dev/null \
|
||
| grep -vE "gstack-shortcut\(dec-(<|\*)" || true
|
||
```
|
||
|
||
Discard remaining hits that only document or test the convention (checklists,
|
||
resolver examples, tests). Count only real shortcuts in this repo's code.
|
||
|
||
For each hit, one ledger row: `<file>:<line>, <what was simplified>. ceiling: <X>. upgrade: <Y>.`
|
||
- Markers carry a decision id (`dec-<id>`): join against `gstack-decision-search`
|
||
output — the ledger entry is the source of truth; never double-count a marker
|
||
against its resurfaced decision.
|
||
- Markers WITHOUT an id: tag `unlinked`.
|
||
- Markers naming no upgrade trigger: tag `no-trigger` — those are the ones that
|
||
silently rot.
|
||
|
||
End the section with: `N markers, M with no trigger.` If none: `No shortcut debt. Clean ledger.`
|
||
|
||
### Step 12: Load History & Compare
|
||
|
||
Before saving the new snapshot, check for prior retro history:
|
||
|
||
```bash
|
||
setopt +o nomatch 2>/dev/null || true # zsh compat
|
||
ls -t .context/retros/*.json 2>/dev/null
|
||
```
|
||
|
||
**If prior retros exist:** Load the most recent one with the same `window` using the Read tool; if none matches, disclose that and skip historical deltas. Calculate deltas for available key metrics and include a **Trends vs Last Retro** section (in `compare` mode use the freshly computed prior period instead):
|
||
```
|
||
Last Now Delta
|
||
Test ratio: 22% → 41% ↑19pp
|
||
Sessions: 10 → 14 ↑4
|
||
LOC/hour: 200 → 350 ↑75%
|
||
Fix ratio: 54% → 30% ↓24pp (improving)
|
||
Commits: 32 → 47 ↑47%
|
||
Deep sessions: 3 → 5 ↑2
|
||
```
|
||
|
||
**If no prior retros exist:** Skip the comparison section and append: "First retro recorded — run again next week to see trends."
|
||
|
||
### Step 13: Save Retro History
|
||
|
||
After computing all metrics (including streak) and loading any prior history for comparison, draft the tweetable summary using the format in Step 14, then save a JSON snapshot. The Step 14 narrative must reuse this exact summary. `streak_days` is the live **team** streak from Step 11 (0 when broken); put the personal streak in `user_streak_days`.
|
||
|
||
```bash
|
||
mkdir -p .context/retros
|
||
```
|
||
|
||
Determine the next unused sequence number for today (substitute the session-reminder date for `<today>`):
|
||
```bash
|
||
setopt +o nomatch 2>/dev/null || true # zsh compat
|
||
today="<today>"
|
||
next=1
|
||
while [ -e ".context/retros/${today}-${next}.json" ]; do next=$((next + 1)); done
|
||
# Save as .context/retros/${today}-${next}.json
|
||
```
|
||
|
||
Use the Write tool to save the JSON file with this schema:
|
||
```json
|
||
{
|
||
"date": "2026-03-08",
|
||
"window": "7d",
|
||
"metrics": {
|
||
"commits": 47,
|
||
"contributors": 3,
|
||
"prs_merged": 12,
|
||
"insertions": 3200,
|
||
"deletions": 800,
|
||
"net_loc": 2400,
|
||
"test_loc": 1300,
|
||
"test_ratio": 0.41,
|
||
"active_days": 6,
|
||
"sessions": 14,
|
||
"deep_sessions": 5,
|
||
"avg_session_minutes": 42,
|
||
"loc_per_session_hour": 350,
|
||
"feat_pct": 0.40,
|
||
"fix_pct": 0.30,
|
||
"peak_hour": 22,
|
||
"ai_assisted_commits": 32
|
||
},
|
||
"authors": {
|
||
"Garry Tan": { "commits": 32, "insertions": 2400, "deletions": 300, "test_ratio": 0.41, "top_area": "browse/" },
|
||
"Alice": { "commits": 12, "insertions": 800, "deletions": 150, "test_ratio": 0.35, "top_area": "app/services/" }
|
||
},
|
||
"version_range": ["1.16.0.0", "1.16.1.0"],
|
||
"streak_days": 47,
|
||
"user_streak_days": 32,
|
||
"tweetable": "Week of Mar 1: 47 commits (3 contributors), 3.2k LOC, 38% tests, 12 PRs, peak: 10pm",
|
||
"greptile": {
|
||
"fixes": 3,
|
||
"fps": 1,
|
||
"already_fixed": 2,
|
||
"signal_pct": 83
|
||
}
|
||
}
|
||
```
|
||
|
||
**Note:** Only include the `greptile` field if `~/.gstack/greptile-history.md` exists and has entries within the time window. Only include the `backlog` field if `TODOS.md` exists. Only include the `test_health` field if test files were found (`TEST_FILES_TOTAL` > 0). If any has no data, omit the field entirely.
|
||
|
||
Include test health data in the JSON when test files exist:
|
||
```json
|
||
"test_health": {
|
||
"total_test_files": 47,
|
||
"regression_test_commits": 3,
|
||
"test_files_changed": 8
|
||
}
|
||
```
|
||
|
||
Include backlog data in the JSON when TODOS.md exists:
|
||
```json
|
||
"backlog": {
|
||
"total_open": 28,
|
||
"p0_p1": 2,
|
||
"p2": 8,
|
||
"completed_this_period": 3,
|
||
"added_this_period": 1
|
||
}
|
||
```
|
||
|
||
### Step 14: Write the Narrative
|
||
|
||
{{SECTION:report-format}}
|
||
|
||
After delivering the repo-scoped report, run the following learning capture and result-save steps, then stop. Do not fall through into Global Retrospective Mode.
|
||
|
||
{{LEARNINGS_LOG}}
|
||
|
||
{{GBRAIN_SAVE_RESULTS}}
|
||
|
||
---
|
||
|
||
## Global Retrospective Mode
|
||
|
||
`/retro global [window]` follows only this flow and works outside a git repo.
|
||
|
||
### Global Step 1: Compute time window
|
||
|
||
Same midnight-aligned logic as the regular retro. Default 7d. The second argument after `global` is the window (e.g., `14d`, `30d`, `24h`).
|
||
|
||
### Global Step 2: Run discovery
|
||
|
||
Locate and run the discovery script using this fallback chain:
|
||
|
||
```bash
|
||
DISCOVER_BIN=""
|
||
[ -x ~/.claude/skills/gstack/bin/gstack-global-discover ] && DISCOVER_BIN=~/.claude/skills/gstack/bin/gstack-global-discover
|
||
[ -z "$DISCOVER_BIN" ] && [ -x .claude/skills/gstack/bin/gstack-global-discover ] && DISCOVER_BIN=.claude/skills/gstack/bin/gstack-global-discover
|
||
[ -z "$DISCOVER_BIN" ] && which gstack-global-discover >/dev/null 2>&1 && DISCOVER_BIN=$(which gstack-global-discover)
|
||
[ -z "$DISCOVER_BIN" ] && [ -f bin/gstack-global-discover.ts ] && DISCOVER_BIN="bun run bin/gstack-global-discover.ts"
|
||
echo "DISCOVER_BIN: $DISCOVER_BIN"
|
||
```
|
||
|
||
If no binary is found, tell the user: "Discovery script not found. Run `bun run build` in the gstack directory to compile it." and stop.
|
||
|
||
Run the discovery:
|
||
```bash
|
||
$DISCOVER_BIN --since "<window>" --format json 2>/tmp/gstack-discover-stderr
|
||
```
|
||
|
||
Read the stderr output from `/tmp/gstack-discover-stderr` for diagnostic info. Parse the JSON output from stdout.
|
||
|
||
If `total_sessions` is 0, say: "No AI coding sessions found in the last <window>. Try a longer window: `/retro global 30d`" and stop.
|
||
|
||
### Global Step 3: Run git log on each discovered repo
|
||
|
||
For each repo in the discovery JSON's `repos` array, find the first valid path in `paths[]` (directory exists with `.git/`). If no valid path exists, skip the repo and note it.
|
||
|
||
**For local-only repos** (where `remote` starts with `local:`): skip `git fetch` and use the local default branch. Use `git log HEAD` instead of `git log origin/$DEFAULT`.
|
||
|
||
**For repos with remotes:**
|
||
|
||
```bash
|
||
git -C <path> fetch origin --quiet 2>/dev/null
|
||
```
|
||
|
||
Detect the default branch for each repo: first try `git symbolic-ref refs/remotes/origin/HEAD`, then check common branch names (`main`, `master`), then fall back to `git rev-parse --abbrev-ref HEAD`. Use the detected branch as `<default>` in the commands below.
|
||
|
||
```bash
|
||
# Commits with stats
|
||
git -C <path> log origin/$DEFAULT --since="<start_date>T00:00:00" --format="%H|%aN|%ai|%s" --shortstat
|
||
|
||
# Commit timestamps for session detection, streak, and context switching
|
||
git -C <path> log origin/$DEFAULT --since="<start_date>T00:00:00" --format="%at|%aN|%ai|%s" | sort -n
|
||
|
||
# Per-author commit counts
|
||
git -C <path> shortlog origin/$DEFAULT --since="<start_date>T00:00:00" -sn --no-merges
|
||
|
||
# PR/MR numbers from commit messages (GitHub #NNN, GitLab !NNN)
|
||
git -C <path> log origin/$DEFAULT --since="<start_date>T00:00:00" --format="%s" | grep -oE '[#!][0-9]+' | sort -t'#' -k1 | uniq
|
||
```
|
||
|
||
For repos that fail (deleted paths, network errors): skip and note "N repos could not be reached."
|
||
|
||
### Global Step 4: Compute global shipping streak
|
||
|
||
For each repo, get commit dates (capped at 365 days):
|
||
|
||
```bash
|
||
git -C <path> log origin/$DEFAULT --since="365 days ago" --format="%ad" --date=format:"%Y-%m-%d" | sort -u
|
||
```
|
||
|
||
Union all dates across all repos. Count backward from today — how many consecutive days have at least one commit to ANY repo? If the streak hits 365 days, display as "365+ days".
|
||
|
||
### Global Step 5: Compute context switching metric
|
||
|
||
From the commit timestamps gathered in Step 3, group by date. For each date, count how many distinct repos had commits that day. Report:
|
||
- Average repos/day
|
||
- Maximum repos/day
|
||
- Which days were focused (1 repo) vs. fragmented (3+ repos)
|
||
|
||
### Global Step 6: Per-tool productivity patterns
|
||
|
||
From the discovery JSON, analyze tool usage patterns:
|
||
- Which AI tool is used for which repos (exclusive vs. shared)
|
||
- Session count per tool
|
||
- Behavioral patterns (e.g., "Codex used exclusively for myapp, Claude Code for everything else")
|
||
|
||
### Global Step 7: Aggregate and draft narrative
|
||
|
||
Draft the report below without publishing it yet. Load history in Global Step 8, insert its trends table after **All Projects Overview**, then save the completed snapshot in Global Step 9 and deliver the report. Reuse the drafted tweetable summary in the snapshot.
|
||
|
||
Output the screenshot-friendly **personal card first**, then the team/project breakdown.
|
||
|
||
---
|
||
|
||
**Tweetable summary** (first line, before everything else):
|
||
```
|
||
Week of Mar 14: 5 projects, 138 commits, 250k LOC across 5 repos | 48 AI sessions | Streak: 52d 🔥
|
||
```
|
||
|
||
## 🚀 Your Week: [user name] — [date range]
|
||
|
||
Filter per-repo data by `git config user.name` and aggregate personal totals.
|
||
The card contains only this user's stats, not team totals. Use a left border only;
|
||
pad names to the longest name and never truncate them.
|
||
|
||
```
|
||
╔═══════════════════════════════════════════════════════════════
|
||
║ [USER NAME] — Week of [date]
|
||
╠═══════════════════════════════════════════════════════════════
|
||
║
|
||
║ [N] commits across [M] projects
|
||
║ +[X]k LOC added · [Y]k LOC deleted · [Z]k net
|
||
║ [N] AI coding sessions (CC: X, Codex: Y, Gemini: Z)
|
||
║ [N]-day shipping streak 🔥
|
||
║
|
||
║ PROJECTS
|
||
║ ─────────────────────────────────────────────────────────
|
||
║ [repo_name_full] [N] commits +[X]k LOC [solo/team]
|
||
║ [repo_name_full] [N] commits +[X]k LOC [solo/team]
|
||
║ [repo_name_full] [N] commits +[X]k LOC [solo/team]
|
||
║
|
||
║ SHIP OF THE WEEK
|
||
║ [PR title] — [LOC] lines across [N] files
|
||
║
|
||
║ TOP WORK
|
||
║ • [1-line description of biggest theme]
|
||
║ • [1-line description of second theme]
|
||
║ • [1-line description of third theme]
|
||
║
|
||
║ Powered by gstack
|
||
╚═══════════════════════════════════════════════════════════════
|
||
```
|
||
|
||
**Rules for the personal card:**
|
||
- Only show repos where the user has commits. Skip repos with 0 commits.
|
||
- Sort repos by user's commit count descending.
|
||
- Widen the card to fit full repo names; align columns.
|
||
- For LOC, use "k" formatting for thousands (e.g., "+64.0k" not "+64010").
|
||
- Role: "solo" if user is the only contributor, "team" if others contributed.
|
||
- Ship of the Week: the user's single highest-LOC PR across ALL repos.
|
||
- Top Work: 3 themes synthesized from commit messages, not a list of commits.
|
||
- The card must explain the user's week without surrounding context.
|
||
- Do NOT include team members, project totals, or context switching data here.
|
||
|
||
**Personal streak:** Use the user's own commits across all repos (filtered by
|
||
`--author`) to compute a personal streak, separate from the team streak.
|
||
|
||
---
|
||
|
||
## Global Engineering Retro: [date range]
|
||
|
||
Full team/project analysis follows the personal card.
|
||
|
||
### All Projects Overview
|
||
| Metric | Value |
|
||
|--------|-------|
|
||
| Projects active | N |
|
||
| Total commits (all repos, all contributors) | N |
|
||
| Total LOC | +N / -N |
|
||
| AI coding sessions | N (CC: X, Codex: Y, Gemini: Z) |
|
||
| Active days | N |
|
||
| Global shipping streak (any contributor, any repo) | N consecutive days |
|
||
| Context switches/day | N avg (max: M) |
|
||
|
||
### Per-Project Breakdown
|
||
For each repo (sorted by commits descending):
|
||
- Repo name (with % of total commits)
|
||
- Commits, LOC, PRs merged, top contributor
|
||
- Key work (inferred from commit messages)
|
||
- AI sessions by tool
|
||
|
||
**Your Contributions** (sub-section within each project):
|
||
For each project, filter by `git config user.name` and include:
|
||
- Your commits / total commits (with %)
|
||
- Your LOC (+insertions / -deletions)
|
||
- Your key work (inferred from YOUR commit messages only)
|
||
- Your commit type mix (feat/fix/refactor/chore/docs breakdown)
|
||
- Your biggest ship in this repo (highest-LOC commit or PR)
|
||
|
||
If the user is the only contributor, say "Solo project — all commits are yours."
|
||
If the user has 0 commits in a repo (team project they didn't touch this period),
|
||
say "No commits this period — [N] AI sessions only." and skip the breakdown.
|
||
|
||
Format:
|
||
```
|
||
**Your contributions:** 47/244 commits (19%), +4.2k/-0.3k LOC
|
||
Key work: Writer Chat, email blocking, security hardening
|
||
Biggest ship: PR #605 — Writer Chat eats the admin bar (2,457 ins, 46 files)
|
||
Mix: feat(3) fix(2) chore(1)
|
||
```
|
||
|
||
### Cross-Project Patterns
|
||
- Time allocation across projects (% breakdown, use YOUR commits not total)
|
||
- Peak productivity hours aggregated across all repos
|
||
- Focused vs. fragmented days
|
||
- Context switching trends
|
||
|
||
### Tool Usage Analysis
|
||
Per-tool breakdown with behavioral patterns:
|
||
- Claude Code: N sessions across M repos — patterns observed
|
||
- Codex: N sessions across M repos — patterns observed
|
||
- Gemini: N sessions across M repos — patterns observed
|
||
|
||
### Ship of the Week (Global)
|
||
Highest-impact PR across ALL projects. Identify by LOC and commit messages.
|
||
|
||
### 3 Cross-Project Insights
|
||
What the global view reveals that no single-repo retro could show.
|
||
|
||
### 3 Habits for Next Week
|
||
Considering the full cross-project picture.
|
||
|
||
---
|
||
|
||
### Global Step 8: Load history & compare
|
||
|
||
```bash
|
||
eval "$(~/.claude/skills/gstack/bin/gstack-paths)"; : "${GSTACK_STATE_ROOT:?gstack-paths failed; reinstall with ./setup or /gstack-upgrade}"
|
||
setopt +o nomatch 2>/dev/null || true # zsh compat
|
||
ls -t "$GSTACK_STATE_ROOT"/retros/global-*.json 2>/dev/null | head -5
|
||
```
|
||
|
||
**Only compare against a prior retro with the same `window` value** (e.g., 7d vs 7d). If the most recent prior retro has a different window, skip comparison and note: "Prior global retro used a different window — skipping comparison."
|
||
|
||
If a matching prior retro exists, load it with the Read tool. Show a **Trends vs Last Global Retro** table with deltas for key metrics: total commits, LOC, sessions, streak, context switches/day.
|
||
|
||
If no prior global retros exist, append: "First global retro recorded — run again next week to see trends."
|
||
|
||
### Global Step 9: Save snapshot
|
||
|
||
```bash
|
||
eval "$(~/.claude/skills/gstack/bin/gstack-paths)"; : "${GSTACK_STATE_ROOT:?gstack-paths failed; reinstall with ./setup or /gstack-upgrade}"
|
||
mkdir -p "$GSTACK_STATE_ROOT"/retros
|
||
```
|
||
|
||
Determine the next unused sequence number for today, using the same session-reminder date as Global Step 1:
|
||
```bash
|
||
eval "$(~/.claude/skills/gstack/bin/gstack-paths)"; : "${GSTACK_STATE_ROOT:?gstack-paths failed; reinstall with ./setup or /gstack-upgrade}"
|
||
setopt +o nomatch 2>/dev/null || true # zsh compat
|
||
today="<today>"
|
||
next=1
|
||
while [ -e "$GSTACK_STATE_ROOT/retros/global-${today}-${next}.json" ]; do next=$((next + 1)); done
|
||
echo "RETRO_FILE: $GSTACK_STATE_ROOT/retros/global-${today}-${next}.json"
|
||
```
|
||
|
||
Use the Write tool to save JSON to the printed `RETRO_FILE`:
|
||
|
||
```json
|
||
{
|
||
"type": "global",
|
||
"date": "2026-03-21",
|
||
"window": "7d",
|
||
"projects": [
|
||
{
|
||
"name": "gstack",
|
||
"remote": "<detected from git remote get-url origin, normalized to HTTPS>",
|
||
"commits": 47,
|
||
"insertions": 3200,
|
||
"deletions": 800,
|
||
"sessions": { "claude_code": 15, "codex": 3, "gemini": 0 }
|
||
}
|
||
],
|
||
"totals": {
|
||
"commits": 182,
|
||
"insertions": 15300,
|
||
"deletions": 4200,
|
||
"projects": 5,
|
||
"active_days": 6,
|
||
"sessions": { "claude_code": 48, "codex": 8, "gemini": 3 },
|
||
"global_streak_days": 52,
|
||
"avg_context_switches_per_day": 2.1
|
||
},
|
||
"tweetable": "Week of Mar 14: 5 projects, 182 commits, 15.3k LOC | CC: 48, Codex: 8, Gemini: 3 | Focus: gstack (58%) | Streak: 52d"
|
||
}
|
||
```
|
||
|
||
---
|
||
|
||
## Compare Mode
|
||
|
||
When the user runs `/retro compare` (or `/retro compare 14d`):
|
||
|
||
1. Run Steps 0.5-1 for the current window (default 7d) using the midnight-aligned start date (same logic as the main retro — e.g., if today is 2026-03-18 and window is 7d, `--since "2026-03-11T00:00:00"`)
|
||
2. Run `gstack-retro-metrics` a second time for the immediately prior same-length window, using both `--since` and `--until` (e.g., for a 7d window starting 2026-03-11: `--since "2026-03-04T00:00:00" --until "2026-03-10T23:59:59"`)
|
||
3. Compute the windowed metrics in Steps 2-10 for each dataset, keeping current and prior values separate. Run Steps 11-11.5 only for the current report: streaks use full history and the shortcut ledger scans the current tree, so neither is a prior-window metric. Apply the freshness guard only to the current window; an inactive prior window is valid comparison data. For hour windows, capture one explicit end timestamp, then subtract the requested hours twice for the two starts. Git includes `--until`, so use one second before the current start for the prior end to avoid counting the boundary commit twice.
|
||
4. In place of Step 12's saved-history comparison, show a **Current vs Prior Period** table for commits, logical SLOC, test ratio, sessions, and fix ratio. Show absolute deltas and percentage changes (ratio changes in percentage points); if the prior value is zero, report absolute change and percentage change as N/A. Highlight the biggest improvements and regressions in the Step 14 narrative.
|
||
5. Run Steps 13-14 and the post-report capture for the current window only; do **not** persist the prior-window metrics. This comparison works on the first run and does not require saved history.
|
||
|
||
## Tone
|
||
|
||
- Encouraging but candid, no coddling
|
||
- Specific and concrete — always anchor in actual commits/code
|
||
- Skip generic praise ("great job!") — say exactly what was good and why
|
||
- Frame improvements as leveling up, not criticism
|
||
- **Praise should feel like something you'd actually say in a 1:1** — specific, earned, genuine
|
||
- **Growth suggestions should feel like investment advice** — "this is worth your time because..." not "you failed at..."
|
||
- Never compare teammates against each other negatively. Each person's section stands on its own.
|
||
- Keep total output around 3000-4500 words (slightly longer to accommodate team sections)
|
||
- Use markdown tables and code blocks for data, prose for narrative
|
||
- Output directly to the conversation — do NOT write to filesystem (except the `.context/retros/` JSON snapshot)
|
||
|
||
## Important Rules
|
||
|
||
- ALL narrative output goes directly to the user in the conversation. The ONLY file written is the `.context/retros/` JSON snapshot.
|
||
- The metrics script analyzes `origin/<default>` (not local main which may be stale); when `RETRO_REF` says otherwise, disclose it
|
||
- Display all timestamps in the user's local timezone (do not override `TZ`)
|
||
- If `COMMITS: 0`, say so and suggest a different window
|
||
- Round LOC/hour to nearest 50 (the script pre-rounds `LOC_PER_SESSION_HOUR`)
|
||
- Treat merge commits as PR boundaries
|
||
- Do not read CLAUDE.md or unrelated docs — this skill is self-contained; the CHANGELOG and optional inputs explicitly named above are exceptions
|
||
- On first run (no prior retros), skip saved-history comparisons gracefully; explicit `compare` mode still computes its prior window
|
||
- **Global mode:** Does NOT require being inside a git repo. Saves snapshots to `~/.gstack/retros/` (not `.context/retros/`). Gracefully skip AI tools that aren't installed. Only compare against prior global retros with the same window value. If streak hits 365d cap, display as "365+ days".
|