mirror of
https://github.com/garrytan/gstack.git
synced 2026-10-04 18:36:54 +02:00
7fca42ad8b6c707b8a38f579f72bf3c4f7de6d85
10
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
7fca42ad8b |
v1.91.12.0 v1.91.12.0: audit fix wave, ~11-minute paid eval lanes, eval reliability policy (#2999)
* test: delete test-infrastructure dead code (G) - exit-propagation drives the runner's real strict verdict (BunTestOutputClassifier + strictTestExitCode); delete the unused shardRunLooksTruncated predicate. - delete skill-coverage-matrix registry + its gate (nothing reads it; the floor already iterates skillCensus()). - delete touchfiles-facade export-parity tests (Bun fails missing imports at link time) and the duplicated E2E_TIERS tier-value test. - delete brain-cache-spec TRANSPORT_DEFAULT_POLICY, SKILL_RUN_RETENTION_DAYS and the now-unused BrainTrustPolicy type with their literal tests. AUTOPLAN_PREFLIGHT_BUDGET_BYTES stays: skill-preflight-budget enforces it against real resolver output. - delete audit-compliance's JSDoc-comment grep. * test: replace product tests that fake the product with real-boundary tests (F) - design: serve.test.ts drove an inline mirror server; now two tests run the real serve() on an ephemeral port (reload confinement, submit exit 0). - setup-gbrain: rollback + voyage tests execute the template-extracted init blocks (3 sites) instead of drifted local bash copies. - terminal-agent: internalHandler source greps replaced by a behavioral /internal/grant + /internal/revoke auth matrix (no/wrong/valid token). - /health: server-security-surface and the server-auth / security-audit-r2 / sidebar-tabs source greps fold into one liveness-only check on the real body; the L4 sidecar wiring gets a behavioral /pty-inject-scan test. - delete tautologies (browser-manager onDisconnect, memory-command #12), ios swiftui tap fixture self-check, memory-ingest put_page grep, detach source greps, sidebar-agent absence pins, dead-CSS pins + the dead CSS, security-audit-r2 Task 1 + the test-only meta-commands re-export, duplicate generated-SKILL.md checks. - make-pdf coverage-gaps cases move into their owner test files. * test: delete tests of dead eval code (A) - A1: the retired Eng lexical oracle (evaluateEngSeedCoverage, isEngSeedDecisionAUQ), the completion-handoff detector and the retained corpus had no paid caller since v1.87.6; delete their 26 replay files, ~2.6k helper LOC and fixtures, and the dead blocks in 8 mixed files (live hasNativePlanTerminal / batching assertions stay). - A2: dead viewport approvers in autoplan-artifact-permission and their 11 replay files + fixtures; recorder/launcher cases stay. - A3: never-wired oracles and seeders (autoplan-phase-order, eng-finding-fixture, ceo-paired-fixture, design-ui-scope, plan-skill-completion, pty-current-screen, required-reads, transcript-section-logger); plan-seed-submission now decodes through the production createPtyScreen; section manifests name their actual guard. - A4: zero-reference helper exports, plus execGit and invokeAndObserve found by the reachability pass. - 52 fixtures orphaned by the deletions; touchfile and selection-table entries for every deleted path. * test: clean up the paid eval lane (B1-B4, B6, B7) - B1: delete paid files that assert nothing or cannot pass meaningfully: skill-llm-eval-spec and skill-e2e-spec-execute (test.todo), gemini-e2e (+ gemini-session-runner; no gemini CLI in CI), ship-idempotency (red since v1.63), the two opus-4-7 *-sonnet overlay wrappers, conductor-prose (+ its source-evaluation replay), codex-e2e-plan-format; drop their keys, scripts and census rows. - B2: skill-llm-eval grades browse/sections/command-list.md with one union judge that also carries the baseline score pin; regression-vs-baseline deleted (paid run: pass, c4/c4/a4). - B3: memory-pipeline, ios-qa, ios-qa-swift-build and plan-tune-cathedral make no model calls; renamed out of the paid glob so they run on every PR. Swift builds need GSTACK_TEST_SWIFT=1; device stub deleted. - B4: codex-e2e*, outside-voice, aside and ios-device cannot run in the CI image; excluded from the weekly lane with a tracked re-entry condition. - B6: fold opus-47's negative routing controls into skill-routing-e2e journey-negatives (paid run: 3/3 unrouted) and delete the file. - B7: delete the never-green brain-privacy-gate eval; a free gstack-skill-start test now proves consent precedes artifacts egress. * test: retire the finding-count cluster and trim its helpers (C) - C0/C1: the five never-green evals (skill-e2e-autoplan-chain and skill-e2e-plan-{ceo,eng,design,devex}-finding-count) failed on harness and budget, never on skill behavior; delete them, their touchfile/tier ids, AUTOPLAN_CHAIN_BUDGET and the dedicated eighth periodic slice (--slices 7). - C2: delete the helper groups whose only paid consumers were those files (11 modules), trim claude-pty-runner and eng-seeded-coverage to the paid closure, and delete the free replay tests whose assertions exercised only that dead code (89 files, 135 orphaned fixtures). Blocks that used dead code only as input for a live subject keep their assertions: the multiSelect default moved to plan-review-decisions, runner PTY tests use inline caller policies, and the timer-safe budget checks moved to eng-finding-retry-budget. - The eight production-touching files stay except ceo-current-decision-record (its template read only feeds the retired counter). - CARVE_GUARDS.autoplan is behavioral 'none'; TODOS records the lost chain and per-finding cadence coverage with their re-entry tests. * test: fold per-incident replay series into their detector owners (D) Twelve detector families move into one owner test each: 73 incident files become describe blocks in ceo-section-loading-fixture (stale-fill race), model-overlays, coverage-audit-evidence, autoplan-phase-observer, native-auto-decide, outside-voice-evidence, eng-first-review, plan-count-completion, plan-count-file-permission, ceo-mode-option, plan-scope-selection and plan-count-prerequisite. Each block keeps its original code and fixture, so every case still runs; only tests asserting the incident file's own touchfile registration are dropped (41). Touchfile lists that named an incident now name its owner. * test: start the plan-count history PTY on its readiness marker (H) The fake CLI prints a startup marker and the runner waits for it instead of the fixed 8 s startup sleep (8.6 s -> 0.9 s locally). eng-semantic-terminal's sleeping registration cases went with C; plan-count-timeout keeps the fixed wait because it asserts deadline behavior. * test: derive paid touchfiles from each eval's static closure (E) touchfiles.test.ts now checks, per key, that the paid file's static test/helpers and test/fixtures closure (plus fixture paths it names in string literals) is covered, and names the file, path, chain and key to fix when it is not. Free *.test.ts files are no longer touchfiles, so editing a free replay test stops selecting paid evals: 950 entries removed, 653 real closure paths added. The hand-copied inventories go: periodic-fixture-selection, fake-impeccable-touchfiles and 45 per-file selection examples. Selection for the sample edits (plan-eng-review template, claude-pty-runner, plan-count-fixture, gstack-config) loses no case under either profile. CONTRIBUTING documents the rule and its lower bound. * test: skip hollow tier shards and census judges in the paid planner (B5) A paid file is now skipped for a tier lane only when every E2E id it registers is known statically and none has that tier; ids come from the touchfile registrations and literal testName/*IfSelected arguments, so a comment or skill path that quotes another id cannot unschedule it, and computed names keep today's scheduling. --list and the manifest show each skip as "skipped: no E2E_TIERS id has tier <tier>". The weekly gate census drops the LLM judges (--skip-judges); they still run in the periodic census and PR gate lanes. Gate lane 52 -> 42 files, census 41; periodic 77 -> 69. * test: run seven paid evals on the current default capture model (B8) skill-e2e-{auq-matrix,plan-format,qa-bugs,retro,workflow} pinned claude-opus-4-7 and skill-e2e-office-hours plus -brain-writeback pinned claude-sonnet-4-6; none tests a historical model, so they now capture with resolveEvalModel('capture'), and the free harness tests that execute these registrations receive the same resolver. The paid re-pin run passed all of them. skill-e2e-{design,office-hours-phase4,plan-prosons,plan} keep claude-opus-4-7: six of their cases failed on the default model (three timeouts, a missing report file, a format miss and a posture score of 3), so per the plan's fallback they keep their pins with a TODOS entry. The pre-spend estimate and drop threshold are in docs/test-audit-2026-09.md. * test: guard the reduced suite against new test-of-test files - test/test-of-test-ratchet.test.ts records the 228 free tests that import only test/ code and fails on a new one, naming the owner test to extend instead; a stale baseline entry fails with the remove instruction. - test/helpers/resolve-repo-path.ts is the one specifier/literal resolver for the ratchet and the touchfile closure invariant, with its own unit tests. - CONTRIBUTING "Test tiers" describes the paid-failure workflow (fix, then one row in the detector's owner test) and the ratchet; TEST_PORTFOLIO gains the detector -> owner-test table and no longer claims an Autoplan chain eval. - TODOS: automatic exclusion policy for chronically red periodic files (P3), the deferred native-completion table collapse, the unused CEO payment seeder; the PTY readiness item is narrowed to the paid runner. - docs/test-audit-2026-09.md collects the triage, security mapping, inventories, selection proof, behavior-commit decisions and retained false positives. * v1.91.8.0 test: smaller suite, derived paid selection, retired never-green evals Release metadata for the test-reduction branch: VERSION 1.91.8.0 (1.91.7.0 is claimed by #2983), CHANGELOG with the measured before/after table and a contributor section, durations re-recorded on Ubicloud standard-16 (857 files, 0 failures), the agents digest, CONTRIBUTING's after-measurement row, the B8 fallback TODOS entry, and the after metrics, kept-vs-plan notes, B8 run and census estimate in docs/test-audit-2026-09.md. * fix(ubicloud): skip retrieval globs that match nothing instead of reporting a failed pull * test: pin DISABLE_AUTOUPDATER in hermetic env and capture corrupt-seed warning Both EVALS_HERMETIC branches of buildHermeticEnv now carry DISABLE_AUTOUPDATER=1 (the allowlist scrubbed the workflow's copy, so every PTY screen showed the updater's npm-prefix failure). Per-test overrides still win. The corrupt durations-seed test now captures its expected warning and restores the console spy. * style(cso): format lib/cso TypeScript with pinned Prettier Mechanical reformat only. Minified transpile output is byte-identical for 21 of 22 files; witness.ts differs only in three regex flag orders (/mi -> /im), which JavaScript canonicalizes. Source-text assertions over lib/cso now compare whitespace-insensitively with the same tokens. * fix(cso): import join for compiled-launcher assertion witnesses Compiled installs always take the non-Bun branch, which called an unimported join and threw before any runtime-tested assertion could be witnessed. The child command selection is now a pure, platform-aware function; a missing sibling launcher fails with its expected path. * fix(browse): make connect --supervise actually respawn a crashed server The supervisor respawned with a block-scoped env that no longer existed, so every attempt threw and the loop gave up after five tries. The headed env is now one pure helper used by connect and respawn, the loop is an injectable runHeadedSupervisor with behavioral tests, failures name the daemon log and relaunch command, and connect's usage advertises --supervise. * test: one finite PR world for the shared-libs fixture; name dual-voice probe evidence The shared-libs shim served 2 PRs for pulls?state=all and endless full pages for state=open. gh pr list, pulls?state=open|all|closed (per_page/page, short last page, direction) and search/issues now page one deterministic table: PR 7, 600 older open PRs, PR 42 and 3 closed PRs, so five 100-item open-metadata pages still leave older open PRs unchecked. The Contents API lists pinned directories (the captured attempt got 404 for contents/ and contents/src while files resolved, then fell back to a raw host), unknown endpoints return 404 instead of repo metadata, and the read-only detector is unchanged. Free tests cover view agreement, the budget bound, gh/curl agreement and the empty world. Dual-voice outside-voice failures now report probeToolUseId, probeMode and the canonical-match result with the reason the probe output was rejected. * feat: require a zero-error product typecheck and a test type-debt ratchet Adds tsconfig.json (strict) over product code, fixes its remaining 90 diagnostics (type-only, interface corrections, and explicit narrowing), and adds a typecheck job to the required free-tests aggregate running bun run typecheck, the test-code ratchet (identity -> count baseline, fails on new, repeated, or unlocked fixed diagnostics), and the lib/cso format check. Reuses fixes from #2447 where they still applied. * test: follow the headed env helper and the typecheck gate in source-shape checks * fix(test): pin the package.json change kind in shared-input selection tests computePaidCaseSelection read the version-only exemption from git even when changed files were injected, so the shared-input test failed on main and on version-only branches. The exemption is now an optional input; the test pins a real package.json change and covers the version-only case. * test: judge plan-count completion on structured evidence, not wording Replaying run 36385945043's two Design attempts showed the existing routes rejected correct endings: attempt 1 at the typed-completion path field ('- Reviewed plan written to …' is not a 'Plan written to' line), attempt 2 at the leading-fence veto (its final message opens with the dashboard). nativePlanTerminalPreconditions is the structural prefix of hasNativePlanTerminal (behavior unchanged). structuredPlanCompletion adds, inside the existing nativeSummary branch: a complete report (Design binding for Design), a completed review-log row for the expected skill appended during this attempt under the child's GSTACK_HOME/project slug (resolved with bin/gstack-slug) and stamped with the fixture commit, timed between the report/last answer (second resolution) and the final native message, a final message with stop_reason end_turn (now carried on public transcript messages), and no visible question or permission prompt. Timeout summaries add idleFor and lastTerminalCandidate. Terminal and throw captures copy the plan file and review-log rows into the artifact directory; copies are best-effort and recorded in evidence-copy.json. Free regressions: both captured Design endings (trimmed fixture with provenance; report, row and end_turn reconstructed and labelled), the negative controls, and real-PTY completion/timeout runs through the real review logger. * test: structural Design count boundary; TODO proposals are not findings Replaying run 36385945043 through the Design count predicates: routing, focus and learnings setup was not recognized as setup, Issue 1 was counted pre-review in both attempts (the boundary fired on it), and attempt 2 counted the Font TODO proposal as a finding (review=4 and review=5 for five issues). The paid caller now starts review at the first answered native decision that is not setup (recognized packet, or setup header/question ID), a completion handoff, artifact rendering or a TODO proposal (the review's Add to TODOS.md / Skip / Build it now menu). TODO proposals are recorded as administrative extra decisions. The replay asserts each counted call: both attempts review=5 (Issues 1-5). isDesignCountFirstReview and its controls are unchanged. * test: CEO classifier throws name the question and matched predicates Replaying run 36385945043's FAN-1 and ERR-1 throws (ledger rows reconstructed from rendered diffs) through ceoPaymentFinding: the email obligation's row, subject, option and proposal predicates pass and the ELI10 explanation-defect predicate fails first ('lets that exception fly out', 'the error bubbles up'). Binding the defect to the named ledger row instead (the planned fix) was tried and reverted: scoped to the email seed it flips 30+ existing cf74 still-rejects replays, which require a vocabulary-free, ledger-bound email question to earn credit only through a complete saved comparison. With FAN-1's rendered currentDecision payload reconstructed, the recorded- decision path counts it, so the real saved plan (not uploaded) must have differed; failure artifacts now retain it. The classifier stays fail-closed and unchanged. Its throw now prints the header, the first 200 question characters and each obligation's predicate results. Free regressions with provenance and negative controls: an unrelated question, an email question whose row says it is already rescued, and a ledger ID whose row belongs to another seed. * chore: regenerate the test type-debt baseline on top of #2994 * fix(typecheck): strip the checkout root from ratchet diagnostic identities * fix(test): recognize ledger row-ID split candidates so collection stops at the last ACK Run 36385945043's split-overflow case asked all five candidate decisions by 8m55s, but the live candidate check required the question to open with "E1:" and every option to be a known disposition. The skill cited ledger row IDs ("D2.1 — R-E1: …") and offered "Hold, discuss first", so no candidate was recognized and the attempt ran the whole review (1302s). Identity now comes from the native header; the question must open with that candidate's ledger reference, name only that candidate, and offer exactly one include, defer and cut disposition. The selected answer must still be one of those three. The semantic evaluator and every existing negative control are unchanged; a trimmed capture from the run adds the positive case and four row-ID negative controls. * fix(test): stop the eng batching eval once its floor is proven The case's only verdict is reviewCount >= FLOOR (3). Run 36385945043 had three distinct acknowledged review decisions at 6m41s but kept answering until the ceiling (7) at 12m13s. The registration now passes the runner's existing isCollectionComplete stop once FLOOR non-setup, non-administrative review decisions are acknowledged; the floor check, ceiling, budget and counter are unchanged. A child-process registration test proves the stop predicate and that below-floor and timeout outcomes still fail. * test: add the non-blocking 'marathon' E2E tier Full start-to-finish flows move out of the blocking lanes. E2E_TIERS and E2ETier gain 'marathon'; describeE2ETier('marathon') is enabled only when EVALS_TIER=marathon, so the gate/PR and periodic lanes (and the gate census) never run those cases. The PR profile accepts marathon ids as scheduled elsewhere and defers them with their own reason, even on full fallback. * test: move the full office-hours workflow to marathon; add a periodic design-draft checkpoint The full startup workflow runs 1–3 real spec-review rounds (~280s each) and hit its 1200s capture in run 36385945043 at finalize. Review depth is the product's loop, so the case cannot fit a blocking lane without cutting rounds. It is now marathon tier with every assertion unchanged. skill-e2e-office-hours-design-draft.test.ts (periodic) runs the same fixed interview only through the Write that creates the design (269s in that run) and applies the full validator's design-draft checks, the required section reads and the launch/foreign-skill-read guards. validateOfficeHoursDesignDraft is extracted from validateOfficeHoursCompletion, which still applies it. Selection: office-hours-design-draft is registered periodic; the marathon-only file is already excluded from the gate and periodic plans by the B5 planner rule. Tier-alignment regexes and the valid-tier check accept 'marathon'. A type-only cast in plan-scope-selection.test.ts removes a diagnostic whose union print order made the ratchet identity unstable; baseline tightened. * test: supply the split-overflow fixture's HOLD SCOPE mode as a prerequisite The split actor always answered 0E's mode question with HOLD SCOPE. The skill skips that question on an explicit choice, so the fixture now states it and the attempt starts at the five candidate decisions (about 1.5 min earlier in run 36385945043). Candidates, actor policy, floor and semantic evaluation are unchanged; the fixture test pins the supplied choice. * test: start the eng batching eval with its setup prerequisites supplied Routing setup and cross-project learnings (D1/D2 in run 36385945043) are never counted and are not what the case measures. The registration now uses the runner's existing preconfiguredReviewActor so the attempt starts at the review; engSetupAUQ still vetoes any late setup question. The registration test pins the option. * test: count the design-draft paid file and defer marathon ids in PR selection pins The discovered paid-file census grows by one (skill-e2e-office-hours-design-draft). Full-fallback PR selection defers every non-gate id; the shared-input pins now expect periodic and marathon ids there. * fix(review): resolve the judged revalidation, setup-authority, plan-gate and findings-record ambiguities The census review workflow judge scored clarity/actionability 3 on both attempts: smoke-clock limits appeared to forbid post-repair revalidation, the caller deadline was undefined, 'ask for setup' conflicted with the report-only browser rule, fallback-sourced HIGH discrepancies had no gate decision, and the Step 5.8 record omitted adversarial findings. * fix(office-hours): load the builder section for every builder-mode reply Both census builder-wildness attempts answered a direct request for adjacent unlocks without reading phase-2b-builder-brainstorm.md, whose trigger read as applying only to the generative questions. * fix(sync-gbrain): define Step 4 helper args and one atomic write path Both census read-ready attempts spent turns reading the helper source to resolve <user-args>, inspecting fixture internals kept inside the repo, and reconciling 'Read + Edit' with the tmp+mv atomic write, then hit max turns before the verdict. * refactor(evals): share the import-closure walker and add the E2E shard reuse identity sourceDependencyClosure moves from the workflow-judge adapter into scripts/eval-input-cache.ts unchanged, so judge keys stay byte-identical. scripts/e2e-shard-reuse.ts builds the consumed-input identity of one PR-lane E2E shard (test import closure, every registered case's touchfiles, globals, runner/workflow/setup actions, child env pins, CI image, Claude CLI) and fails closed on anything unknown. Marathon joins the always-fresh purposes. * feat(evals): ~12-minute blocking paid lanes and a non-blocking marathon lane - Planner budget mode (--slice-budget S --jobs J): recorded per-tier wall times pack into as many ~9-minute executors as the work needs; the plan records per-slice estimates and the CI job timeout (supervised worst case + 20 min). evals.yml and evals-periodic.yml derive matrix size and timeout-minutes from it; max-parallel covers every slice at once. - Case shards: plan/design/review-army/shared-libs(-paths) run one registered case per process (<file>#<case id>, exact name pattern, exactly one case). - Retry rule: a timed-out attempt is a verdict. Only files whose every case budget is CAPTURE tier or shorter keep one retry; walls shrink to match. - Marathon tier: positive selection, excluded from gate/periodic planners, run by the new evals-marathon.yml (weekly + dispatch, fresh, own report). - PR-lane E2E reuse of verified first-attempt passes on identical inputs; the report rejects reuse outside the fast PR profile. - Duration seed from census run 36385945043, per tier and per case shard. * docs: blocking lane budget, marathon lane, retry policy and E2E reuse * chore(typecheck): lock in two fixed test diagnostics * fix(ci): drop a duplicated env/jobs block in evals-marathon.yml * test(ship-docsync): shard the doc-sync lifecycle by case and drop the duplicate dispatch-only case ship-docsync ran the same fixture and prompt as ship-docsync-completion and asserted a subset of it. The file now runs one case per process, so its lane wall is its longest case instead of half the sum of thirteen. * fix(evals): plan CI-unrunnable cases as excluded entries, not empty case shards design-review-fix drives the Aside browser and registers test.skip on Linux runners, so its case shard executed zero cases and failed the exact-one-case check in proof census 36597762183 (eval-slices 6). CASE_CI_EXCLUDE (reason + tracking, beside PERIODIC_CI_EXCLUDE) now turns such cases into excluded manifest entries that --list and the manifest surface; every planned case shard still must execute exactly its case. * docs(todos): list the case-level Aside exclusion with the CI-unrunnable evals * fix(plan-ceo-review): restore experience-first expansion framing, require the mode handoff, skip pacing menus Census 36597762183: both mode-routing runs logged provenance and moved on without the mandated handoff chat; the EXPANSION run asked an unauthorized batch/narrow pacing menu instead of the first per-addition question; the expansion-energy proposals led with the spec because v1.87.6.0 dropped 'lead with the felt experience'. The HOLD review detector also rejected a decision whose grounding line named no plan file although the owned source Read binds it. * test(outside-plan-disabled): bind quoted prior-record values by their sentence, not phrase order The parent obeyed the off switch and twice named the seeded completed record as pre-existing, once with the quotation after its owner and once with slash separators; the order-specific stripper counted both as current completion. Timestamp, location, current-claim and value-match controls still reject. * test(outside-plan-disabled): compare named record timestamps as instants; negated authorship is not a current claim The repair rerun named the seeded record by its ISO second (2026-09-29T16:58:52Z vs .727Z) and said 'I did not write'; both were misread as a foreign timestamp and a current write. * test(ceo-section-loading): recognize an arrow-ordered stale-fill execution by event roles The census review traced the seeded race as 'R1 miss -> R1 store read (v1) -> W commit v2 -> W cache.delete -> W fulfills -> R1 cache.set(v1) -> R2 (begun after W) hits v1', but the in-flight gate only accepted race vocabulary or fixed sentence shapes. Order, actor, version and dismissal mutations still fail. * test(design-floor): answer the seed-declared all-seven 0D focus menu while it is pending The actor declares 'Design: review all seven dimensions', but its picker reused designReviewSetupAUQ, which only matches already-answered calls (and a narrower header/label set), so the pending D1 focus menu was never answered and the case waited out its 609 s deadline. The skill's Step 0D requires asking; the fixture now answers it. * test(ceo-mode-routing): accept the skill-mandated Note form and Recommendation reason as HOLD posture HOLD Defer/Keep briefs must use 'Note: options differ in kind' (preamble), but the answered-HOLD path demanded a Completeness score, rejected a one-line Net with a semicolon, and read posture only from ELI10. The rerun's brief applied HOLD SCOPE in its Recommendation reason. Revert the ineffective 'always'/'handoff chat' wording: two runs still skipped the mode handoff. * test(qa-bugs): keep claude-opus-4-7 after qa-b6-static stalled on the default model qa-b6-static timed out on claude-fable-5-1 in census 36597762183 and in one of two targeted reruns. Both times the stream stopped mid-message with no pending tool, right after the model found the disabled submit button, and stayed silent until the 300 s deadline. Per the B8 fallback, re-pin with a TODOS entry; budgets and retries are unchanged. A rerun on opus-4-7 passed (125 s, 5/5 detected). * test(evals): add E2E_KINDS, BEHAVIOR_WHY, EVAL_POLICY and CASE_QUARANTINE skeletons Every E2E_TIERS and LLM_JUDGE_TOUCHFILES key starts as 'rule'; BEHAVIOR_WHY and CASE_QUARANTINE start empty. EVAL_POLICY pre-registers the approved panel (3, majority 2), quarantine entry 0.95/10 and exit 0.97/10, 10% cap, 8-weekly-run expiry, Fisher drift alarm and one INFRA re-dispatch. * test(evals): add trial records, panelVerdict, expectContract and trial-outcomes JSONL EvalTestEntry gains case_id, kind, trial, panel, failure_class and policy_version, stamped from the runner's TRIAL_ENV on isolated trial shards. panelVerdict() is the single verdict function (INCOMPLETE on missing or duplicate trials, contract veto at any count, quarantine hard-break rule, INFRA/INCOMPLETE machine classification). expectContract() records failure_class 'contract' on the collector entry and a sidecar before throwing. trial-outcomes JSONL has a fail-closed writer and a data-only reader. * test(evals): pin the fail-closed rule-shard gate through the real --report path Synthetic slice artifacts for rule fail, timeout, missing slice, unreported entry, hollow, never-started, collector failure and wrong-slice reports all exit red before the panel-verdict gate change lands. * test(evals): retire every paid automatic retry Paid evals never retry (approved 2026-09-29): delete SHORT_CASE_RETRY_FILES and retriesWithinCaseCap, drop the retry fields from the registered wall rows (walls now cover one run plus reserve), make retriesForFiles return 0, pass --retry 0 explicitly, and drop --retry 1 from the package.json paid scripts. Add the eval:pass-rates alias. Tests that pinned the old retry allowance are updated as a policy change; review-finalization-budget now proves late-result recording under the production zero-retry arguments. * test(llm-judge): sample every judge as a pre-registered 3-sample panel Each of the 24 skill-llm-eval judges now draws EVAL_POLICY.judge.samples independent samples of the same prompt concurrently inside the unchanged JUDGE_MS budget. Numeric dimensions gate on the per-dimension panel mean against the unchanged threshold; booleans (would_browse, consistent) on a strict majority. An erroring sample fails the whole panel and is never resampled; a refusal is an unscored panel only when every sample refused. callJudge's 429 backoff stays: it is transport before any model output. The workflow-judge cache stores and validates only complete panels, and its identity now records the panel and zero file retries. Harness tests that pinned one provider call per case now pin the panel size. * test(evals): classify every live case and re-select a case when its kind changes E2E_KINDS: rule by default (191 E2E ids), 22 behavior cases whose verdict is a live model choice with an acceptable sub-100% per-trial rate, each with a BEHAVIOR_WHY tolerance, and 25 judge entries (the 24 workflow judges plus the fixed-fixture llm-judge-recommendation rubric check). Contract-shaped cases (ask-before-decide, plan-mode no-writes, mandated steps, secrets, the batching floor) stay rule. Behavior requires a known literal registration and an exact Bun test name so the case runs as its own trial shard. Map-diff selection now diffs E2E_KINDS and BEHAVIOR_WHY per key, and a base revision without them selects every key, so a kind flip runs the panel it introduces. test/eval-kinds.test.ts enforces coverage, tolerances, isolatability and the reviewed counts, printing the literal to add. * feat(evals): per-case pass rates with Wilson intervals, identity series and quarantine policy scripts/eval-flake-rank.ts becomes eval:pass-rates (eval:flake-rank stays an alias, and the legacy aggregate stays exported). It reads eval-store's trial-outcomes JSONL from the last N completed evals-periodic runs on this branch and main (gh, downloading only the trial-outcomes artifact, cached and size-capped, parsed as data), plus local eval dirs, and prints per-case per-trial pass rates with 95% Wilson intervals. A series is a case's own touchfiles minus GLOBAL_TOUCHFILES (caseSeriesIdentities, for the report job to stamp), per model, CLI version and policy version. Labels: INCONCLUSIVE, BROKEN, FLAKY, FAILING, PASSING. --backfill imports legacy slice artifacts as pre-policy trials (first attempt only, attributed by registry id, never guessed) for display only. --gate fails with ACTION REQUIRED on post-policy evidence only: drift below the quarantine entry rule, a rule case behaving like behavior, a one-sided Fisher drop against the previous identity (Holm-controlled), and quarantine entries that met their exit rule, expired after 8 weekly runs, broke the 10% tier cap or are invalid. CASE_QUARANTINE entries now carry a failureClass (detector, harness or model-latency); a product defect has no class and is never quarantined. The policy test pins EVAL_POLICY's approved constants. * feat(eval-pass-rates): attribute legacy records by the exact slug of their display name * ci(image): pin Claude Code 2.1.284 so the eval model is recognized 2.1.251 logs [claude-code:unrecognized_model] for claude-fable-5-1, the eval capture/judge default. 2.1.284 does not. The gate PTY smoke subset (plan-ceo/plan-devex plan-mode, plan-mode-no-op) parses on the new TUI; plan-design-review-plan-mode passed at 293 s on 2.1.284 and timed out at 300 s on 2.1.251 on the same tree. * test(eng-batching): grade the floor once the review report is complete A completed GSTACK REVIEW REPORT ends the review, so the review-question count is final there. Run 36606688266 wrote its report at 1,248 s and closed the session at 1,318 s; the case now stops collection and applies the unchanged floor at the report instead of waiting out the session. No budget changes. * test(eng-batching): bind unsourced native briefs through the report's target Run 36606688266 asked ten separate native review questions (D1-D9 bound to ledger records R1-R9) and failed reviewCount=0 < FLOOR=3: its briefs named the plan by title instead of citing PLAN.md, its report declared 'Review target (fixed): PLAN.md' under '# Engineering review: <plan>', and it kept an unfenced copy of the plan's own H1. The named-source route now accepts those spellings and non-inline ledger briefs. The same replay rejects a foreign, mixed, duplicate or missing target, another plan's title or copied H1, a brief naming another plan or file, a mismatched saved brief, and re-asks. The run-36597762183 capture still counts 3. * fix(plan-design-review): treat a designer with no API key as unavailable Both proof runs (36597762183, 36606688266) printed DESIGN_READY, hit 'No OpenAI API key found' on the first $D variants call, then hand-built HTML/CSS wireframes, screenshots and a comparison board for ~195-245 s before the first review question; the second run timed out at 600 s. A failed first generation now takes the existing text-only path, and the skill forbids substituting hand-built mockups. * fix(deslop-shared-libs): read related sources together within the turn limit Run 36606688266's opportunity audit read sixteen sources one per turn and stopped at error_max_turns; the passing run 36597762183 read the same files in three batched commands. The skill now says turns are bounded and asks for parallel reads or one read-only command per step. * test(ceo-mode-routing): submit a mode review that scrolled past the viewport Run 36606688266 bundled routing, learnings and the mode choice into one native call. Its review panel was taller than the terminal, so the tab bar scrolled off, ceoModeSubmissionInput returned null for 240 s and HOLD SCOPE was never submitted ('no posture match'). With no bar on screen the viewport must still end at the focused Submit prompt, and the accumulated screen text supplies the one complete panel, authenticated exactly as before. Replay controls reject another mode, an unoffered answer, an altered question, a quoted panel, trailing output, a moved cursor and an answered or changed call. * docs(evals): document the pre-registered verdict policy, quarantine, pass-rate history and arithmetic AGENTS.md replaces the retry rule with the approved policy text (no retries; kind fixes trials; no added trials, samples or dispatches after a result; quarantine by CASE_QUARANTINE only; one INFRA/INCOMPLETE re-dispatch) and notes that a pre-registered fixed panel is not rejudging. CONTRIBUTING gains the kind rules, the judge panel, eval:pass-rates and an 'Add a paid eval' checklist. TESTING_INTERNALS describes verdicts, quarantine, history and the arithmetic, including the rule term: 1 trial vs 2-of-3 red rates at p = 0.99/0.95/0.90/0.70/0.30 and lane all-green probabilities for the current 191 rule / 22 behavior / 25 judge registry. * feat(evals): trial planner, slice exit split and panel-verdict report Planner: behavior and quarantined cases become panels of isolated trial shards (<file>#<id>~t<N>) bound by EVALS_SELECTION_JSON=[id] and the exact test name; the file shard excludes them by name. Trials of one case never share a slice, result slugs are unique, panels are validated whole, unknown registrations throw, and the planner prints a capacity preflight. Executor: each trial shard gets its TRIAL_ENV identity and a trial record (outcome, failure class, cause, cost); every shard writes a JUnit report. The slice exit now means execution completeness: a failed rule shard or a trial without a record reds the runner, a failed trial does not. Report: panelVerdict() decides every panel of the first run attempt (later attempts are reported, never replacing it); rule shards keep the unchanged fail-closed checks; collector records all count (no last-attempt wins); census runs enforce the quarantine cap and expiry. It writes collector-outcomes v2, trial-outcomes.jsonl (trials plus JUnit rule/judge cases), report-summary.md, and one headline + failure block with rerun commands, and flags INFRA/INCOMPLETE-only reds for the one re-dispatch. The fail-open suite gains the panel cases: behavior 1/3 red, 2/3 green with its failed trial shown, missing trial INCOMPLETE, contract at 2/3 red, quarantined 1/3 green, 0/3 and contract red, missing slice red, and a later attempt never replacing the first. * chore(evals): refresh paid duration seeds from proof runs 36597762183 and 36606688266 Both tiers, merged in run order (the later run wins). Notable: split-overflow 1332s -> 504s, section-loading 604s -> 342s, mode-routing 575s -> 444s; multi-finding-batching 734s -> 1318s (its red path in run 36606688266). * feat(evals): stamp trial series identities and fit panels to the live registry - scripts/eval-trial-series.ts stamps series_identity (eval-flake-rank's caseSeriesIdentities) on a report's trial-outcomes JSONL as its own step, keeping the history tool out of the paid runner's closure; TrialOutcomeRecord gains the optional series_identity field. - Slice-count plans let a registered trial spill into an ordinary lane when its siblings hold every long lane, so panels never share a runner. - Re-audited test-selection.ts (Stream B added the E2E_KINDS/BEHAVIOR_WHY map-diff; no new module loading) and repinned its hash. - Detach and release floors now count trial shards (66 periodic trials in 22 panels): periodic floor 33,821s, still under eval:bg:periodic's 67,380s. - Coordination fixtures supply the executor's trial records. * ci(evals): attempt-scoped artifacts, verdict-v2 PR comment, weekly pass-rate gate and one INFRA re-dispatch - Slice, census and marathon artifacts carry -a<run_attempt>; reports download them per artifact (no merge), so records never overwrite and a re-run never replaces the first attempt's verdict. - Planners pass --max-parallel for the capacity preflight (24/16 unchanged: the refreshed periodic plan needs 24 slices, the gate census 12). - PR comment: jq-only job reads collector-outcomes v2 (headline, sanitized failure block); the group_by(.name)|last recomputation is gone. - Reports stamp series identities, upload trial-outcomes-* for history, and shard logs upload always (a failed trial no longer reds its runner). - Weekly report: headline + failure block of both lanes in the issue body, the eval:pass-rates --gate step (fails closed without history), close the issue on a green run, and UC-E1: when every red is machine-classified INFRA/INCOMPLETE, one re-dispatch as a new run in its own concurrency group (redispatch_of), both runs reported. * feat(evals): planner-side whole-panel reuse and negative receipts The planner job restores this PR's receipt store once and ships a single filtered set with the plan: a pass or panel receipt with a same-or-newer FAIL for its input identity is dropped, and a panel receipt ships only as a whole PASS panel (re-verified with panelVerdict) from one run. Executors read only that set (no per-slice cache restore or save), so every trial of a panel sees the same receipts; a trial reuses its own record from the panel receipt, keeping a split PASS's failed trial. Trial identities drop the trial index (run-scoped) and bind the panel policy. Executed shards carry their input identity; the report turns a whole fresh PASS panel into a panel receipt and a FAIL panel or failed rule shard into a negative receipt, and marks a panel that mixes reused and fresh trials INCOMPLETE. The report job merges plan, slice and report receipts (newest per file) and saves one store per run. Also fixes two TS2352 casts in browse/test/dia-macos-qualification.test.ts whose diagnostic text drifted with program order (baseline locked, fix only). * feat(evals): --case/--trials local diagnosis and panels in local sharded runs bun run scripts/test-paid-shards.ts --case <id> [--trials N] runs N independent trials of one case through the CI panel runner (trial shards, TRIAL_ENV identity, name-pattern isolation) and prints its panelVerdict(); N defaults to the case's policy panel and CI never reads it. The local sharded path (test:gate:sharded, test:periodic:sharded) now plans the same trial shards and exclusions as CI and exits on execution completeness plus panel verdicts. * test(pty): grant an owned Create pane whose title row is cropped The targeted batching rerun on Claude Code 2.1.284 left its first report Write unanswered for 1,372 s and timed out: the viewport began at the pane's relative file row and rule, with the 'Create file' title cropped above, so the preview parser rejected the file row as foreign. That row must now resolve to the owned path and is skipped before the unchanged line-by-line preview match. Replay controls reject another file, another directory and an edited preview row. * fix(evals): tsx-safe generics in eval-flake-rank, legacy artifact names, no-retry wall docs * test(evals): record the read-only and detector-row invariants as contracts shared-libs-opportunity-judgment and review-design-lite are behavior cases: their recommendation and checklist judgments may vary, but the read-only invariant (commands, provider requests, fixture bytes, hooks, state) and the deterministic fake-engine detector rows are contracts. Both now go through expectContract, so any failure vetoes the panel. * test(judges): sample the recommendation rubric as a panel; never re-ask armJudge llm-judge-recommendation is a judge case: each fixture now draws a 3-sample judgePanel, gates reason_substance on the panel mean and the present/commits/has_because checks on a 2-of-3 majority, thresholds unchanged. armJudge no longer re-asks on a malformed verdict; it is a failed sample, as the judge policy requires. * test(evals): record a pre-turn API or CLI failure as infra recordE2E sets failure_class 'infra' on a failed session whose runner reports error_api, timeout_startup, error_output_stream or a non-zero CLI exit with zero turns and no assistant event. A model refusal, a timeout after model work, max turns, or an explicit caller pass/class keeps its ordinary classification. * test: pin every-record outcome counts and the twelve doc-sync callbacks * test(eng-batching): read the report target as a field, not a spelling The next targeted rerun (Claude Code 2.1.284) again asked eleven separate native questions and again counted zero: its briefs named no plan and its report declared '- **Review target (fixed):** `/abs/PLAN.md`' under '# Eng Review — PLAN.md: <plan>'. An unsourced brief now inherits the one current target field that names a PLAN.md file, whatever its list or emphasis markup; its ledger record still supplies the cited finding and must reproduce the brief exactly. A brief that names its plan must still match the report title. Replays of all three captures count 9, 9 and 3; controls reject a foreign, duplicate or missing target and an archived title. * fix(evals): --case list mode and name precheck; case-shard qa-callers; refresh batching and design-with-ui seeds * chore(release): v1.91.9.0 * test: settle the post-response composer before seeding; give the TPA recorder adapter its infra helper submitPlanSeed accepted a stale empty composer when the transcript recorded end_turn before the CLI repainted (late-repaint-typed-current fails 5/5 on the old helper, passes 5/5 now). The TPA recording fixture extracted recordE2E without isPreTurnInfraFailure, so every failed case threw before recording. * test(autoplan-dual-voice): unwrap Claude Code 2.1.284 subagent hand-back frames; accept read-only probe diagnostics; record before asserting Census run 36626737820: the native CEO report arrived framed and indented, so its INPUT line never matched, and the model's exact probe plus two variable echoes was not canonical. A column-zero line inside a frame, command substitution, backticks, redirects, assignments, CODEX_MODE echoes and output line-count mismatches stay rejected. The failure now records before asserting. * ci(image): keep Claude Code 2.1.251; test(ceo-mode-routing): keep HOLD's own deferrals in scope before assessing its rigor decision 2.1.284 enables per-turn effort for claude-fable-5-1: in gate census 36626737820, 66 of 84 sessions ran longer than on 2.1.251 (+20% session time, +32% thinking tokens) and 11 cases timed out on unchanged budgets. HOLD SCOPE's 0G step asks its own defer/keep menu; the actor answered it Defer and the assessment then judged that scope question as the rigor decision. The actor now answers that menu Keep and assesses the next one. * test: attribute quoted prior-record field lists, state the judge reason bound in its schema, move split-overflow to marathon Census 36629958451 reds: - outside-plan-disabled-no-fallback: the model quoted the pre-existing record as a parenthesized field list with its exact timestamp; attribution now requires that exact timestamp and the record's own field values. - plan-devex-peer-comparison-classification: the judge correctly returned missing but wrote a 1069-character reason, voiding the judgment; structured outputs cannot enforce maxLength, so the bound is stated on the field. - plan-ceo-split-overflow ran 504-1188 s as one PTY flow and set the periodic lane's wall clock; it now runs weekly in the marathon lane. * test: supply holdDeferKeepIndex to the CEO routing mocks and follow split-overflow into the marathon lane The registered-callback fixtures mock ceo-mode-option and lacked the new export; the split fixtures asserted the periodic tier; the registered-budget check looked for split-overflow only in the periodic manifest. * fix(qa): checkpoint receipts print the report link for their exploration file qa-functional-webhook-report failed in two of three censuses because the report linked .qa-evidence/NNN capture folders as "checkpoints" and never linked exploration-NNN.json. The checkpoint receipt now prints link: [checkpoint NNN](exploration-NNN.json), and the functional report template says capture folders are not checkpoints. * docs: final census numbers in the v1.91.9.0 entry; file the paid-eval follow-ups * ci(evals): name the PR-comment loop's unused fields so shellcheck passes (SC2034) * fix(plan-ceo-review): tighten expansion pacing wording to fit the skeleton cap after the main merge The merged skeleton measured 80,166 bytes against its unchanged 80,150 cap. Same instructions: ask separately for each addition, in turn, with no pacing menu; lead each proposal with the felt experience, then shape, effort and impact. * fix(eval-pass-rates): match trial-outcome files by basename so Windows backslash paths are read * fix(evals): repair proof-run reds in design-consultation, document-release, design and QA fixtures - design-consultation Phase 1 asks one brief that confirms context and decides research; the confirm-only first question scored substance 2. - document-release defines ship-owned inputs, exact steps and the JSON result, and drops stale spawned-from-/ship text (judge actionability 3.67 -> 4/4/4). - plan-design-with-ui accepts the Step 0D focus menu the same way the shared picker does ("focus on specific ones?"). - plan-design-review plan-mode saves in three Edits instead of one final Write. - QA functional annotations ask for the full 40-character revision. - Outside-disabled attribution judges quoted prior-record data by its exact timestamp or a dated, pre-existing-record sentence; four captured phrasings replay clean and current claims still fail. - --case can select autoplan-dual-voice by its literal test name. * test(design): revert the three-Edit plan-mode flow A focused paid run still timed out at 300 s: the first three passes alone took 150 s of thinking. The case stays a named timeout red rather than cutting review depth. * test: accept 'review mode = X' auto-decide declarations and parenthetical scope exclusions in the shared-libs actor auto-decide-preserved: the product auto-decided HOLD SCOPE and said "Decision: review mode = HOLD SCOPE"; the grammar knew only "is" and ":". shared-libs-plan-callers: the recommended option said "(no hardening)" and the actor read "hardening" as an expansion. Both replay the captured text, keep negative controls, and passed focused paid runs. * fix(review): pass Review Army checklists by path, run research alongside dispatch, always probe the design detector; state review-log invocation and statuses in the caller fixture - review-army-perf-n-plus-one: the parent copied full checklists into agent prompts and ran web research before dispatch (290 s on a 12-line diff); 212 s now. - review-design-lite: 5 of 6 captured trials reported the detector absent without probing; the probe is mandatory and its first line is reported, and the contract credits only fake-engine rule ids the checklist never names. - review-exploratory-small-cli: the fixture never gave review-log's direct invocation or status vocabulary; the model ran it through bun and wrote status "blocked". The prompt states both and the validator rejects out-of-vocabulary review statuses. Each case passed a focused paid run after repair. * docs(changelog): proof-run product fixes * fix(ship): always run the design-lite detector probe; test(shared-libs): credit a failed first file view and deferred-reuse Skip wording - /ship design-lite: the probe is mandatory and any non-ready first line is stated, matching /review (5 of 6 captured /review trials had skipped it). - shared-libs-pr-coverage: the first PR 42 page-1 read printed only a jq error, so the one refetch is a legitimate recovery, charged to the same budget. - shared-libs-review-prior-coverage: the Skip option said a future review can "reuse it once snapshot coverage holds"; a conditional tail on the recorded decision is not product work. Captured-text regressions and negative controls. * fix(ship,qa,document-release): repair proof-run regressions and fixture gaps - ship-docsync-completion: yesterday's audit-scope result dropped the section's status, so /ship spliced one in; the section now opens with **Status:**. - ship-docsync-missing-asset: a missing section or old Ship-owned mode blocks before launch. - ship-docsync-late-result: the invocation record says prepare already saves the candidate selection (no extra Read; budget unchanged). - qa exploratory: await the method Reads before the first probe. - qa-callers fixture: quote the real review-log record template; allow the git log command plan-completion prescribes. - qa functional observer: a receipt caught mid-link(2) is checked at stop instead of failing with ENOENT (reproduced from CI). Each repaired case passed a focused paid run. * ci(image): pin Claude Code 2.1.284, the version users run Request-body capture shows both 2.1.251 and 2.1.284 send effort "high" to claude-fable-5-1; 2.1.284 adds the model's own profile. The slower 2.1.284 census was mostly API latency: its SDK-only judges were 25% slower too. Nine previously slow cases pass on 2.1.284 within unchanged budgets. * test: one owner per case id, a structural devex 0B setup rule, and correct design/gbrain actors - plan-design-review-plan-mode was registered by two files; the PTY smoke is now plan-design-review-plan-mode-smoke, and a registry test requires one owner per case in case-sharded files. - plan-devex-finding-floor: the template's 0B narrative-confirmation question is classified as setup structurally instead of timing out a Haiku assessor. - setup-gbrain-remote: the actor accepted 'skip' on the MCP-registration question the test asserts; it now accepts that question and declines others. - design-review-plugin-handoff: the fake engine cited a file absent from the fixture repo and index.html linked a missing styles.css. Captured-question regressions with negative controls; each case passed a focused paid run. * test: PTY harness handles clipped reviews and bundled setup tabs; AUQ judge uses structured output; design-consultation carve declines optional outside voices - ceo mode routing: a Submit review taller than the viewport, a setup tab bundled after the mode tab, and a clip through the mode question each hung or misread the run; the native answer is still verified after Submit. - judgeRecommendation requests a 1-5 enum schema; a malformed Haiku reply had scored substance 0 for a 4/5 brief. Judge failures now propagate. - carve section-loading for design-consultation declines the optional outside voices (a supported path) and treats DESIGN.md as the report; timeout unchanged. The Step 0E handoff defect is not fixed (0/15 samples across four wordings, none shipped) and is filed in TODOS. * test: fold the design-consultation completion replay into carve-section-sharding (test-of-test ratchet) * docs(todos): record the pre-push hook shard-order hang * test(qa-callers): disable git auto maintenance in the fixture repo (same guard as shared-libs; from #3002) * test(office-hours-attempt): the fake judge SDK response carries stop_reason like the real API (structured judge requires end_turn) * fix(qa): the caller STOP line says to await the method Reads before any probe ship-exploratory-plan-checks: the model read exploratory.md and sent a capture in the same response, before seeing the section's own await rule. * fix(qa): number the qa value-bar questions from 1 and say reproduced bugs already answer the first two * fix(qa): define evidence.json where it is built, point the preparation gate at the next section, name measured command durations in the report template Recurring qa/qa-only workflow-judge complaints in CI (clarity/actionability 3.33). * fix(plan-eng-review,review): a disallowed question tool is not headless; report kept tests only when some were skipped * fix(plan-eng-review): keep the headless-rule contract phrases adjacent * fix(evals): cut path variance at its measured sources - gstack-qa-evidence capture prints startedAt/completedAt/durationMs and, for --deadline captures, remainingMs; the functional report takes durations from them. The section clock notice asks for one clock read up front instead of one after every checkpoint (QA runs spent 7-14% of tool calls on date -u). - ship plan-completion: skip the audit dispatch when discovery already found no plan (the dispatch-vs-skip conflict produced an optional 60-100 s subagent). - materialize/checkpoint validation errors state the expected schema, so a rejected annotations file is fixable in one call instead of blocking the phase. - session-runner counts turns from the transcript when a run times out, so timeouts stop reporting 'turn 0'. * fix(evals): count timeout turns only from object transcript events * test(qa-callers): deterministic child transport, completion-time handoff reads, compact phase report The exploratory caller cases exist to prove the caller starts and bounds exploratory QA. Their native adversarial reviewer (review) and plan audit (ship plan-checks) now come from recorded child outputs instead of a live subagent, handoff freshness reads are required before completion records rather than every bookkeeping log, and the phase report is compact. Measured: 194-257 s per case against 208-284 s before, no subagent calls. * test(ship-docsync): seed fault cases at their gate instead of replaying attempt 1 The post-dispatch fault cases (missing-marker, launch-failure, timeout-unsettled, late-result, stale-before, stale-after, recovery) now start from a fixture-owned attempt 1: the real actor prepares and dispatches it, its verbatim output is saved once, and the invocation journal carries its pre-dispatch entry with the child asset hashes. The model resumes at Parent processing with a trimmed read list, inspect named as the authoritative repository observation, and recovery's intermediate checkpoint folded into the next attempt's pre-dispatch entry. Assertions count only parent-issued transport events and require a read of the saved attempt-1 output; missing-asset and the legacy failure case keep the full model-driven first attempt, and their prompts are byte-identical. * test(ship-docsync): name the seeded read list and cap journal/report length The first seeded stale-before run spent calls locating documentation.md (two ls sweeps), reading through cat and re-Reading the record before Edit, and ~40 s composing 1.5-2.2 KB entries and report. Name every seeded read path, ask for native Read, and bound entry/report length. * test(ship-docsync): trim the seeded parent's measured model time Measured on the seeded runs: one read the 78 KB ship/SKILL.md, the post-child freshness comparison spent 18-32 s of thinking over full inspect contents, and the final response restated the report (~1.1 KB). Say the phase excerpt stands in for ship/SKILL.md, compare hashes first and read content only for changed paths, and end with one status line. * feat(qa-evidence): enforce the checkpoint sequence and fill report bookkeeping in code - capture refuses to run another probe until a checkpoint anchored on the latest complete capture names this capture as its next command, and every complete capture prints that requirement. - materialize fills revision, runtime, cwd and learning (checkpoints whose next native command differs) when omitted and prints the reportLinks the report must include; the QA section shrinks accordingly. * test(qa-callers): hand the caller phase its invocation-start observations and review token; fix(next-version): fetch without auto maintenance - Every caller case receives the diff, status, log, untracked list, HEAD and an already-captured review start token, so the phase spends its budget on the contract under test instead of re-running setup reads. - gstack-next-version's fetches pass --no-auto-maintenance. On git 2.55 a completed fetch forks detached maintenance in the caller's repository; the free suite's live smoke test ran it inside the CI checkout, and every shard-12 pre-push hook hang so far followed a completed smoke fetch. * feat(deslop-shared-libs): route every Git read through bin/gstack-safe-git The skill made the model retype a long safe-Git prefix on each call and a dropped flag failed shared-libs-read-only. bin/gstack-safe-git applies the fixed env + flag prefix, adds --no-ext-diff --no-textconv to log/show/diff, allows diff only between two explicit object IDs and ls-files only in the NUL-delimited overlay form, and refuses every other shape with one line naming the allowed forms. The template now points at the installed helper (host global runtime via {{SAFE_GIT}}) and drops the prose it enforces. Fixtures resolve the helper to this checkout, the git shim records the safety environment, and isGuardedGitRequest requires the complete prefix (env included) for every repository read. * test(shared-libs): tee to a discard device is not a file write Paid shared-libs-opportunity-judgment t1 on 1213b01 failed read-only on '... | tee /dev/null | sha256sum'. The detector flagged any tee operand while the same devices are allowed for redirection. tee now fails only when an operand is a real file; tee to a file, -a file and -- -a stay violations. * fix(qa-evidence,observer): reject placeholder metadata and replay-only learning; declare the docs atomic-write target - materialize measures revision, runtime and cwd itself and rejects supplied values that differ (CI run wrote revision "HEAD" and runtime "bun"), and refuses learning checkpoints that replay the same probe, naming the fix. - The docs write observer treats Claude Code's atomic temp for the authorized doc target as transient, so a temp renamed before its per-file watch no longer marks the observation incomplete (ship-docsync-completion flake). Per-file monitoring outside declared targets stays fail-closed. * test(qa-functional): fix mode requires only the happy scenario from the model (carried byte-identical from #3002 183b01f4..3e6074b4) verifyQANativeRegression already reruns all eight webhook scenarios on the repaired source, so the model-side eight-scenario requirement in fix mode duplicated harness coverage and pushed qa-functional-webhook-fix past its budget. qa-only still requires every scenario. * fix(deslop-shared-libs): probe the audited repository with -C <repo> A CI run probed safe-git from the session directory above the target repo, so the capability probe never touched the repository and the run fell back to the API without a local attempt. The probe (and any call from elsewhere) now names the audited repository. * test(qa-deadline): never attach a reader to the full-pipe fixture's stdout The full-pipe receipt test attached a 'data' listener (flowing mode) and then paused; on CI the reader could drain the 2 MB write before the pause, so the receipt write never blocked and the helper exited 0 in ~126 ms. The stdout pipe now stays unread until the assertion, which is what the test means to model. * feat(qa): helpers answer --help, and the QA eval interfaces declare it Approved by Garry: asking gstack-qa-evidence or gstack-qa-deadline for usage is read-only, so both helpers print usage and exit 0 on --help (the evidence usage now names the annotation shape), and the functional and caller command allowlists accept exactly 'bun <path>/bin/gstack-qa-{evidence,deadline} --help'. Two CI runs failed only on that call. * fix(qa): after an input change, a probe is affected unless shown otherwise CI late-input run finished in time but revalidated only the happy probe after the locale input changed and reported the stale adverse probe green. The revalidation step now treats any probe not shown to be unaffected as affected. * test(shared-libs): seed the lifecycle replay's first Step 3 pass instead of replaying it shared-libs-review-lifecycle ran ~88% of its 300 s session budget (12-run census median 265 s, 4/24 sessions timed out). The fixture now executes pass 1's Step 3 once with the real logger and Git: a real unused REVIEW_START, then the diff, inventories, attributes/config/index flags, gstack-review-read output and every file's bytes and sha256, saved to one observation. The model resumes at Step 4 with an exact four-file first read, the observation named as the authoritative pass-1 repository read, one post-fix verification, an explicit pass-2 read list and a twelve-line summary. Pass 2 still runs its own --start, diff, reads, fingerprint and stage actor before --finish. The actor scope now states that a current settled final-pass actor result supplies the replaced QA/adversarial prerequisites and that the no-credit disclosure is a reporting label: one r1 session persisted completed:false from that ambiguity. New assertions: the final binding never uses the seeded token's start or tree, and the observation was read; free controls finish the seeded token (binding changed) and omit the observation read, and both fail. * test(shared-libs): trim the resumed review replays' setup and report Every sibling review session (revalidation, path-eligibility, index-flags, prior-coverage) loaded qa/sections/exploratory.md and often scope.md although its QA and native adversarial results are supplied synthetic inputs, then spent a second request on shared-code-reuse.md and base metadata. The resumed scope now states that the supplied results replace Step 4's QA method loading; the revalidation contract names one first response (workflow, checklist, finding, prerequisites, shared-code-reuse.md, base metadata) and caps the summary at twelve lines. Receipt order, direct source reads, the checker, the question and final persistence are unchanged. * fix(review): define what a Step 5c Skip option says Step 5c named "B) Skip" without saying what its description may claim. Two CI captures (path-eligibility on |
||
|
|
730a1017d1 |
v1.89.1.0 fix: remove continuous checkpoint commits (#2970)
* v1.89.1.0 fix: remove continuous checkpoint commits and repair validation blockers * fix: clarify shipping and engineering review recovery * fix: interpret native no-change review descriptions * test: separate descendant readiness from timeout delivery |
||
|
|
b9706f3635 |
v1.88.1.0 fix: harden credential boundaries and owned state (#2942)
* fix(settings): preserve symlinked settings targets
Resolve the selected target for locking, mutation, backup, and rollback; refuse target changes and preserve private modes. Addresses #2830.
* fix(redact): bind masking to original detected spans
Inspired by #2929's anchored-span diagnosis; independently implemented using normalization offsets. Addresses #2930 and the relocation portion of #2912 without changing detection sensitivity.
* fix(evals): exclude operator credentials from prefix admission
Adapts the credential-suffix screen proposed in #2636, with real launched-child regression coverage and deliberate provider-auth exceptions.
* fix(artifacts): retain custom allowlist rules on reinitialization
Preserve the exact user-owned suffix and publish only a successfully assembled replacement. Independently implements the repair reported in #2907.
* test(cso): verify exact masked reads and unmaskable payload refusal
* fix(cso): preserve exact filesystem identities through lease recovery
Preserve 64-bit device/inode identity and nanosecond race checks. Add native NTFS lifecycle coverage for #2927; retain ambiguous legacy-state refusal without claiming Windows PID-reuse recovery is resolved.
* fix(redact): bind pre-push scans to destination and preserve seam context
Uses #2935 (
|
||
|
|
636175d349 |
v1.87.6.0 fix: make checks reliable and everyday validation faster (#2898)
* fix: acknowledge seeded plans before invoking review skills * fix: distinguish current plan input from conversation history * fix: keep hermetic plan reviews on manual permissions * fix: distinguish tool discovery from file permission ownership * fix: preserve initial plan mode in observation tests * fix: wait for scope decisions before writing review findings * fix: carry autoplan decisions consistently into review artifacts * test: retain native failure context in periodic assertions * fix: advance active file permissions before queued questions * fix: finish red-team attempts before retry and cleanup * fix: finalize plan format captures and judges before retry * fix: cancel setup-gbrain SDK attempts before fixture cleanup * test: select periodic consumers of the bounded attempt helper * fix native Bash permission cards and queued questions * fix: preserve independent decisions and review scope Keep CEO approach, engineering scope and outside-review choices from approving independent remedies together. Carry declared contracts through DX polish and resolve new gaps before editing the plan. Regenerate every host and retain existing stop boundaries. Validation: 654 focused tests passed across nine files; all-host generation passed. Full free and periodic validation pending. Co-Authored-By: OpenAI Codex <noreply@openai.com> * fix: require approval before design plan amendments Align the Design review philosophy and rating recipe with its section protocol: resolve one proposed fix, then apply only that approved decision and retain honest scores for declined fixes. Validation: 469 focused tests passed across four files; all-host generation passed. Co-Authored-By: OpenAI Codex <noreply@openai.com> * fix: observe native question completion before transcript persistence Match owned completion hooks to submitted choices, reject conflicting or late answers, and retain bounded failure evidence. * test: recognize review posture in acknowledged native questions Require the selected mode acknowledgement, a completed follow-up question, and its current decoded display while preserving existing posture assertions. * fix: preserve settled CEO choices and isolate pending remedies Resolve established approach gates with cited authority and keep independent fixes out of unrelated option commitments and plan amendments. * fix: carry approved DX choices through later review steps Choose documentation approaches within the accepted scope and map resolved confusion points without reopening them through a bulk menu. * test: handle native settings-file edit prompts Keep one-time owned-file approvals and retain the actual sampled Autoplan permission frame with its matching barrier state. * test: accept standard CEO reply directives with tuning footers Recognize the exact trailing preference footer and letter-list directive while preserving current-display and exact acknowledgement checks. * test: scope split reviewers to their generated plan artifacts * test: observe native Bash permissions and invocation results * test: handle owned Bash prompts during mode preference checks * test: preserve synchronous subprocess rejection in Codex fixture * Fix periodic review handoff navigation Recognize review-first and explicit manual-next-step labels while preserving exact action families, manual preference, and ambiguous-menu rejection. Co-authored-by: OpenAI Codex <noreply@openai.com> * Bind pending file permissions to distinct current targets Allow one captured file request to own the complete current dialog while unrelated file work is pending. Preserve same-path ambiguity, exact input ownership, and one-time grant checks. Co-authored-by: OpenAI Codex <noreply@openai.com> * Make paired CEO verification choices genuinely unresolved Start the positive control with proposed manual checks so its unchanged oracle measures two new coverage decisions. Preserve runtime contracts, targets, count bounds, and all assertions. Co-authored-by: OpenAI Codex <noreply@openai.com> * Keep CEO review options and verification within approved scope Audit every offered option for independent add-ons and keep new verification depth pending until accepted. Preserve already requested coverage and trace plan changes to the actual decision. Co-authored-by: OpenAI Codex <noreply@openai.com> * Assemble DX review artifacts before appending the final report Keep early DX evidence above decisions, update artifact sections in place, and append the report using the actual current file suffix. Re-read after deleting an existing report before choosing the append anchor. Co-authored-by: OpenAI Codex <noreply@openai.com> * Keep outside plan reviews exclusive and invocation-owned Follow one preflight-selected backend, terminate failed Codex work before fallback, and allocate extra prompt/output files uniquely. Consume only the current invocation’s completed output. Co-authored-by: OpenAI Codex <noreply@openai.com> * Select periodic completion evaluations for report writer changes Register the shared review resolver for eight missing consumers and regress selection for all nine completion cases without changing their IDs or tiers. Co-authored-by: OpenAI Codex <noreply@openai.com> * Keep permission ambiguity fixtures on the same normalized target Use distinct raw spellings of one target in the four negative fixtures so they exercise the normalized duplicate-owner guard after exact current-file disambiguation. Preserve the existing exception, no-input, diagnostic and cleanup assertions. Co-authored-by: OpenAI Codex <noreply@openai.com> * Clarify preserved contracts in engineering review fixture Co-authored-by: OpenAI Codex <noreply@openai.com> * Recognize the offered DX follow-up handoff Co-authored-by: OpenAI Codex <noreply@openai.com> * Check independent commitments before presenting review options Co-authored-by: OpenAI Codex <noreply@openai.com> * Keep Codex review output and status in one shell invocation Co-authored-by: OpenAI Codex <noreply@openai.com> * Distinguish seeded plans from reports written by a test attempt Co-authored-by: OpenAI Codex <noreply@openai.com> * Recover clipped Autoplan file approvals with bounded viewport resizing Co-authored-by: OpenAI Codex <noreply@openai.com> * Recover clipped Bash approvals before binding the complete command Co-authored-by: OpenAI Codex <noreply@openai.com> * Isolate setup message tests from the shared checkout Run the real installer in a temporary payload with private config, require successful completion, and guard source and binary contents and mtimes. Co-authored-by: OpenAI Codex <noreply@openai.com> * Fix periodic native permission and report completion handling Match the pinned CLI's soft wraps and clipped headings without granting from incomplete frames. Retire completed file requests, retain mode annotations, and ask section captures for a short final acknowledgement after their full report is saved. Co-authored-by: OpenAI Codex <noreply@openai.com> * Preserve review approvals and validate DX comparison artifacts Keep independent remedies and approved amendments explicit. Give the synthetic DX review its existing documentation and validate peer comparison as required analysis alongside four native decisions. Add positive and negative semantic calibrations while preserving review counts, model budgets and prompt size limits. Co-authored-by: OpenAI Codex <noreply@openai.com> * Make the five-finding CEO fixture's application boundary explicit Materialize the request adapter and service composition used by the synthetic payment application. Explicitly declare the revised unregistered-event and mail-telemetry assumptions while preserving uncaught handler errors, the original invoice path and all five unresolved findings. Co-authored-by: OpenAI Codex <noreply@openai.com> * Keep CEO state-path checks scoped to directory preparation Co-authored-by: OpenAI Codex <noreply@openai.com> * Use checked ports and bounded cleanup in pair-agent tests Discover the daemon port from its owned state file, retain startup diagnostics, and await failed-start cleanup. Add occupied-port, early-exit, deadline, and foreign-state regressions while preserving the existing HTTP assertions and hook budgets. Co-authored-by: Codex <noreply@openai.com> * Preserve queued edit identity and recover clipped Bash permissions Distinguish separately queued unfinished edits from mutation of one native tool ID. Keep grants bound to an exact owned request and reject reused IDs, ambiguous inputs, and competing owners. Support the pinned renderer's literal em dash and request a repaint when only the Bash card's top rule is clipped. Grants still require the complete fresh card and an exact native acknowledgment. Validation: 413 integrated parser/event tests passed; private repaint controls and joint source review passed. Full canonical suite and native periodic rerun remain pending. Co-authored-by: Codex <noreply@openai.com> * Keep periodic reviews within their approved contracts and deliverables Carry exact approvals through engineering review, preserve declared contracts when amending CEO plans, and keep prioritization at the requested decision level. Materialize the revised synthetic SDK reference contract while retaining the five original documentation gaps. Accept the observed semicolon in the finite DX handoff menu and register the direct source dependencies used by the engineering cases. Regenerate canonical review documents without changing model budgets, retries, count bands, or native completion assertions. Validation: all-host generation and 275 review, fixture, selection and parity tests passed. Full free-suite and native periodic validation remain pending. Co-authored-by: Codex <noreply@openai.com> * Keep Eng approval cadence and independence guards explicit * Accept ordinary punctuation in manual review handoffs * Recover file permissions alongside queued Bash calls * Carry approved DX work through later review findings * Clarify the synthetic auth internal failure decision * Bound the periodic DX fixture to onboarding changes * Recognize native Design review handoff labels * Hold scope in the integration-choice review fixture * Carry approved Design decisions through review evidence * Capture listener state when feedback reload fails * Exclude workspace caches before checking deprecated flags * Verify Design UI scope against a seeded review plan * Clarify plan review decisions and outside-voice approval flow * Reject setup menus in the Design UI gate * docs: require focused repair validation before final acceptance * fix: separate review commitments within existing prompt budgets * docs: align generation and contributor validation guidance * fix: advance native review prompts and count acknowledged findings * chore: bump version and changelog (v1.87.1.0) Co-Authored-By: OpenAI Codex <noreply@openai.com> * chore: enforce cheap checks and side-effect-free validation previews * fix: handle owned Fetch permissions and oversized native cards * test: ground review fixtures in independent executable contracts * fix: preserve review decisions and verify reports before completion * test: construct the synthetic credential URL without a scanner false positive * test: materialize DX examples and verify their actual local behavior * fix: clarify CEO review decisions and execution order * fix: clarify review workflow ordering and select Design quality checks * Fix review decision gates and incomplete evaluation fixtures Persist CEO and engineering commitment ledgers before menus, preserve exact approvals, and distinguish implementation structure from feature scope. Route Autoplan through the canonical CEO Step 0 ordering. Classify DX findings before requesting approval and ground runtime claims in actual evidence. Complete neutral non-target fixture contracts and accept the captured Design handoff purpose without relaxing its ownership or acknowledgment checks. Record runtime-capability verification in AGENTS.md validation discipline. Validation: 1,335 focused tests passed across 21 files; build, all-host freshness, skill validation (647 artifacts / 107 tracked), and credential checks passed. Prior paid failures are preserved; behavioral acceptance remains pending. * Fix review decision boundaries and owned Read prompts Preserve exact approvals across review options, compare consistent DX milestones, and keep proposed implementation separate from review evidence. Bind modern Read prompts to one immutable native request and wait for its result. Retain captured regression verdicts, correct fixture error names, improve import probe diagnostics, and record focused-first validation discipline in AGENTS.md. * Clarify CEO and engineering review decisions Use explicit decision steps, one engineering ledger, and clear scope/write transitions. Preserve exact approvals and distinguish pending test requirements. Keep unrelated generated content unchanged. * Fix review decision ordering and native evaluation interactions * Clarify engineering decisions and test artifact order * Clarify pending choices and approvals in CEO reviews * Make CEO review phases sequential and clarify completion * Fix Design board submission intent matching * Seed an existing browser test baseline for Autoplan * Document decision-log payloads before state initialization * Preserve exact review scope and decide one change before drafting options * Require input identity before repeating passing model judges * Honor permitted storage throughout CEO review completion * Match complete native permission text within the pinned renderer contract * Align review approvals, independent choices, and bounded validation * fix: preserve reopened approvals and declare fixture interfaces * fix: isolate review artifacts and audit complete questions * fix: match detector artifact permissions to configured storage * fix: complete native permissions and review fixture workflows * fix: order CEO review work and separate engineering guarantees * fix: preserve native validation and separate review choices * fix: clarify review decisions and judge complete report context * fix: constrain review judgments and retain parse failures * fix: compare each affected value before review decisions * fix: make engineering review decisions and completion order explicit * fix: give the complete Autoplan evaluation a bounded chain budget * fix(cso): diagnose forbidden Docker endpoints before tool lookup * fix(reviews): reconcile workflow contracts and generated artifacts after main integration * fix(evals): migrate retained regressions to the native review harness * fix(tests): close native harness and workflow integration regressions * fix(evals): preserve complete permission context and native menu contracts * fix(tests): capture synchronous command output without pipe drain stalls * fix(reviews): clarify decision and completion ordering * fix(reviews): separate decision readiness from final completion checks * refactor(reviews): consolidate decision rules and completion branches * fix(plan-eng-review): order preparation and clarify decision routing * fix(plan-eng-review): restore size and question-format guard parity * fix(plan-eng-review): clarify scope phases and blocked completion * fix(plan-eng-review): unify review flow and report destination * fix(plan-eng-review): define bootstrap and question stage ownership * fix(plan-eng-review): clarify review structure and design lookup * fix(plan-eng-review): render report examples and show saved decisions * fix: consolidate Eng review decisions and select their evaluations * test: cover overlapping terminal attachments and clean merged runner type * fix: preserve Office Hours relationship closings during review updates * fix: retain pasted review targets across slash invocations * docs: preserve validation traces and correct release scope * test: cover pasted targets in both review skills * fix: validate report artifacts before recording success * fix: redact source roots at CSO report boundaries * fix: bind native Design questions before answering * test: select report privacy and native recovery regressions * test: bind rejection predicate in extracted observers * fix: bind complete boxed native questions * test: keep the Design UI fixture on native review * fix: preserve review decisions and evaluation completion outcomes * fix: clarify CEO approval and report completion order * fix: align native review evaluation ownership and completion * fix: bind review evaluators to native decisions and owned artifacts * fix: validate review decisions against native outcomes * fix: preserve review evidence and Autoplan phase handoffs * test: bind review evidence to owned decisions and completion * fix: retain owned native history across compaction * fix(evals): validate current review decisions and setup choices * fix: bind Autoplan reviews and phase completion to current amended input * fix: reconcile native review evidence and close Autoplan phases * test: recognize owned whole-candidate complexity decisions * test: preserve report freshness for approved investigation handoffs * fix: recognize scoped review findings and isolate dual voice fixtures * fix: make review handoffs and question dispatch self-contained * test: recognize complete CEO decisions and procedural pauses * fix: bind current CEO comparison options and risk intervals * test: bind engineering decisions and completion to owned evidence * fix: publish Autoplan phase reports before continuing tools * test: verify actual Autoplan dual-review dispatch evidence * test: select dual review when shared evidence fixtures change * fix: clarify plan review decisions and completion gates * fix: make CEO review decisions and return paths explicit * test: keep Autoplan prompt files inside attempt state * test: preserve source whitespace across permission dialog wraps * fix: publish Autoplan phase reports before continuing * test: recognize current CEO comparisons and reject inactive records * fix: reconcile engineering decision states before completion * test: recognize complete Design decisions and reports * test: verify current engineering decisions before navigation * Recognize source-owned component reduction choices * fix: recognize current CEO ledger and commitment grids * test: supply RequestPolicy context to Eng count fixture * fix: save complete engineering decisions before asking * fix: bind Autoplan publication to the complete phase readback * chore: prepare 1.87.5.0 reliability release * fix: clarify engineering review completion and preserve log failures * fix: bind CEO saved choices and current section ancestry * fix(evals): bind review execution and completion evidence * fix(plan-ceo-review): verify complete decisions before asking * fix(evals): preserve complete engineering choice records * fix(evals): preserve complete review outcomes and bounded fixtures * fix(autoplan): publish phase reports before advancing * fix(plan-ceo-review): validate option fields before asking * fix(plan-eng-review): verify current decisions after answers * fix(evals): bind review decisions and bound fixture scope * fix(plan-ceo-review): verify decision rows and edit saved checkpoints * fix(evals): bind review evidence and scope document lookup * fix(plan-eng-review): update resolution state with its answer * fix(reviews): preserve complete questions through dispatch * fix(evals): recognize completed mode declarations * fix(evals): define cache consistency at wrapper completion * fix(evals): validate owned initial scope and completed review handoffs * fix: assemble complete CEO decision fields before saving * fix: authenticate automatic mode decisions without guessing selectors * fix: bind engineering coverage to approved regression contracts * fix(evals): supply review helpers to native Eng capture * fix(plan-eng-review): preserve the full selected option scope * fix(evals): recognize owned engineering seed and regression evidence * fix(evals): bind engineering retry reports to native approvals * docs: clarify release guarantees (v1.87.5.0) Co-Authored-By: OpenAI Codex <noreply@openai.com> * fix(evals): recognize owned engineering decisions and handoffs * fix(evals): bind engineering decisions and completion evidence * fix(tests): align review contracts and selection fixtures * fix(skills): restore review prompt size limits * fix(plan-eng-review): clarify review execution and completion * fix(evals): preserve configured retries through all supervision layers * Clarify Engineering decisions and report completion * Keep native decision assertions within their source boundary * fix: recognize owned engineering decisions and completed navigation * fix: bind completed auto decisions to their current review * fix: recognize explicit CEO source attribution * fix: dispatch verified CEO decisions without recomposing fields * test: expose existing execution deadlines to review actors * fix: distinguish CEO decision records from incidental headings * test: bind split-scope choices to the registered native actor * test: connect reviewed regressions to required evaluation coverage * Clarify CEO decision routing and completion stages * test: expose existing section review deadlines to fixture actors * test: recognize complete native CEO pacing inventories * test: exclude answered history from current CEO payloads * test: detect phase entry through owned skill HOME aliases * test: validate native review completion and owned report permissions * fix: make Autoplan close packets carry the parent handoff steps * test: assess source-bound HOLD decisions within the existing deadline * fix: keep CEO native decision fields under one formatting authority * test: register integrated review and permission dependencies * test: align native review adapters and finding coverage Preserve explicit AUTO decisions, apply native single-select defaults, and bind complete cropped questions and report permissions to their owned requests. Require seeded review findings instead of crediting setup menus. Keep captured failure controls and additive selection dependencies. The integrated candidate passed 3,099 focused tests across 65 files; affected paid validation remains required before publication. * fix(autoplan): require phase reports before advancing * fix(evals): bind setup and evidence to complete attempts * fix(evals): bind native answers and pending writes to fixture scope Preserve complete option rows when native descriptions wrap, retain current owned Write arguments before journal publication, and keep engineering and DX answers within their declared fixture interfaces. Add captured free regressions without increasing model budgets or relaxing completion checks. * fix(autoplan): verify phase reports across native tool paths Guard owned methodology reads and reviewer dispatches, detect complete driver loads through Bash, and distinguish report-only edits from implementation changes. Follow authenticated native UUID ancestry when journal writes arrive out of order and verify earlier native content for cached phase reads. Keep current close acknowledgment and parent publication in order, require CEO entry before later phases, and register captured failure regressions. * fix(evals): honor native input and collection lifecycles Match complete native Edit panes and truncated question borders, reject stderr close before EOF, and stop the CEO split fixture once its acknowledged scope decisions are collected. Keep semantic validation, process failures, report requirements, and absolute deadlines authoritative. Add captured-event and real-process regressions with selection dependencies. Focused checks pass; final integrated paid and full-suite acceptance remain pending. * fix(autoplan): retain native session ownership across directory changes Recover missed native UUID ancestry through the existing strict graph while preserving ordinary event order and legacy scoping. Bind publication hooks to Claude's original project directory while retaining current cwd for requested file paths. Captured public-event regressions, existing caller checks, and a pinned native CLI loopback verify both fixes. Preserve failed attempts and require fresh paid and final full-suite acceptance. * docs: align evaluation limits and completion version * fix(autoplan): allow authenticated phase reads during journal streaming * fix(evals): bind clipped native questions and owned edit dialogs * fix: preserve overlay retries and bounded cleanup * fix: recognize owned planning preludes in native questions * docs: explain overlay scheduling and cleanup guarantees * fix: require fresh publication after Autoplan phase reruns * Release gstack 1.87.6 * fix: preserve CI paths, process identity, and test deadlines * fix: keep informational setup commands independent of install probes * fix: clarify plan review decisions and bound source audit reports * Fix remaining Windows identity and native path CI failures * Clarify CEO review decision and reviewer-result routing * test: accept no-install planner in retry supervision * fix(ceo-review): make review decisions and report completion explicit * perf(test): add fast PR gates, input-keyed judge reuse and isolated free shards * fix(test): start isolated CEO smoke from its existing project plan * fix(test): repair CI fixture races and preserve retry evidence * fix(ceo-review): clarify approvals, depth and saved completion --------- Co-authored-by: OpenAI Codex <noreply@openai.com> |
||
|
|
9f81911136 |
v1.86.0.0 feat: route outside reviews by harness (#2850)
* feat: add a restricted and supervised Claude Code runner Preserve configured authentication and models while enforcing tool access, strict completion JSON, bounded output and process cleanup. Cover argv, failure handling, session metadata and Windows process containment. * feat: route outside reviews by harness and migrate wrapper installs Use Claude Code from Codex and Codex from other supported hosts, with shared invocation rendering, positive gate validation and per-phase provenance. Rename /claude to /claude-code, repair managed shared and copied installations safely, and generate native Kiro skills. Add installed-workflow, failure-injection and live cross-harness regression coverage. * test: recognize CEO mode labels without terminal spacing The paid workflow rendered SCOPEEXPANSION at option 4, but its driver required a literal space. Match the leading mode title without cursor-spacing artifacts and ignore adjacent preview text. Preserve missing-target failures and downstream posture assertions. * test: isolate plan-count fixtures before starting review workflows Seed the complete test plan in a private git repository before launching Claude, so a bare slash command cannot review the live workspace while a delayed fixture message remains queued. Preserve count thresholds, parsers and budgets. Add initial-context and installed-discovery tests, and retain startup/terminal diagnostics on failed evaluations. * test: stabilize review fixtures and Claude eval startup Preserve source boundaries in workflow judge inputs, isolate CEO mode plans, and wait for interactive trust input readiness. Keep startup failure evidence and retain existing models, budgets, and assertions. Co-Authored-By: OpenAI Codex <noreply@openai.com> * test: classify collapsed review modes and isolate seeded findings Keep review questions out of the setup count when terminal cursor positioning removes spaces. State existing webhook safeguards so the five-finding control measures its seeded defects without accidental extra security and concurrency gaps. Preserve question bands and the paired control. Co-Authored-By: OpenAI Codex <noreply@openai.com> * test: isolate browser daemon state across free shards Co-Authored-By: OpenAI Codex <noreply@openai.com> * test: stabilize native review counting and interactive navigation Co-Authored-By: OpenAI Codex <noreply@openai.com> * chore: prepare v1.82.0.0 release Co-Authored-By: OpenAI Codex <noreply@openai.com> * fix: eliminate browser and process-cleanup test flakes Pin every CI surface to Bun 1.4.0 to avoid extra-stdio finalizers closing reused live sockets. Add an isolated GC/listener regression that fails on Bun 1.3.13, and prevent coordinated rollback to an affected CI runtime. Check renderer cleanup against the render's own staging directory so concurrent renders cannot invalidate the assertion. Make the no-pgrep process-tree walk tolerate disappearing /proc entries, and synchronize its test fixture through child readiness and pipe EOF instead of sleeps. Validation: 9,157 passed, 31 skipped, zero failures across 556 files with retries disabled. Build, all-host generation freshness, and skill checks passed. All three races have failing-before/passing-after regressions. * fix: count completed native review questions in evals * fix: drive review navigation from confirmed native choices * fix: require complete section-loading eval reports * test: isolate telemetry HTTP transport from local assertions * fix: keep review input on the active native question * test: let tunnel revocation daemon choose an available port * test: allocate available ports for pairing and watchdog fixtures * fix: stabilize planning eval navigation and phase reporting * test: isolate installed runtime paths in planning evals * test: stabilize review evidence and concurrent refresh fixtures * fix: resolve design findings before editing the plan * fix: honor and persist disabled outside plan reviews * fix: preserve planning decisions and terminal evidence Load installed host reviews at autoplan phase entry and wait for completed reviewers and saved artifacts. Reuse approved remedies while preserving individual finding decisions. Drive interactive evals from the current terminal viewport, bind native questions across scrolling, and require complete native report evidence. Cover captured stale menus, permission lifecycles, setup classification, and disabled-review tool availability with deterministic regressions. Advance release metadata and the upgrade migration to the unclaimed 1.83.0.0 slot. * fix: drive native review questions and preserve current plans Use the native single-choice keyboard protocol and current terminal viewport, with per-question navigation inside packets and completed-call coverage. Keep permissions, multi-select menus, and Submit controls distinct. Send Autoplan reviewers the amended implementation plan, keep its review record separate, and supply retained application contracts in the chain fixture. Clarify individual DevEx decisions and complete CEO fix options; use one active plan destination for the section-loading report. * fix: preserve complete plan-review decisions * fix: recognize native plan dialogs and reviewer controls * fix: preserve review decisions and phase completion * fix: recognize completed reviews without losing findings * fix: preserve review continuity and native eval completion * test: fix native review completion and eval retry isolation * test: handle native review menus and complete eval fixtures * test: fix native review setup, completion, and isolation failures * test: limit native skill discovery to runtime assets * fix: bind Autoplan reviews to full ordered phase inputs * test: fix planning eval routing, counting, and timeout handling * chore: advance queued release to v1.84.0.0 * fix: preserve complete review inputs and planning decisions * fix: reconcile review approvals and preserve phase obligations * fix: preserve review obligations and unblock eval permissions Carry recorded Autoplan requirements into blind phase inputs, require Eng review approvals before exit, and exercise combined asynchronous flows in CEO reviews. Correct native finding and handoff classification and unblock repeated report edits using scoped request identities. * fix: retain plan requirements and complete native review dialogs * fix: complete native review prompts and retain plan references * fix: preserve review inputs and classify native eval evidence * fix: check competing completion orders in CEO reviews * fix: recognize review decisions and require phase methodology Require the current phase methodology before Autoplan snapshots. Correct substantive decision, closed handoff, and cache-finding classification, and honor the recommended implementation approach in native review dialogs. Add captured-transcript regressions without changing review thresholds, provider models, retries, or deadlines. * test: bind native review decisions and close completed handoffs * fix: complete review dialogs and verify methodology delivery * fix: preserve review evidence and unblock native eval prompts * fix: handle native review question completions * fix: recognize native review narration and controls * fix: count native review decisions and isolate eval fixtures * test: verify seeded review coverage and current artifact permissions * test: isolate model and brain-aware skill renders * fix: repair native workflow evaluation and clarify review steps * fix: stabilize workflow eval evidence and review guidance * test: repair native workflow observation and fixture isolation * fix: recognize completed workflow evidence and owned skill reads * test: repair seeded workflow delivery and completion evidence * test: recognize current review evidence across native forms * test: handle native review variants and permission redraws * fix: honor review preferences and recognize native eval evidence * test: recognize completed review decisions and queued permissions * test: match current review contracts and partial-line edits * test: recognize completed workflow evidence and bounded human waits * fix: preserve review entry gates and native eval interactions * fix: recognize native workflow evidence and preserve review gates * test: recognize current review evidence and preconfigure workflow fixtures * test: recognize completed review findings and scoped artifact permissions * fix: stabilize native workflow review and permission evidence * fix: recognize current review evidence and scoped edit confirmations Clarify Design and engineering review entry instructions and Design scoring. Recognize required legacy coverage and public Autoplan completion recaps. Bind the pending Edit confirmation to its exact file, ordered digest, and one-request approval when a preceding command display remains visible. Keep reviews within their existing size limits and preserve scope gates when extracting workflow fixtures from either supported preamble header. Keep failure outcomes, review thresholds, provider choices, and eval budgets. * fix: recover review workflow progress and eval evidence * fix: recognize valid review evidence and scope selection * test: fix review evidence parsing and repeated artifact prompts * test: recognize valid review decisions and pending native cards * fix(plan-eng-review): keep final navigation consistent with approved tasks * test: recognize valid review evidence and bind legacy diff requests * fix: stabilize review eval evidence and harness repair guidance * docs: update project documentation for v1.85.0.0 Co-Authored-By: OpenAI Codex <noreply@openai.com> * test: fix Windows CI fixtures and credential scan Rebase captured JSON values and filesystem evidence using the appropriate path convention. Compile native fake CLIs on Windows and synchronize pipe holder readiness, with cleanup retained when assertions fail. Assemble synthetic credential fixtures at runtime so the added-line scan keeps enforcing the same gate without flagging its own rejection controls. Discover generated skills directly for the empty-find regression check, avoiding a recursive scan through saved evaluation artifacts and dependencies. * fix: preserve source renders on Windows Compare canonical generator paths using native separators so an output sidecar pointing at the source cannot overwrite its skill or metadata. Keep the regression fixture isolated from the real checkout and expose freshness diagnostics before asserting subprocess status. Detach Windows drain-test pipe holders from the fake provider's automatic child cleanup while preserving the enclosing runner job and its assertions. * fix: clarify outside review fallback and CEO decisions Render one applicable own-harness fallback path and retain native review, disabled policy, and missing-coverage semantics. Align report field names and mode labels, and make the existing per-cut scope approval explicit. Regenerate skill outputs and keep the workflow judge's model, thresholds, and retry policy unchanged. * chore: move release to free version slot (v1.86.0.0) PR #2852 now claims v1.85.0.0. Align the release metadata and rename migration so upgrades from that version still receive it. Co-Authored-By: OpenAI Codex <noreply@openai.com> * fix: include engineering review prerequisites and restore branch context * fix: recognize coverage diagrams and clarify design review instructions * fix: preserve file identities and join Windows test processes --------- Co-authored-by: OpenAI Codex <noreply@openai.com> |
||
|
|
0530392821 |
v1.81.0.0 feat: Aside is the browser gstack drives first; every browsing skill, the PDF/diagram renderer, and web research; the bundled browser stays the automatic fallback (#2810)
* feat(aside): browser-driver contract, cookbook, research and fallback resolvers
{{ASIDE_SETUP}} (readiness probe + ten rules for driving the user's real browser), {{ASIDE_COOKBOOK}} (script shapes verified live against Aside CLI 1.26: one flow per aside repl script, CDP console hook before navigation, evidence lines, session-directory artifact handoff, GSTACK_STEP_OK sentinel), {{ASIDE_RESEARCH}} (research through aside exec, WebSearch when Aside is absent, knowledge otherwise) and {{BROWSE_FALLBACK}} (the fifteen-row Aside-step to $B-command table plus the rules that differ, so every browsing skill keeps working on gstack's own headless browser). test/aside-driver.test.ts pins the sentences and asserts every browsing skill carries the Aside block followed by the fallback; test/helpers/aside-available.ts is the shared live-Aside probe.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* feat(render): Aside-first local-HTML renderer with the bundled browser as fallback
lib/aside-render.ts serves the HTML's directory on loopback (Aside refuses file:// URLs), opens it with waitUntil load, prints through CDP Page.printToPDF so tagged output, outlines, header/footer templates and page numbers survive, emulates device metrics for sized screenshots, and writes in-page evaluations to files; when Aside is absent it runs the same spec through the browse daemon (newtab, load, js, pdf, screenshot, closetab) and reports ENGINE=aside|browse. bin/gstack-render.ts is the CLI skill templates call. lib/claude-bin.ts and lib/error-handling.ts become the canonical copies (browse/src re-exports them).
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* refactor(browse): /browse drives Aside first, with the $B reference behind the fallback
Contract, cookbook, mode choice (aside repl by default, aside exec for reading), report format, the fallback section, and the full command reference carved on demand.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* refactor(qa): /qa and /qa-only drive Aside, fall back to $B
QA_METHODOLOGY runs every phase as Aside scripts (orient, explore, document, re-test, mobile viewport via CDP emulation, links via HEAD fetch); the authenticate phase is 'you are already signed in'; a 13th rule requires consent before mutating actions on non-local targets; the fallback section translates each step onto $B. The qa E2E tests run on whichever engine is present.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* refactor(design): design-review, design-consultation, design-shotgun, plan-design-review, design-html drive Aside
Design-system extraction is one script printing FONTS/COLORS/HEADINGS/TOUCH_TARGETS/NAV; competitor research confirms the exact URLs before opening them in the real browser and runs on the bundled browser when Aside is absent; design-html's viewport screenshots, sketches and comparison boards render through gstack-render.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* refactor(deploy): benchmark, canary, land-and-deploy Step 7, devex-review drive Aside
One aside repl script per page prints NAV/PAINT/LCP/RESOURCES/SCRIPTS/CSS/SUMMARY (benchmark), CONSOLE_ERRORS/NAV/TEXT + screenshot (canary, re-run every 60s), and the post-deploy check reads responseStatus from the navigation entry; each carries the $B fallback.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* feat(third-party-actions): Aside is the recommended driver; gstack's visible browser stays the fallback
The readiness probe is lifted from {{ASIDE_SETUP}} at gen time (byte-identity pinned) and rule 3 points at browse/SKILL.md for how to drive; the consent question offers Aside first and gstack's own visible browser (handoff/resume for sign-in) as the fallback, as v1.72 framed it.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* refactor(scrape): /scrape reads pages through Aside; the browser-skills runtime rides the fallback
Look-then-extract scripts build the JSON inside the page and print it between JSON_START/JSON_END; aside exec for fuzzy intents; on the $B fallback the browser-skills match/prototype flow and /skillify apply as before.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* refactor(make-pdf): print through Aside first, the bundled browser otherwise
asideClient.ts replaces the direct $B client with one render() call per PDF (the exact option mapping the browse pdf command had: paper, margins, header/footer/page numbers, tagged, outline, printBackground, preferCSSPageSize, Paged.js wait); the diagram pre-pass, oversized-image downscale and DOCX rasters each run as one render script with per-fence try/catch; exit 4 now means no browser is available and names both remedies; $P setup reports which engine it found. The e2e gates run on whichever engine is present, so the Linux lane exercises the fallback.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* refactor(diagram): the triplet is one gstack-render call
SVG, PNG and excalidraw from one invocation over the content-addressed bundle staged under /tmp/gstack-render; every diagram type gets an excalidraw export; gstack-render picks the engine and prints ENGINE=; the diagram E2E gates on either engine.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* feat(research): web research runs in Aside first, WebSearch second
The planning, review, design, security and investigate skills research through {{ASIDE_RESEARCH}}; WebSearch stays in allowed-tools as the fallback; testing.ts's bootstrap step follows; skeleton ceilings ratcheted for the research block.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* feat(setup,gen-skill-docs): prune renders of skills that no longer exist
setup gains _prune_stale_generated for every host tree and the doc generator removes gstack-* output dirs it did not write, so a skill removed from the source tree can never linger in an install.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* test: registries, budgets and suite reconciled for Aside-first with the $B fallback
Touchfiles + E2E tiers gain the Aside keys, coverage matrix and eval baselines updated, size budget re-baselined to parity-baseline-v1.80.0.0.json (the contract plus fallback ride in every browsing skill), parity ceilings ratcheted with measured values, LLM-judge prompts and the E2E fixtures speak Aside-first, browse-fallback.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* docs: Aside first, gstack browser fallback
README, BROWSER.md, docs/, CONTRIBUTING, CLAUDE.md, ARCHITECTURE, AGENTS.md, TODOS and the root router describe the one product story: Aside is the browser gstack drives first; the bundled headless browser is the automatic fallback (Linux, Windows, app closed) where cookie import, GStack Browser, pair-agent and browser-skills still apply.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* chore: regenerate SKILL.md docs, llms.txt, agents digest, ship goldens, context-budget fixture
bun run gen:skill-docs over the templates; goldens re-rendered; context-budget ceilings recaptured.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* v1.80.0.0: Aside is the browser gstack drives first; the bundled browser is the fallback
MINOR: new capability across ten skills, the renderer and research; nothing removed. CHANGELOG release summary + itemized changes; VERSION 1.80.0.0; package.json 1.80.0.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* docs(todos): file non-Claude host ownership-gate and version-heading pin follow-ups
Two follow-ups from the /plan-ceo-review + /plan-eng-review pass on merging
PR #2804 with main's v1.80.0.0 ownership gate: bring the Codex/Factory/
OpenCode/Cursor/Kiro copy loops and the stale-render prune under the
.gstack-owned marker rule, and a free test pinning that the CHANGELOG top
heading equals VERSION (the collision that git cannot see).
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* fix: pre-landing review fixes for the Aside-first branch
Review army + adversarial passes (Claude and Codex) on the merged branch:
setup
- _prune_stale_generated scans the host dirs too (the generator already
removed the render before setup ran, so the host branch was dead), skips
symlinks in the render tree (rm -rf on a slash-terminated link empties its
target), removes a host symlink only when it resolves into gstack, cleans a
bannered real dir through _cleanup_weak_dir, recognizes frontmatter-renamed
skills, and logs through log. The always-run codex render passes every host
dir that may link to it.
- NEEDS_BUILD checks all three binaries (with $_EXE) and lib/ sources; the
browser hint and the bootstrap summary honor GSTACK_SKIP_ASIDE, treat a
requested skip as a request, and derive one skill list.
lib/aside-render.ts + bin/gstack-render.ts
- The loopback server carries a per-render secret path, checks containment on
the real path (symlink escapes are 403), and rejects malformed encoding.
- Inline eval results are one base64 line, so page text cannot forge
ASIDE_DIR= or the sentinel; the last ASIDE_DIR wins.
- runProc escalates SIGTERM to SIGKILL, bounds every wait, and clears every
timer (an uncleared one kept gstack-render alive after printing OK).
- renderTmpDir refuses a shared /tmp name owned by someone else; the work dir
and server are created inside try; goto's budget follows the render budget.
- probeAside classifies a present-but-failing CLI as ASIDE_NOT_RUNNING like
the skills' bash probe; render() retries on gstack's own browser when Aside
could not start or its private CDP bridge is gone (never on a page error
or a timeout of a running script); the CLI reports the engine that actually
rendered, exits 0 on --help, rejects non-numeric flags, documents
--wait-timeout, fences EVAL/PAGE_ERRORS as untrusted content, and names the
daemon's cookie-import JS lock remedy.
- The browse path passes --scale only when asked (a scale change rebuilds
the daemon context) and restores the viewport after a sized screenshot.
resolvers / templates
- The bash probe honors GSTACK_SKIP_ASIDE and has a perl deadline on stock
macOS; .local is no longer LOCAL (mDNS); same-origin filters compare parsed
origins; link status is HEAD-checked only on LOCAL targets; every
aside exec goes through the receipted _aside_exec prelude
({{ASIDE_EXEC_PRELUDE}}), including nine template blocks that called it
bare; the design sketch and diagram staging use private directories.
- The generator prunes only bannered renders and never a host whose
generation failed.
Docs, stale comments and dead code cleaned; goldens re-rendered; tests
updated and added for every behavior above.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* test: coverage for the render CLI, setup rebuild check, make-pdf exit codes, and prose $B spans
New free tests from the ship coverage audit: test/gstack-render-cli.test.ts
(argv guards, --help, output contract with a fake daemon, failure and
serve-root paths, no-browser case, prompt exit), test/setup-needs-build.test.ts
(every binary and source set flips NEEDS_BUILD, Windows suffixes),
make-pdf/test/cli-exit-codes.test.ts and setup-smoke.test.ts (error to exit
code mapping, runSetup stages, renderPdf's engine), and prose-span cases for
extractBrowseCommands in test/skill-parser.test.ts.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* docs: CHANGELOG and TODOS cover the review fixes (v1.81.0.0)
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* docs: sync project docs with the v1.81.0.0 review fixes
BROWSER.md, ARCHITECTURE.md, CONTRIBUTING.md, README.md, CLAUDE.md,
docs/TESTING_INTERNALS.md and docs/PROJECT_STRUCTURE.md now describe the
shipped renderer and setup: the loopback render server's per-render secret
path and real-path containment, ENGINE= naming the engine that actually
rendered (mid-run retry on gstack's own browser), EVAL/PAGE_ERRORS fenced as
untrusted content, --wait-timeout and the CLI's argv guards, the receipted
_aside_exec prelude ({{ASIDE_EXEC_PRELUDE}} in the placeholder table), the
LOCAL host rule without .local, LOCAL-only HEAD checks in the links script,
GSTACK_SKIP_ASIDE across probe/renderer/setup, the ownership-gated
retired-skill prune, the widened NEEDS_BUILD check, and the new free tests
(gstack-render-cli, setup-prune-stale-generated, setup-browser-hint,
setup-needs-build, make-pdf cli-exit-codes and setup-smoke).
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* docs: CHANGELOG states the precise mid-run retry rule
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* fix(test): skill-e2e-bws slices the $B setup block from the Browser fallback section
browse/SKILL.md no longer has '## SETUP' / '## Core QA Patterns' (Aside is the
primary driver; the $B block moved under 'Browser fallback'), so the gate test
sliced an empty block and handed the agent nothing to run. Anchor on
'### Find the `$B` binary' up to the next heading. 7/7 pass.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* fix(test): gate POSIX-only fixtures off Windows
windows-free-tests: the gstack-render CLI tests drive a shebang fake browse
that CreateProcess cannot exec, and two NEEDS_BUILD cases assert an execute
bit and a bare-name miss that MSYS bash does not have (test -x ignores mode
bits and resolves design -> design.exe). Those describes and cases now
self-skip on win32; argument guards, --help, the no-browser case, and every
other rebuild-check case still run there.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* fix(render): runProc waits for the exit code until the kill deadline; newtab retries once on a cold daemon
A process whose pipes have reached EOF is exiting, but runProc gave the exit
code only five seconds to arrive and then returned null, which run() reports
as a failed command. Under CI's six-shard load one such render failed with the
artifact already written. The SIGTERM/SIGKILL timers already bound the wait,
so the exit race now runs to the kill deadline.
The first CLI call auto-starts the browse daemon; on a cold start it can
answer 'Unable to connect' once while the server is still coming up. That
single case is retried after 1.5s; every other newtab failure is not.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* test(aside-render): warm the daemon before live fallback cases; failures name the render error
- Live fallback cases run 'goto about:blank' up to twice before asserting and
skip (never fail) when the daemon cannot come up.
- expectOk() puts r.error and the browse transcript into the assertion so a
failed render is diagnosable from the CI log.
- The argv-contract cases dump the fake's log on a miss.
- File default timeout is 30s: the subject is the CLI contract, not latency.
- Two cases pin the cold-daemon newtab retry and that other errors are not
retried.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* docs: CHANGELOG notes the cold-start tolerance of the bundled-browser renderer
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
---------
Co-authored-by: Sina <sdroid674+github@gmail.com>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
|
||
|
+8 |
702a1a9b69 |
v1.78.0.0 fix: the two-red-lanes wave — AUQ collapse rooted, OSV green from 105, 18 community PRs absorbed, upgrade path can't eat installs (#2752)
* fix(auq): spawned trigger is objective — explicit declaration or STATUS echo, never inference (periodic-lane AUQ collapse)
The v1.76 spawned rule's parenthetical '(or your dispatch prompt marks this
session as spawned)' let the model INFER spawned status from a scripted-looking
prompt in a CI-looking session and silently auto-choose every review-phase
question: reviewCount=0 across the plan-review periodic E2Es (weekly run
33363624506, 9 of 14 failed shards; reproduced locally, zero AUQ fingerprints).
Env and hook paths were excluded by inspection: hermetic children echo
SESSION_KIND: interactive (CLAUDE_CODE_ENTRYPOINT=cli beats CI markers) and the
question-preference hook isn't installed there.
The trigger is now objective: the echoed SESSION_KIND: spawned STATUS line, or
an EXPLICIT dispatch-prompt declaration ("you are a SPAWNED subagent") —
declared, never inferred — with an absence-safe interactive fence: CI env vars,
scripted-looking or pasted prompts, and write-to-this-exact-file instructions
are NOT spawned markers. The prose channel stays because Task-tool subagents
inherit the parent env (no spawned prefix) — their dispatch prompt is the only
signal; #2733's env-prefix channel is untouched.
19 carve skeleton ceilings re-pinned with measured values (+~440 bytes/skill);
ship goldens refreshed for all three hosts; resolver pins extended with the
no-inference regression tests.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix: mktemp failure aborts loudly at all three skill-content sites; failed upgrade swap restores the backup (#2679)
An empty $(mktemp) result silently disabled the redaction pass (redact-doc
resolver, ship pr-body) and made /gstack-upgrade's vendored path destructive:
clone lands at "/gstack", the swap mv fails, and rm -rf then deletes BOTH the
live install's backup and "". All three sites now guard the assignment with a
loud exit; the vendored block additionally restores the backup when the swap
fails (same failure class — backup deletion after a failed mv) and the GitLab
MR path sends the SCANNED file's bytes instead of re-rendering an unscanned
heredoc. bin/gstack-redact rejects an explicit empty --from-file path instead
of silently falling through to stdin.
Receipts: 6 of 8 new regression checks fail on a v1.77.0.0 scratch worktree.
Fixes #2679
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(auq): the interactive fence classifies the session — it never nudges ask-count
Burn-in run 1 of the periodic repro overshot the review band (reviewCount=8 >
CEILING=7) with the fence's 'when unsure, ask' tail: that phrasing is a quota
nudge, not a classification default. The fence now states it only classifies
the session and never changes how many questions the skill asks. Pin added.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(ci): OSV suppression config actually loads — explicit global --config + expiring, reasoned ignores
The ignore file was inert from v1.65.0.0: OSV-Scanner only auto-discovers
configs named osv-scanner.toml (no leading dot) and applies them
per-directory, so the root config never covered lib/diagram-render/bun.lock
either way. The workflow now passes --config=.osv-scanner.toml globally.
Every IgnoredVulns entry carries a reason with an upgrade trigger and an
ignoreUntil expiry (~90 days) so suppressions must be re-justified. A wiring
test pins flag ↔ filename ↔ entry hygiene so the file can never silently go
inert again.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(deps): dependency wave — 105 OSV advisories → 3 reasoned suppressions, all lanes verified on the pinned scanner
Root: overrides pin ip-address 10.3.1 (defeats BOTH nested nodes — socks'
range pull and express-rate-limit's exact 10.1.0 pin, which a top-level bump
provably cannot reach) and sharp 0.35.0 (GHSA-f88m, HIGH; transformers still
pins ^0.34 upstream — smoke-tested round-trip); marked ^18.0.11; full in-range
lockfile refresh clears hono, fast-uri, protobufjs, qs, body-parser, nanoid,
uuid, immutable and friends.
lib/diagram-render (via its own build-script contract: exact pins edited,
fresh lock, dist rebuilt): mermaid 11.16.1, @excalidraw/excalidraw 0.18.1,
@excalidraw/mermaid-to-excalidraw 1.1.2 → 2.2.2 — the 1.x line exact-pinned
mermaid 10.9.x and dragged the entire duplicate mermaid-10 advisory chain
(dompurify 3.1.6, nanoid 3.3.3, lodash-es); the bundle shrinks 9.96 → 7.59 MB
with the duplicate mermaid gone. Nested exact pins that survived get scoped
overrides (nanoid 5.1.16, lodash-es 4.18.1).
Verification: clean-worktree frozen-lockfile installs (root + nested) + the
SAME osv-scanner release the action pins (v2.3.8) with the workflow's exact
scan-args → exit 0, 'No issues found'. Smoke tests cover the override
surfaces (sharp round-trip, ip-address lockfile assertion, marked parse);
socks + diagram-drift suites already pin the rest.
Supersedes #2695 (its own lockfile kept socks/ip-address@10.2.0; @anupamme's
report credited for the parallel diagnosis).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* test(gbrain-sync): stub pgrep so the pin case is hermetic
The only non-dry-run --code-only child hits #1734's PATH-resolved
autopilot probe. A live host daemon is a correct refuse; the test
cannot inject processRunning. Neutralize pgrep in the fixture bindir
instead of adding a production env hatch.
Co-authored-by: Cursor <cursoragent@cursor.com>
* test(gbrain-sync): blank inherited GBRAIN_HOME in the pin child
Lock paths are checked before pgrep. Spreading process.env let a runner
GBRAIN_HOME with a live lock refuse the case before the stub ran.
Co-authored-by: Cursor <cursoragent@cursor.com>
* fix: point ship design-checklist at installed gstack/review path
The /ship Design Review step skipped the checklist because the generated path omitted the gstack/ install segment. Sync the generated skill doc and pin a regression assertion.
Co-authored-by: Cursor <cursoragent@cursor.com>
Wave-amended: goldens regenerated against the wave tree (author's golden commit
|
||
|
|
394db326f2 |
v1.71.0.0 feat: token-load reduction — preamble runtime scripts, gated onboarding, 20 skill carves, CLAUDE.md trim (#2691)
* feat(gen): strip gen-time-only frontmatter keys from Claude renders
interactive + benefits-from are read from the .tmpl by buildContext at
generation time; no runtime, host, or test reader consumes them from the
generated SKILL.md (e2e-harness-audit reads .tmpl; benefits-from tests
assert rendered prose). gbrain: stays (bin/gstack-brain-context-load reads
it from the installed render); hooks: stays (Claude Code host wires
PreToolUse from it).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* chore(gen): regenerate SKILL.md — dead frontmatter keys removed
Mechanical regen after hosts/claude.ts stripFields change.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* feat(test): context-budget ratchet — CI ceilings on always-on + eager token ledgers
New free test grades the two ledgers nothing else guards: the full-frontmatter
always-on catalog (aggregate) and per-skill eager tokens (SKILL.md +
forced-read refs), via checkBudget from lib/context-bill.ts. Ceilings live in
test/fixtures/context-budget.json with x1.05/x1.10 headroom; regenerate with
bun test/helpers/capture-context-budget.ts. New skills fail until consciously
budgeted; removed skills fail until the fixture is refreshed; reductions
ratchet the ceilings down so wins lock in.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* docs(todos): file output-template carve wave + plan-ceo doctrine revisit; mark preamble-carve P3 in flight
Two follow-ups deferred from the approved token-reduction program (CEO review
'NOT in scope' list), filed with full context per TODOS format. The existing
P3 preamble-carve entry gets a status update pointing at the program that
supersedes it.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(test): review findings — Windows path normalization, full totals rebuild, ratchet coverage
Pre-landing review (5 specialists) found one critical: the ratchet test runs
in the curated Windows lane, where path.relative yields backslash skill names
that miss the test/ filter and mismatch every POSIX fixture key. Names are now
normalized once in buildRatchetBill (toPosixName) and the fixture filter is
tightened to test/fixtures/. All eight Bill.totals fields are rebuilt from the
filtered list (no fixture-polluted perInvocation/totalMd numbers for future
consumers). New coverage: Windows-separator normalization pins, a
captureContextBudget round-trip against tree-a (headroom math exact), a
stripFields regression pin (interactive/benefits-from absent from renders,
hooks/gbrain preserved), and the ceilings test no longer double-reports
stale-fixture entries.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(test): adversarial findings — stable root key, symlink-alias dedupe, fixture-shape guard
Adversarial review (Claude subagent) verified the fixture's root-skill key was
the capture machine's checkout dirname: any non-gstack-named clone (every
Conductor worktree) failed the free suite, and the documented re-run-the-capture
recovery baked the local dirname into the committed fixture — silent corruption
through the tool's own protocol. The root skill is now pinned to ROOT_SKILL_KEY
('gstack', its frontmatter name). Symlink aliases are realpath-deduped (census
precedent): connect-chrome no longer gets its own ceiling, so Windows checkouts
that materialize the symlink as a plain file can't fail the stale-ceiling
set-equality test. New guards: fixture-shape validation (a string alwaysOnTotal
can no longer silently disable the ceiling), a mutation pin that the filter
shrinks the always-on ledger vs the raw bill, an alwaysOnTotal violation test
(the branch was load-bearing with only under-budget coverage), and an atomic
temp+rename fixture write. Fixture regenerated: 59 ceilings, alwaysOnTotal 6344.
Deferred with a TODO: anchoring transformFrontmatter's denylist strip to the
frontmatter block (latent, zero live collisions, pre-existing path).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* chore: bump version and changelog (v1.69.1.0)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* docs: update project documentation for v1.69.1.0
CLAUDE.md: Token ceiling section documents the context-budget ratchet as
the third guard (test file, fixture, new-skill budgeting, capture command).
CONTRIBUTING.md: Tier 1 guard list gains a Context-budget ratchet bullet;
the Adding-a-new-skill checklist gains the budget-capture step.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* docs: pin exact guard semantics for the context-budget ratchet in CLAUDE.md
Doc-review finding: "a third enforced ceiling" undercounted the guard
family (skill-size-budget floors and parity ratios also watch these
ledgers, relatively). Rephrased to match the ratchet test's own header:
absolute ceilings vs relative floors/ratios.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* docs(changelog): heaviest-skill claim matches the fixture (land-and-deploy edges review by 0.2%)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* feat(bin): gstack-skill-start + gstack-skill-end — the preamble runtime, consolidated
Absorbs the ~13KB of bash every tier-2+ SKILL.md inlined twice over (bootstrap
fence + artifacts-sync fence) and the skill-end telemetry/sync fences. Same
KEY: value STATUS-line contract the prose interprets, plus SKILL_START_PROTO
handshake (OV5), SESSION_ID/TEL_START echoes, GSTACK_HOME-normalized state
paths (EOV7), --parent-pid session identity (EOV5: $PPID inside the script is
the ephemeral tool-call shell), OV4 sanitization of passthrough output, and a
receipted daily artifacts pull (_receipted_git, brain-sync class, fail-closed).
Per-line || true error style throughout (F3) — a mid-script failure never drops
later STATUS lines.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* feat(gen): preamble resolvers emit a script invocation fence instead of inline bash
generate-preamble-bash: ~6.3KB fence -> 4-line gstack-skill-start invocation
(quoted-tilde pitfall handled: leading ~ interpolates through $HOME; env-var
hosts keep $GSTACK_BIN) + degraded-mode prose (F1/EOV8: safe defaults, consent
gates deferred-never-lost; OV5: proto rule). generate-brain-sync-block: ~6.8KB
bash -> interpretation prose + the privacy stop-gate (stays inline until
Phase 2's gated emission). generate-completion-status: telemetry fence -> one
gstack-skill-end call with SESSION_ID/TEL_START handoff.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* chore(gen): regenerate all skills + golden fixtures — inline preamble bash removed
Mechanical regen after the resolver change: −12,628 lines across 52 renders
(corpus 952K -> 806K render tokens; tier-2 skills −11-13KB each). Golden
per-host ship fixtures refreshed from the fresh claude/codex/factory renders.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* test: skill-start contract suite + preamble A/B eval + touchfiles registration
test/gstack-skill-start.test.ts (11 free tests): STATUS-key contract vs the
prose (F2), per-host fence resolution shapes (E1), proto-first, OV4 marker
sanitization, --parent-pid identity, headless suppression, skill-end duration
math + pending cleanup. test/skill-e2e-preamble-script-ab.test.ts (gate tier,
OV7): inline-bash render (pinned from
|
||
|
|
1d41ee3ab3 |
v1.63.0.0 feat: GStack 2 fork port wave — egress receipts, context-bill, sharded gate, /health fix (#2541)
* test(helpers): shared skill-census helper with three explicit counts
physicalSkillFiles (symlinked dirs included, root router included),
authoredSkills (realpath-deduped, router excluded), registryEntries
(what ./setup registers: unique frontmatter names + _gstack-command).
One counting authority for the hermetic seeder, context-bill ground
truth, and the catalog-budget test — connect-chrome's dir symlink and
the root router otherwise produce three subtly different hand-rolled
censuses. Ported-wave foundation (C11).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(evals): stop the harness grading itself
findPreviousRun excluded only the file being written, by name, so every
suite compared against _partial-e2e.json — the current run's own
accumulator, relabelled with the current tier just before each flush.
That is why every block read '+$0.00, +0s, Stable run, no regressions.'
This harness has never been able to detect a regression, and reassuring
output that cannot fail is worse than none. In-progress runs are now
excluded by role, and a run with nothing to compare against says NO
BASELINE instead of claiming stability.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
(cherry picked from commit f3140b5245221fff7fb9411c7ec07c2ca11587b5)
* refactor(evals): shared partial-run predicate + finalized-run lookup
isPartialEval(data, filename) is the one place that decides what counts
as an in-progress accumulator (the _partial flag OR a _partial-prefixed
filename), and findLatestFinalizedRun(evalDir, tier) is the one place
that finds the newest real run — scanning the eval dir plus one level of
shards/<slug>/ subdirs, where the sharded paid runner points each
shard's collector. skill-budget-regression.test.ts's hand-rolled
findLatestRun (flag-blind: a flagged-but-renamed accumulator passed its
name check) is replaced by the shared helper.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
(cherry picked from commit b55fcf6966366fd21a8cdc46de61aab6e1b1d100)
* feat(evals): register shipped skills for hermetic PTY children
Hermetic children get a config dir that deliberately seeds no skills —
right for children that install their own, fatal for the PTY family that
TYPES /office-hours or /plan-ceo-review: claude rejects the command as
Unknown before any model turn, so the plan-family gate smokes measure
nothing. hermeticSkillsConfigDir() is a second, opt-in config dir under
the same runRoot that mirrors ./setup's registration exactly (real dir
per registry name, SKILL.md + sections/ symlinks, frontmatter-name
resolution, _gstack-command root alias), driven by the shared
skill-census so connect-chrome's dir symlink collapses the same way
setup's idempotent overwrite does.
Ported from fork commit 03c4eca2, tree walk rewritten for the upstream
layout (top-level <skill>/SKILL.md dirs, no skills/ tree). Unit tests
are new: seed shape, census parity, symlink resolution, connect-chrome
collapse, idempotence, no-API-key seed.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
(cherry picked from commit 93dae6107b30ce453a07c2d342b60262bba6ce0b)
* feat(evals): seedSkills opt-in for PTY slash-command tests + tripwire
Wire ClaudePtyOptions.seedSkills through launchClaudePty: when set (and
hermetic, and no per-test CLAUDE_CONFIG_DIR override), the child gets
hermeticSkillsConfigDir() so typed /skill slash commands resolve instead
of dying as Unknown command before any model turn. Opted in at the three
runPlanSkill* helpers and the four direct-launch slash-command tests
(plan-design-with-ui, plan-ceo-mode-routing, autoplan-chain,
ship-idempotency).
New static tripwire (test/pty-skill-seeding-wiring.test.ts): any test
file that sends a slash command over the PTY must route through a
runPlanSkill* helper or pass seedSkills: true — an unseeded slash-command
test spends money and measures nothing. hermetic-wiring.test.ts now
blesses the repo-tree seeding path explicitly (config dir under runRoot,
symlinks into the repo checkout, never operator ~/.claude).
The CI "Register gstack skills for PTY smoke" step keeps a keep-me note:
container cross-mount symlinks defeat the TUI scanner and HOME is not
hermeticized, so the real-file copies there must survive this change.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
(cherry picked from commit 63c52269daaffb833b3105ea9b4b99be6df8fec7)
* refactor(evals): single shared paid-test-set module
test/helpers/paid-test-set.ts is now the one definition of which test
files are paid (the exact globs package.json's test:gate expands).
scripts/test-free-shards.ts derives its free/paid exclusion from it
instead of a private regex list, dropping the dead
browse/test/security-review-fullstack.test.ts pattern (file no longer
exists). The sharded paid runner derives its enumeration from the same
module, so a file added to one list can no longer silently miss the
other.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
(cherry picked from commit a7f36479a6a1f3656452370f5883371f3cb65623)
* feat(evals): env-driven lazy eval dir + shard-aware store and tooling
Importing eval-store no longer spawns the gstack-slug subprocess: the
module-level DEFAULT_EVAL_DIR constant is now a memoized defaultEvalDir()
resolved at collector construction. Resolution order: explicit
constructor arg, then GSTACK_EVAL_DIR, then slug detection — so the
sharded paid runner can point each shard child at its own
<evalDir>/shards/<slug>/ dir with plain env, no --preload.
Runs collected under a shards/ subdir record their slug in the eval
JSON (EvalResult.shard). findPreviousRun scans one shards/<slug>/ level
and prefers same-slug priors, so each shard baselines against its own
history instead of whichever shard flushed last. eval:list,
eval:summary, and eval:compare enumerate the same one level of shard
subdirs; eval:compare's no-arg mode also stops picking an in-progress
accumulator as the after-run.
eval-watch stays flat (documented follow-up): it tails a single dir for
live progress and gains nothing from per-shard baselines until the
runner emits a merged stream.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
(cherry picked from commit e1f53f7d9c7fe6b65877d843f2e25bd2e2d12ffd)
* feat(evals): sharded paid tier runner
scripts/test-paid-shards.ts runs the gate/periodic tier one Bun process
per test file, with an EXTERNAL wall-clock timeout that SIGKILLs the
shard's detached process group and an aggregate that distinguishes
passed / failed / timed-out / never-started — partial execution can no
longer read as a pass. Bun's native --shard/--isolate covers none of
this: no process-group kill (hung claude/codex PTY grandchildren
survive in-process isolation), no never-started taxonomy, no per-shard
env. Each shard child gets GSTACK_EVAL_DIR=<evalDir>/shards/<slug>/
(slug = test filename sans extension, stable across runs) so shard
baselines compare against their own prior runs.
Output classification lives in scripts/test-strict-output.ts (strict
exit-code derivation, incremental fail-line classifier, child signal
forwarding) so the runner and any future strict bun-test wrapper share
one implementation. Enumeration derives from the shared paid-test-set
module; tier exclusion fires only on an explicit whole-file
EVALS_TIER === '<other>' guard.
package.json gains test:gate:sharded / test:periodic:sharded, and
eval:bg:gate / eval:bg:periodic now run the sharded scripts with detach
timeouts sized to the worst case (gate: 49 shards x 30min / 4 jobs ~
6.2h -> 25200s; periodic: 59 -> 28800s).
test/paid-shards.test.ts pins enumeration, tier classification, and the
kill-and-continue property with a real busy-loop shard.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
(cherry picked from commit 5e76bd5931836257f896cedfe4e93912cb759c70)
* feat(security): hash-chained egress receipt ledger (core)
Port lib/egress-receipt from the v2 fork as TypeScript: writeReceipt
(sync, fail-closed via typed EGRESS_RECEIPT_FAILED), best-effort
writeOutcome, readLedger/listReceipts/verifyLedger, GSTACK_HOME ->
GSTACK_STATE_DIR -> ~/.gstack resolution, 0600 ledger under a 0700
security dir, and an mkdir spin lock (2.5s budget) with documented
>10s-mtime stale-lock reclaim.
Changes vs the fork:
- lastRawLine tail-reads the final 4KB instead of loading the whole
ledger, so appends stay O(1) as the file grows.
- WARN-at-size: past 25MB writeReceipt emits one self-explanatory
stderr warning per process (what the ledger is, how to inspect it,
rotation TODO); verifyLedger gains a sizeWarning field. Rotation
TODO carries the chain-genesis sketch (new generation's first record
embeds the prior file's tail hash).
bin/gstack-egress-receipt is a bun script bridging shell callers:
write|outcome subcommands, exit 3 + EGRESS_RECEIPT_FAILED on stderr on
failure; --no-payload records sha256:null for git-class ops.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
(cherry picked from commit 619726a3d77d987a2e50151a5727b3faaaf5fc6a)
* chore(bin): delete dead brain-consumer/reader scripts
bin/gstack-brain-consumer and bin/gstack-brain-reader are byte-identical
dead scripts that POST the repo URL + a Bearer token to a /ingest-repo
endpoint gbrain removed (docs/gbrain-sync.md already documents the
removal in past tense). No live references remain; CHANGELOG mentions
are historical.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
(cherry picked from commit 254ddc69fc5a0270fcc973e36b6a81766d835d2d)
* feat(security): shared shell receipt helpers
bin/gstack-egress-lib.sh (sourced library, gstack-gbrain-lib.sh
precedent) provides _receipted_curl and _receipted_git: write the
egress receipt BEFORE the send via gstack-egress-receipt, hand curl the
SAME payload file via --data-binary @file so the receipt hash matches
the wire bytes exactly, then append a best-effort outcome. Per-call
fail policy: 'closed' refuses the send (return 3, problem/cause/fix
message on stderr) and 'open' warns and proceeds. Payload temp files
are consumed immediately per call — no EXIT traps, since callers like
gstack-telemetry-sync own their own EXIT trap and a sourced trap would
clobber it.
Tested end-to-end against a local Bun.serve listener: receipt sha256
equals the sha256 of the bytes the listener received, fail-closed
refusal never touches the network and carries the problem/cause/fix
stderr shape, fail-open warns and proceeds.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
(cherry picked from commit 6d067dce2d4c8815dec98be551763c85a3671357)
* feat(security): receipt core shell sinks
Wire the three core bash egress sinks through gstack-egress-lib.sh:
- gstack-telemetry-sync: the batch POST now writes the payload to a
temp file, receipts those exact bytes fail-closed, and hands curl the
SAME file. On refusal nothing is sent and the cursor does not
advance, so the batch stays buffered for the next run. The HTTP
status is recorded as the receipt outcome.
- gstack-update-check: fail-open receipts (warn + proceed) on the
Supabase ping POST, both VERSION curls (via a local
_receipted_version_fetch helper that skips non-network schemes), and
git ls-remote. The ping receipt is written inside the backgrounded
subshell, so it can never block the script's exit.
- gstack-brain-sync: fail-closed git-class receipts. The push receipt
is written BEFORE the commit consumes the queue, so a refused receipt
leaves the queue intact and the next run retries the whole drain
(pinned by a new queue-intact-on-refusal test, including the
problem/cause/fix refusal message shape). The retry-path fetch and
retry push carry their own fail-closed receipts.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
(cherry picked from commit 3c60f699acceaf1c92a218874711e05fc17dca5d)
* feat(security): receipt TS module sinks + tunnel
writeReceipt (fail-closed, sha256:null — a subprocess or SDK owns the
wire bytes) before every TS-module network-bearing operation:
- bin/gstack-gbrain-sync.ts: before the gbrain code walk that ships
repo content to the user's gbrain DB (may be remote Postgres). A
refused receipt fails the stage with status refused-egress-receipt.
- bin/gstack-memory-ingest.ts: before the gbrain batch import of
transcript pages. A refused receipt returns a system_error verdict
without spawning the import.
- browse/src/server.ts: before both ngrok.forward call sites (start-up
BROWSE_TUNNEL=1 path and the /tunnel/start endpoint). A receipt
failure lands in the existing catch that tears the tunnel listener
back down and refuses the start.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
(cherry picked from commit 5677d618a48fcd0ae2b068bf868781d90f809cb5)
* feat(design): receipted fetch for OpenAI calls
design/src/receipted-fetch.ts wraps every api.openai.com call: a
content-free egress receipt (sink design-openai, sha256 of the JSON
body — hash only, never the body) is written BEFORE the send. Polarity
is FAIL-OPEN: user-facing generation must not die because an audit log
hiccuped, so a receipt failure warns on stderr and the call proceeds.
Streams pass through untouched (response bodies returned as-is;
non-string request bodies receipted as sha256:null rather than drained
to hash).
All ten call sites converted with per-command payload classes:
generate, variants (injected fetchFn passes through), iterate (both
threaded and fresh paths), evolve (image + screenshot analysis), check,
diff, design-to-code, memory.
Unit-tested with injected fetch: receipt-before-send ordering, stream
passthrough, and fail-open on an unwritable ledger.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
(cherry picked from commit c0e5ff6639414ac2fd98e8ac3affb51401746b55)
* feat(security): receipt admin scripts + user git-ops (zero exceptions)
Wire the remaining shell egress through gstack-egress-lib.sh:
- gstack-gbrain-mcp-verify: both JSON-RPC probe POSTs (initialize +
tools/list) receipted fail-closed via payload files (hash == wire
bytes). A refused receipt lands in the NETWORK class — no send.
- gstack-security-dashboard / gstack-community-dashboard: the
community-pulse GETs receipted fail-open (read-only stats must not
break over an audit hiccup).
- gstack-gbrain-supabase-provision: api_call receipted fail-closed.
Each retry attempt hands the helper a fresh copy of the body file
(the helper consumes its payload). The receipt hashes the request
body only — the PAT never reaches the ledger or any log. Refusal
exits 8 without retrying.
- git-class sha256:null receipts, fail-open: gstack-artifacts-init
(ls-remote, initial push, fetch/pull recovery, retry push),
gstack-brain-restore (staging clone, existing-repo fetch),
gstack-session-update (self-update pull).
gstack-team-init needs no wiring: every git clone in it is inside an
echoed instruction string, not an executed command.
The lib now self-locates with shell builtins only (no dirname), so
sourcing works under the whitelist-PATH test harnesses.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
(cherry picked from commit b8c5e2055b21ab72878b3e46f8047782ee65a11c)
* test(security): egress wiring tripwire + polarity contract
Static-grep tripwire pinning the egress-receipt wiring (threat model in
the header: the ledger is forensic observability of ATTEMPTED egress,
not an exfiltration control):
- Per-sink assertions: every wired TS module imports egress-receipt and
calls writeReceipt; every wired shell sink sources
gstack-egress-lib.sh with each network op under a receipt;
ngrok-proximity check for server.ts; every design api.openai.com call
routes through receiptedFetch.
- Absence assertions: the dead brain-consumer/reader scripts stay
deleted (lstat, so a dangling symlink also fails).
- Polarity table pinned as data (fail-closed: brain-sync,
memory-ingest, gbrain-sync, telemetry-sync, ngrok, mcp-verify,
supabase-provision; fail-open: design-openai, update-check,
dashboards, git-class user ops, context-bill --exact) plus per-file
polarity spot-checks.
- NEW-SINK SCANNER with zero KNOWN_UNWIRED: sweeps bin/, lib/,
scripts/, design/src, browse/src for curl, absolute-URL fetch(, and
git remote ops (never local rev-parse/get-url; heredoc bodies and
message strings excluded) and requires every hit to be receipted or
in a REASONED exemption list where each entry carries its why.
Preamble-generated skill prose documented out-of-scope in the header.
- Shebang tripwire: no bin/gstack-* file may carry a node shebang.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
(cherry picked from commit ff69ceeafaf9c017d539b6ad77ff8f95b680b979)
* feat(cli): gstack-egress reader
bin/gstack-egress (bun) — the auditor's view of the receipts ledger:
- list: one row per receipt (what gstack ATTEMPTED to send), with
--since/--host/--sink filters and --json.
- verify: recompute the hash chain; exit 3 on tamper naming the first
broken line; prints the sizeWarning when the ledger passes 25MB.
- grants: what CAN leave, built on the upstream config keys only
(telemetry, artifacts_sync_mode, redact_repo_visibility,
redact_prepush_hook via gstack-config get) — each grant names its
file, key, and the exact revoke command.
CLI smoke tests spawn the real bin against a temp GSTACK_HOME,
including a broken-chain fixture asserting exit 3.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
(cherry picked from commit 9e24eca0f1069fea2ea69e7df4e9b256e93d59a3)
* feat(cli): context-bill — token bill-of-materials (stripped port)
lib/context-bill.ts, ported from the v2 fork and STRIPPED to the tiers
this repo's skills can exercise: ALWAYS-ON (per-skill frontmatter bytes
with dead-key and foreign-host-file flags), EAGER (SKILL.md + any
forced 'for every invocation' references), on-disk totals, --diff,
--budget, and --exact with the calibration table. The fork's
CONDITIONAL/TRANSITIVE/LAZY/FAST-PATH parsers understand only its
dispatcher layout and were dropped; the tier fields stay in the report
shape (empty/zero/null) so re-adding a parser is additive.
TOKEN_DIVISORS and their provenance docblock kept; --help notes
recalibration via --exact's calibration block.
Three upstream fixes over the fork:
(a) findSkillDirs treats the walk ROOT as a container — the repo root's
router SKILL.md is billed AND its children are walked (the fork
short-circuited and billed one skill); walkMd skips node_modules
and dot-directories.
(b) installed-tree layout: subdirs that are their own repo checkout
(a gstack/ clone inside ~/.claude/skills, detected by .git) are
skipped, and directory symlinks (connect-chrome) are followed with
a container-recursion cycle guard.
(c) ROUTER_KEYS widened to the upstream frontmatter contract {name,
description, version, allowed-tools, triggers, preamble-tier}.
--exact writes an egress receipt (sink context-bill-exact, host
api.anthropic.com) BEFORE any count_tokens POST; if the receipt cannot
be written the run degrades to the offline estimate with a warning —
nothing is sent unrecorded. bin/gstack-context-bill is the bun shim.
Tests: fixture-tree ledgers, the three fixes, --diff/--budget exit
codes, --exact with injected fetch (envelope subtraction, receipt
ordering, fail-open degradation), CLI smoke test, and ground truth
against THIS repo via test/helpers/skill-census.ts.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
(cherry picked from commit 675c19876b87ec927b555f5f64c7f93130b3de90)
* test(catalog): aggregate discovery-surface budget with ratchet protocol
Every host loads every skill's frontmatter name + description at
discovery, every session. applyCatalogTrim in scripts/gen-skill-docs.ts
shapes each description and the 160KB per-file warn covers body size,
but nothing capped the aggregate frontmatter — the catalog could grow
one reasonable-looking description at a time. This test is that
enforcement layer.
Measures the catalog via test/helpers/skill-census.ts authoredSkills
(symlink-deduped, root router counted separately as the _gstack-command
alias line item): 53 skills + router = 4,420 bytes = 1,105
token-equivalents today, asserted <= 1,150 (~4% headroom). Per-skill
sub-cap of 260 bytes (largest today: design-consultation at 229), plus
a non-empty-description check.
Failure messages are self-service ratchets: they print the new total,
the delta, and the update protocol (bump the constant AND the
derivation comment in the same commit; trim instead of grow for
existing descriptions). Parser handles folded block scalars
(description: >-) for fork parity; import-free by design so it
survives generator refactors.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
(cherry picked from commit c106fb36f768181b80c257e5cff1cde4f435f9c0)
* fix(browse): extension token bootstrap moves to pinned-origin POST; /health carries no token
GET /health is now liveness/status only in every mode — both token
carve-outs (headed-mode disjunct AND chrome-extension:// Origin
disjunct) are removed. Token bootstrap is POST /extension-token on the
local listener: the Origin header must be exactly
chrome-extension://<GSTACK_EXTENSION_ID> and the Host header's hostname
must parse to 127.0.0.1 or localhost (parsed via new URL, never literal
equality — Host arrives as '127.0.0.1:34567'). Wrong origin/host → 403
with no detail. The tunnel surface 404s the endpoint (not in
TUNNEL_PATHS, verified by test).
The extension ID is pinned by a new "key" field (RSA public key) in
extension/manifest.json; browse/scripts/extension-id.ts reproduces the
ID derivation (first 16 bytes of SHA-256 of the DER public key, hex
mapped 0-9a-f → a-p). The private key is not committed anywhere —
unpacked/baked-in loads only need the public key.
Extension side: background.js bootstraps and refreshes the token via
POST /extension-token (403 → disconnected state); sidepanel.js direct
connect path does the same; sidepanel-terminal.js's dead /health token
fallback (read AUTH_TOKEN/authToken keys the server never sent,
hardcoded port) is replaced with the window.gstackAuthToken path.
MIGRATION NOTE: the manifest key pins the extension ID, so existing
installs' side-panel local state (saved port, snoozes) resets once —
explained in-product via a one-time notice (flag
gstack_id_migrated_v162). After upgrading the server, restart the
browser so the old service worker stops polling for a token GET /health
no longer serves.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
(cherry picked from commit e9a0b6847a2d17fe6656a4686b4efd0c8380eb09)
* docs: correct stale compiled-binaries claim; file three egress/eval follow-ups
CLAUDE.md's compiled-binaries section claimed browse/dist binaries are
tracked by git and appear as modified in git status — false since
|
||
|
|
c7ae63201a |
v1.58.1.0 feat: hermetic local E2E + Conductor prose AskUserQuestion (#2004)
* feat: add shared call-time isConductor() helper
Single source of truth for Conductor host detection in TS consumers
(CONDUCTOR_WORKSPACE_PATH / CONDUCTOR_PORT). Reads the passed env at
call time, not a module-load snapshot, so unit tests can pin the env
inline without Bun --preload (esm-hoist-breaks-env-pin-bootstrap).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* test: harden question-preference-hook harness against ambient Conductor env
runHook copied all of process.env into the hook subprocess, so running the
suite inside Conductor (CONDUCTOR_WORKSPACE_PATH/PORT set) would leak those
markers. Strip them so the existing cases deterministically characterize
NON-Conductor behavior before the Conductor branch lands. Baseline: 15 pass.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* feat: PreToolUse hook denies AskUserQuestion in Conductor, redirects to prose
Conductor disables native AskUserQuestion and routes through a flaky MCP
variant that returns '[Tool result missing due to internal error]'. The
hook now denies any AUQ call in a Conductor session and instructs the model
to render a prose decision brief instead (transport avoidance, not preference
enforcement) — firing for one-way doors too, with a typed-confirmation
requirement for destructive paths.
Precedence: never-ask auto-decide still wins (user already settled those);
Conductor prose is the fallback for everything else; non-Conductor behavior
is byte-for-byte unchanged. Restructured the per-question loop to compute
eligibility without early-returning so the Conductor branch can run as the
fallback while preserving memoryContext on every exit.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* feat: Conductor renders AskUserQuestion decisions as prose by default
In Conductor, native AskUserQuestion is disabled and the MCP variant is
flaky, so skills now render every decision as a plain-text prose brief the
user answers by typing a letter — proactively, not as a failure reaction.
- Preamble emits CONDUCTOR_SESSION, gated on != headless so eval/CI inside
Conductor still BLOCKs instead of rendering prose to nobody.
- AskUserQuestion Format gains a Conductor-default-prose rule (auto-decide
preferences still apply first; prose decisions log via gstack-question-log
since PostToolUse never fires), a one-way/destructive typed-confirmation
rule, and a typed-reply continuation protocol for split chains.
- Regenerated all SKILL.md + ship golden fixtures; bumped affected carve
skeleton caps to absorb the always-loaded additions.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* feat: deploy the Conductor AskUserQuestion hook (setup + upgrade migration)
The PreToolUse hook only delivers its Conductor-prose guarantee if it's
installed, but setup skips hook registration in non-interactive (conductor/CI)
setups. Two fixes so layer 3 actually deploys:
- setup: treat a Conductor workspace as an implicit opt-in for the PreToolUse
hook on the silent fall-through (never overriding an explicit opt-out).
- migration v1.58.0.0: re-register the hook for existing Conductor installs on
/gstack-upgrade, idempotent and respecting plan_tune_hooks=no.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* test: E2E for Conductor prose + fix auto-decide-preserved GSTACK_HOME bug
- New skill-e2e-conductor-prose (periodic): Conductor env + plan-eng-review
surfaces a prose decision brief, not a silent skip. Header documents this is
end-to-end behavior coverage; the deterministic Conductor guard is the
question-preference-hook unit test (the PTY harness can't register the MCP
variant — Codex #10).
- Fix the pre-existing bug in auto-decide-preserved: it seeded the never-ask
preference under GSTACK_HOME=tmpHome but never passed GSTACK_HOME into the
PTY run, so the spawned claude read the real ~/.gstack and the preference
was inert (Codex #9). Now passes GSTACK_HOME + CONDUCTOR_WORKSPACE_PATH to
prove auto-decide still wins over the Conductor prose redirect.
- Register both in touchfiles (periodic tier).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* v1.58.0.0 feat: Conductor renders AskUserQuestion decisions as prose
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* test: strip ambient Conductor env in memory-cache-injection hook harness
Same dev-in-Conductor leak fixed for question-preference-hook: this suite's
runHook copies process.env, so running it inside Conductor flipped the
defer-path memoryContext assertions into the [conductor] prose deny. Strip
CONDUCTOR_* so the cases characterize non-Conductor behavior. (CI is headless,
so this only bit local Conductor runs.)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* feat: gstack-detach — run agent eval/bench jobs in their own session
Long agent-run jobs (30-60 min evals, benchmarks) die when the harness sends
SIGTERM to a background task's process group on turn boundaries / monitor
stops / interruptions (observed: 'script test:gate terminated by signal
SIGTERM'). gstack-detach runs the command in a fresh session (python3
os.setsid, or setsid on Linux, nohup fallback) so a group SIGTERM can't reach
it, and wraps it in caffeinate -i on macOS so idle-sleep can't kill it either.
Returns immediately; caller polls the logfile. Secrets stay in env, never argv.
The guard test pins the contract: the command runs in a different process
group than the caller and outlives the launching shell.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* feat: eval:bg* scripts — detached eval runs for agents
Agent-facing convenience scripts that launch the eval suites through
gstack-detach so a harness SIGTERM can't kill a long run. eval:bg (diff-based),
eval:bg:all, eval:bg:gate, eval:bg:periodic — each returns immediately and
streams to /tmp/gstack-evals.log for polling. The plain test:evals / test:e2e
scripts stay foreground for humans.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* docs: CLAUDE.md — agents must run long evals via gstack-detach
Codifies the detached-execution default: agent-launched eval/benchmark runs go
through bin/gstack-detach (or the eval:bg* scripts) so a harness SIGTERM or
macOS idle-sleep can't kill a 30-60 min run, then poll the log with a
death-aware watcher. Humans keep foreground scripts.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* feat: harden gstack-detach against all four eval-infra killers
The basic bash detach fixed SIGTERM but a real run on a shared dev box hit
three more killers: cross-worktree API saturation (15-way concurrency x a
sibling worktree mass-timed-out the suite), a silent hang (periodic bun died
with no exit marker), and shared-/tmp log contamination (a concurrent
worktree's agent output bled into the log). Rewrite as a portable python3 tool
that bakes in all four fixes:
- fork + setsid: SIGTERM-proof (own session, survives harness polite-quit)
- caffeinate -i on macOS: no idle-sleep death
- --lock NAME (fcntl, machine-wide): concurrent worktrees SERIALIZE instead of
saturating the shared model API
- run-scoped default log (~/.gstack-dev/eval-runs/<label>-<slug>-<branch>-<ts>-<pid>):
no cross-worktree collision/contamination
- --timeout watchdog + a guaranteed '### gstack-detach EXIT=<code> ###' sentinel
on every terminal path: no silent hang, finished-vs-died always detectable
Guard test pins all four: detached pgid differs + outlives launcher, run-scoped
log path, watchdog EXIT=timeout, and lock serialization (second run WAITS).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* feat: eval:bg* use run-scoped logs + machine lock + watchdog
Drop the shared /tmp/gstack-evals.log path (the cross-worktree collision that
contaminated a live run) for gstack-detach's run-scoped default, and add the
machine-wide gstack-evals lock (concurrent worktrees serialize, no API
saturation) plus per-tier watchdog timeouts (60/90/120 min). Each eval:bg*
prints its run-scoped log path to poll.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* docs: wire detached-eval guidance into /ship + correct CLAUDE.md flags
- /ship eval step (sections/tests.md): long eval suites launch via gstack-detach
(own session, machine lock, EXIT sentinel) so a turn boundary can't kill a
30+ min run mid-ship — the exact failure observed during this branch's ship.
- CLAUDE.md: correct the now-stale /tmp reference; document the --lock (serialize
worktrees, no API saturation), --timeout watchdog, run-scoped log, and the
guaranteed EXIT sentinel the poller breaks on.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* refactor: extract pure promotedEnv() from conductor-env-shim
Single source of truth for GSTACK_* key promotion semantics. The ambient
promoteConductorEnv() becomes a wrapper; behavior-preserving. Needed by the
hermetic env builder which must not mutate process.env.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* feat: hermetic child-env builder for E2E runners
Allowlist scrub (basics/network/named-auth kept; CONDUCTOR_*, CLAUDE_*,
GSTACK_*, MCP_*, GBRAIN_*, operator credentials dropped), per-runner
extraAllow, overrides merge last, EVALS_HERMETIC=0 byte-identical escape
hatch read at call time (ESM-hoist safe). Sync memoized singleton temp dirs
(<runRoot>/.claude keeps the extractPlanFilePath contract), seeded
.claude.json for non-interactive first run, pid-aware GC of crashed runs.
19 free unit tests.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* feat: session-runner spawns hermetic children + isolation canaries
claude -p children now get the allowlist-scrubbed env and a gated
--strict-mcp-config (EVALS_HERMETIC=0 restores operator env AND args).
Two gate-tier canaries make the clean room falsifiable: hermetic-canary
asserts env redirect + scrub + zero MCP servers + nonzero API-key cost
from the Bash tool_result (never model prose); hermetic-sentinel plants a
poisoned operator config (user CLAUDE.md + MCP server) and proves the
child cannot see it. Empirically verified on claude 2.1.175: print mode
needs no seed config (the seed serves the PTY path); the child CLI sets
CLAUDECODE for its own tools, so that scrub is pinned in unit tests, not
E2E. hermetic-env.ts joins GLOBAL_TOUCHFILES.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* feat: PTY runner spawns hermetic claude sessions
launchClaudePty children get the allowlist-scrubbed env, a gated
--strict-mcp-config, and the session exposes hermeticConfigDir for
forensics (hermetic plan files live under <dir>/plans/ and still match
extractPlanFilePath via the /.claude dir-name contract). Seeded trust
state covers repo-cwd sessions; the 15s trust-watcher stays as fallback.
Verified foreground via the plan-mode-no-op gate test.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* feat: codex/gemini runners spawn hermetic children
Same allowlist scrub as the claude runners, with each provider's auth
surface re-admitted via extraAllow (codex: OPENAI_API_KEY/CODEX_* plus
its tempHome .codex copy; gemini: GEMINI_*/GOOGLE_* with real HOME for
~/.gemini auth). The gemini spawn previously inherited the full operator
env with no env property at all.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* feat: agent-sdk-runner spawns hermetic children via complete Options.env
The historical 'env: breaks SDK auth' failure was partial-env replacement:
Options.env replaces the child's entire environment, so objects lacking
ANTHROPIC_API_KEY killed auth. Passing the complete hermetic env (key +
PATH + redirected CLAUDE_CONFIG_DIR/GSTACK_HOME) works — validated live
via query() with a Bash tool call (success, real cost, Conductor vars
scrubbed). Per-test opts.env merges last; ambient key mutation still
works because the builder reads process.env at call time.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* test: static tripwire pins hermetic wiring in all five runners
Free-tier invariants: every runner builds child env via hermeticChildEnv,
no raw ...process.env spread at any spawn site, --strict-mcp-config gated
on isHermeticEnabled in both claude runners, and no test callsite passes
the operator env into a runner's override parameter (scoped to runner
calls — unit tests spawning gstack bin scripts directly are exempt).
Mirrors the terminal-agent-pid-identity / server-embedder-terminal-port
tripwire idiom.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* test: refresh codex/factory ship goldens with detached-eval block
|