Commit Graph
186 Commits
Author SHA1 Message Date
garrytan f4ab5ee75f test(judges): sample the recommendation rubric as a panel; never re-ask armJudge
llm-judge-recommendation is a judge case: each fixture now draws a
3-sample judgePanel, gates reason_substance on the panel mean and the
present/commits/has_because checks on a 2-of-3 majority, thresholds
unchanged. armJudge no longer re-asks on a malformed verdict; it is a
failed sample, as the judge policy requires.
2026-09-29 19:51:35 +00:00
garrytan b2ca207cf0 Merge remote-tracking branch 'origin/capy/rel-a' into capy/rel-c 2026-09-29 19:47:17 +00:00
garrytan 74c3184dcc test(pty): grant an owned Create pane whose title row is cropped
The targeted batching rerun on Claude Code 2.1.284 left its first report
Write unanswered for 1,372 s and timed out: the viewport began at the
pane's relative file row and rule, with the 'Create file' title cropped
above, so the preview parser rejected the file row as foreign. That row
must now resolve to the owned path and is skipped before the unchanged
line-by-line preview match. Replay controls reject another file, another
directory and an edited preview row.
2026-09-29 19:47:15 +00:00
garrytan 62fb9a255d feat(evals): stamp trial series identities and fit panels to the live registry
- scripts/eval-trial-series.ts stamps series_identity (eval-flake-rank's
  caseSeriesIdentities) on a report's trial-outcomes JSONL as its own step,
  keeping the history tool out of the paid runner's closure;
  TrialOutcomeRecord gains the optional series_identity field.
- Slice-count plans let a registered trial spill into an ordinary lane when
  its siblings hold every long lane, so panels never share a runner.
- Re-audited test-selection.ts (Stream B added the E2E_KINDS/BEHAVIOR_WHY
  map-diff; no new module loading) and repinned its hash.
- Detach and release floors now count trial shards (66 periodic trials in
  22 panels): periodic floor 33,821s, still under eval:bg:periodic's 67,380s.
- Coordination fixtures supply the executor's trial records.
2026-09-29 19:30:43 +00:00
garrytan a385e5de18 Merge remote-tracking branch 'origin/capy/rel-b' into capy/rel-a 2026-09-29 19:24:03 +00:00
garrytan 49761c97d6 test(ceo-mode-routing): submit a mode review that scrolled past the viewport
Run 36606688266 bundled routing, learnings and the mode choice into one
native call. Its review panel was taller than the terminal, so the tab
bar scrolled off, ceoModeSubmissionInput returned null for 240 s and HOLD
SCOPE was never submitted ('no posture match'). With no bar on screen the
viewport must still end at the focused Submit prompt, and the accumulated
screen text supplies the one complete panel, authenticated exactly as
before. Replay controls reject another mode, an unoffered answer, an
altered question, a quoted panel, trailing output, a moved cursor and an
answered or changed call.
2026-09-29 19:16:01 +00:00
garrytan 8bd53faa6c test(eng-batching): bind unsourced native briefs through the report's target
Run 36606688266 asked ten separate native review questions (D1-D9 bound
to ledger records R1-R9) and failed reviewCount=0 < FLOOR=3: its briefs
named the plan by title instead of citing PLAN.md, its report declared
'Review target (fixed): PLAN.md' under '# Engineering review: <plan>', and
it kept an unfenced copy of the plan's own H1. The named-source route now
accepts those spellings and non-inline ledger briefs. The same replay
rejects a foreign, mixed, duplicate or missing target, another plan's
title or copied H1, a brief naming another plan or file, a mismatched
saved brief, and re-asks. The run-36597762183 capture still counts 3.
2026-09-29 19:15:06 +00:00
garrytan 5269452983 test(eng-batching): grade the floor once the review report is complete
A completed GSTACK REVIEW REPORT ends the review, so the review-question
count is final there. Run 36606688266 wrote its report at 1,248 s and
closed the session at 1,318 s; the case now stops collection and applies
the unchanged floor at the report instead of waiting out the session.
No budget changes.
2026-09-29 19:15:06 +00:00
garrytan d6b3559b78 feat(evals): per-case pass rates with Wilson intervals, identity series and quarantine policy
scripts/eval-flake-rank.ts becomes eval:pass-rates (eval:flake-rank stays an
alias, and the legacy aggregate stays exported). It reads eval-store's
trial-outcomes JSONL from the last N completed evals-periodic runs on this
branch and main (gh, downloading only the trial-outcomes artifact, cached and
size-capped, parsed as data), plus local eval dirs, and prints per-case
per-trial pass rates with 95% Wilson intervals.

A series is a case's own touchfiles minus GLOBAL_TOUCHFILES
(caseSeriesIdentities, for the report job to stamp), per model, CLI version
and policy version. Labels: INCONCLUSIVE, BROKEN, FLAKY, FAILING, PASSING.
--backfill imports legacy slice artifacts as pre-policy trials (first
attempt only, attributed by registry id, never guessed) for display only.

--gate fails with ACTION REQUIRED on post-policy evidence only: drift below
the quarantine entry rule, a rule case behaving like behavior, a one-sided
Fisher drop against the previous identity (Holm-controlled), and quarantine
entries that met their exit rule, expired after 8 weekly runs, broke the
10% tier cap or are invalid. CASE_QUARANTINE entries now carry a
failureClass (detector, harness or model-latency); a product defect has no
class and is never quarantined. The policy test pins EVAL_POLICY's approved
constants.
2026-09-29 19:11:27 +00:00
garrytan 0292ee3f6f test(evals): classify every live case and re-select a case when its kind changes
E2E_KINDS: rule by default (191 E2E ids), 22 behavior cases whose verdict is
a live model choice with an acceptable sub-100% per-trial rate, each with a
BEHAVIOR_WHY tolerance, and 25 judge entries (the 24 workflow judges plus the
fixed-fixture llm-judge-recommendation rubric check). Contract-shaped cases
(ask-before-decide, plan-mode no-writes, mandated steps, secrets, the batching
floor) stay rule. Behavior requires a known literal registration and an exact
Bun test name so the case runs as its own trial shard.

Map-diff selection now diffs E2E_KINDS and BEHAVIOR_WHY per key, and a base
revision without them selects every key, so a kind flip runs the panel it
introduces. test/eval-kinds.test.ts enforces coverage, tolerances,
isolatability and the reviewed counts, printing the literal to add.
2026-09-29 19:11:04 +00:00
garrytan 3f68572cab test(llm-judge): sample every judge as a pre-registered 3-sample panel
Each of the 24 skill-llm-eval judges now draws EVAL_POLICY.judge.samples
independent samples of the same prompt concurrently inside the unchanged
JUDGE_MS budget. Numeric dimensions gate on the per-dimension panel mean
against the unchanged threshold; booleans (would_browse, consistent) on a
strict majority. An erroring sample fails the whole panel and is never
resampled; a refusal is an unscored panel only when every sample refused.
callJudge's 429 backoff stays: it is transport before any model output.

The workflow-judge cache stores and validates only complete panels, and its
identity now records the panel and zero file retries. Harness tests that
pinned one provider call per case now pin the panel size.
2026-09-29 19:11:04 +00:00
garrytan b1f5bc0032 test(evals): retire every paid automatic retry
Paid evals never retry (approved 2026-09-29): delete SHORT_CASE_RETRY_FILES
and retriesWithinCaseCap, drop the retry fields from the registered wall rows
(walls now cover one run plus reserve), make retriesForFiles return 0, pass
--retry 0 explicitly, and drop --retry 1 from the package.json paid scripts.
Add the eval:pass-rates alias. Tests that pinned the old retry allowance are
updated as a policy change; review-finalization-budget now proves late-result
recording under the production zero-retry arguments.
2026-09-29 19:09:51 +00:00
garrytan adced7e046 test(evals): add trial records, panelVerdict, expectContract and trial-outcomes JSONL
EvalTestEntry gains case_id, kind, trial, panel, failure_class and
policy_version, stamped from the runner's TRIAL_ENV on isolated trial
shards. panelVerdict() is the single verdict function (INCOMPLETE on
missing or duplicate trials, contract veto at any count, quarantine
hard-break rule, INFRA/INCOMPLETE machine classification). expectContract()
records failure_class 'contract' on the collector entry and a sidecar
before throwing. trial-outcomes JSONL has a fail-closed writer and a
data-only reader.
2026-09-29 18:56:27 +00:00
garrytan 3c441fe311 test(evals): add E2E_KINDS, BEHAVIOR_WHY, EVAL_POLICY and CASE_QUARANTINE skeletons
Every E2E_TIERS and LLM_JUDGE_TOUCHFILES key starts as 'rule'; BEHAVIOR_WHY
and CASE_QUARANTINE start empty. EVAL_POLICY pre-registers the approved
panel (3, majority 2), quarantine entry 0.95/10 and exit 0.97/10, 10% cap,
8-weekly-run expiry, Fisher drift alarm and one INFRA re-dispatch.
2026-09-29 18:56:27 +00:00
garrytan 1c32c7d16a test(ceo-mode-routing): accept the skill-mandated Note form and Recommendation reason as HOLD posture
HOLD Defer/Keep briefs must use 'Note: options differ in kind' (preamble),
but the answered-HOLD path demanded a Completeness score, rejected a
one-line Net with a semicolon, and read posture only from ELI10. The rerun's
brief applied HOLD SCOPE in its Recommendation reason. Revert the
ineffective 'always'/'handoff chat' wording: two runs still skipped the
mode handoff.
2026-09-29 17:10:33 +00:00
garrytan 41dca7609d test(design-floor): answer the seed-declared all-seven 0D focus menu while it is pending
The actor declares 'Design: review all seven dimensions', but its picker
reused designReviewSetupAUQ, which only matches already-answered calls
(and a narrower header/label set), so the pending D1 focus menu was never
answered and the case waited out its 609 s deadline. The skill's Step 0D
requires asking; the fixture now answers it.
2026-09-29 17:07:38 +00:00
garrytan 3a888896db test(ceo-section-loading): recognize an arrow-ordered stale-fill execution by event roles
The census review traced the seeded race as 'R1 miss -> R1 store read (v1)
-> W commit v2 -> W cache.delete -> W fulfills -> R1 cache.set(v1) -> R2
(begun after W) hits v1', but the in-flight gate only accepted race
vocabulary or fixed sentence shapes. Order, actor, version and dismissal
mutations still fail.
2026-09-29 17:05:15 +00:00
garrytan 66d49280a5 test(outside-plan-disabled): compare named record timestamps as instants; negated authorship is not a current claim
The repair rerun named the seeded record by its ISO second
(2026-09-29T16:58:52Z vs .727Z) and said 'I did not write'; both were
misread as a foreign timestamp and a current write.
2026-09-29 17:02:31 +00:00
garrytan 97f0eee33d Merge remote-tracking branch 'origin/capy/audit-fix-wave' into capy/fixwave-baseline-repairs 2026-09-29 16:57:56 +00:00
garrytan 2bf97ab077 test(outside-plan-disabled): bind quoted prior-record values by their sentence, not phrase order
The parent obeyed the off switch and twice named the seeded completed
record as pre-existing, once with the quotation after its owner and once
with slash separators; the order-specific stripper counted both as current
completion. Timestamp, location, current-claim and value-match controls
still reject.
2026-09-29 16:57:54 +00:00
garrytan 75b22463f3 fix(plan-ceo-review): restore experience-first expansion framing, require the mode handoff, skip pacing menus
Census 36597762183: both mode-routing runs logged provenance and moved on
without the mandated handoff chat; the EXPANSION run asked an unauthorized
batch/narrow pacing menu instead of the first per-addition question; the
expansion-energy proposals led with the spec because v1.87.6.0 dropped
'lead with the felt experience'. The HOLD review detector also rejected a
decision whose grounding line named no plan file although the owned source
Read binds it.
2026-09-29 16:57:54 +00:00
garrytan 0554f3d959 fix(evals): plan CI-unrunnable cases as excluded entries, not empty case shards
design-review-fix drives the Aside browser and registers test.skip on Linux
runners, so its case shard executed zero cases and failed the exact-one-case
check in proof census 36597762183 (eval-slices 6). CASE_CI_EXCLUDE (reason +
tracking, beside PERIODIC_CI_EXCLUDE) now turns such cases into excluded
manifest entries that --list and the manifest surface; every planned case
shard still must execute exactly its case.
2026-09-29 16:55:55 +00:00
garrytan 45cdfab2a0 test(ship-docsync): shard the doc-sync lifecycle by case and drop the duplicate dispatch-only case
ship-docsync ran the same fixture and prompt as ship-docsync-completion and
asserted a subset of it. The file now runs one case per process, so its lane
wall is its longest case instead of half the sum of thirteen.
2026-09-29 16:47:24 +00:00
garrytan bfd1783fbf Merge remote-tracking branch 'origin/capy/fixwave-baseline-repairs' into capy/audit-fix-wave 2026-09-29 16:43:05 +00:00
garrytan 0023d011a3 feat(evals): ~12-minute blocking paid lanes and a non-blocking marathon lane
- Planner budget mode (--slice-budget S --jobs J): recorded per-tier wall
  times pack into as many ~9-minute executors as the work needs; the plan
  records per-slice estimates and the CI job timeout (supervised worst case
  + 20 min). evals.yml and evals-periodic.yml derive matrix size and
  timeout-minutes from it; max-parallel covers every slice at once.
- Case shards: plan/design/review-army/shared-libs(-paths) run one registered
  case per process (<file>#<case id>, exact name pattern, exactly one case).
- Retry rule: a timed-out attempt is a verdict. Only files whose every case
  budget is CAPTURE tier or shorter keep one retry; walls shrink to match.
- Marathon tier: positive selection, excluded from gate/periodic planners,
  run by the new evals-marathon.yml (weekly + dispatch, fresh, own report).
- PR-lane E2E reuse of verified first-attempt passes on identical inputs;
  the report rejects reuse outside the fast PR profile.
- Duration seed from census run 36385945043, per tier and per case shard.
2026-09-29 16:19:52 +00:00
garrytan acb7bc02b1 refactor(evals): share the import-closure walker and add the E2E shard reuse identity
sourceDependencyClosure moves from the workflow-judge adapter into
scripts/eval-input-cache.ts unchanged, so judge keys stay byte-identical.
scripts/e2e-shard-reuse.ts builds the consumed-input identity of one PR-lane
E2E shard (test import closure, every registered case's touchfiles, globals,
runner/workflow/setup actions, child env pins, CI image, Claude CLI) and fails
closed on anything unknown. Marathon joins the always-fresh purposes.
2026-09-29 16:19:51 +00:00
garrytan 39fc663894 fix(sync-gbrain): define Step 4 helper args and one atomic write path
Both census read-ready attempts spent turns reading the helper source to
resolve <user-args>, inspecting fixture internals kept inside the repo,
and reconciling 'Read + Edit' with the tmp+mv atomic write, then hit
max turns before the verdict.
2026-09-29 16:19:39 +00:00
garrytan 4973f96d98 test: move the full office-hours workflow to marathon; add a periodic design-draft checkpoint
The full startup workflow runs 1–3 real spec-review rounds (~280s each) and
hit its 1200s capture in run 36385945043 at finalize. Review depth is the
product's loop, so the case cannot fit a blocking lane without cutting
rounds. It is now marathon tier with every assertion unchanged.

skill-e2e-office-hours-design-draft.test.ts (periodic) runs the same fixed
interview only through the Write that creates the design (269s in that run)
and applies the full validator's design-draft checks, the required section
reads and the launch/foreign-skill-read guards. validateOfficeHoursDesignDraft
is extracted from validateOfficeHoursCompletion, which still applies it.

Selection: office-hours-design-draft is registered periodic; the marathon-only
file is already excluded from the gate and periodic plans by the B5 planner
rule. Tier-alignment regexes and the valid-tier check accept 'marathon'.
A type-only cast in plan-scope-selection.test.ts removes a diagnostic whose
union print order made the ratchet identity unstable; baseline tightened.
2026-09-29 15:29:46 +00:00
garrytan a41cdb7d7d test: add the non-blocking 'marathon' E2E tier
Full start-to-finish flows move out of the blocking lanes. E2E_TIERS and
E2ETier gain 'marathon'; describeE2ETier('marathon') is enabled only when
EVALS_TIER=marathon, so the gate/PR and periodic lanes (and the gate census)
never run those cases. The PR profile accepts marathon ids as scheduled
elsewhere and defers them with their own reason, even on full fallback.
2026-09-29 15:21:17 +00:00
garrytan dc5934ee54 fix(test): recognize ledger row-ID split candidates so collection stops at the last ACK
Run 36385945043's split-overflow case asked all five candidate decisions by
8m55s, but the live candidate check required the question to open with
"E1:" and every option to be a known disposition. The skill cited ledger
row IDs ("D2.1 — R-E1: …") and offered "Hold, discuss first", so no
candidate was recognized and the attempt ran the whole review (1302s).

Identity now comes from the native header; the question must open with that
candidate's ledger reference, name only that candidate, and offer exactly one
include, defer and cut disposition. The selected answer must still be one of
those three. The semantic evaluator and every existing negative control are
unchanged; a trimmed capture from the run adds the positive case and four
row-ID negative controls.
2026-09-29 15:18:46 +00:00
garrytan 2e1dff825d Merge capy/audit-fix-wave (#2994) into the harness branch
#2994 deletes the plan-*-finding-count evals, ceo-payment-findings.ts and
design-count-review.ts. Drop the CEO throw diagnostics and Design boundary
work with them, and drop the structured completion predicate, stopReason,
review-log binding and plan/review-log evidence copy: no surviving
runPlanSkillCounting caller passes expectedPlanPath, so they would be dead
code. Keep idleFor in timeout summaries (every counting caller can time
out), asserted in the existing timeout test. W7 and W8 are unchanged.
2026-09-29 15:14:06 +00:00
garrytan 77aa845088 test: CEO classifier throws name the question and matched predicates
Replaying run 36385945043's FAN-1 and ERR-1 throws (ledger rows
reconstructed from rendered diffs) through ceoPaymentFinding: the email
obligation's row, subject, option and proposal predicates pass and the
ELI10 explanation-defect predicate fails first ('lets that exception fly
out', 'the error bubbles up').

Binding the defect to the named ledger row instead (the planned fix) was
tried and reverted: scoped to the email seed it flips 30+ existing cf74
still-rejects replays, which require a vocabulary-free, ledger-bound email
question to earn credit only through a complete saved comparison. With
FAN-1's rendered currentDecision payload reconstructed, the recorded-
decision path counts it, so the real saved plan (not uploaded) must have
differed; failure artifacts now retain it.

The classifier stays fail-closed and unchanged. Its throw now prints the
header, the first 200 question characters and each obligation's predicate
results. Free regressions with provenance and negative controls: an
unrelated question, an email question whose row says it is already
rescued, and a ledger ID whose row belongs to another seed.
2026-09-29 15:04:36 +00:00
garrytan d05097b157 test: structural Design count boundary; TODO proposals are not findings
Replaying run 36385945043 through the Design count predicates: routing,
focus and learnings setup was not recognized as setup, Issue 1 was counted
pre-review in both attempts (the boundary fired on it), and attempt 2
counted the Font TODO proposal as a finding (review=4 and review=5 for five
issues). The paid caller now starts review at the first answered native
decision that is not setup (recognized packet, or setup header/question ID),
a completion handoff, artifact rendering or a TODO proposal (the review's
Add to TODOS.md / Skip / Build it now menu). TODO proposals are recorded as
administrative extra decisions. The replay asserts each counted call: both
attempts review=5 (Issues 1-5). isDesignCountFirstReview and its controls
are unchanged.
2026-09-29 15:04:28 +00:00
garrytan f1eeec384e test: judge plan-count completion on structured evidence, not wording
Replaying run 36385945043's two Design attempts showed the existing routes
rejected correct endings: attempt 1 at the typed-completion path field
('- Reviewed plan written to …' is not a 'Plan written to' line), attempt 2
at the leading-fence veto (its final message opens with the dashboard).

nativePlanTerminalPreconditions is the structural prefix of
hasNativePlanTerminal (behavior unchanged). structuredPlanCompletion adds,
inside the existing nativeSummary branch: a complete report (Design
binding for Design), a completed review-log row for the expected skill
appended during this attempt under the child's GSTACK_HOME/project slug
(resolved with bin/gstack-slug) and stamped with the fixture commit, timed
between the report/last answer (second resolution) and the final native
message, a final message with stop_reason end_turn (now carried on public
transcript messages), and no visible question or permission prompt.

Timeout summaries add idleFor and lastTerminalCandidate. Terminal and throw
captures copy the plan file and review-log rows into the artifact
directory; copies are best-effort and recorded in evidence-copy.json.
Free regressions: both captured Design endings (trimmed fixture with
provenance; report, row and end_turn reconstructed and labelled), the
negative controls, and real-PTY completion/timeout runs through the real
review logger.
2026-09-29 15:04:28 +00:00
garrytan a0aa1e8f39 test: one finite PR world for the shared-libs fixture; name dual-voice probe evidence
The shared-libs shim served 2 PRs for pulls?state=all and endless full pages
for state=open. gh pr list, pulls?state=open|all|closed (per_page/page,
short last page, direction) and search/issues now page one deterministic
table: PR 7, 600 older open PRs, PR 42 and 3 closed PRs, so five 100-item
open-metadata pages still leave older open PRs unchecked. The Contents API
lists pinned directories (the captured attempt got 404 for contents/ and
contents/src while files resolved, then fell back to a raw host), unknown
endpoints return 404 instead of repo metadata, and the read-only detector
is unchanged. Free tests cover view agreement, the budget bound, gh/curl
agreement and the empty world.

Dual-voice outside-voice failures now report probeToolUseId, probeMode and
the canonical-match result with the reason the probe output was rejected.
2026-09-29 14:28:11 +00:00
garrytan 049fdc315b test: pin DISABLE_AUTOUPDATER in hermetic env and capture corrupt-seed warning
Both EVALS_HERMETIC branches of buildHermeticEnv now carry
DISABLE_AUTOUPDATER=1 (the allowlist scrubbed the workflow's copy, so every
PTY screen showed the updater's npm-prefix failure). Per-test overrides
still win. The corrupt durations-seed test now captures its expected
warning and restores the console spy.
2026-09-29 14:22:05 +00:00
garrytan b421bba2c9 Merge origin/main (v1.91.7.0) into test-audit-reduction
Keep both intents: v1.91.7.0's functional QA, docsync and exploratory
paid cases and their free owners stay; this branch's deletions stay
deleted. main's new paid keys follow the derived-closure touchfile rule
(free *.test.ts paths dropped, static helper/fixture closure added), its
new helper-only tests join the ratchet baseline, and its free selection
examples that named free test files now assert the derived selection.

Periodic CI keeps seven slices without the retired Autoplan slice; the
gate census keeps seven single-worker slices with --skip-judges. Wall
and census literals are recomputed from the merged planner, durations
are re-recorded on Ubicloud, and VERSION stays 1.91.8.0 above 1.91.7.0.
2026-09-29 13:48:02 +00:00
Garry Tan dcaea52800 v1.91.7.0 feat: add functional QA and pre-publication docs checks (#2983)
* feat: add surface-aware exploratory QA and ship documentation gates

* test: preserve delegated QA setup authority after main integration

* fix(qa): clarify exploration order and preserve report artifacts

* test(qa): follow the shared setup reference directly

* refactor(ship): make verification and recovery routes explicit

* test(ship): align evidence and review guards with explicit routes

* fix(workflows): clarify ship recovery and functional QA evidence

* fix(workflows): clarify approval recovery and full QA coverage

* refactor(workflows): order review transactions and clarify ship state

* fix(ship): clarify final verification and fail closed at publication

* fix(evals): attribute native atomic documentation writes

* fix(ship): clarify recovery and documentation lifecycle guidance

* fix(test): preserve observed native placeholder styling in CI

* fix(codex): report watchdog timeouts without a process-exit race

* Checkpoint functional QA implementation and workflow validation repairs

* Fix documentation and shared-review fixture contracts

* docs: clarify judge reuse and evaluation supervision

* test: align review evidence and selected case contracts

* test: verify append-only documentation checkpoints and recovery

* fix: qualify QA workflows and CI validation repairs

* fix: launch shared-libs fixture scripts on Windows

* fix: qualify QA deadlines, fixture isolation, and shard cleanup

* fix: preserve qualified QA and cancellation repairs

* fix: enforce functional fixture authority and share strict event decoding

* fix: retain free-test evidence and explain recovery

* fix: reject malformed native evidence after decoder consolidation

* test: use reliable capture for telemetry privacy filters

* test: refresh measured quick coverage and document validation costs

* Fix native fixture receipts and preserve VM validation evidence

* Align negative judge controls with upstream clarity policy

* Fix report-only QA preparation and public evidence handling

* Clarify QA-only preparation and current-report preservation

* Stream Ship quality judgments with an explicit 64k response contract

* Validate compact judge reasoning locally with supported wire schema

* Align functional QA fixture instructions with evidence acceptance

* Bind native browser diagnostics to execution evidence and align review verdicts

* Preserve native diagnostic line boundaries

* Serialize functional QA evidence from native captures

* Keep large QA evidence fixture payload out of Windows argv
2026-09-29 06:07:35 -07:00
garrytan 13e4f37201 test: guard the reduced suite against new test-of-test files
- test/test-of-test-ratchet.test.ts records the 228 free tests that import only
  test/ code and fails on a new one, naming the owner test to extend instead;
  a stale baseline entry fails with the remove instruction.
- test/helpers/resolve-repo-path.ts is the one specifier/literal resolver for
  the ratchet and the touchfile closure invariant, with its own unit tests.
- CONTRIBUTING "Test tiers" describes the paid-failure workflow (fix, then one
  row in the detector's owner test) and the ratchet; TEST_PORTFOLIO gains the
  detector -> owner-test table and no longer claims an Autoplan chain eval.
- TODOS: automatic exclusion policy for chronically red periodic files (P3),
  the deferred native-completion table collapse, the unused CEO payment
  seeder; the PTY readiness item is narrowed to the paid runner.
- docs/test-audit-2026-09.md collects the triage, security mapping, inventories,
  selection proof, behavior-commit decisions and retained false positives.
2026-09-29 09:16:19 +00:00
garrytan 7cbb5e04e4 test: derive paid touchfiles from each eval's static closure (E)
touchfiles.test.ts now checks, per key, that the paid file's static
test/helpers and test/fixtures closure (plus fixture paths it names in string
literals) is covered, and names the file, path, chain and key to fix when it is
not. Free *.test.ts files are no longer touchfiles, so editing a free replay
test stops selecting paid evals: 950 entries removed, 653 real closure paths
added. The hand-copied inventories go: periodic-fixture-selection,
fake-impeccable-touchfiles and 45 per-file selection examples. Selection for
the sample edits (plan-eng-review template, claude-pty-runner,
plan-count-fixture, gstack-config) loses no case under either profile.
CONTRIBUTING documents the rule and its lower bound.
2026-09-29 07:00:28 +00:00
garrytan 0489bc9fe8 test: fold per-incident replay series into their detector owners (D)
Twelve detector families move into one owner test each: 73 incident files
become describe blocks in ceo-section-loading-fixture (stale-fill race),
model-overlays, coverage-audit-evidence, autoplan-phase-observer,
native-auto-decide, outside-voice-evidence, eng-first-review,
plan-count-completion, plan-count-file-permission, ceo-mode-option,
plan-scope-selection and plan-count-prerequisite. Each block keeps its original
code and fixture, so every case still runs; only tests asserting the incident
file's own touchfile registration are dropped (41). Touchfile lists that named
an incident now name its owner.
2026-09-29 06:37:21 +00:00
garrytan 6415690a18 test: retire the finding-count cluster and trim its helpers (C)
- C0/C1: the five never-green evals (skill-e2e-autoplan-chain and
  skill-e2e-plan-{ceo,eng,design,devex}-finding-count) failed on harness and
  budget, never on skill behavior; delete them, their touchfile/tier ids,
  AUTOPLAN_CHAIN_BUDGET and the dedicated eighth periodic slice (--slices 7).
- C2: delete the helper groups whose only paid consumers were those files
  (11 modules), trim claude-pty-runner and eng-seeded-coverage to the paid
  closure, and delete the free replay tests whose assertions exercised only
  that dead code (89 files, 135 orphaned fixtures). Blocks that used dead code
  only as input for a live subject keep their assertions: the multiSelect
  default moved to plan-review-decisions, runner PTY tests use inline caller
  policies, and the timer-safe budget checks moved to eng-finding-retry-budget.
- The eight production-touching files stay except ceo-current-decision-record
  (its template read only feeds the retired counter).
- CARVE_GUARDS.autoplan is behavioral 'none'; TODOS records the lost chain
  and per-finding cadence coverage with their re-entry tests.
2026-09-29 06:08:49 +00:00
garrytan 53e7f3212f test: clean up the paid eval lane (B1-B4, B6, B7)
- B1: delete paid files that assert nothing or cannot pass meaningfully:
  skill-llm-eval-spec and skill-e2e-spec-execute (test.todo), gemini-e2e
  (+ gemini-session-runner; no gemini CLI in CI), ship-idempotency (red
  since v1.63), the two opus-4-7 *-sonnet overlay wrappers, conductor-prose
  (+ its source-evaluation replay), codex-e2e-plan-format; drop their keys,
  scripts and census rows.
- B2: skill-llm-eval grades browse/sections/command-list.md with one union
  judge that also carries the baseline score pin; regression-vs-baseline
  deleted (paid run: pass, c4/c4/a4).
- B3: memory-pipeline, ios-qa, ios-qa-swift-build and plan-tune-cathedral
  make no model calls; renamed out of the paid glob so they run on every
  PR. Swift builds need GSTACK_TEST_SWIFT=1; device stub deleted.
- B4: codex-e2e*, outside-voice, aside and ios-device cannot run in the CI
  image; excluded from the weekly lane with a tracked re-entry condition.
- B6: fold opus-47's negative routing controls into skill-routing-e2e
  journey-negatives (paid run: 3/3 unrouted) and delete the file.
- B7: delete the never-green brain-privacy-gate eval; a free
  gstack-skill-start test now proves consent precedes artifacts egress.
2026-09-29 06:08:49 +00:00
garrytan 5d032ef299 test: delete tests of dead eval code (A)
- A1: the retired Eng lexical oracle (evaluateEngSeedCoverage,
  isEngSeedDecisionAUQ), the completion-handoff detector and the retained
  corpus had no paid caller since v1.87.6; delete their 26 replay files,
  ~2.6k helper LOC and fixtures, and the dead blocks in 8 mixed files
  (live hasNativePlanTerminal / batching assertions stay).
- A2: dead viewport approvers in autoplan-artifact-permission and their 11
  replay files + fixtures; recorder/launcher cases stay.
- A3: never-wired oracles and seeders (autoplan-phase-order,
  eng-finding-fixture, ceo-paired-fixture, design-ui-scope,
  plan-skill-completion, pty-current-screen, required-reads,
  transcript-section-logger); plan-seed-submission now decodes through the
  production createPtyScreen; section manifests name their actual guard.
- A4: zero-reference helper exports, plus execGit and invokeAndObserve
  found by the reachability pass.
- 52 fixtures orphaned by the deletions; touchfile and selection-table
  entries for every deleted path.
2026-09-29 05:09:36 +00:00
Garry Tan 080a655791 v1.91.5.0 feat: balanced free-suite shards, test:ubicloud, and faster PR eval lane (#2989)
* v1.91.5.0 feat: balanced free-suite shards, 16-way Linux runs, and bun run test:ubicloud

* chore: regenerate agents digest for v1.91.5.0

* ci: serial flaky retry and flake ledger for the Windows free lane

* ci: cancel superseded eval runs; relax LLM-judge clarity bar to 3
2026-09-28 16:00:12 -07:00
Garry TanandMatt Van Horn d2a0bbcf4c v1.91.4.0 fix: Windows Opera cookie import wave (#2980, refs #2957) (#2987)
* fix: support Windows Opera and Opera GX cookie imports

Fixes #2957

* fix: repair Windows cookie decryption and make Opera import failures actionable

- strip the SHA-256(host_key) prefix Chromium adds to v10 values (DB meta v24+) on Windows
- keep receipts and name `$B handoff` recovery for App-Bound rows in browsers without native extraction
- explain missing browsers, missing profiles and ambiguous profile selection with next steps
- add windowsNative/resolveBrowserInfo, the operagx alias and sorted failure reasons in CLI output
- cover the Node server runtime with a real-DPAPI Windows test

* docs: document Windows Opera cookie import and guard browser lists against drift

* chore: file cookie-import follow-ups from the Opera fix wave review

* test: keep Opera receipt tests independent of the shared key cache

* test: declare generated gstack/llms.txt as a command-reference input for PR selection

* chore: release v1.91.4.0

* test: reconstruct the historical cookie-workflow approval after the Opera BROWSER.md additions

* test: run the real-DPAPI Windows check with the runner's environment

PowerShell launched with a stripped environment took ~18-21s on the Windows
runner (measured on a throwaway diagnostics run), past dpapiDecrypt's 10s
deadline; with the full environment it returns in ~0.3s. Only APPDATA is
redirected to the fixture's Opera root.

---------

Co-authored-by: Matt Van Horn <455140+mvanhorn@users.noreply.github.com>
2026-09-28 13:09:52 -07:00
01593aa67c v1.91.2.0 fix: consolidate gstack reliability wave (#2959)
* fix(memory-ingest): --scan-secrets scans the rendered page and fails closed

--scan-secrets ran gitleaks on the raw transcript .jsonl, then imported a
page rendered from it. gitleaks' assignment rules don't match across a
JSON-escaped quote (KEY=\"v\" on disk), so a secret the rendered page
shows as KEY="v" was imported unflagged. And the gate skipped a file only
on scanner "gitleaks" with findings, so a scan that errored (non-zero
exit, 16MB maxBuffer overflow on a file with many findings, unparseable
report) or could not run (gitleaks missing, slow-probe cooldown) imported
the file unscanned.

Scan the rendered page body, the exact bytes writeStaged() writes, via a
new secretScanText() helper, and skip the file whenever the scan did not
complete. Skipped files stay out of the state file, so the next run
retries them. Reword the helper warnings and setup-gbrain/memory.md,
which described the fail-open as intended.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(test): reconcile Bun failure markers and footer counts

* fix(sync-gbrain): verify source-scoped reads without mutation

* fix(test): recognize grounded TTHW target choices structurally

* fix(aside): make the readiness probe work under zsh and report why it failed

The probe built its deadline into `_T` and expanded it unquoted, so
`$_T aside repl …` only worked in a shell that word-splits. zsh does not: it
looked for a command literally named "gtimeout 30", the probe answered
ASIDE_NOT_RUNNING with Aside installed and ready, and every browsing skill
fell back to the bundled Chromium in silence. zsh is the macOS default and
Aside is macOS-only, so on a stock Mac the probe could never report READY.

The deadline becomes a function, `_gs_d`. It receives the command as "$@",
already split, so sh, bash and zsh all behave the same, and the gtimeout →
timeout → perl alarm chain is unchanged. A 4th arm runs the call unbounded
when none of the three is present, which is what the empty `_T` did before.
Not `eval`: it re-parses the string, so the parens and `;` of the perl arm
become syntax and that arm dies in bash *and* zsh — on a stock Mac, the arm
that actually runs.

On failure the probe now prints the CLI's reason after ASIDE_NOT_RUNNING:,
the shape gstack-render already uses: the first line that starts with a
capital letter, i.e. the CLI's own sentence or Node's `Error:` line below its
loader frame. "Not running" covers states with different fixes — no window
open for the profile, a NODE_OPTIONS preload that kills the CLI — and a bare
verdict sent all of them to "open the Aside app". The BROWSER SETUP prose
quotes that reason before asking the user to open the app.

The text pin asserted the broken invocation verbatim, so it now pins the
function and asserts neither `$_T aside repl` nor an eval form comes back. A
second test executes the rendered probe in sh, bash and zsh on each of the
four deadline arms with stubbed binaries on a narrowed PATH, plus two failing
CLIs: one that prints its own sentence, one that crashes like Node with the
useful line below the frame.

The deadline function costs zero bytes against the lines it replaces; the
reason costs 53 per copy of the probe (44 where the reworded BROWSER SETUP
line gives 9 back). That moves four guards by the measured amount:
plan-devex-review's skeleton cap to 68,550 (measured 68,544), plan-ceo-review's
skeleton cap to 80,150 (measured 80,111) and union ratio to 1.081 (measured
1.0803), and plan-eng-review's union ratio to 1.151 (measured 1.1504).

Fixes #2842, #2941.

* Clarify engineering review startup and decision flow

* Fix Windows readiness fixture PATH and command shim

* fix(test): recognize grounded TTHW target choices structurally

* Clarify engineering review startup and decision flow

* fix(test): restrict QA-only fixture tools to its no-Edit contract

* v1.90.0.0 fix(sync-gbrain): guard readiness verdicts and refresh metadata

* fix(browse): validate canonical upload targets

* fix(gbrain): classify structured PGLite busy response

* fix(browse): preserve native extension runtime APIs

* Fix displayless browser handoff ownership

* Accept unique installed autoplan methodology aliases

* fix(skills): preserve positional literals during installation

* fix(browse): checksum installer contents through stdin

* fix(test): normalize Windows checksum fixture paths

* test: emulate unavailable shasum in Windows checksum fixture

* fix(investigate): preserve owned freeze lifecycle

* fix(review): preserve N+1 retry and Red Team completion

* fix: bound Aside readiness and preserve safe fallback

* test: exercise setup and Chromium on native ARM

* fix: preserve install ownership and ARM browser selection

* Fix gbrain ingest scan boundaries and seed observation

* Refresh managed ship hooks and supervise expanded paid census

* Reject resumed gbrain pages excluded by current policy

* Recover zombie agent locks safely and enable CI Python venv

* Repair paid actor declarations and Aside pitch assertions

* Bump consolidated wave to next free minor release

* Clarify CEO review admin choices and option tradeoffs

* Preserve CEO mode handoff anchors in clarified workflow

* Make Windows portability fixtures use shell-native paths

* Restore ARM Bun alias and clarify ship review gates

* Refresh ship workflow golden snapshots

* Fix Windows DX documentation controls without piped stdin

* Decode Codex child pipes without Bun's encoded-stream stall

* Bound DX pre-review audit before product questions

* Clarify trusted review-start read in paid revalidation

* Bump consolidated wave to next free minor release

* Clarify CEO review admin choices and option tradeoffs

* Preserve CEO mode handoff anchors in clarified workflow

* Make Windows portability fixtures use shell-native paths

* Restore ARM Bun alias and clarify ship review gates

* Refresh ship workflow golden snapshots

* Fix Windows DX documentation controls without piped stdin

* Decode Codex child pipes without Bun's encoded-stream stall

* Bound DX pre-review audit before product questions

* Clarify trusted review-start read in paid revalidation

* Reconcile new main planning flow and paid judge census

* fix: reconcile rebased planning and source-bound validation

* test: pin cookie workflow judge to scored Sonnet model

* fix: keep terminal agent boot out of module imports

* fix: preserve pending-question uncertainty in engineering review

* fix: stabilize Windows reliability-wave fixtures

* fix: clarify design consultation research workflow

* fix: preserve independent design consultation inputs

* fix: resolve design taste scope and browser research guidance

* fix: make consultation opt-in preflight unambiguous

* test: await native Edge owner readiness or terminal result

---------

Co-authored-by: Bruce Krysiak <brucek@alum.mit.edu>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: Antonio Vitalic <antoninte99@gmail.com>
2026-09-26 18:57:53 -04:00
Garry TanandSom Samantray 2a113ae7e6 v1.91.1.0 fix: harden Impeccable plugin discovery (#2978)
* fix(design-detect): find impeccable installed as a Claude Code plugin

The design-detector probe only ever checked <root>/<SKILL_ROOTS>/skills/impeccable/,
never the Claude Code plugin-cache layout
(<root>/.claude/plugins/cache/<marketplace>/<plugin>/<version>/skills/impeccable/).
A plugin-installed impeccable was therefore invisible: IMPECCABLE_SKILL stayed
absent, the launcher was never found (so the NOT_CACHED run hint never fired),
and its sibling engine was never considered.

Add a plugin-cache walk alongside the existing SKILL_ROOTS walk, sharing the
same presence/launcher/repo-local-exclusion/sibling-engine logic via an
extracted checkSkillDir() helper so both paths stay behaviorally identical.

Fixes #2838

* refactor(design-detect): consolidate newestSemverDir onto safeReaddir

Both did the identical try/catch-around-readdirSync; newestSemverDir now
reuses the new safeReaddir helper instead of duplicating it.

* fix: harden Impeccable plugin discovery and regression fixtures

* test: supply eval mode to the integrated detector callback adapter

---------

Co-authored-by: Som Samantray <som.samantray@gmail.com>
2026-09-25 14:32:12 -04:00
Garry Tan 7b534d3e90 v1.90.2.0 perf: halve local free-suite time and preserve coverage (#2972)
* perf: remove repeated test work and preserve AUQ execution budgets

* fix: validate native evaluation fixture evidence at its actual boundaries

* fix: clarify deployment approval and recovery state transitions

* chore: document coverage and release v1.90.2.0

* test: preserve Windows scheduling and native no-change consent

* test: recognize verified reads through fixture symlinks

* test: isolate alias-name installation from runtime assets
2026-09-25 13:27:07 -04:00
Garry Tan a84b0b5b6d v1.90.0.0 feat: make browser cookie imports explicit and safe (#2964)
* fix(browse): prepare reliable cookie import wave for validation

* ci: sequence quality and behavior for validation branch

* fix(browse): isolate Windows qualification and preserve native diagnostics

* test(browse): cover cookie workflow quality and isolate Windows user paths

* test(browse): trace native member startup and initialize fresh folders

* fix(browse): keep Windows member stdin alive through EOF

* fix(browse): latch native timeouts and compare contained Edge startup

* test(browse): verify native version metadata and actual Windows argv

* test(browse): qualify Dia import on isolated macOS CI

* fix(browse): require picker origin for session mutations

* fix(browse): bound credential reads through stream completion

* test(browse): inspect owned Windows process arguments natively

* test(evals): preserve passing coverage during cookie repair reruns

* test(browse): isolate Dia qualification in a fresh macOS account

* test(browse): pass bounded integer timeouts to native Mac probes

* test(browse): distinguish Windows profile initialization from containment

* test(browse): await descendant pipe readiness before parent exit

* test(browse): initialize and restore isolated macOS Keychain state

* test(browse): initialize Windows fixture folders before qualification

* test(ci): pin the same Node runtime across Windows checks

* test(browse): distinguish native macOS browser preflight stages

* test(browse): isolate Windows descendant console lifetime

* test(browse): preserve native receipts and identify fixture lock holders

* test(browse): prepare dependency resolution before native Mac worker startup

* test(ci): include lock and close checks in native diagnostics

* test(browse): preserve native owner probe stages and subprocess deadlines

* fix(browse): classify Chromium profile-in-use exit precisely

* test(browse): retain Mac qualification evidence through cleanup failures

* test(browse): bound Mac fixture paths and retire its owned user domain

* test(browse): accept vanished fixture entries without weakening cleanup

* test(browse): identify probe-created macOS user domains safely

* test(browse): observe Mac user domains without targeting them first

* test(browse): use passive fresh-user ownership throughout Mac qualification

* test(browse): distinguish profile and registered-home Keychain lookups

* test(browse): qualify Dia under one registered account home

* test(browse): identify Dia startup and owned process-group failures

* test(browse): classify bounded Dia startup diagnostics without leaking output

* fix(test): preserve native Mac sandboxing and reap owned browser children

* fix(browse): preserve Chromium sandboxing for native profile imports

* test(browse): inspect signed Mach-O architecture without launching Xcode tools

* test(browse): sample pending Dia startup and reap on all cleanup paths

* test(browse): compare protected Dia launches in fresh Bun and Node accounts

* test(browse): inspect isolated Mac GUI readiness without browser access

* v1.90.0.0 fix: bind cookie picker actions to their document

* test: validate cookie guards and fit nested launch fixtures

* ci: configure the bundled Chromium sandbox helper

* fix(browse): classify Playwright authentication timeouts

* test: retain bounded Windows lifecycle diagnostics

* test(cso): reuse bounded NTFS precision candidates

* test(review): handle explicit preservation choices safely

* test(browse): remove owned fixture directories with explicit primitives

* test(review): distinguish descriptive reuse from edit commitments

* test: admit only the approved unscored cookie workflow refusal

* test: keep the Office Hours judge mock export-complete

* fix: keep dependency-free CI planners independent of the model SDK

* test: observe the exact holder after a native fixture unlink failure

* fix: start seeded PTY observations at owned readiness

* test: acquire identity-bound Windows deletion admission before profile resets

* test: preserve qualified Git index bits without authorizing mutations
2026-09-25 12:06:45 -04:00