Commit Graph
100 Commits
Author SHA1 Message Date
garrytan 7eaefdec59 fix(plan-ceo-review): name the mode preference command and the exact handoff line
auto-decide-preserved at 6fcb0981: the model never ran the preference check,
read 'check ... through the preamble' as already done, auto-selected 'per your
preference setting', and wrote 'Selected mode: HOLD SCOPE, auto-decided from
your tuned preference' instead of the AUTO_DECIDE handoff line. At 9a7a7e54 it
ran the check but wrote 'Decision: HOLD SCOPE is the review mode for ...'.
Neither matched the handoff template the observer recognizes. Name
gstack-question-preference --check at the point of use and say the handoff
begins with the exact matching line. Collapse the audit block's comment
padding to stay within the unchanged 80150-byte skeleton cap.
2026-09-30 23:23:18 +00:00
garrytan 6fcc30ddc1 fix(office-hours): a forcing question's recommendation takes the position the founder's words support
auq-matrix office-hours asked D1 Demand as options about the founder's own
evidence and, with no rule for that shape, recommended 'answer whichever is
TRUE — A is marked recommended only because it is the strongest position'
(substance 2). Say what such a recommendation is: the option the founder's own
words support, why it matters for the next step, and what would change it.
2026-09-30 23:23:18 +00:00
garrytan 09b686d4ec test(sync-gbrain-readiness): a negation earlier in the claim clause is not a search/write readiness claim
The census unknown actor wrote 'nothing about read, search, or write capability
is confirmed either way' after a YELLOW/WARN verdict. The claim window started
at 'write', so the leading 'nothing' was outside it. Check the clause subject for
nothing/neither/none/no; keep the original in-claim negations. Replay of the
captured output passes; positive controls still flag an unnegated claim.
2026-09-30 23:23:18 +00:00
garrytan 6fcb098129 fix(plan-eng-review): show the accepted dedicated read form for coverage-diagram sources
CI plan-eng-coverage-audit mixed package/config and git diff into the source
read; the review variant, whose prompt shows the && display form, does not.
The plan trace step now shows it too, within the unchanged size cap.
2026-09-30 22:42:29 +00:00
garrytan d0c5357763 test(ship-docsync): seeded attempt 1 counts toward the limit; transport counts ignore calls that never reached the state file
- CI launch-failure retried after the seeded attempt 1 as if that attempt were
  the fixture's; the seeded prompt now says attempt 1 is this invocation's and
  a further attempt needs what Blocked recovery requires.
- A late-result run typo'd the state path once (ENOENT, the actor never ran),
  then repeated the call correctly; the per-action count compared both calls
  with one actor event. Only calls naming the real state file are counted.
2026-09-30 22:37:12 +00:00
garrytan 4dfed0b783 fix(qa-only,qa): name the exploratory read point and finalization order; judge qa with its browser assets
qa-only judges cited 'next section' pointing at the wrong heading, an exploratory
trigger that contradicted its read point, clock ownership in mixed runs and the
unstated order of exploratory section 4 vs reporting. The qa judge penalized the
absent qa-report-template and issue-taxonomy that qa-patterns loads; with them
in, it found issue-taxonomy's dangling 'rule 13' (the consent rule is browser rule 3).
2026-09-30 22:07:25 +00:00
garrytan 23e3636e4c test(shared-libs): the plan actor reads scheduler parity and unchanged-scope lists
Census 36776104571's question preserved the contract ('behaving exactly like the
scheduler', 'scheduler parity holds by construction') and excluded work with
'Existing copies and helper hardening stay unchanged'. Accept exactly/parity as
preservation (negated forms refuse) and a bare noun list that stays unchanged as
an exclusion for the expansion scan only; verb-led clauses still refuse.
2026-09-30 22:07:25 +00:00
garrytan 746f9b9cf1 fix(review): design-lite rows keep the detector's [rule-id]; the e2e detector rows point at the diff
The output template had no rule-id slot, so rows merged with checklist items
dropped the detector id (census t2, local t1). Rows now carry [rule-id]. The
fake engine's sample rows named a foreign fixture path at line 0; the e2e remaps
them to landing.html/styles.css so trials stop spending turns reconciling it.
2026-09-30 22:07:25 +00:00
garrytan a487a09bf1 test(office-hours): a fork whose outer options are the seeded shapes is the Phase 4 question
Census trials 1-2 captured complete Phase 4 forks (A) Server-side B) Client-side
C) Hybrid, recommendation with because) whose prose used none of the vocabulary
words. Accept two seeded shapes as outer options as Phase 4 specificity; the
earlier-phase, nested, fenced and single-shape controls still fail.
2026-09-30 22:07:25 +00:00
garrytan 1751985a21 test(coverage-audit): a fenced plain-word caption in a successful && read chain is display only
Census 36776104571 plan-eng capture read both owned files with cat -n in one
successful && chain; the caption 'echo "=== git diff main --stat ==="' fell
outside the two-token caption grammar, so both reads lost credit. Accept a fenced
caption of plain words; unfenced command strings, expansions, redirection,
-e escapes and ; / || tails stay rejected.
2026-09-30 22:07:25 +00:00
garrytan 190e650d97 test(autoplan-dual-voice): the fixture declares that delivered bash blocks run alone, diagnostics separately 2026-09-30 22:02:15 +00:00
garrytan 2734e2035c test(plan-ceo floor): scope preservation approves no premise, approach or remedy 2026-09-30 22:02:15 +00:00
garrytan 0b62c51ddb test(auto-decide): ask for the selected mode in the skill's mode handoff line, not a separate public decision 2026-09-30 22:02:15 +00:00
garrytan 005352ddc7 test(section-loading): credit a Bash print that contains every line of the carved section 2026-09-30 22:02:15 +00:00
garrytan 5530a7ad1e test(autoplan-dual-voice): unwrap the framed native report before Claude Code 2.1.284's agentId/usage trailer 2026-09-30 22:02:15 +00:00
garrytan 86191bbc23 fix(qa-evidence): accept a single limits string; test(qa-callers): read the handoff first when a probe snapshot changes
CI late-input spent a turn rewriting limits as an array after materialize
refused a string, and a ten-read sweep hunting for the changed input before it
read reports/HANDOFF.md, then timed out at 300 s.
2026-09-30 21:24:01 +00:00
garrytan 73e3bae953 test(strict-output): give the spool-prefix child time to finish before the pending stream times out
windows-free-tests failed on 9a7a7e54: the 150 ms shared deadline raced Bun
startup on Windows, so the child was killed mid-write and the spool held a
partial payload. Only the never-released extra stream should time out; the
child now has 3 s.
2026-09-30 21:07:08 +00:00
garrytan 9a7a7e5499 docs(changelog): v1.91.10.0 records the flake census and its repairs 2026-09-30 20:56:12 +00:00
garrytan 2bd4651cde test(review-army): share the recorded Step 4.5 staging with consensus and supply its Red Team
review-army-consensus (periodic) timed out in 2 of 13 census sessions; passing
runs took 213-297 s of 300. Like N+1 it spent ~30-50 s reading the whole
extracted SKILL, checklist and every specialist file, sometimes dispatched an
unrequested Maintainability specialist, then ran a Red Team (60-70 s) and a
second merge before writing a 9-15 KB report.

The N+1 staging and scope text move into stageReviewArmySession /
reviewArmyScope / reviewArmyChecklists (the N+1 prompt renders byte-identical).
Consensus now records its detect-scope, stats, learnings and diff, stages the
Review Army section with the security and testing checklists, forces
--security --testing, and caps the report like N+1. Its Red Team is outside
the multi-specialist contract, so the fixture supplies a labeled synthetic
NO FINDINGS result instead of a dispatch. The existing SQL-finding and
browser-error assertions are unchanged; the lifecycle adapter's spawnSync now
returns the git output the staging reads.
2026-09-30 20:54:35 +00:00
garrytan 17ee2e5428 test(review-army): record N+1's pre-dispatch stages and scope the session to Step 4.5
review-army-perf-n-plus-one timed out in 7 of 13 CI runs on this branch (passing
245-280 s of 300). Each session spent ~95 s on setup (the full extracted SKILL,
checklist, section greps, exploratory.md, diff-scope/stats/learnings, tooling
checks), ran Step 4's core pass, a search-before-recommending WebSearch, and
wrote a 10-16 KB report (~100 s after the Red Team returned).

The fixture now stages only review/sections/review-army.md plus the performance
and red-team checklists, and hands the session the recorded detect-scope,
specialist-stats and learnings outputs and the diff. The caller passes
--performance (every CI parent already treated the prompt as that force flag
against the <50-line skip), declares the core pass, QA, adversarial review, web
research, Fix-First and persistence out of scope, and caps the report at the
selection line, the SPECIALIST REVIEW block and the Red Team result (30 lines).
The Performance specialist and the conditional Red Team are still real
foreground subagents, and the report still has to surface the N+1.

New assertion: a foreground Performance specialist dispatch precedes the Red
Team dispatch. Free controls omit the Performance dispatch or background it, and
both fail; the budget lifecycle adapter supplies the current result shape.
Touchfiles now include the .rb fixture the case reads.
2026-09-30 20:54:35 +00:00
garrytan 90203acda3 feat(qa-evidence): captures list the caller's declared-but-unrun required probes
GSTACK_QA_REQUIRED_PROBES (a JSON array of native child commands) makes every
capture print requiredRemaining; it never judges pass or fail. The functional
eval passes the webhook list from QA_WEBHOOK_REQUIRED_SCENARIOS, which the
verdict now reads too, so the nudge and the verdict share one source (agreed
with #3002's owner). CI webhook-report kept stopping with scenarios unrun.
2026-09-30 20:42:29 +00:00
garrytan a615303213 test(qa-functional): point the fixture at the helper's --help instead of its source
A CI webhook-fix run spent three turns reading lib/qa-evidence.ts to learn the
interface and timed out just before materialize (agreed with #3002's owner).
2026-09-30 20:34:37 +00:00
garrytan 1421c641e8 feat(qa-evidence): refuse evidence observed on an older input snapshot than the latest capture
When native probe output declares a top-level input snapshot, materialize
compares each evidence row with the latest capture's snapshot and refuses
stale rows unless they are classified superseded, naming the captures to
rerun. ship-exploratory-late-input kept reporting a pre-change adverse probe
green after the input changed.
2026-09-30 20:04:19 +00:00
garrytan 00dacee8bd test(qa-observer): fix mode treats atomic temps of authorized src/test writes as transient
CI webhook-fix failed with 'Could not watch test/worker.regression-1.test.ts.tmp...':
Claude Code's Write renamed its temp before the per-file watch was added. The
functional eval now tells the observer its mode, and a temp whose target that
mode may write is observed through its directory watch. Report-only mode and
undeclared paths keep failing closed.
2026-09-30 19:49:54 +00:00
garrytan b541f28ddb fix(qa-evidence): every complete capture needs an evidence row; test(tpa): accept the hyphenated app-specific-password spelling
- materialize refuses when a complete capture has no evidence row and is not
  named in limits (CI cli-report omitted capture 004), naming the missing IDs.
- tpa-apple-ban's detector required 'app-specific password' with a space; the
  CI answer said 'app-specific-password path' and was otherwise correct.
2026-09-30 19:48:40 +00:00
garrytan f02636f05e fix(review): define what a Step 5c Skip option says
Step 5c named "B) Skip" without saying what its description may claim. Two
CI captures (path-eligibility on 131d43be, index-flags on 4643cb85) offered a
Skip whose description added effects beyond declining: "The extraction can be
applied in a later editing review pass" and "replacing the invalidated prior
Skip". Those read as change commitments, so the no-change actor refused both.
Step 5c now says to describe Skip only as no code/index change with the Skip
recorded; adjacent lines are compacted so the review parity caps hold
unchanged. Both exact packets are kept as a free regression: still refused,
and accepted once Skip follows the rule. The actor's classifier is unchanged.
2026-09-30 19:36:02 +00:00
garrytan cc044e5f89 test(shared-libs): trim the resumed review replays' setup and report
Every sibling review session (revalidation, path-eligibility, index-flags,
prior-coverage) loaded qa/sections/exploratory.md and often scope.md although
its QA and native adversarial results are supplied synthetic inputs, then spent
a second request on shared-code-reuse.md and base metadata. The resumed scope
now states that the supplied results replace Step 4's QA method loading; the
revalidation contract names one first response (workflow, checklist, finding,
prerequisites, shared-code-reuse.md, base metadata) and caps the summary at
twelve lines. Receipt order, direct source reads, the checker, the question and
final persistence are unchanged.
2026-09-30 19:36:02 +00:00
garrytan ea7fbbcb7c test(shared-libs): seed the lifecycle replay's first Step 3 pass instead of replaying it
shared-libs-review-lifecycle ran ~88% of its 300 s session budget (12-run
census median 265 s, 4/24 sessions timed out). The fixture now executes pass 1's
Step 3 once with the real logger and Git: a real unused REVIEW_START, then the
diff, inventories, attributes/config/index flags, gstack-review-read output and
every file's bytes and sha256, saved to one observation. The model resumes at
Step 4 with an exact four-file first read, the observation named as the
authoritative pass-1 repository read, one post-fix verification, an explicit
pass-2 read list and a twelve-line summary. Pass 2 still runs its own --start,
diff, reads, fingerprint and stage actor before --finish.

The actor scope now states that a current settled final-pass actor result
supplies the replaced QA/adversarial prerequisites and that the no-credit
disclosure is a reporting label: one r1 session persisted completed:false
from that ambiguity.

New assertions: the final binding never uses the seeded token's start or tree,
and the observation was read; free controls finish the seeded token (binding
changed) and omit the observation read, and both fail.
2026-09-30 19:36:02 +00:00
garrytan ac177337a1 fix(qa): after an input change, a probe is affected unless shown otherwise
CI late-input run finished in time but revalidated only the happy probe after
the locale input changed and reported the stale adverse probe green. The
revalidation step now treats any probe not shown to be unaffected as affected.
2026-09-30 19:29:54 +00:00
garrytan ee929ff710 feat(qa): helpers answer --help, and the QA eval interfaces declare it
Approved by Garry: asking gstack-qa-evidence or gstack-qa-deadline for usage
is read-only, so both helpers print usage and exit 0 on --help (the evidence
usage now names the annotation shape), and the functional and caller command
allowlists accept exactly 'bun <path>/bin/gstack-qa-{evidence,deadline} --help'.
Two CI runs failed only on that call.
2026-09-30 19:14:59 +00:00
garrytan 4643cb8550 test(qa-deadline): never attach a reader to the full-pipe fixture's stdout
The full-pipe receipt test attached a 'data' listener (flowing mode) and then
paused; on CI the reader could drain the 2 MB write before the pause, so the
receipt write never blocked and the helper exited 0 in ~126 ms. The stdout pipe
now stays unread until the assertion, which is what the test means to model.
2026-09-30 18:57:26 +00:00
garrytan 07b21ea133 fix(deslop-shared-libs): probe the audited repository with -C <repo>
A CI run probed safe-git from the session directory above the target repo, so
the capability probe never touched the repository and the run fell back to the
API without a local attempt. The probe (and any call from elsewhere) now names
the audited repository.
2026-09-30 18:52:14 +00:00
garrytan 271c14078b test(qa-functional): fix mode requires only the happy scenario from the model (carried byte-identical from #3002 183b01f4..3e6074b4)
verifyQANativeRegression already reruns all eight webhook scenarios on the
repaired source, so the model-side eight-scenario requirement in fix mode
duplicated harness coverage and pushed qa-functional-webhook-fix past its
budget. qa-only still requires every scenario.
2026-09-30 18:34:36 +00:00
garrytan 1643cd94de fix(qa-evidence,observer): reject placeholder metadata and replay-only learning; declare the docs atomic-write target
- materialize measures revision, runtime and cwd itself and rejects supplied
  values that differ (CI run wrote revision "HEAD" and runtime "bun"), and
  refuses learning checkpoints that replay the same probe, naming the fix.
- The docs write observer treats Claude Code's atomic temp for the authorized
  doc target as transient, so a temp renamed before its per-file watch no
  longer marks the observation incomplete (ship-docsync-completion flake).
  Per-file monitoring outside declared targets stays fail-closed.
2026-09-30 18:24:46 +00:00
garrytan 131d43be0a test(shared-libs): tee to a discard device is not a file write
Paid shared-libs-opportunity-judgment t1 on 1213b01 failed read-only on
'... | tee /dev/null | sha256sum'. The detector flagged any tee operand while
the same devices are allowed for redirection. tee now fails only when an
operand is a real file; tee to a file, -a file and -- -a stay violations.
2026-09-30 17:59:24 +00:00
garrytan a367a1f265 feat(deslop-shared-libs): route every Git read through bin/gstack-safe-git
The skill made the model retype a long safe-Git prefix on each call and a
dropped flag failed shared-libs-read-only. bin/gstack-safe-git applies the
fixed env + flag prefix, adds --no-ext-diff --no-textconv to log/show/diff,
allows diff only between two explicit object IDs and ls-files only in the
NUL-delimited overlay form, and refuses every other shape with one line
naming the allowed forms. The template now points at the installed helper
(host global runtime via {{SAFE_GIT}}) and drops the prose it enforces.

Fixtures resolve the helper to this checkout, the git shim records the safety
environment, and isGuardedGitRequest requires the complete prefix (env
included) for every repository read.
2026-09-30 17:59:23 +00:00
garrytan 157a5ff520 test(qa-callers): hand the caller phase its invocation-start observations and review token; fix(next-version): fetch without auto maintenance
- Every caller case receives the diff, status, log, untracked list, HEAD and an
  already-captured review start token, so the phase spends its budget on the
  contract under test instead of re-running setup reads.
- gstack-next-version's fetches pass --no-auto-maintenance. On git 2.55 a
  completed fetch forks detached maintenance in the caller's repository; the
  free suite's live smoke test ran it inside the CI checkout, and every
  shard-12 pre-push hook hang so far followed a completed smoke fetch.
2026-09-30 17:59:23 +00:00
garrytan a7872aaae0 feat(qa-evidence): enforce the checkpoint sequence and fill report bookkeeping in code
- capture refuses to run another probe until a checkpoint anchored on the
  latest complete capture names this capture as its next command, and every
  complete capture prints that requirement.
- materialize fills revision, runtime, cwd and learning (checkpoints whose next
  native command differs) when omitted and prints the reportLinks the report
  must include; the QA section shrinks accordingly.
2026-09-30 17:40:33 +00:00
garrytan 6ce10ff7c4 test(ship-docsync): trim the seeded parent's measured model time
Measured on the seeded runs: one read the 78 KB ship/SKILL.md, the post-child
freshness comparison spent 18-32 s of thinking over full inspect contents, and
the final response restated the report (~1.1 KB). Say the phase excerpt stands
in for ship/SKILL.md, compare hashes first and read content only for changed
paths, and end with one status line.
2026-09-30 16:30:43 +00:00
garrytan e6ac813ddb test(ship-docsync): name the seeded read list and cap journal/report length
The first seeded stale-before run spent calls locating documentation.md (two ls
sweeps), reading through cat and re-Reading the record before Edit, and ~40 s
composing 1.5-2.2 KB entries and report. Name every seeded read path, ask for
native Read, and bound entry/report length.
2026-09-30 16:30:43 +00:00
garrytan fb52689822 test(ship-docsync): seed fault cases at their gate instead of replaying attempt 1
The post-dispatch fault cases (missing-marker, launch-failure, timeout-unsettled,
late-result, stale-before, stale-after, recovery) now start from a fixture-owned
attempt 1: the real actor prepares and dispatches it, its verbatim output is saved
once, and the invocation journal carries its pre-dispatch entry with the child
asset hashes. The model resumes at Parent processing with a trimmed read list,
inspect named as the authoritative repository observation, and recovery's
intermediate checkpoint folded into the next attempt's pre-dispatch entry.
Assertions count only parent-issued transport events and require a read of the
saved attempt-1 output; missing-asset and the legacy failure case keep the full
model-driven first attempt, and their prompts are byte-identical.
2026-09-30 16:30:43 +00:00
garrytan dd5708e6e8 test(qa-callers): deterministic child transport, completion-time handoff reads, compact phase report
The exploratory caller cases exist to prove the caller starts and bounds
exploratory QA. Their native adversarial reviewer (review) and plan audit
(ship plan-checks) now come from recorded child outputs instead of a live
subagent, handoff freshness reads are required before completion records
rather than every bookkeeping log, and the phase report is compact. Measured:
194-257 s per case against 208-284 s before, no subagent calls.
2026-09-30 16:19:39 +00:00
garrytan 3e5ec6df17 fix(evals): count timeout turns only from object transcript events 2026-09-30 16:04:19 +00:00
garrytan 79092b22ea fix(evals): cut path variance at its measured sources
- gstack-qa-evidence capture prints startedAt/completedAt/durationMs and, for
  --deadline captures, remainingMs; the functional report takes durations from
  them. The section clock notice asks for one clock read up front instead of one
  after every checkpoint (QA runs spent 7-14% of tool calls on date -u).
- ship plan-completion: skip the audit dispatch when discovery already found no
  plan (the dispatch-vs-skip conflict produced an optional 60-100 s subagent).
- materialize/checkpoint validation errors state the expected schema, so a
  rejected annotations file is fixable in one call instead of blocking the phase.
- session-runner counts turns from the transcript when a run times out, so
  timeouts stop reporting 'turn 0'.
2026-09-30 15:59:58 +00:00
garrytan a8a32e3216 fix(plan-eng-review): keep the headless-rule contract phrases adjacent 2026-09-30 14:34:59 +00:00
garrytan 0d22bae39a fix(plan-eng-review,review): a disallowed question tool is not headless; report kept tests only when some were skipped 2026-09-30 14:33:41 +00:00
garrytan f4fa38ac68 fix(qa): define evidence.json where it is built, point the preparation gate at the next section, name measured command durations in the report template
Recurring qa/qa-only workflow-judge complaints in CI (clarity/actionability 3.33).
2026-09-30 13:59:06 +00:00
garrytan 53c5b7505d fix(qa): number the qa value-bar questions from 1 and say reproduced bugs already answer the first two 2026-09-30 13:31:29 +00:00
garrytan 2b76634831 fix(qa): the caller STOP line says to await the method Reads before any probe
ship-exploratory-plan-checks: the model read exploratory.md and sent a capture
in the same response, before seeing the section's own await rule.
2026-09-30 13:19:52 +00:00
garrytan 6bdb0bf1be test(office-hours-attempt): the fake judge SDK response carries stop_reason like the real API (structured judge requires end_turn) 2026-09-30 13:04:59 +00:00
garrytan a0a6df2a8e test(qa-callers): disable git auto maintenance in the fixture repo (same guard as shared-libs; from #3002) 2026-09-30 12:59:59 +00:00
garrytan 6746ffbc4b docs(todos): record the pre-push hook shard-order hang 2026-09-30 12:58:44 +00:00
garrytan e3217b6d09 test: fold the design-consultation completion replay into carve-section-sharding (test-of-test ratchet) 2026-09-30 12:58:14 +00:00
garrytan a27365bffe test: PTY harness handles clipped reviews and bundled setup tabs; AUQ judge uses structured output; design-consultation carve declines optional outside voices
- ceo mode routing: a Submit review taller than the viewport, a setup tab
  bundled after the mode tab, and a clip through the mode question each hung
  or misread the run; the native answer is still verified after Submit.
- judgeRecommendation requests a 1-5 enum schema; a malformed Haiku reply had
  scored substance 0 for a 4/5 brief. Judge failures now propagate.
- carve section-loading for design-consultation declines the optional outside
  voices (a supported path) and treats DESIGN.md as the report; timeout unchanged.
The Step 0E handoff defect is not fixed (0/15 samples across four wordings,
none shipped) and is filed in TODOS.
2026-09-30 12:52:56 +00:00
garrytan dfe5e733fb test: one owner per case id, a structural devex 0B setup rule, and correct design/gbrain actors
- plan-design-review-plan-mode was registered by two files; the PTY smoke is
  now plan-design-review-plan-mode-smoke, and a registry test requires one
  owner per case in case-sharded files.
- plan-devex-finding-floor: the template's 0B narrative-confirmation question
  is classified as setup structurally instead of timing out a Haiku assessor.
- setup-gbrain-remote: the actor accepted 'skip' on the MCP-registration
  question the test asserts; it now accepts that question and declines others.
- design-review-plugin-handoff: the fake engine cited a file absent from the
  fixture repo and index.html linked a missing styles.css.
Captured-question regressions with negative controls; each case passed a
focused paid run.
2026-09-30 12:24:05 +00:00
garrytan 8cf87d4729 ci(image): pin Claude Code 2.1.284, the version users run
Request-body capture shows both 2.1.251 and 2.1.284 send effort "high" to
claude-fable-5-1; 2.1.284 adds the model's own profile. The slower 2.1.284
census was mostly API latency: its SDK-only judges were 25% slower too. Nine
previously slow cases pass on 2.1.284 within unchanged budgets.
2026-09-30 12:14:37 +00:00
garrytan 77cce3bec4 fix(ship,qa,document-release): repair proof-run regressions and fixture gaps
- ship-docsync-completion: yesterday's audit-scope result dropped the section's
  status, so /ship spliced one in; the section now opens with **Status:**.
- ship-docsync-missing-asset: a missing section or old Ship-owned mode blocks
  before launch.
- ship-docsync-late-result: the invocation record says prepare already saves
  the candidate selection (no extra Read; budget unchanged).
- qa exploratory: await the method Reads before the first probe.
- qa-callers fixture: quote the real review-log record template; allow the
  git log command plan-completion prescribes.
- qa functional observer: a receipt caught mid-link(2) is checked at stop
  instead of failing with ENOENT (reproduced from CI).
Each repaired case passed a focused paid run.
2026-09-30 12:11:34 +00:00
garrytan 4a87fa9d59 fix(ship): always run the design-lite detector probe; test(shared-libs): credit a failed first file view and deferred-reuse Skip wording
- /ship design-lite: the probe is mandatory and any non-ready first line is
  stated, matching /review (5 of 6 captured /review trials had skipped it).
- shared-libs-pr-coverage: the first PR 42 page-1 read printed only a jq error,
  so the one refetch is a legitimate recovery, charged to the same budget.
- shared-libs-review-prior-coverage: the Skip option said a future review can
  "reuse it once snapshot coverage holds"; a conditional tail on the recorded
  decision is not product work. Captured-text regressions and negative controls.
2026-09-30 11:35:24 +00:00
garrytan 0748063aba docs(changelog): proof-run product fixes 2026-09-29 22:49:27 +00:00
garrytan a06d22e52a fix(review): pass Review Army checklists by path, run research alongside dispatch, always probe the design detector; state review-log invocation and statuses in the caller fixture
- review-army-perf-n-plus-one: the parent copied full checklists into agent
  prompts and ran web research before dispatch (290 s on a 12-line diff); 212 s now.
- review-design-lite: 5 of 6 captured trials reported the detector absent
  without probing; the probe is mandatory and its first line is reported, and
  the contract credits only fake-engine rule ids the checklist never names.
- review-exploratory-small-cli: the fixture never gave review-log's direct
  invocation or status vocabulary; the model ran it through bun and wrote
  status "blocked". The prompt states both and the validator rejects
  out-of-vocabulary review statuses.
Each case passed a focused paid run after repair.
2026-09-29 22:49:12 +00:00
garrytan a18e6cf655 test: accept 'review mode = X' auto-decide declarations and parenthetical scope exclusions in the shared-libs actor
auto-decide-preserved: the product auto-decided HOLD SCOPE and said
"Decision: review mode = HOLD SCOPE"; the grammar knew only "is" and ":".
shared-libs-plan-callers: the recommended option said "(no hardening)" and the
actor read "hardening" as an expansion. Both replay the captured text, keep
negative controls, and passed focused paid runs.
2026-09-29 22:46:30 +00:00
garrytan 12ab3b3c54 test(design): revert the three-Edit plan-mode flow
A focused paid run still timed out at 300 s: the first three passes alone took
150 s of thinking. The case stays a named timeout red rather than cutting review depth.
2026-09-29 22:40:01 +00:00
garrytan a111225e78 fix(evals): repair proof-run reds in design-consultation, document-release, design and QA fixtures
- design-consultation Phase 1 asks one brief that confirms context and decides
  research; the confirm-only first question scored substance 2.
- document-release defines ship-owned inputs, exact steps and the JSON result,
  and drops stale spawned-from-/ship text (judge actionability 3.67 -> 4/4/4).
- plan-design-with-ui accepts the Step 0D focus menu the same way the shared
  picker does ("focus on specific ones?").
- plan-design-review plan-mode saves in three Edits instead of one final Write.
- QA functional annotations ask for the full 40-character revision.
- Outside-disabled attribution judges quoted prior-record data by its exact
  timestamp or a dated, pre-existing-record sentence; four captured phrasings
  replay clean and current claims still fail.
- --case can select autoplan-dual-voice by its literal test name.
2026-09-29 22:36:28 +00:00
garrytan aba80c8fb6 fix(eval-pass-rates): match trial-outcome files by basename so Windows backslash paths are read 2026-09-29 22:16:52 +00:00
garrytan bfabda7419 fix(plan-ceo-review): tighten expansion pacing wording to fit the skeleton cap after the main merge
The merged skeleton measured 80,166 bytes against its unchanged 80,150 cap.
Same instructions: ask separately for each addition, in turn, with no pacing
menu; lead each proposal with the felt experience, then shape, effort and impact.
2026-09-29 22:07:18 +00:00
garrytan 18ea34b949 ci(evals): name the PR-comment loop's unused fields so shellcheck passes (SC2034) 2026-09-29 22:02:03 +00:00
garrytan 58b5e3f977 Merge origin/main (v1.91.9.0, #2998) into capy/audit-fix-wave; release as v1.91.10.0
Main shipped v1.91.9.0, so this wave becomes v1.91.10.0. Conflicts kept this
branch's planner-derived assertions. Main's new test-value eval gets rule
kinds for its three gate cases and its measured 332 s duration from #2998's
CI; census counts and the PR fallback floor follow the new file. The packing
test now weighs each tier's recorded durations the way the planner does.
2026-09-29 22:00:14 +00:00
garrytan bfad7fd37d docs: final census numbers in the v1.91.9.0 entry; file the paid-eval follow-ups 2026-09-29 21:50:57 +00:00
garrytan 8dd3c44fc6 fix(qa): checkpoint receipts print the report link for their exploration file
qa-functional-webhook-report failed in two of three censuses because the
report linked .qa-evidence/NNN capture folders as "checkpoints" and never
linked exploration-NNN.json. The checkpoint receipt now prints
link: [checkpoint NNN](exploration-NNN.json), and the functional report
template says capture folders are not checkpoints.
2026-09-29 21:49:46 +00:00
Garry Tan 96764e80a6 v1.91.9.0 feat: test value bar in plan-eng-review, review, qa and ship, plus /test-audit (#2998) 2026-09-29 14:35:00 -07:00
garrytan 958d1ceba6 test: supply holdDeferKeepIndex to the CEO routing mocks and follow split-overflow into the marathon lane
The registered-callback fixtures mock ceo-mode-option and lacked the new
export; the split fixtures asserted the periodic tier; the registered-budget
check looked for split-overflow only in the periodic manifest.
2026-09-29 21:33:13 +00:00
garrytan bf4667146c test: attribute quoted prior-record field lists, state the judge reason bound in its schema, move split-overflow to marathon
Census 36629958451 reds:
- outside-plan-disabled-no-fallback: the model quoted the pre-existing record
  as a parenthesized field list with its exact timestamp; attribution now
  requires that exact timestamp and the record's own field values.
- plan-devex-peer-comparison-classification: the judge correctly returned
  missing but wrote a 1069-character reason, voiding the judgment; structured
  outputs cannot enforce maxLength, so the bound is stated on the field.
- plan-ceo-split-overflow ran 504-1188 s as one PTY flow and set the
  periodic lane's wall clock; it now runs weekly in the marathon lane.
2026-09-29 21:26:37 +00:00
garrytan f63e1fb7cb ci(image): keep Claude Code 2.1.251; test(ceo-mode-routing): keep HOLD's own deferrals in scope before assessing its rigor decision
2.1.284 enables per-turn effort for claude-fable-5-1: in gate census
36626737820, 66 of 84 sessions ran longer than on 2.1.251 (+20% session time,
+32% thinking tokens) and 11 cases timed out on unchanged budgets.

HOLD SCOPE's 0G step asks its own defer/keep menu; the actor answered it
Defer and the assessment then judged that scope question as the rigor
decision. The actor now answers that menu Keep and assesses the next one.
2026-09-29 20:57:04 +00:00
garrytan d05d161713 test(autoplan-dual-voice): unwrap Claude Code 2.1.284 subagent hand-back frames; accept read-only probe diagnostics; record before asserting
Census run 36626737820: the native CEO report arrived framed and indented, so
its INPUT line never matched, and the model's exact probe plus two variable
echoes was not canonical. A column-zero line inside a frame, command
substitution, backticks, redirects, assignments, CODEX_MODE echoes and output
line-count mismatches stay rejected. The failure now records before asserting.
2026-09-29 20:48:37 +00:00
garrytan 3fc05932b7 test: settle the post-response composer before seeding; give the TPA recorder adapter its infra helper
submitPlanSeed accepted a stale empty composer when the transcript recorded
end_turn before the CLI repainted (late-repaint-typed-current fails 5/5 on the
old helper, passes 5/5 now). The TPA recording fixture extracted recordE2E
without isPreTurnInfraFailure, so every failed case threw before recording.
2026-09-29 20:43:00 +00:00
garrytan 80c92c86ed chore(release): v1.91.9.0 2026-09-29 20:29:34 +00:00
garrytan 581e603313 fix(evals): --case list mode and name precheck; case-shard qa-callers; refresh batching and design-with-ui seeds 2026-09-29 20:28:31 +00:00
garrytan 2dae4bf944 Merge remote-tracking branch 'origin/capy/rel-c' into capy/audit-fix-wave 2026-09-29 20:23:51 +00:00
garrytan 475667ff5d test(eng-batching): read the report target as a field, not a spelling
The next targeted rerun (Claude Code 2.1.284) again asked eleven separate
native questions and again counted zero: its briefs named no plan and its
report declared '- **Review target (fixed):** `/abs/PLAN.md`' under
'# Eng Review — PLAN.md: <plan>'. An unsourced brief now inherits the one
current target field that names a PLAN.md file, whatever its list or
emphasis markup; its ledger record still supplies the cited finding and
must reproduce the brief exactly. A brief that names its plan must still
match the report title. Replays of all three captures count 9, 9 and 3;
controls reject a foreign, duplicate or missing target and an archived
title.
2026-09-29 20:12:02 +00:00
garrytan dc640acd8a test: pin every-record outcome counts and the twelve doc-sync callbacks 2026-09-29 19:55:37 +00:00
garrytan e758fbe96d test(evals): record a pre-turn API or CLI failure as infra
recordE2E sets failure_class 'infra' on a failed session whose runner
reports error_api, timeout_startup, error_output_stream or a non-zero CLI
exit with zero turns and no assistant event. A model refusal, a timeout
after model work, max turns, or an explicit caller pass/class keeps its
ordinary classification.
2026-09-29 19:51:35 +00:00
garrytan f4ab5ee75f test(judges): sample the recommendation rubric as a panel; never re-ask armJudge
llm-judge-recommendation is a judge case: each fixture now draws a
3-sample judgePanel, gates reason_substance on the panel mean and the
present/commits/has_because checks on a 2-of-3 majority, thresholds
unchanged. armJudge no longer re-asks on a malformed verdict; it is a
failed sample, as the judge policy requires.
2026-09-29 19:51:35 +00:00
garrytan 2fb55e4c7a test(evals): record the read-only and detector-row invariants as contracts
shared-libs-opportunity-judgment and review-design-lite are behavior
cases: their recommendation and checklist judgments may vary, but the
read-only invariant (commands, provider requests, fixture bytes, hooks,
state) and the deterministic fake-engine detector rows are contracts.
Both now go through expectContract, so any failure vetoes the panel.
2026-09-29 19:51:35 +00:00
garrytan 1e640bbfb7 fix(evals): tsx-safe generics in eval-flake-rank, legacy artifact names, no-retry wall docs 2026-09-29 19:47:44 +00:00
garrytan b2ca207cf0 Merge remote-tracking branch 'origin/capy/rel-a' into capy/rel-c 2026-09-29 19:47:17 +00:00
garrytan 74c3184dcc test(pty): grant an owned Create pane whose title row is cropped
The targeted batching rerun on Claude Code 2.1.284 left its first report
Write unanswered for 1,372 s and timed out: the viewport began at the
pane's relative file row and rule, with the 'Create file' title cropped
above, so the preview parser rejected the file row as foreign. That row
must now resolve to the owned path and is skipped before the unchanged
line-by-line preview match. Replay controls reject another file, another
directory and an edited preview row.
2026-09-29 19:47:15 +00:00
garrytan 32772f535d feat(evals): --case/--trials local diagnosis and panels in local sharded runs
bun run scripts/test-paid-shards.ts --case <id> [--trials N] runs N
independent trials of one case through the CI panel runner (trial shards,
TRIAL_ENV identity, name-pattern isolation) and prints its panelVerdict();
N defaults to the case's policy panel and CI never reads it. The local
sharded path (test:gate:sharded, test:periodic:sharded) now plans the same
trial shards and exclusions as CI and exits on execution completeness plus
panel verdicts.
2026-09-29 19:44:55 +00:00
garrytan 3bae8e33da feat(evals): planner-side whole-panel reuse and negative receipts
The planner job restores this PR's receipt store once and ships a single
filtered set with the plan: a pass or panel receipt with a same-or-newer
FAIL for its input identity is dropped, and a panel receipt ships only as
a whole PASS panel (re-verified with panelVerdict) from one run. Executors
read only that set (no per-slice cache restore or save), so every trial of
a panel sees the same receipts; a trial reuses its own record from the
panel receipt, keeping a split PASS's failed trial.

Trial identities drop the trial index (run-scoped) and bind the panel
policy. Executed shards carry their input identity; the report turns a
whole fresh PASS panel into a panel receipt and a FAIL panel or failed rule
shard into a negative receipt, and marks a panel that mixes reused and
fresh trials INCOMPLETE. The report job merges plan, slice and report
receipts (newest per file) and saves one store per run.

Also fixes two TS2352 casts in browse/test/dia-macos-qualification.test.ts
whose diagnostic text drifted with program order (baseline locked, fix only).
2026-09-29 19:41:39 +00:00
garrytan 8622535b90 ci(evals): attempt-scoped artifacts, verdict-v2 PR comment, weekly pass-rate gate and one INFRA re-dispatch
- Slice, census and marathon artifacts carry -a<run_attempt>; reports
  download them per artifact (no merge), so records never overwrite and a
  re-run never replaces the first attempt's verdict.
- Planners pass --max-parallel for the capacity preflight (24/16 unchanged:
  the refreshed periodic plan needs 24 slices, the gate census 12).
- PR comment: jq-only job reads collector-outcomes v2 (headline, sanitized
  failure block); the group_by(.name)|last recomputation is gone.
- Reports stamp series identities, upload trial-outcomes-* for history, and
  shard logs upload always (a failed trial no longer reds its runner).
- Weekly report: headline + failure block of both lanes in the issue body,
  the eval:pass-rates --gate step (fails closed without history), close the
  issue on a green run, and UC-E1: when every red is machine-classified
  INFRA/INCOMPLETE, one re-dispatch as a new run in its own concurrency
  group (redispatch_of), both runs reported.
2026-09-29 19:34:33 +00:00
garrytan 62fb9a255d feat(evals): stamp trial series identities and fit panels to the live registry
- scripts/eval-trial-series.ts stamps series_identity (eval-flake-rank's
  caseSeriesIdentities) on a report's trial-outcomes JSONL as its own step,
  keeping the history tool out of the paid runner's closure;
  TrialOutcomeRecord gains the optional series_identity field.
- Slice-count plans let a registered trial spill into an ordinary lane when
  its siblings hold every long lane, so panels never share a runner.
- Re-audited test-selection.ts (Stream B added the E2E_KINDS/BEHAVIOR_WHY
  map-diff; no new module loading) and repinned its hash.
- Detach and release floors now count trial shards (66 periodic trials in
  22 panels): periodic floor 33,821s, still under eval:bg:periodic's 67,380s.
- Coordination fixtures supply the executor's trial records.
2026-09-29 19:30:43 +00:00
garrytan d5876efa5d chore(evals): refresh paid duration seeds from proof runs 36597762183 and 36606688266
Both tiers, merged in run order (the later run wins). Notable: split-overflow
1332s -> 504s, section-loading 604s -> 342s, mode-routing 575s -> 444s;
multi-finding-batching 734s -> 1318s (its red path in run 36606688266).
2026-09-29 19:30:43 +00:00
garrytan a385e5de18 Merge remote-tracking branch 'origin/capy/rel-b' into capy/rel-a 2026-09-29 19:24:03 +00:00
garrytan 8eb55b6ebf feat(evals): trial planner, slice exit split and panel-verdict report
Planner: behavior and quarantined cases become panels of isolated trial
shards (<file>#<id>~t<N>) bound by EVALS_SELECTION_JSON=[id] and the exact
test name; the file shard excludes them by name. Trials of one case never
share a slice, result slugs are unique, panels are validated whole, unknown
registrations throw, and the planner prints a capacity preflight.

Executor: each trial shard gets its TRIAL_ENV identity and a trial record
(outcome, failure class, cause, cost); every shard writes a JUnit report.
The slice exit now means execution completeness: a failed rule shard or a
trial without a record reds the runner, a failed trial does not.

Report: panelVerdict() decides every panel of the first run attempt (later
attempts are reported, never replacing it); rule shards keep the unchanged
fail-closed checks; collector records all count (no last-attempt wins);
census runs enforce the quarantine cap and expiry. It writes
collector-outcomes v2, trial-outcomes.jsonl (trials plus JUnit rule/judge
cases), report-summary.md, and one headline + failure block with rerun
commands, and flags INFRA/INCOMPLETE-only reds for the one re-dispatch.

The fail-open suite gains the panel cases: behavior 1/3 red, 2/3 green
with its failed trial shown, missing trial INCOMPLETE, contract at 2/3 red,
quarantined 1/3 green, 0/3 and contract red, missing slice red, and a later
attempt never replacing the first.
2026-09-29 19:23:55 +00:00
garrytan ccb5f3c07c docs(evals): document the pre-registered verdict policy, quarantine, pass-rate history and arithmetic
AGENTS.md replaces the retry rule with the approved policy text (no retries;
kind fixes trials; no added trials, samples or dispatches after a result;
quarantine by CASE_QUARANTINE only; one INFRA/INCOMPLETE re-dispatch) and
notes that a pre-registered fixed panel is not rejudging. CONTRIBUTING gains
the kind rules, the judge panel, eval:pass-rates and an 'Add a paid eval'
checklist. TESTING_INTERNALS describes verdicts, quarantine, history and the
arithmetic, including the rule term: 1 trial vs 2-of-3 red rates at
p = 0.99/0.95/0.90/0.70/0.30 and lane all-green probabilities for the
current 191 rule / 22 behavior / 25 judge registry.
2026-09-29 19:17:12 +00:00
garrytan 49761c97d6 test(ceo-mode-routing): submit a mode review that scrolled past the viewport
Run 36606688266 bundled routing, learnings and the mode choice into one
native call. Its review panel was taller than the terminal, so the tab
bar scrolled off, ceoModeSubmissionInput returned null for 240 s and HOLD
SCOPE was never submitted ('no posture match'). With no bar on screen the
viewport must still end at the focused Submit prompt, and the accumulated
screen text supplies the one complete panel, authenticated exactly as
before. Replay controls reject another mode, an unoffered answer, an
altered question, a quoted panel, trailing output, a moved cursor and an
answered or changed call.
2026-09-29 19:16:01 +00:00
garrytan 69cb4c8e9d fix(deslop-shared-libs): read related sources together within the turn limit
Run 36606688266's opportunity audit read sixteen sources one per turn and
stopped at error_max_turns; the passing run 36597762183 read the same
files in three batched commands. The skill now says turns are bounded and
asks for parallel reads or one read-only command per step.
2026-09-29 19:16:00 +00:00
garrytan 65f28fa9e4 fix(plan-design-review): treat a designer with no API key as unavailable
Both proof runs (36597762183, 36606688266) printed DESIGN_READY, hit
'No OpenAI API key found' on the first $D variants call, then hand-built
HTML/CSS wireframes, screenshots and a comparison board for ~195-245 s
before the first review question; the second run timed out at 600 s.
A failed first generation now takes the existing text-only path, and the
skill forbids substituting hand-built mockups.
2026-09-29 19:16:00 +00:00
garrytan 8bd53faa6c test(eng-batching): bind unsourced native briefs through the report's target
Run 36606688266 asked ten separate native review questions (D1-D9 bound
to ledger records R1-R9) and failed reviewCount=0 < FLOOR=3: its briefs
named the plan by title instead of citing PLAN.md, its report declared
'Review target (fixed): PLAN.md' under '# Engineering review: <plan>', and
it kept an unfenced copy of the plan's own H1. The named-source route now
accepts those spellings and non-inline ledger briefs. The same replay
rejects a foreign, mixed, duplicate or missing target, another plan's
title or copied H1, a brief naming another plan or file, a mismatched
saved brief, and re-asks. The run-36597762183 capture still counts 3.
2026-09-29 19:15:06 +00:00
garrytan 5269452983 test(eng-batching): grade the floor once the review report is complete
A completed GSTACK REVIEW REPORT ends the review, so the review-question
count is final there. Run 36606688266 wrote its report at 1,248 s and
closed the session at 1,318 s; the case now stops collection and applies
the unchanged floor at the report instead of waiting out the session.
No budget changes.
2026-09-29 19:15:06 +00:00
garrytan 392d63a537 ci(image): pin Claude Code 2.1.284 so the eval model is recognized
2.1.251 logs [claude-code:unrecognized_model] for claude-fable-5-1, the
eval capture/judge default. 2.1.284 does not. The gate PTY smoke subset
(plan-ceo/plan-devex plan-mode, plan-mode-no-op) parses on the new TUI;
plan-design-review-plan-mode passed at 293 s on 2.1.284 and timed out at
300 s on 2.1.251 on the same tree.
2026-09-29 19:14:54 +00:00