GSTACK_QA_REQUIRED_PROBES (a JSON array of native child commands) makes every
capture print requiredRemaining; it never judges pass or fail. The functional
eval passes the webhook list from QA_WEBHOOK_REQUIRED_SCENARIOS, which the
verdict now reads too, so the nudge and the verdict share one source (agreed
with #3002's owner). CI webhook-report kept stopping with scenarios unrun.
When native probe output declares a top-level input snapshot, materialize
compares each evidence row with the latest capture's snapshot and refuses
stale rows unless they are classified superseded, naming the captures to
rerun. ship-exploratory-late-input kept reporting a pre-change adverse probe
green after the input changed.
- materialize refuses when a complete capture has no evidence row and is not
named in limits (CI cli-report omitted capture 004), naming the missing IDs.
- tpa-apple-ban's detector required 'app-specific password' with a space; the
CI answer said 'app-specific-password path' and was otherwise correct.
Approved by Garry: asking gstack-qa-evidence or gstack-qa-deadline for usage
is read-only, so both helpers print usage and exit 0 on --help (the evidence
usage now names the annotation shape), and the functional and caller command
allowlists accept exactly 'bun <path>/bin/gstack-qa-{evidence,deadline} --help'.
Two CI runs failed only on that call.
- materialize measures revision, runtime and cwd itself and rejects supplied
values that differ (CI run wrote revision "HEAD" and runtime "bun"), and
refuses learning checkpoints that replay the same probe, naming the fix.
- The docs write observer treats Claude Code's atomic temp for the authorized
doc target as transient, so a temp renamed before its per-file watch no
longer marks the observation incomplete (ship-docsync-completion flake).
Per-file monitoring outside declared targets stays fail-closed.
- capture refuses to run another probe until a checkpoint anchored on the
latest complete capture names this capture as its next command, and every
complete capture prints that requirement.
- materialize fills revision, runtime, cwd and learning (checkpoints whose next
native command differs) when omitted and prints the reportLinks the report
must include; the QA section shrinks accordingly.
- gstack-qa-evidence capture prints startedAt/completedAt/durationMs and, for
--deadline captures, remainingMs; the functional report takes durations from
them. The section clock notice asks for one clock read up front instead of one
after every checkpoint (QA runs spent 7-14% of tool calls on date -u).
- ship plan-completion: skip the audit dispatch when discovery already found no
plan (the dispatch-vs-skip conflict produced an optional 60-100 s subagent).
- materialize/checkpoint validation errors state the expected schema, so a
rejected annotations file is fixable in one call instead of blocking the phase.
- session-runner counts turns from the transcript when a run times out, so
timeouts stop reporting 'turn 0'.
qa-functional-webhook-report failed in two of three censuses because the
report linked .qa-evidence/NNN capture folders as "checkpoints" and never
linked exploration-NNN.json. The checkpoint receipt now prints
link: [checkpoint NNN](exploration-NNN.json), and the functional report
template says capture folders are not checkpoints.