Files
gstack/.github/workflows/evals.yml
T
Garry Tan a84b0b5b6d v1.90.0.0 feat: make browser cookie imports explicit and safe (#2964)
* fix(browse): prepare reliable cookie import wave for validation

* ci: sequence quality and behavior for validation branch

* fix(browse): isolate Windows qualification and preserve native diagnostics

* test(browse): cover cookie workflow quality and isolate Windows user paths

* test(browse): trace native member startup and initialize fresh folders

* fix(browse): keep Windows member stdin alive through EOF

* fix(browse): latch native timeouts and compare contained Edge startup

* test(browse): verify native version metadata and actual Windows argv

* test(browse): qualify Dia import on isolated macOS CI

* fix(browse): require picker origin for session mutations

* fix(browse): bound credential reads through stream completion

* test(browse): inspect owned Windows process arguments natively

* test(evals): preserve passing coverage during cookie repair reruns

* test(browse): isolate Dia qualification in a fresh macOS account

* test(browse): pass bounded integer timeouts to native Mac probes

* test(browse): distinguish Windows profile initialization from containment

* test(browse): await descendant pipe readiness before parent exit

* test(browse): initialize and restore isolated macOS Keychain state

* test(browse): initialize Windows fixture folders before qualification

* test(ci): pin the same Node runtime across Windows checks

* test(browse): distinguish native macOS browser preflight stages

* test(browse): isolate Windows descendant console lifetime

* test(browse): preserve native receipts and identify fixture lock holders

* test(browse): prepare dependency resolution before native Mac worker startup

* test(ci): include lock and close checks in native diagnostics

* test(browse): preserve native owner probe stages and subprocess deadlines

* fix(browse): classify Chromium profile-in-use exit precisely

* test(browse): retain Mac qualification evidence through cleanup failures

* test(browse): bound Mac fixture paths and retire its owned user domain

* test(browse): accept vanished fixture entries without weakening cleanup

* test(browse): identify probe-created macOS user domains safely

* test(browse): observe Mac user domains without targeting them first

* test(browse): use passive fresh-user ownership throughout Mac qualification

* test(browse): distinguish profile and registered-home Keychain lookups

* test(browse): qualify Dia under one registered account home

* test(browse): identify Dia startup and owned process-group failures

* test(browse): classify bounded Dia startup diagnostics without leaking output

* fix(test): preserve native Mac sandboxing and reap owned browser children

* fix(browse): preserve Chromium sandboxing for native profile imports

* test(browse): inspect signed Mach-O architecture without launching Xcode tools

* test(browse): sample pending Dia startup and reap on all cleanup paths

* test(browse): compare protected Dia launches in fresh Bun and Node accounts

* test(browse): inspect isolated Mac GUI readiness without browser access

* v1.90.0.0 fix: bind cookie picker actions to their document

* test: validate cookie guards and fit nested launch fixtures

* ci: configure the bundled Chromium sandbox helper

* fix(browse): classify Playwright authentication timeouts

* test: retain bounded Windows lifecycle diagnostics

* test(cso): reuse bounded NTFS precision candidates

* test(review): handle explicit preservation choices safely

* test(browse): remove owned fixture directories with explicit primitives

* test(review): distinguish descriptive reuse from edit commitments

* test: admit only the approved unscored cookie workflow refusal

* test: keep the Office Hours judge mock export-complete

* fix: keep dependency-free CI planners independent of the model SDK

* test: observe the exact holder after a native fixture unlink failure

* fix: start seeded PTY observations at owned readiness

* test: acquire identity-bound Windows deletion admission before profile resets

* test: preserve qualified Git index bits without authorizing mutations
2026-09-25 12:06:45 -04:00

545 lines
28 KiB
YAML

name: E2E Evals
on:
pull_request:
branches: [main]
workflow_dispatch:
inputs:
evals_all:
description: 'Run ALL gate tests in the sliced lane (bypass diff selection; also arms the hollow-shard guard)'
type: boolean
default: true
validation_phase:
description: 'Validation branch phase; run quality before behavior on unchanged inputs'
type: choice
options: [all, quality, cookie-quality, behavior, cookie-behavior]
default: all
concurrency:
group: evals-${{ github.event.pull_request.number || github.run_id }}
cancel-in-progress: true
env:
IMAGE: ghcr.io/${{ github.repository }}/ci
# PRs run changed fast probes; manual runs retain the complete gate census.
EVALS_PROFILE: ${{ github.event_name == 'pull_request' && 'pr' || 'full' }}
EVALS_FRESH: ${{ github.event_name == 'workflow_dispatch' && '1' || '' }}
jobs:
# Build Docker image with pre-baked toolchain (cached — only rebuilds on Dockerfile/lockfile change)
build-image:
# Dependabot-triggered pull_request runs get a read-only GITHUB_TOKEN, so
# a lockfile bump = new hash = failed ghcr push = permanently red check
# (EV6, fork port wave 2). Skip the build for dependabot; the evals job's
# explicit actor guard mirrors it because no eval test selects on a lockfile-only
# diff — a maintainer's next push rebuilds the image with real perms.
if: github.actor != 'dependabot[bot]'
runs-on: ubicloud-standard-8
timeout-minutes: 15
permissions:
contents: read
packages: write
outputs:
image-tag: ${{ steps.meta.outputs.tag }}
runtime-id: ${{ steps.runtime.outputs.id }}
steps:
- uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7
- id: meta
# Key on Dockerfile + lockfile only. package.json is deliberately NOT
# hashed: its version field changes on every ship (60/60 recent commits),
# which rebuilt the image each time for a dependency set that only
# bun.lock determines. A stale baked package.json is harmless — checkout
# overwrites /workspace and node_modules comes from the lockfile.
run: echo "tag=${{ env.IMAGE }}:${{ hashFiles('.github/docker/Dockerfile.ci', 'bun.lock', 'patches/**') }}" >> "$GITHUB_OUTPUT"
- uses: docker/login-action@dbcb813823bdd20940b903addbd779551569679f # v4
with:
registry: ghcr.io
username: ${{ github.actor }}
password: ${{ secrets.GITHUB_TOKEN }}
- name: Check if image exists
id: check
run: |
if docker manifest inspect ${{ steps.meta.outputs.tag }} > /dev/null 2>&1; then
echo "exists=true" >> "$GITHUB_OUTPUT"
else
echo "exists=false" >> "$GITHUB_OUTPUT"
fi
- if: steps.check.outputs.exists == 'false'
run: cp package.json bun.lock .github/docker/ && cp -R patches .github/docker/patches
# A fork PR's GITHUB_TOKEN only has `packages: read`, so pushing fails.
# Still BUILD (validates Dockerfile.ci changes), just don't publish. This
# job intentionally keeps no `if:` so fork PRs still get one real, honest
# green check here instead of a run where every job is grey.
# Registry cache export needs a docker-container builder — the default
# `docker` driver hard-errors on cache-to (first live run of the trio).
- if: steps.check.outputs.exists == 'false'
uses: docker/setup-buildx-action@37fe631027851001ddb9b187196cc803df7f5f0e # v4
- if: steps.check.outputs.exists == 'false'
uses: docker/build-push-action@53b7df96c91f9c12dcc8a07bcb9ccacbed38856a # v7
with:
context: .github/docker
file: .github/docker/Dockerfile.ci
push: ${{ github.event_name != 'pull_request' || github.event.pull_request.head.repo.full_name == github.repository }}
# Registry layer cache: reads are safe everywhere; the export is gated
# to same-repo runs because a fork PR's token can't write GHCR.
cache-from: type=registry,ref=${{ env.IMAGE }}:buildcache
cache-to: ${{ (github.event_name != 'pull_request' || github.event.pull_request.head.repo.full_name == github.repository) && format('type=registry,ref={0}:buildcache,mode=max', env.IMAGE) || '' }}
tags: |
${{ steps.meta.outputs.tag }}
${{ env.IMAGE }}:latest
- name: Identify the installed eval runtime
id: runtime
if: github.event_name != 'pull_request' || github.event.pull_request.head.repo.full_name == github.repository
env:
EVAL_IMAGE: ${{ steps.meta.outputs.tag }}
run: |
docker manifest inspect "$EVAL_IMAGE" > /tmp/eval-runtime-manifest.json
echo "id=$(sha256sum /tmp/eval-runtime-manifest.json | cut -d ' ' -f1)" >> "$GITHUB_OUTPUT"
# ── Sliced lane (the ONLY paid lane; legacy 17-row matrix deleted) ──────────
# One PLANNER computes diff selection + the slice plan ONCE (killing
# per-slice selector divergence); K executors consume the manifest; the
# report reconciles results against it FAIL-CLOSED (a slice whose artifact
# never landed is a failure, a planned shard nobody reported is a failure —
# hollow lanes cannot aggregate green). Engine: scripts/test-paid-shards.ts —
# the same runner local eval:bg:gate uses, so CI and local share one
# selection engine, and every gate-tier file is in the census by
# construction (no hand-enumerated rows to drift). The legacy matrix ran
# 18 enumerated files for 22.6 min/$21 per PR serialized AHEAD of this
# lane's 49-file diff-selected census; parity was demonstrated (sliced
# census ⊇ matrix files) and the matrix deleted — one revert restores it.
#
# Fork PRs never receive repository secrets (ANTHROPIC_API_KEY et al), so
# every API-calling eval fails at SDK auth before a model runs. Skip
# deterministically; fork work gets real coverage via a trusted base-repo
# branch (see CLAUDE.md's garrytan-agents workflow).
plan-slices:
runs-on: ubicloud-standard-8
if: github.actor != 'dependabot[bot]' && (github.event_name != 'pull_request' || github.event.pull_request.head.repo.full_name == github.repository)
timeout-minutes: 10
permissions:
contents: read
steps:
- uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7
with:
# Preserve full history for merge-base diff selection. Moving the
# planner off the eval image must not change its selection inputs.
fetch-depth: 0
persist-credentials: false
- uses: oven-sh/setup-bun@0c5077e51419868618aeaa5fe8019c62421857d6 # v2
with:
bun-version: 1.4.0
- name: Emit run manifest
if: github.event_name != 'workflow_dispatch' || inputs.validation_phase == 'all'
env:
EVALS_ALL: ${{ (github.event_name == 'workflow_dispatch' && inputs.evals_all) && '1' || '' }}
run: EVALS_TIER=gate bun --no-install run scripts/test-paid-shards.ts --tier gate --emit-plan /tmp/paid-plan/manifest.json --slices 6
- name: Emit validation-phase manifest
if: github.event_name == 'workflow_dispatch' && inputs.validation_phase != 'all'
env:
VALIDATION_PHASE: ${{ inputs.validation_phase }}
EVALS_ALL: ${{ inputs.evals_all && '1' || '' }}
EVALS_TIER: gate
run: |
bun --no-install -e '
import { mkdirSync, writeFileSync } from "node:fs";
import { buildRunManifest, collectPaidTestFiles } from "./scripts/test-paid-shards.ts";
const phase = process.env.VALIDATION_PHASE;
if (!["quality", "cookie-quality", "behavior", "cookie-behavior"].includes(phase)) throw new Error("Invalid validation phase");
const cookieBehavior = phase === "cookie-behavior";
const discovered = phase === "cookie-quality" ? ["test/skill-llm-eval.test.ts"]
: cookieBehavior ? ["test/skill-e2e-bws.test.ts", "test/skill-e2e-qa-workflow.test.ts", "test/skill-e2e-design.test.ts", "test/skill-e2e-diagram.test.ts", "test/skill-e2e-deploy.test.ts"]
: collectPaidTestFiles().filter(file => file.startsWith("test/skill-llm-eval") === (phase === "quality"));
const manifest = buildRunManifest({ tier: "gate", profile: "full", sliceCount: 6, evalsAll: !cookieBehavior && process.env.EVALS_ALL === "1", discovered,
...(cookieBehavior ? { changedFiles: ["browse/src/cookie-picker-routes.ts", "browse/src/cookie-import-browser.ts", "browse/src/bun-polyfill.cjs"], env: { ...process.env, EVALS_ALL: "" } } : {}) });
if (phase === "cookie-quality") manifest.selection = { e2e: [], judges: ["setup-browser-cookies/SKILL.md workflow"] };
if (cookieBehavior) manifest.selection = { e2e: ["browse-basic", "browse-snapshot", "qa-quick", "qa-only-no-fix", "design-review-detector-shim-dom", "diagram-triplet", "canary-workflow", "benchmark-workflow"], judges: [] };
manifest.selectionReason = phase + " validation subset; " + manifest.selectionReason;
mkdirSync("/tmp/paid-plan", { recursive: true });
writeFileSync("/tmp/paid-plan/manifest.json", JSON.stringify(manifest, null, 2) + "\n");
console.log(phase + ": " + manifest.entries.filter(entry => entry.status === "planned").length + " planned shards");
'
- uses: actions/upload-artifact@043fb46d1a93c77aae656e7c1c64a875d1fc6a0a # v7
with:
name: paid-plan
path: /tmp/paid-plan/manifest.json
retention-days: 30
eval-slices:
runs-on: ubicloud-standard-8
needs: [build-image, plan-slices]
if: always() && needs.build-image.result == 'success' && needs.plan-slices.result == 'success'
# Aggregate spawn-concurrency budget: 6 slices x EVALS_JOBS=2 x
# EVALS_CONCURRENCY=2 = 24 concurrent tests lane-wide (the old matrix's
# 40-way per row queued claude session STARTUP behind 39 siblings and ate
# per-test budgets — the documented timeout-flake family). Tune with
# parity data before raising.
# The complete gate census needs at most 201 minutes per slice; keep
# 20 minutes for setup/upload without preempting configured retries.
timeout-minutes: 221
permissions:
contents: read
packages: read
container:
image: ${{ needs.build-image.outputs.image-tag }}
credentials:
username: ${{ github.actor }}
password: ${{ secrets.GITHUB_TOKEN }}
options: --user runner
strategy:
fail-fast: false
matrix:
slice: [1, 2, 3, 4, 5, 6]
steps:
- uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7
with:
# Full history: files with SELF-derived selection (the LLM-judge
# map, routing) walk git at module load, and selection is
# fail-closed on git errors — a shallow checkout crashed those
# shards on the lane's first live run ("ambiguous argument
# 'main...HEAD'"). The manifest still governs WHICH shards run.
fetch-depth: 0
persist-credentials: false
- name: Fix bun temp
uses: ./.github/actions/fix-bun-temp
- name: Restore deps
uses: ./.github/actions/restore-deps
- run: bun run build
# Any slice can host a PTY smoke, so the seed/registration steps run
# UNCONDITIONALLY (both are idempotent) — the old matrix keyed them on
# matrix.suite.name, which a sliced lane cannot do. The register
# composite carries the fail-fast dangling-symlink/frontmatter
# verification loop, so a moved skill target fails HERE in seconds,
# not as a wedged PTY session at the shard wall.
- name: Seed claude interactive config
uses: ./.github/actions/seed-claude-config
with:
anthropic-api-key: ${{ secrets.ANTHROPIC_API_KEY }}
- name: Register gstack skills for PTY smokes
uses: ./.github/actions/register-gstack-skills
- uses: actions/download-artifact@3e5f45b2cfb9172054b4087a40e8e0b5a5461e7c # v8
with:
name: paid-plan
path: /tmp/paid-plan
# Only this PR's receipts are eligible. No base-branch or cross-PR restore
# prefix; every receipt also verifies exact inputs and its original age.
- name: Restore this PR's verified judge results
if: github.event_name == 'pull_request'
uses: actions/cache/restore@55cc8345863c7cc4c66a329aec7e433d2d1c52a9 # v6
with:
path: /tmp/gstack-eval-input-cache
key: eval-input-v1-${{ github.repository_id }}-pr-${{ github.event.pull_request.number }}-${{ github.run_id }}-${{ github.run_attempt }}-${{ matrix.slice }}
restore-keys: eval-input-v1-${{ github.repository_id }}-pr-${{ github.event.pull_request.number }}-
- name: Run slice ${{ matrix.slice }}/6
env:
ANTHROPIC_API_KEY: ${{ secrets.ANTHROPIC_API_KEY }}
OPENAI_API_KEY: ${{ secrets.OPENAI_API_KEY }}
GEMINI_API_KEY: ${{ secrets.GEMINI_API_KEY }}
PLAYWRIGHT_BROWSERS_PATH: /opt/playwright-browsers
EVALS_JOBS: "2"
EVALS_CONCURRENCY: "2"
GSTACK_EVAL_DIR: /tmp/paid-slice-results
EVALS_CACHE_DIR: /tmp/gstack-eval-input-cache
EVALS_CACHE_REPOSITORY: ${{ github.repository }}
EVALS_CACHE_PR: ${{ github.event.pull_request.number }}
EVALS_CACHE_RUNTIME_ID: ${{ needs.build-image.outputs.runtime-id }}
run: EVALS_TIER=gate bun run scripts/test-paid-shards.ts --tier gate --plan /tmp/paid-plan/manifest.json --slice ${{ matrix.slice }}
- name: Find finalized passing receipts
id: receipts
if: ${{ !cancelled() && github.event_name == 'pull_request' }}
run: |
# Only a producer publishes. A later reuse-only slice must not become
# the newest prefix match and hide another slice's newly earned pass.
for receipt in /tmp/gstack-eval-input-cache/*.json; do
[ -f "$receipt" ] || continue
if jq -e --arg run "$GITHUB_RUN_ID/$GITHUB_RUN_ATTEMPT" '.proof.source.runId == $run' "$receipt" >/dev/null 2>&1; then
echo 'present=true' >> "$GITHUB_OUTPUT"
break
fi
done
# An unrelated failing case does not discard already verified passes.
# Failed/retried/partial attempts never become receipts in the first place.
- name: Save verified judge results for this PR
if: ${{ !cancelled() && steps.receipts.outputs.present == 'true' }}
uses: actions/cache/save@55cc8345863c7cc4c66a329aec7e433d2d1c52a9 # v6
with:
path: /tmp/gstack-eval-input-cache
key: eval-input-v1-${{ github.repository_id }}-pr-${{ github.event.pull_request.number }}-${{ github.run_id }}-${{ github.run_attempt }}-${{ matrix.slice }}
- name: Upload slice results
if: always()
uses: actions/upload-artifact@043fb46d1a93c77aae656e7c1c64a875d1fc6a0a # v7
with:
name: paid-slice-${{ matrix.slice }}
path: /tmp/paid-slice-results
retention-days: 90
# The spooled per-shard full logs — a red weekly/PR lane three weeks
# later needs more than a summary line.
- name: Upload shard logs on failure
if: failure()
uses: actions/upload-artifact@043fb46d1a93c77aae656e7c1c64a875d1fc6a0a # v7
with:
name: paid-slice-${{ matrix.slice }}-logs
include-hidden-files: true
# The Fix-bun-temp step points TMPDIR at /home/runner/.cache, so the
# runner's spool lands THERE, not /tmp — the original /tmp glob
# uploaded nothing and a red slice's diagnostics were unreachable.
path: |
/home/runner/.cache/gstack-paid-shard-*.log
/tmp/gstack-paid-shard-*.log
if-no-files-found: ignore
retention-days: 30
slices-report:
runs-on: ubicloud-standard-2
needs: [plan-slices, eval-slices]
# always(): the report must run (and FAIL) when an executor died — a
# missing slice artifact reading as green is the class this lane kills.
if: always() && needs.plan-slices.result == 'success'
timeout-minutes: 5
# contents:read ONLY — this job executes the PR-authored reconcile
# runner from the PR checkout, so it
# must never hold a write-scoped token. The PR comment lives in the
# separate slices-comment job below, which runs NO repo code: a
# $GITHUB_ENV/BASH_ENV persistence trick is job-scoped, so the split is
# the trust boundary (codex adversarial finding, 2026-08-31 — the old
# matrix-era report job had this separation and the consolidation had
# regressed it).
permissions:
contents: read
outputs:
reconcile-exit: ${{ steps.reconcile.outputs.exit }}
steps:
- uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7
with:
persist-credentials: false
- uses: oven-sh/setup-bun@0c5077e51419868618aeaa5fe8019c62421857d6 # v2
with:
bun-version: 1.4.0
- uses: actions/download-artifact@3e5f45b2cfb9172054b4087a40e8e0b5a5461e7c # v8
with:
name: paid-plan
path: /tmp/paid-report
- uses: actions/download-artifact@3e5f45b2cfb9172054b4087a40e8e0b5a5461e7c # v8
with:
pattern: paid-slice-[0-9]*
path: /tmp/paid-report
merge-multiple: true
- name: Reconcile slices against the manifest (fail-closed)
id: reconcile
run: |
set +e
EVALS_TIER=gate bun --no-install run scripts/test-paid-shards.ts --tier gate --report /tmp/paid-report | tee /tmp/report.txt
# PIPESTATUS[0], NOT $?: GitHub's default run-step shell is
# `bash -e {0}` with NO pipefail, so $? after the pipe is tee's
# exit (always 0) — the fail-closed gate was silently fail-open
# (caught by the ship review army; the wiring test now pins this).
echo "exit=${PIPESTATUS[0]}" >> "$GITHUB_OUTPUT"
- name: Upload reconciliation output for the comment job
if: always()
uses: actions/upload-artifact@043fb46d1a93c77aae656e7c1c64a875d1fc6a0a # v7
with:
name: report-verdict
path: |
/tmp/report.txt
/tmp/paid-report/collector-outcomes.json
if-no-files-found: ignore
retention-days: 30
- name: Fail the workflow when reconciliation failed
if: steps.reconcile.outputs.exit != '0'
run: exit 1
# PR comment in its OWN job with the write token and ZERO repo code: no
# checkout, no bun install — only downloaded artifacts, jq, and gh. See the
# trust-boundary note on slices-report.
slices-comment:
runs-on: ubicloud-standard-2
needs: slices-report
if: always() && github.event_name == 'pull_request' && github.event.pull_request.head.repo.full_name == github.repository && needs.slices-report.result != 'skipped'
timeout-minutes: 5
permissions:
pull-requests: write
# The comment upsert calls the REST `/issues/{n}/comments` endpoints
# (gh api ... issues/comments). With GITHUB_TOKEN those are gated by the
# `issues` permission, not `pull-requests` (#1802 CI fix).
issues: write
steps:
- uses: actions/download-artifact@3e5f45b2cfb9172054b4087a40e8e0b5a5461e7c # v8
with:
name: paid-plan
path: /tmp/paid-report
- uses: actions/download-artifact@3e5f45b2cfb9172054b4087a40e8e0b5a5461e7c # v8
with:
pattern: paid-slice-[0-9]*
path: /tmp/paid-report
merge-multiple: true
- uses: actions/download-artifact@3e5f45b2cfb9172054b4087a40e8e0b5a5461e7c # v8
with:
name: report-verdict
path: /tmp/verdict
continue-on-error: true
# Verified counts come from the read-only report job, not repo code in
# this write-token job. Keeps the
# "## E2E Evals" marker so the upsert keeps updating the same comment.
# Runs even when reconciliation failed — a red lane on the PR is the point.
- name: Post PR comment
env:
GH_TOKEN: ${{ secrets.GITHUB_TOKEN }}
RECONCILE_EXIT: ${{ needs.slices-report.outputs.reconcile-exit }}
run: |
# shellcheck disable=SC2086,SC2059
RESULTS=$(find /tmp/paid-report -name '*.json' ! -name 'manifest.json' ! -name 'slice-*.json' ! -name '_partial*' 2>/dev/null | sort)
TOTAL=0; PASSED=0; FAILED=0; MANUAL=0; FLAKY=0; EXECUTED=0; REUSED=0; COST="0"
SUITE_LINES=""
VERIFIED=/tmp/verdict/paid-report/collector-outcomes.json
if ! jq -e '
. as $summary |
.version == 1 and (.files | type == "array") and (.totals | type == "object") and
([.files[] | .total == (.passed + .failed + .manual_accepted) and
(.total == (.executed + .reused)) and
([.total,.passed,.failed,.manual_accepted,.executed,.reused,.attempts,.flaky] | all(. >= 0 and (floor == .))) ] | all) and
(.totals | .total == (.passed + .failed + .manual_accepted) and .total == (.executed + .reused)) and
(["total","passed","failed","manual_accepted","executed","reused","attempts","flaky"] |
all(. as $key | ([$summary.files[] | .[$key]] | add // 0) == $summary.totals[$key]))
' "$VERIFIED" >/dev/null 2>&1; then
VERIFIED=""
echo 'Verified collector summary unavailable; manual acceptance is unavailable/unverified.'
fi
if [ -n "$VERIFIED" ]; then
while IFS=$'\t' read -r f T P F M FL EX RE _ATTEMPTS C TIER SHARD; do
[ "$T" -eq 0 ] && continue
TOTAL=$((TOTAL + T)); PASSED=$((PASSED + P)); FAILED=$((FAILED + F))
MANUAL=$((MANUAL + M)); FLAKY=$((FLAKY + FL))
EXECUTED=$((EXECUTED + EX)); REUSED=$((REUSED + RE))
COST=$(echo "$COST + $C" | bc)
STATUS_ICON="✅"
[ "$M" -gt 0 ] && STATUS_ICON="⚠ manual/unscored"
[ "$F" -gt 0 ] && STATUS_ICON="❌"
[ "$F" -eq 0 ] && [ "$M" -eq 0 ] && [ "$FL" -gt 0 ] && STATUS_ICON="✅⚠"
SUITE_LINES="${SUITE_LINES}| ${TIER}/${SHARD} | ${P}/${T} | ${M} | ${EX} | ${RE} | ${STATUS_ICON} | \$${C} |\n"
done < <(jq -r '.files[] | [.file,.total,.passed,.failed,.manual_accepted,.flaky,.executed,.reused,.attempts,.cost,.tier,.shard] | @tsv' "$VERIFIED")
else
for f in $RESULTS; do
if ! jq -e '.total_tests' "$f" >/dev/null 2>&1; then
echo "Skipping malformed JSON: $f"
continue
fi
# FINAL-attempt accounting: eval-store keeps EVERY retry attempt
# as its own record (that's the flake telemetry), so counting raw
# records marks a pass-on-retry as a failure and inflates totals.
# Group by test name and judge the LAST record. Retry metadata
# includes both passing and failing final outcomes; show it separately.
# Guarded: a file with total_tests but a null/non-array `tests`
# passes the -e probe, the group_by then fails, and an empty $T
# would abort the whole step under bash -e ([ "" -eq 0 ] is an
# error) — killing the comment on exactly the corrupted-artifact
# runs where the red evidence matters (claude adversarial).
STATS=$(jq -r '[.tests | group_by(.name)[] | last] as $final | "\($final | length) \([$final[] | select(.passed)] | length) \([$final[] | select(.passed | not)] | length) \(.flaky_retries // [] | length) \([$final[] | select(.execution != "reused")] | length) \([$final[] | select(.execution == "reused")] | length)"' "$f" 2>/dev/null) || { echo "Skipping malformed tests[] in: $f"; continue; }
read -r T P F FL EX RE <<< "$STATS"
[ -z "$T" ] && { echo "Skipping malformed tests[] in: $f"; continue; }
C=$(jq -r '.total_cost_usd // 0' "$f")
TIER=$(jq -r '.tier // "unknown"' "$f")
SHARD=$(jq -r '.shard // "-"' "$f")
[ "$T" -eq 0 ] && continue
TOTAL=$((TOTAL + T))
PASSED=$((PASSED + P))
FAILED=$((FAILED + F))
FLAKY=$((FLAKY + FL))
EXECUTED=$((EXECUTED + EX))
REUSED=$((REUSED + RE))
COST=$(echo "$COST + $C" | bc)
STATUS_ICON="✅"
[ "$F" -gt 0 ] && STATUS_ICON="❌"
[ "$F" -eq 0 ] && [ "$FL" -gt 0 ] && STATUS_ICON="✅⚠"
SUITE_LINES="${SUITE_LINES}| ${TIER}/${SHARD} | ${P}/${T} | unverified | ${EX} | ${RE} | ${STATUS_ICON} | \$${C} |\n"
done
fi
COVERAGE=$(jq -r '"Profile: \(.profile // "full") / \(.prCoverage.mode // "broad"); selected behaviors: \(.selection.e2e | if . == null then "all" else length end), judges: \(.selection.judges | if . == null then "all" else length end). Deferred to scheduled/release coverage: \(.prCoverage.deferred // [] | length) behaviors and \(.prCoverage.deferredPromptFiles // [] | length) changed prompt files. Deferred checks did not run and receive no PR-pass credit."' /tmp/paid-report/manifest.json) || COVERAGE='Coverage manifest unavailable; no coverage claim.'
STATUS="✅ PASS"
if [ "${RECONCILE_EXIT:-1}" != "0" ] || [ "$FAILED" -gt 0 ]; then STATUS="❌ FAIL"; fi
if [ "$STATUS" = '✅ PASS' ] && [ "$MANUAL" -gt 0 ]; then STATUS='⚠ MANUAL ACCEPTED (unscored)'; fi
if [ -z "$VERIFIED" ]; then STATUS='❌ FAIL (manual acceptance unavailable/unverified)'; fi
BODY="## E2E Evals: ${STATUS}
**${PASSED} automated passed / ${TOTAL} final results** | **${FAILED} failed, ${MANUAL} manual accepted (unscored; no score-cache credit)** | **${EXECUTED} executed, ${REUSED} reused** | **\$${COST}** total cost | reconcile exit: ${RECONCILE_EXIT:-missing}$([ "$FLAKY" -gt 0 ] && printf ' | ⚠ %s cases with multiple attempts' "$FLAKY")
${COVERAGE}
| Shard | Automated result | Manual/unscored | Executed | Reused | Status | Cost |
|-------|------------------|-----------------|----------|--------|--------|------|
$(echo -e "$SUITE_LINES")
<details><summary>Fail-closed reconciliation</summary>
\`\`\`
$(tail -c 4000 /tmp/verdict/report.txt 2>/dev/null || echo '(no reconciliation output)')
\`\`\`
</details>
---
*Sliced lane: declared PR profile or broad fallback via scripts/test-paid-shards.ts (planner → 6 executors → fail-closed report). Reused scores retain their original provenance and expiry.*"
if [ "$FAILED" -gt 0 ]; then
FAILURES=""
for f in $RESULTS; do
if ! jq -e '.failed' "$f" >/dev/null 2>&1; then continue; fi
if [ -n "$VERIFIED" ]; then
FAILS=$(jq -r '[.tests | group_by(.name)[] | last | select(.passed == false and (has("manual_review") | not))][] | "- ❌ \(.name): \(.exit_reason // "unknown")"' "$f" 2>/dev/null || echo "- ⚠️ parse error")
else
FAILS=$(jq -r '[.tests | group_by(.name)[] | last | select(.passed == false)][] | "- ❌ \(.name): \(.exit_reason // "unknown")"' "$f" 2>/dev/null || echo "- ⚠️ parse error")
fi
FAILURES="${FAILURES}${FAILS}\n"
done
BODY="${BODY}
### Failures
$(echo -e "$FAILURES")"
fi
COMMENT_ID=$(gh api repos/${{ github.repository }}/issues/${{ github.event.pull_request.number }}/comments \
--jq '.[] | select(.body | startswith("## E2E Evals")) | .id' | tail -1)
if [ -n "$COMMENT_ID" ]; then
gh api "repos/${{ github.repository }}/issues/comments/${COMMENT_ID}" \
-X PATCH -f body="$BODY"
else
# REST, not gh's pr-comment subcommand: this job runs with NO
# checkout (the token/exec split), and that subcommand resolves
# the repo FROM git — it dies with "not a git repository" here.
gh api "repos/${{ github.repository }}/issues/${{ github.event.pull_request.number }}/comments" \
-X POST -f body="$BODY"
fi