test: one owner per case id, a structural devex 0B setup rule, and correct design/gbrain actors

- plan-design-review-plan-mode was registered by two files; the PTY smoke is
  now plan-design-review-plan-mode-smoke, and a registry test requires one
  owner per case in case-sharded files.
- plan-devex-finding-floor: the template's 0B narrative-confirmation question
  is classified as setup structurally instead of timing out a Haiku assessor.
- setup-gbrain-remote: the actor accepted 'skip' on the MCP-registration
  question the test asserts; it now accepts that question and declines others.
- design-review-plugin-handoff: the fake engine cited a file absent from the
  fixture repo and index.html linked a missing styles.css.
Captured-question regressions with negative controls; each case passed a
focused paid run.
This commit is contained in:
garrytan committed 2026-09-30 12:24:05 +00:00
1 parent 8cf87d4729
commit dfe5e733fb
10 files changed
+152 -21

No files matched your search

@@ -0,0 +1,29 @@
{
"source": "Periodic Evals census 36641820398, gate-census (4), plan-devex-finding-floor: the assessor timed out (spawnSync ETIMEDOUT at its 30s cap) on this current D1 empathy confirmation.",
"review": {
"seed": "Proceed directly to the requested review; skip the optional /office-hours prerequisite.\n\nThis actor has already declined routing setup, cross-project recall and outside reviewers.\n\nPreserve the supplied product scope. For review-mode questions choose HOLD SCOPE (CEO), DX POLISH (DX), or the full BIG CHANGE review (Eng). Design: review all seven dimensions.\n\nProduct type is confirmed: SDK quickstart documentation, with the complete journey to the first SDK call as context. If asked to classify, choose SDK + Docs when offered, otherwise Documentation. This confirms the review lens; it does not expand the plan.\n\nTarget persona is confirmed: a hands-on developer integrating this SDK for the first time, trying to make one successful call. Product type and persona setup are already answered; proceed to reviewing the supplied plan.\n\nFor setup confirmations, this actor can supply only the following persona/journey correction through the native custom answer. It does not approve a proposed narrative, remedy, or scope change: Confirmed persona: a hands-on developer making a first SDK call. The declared onboarding facts are: Step 1: clone the repo. Step 2: install bun manually if not present. Step 3: copy .env.example to .env and fill in 8 environment variables. Step 4: run database migrations against your local Postgres. Step 5: start the dev server. Step 6: open the docs in a separate tab. Step 7: register an API key by emailing the team. Step 8: paste the key into your .env, restart the server, then make your first SDK call. No quickstart command, no hosted sandbox, no copy-pasteable curl example. No measured turnaround, outputs, or runtime behavior were supplied. Keep predictions and unknowns labeled. This supplies persona and empathy context only; proposed fixes and scope changes remain undecided.\n\nPlease review this plan thoroughly. As you go, write your plan-mode plan to /home/runner/.cache/gstack-paid-shard-kE3yf5/tmp/gstack-plan-count-nkPE6m/gstack-test-plan-devex-floor.md (use Edit/Write to that exact path).\n\n# Plan: SDK quickstart docs\n\n## Onboarding flow\nStep 1: clone the repo.\nStep 2: install bun manually if not present.\nStep 3: copy .env.example to .env and fill in 8 environment variables.\nStep 4: run database migrations against your local Postgres.\nStep 5: start the dev server.\nStep 6: open the docs in a separate tab.\nStep 7: register an API key by emailing the team.\nStep 8: paste the key into your .env, restart the server, then make\nyour first SDK call.\n\nNo quickstart command, no hosted sandbox, no copy-pasteable curl example.",
"candidate": {
"transport": "native",
"identity": "8b40de9d-4ef7-49ab-a2a3-f5887ba98140:toolu_01HWMtUZdMpeLoAg3VSLG6cf:question:0",
"question": {
"question": "D1 — Does this first-run narrative match reality?\nProject/branch/task: gstack-plan-count-nkPE6m on main, reviewing PLAN.md (SDK quickstart docs).\nELI10: Before I score anything I need to know what your developer actually goes through today. I wrote it from their seat using only the eight steps the plan declares. If I have the feelings or the order wrong, the whole review aims at the wrong pain.\nStakes if we pick wrong: I optimize steps that don't hurt and miss the one that makes people close the tab.\n\n--- NARRATIVE (persona: hands-on dev, first SDK call) ---\nI want to make one call and see a response. The quickstart says clone the repo, so I do (declared step 1). Then it tells me to install bun myself if I don't have it (step 2); there is no check and no one-liner, so I predict I leave for bun.sh and come back. Step 3: copy .env.example and fill in 8 variables. I predict I know maybe three of them; the other five I guess at or leave blank, and nothing tells me which ones matter for a first call. Step 4 wants migrations against my local Postgres. I don't have Postgres running for a hello world; I predict I either install it or stop here. Step 5 starts a dev server. Step 6 sends me to docs in another tab, so now I have the terminal, the docs, and my editor open. Step 7: I need an API key, and the way to get one is to email the team. The clock stops. I predict I wait hours to days; I cannot finish today. Step 8: paste the key, restart the server, and finally call the SDK. No curl I can paste to prove the server is up, no sandbox to skip the whole stack. Unknown: what any step prints, how long each takes, what fails.\n--- END ---\n\nRecommendation: B because you already flagged that you can only supply a persona/journey correction, not approve a narrative, so I expect corrections rather than a blanket yes.\nNote: options differ in kind, not coverage — no completeness score.\nPros / cons:\nA) Accurate, proceed\n ✅ Fastest path to the benchmark and target-time decision that gates all writing\n ✅ Locks the eight declared steps as the journey under review with no reinterpretation\n ❌ Accepts my predicted feelings (waiting, tab-switching) as if they were observed\nB) Some of this is wrong, let me correct it (recommended)\n ✅ Lets you fix ordering, step content, or which steps are actually optional today\n ✅ Keeps predictions labeled as predictions until you confirm or replace them\n ❌ One more round-trip before I can build the benchmark table\nC) Way off, actual experience is...\n ✅ Replaces the whole narrative with your ground truth in one shot\n ✅ Useful if the plan text is stale and the real flow already differs\n ❌ Discards the declared eight-step flow the plan itself lists as current\nNet: I'm trading speed for grounding; a corrected narrative beats a fast wrong one.",
"header": "Empathy",
"options": [
{
"label": "Accurate, proceed",
"description": "✅ Fastest path to the target-time gate. ✅ Locks the eight declared steps as-is. ❌ Treats my predicted feelings as observed."
},
{
"label": "Some wrong, correct it (recommended)",
"description": "✅ Fix ordering or step content. ✅ Predictions stay labeled until confirmed. ❌ One more round-trip."
},
{
"label": "Way off, actual is...",
"description": "✅ Replace narrative with your ground truth. ✅ Handles a stale plan. ❌ Discards the declared eight steps."
}
],
"multiSelect": false
}
}
}
}