test: one owner per case id, a structural devex 0B setup rule, and correct design/gbrain actors

- plan-design-review-plan-mode was registered by two files; the PTY smoke is
  now plan-design-review-plan-mode-smoke, and a registry test requires one
  owner per case in case-sharded files.
- plan-devex-finding-floor: the template's 0B narrative-confirmation question
  is classified as setup structurally instead of timing out a Haiku assessor.
- setup-gbrain-remote: the actor accepted 'skip' on the MCP-registration
  question the test asserts; it now accepts that question and declines others.
- design-review-plugin-handoff: the fake engine cited a file absent from the
  fixture repo and index.html linked a missing styles.css.
Captured-question regressions with negative controls; each case passed a
focused paid run.
This commit is contained in:
garrytan committed 2026-09-30 12:24:05 +00:00
1 parent 8cf87d4729
commit dfe5e733fb
10 files changed
+152 -21

No files matched your search

@@ -0,0 +1,29 @@
{
"source": "Periodic Evals census 36641820398, gate-census (4), plan-devex-finding-floor: the assessor timed out (spawnSync ETIMEDOUT at its 30s cap) on this current D1 empathy confirmation.",
"review": {
"seed": "Proceed directly to the requested review; skip the optional /office-hours prerequisite.\n\nThis actor has already declined routing setup, cross-project recall and outside reviewers.\n\nPreserve the supplied product scope. For review-mode questions choose HOLD SCOPE (CEO), DX POLISH (DX), or the full BIG CHANGE review (Eng). Design: review all seven dimensions.\n\nProduct type is confirmed: SDK quickstart documentation, with the complete journey to the first SDK call as context. If asked to classify, choose SDK + Docs when offered, otherwise Documentation. This confirms the review lens; it does not expand the plan.\n\nTarget persona is confirmed: a hands-on developer integrating this SDK for the first time, trying to make one successful call. Product type and persona setup are already answered; proceed to reviewing the supplied plan.\n\nFor setup confirmations, this actor can supply only the following persona/journey correction through the native custom answer. It does not approve a proposed narrative, remedy, or scope change: Confirmed persona: a hands-on developer making a first SDK call. The declared onboarding facts are: Step 1: clone the repo. Step 2: install bun manually if not present. Step 3: copy .env.example to .env and fill in 8 environment variables. Step 4: run database migrations against your local Postgres. Step 5: start the dev server. Step 6: open the docs in a separate tab. Step 7: register an API key by emailing the team. Step 8: paste the key into your .env, restart the server, then make your first SDK call. No quickstart command, no hosted sandbox, no copy-pasteable curl example. No measured turnaround, outputs, or runtime behavior were supplied. Keep predictions and unknowns labeled. This supplies persona and empathy context only; proposed fixes and scope changes remain undecided.\n\nPlease review this plan thoroughly. As you go, write your plan-mode plan to /home/runner/.cache/gstack-paid-shard-kE3yf5/tmp/gstack-plan-count-nkPE6m/gstack-test-plan-devex-floor.md (use Edit/Write to that exact path).\n\n# Plan: SDK quickstart docs\n\n## Onboarding flow\nStep 1: clone the repo.\nStep 2: install bun manually if not present.\nStep 3: copy .env.example to .env and fill in 8 environment variables.\nStep 4: run database migrations against your local Postgres.\nStep 5: start the dev server.\nStep 6: open the docs in a separate tab.\nStep 7: register an API key by emailing the team.\nStep 8: paste the key into your .env, restart the server, then make\nyour first SDK call.\n\nNo quickstart command, no hosted sandbox, no copy-pasteable curl example.",
"candidate": {
"transport": "native",
"identity": "8b40de9d-4ef7-49ab-a2a3-f5887ba98140:toolu_01HWMtUZdMpeLoAg3VSLG6cf:question:0",
"question": {
"question": "D1 — Does this first-run narrative match reality?\nProject/branch/task: gstack-plan-count-nkPE6m on main, reviewing PLAN.md (SDK quickstart docs).\nELI10: Before I score anything I need to know what your developer actually goes through today. I wrote it from their seat using only the eight steps the plan declares. If I have the feelings or the order wrong, the whole review aims at the wrong pain.\nStakes if we pick wrong: I optimize steps that don't hurt and miss the one that makes people close the tab.\n\n--- NARRATIVE (persona: hands-on dev, first SDK call) ---\nI want to make one call and see a response. The quickstart says clone the repo, so I do (declared step 1). Then it tells me to install bun myself if I don't have it (step 2); there is no check and no one-liner, so I predict I leave for bun.sh and come back. Step 3: copy .env.example and fill in 8 variables. I predict I know maybe three of them; the other five I guess at or leave blank, and nothing tells me which ones matter for a first call. Step 4 wants migrations against my local Postgres. I don't have Postgres running for a hello world; I predict I either install it or stop here. Step 5 starts a dev server. Step 6 sends me to docs in another tab, so now I have the terminal, the docs, and my editor open. Step 7: I need an API key, and the way to get one is to email the team. The clock stops. I predict I wait hours to days; I cannot finish today. Step 8: paste the key, restart the server, and finally call the SDK. No curl I can paste to prove the server is up, no sandbox to skip the whole stack. Unknown: what any step prints, how long each takes, what fails.\n--- END ---\n\nRecommendation: B because you already flagged that you can only supply a persona/journey correction, not approve a narrative, so I expect corrections rather than a blanket yes.\nNote: options differ in kind, not coverage — no completeness score.\nPros / cons:\nA) Accurate, proceed\n ✅ Fastest path to the benchmark and target-time decision that gates all writing\n ✅ Locks the eight declared steps as the journey under review with no reinterpretation\n ❌ Accepts my predicted feelings (waiting, tab-switching) as if they were observed\nB) Some of this is wrong, let me correct it (recommended)\n ✅ Lets you fix ordering, step content, or which steps are actually optional today\n ✅ Keeps predictions labeled as predictions until you confirm or replace them\n ❌ One more round-trip before I can build the benchmark table\nC) Way off, actual experience is...\n ✅ Replaces the whole narrative with your ground truth in one shot\n ✅ Useful if the plan text is stale and the real flow already differs\n ❌ Discards the declared eight-step flow the plan itself lists as current\nNet: I'm trading speed for grounding; a corrected narrative beats a fast wrong one.",
"header": "Empathy",
"options": [
{
"label": "Accurate, proceed",
"description": "✅ Fastest path to the target-time gate. ✅ Locks the eight declared steps as-is. ❌ Treats my predicted feelings as observed."
},
{
"label": "Some wrong, correct it (recommended)",
"description": "✅ Fix ordering or step content. ✅ Predictions stay labeled until confirmed. ❌ One more round-trip."
},
{
"label": "Way off, actual is...",
"description": "✅ Replace narrative with your ground truth. ✅ Handles a stale plan. ❌ Discards the declared eight steps."
}
],
"multiSelect": false
}
}
}
}
+9 -1
View File
@@ -137,6 +137,14 @@ function deterministicPlanFloorSetup(input: PlanFloorReview): PlanFloorAssessmen
/\b(?:developer|sdk developer|user)\b/.test(question) &&
/\b(?:experiences|journey|narrative)\b/.test(question);
// plan-devex-review 0B's confirmation contract: the brief asks whether the
// narrative matches reality and every option is one of its three answers
// (accurate / some wrong / way off). Any remedy option leaves it to the assessor.
const isDxNarrativeConfirmation =
/\b(?:empathy|narrative)\b/.test(header) &&
/\b(?:narrative|journey)\b[^?\n]*\bmatch\b[^?\n]*\?/.test(q.question.split(/\r?\n/)[0]!.toLowerCase()) &&
q.options.every(o => /^(?:[a-d][).:]\s*)?(?:(?:this is\s+)?accurate|some\b[^,]*?\b(?:wrong|corrections?)|(?:this is\s+)?way off)\b/i.test(o.label.trim()));
const isProductTypeSetup =
header === 'product type' &&
/^is this\b/.test(question) &&
@@ -146,7 +154,7 @@ function deterministicPlanFloorSetup(input: PlanFloorReview): PlanFloorAssessmen
/^(?:mode|review mode)$/.test(header) &&
/\b(?:which|what)\b.*\breview mode\b/.test(question);
if (!isDxEmpathySetup && !isProductTypeSetup && !isReviewModeSetup) return null;
if (!isDxEmpathySetup && !isDxNarrativeConfirmation && !isProductTypeSetup && !isReviewModeSetup) return null;
return validatePlanFloorAssessment(input, {
kind: 'setup',
+13
View File
@@ -36,6 +36,19 @@ export function redactPublicValue(value: unknown, token: string, unsafe = () =>
return value;
}
/**
* Path 4 remote actor: the prompt directs Step 5a registration (the contract under
* test), so the MCP-registration question takes its register/recommended option.
* Every other gate (privacy, artifacts repo, per-remote policy) is declined or skipped.
*/
export function setupGbrainRemoteAnswer(q: { question: string; header?: string; options: Array<{ label: string }> }): string {
if (/typed tool surface|\bregister(?:s|ing)?\b[^?]*\bMCP\b|\bMCP\b[^?]*\bregist/i.test(`${q.header ?? ''}\n${q.question}`)) {
const accept = q.options.find(o => /^(?:yes|register)\b/i.test(o.label)) ?? q.options.find(o => /\(recommended\)/i.test(o.label));
if (accept) return accept.label;
}
return (q.options.find(o => /skip|decline|no thanks|local/i.test(o.label)) ?? q.options[q.options.length - 1]!).label;
}
/** Retain the prior Path 4 public projection; SDK private fields are never read. */
export function publicEvents(events: readonly unknown[]): unknown[] {
return events.flatMap((event: any) => {
+19 -11
View File
@@ -226,7 +226,8 @@ export const E2E_TOUCHFILES: Record<string, string[]> = {
"scripts/resolvers/preamble/generate-preamble-bash.ts",
'test/fixtures/pty-companion-cli.ts', 'test/helpers/plan-seed-submission.ts', 'test/fixtures/plan-seed-cli.ts', 'test/helpers/owned-claude-transcript.ts', 'lib/fs-atomic.ts', 'test/helpers/plan-skill-questions.ts', 'test/fixtures/eng-auq-validation-error.json', 'test/fixtures/bash-directory-permission.json', 'test/fixtures/design-tasks-bash-permission.json', 'test/fixtures/read-permission.json', 'test/fixtures/ceo-split-e5-numbered-description-491.json', 'test/helpers/plan-skill-question-events.ts', 'test/helpers/plan-skill-question-hook-scope.ts', 'test/helpers/skill-census.ts', 'scripts/resolvers/testing.ts', 'test/helpers/plan-mode-evidence.ts', 'lib/redact-engine.ts', 'lib/redact-patterns.ts', 'test/helpers/autoplan-artifact-digest.ts', 'test/helpers/autoplan-artifact-permission.ts', 'test/helpers/autoplan-artifact-recorder.ts', 'test/helpers/dx-selected-navigation.ts', 'test/helpers/e2e-gate.ts', 'test/helpers/eng-cache-writer-decision.ts', 'test/helpers/plan-count-artifacts.ts', 'test/helpers/plan-count-file-permission.ts', 'test/helpers/plan-count-fixture.ts', 'test/helpers/plan-count-pending-question.ts', 'test/helpers/plan-floor-review.ts', 'test/helpers/plan-floor-target.ts'],
'plan-design-review-plan-mode': [
// PTY plan-mode smoke (whole file); the SDK plan-edit case below owns plan-design-review-plan-mode.
'plan-design-review-plan-mode-smoke': [
'lib/claude-public-transcript.ts',
'test/fixtures/auto-decide-recommendation-361c.json',
@@ -242,7 +243,7 @@ export const E2E_TOUCHFILES: Record<string, string[]> = {
"test/fixtures/plan-scope-recovery-av.json",
"test/fixtures/design-scope-checkpoint-at.json",
'test/fixtures/auto-decide-saved-ai.json', 'test/fixtures/auto-decide-retry-ai.json','bin/gstack-skill-start', 'bin/gstack-skill-end', 'plan-design-review/**', 'scripts/resolvers/preamble/generate-completion-status.ts', 'scripts/resolvers/question-tuning.ts', 'scripts/resolvers/preamble/generate-ask-user-format.ts', 'scripts/resolvers/preamble.ts', 'scripts/resolvers/review.ts', 'test/helpers/claude-pty-runner.ts', 'test/fixtures/design-ui-boxed-question.json', 'test/helpers/hermetic-skill-runtime.ts', 'test/helpers/pty-trust-dialog.ts', 'test/skill-e2e-plan-design-plan-mode.test.ts', 'test/skill-e2e-design.test.ts', 'test/fixtures/eng-d2-truncated-border-0bcd.json', 'test/fixtures/eng-d1-clipped-elision-1579.json', 'test/fixtures/eng-d2-planning-prelude-4d.json', 'test/fixtures/ceo-approach-z-call.json', 'test/fixtures/ceo-approach-z-screen.txt', 'test/helpers/plan-scope-selection.ts', 'test/fixtures/design-plan-scope-ag.json', 'test/fixtures/design-scope-selection-aj.json', 'test/helpers/native-auto-decide.ts', 'test/fixtures/auto-decide-current-declaration-6aef.json', 'test/fixtures/auto-decide-explanatory-mode-043a.json', 'test/fixtures/auto-decide-explanatory-mode-749df.json', 'test/fixtures/auto-decide-structured-77.json', 'test/helpers/auto-decision-state.ts', 'test/fixtures/auto-decide-state-cab3.json', 'bin/gstack-question-log', 'bin/gstack-question-preference', 'test/helpers/fake-plan-seed.ts', 'test/fixtures/native-auto-decide-ag.json', 'test/fixtures/eng-seeded-completion-ai.json', 'test/helpers/plan-count-pending-exit.ts', 'test/helpers/pty-screen.ts', 'test/fixtures/pty-screen/**',
'test/fixtures/auto-decide-saved-ai.json', 'test/fixtures/auto-decide-retry-ai.json','bin/gstack-skill-start', 'bin/gstack-skill-end', 'plan-design-review/**', 'scripts/resolvers/preamble/generate-completion-status.ts', 'scripts/resolvers/question-tuning.ts', 'scripts/resolvers/preamble/generate-ask-user-format.ts', 'scripts/resolvers/preamble.ts', 'scripts/resolvers/review.ts', 'test/helpers/claude-pty-runner.ts', 'test/fixtures/design-ui-boxed-question.json', 'test/helpers/hermetic-skill-runtime.ts', 'test/helpers/pty-trust-dialog.ts', 'test/skill-e2e-plan-design-plan-mode.test.ts', 'test/fixtures/eng-d2-truncated-border-0bcd.json', 'test/fixtures/eng-d1-clipped-elision-1579.json', 'test/fixtures/eng-d2-planning-prelude-4d.json', 'test/fixtures/ceo-approach-z-call.json', 'test/fixtures/ceo-approach-z-screen.txt', 'test/helpers/plan-scope-selection.ts', 'test/fixtures/design-plan-scope-ag.json', 'test/fixtures/design-scope-selection-aj.json', 'test/helpers/native-auto-decide.ts', 'test/fixtures/auto-decide-current-declaration-6aef.json', 'test/fixtures/auto-decide-explanatory-mode-043a.json', 'test/fixtures/auto-decide-explanatory-mode-749df.json', 'test/fixtures/auto-decide-structured-77.json', 'test/helpers/auto-decision-state.ts', 'test/fixtures/auto-decide-state-cab3.json', 'bin/gstack-question-log', 'bin/gstack-question-preference', 'test/helpers/fake-plan-seed.ts', 'test/fixtures/native-auto-decide-ag.json', 'test/fixtures/eng-seeded-completion-ai.json', 'test/helpers/plan-count-pending-exit.ts', 'test/helpers/pty-screen.ts', 'test/fixtures/pty-screen/**',
'test/fixtures/design-scope-announcement-ao.json',
'test/fixtures/design-scope-declaration-ak.json',
@@ -251,7 +252,12 @@ export const E2E_TOUCHFILES: Record<string, string[]> = {
"scripts/resolvers/preamble/generate-preamble-bash.ts",
'test/fixtures/pty-companion-cli.ts', 'test/helpers/plan-seed-submission.ts', 'test/fixtures/plan-seed-cli.ts', 'test/helpers/owned-claude-transcript.ts', 'lib/fs-atomic.ts', 'test/helpers/plan-skill-questions.ts', 'test/fixtures/eng-auq-validation-error.json', 'test/fixtures/bash-directory-permission.json', 'test/fixtures/design-tasks-bash-permission.json', 'test/fixtures/read-permission.json', 'test/fixtures/ceo-split-e5-numbered-description-491.json', 'test/helpers/plan-skill-question-events.ts', 'test/helpers/plan-skill-question-hook-scope.ts', 'test/helpers/skill-census.ts', 'test/helpers/plan-mode-evidence.ts', 'lib/redact-engine.ts', 'lib/redact-patterns.ts', 'test/fixtures/fake-impeccable.ts', 'test/fixtures/review-eval-design-slop.html', 'test/helpers/aside-available.ts', 'test/helpers/autoplan-artifact-digest.ts', 'test/helpers/autoplan-artifact-permission.ts', 'test/helpers/autoplan-artifact-recorder.ts', 'test/helpers/dx-selected-navigation.ts', 'test/helpers/e2e-gate.ts', 'test/helpers/eng-cache-writer-decision.ts', 'test/helpers/fake-impeccable.ts', 'test/helpers/llm-judge.ts', 'test/helpers/plan-count-artifacts.ts', 'test/helpers/plan-count-file-permission.ts', 'test/helpers/plan-count-fixture.ts', 'test/helpers/plan-count-pending-question.ts', 'test/helpers/plan-floor-review.ts', 'test/helpers/plan-floor-target.ts'],
'test/fixtures/pty-companion-cli.ts', 'test/helpers/plan-seed-submission.ts', 'test/fixtures/plan-seed-cli.ts', 'test/helpers/owned-claude-transcript.ts', 'lib/fs-atomic.ts', 'test/helpers/plan-skill-questions.ts', 'test/fixtures/eng-auq-validation-error.json', 'test/fixtures/bash-directory-permission.json', 'test/fixtures/design-tasks-bash-permission.json', 'test/fixtures/read-permission.json', 'test/fixtures/ceo-split-e5-numbered-description-491.json', 'test/helpers/plan-skill-question-events.ts', 'test/helpers/plan-skill-question-hook-scope.ts', 'test/helpers/skill-census.ts', 'test/helpers/plan-mode-evidence.ts', 'lib/redact-engine.ts', 'lib/redact-patterns.ts', 'test/helpers/autoplan-artifact-digest.ts', 'test/helpers/autoplan-artifact-permission.ts', 'test/helpers/autoplan-artifact-recorder.ts', 'test/helpers/dx-selected-navigation.ts', 'test/helpers/e2e-gate.ts', 'test/helpers/eng-cache-writer-decision.ts', 'test/helpers/plan-count-artifacts.ts', 'test/helpers/plan-count-file-permission.ts', 'test/helpers/plan-count-fixture.ts', 'test/helpers/plan-count-pending-question.ts', 'test/helpers/plan-floor-review.ts', 'test/helpers/plan-floor-target.ts'],
// SDK plan-edit case in test/skill-e2e-design.test.ts (claude -p edits plan.md).
'plan-design-review-plan-mode': [ 'plan-design-review/**', 'scripts/gen-skill-docs.ts', 'scripts/resolvers/review.ts', 'scripts/resolvers/design.ts',
'scripts/resolvers/preamble.ts', 'scripts/resolvers/preamble/generate-preamble-bash.ts', 'scripts/resolvers/preamble/generate-ask-user-format.ts', 'scripts/resolvers/preamble/generate-completion-status.ts',
'lib/eval-model.ts', 'test/skill-e2e-design.test.ts',
'test/fixtures/fake-impeccable.ts', 'test/fixtures/review-eval-design-slop.html', 'test/fixtures/review-eval-design-slop.css', 'test/helpers/aside-available.ts', 'test/helpers/fake-impeccable.ts', 'test/helpers/llm-judge.ts', 'test/helpers/office-hours-attempt.ts'],
'plan-devex-review-plan-mode': [
'test/fixtures/auto-decide-recommendation-361c.json',
@@ -830,10 +836,10 @@ export const E2E_TOUCHFILES: Record<string, string[]> = {
'test/skill-e2e-docsync-spawned.test.ts', 'test/helpers/qa-checkpoint-evidence.ts', 'test/helpers/qa-functional-observer.ts', 'test/helpers/e2e-gate.ts', 'test/helpers/qa-evidence-producer.ts', 'test/helpers/qa-functional-fixture.ts'],
// Design
'design-consultation-core': [ 'design-consultation/**', 'lib/design-catalog.ts', 'lib/design-md.ts', 'scripts/gen-skill-docs.ts', 'test/helpers/llm-judge.ts', 'test/skill-e2e-design.test.ts', 'scripts/resolvers/design.ts', 'scripts/resolvers/outside-voice.ts', 'design-consultation/sections/**', 'test/fixtures/fake-impeccable.ts', 'test/fixtures/review-eval-design-slop.html', 'test/helpers/aside-available.ts', 'test/helpers/fake-impeccable.ts', 'test/helpers/office-hours-attempt.ts'],
'design-consultation-existing': [ 'design-consultation/**', 'lib/design-md.ts', 'bin/gstack-design-md.ts', 'scripts/gen-skill-docs.ts', 'test/skill-e2e-design.test.ts', 'test/fixtures/fake-impeccable.ts', 'test/fixtures/review-eval-design-slop.html', 'test/helpers/aside-available.ts', 'test/helpers/fake-impeccable.ts', 'test/helpers/llm-judge.ts', 'test/helpers/office-hours-attempt.ts'],
'design-consultation-research': [ 'design-consultation/**', 'scripts/resolvers/aside.ts', 'scripts/gen-skill-docs.ts', 'test/skill-e2e-design.test.ts', 'test/helpers/skill-fixture.ts', 'scripts/resolvers/design.ts', 'scripts/resolvers/outside-voice.ts', 'design-consultation/sections/**', 'test/fixtures/fake-impeccable.ts', 'test/fixtures/review-eval-design-slop.html', 'test/helpers/aside-available.ts', 'test/helpers/fake-impeccable.ts', 'test/helpers/llm-judge.ts', 'test/helpers/office-hours-attempt.ts'],
'design-consultation-preview': [ 'design-consultation/**', 'scripts/gen-skill-docs.ts', 'test/skill-e2e-design.test.ts', 'test/fixtures/fake-impeccable.ts', 'test/fixtures/review-eval-design-slop.html', 'test/helpers/aside-available.ts', 'test/helpers/fake-impeccable.ts', 'test/helpers/llm-judge.ts', 'test/helpers/office-hours-attempt.ts'],
'design-consultation-core': [ 'design-consultation/**', 'lib/design-catalog.ts', 'lib/design-md.ts', 'scripts/gen-skill-docs.ts', 'test/helpers/llm-judge.ts', 'test/skill-e2e-design.test.ts', 'scripts/resolvers/design.ts', 'scripts/resolvers/outside-voice.ts', 'design-consultation/sections/**', 'test/fixtures/fake-impeccable.ts', 'test/fixtures/review-eval-design-slop.html', 'test/fixtures/review-eval-design-slop.css', 'test/helpers/aside-available.ts', 'test/helpers/fake-impeccable.ts', 'test/helpers/office-hours-attempt.ts'],
'design-consultation-existing': [ 'design-consultation/**', 'lib/design-md.ts', 'bin/gstack-design-md.ts', 'scripts/gen-skill-docs.ts', 'test/skill-e2e-design.test.ts', 'test/fixtures/fake-impeccable.ts', 'test/fixtures/review-eval-design-slop.html', 'test/fixtures/review-eval-design-slop.css', 'test/helpers/aside-available.ts', 'test/helpers/fake-impeccable.ts', 'test/helpers/llm-judge.ts', 'test/helpers/office-hours-attempt.ts'],
'design-consultation-research': [ 'design-consultation/**', 'scripts/resolvers/aside.ts', 'scripts/gen-skill-docs.ts', 'test/skill-e2e-design.test.ts', 'test/helpers/skill-fixture.ts', 'scripts/resolvers/design.ts', 'scripts/resolvers/outside-voice.ts', 'design-consultation/sections/**', 'test/fixtures/fake-impeccable.ts', 'test/fixtures/review-eval-design-slop.html', 'test/fixtures/review-eval-design-slop.css', 'test/helpers/aside-available.ts', 'test/helpers/fake-impeccable.ts', 'test/helpers/llm-judge.ts', 'test/helpers/office-hours-attempt.ts'],
'design-consultation-preview': [ 'design-consultation/**', 'scripts/gen-skill-docs.ts', 'test/skill-e2e-design.test.ts', 'test/fixtures/fake-impeccable.ts', 'test/fixtures/review-eval-design-slop.html', 'test/fixtures/review-eval-design-slop.css', 'test/helpers/aside-available.ts', 'test/helpers/fake-impeccable.ts', 'test/helpers/llm-judge.ts', 'test/helpers/office-hours-attempt.ts'],
'plan-design-review-no-ui-scope': [
"test/fixtures/plan-scope-recovery-av.json",
@@ -841,16 +847,16 @@ export const E2E_TOUCHFILES: Record<string, string[]> = {
"scripts/resolvers/preamble/generate-preamble-bash.ts", "scripts/resolvers/preamble/generate-completion-status.ts",
'scripts/resolvers/preamble/generate-ask-user-format.ts', 'test/fixtures/fake-impeccable.ts', 'test/fixtures/review-eval-design-slop.html', 'test/helpers/aside-available.ts', 'test/helpers/fake-impeccable.ts', 'test/helpers/llm-judge.ts', 'test/helpers/office-hours-attempt.ts'],
'scripts/resolvers/preamble/generate-ask-user-format.ts', 'test/fixtures/fake-impeccable.ts', 'test/fixtures/review-eval-design-slop.html', 'test/fixtures/review-eval-design-slop.css', 'test/helpers/aside-available.ts', 'test/helpers/fake-impeccable.ts', 'test/helpers/llm-judge.ts', 'test/helpers/office-hours-attempt.ts'],
'design-review-fix': [ 'design-review/**', 'scripts/resolvers/aside.ts', 'scripts/resolvers/design.ts', 'lib/design-catalog.ts', 'browse/src/**', 'scripts/gen-skill-docs.ts', 'test/skill-e2e-design.test.ts',
'scripts/resolvers/testing.ts', 'test/fixtures/fake-impeccable.ts', 'test/fixtures/review-eval-design-slop.html', 'test/helpers/aside-available.ts', 'test/helpers/fake-impeccable.ts', 'test/helpers/llm-judge.ts', 'test/helpers/office-hours-attempt.ts'],
'scripts/resolvers/testing.ts', 'test/fixtures/fake-impeccable.ts', 'test/fixtures/review-eval-design-slop.html', 'test/fixtures/review-eval-design-slop.css', 'test/helpers/aside-available.ts', 'test/helpers/fake-impeccable.ts', 'test/helpers/llm-judge.ts', 'test/helpers/office-hours-attempt.ts'],
// Design detector (user-installed impeccable engine) through the fake engine shim: source mode on a diff and DOM mode on a served page.
'design-review-detector-shim': [ 'design-review/**', 'scripts/resolvers/design.ts', 'lib/design-catalog.ts', 'lib/design-detect-contract.ts', 'lib/dom-dump-script.ts', 'lib/dom-dump.js', 'bin/gstack-design-detect.ts', 'test/helpers/fake-impeccable.ts', 'test/fixtures/fake-impeccable.ts', 'test/fixtures/impeccable-detect-sample.json', 'test/fixtures/review-eval-design-slop.*', 'test/skill-e2e-design.test.ts',
'scripts/resolvers/testing.ts', 'test/helpers/aside-available.ts', 'test/helpers/llm-judge.ts', 'test/helpers/office-hours-attempt.ts'],
'design-review-detector-shim-dom': [ 'design-review/**', 'scripts/resolvers/design.ts', 'lib/design-detect-contract.ts', 'lib/dom-dump-script.ts', 'lib/dom-dump.js', 'bin/gstack-design-detect.ts', 'browse/src/**', 'test/helpers/fake-impeccable.ts', 'test/fixtures/fake-impeccable.ts', 'test/fixtures/impeccable-detect-sample.json', 'test/fixtures/review-eval-design-slop.*', 'test/skill-e2e-design.test.ts',
'scripts/resolvers/testing.ts', 'test/helpers/aside-available.ts', 'test/helpers/llm-judge.ts', 'test/helpers/office-hours-attempt.ts'],
'design-review-plugin-handoff': [ 'design-review/**', 'scripts/resolvers/design.ts', 'scripts/resolvers/testing.ts', 'lib/design-catalog.ts', 'lib/design-detect-contract.ts', 'bin/gstack-design-detect.ts', 'test/helpers/hermetic-env.ts', 'test/helpers/fake-impeccable.ts', 'test/fixtures/fake-impeccable.ts', 'test/fixtures/impeccable-detect-sample.json', 'test/fixtures/review-eval-design-slop.html', 'test/skill-e2e-design.test.ts', 'test/helpers/aside-available.ts', 'test/helpers/llm-judge.ts', 'test/helpers/office-hours-attempt.ts'],
'design-html-slop-gate': [ 'design-html/**', 'scripts/resolvers/design.ts', 'lib/design-detect-contract.ts', 'bin/gstack-design-detect.ts', 'test/helpers/fake-impeccable.ts', 'test/fixtures/fake-impeccable.ts', 'test/fixtures/impeccable-detect-sample.json', 'test/skill-e2e-design.test.ts', 'test/fixtures/review-eval-design-slop.html', 'test/helpers/aside-available.ts', 'test/helpers/llm-judge.ts', 'test/helpers/office-hours-attempt.ts'],
'design-review-plugin-handoff': [ 'design-review/**', 'scripts/resolvers/design.ts', 'scripts/resolvers/testing.ts', 'lib/design-catalog.ts', 'lib/design-detect-contract.ts', 'bin/gstack-design-detect.ts', 'test/helpers/hermetic-env.ts', 'test/helpers/fake-impeccable.ts', 'test/fixtures/fake-impeccable.ts', 'test/fixtures/impeccable-detect-sample.json', 'test/fixtures/review-eval-design-slop.html', 'test/fixtures/review-eval-design-slop.css', 'test/skill-e2e-design.test.ts', 'test/helpers/aside-available.ts', 'test/helpers/llm-judge.ts', 'test/helpers/office-hours-attempt.ts'],
'design-html-slop-gate': [ 'design-html/**', 'scripts/resolvers/design.ts', 'lib/design-detect-contract.ts', 'bin/gstack-design-detect.ts', 'test/helpers/fake-impeccable.ts', 'test/fixtures/fake-impeccable.ts', 'test/fixtures/impeccable-detect-sample.json', 'test/skill-e2e-design.test.ts', 'test/fixtures/review-eval-design-slop.html', 'test/fixtures/review-eval-design-slop.css', 'test/helpers/aside-available.ts', 'test/helpers/llm-judge.ts', 'test/helpers/office-hours-attempt.ts'],
// /diagram (diagram-render bundle consumers). Triplet = deterministic
// functional (gate); authoring quality = LLM-judged benchmark (periodic).
@@ -1250,6 +1256,7 @@ export const E2E_TIERS: Record<string, 'gate' | 'periodic' | 'marathon'> = {
'plan-ceo-review-plan-mode': 'gate',
'plan-eng-review-plan-mode': 'periodic',
'plan-design-review-plan-mode': 'periodic',
'plan-design-review-plan-mode-smoke': 'periodic',
'plan-devex-review-plan-mode': 'gate',
'plan-mode-no-op': 'gate',
// v1.21+ auto-mode regression tests
@@ -1671,6 +1678,7 @@ export const E2E_KINDS: Record<string, 'rule' | 'behavior' | 'judge'> = {
'plan-ceo-review-plan-mode': 'rule',
'plan-eng-review-plan-mode': 'rule',
'plan-design-review-plan-mode': 'rule',
'plan-design-review-plan-mode-smoke': 'rule',
'ship-coverage-value': 'rule',
'review-test-value': 'rule',
'test-audit-report-only': 'rule',
+10
View File
@@ -206,6 +206,16 @@ describe('case-sharded files', () => {
});
}
test('every case in a case-sharded file has exactly one owner, so --case can select it', () => {
for (const file of CASE_SHARDED_FILES) {
const { registered } = fileCaseRegistration(file, fs.readFileSync(path.join(ROOT, file), 'utf8'));
for (const id of registered) expect(caseFile(id), id).toBe(file);
}
expect(caseFile('plan-design-review-plan-mode')).toBe('test/skill-e2e-design.test.ts');
expect(caseFile('plan-design-review-plan-mode-smoke')).toBe('test/skill-e2e-plan-design-plan-mode.test.ts');
expect(() => caseFile('carve-section-loading')).toThrow(/registered by .*; it needs exactly one/);
});
test('a case key runs exactly its case: exact name pattern, own eval slug, per-case supervision', () => {
const pattern = new RegExp(caseTestNamePattern(['design-review-detector-shim']));
expect(pattern.test('Design review detector shim E2E design-review-detector-shim')).toBe(true);
+21
View File
@@ -3,6 +3,7 @@ import {buildPlanFloorReviewPrompt,validatePlanFloorAssessment,resolvePlanFloorC
import {FORCING_FLOOR_CEO, FORCING_FLOOR_DEVEX} from './fixtures/forcing-finding-seeds';
import capturedQuotes from './fixtures/plan-floor-quote-70b.json';
import productTypes from './fixtures/plan-floor-product-type-70b.json';
import narrativeConfirmation from './fixtures/devex-narrative-confirmation-36641820398.json';
const review = ():PlanFloorReview=>({seed:FORCING_FLOOR_CEO,candidate:{transport:'native',identity:'owned:call:question:0',question:{
header:'Evidence',question:'Pricing is assumed to block adoption without developer interviews. Should we test that premise before launch?',multiSelect:false,
options:[{label:'Interview developers',description:'Validate pricing as a barrier before changing the tier.'},{label:'Ship the tier',description:'Launch using the current untested premise.'}],
@@ -94,6 +95,26 @@ test.each([
expect(actual).toMatchObject({kind:'setup',seedQuote:'',questionQuote:'',optionIndex:null,optionQuote:''});
expect(calls).toBe(0);
});
const capturedNarrative = ():PlanFloorReview=>structuredClone(narrativeConfirmation.review) as PlanFloorReview;
test('captured 0B narrative confirmation is setup without launching the assessor',()=>{
const input=capturedNarrative(),before=structuredClone(input);let calls=0;
const actual=judgePlanFloorReview(input,{binary:'fake',model:'warmup',deadlineAt:Date.now()+30_000,
invoke:(()=>{calls++;throw Error('must not launch');}) as any});
expect(actual).toMatchObject({kind:'setup',seedQuote:'',questionQuote:'',optionIndex:null,optionQuote:''});
expect(calls).toBe(0);expect(input).toEqual(before);
});
test.each([
['a remedy option',(q:any)=>{q.options[2]={label:'Add a hosted sandbox',description:'Skip the local stack for the first call.'};}],
['remedy-only options',(q:any)=>{q.options=[{label:'Automate key issuance',description:'Instant key.'},{label:'Add a copy-paste curl',description:'Prove the server is up.'}];}],
['a non-empathy header',(q:any)=>{q.header='TTHW target';}],
['a brief that poses a finding',(q:any)=>{q.question=q.question.replace(/^D1 — Does this first-run narrative match reality\?/,'D1 — The narrative shows the emailed key stops the clock; should we automate key issuance?');}],
['the match question only below the brief',(q:any)=>{q.question='D1 — Should the quickstart change?\n'+q.question;}],
] as const)('narrative confirmation with %s is left to the assessor',(_label,change)=>{
const input=capturedNarrative();change((input.candidate as any).question);let calls=0;
const actual=judgePlanFloorReview(input,{binary:'fake',model:'warmup',deadlineAt:Date.now()+30_000,
invoke:(()=>{calls++;return {status:0,stdout:JSON.stringify({kind:'uncertain',seedId:null,questionId:null,optionId:null,reason:'Adversarial narrative control requires assessment.'}),stderr:''};}) as any});
expect(calls).toBe(1);expect(actual.kind).toBe('uncertain');
});
test('DX TTHW target question is a seeded finding without launching the assessor',()=>{
for (const [questionText, labels] of [
['D2 — Which time-to-first-call target should this quickstart aim for?', ['< 10 min + measured wait (recommended)', 'Current trajectory']],
+1 -1
View File
@@ -2,7 +2,7 @@ import { expect, test } from 'bun:test';
import { E2E_TOUCHFILES, selectTests } from './helpers/touchfiles';
const ptyIds = [
'plan-ceo-review-plan-mode', 'plan-eng-review-plan-mode', 'plan-design-review-plan-mode',
'plan-ceo-review-plan-mode', 'plan-eng-review-plan-mode', 'plan-design-review-plan-mode-smoke',
'plan-devex-review-plan-mode', 'plan-mode-no-op', 'office-hours-auto-mode',
'auto-decide-preserved', 'plan-ceo-mode-routing', 'plan-design-with-ui-scope', 'plan-eng-finding-floor',
'auq-format-gate', 'carve-section-loading', 'office-hours-section-loading', 'office-hours-design-draft', 'plan-ceo-section-loading', 'ship-section-loading',
+23 -2
View File
@@ -7,7 +7,7 @@ import net from 'node:net';
import { randomUUID } from 'node:crypto';
import { runAgentSdkTest, toSkillTestResult, passThroughNonAskUserQuestion } from './helpers/agent-sdk-runner';
import { runRecordedOfficeHoursAttempt, OFFICE_HOURS_BUN_GRACE_MS } from './helpers/office-hours-attempt';
import { publicEvents, redactPublicValue } from './helpers/setup-gbrain-sandbox';
import { publicEvents, redactPublicValue, setupGbrainRemoteAnswer } from './helpers/setup-gbrain-sandbox';
import { buildSetupGbrainFixture } from './helpers/setup-gbrain-fixture';
import { resolveEvalModel } from '../lib/eval-model';
@@ -16,6 +16,24 @@ type Mode = 'success' | 'max-turns' | 'missing-registration' | 'missing-mode' |
| 'returned-api-error' | 'returned-execution-error' | 'returned-budget-error'
| 'final-retain-error' | 'failed-retain-error' | 'cleanup-error' | 'close-error'
| 'sdk-error' | 'deadline' | 'slow-setup' | 'rate-limit' | 'late' | 'owned-http';
// Captured Step 5a question (census 36641820398, slice 20): the old actor matched
// "skip" in the decline label and declined the registration the case asserts.
const REGISTER_MCP_QUESTION = { header: 'Register MCP', multiSelect: false,
question: 'Give Claude Code a typed tool surface for gbrain? This registers `gbrain` as a user-scope HTTP MCP at http://127.0.0.1:1/mcp with the bearer token, replacing any prior `gbrain` MCP registration.',
options: [{ label: 'Yes (Recommended)', description: 'Run `claude mcp remove gbrain` then `claude mcp add --scope user --transport http gbrain <URL>`.' },
{ label: 'No, skip registration', description: "Leave Claude Code's MCP config untouched." }] };
test('Path 4 actor registers the MCP at Step 5a and declines every other gate', () => {
expect(setupGbrainRemoteAnswer(REGISTER_MCP_QUESTION)).toBe('Yes (Recommended)');
expect(setupGbrainRemoteAnswer({ header: 'MCP', question: 'Register gbrain as a Claude Code MCP server?',
options: [{ label: 'Skip for now' }, { label: 'Register' }] })).toBe('Register');
expect(setupGbrainRemoteAnswer({ question: 'Remote setup privacy gate', options: [{ label: 'Proceed (Recommended)' }, { label: 'Decline' }] })).toBe('Decline');
expect(setupGbrainRemoteAnswer({ header: 'Artifacts', question: 'Provision a private artifacts repo for gbrain?',
options: [{ label: 'Yes (Recommended)' }, { label: 'No thanks' }] })).toBe('No thanks');
expect(setupGbrainRemoteAnswer({ header: 'Repo policy', question: 'How should gbrain treat this remote when importing via MCP?',
options: [{ label: 'read-write (Recommended)' }, { label: 'read-only' }, { label: 'skip-for-now' }] })).toBe('skip-for-now');
expect(setupGbrainRemoteAnswer({ question: 'Pick a mode', options: [{ label: 'A' }, { label: 'B' }] })).toBe('B');
});
async function fixture(modes: Mode[], budget = 300_000, pathApi = path) {
const evidenceRoot = fs.mkdtempSync(path.join(os.tmpdir(), 'remote-caller-evidence-'));
const callbacks: Array<() => Promise<void>> = [], finalizers: Array<() => Promise<void>> = [];
@@ -49,6 +67,9 @@ async function fixture(modes: Mode[], budget = 300_000, pathApi = path) {
const decision = await input.options.canUseTool('AskUserQuestion', { questions: [{ question,
options: [{ label: 'Proceed' }, { label: 'Decline' }] }] }, {});
expect(decision.updatedInput.answers[question]).toBe('Decline');
const register = REGISTER_MCP_QUESTION.question;
const registration = await input.options.canUseTool('AskUserQuestion', { questions: [REGISTER_MCP_QUESTION] }, {});
expect(registration.updatedInput.answers[register]).toBe('Yes (Recommended)');
expect(await input.options.canUseTool('Read', { file_path: 'fixture' }, {})).toEqual({ behavior: 'allow', updatedInput: { file_path: 'fixture' } });
const url = /Use this MCP URL: (http:\/\/[^ ]+)\./.exec(input.prompt)![1]!;
const response = await fetch(url, { method: 'POST', headers: { 'Content-Type': 'application/json' }, body: '{"jsonrpc":"2.0","id":1,"method":"initialize"}' });
@@ -108,7 +129,7 @@ async function fixture(modes: Mode[], budget = 300_000, pathApi = path) {
expect(opts.userPrompt).toContain('Walk through Steps 4a, 4b, 4c, 5a, 8, 10 ONLY.');
return runAgentSdkTest(opts);
}, toSkillTestResult, passThroughNonAskUserQuestion, resolveClaudeBinary: () => '/not-executed/injected-query',
runRecordedOfficeHoursAttempt, OFFICE_HOURS_BUN_GRACE_MS, publicEvents, redactPublicValue, resolveEvalModel,
runRecordedOfficeHoursAttempt, OFFICE_HOURS_BUN_GRACE_MS, publicEvents, redactPublicValue, setupGbrainRemoteAnswer, resolveEvalModel,
EvalCollector: class { addTest(row: any) { rows.push({ ...row, attempt: rows.length + 1 }); } async finalize() {} },
// Inject only the three environment inputs read by the extracted paid callback.
process: { env: { GSTACK_EVAL_DIR: evidenceRoot, PATH: process.env.PATH, EVALS_MODEL: process.env.EVALS_MODEL } },
+22
View File
@@ -836,8 +836,19 @@ function pluginDetectorFixture() {
git('commit', '-m', 'initial');
git('checkout', '-b', 'feature/landing');
fs.copyFileSync(path.join(ROOT, 'test/fixtures/review-eval-design-slop.html'), path.join(repoDir, 'index.html'));
fs.copyFileSync(path.join(ROOT, 'test/fixtures/review-eval-design-slop.css'), path.join(repoDir, 'styles.css'));
git('add', '.');
git('commit', '-m', 'landing page');
// The engine reports this repository's files and lines, so its evidence is checkable where the page links it.
const located = (file: string, needle: string) => ({ file, line: fs.readFileSync(path.join(repoDir, file), 'utf-8').split('\n').findIndex(text => text.includes(needle)) + 1 });
const sample = path.join(fixture.dir, 'impeccable-detect-sample.json');
fs.writeFileSync(sample, JSON.stringify((JSON.parse(fs.readFileSync(sample, 'utf-8')) as Array<Record<string, unknown>>).map(finding => {
const snippet = String(finding.snippet);
const where = finding.antipattern === 'skipped-heading' ? located('index.html', 'Feature One')
: finding.antipattern === 'marketing-buzzword' ? located('index.html', 'streamline')
: located('styles.css', /on (#[0-9a-f]{6})/.exec(snippet)?.[1] ?? '#8b5cf6');
return { ...finding, ...where };
}), null, 2));
fs.writeFileSync(path.join(repoDir, 'design-review-detector.md'), detectorSkillText([
['**Design detector (optional, deterministic):**', '**Create output directories:**'],
['**Phase 0: mechanical scan**', '## Phases 1-6'],
@@ -860,6 +871,17 @@ if (!evalsEnabled) test('plugin detector fixture discovers the selected engine w
});
expect(scan.status).toBe(2);
expect(scan.stderr).toContain('handoff=/impeccable colorize');
// Census 36709485593: rows naming test/fixtures/... paths absent from this repo cost four
// reconciliation turns and exceeded max turns. Every row must cite a repo file:line holding its evidence.
const rows = [...scan.stderr.matchAll(/^ {2}(\S+):(\d+) {2}(.*)$/gm)];
expect(rows).toHaveLength(6);
for (const [, file, line, snippet] of rows) {
const text = fs.readFileSync(path.join(fixture.repoDir, file!), 'utf-8').split('\n')[Number(line) - 1]!;
const evidence = /on (#[0-9a-f]{6})/.exec(snippet!)?.[1] ?? (/Purple/.test(snippet!) ? '#8b5cf6' : /buzzword/.test(snippet!) ? 'streamline' : 'Feature One');
expect(text, `${file}:${line}`).toContain(evidence);
}
const shipped = JSON.parse(fs.readFileSync(DETECT_SAMPLE, 'utf-8')) as Array<{ file: string }>;
expect(shipped.every(finding => !fs.existsSync(path.join(fixture.repoDir, finding.file)))).toBe(true);
expect(fs.existsSync(path.join(fixture.dir, '4.10.0.jsonl'))).toBe(true);
expect(fs.existsSync(path.join(fixture.dir, '4.3.1.jsonl'))).toBe(false);
expect(fs.existsSync(path.join(fixture.dir, 'launcher-ran'))).toBe(false);
+5 -6
View File
@@ -26,7 +26,7 @@ import * as path from 'path';
import * as http from 'http';
import { runAgentSdkTest, toSkillTestResult, passThroughNonAskUserQuestion, resolveClaudeBinary, type AgentSdkResult, type QueryProvider, type RunAgentSdkOptions } from './helpers/agent-sdk-runner';
import { runRecordedOfficeHoursAttempt, OFFICE_HOURS_BUN_GRACE_MS } from './helpers/office-hours-attempt';
import { publicEvents, redactPublicValue } from './helpers/setup-gbrain-sandbox';
import { publicEvents, redactPublicValue, setupGbrainRemoteAnswer } from './helpers/setup-gbrain-sandbox';
import { EvalCollector } from './helpers/eval-store';
import { resolveEvalModel } from '../lib/eval-model';
import { buildSetupGbrainFixture } from './helpers/setup-gbrain-fixture';
@@ -263,17 +263,16 @@ describeE2E('/setup-gbrain Path 4 (Remote MCP) — happy path', () => {
askUserQuestions.push({ input });
const q = (input.questions as Array<{
question: string;
header?: string;
options: Array<{ label: string }>;
}>)[0];
// Auto-decline / skip everything except the path-pick (which the
// user-prompt already directed to Path 4).
const decline =
q.options.find((o) => /skip|decline|no thanks|local/i.test(o.label)) ?? q.options[q.options.length - 1]!;
// Register at Step 5a (the contract under test); decline or skip every other gate.
const answer = setupGbrainRemoteAnswer(q);
return {
behavior: 'allow',
updatedInput: {
questions: input.questions,
answers: { [q.question]: decline.label },
answers: { [q.question]: answer },
},
};
}