mirror of
https://github.com/garrytan/gstack.git
synced 2026-10-03 18:06:54 +02:00
fix(evals): repair proof-run reds in design-consultation, document-release, design and QA fixtures
- design-consultation Phase 1 asks one brief that confirms context and decides
research; the confirm-only first question scored substance 2.
- document-release defines ship-owned inputs, exact steps and the JSON result,
and drops stale spawned-from-/ship text (judge actionability 3.67 -> 4/4/4).
- plan-design-with-ui accepts the Step 0D focus menu the same way the shared
picker does ("focus on specific ones?").
- plan-design-review plan-mode saves in three Edits instead of one final Write.
- QA functional annotations ask for the full 40-character revision.
- Outside-disabled attribution judges quoted prior-record data by its exact
timestamp or a dated, pre-existing-record sentence; four captured phrasings
replay clean and current claims still fail.
- --case can select autoplan-dual-voice by its literal test name.
This commit is contained in:
1 parent
aba80c8fb6
commit
a111225e78
17 files changed
+180
-92
No files matched your search
@@ -644,15 +644,12 @@ sections. Read a section in full before doing its step; do not work from memory.
|
||||
|
||||
## Phase 1: Product Context
|
||||
|
||||
Confirm product context in Q1, pre-filled from the codebase; then ask the memorable-thing question.
|
||||
**AskUserQuestion Q1 — one brief that confirms context AND decides research.** Never ask a confirm-only question first. In the ELI10, state your pre-filled read (from README, product files or office-hours output): what the product is, who it's for, its space and project type (web app, dashboard, marketing site, editorial, internal tool, etc.). Options:
|
||||
- A) Context right — research what top products in this space do for design first
|
||||
- B) Context right — work from design knowledge only
|
||||
- C) Context wrong or incomplete — I'll correct it
|
||||
|
||||
**AskUserQuestion Q1 — include ALL of these:**
|
||||
1. Confirm what the product is, who it's for, what space/industry
|
||||
2. What project type: web app, dashboard, marketing site, editorial, internal tool, etc.
|
||||
3. "Want me to research what top products in your space are doing for design, or should I work from my design knowledge?"
|
||||
4. **Explicitly say:** "At any point you can just drop into chat and we'll talk through anything — this isn't a rigid form, it's a conversation."
|
||||
|
||||
Pre-fill context from README or office-hours output, then confirm it and the research preference in Q1.
|
||||
Recommend A or B for this product, naming what research buys or costs here versus the other. **Explicitly say:** "At any point you can just drop into chat and we'll talk through anything — this isn't a rigid form, it's a conversation."
|
||||
|
||||
**Memorable-thing forcing question.** Before moving on, ask the user: *"What's the one
|
||||
thing you want someone to remember after they see this product for the first time?"*
|
||||
|
||||
@@ -125,15 +125,12 @@ Phase 5: `DESIGN_READY` uses AI mockups on realistic product screens; `DESIGN_NO
|
||||
|
||||
## Phase 1: Product Context
|
||||
|
||||
Confirm product context in Q1, pre-filled from the codebase; then ask the memorable-thing question.
|
||||
**AskUserQuestion Q1 — one brief that confirms context AND decides research.** Never ask a confirm-only question first. In the ELI10, state your pre-filled read (from README, product files or office-hours output): what the product is, who it's for, its space and project type (web app, dashboard, marketing site, editorial, internal tool, etc.). Options:
|
||||
- A) Context right — research what top products in this space do for design first
|
||||
- B) Context right — work from design knowledge only
|
||||
- C) Context wrong or incomplete — I'll correct it
|
||||
|
||||
**AskUserQuestion Q1 — include ALL of these:**
|
||||
1. Confirm what the product is, who it's for, what space/industry
|
||||
2. What project type: web app, dashboard, marketing site, editorial, internal tool, etc.
|
||||
3. "Want me to research what top products in your space are doing for design, or should I work from my design knowledge?"
|
||||
4. **Explicitly say:** "At any point you can just drop into chat and we'll talk through anything — this isn't a rigid form, it's a conversation."
|
||||
|
||||
Pre-fill context from README or office-hours output, then confirm it and the research preference in Q1.
|
||||
Recommend A or B for this product, naming what research buys or costs here versus the other. **Explicitly say:** "At any point you can just drop into chat and we'll talk through anything — this isn't a rigid form, it's a conversation."
|
||||
|
||||
**Memorable-thing forcing question.** Before moving on, ask the user: *"What's the one
|
||||
thing you want someone to remember after they see this product for the first time?"*
|
||||
|
||||
+10
-10
@@ -425,7 +425,7 @@ Make factual updates directly; ask about risky or subjective decisions in standa
|
||||
|
||||
## Ship-owned documentation mode
|
||||
|
||||
With a ship candidate, require the actual spawned marker and audit-scope rules below.
|
||||
With a ship candidate, follow audit-scope's inputs, steps and JSON result below.
|
||||
Missing marking/inputs/assets returns `blocked`, never standalone execution. Ship
|
||||
authority overrides generic spawned recommendations and standalone steps.
|
||||
|
||||
@@ -438,8 +438,8 @@ authority overrides generic spawned recommendations and standalone steps.
|
||||
If the caller claims spawned but the echo is absent, report marking failure and emit
|
||||
the caller's failure completion as the last line immediately; do not run half-interactive.
|
||||
Otherwise stay interactive without the marker. Outside ship-owned mode, spawned gates
|
||||
auto-choose the RECOMMENDED option, record it in the completion report, and continue:
|
||||
never call AskUserQuestion or stop for a prose answer. The NEVER-do invariants below do
|
||||
auto-choose the RECOMMENDED option, record it in the completion report, and continue
|
||||
through Step 9: never call AskUserQuestion or stop for a prose answer. The NEVER-do invariants below do
|
||||
not relax: skip any recommendation that rewrites CHANGELOG or changes VERSION and
|
||||
record why. Step 8 and cross-model review refer to this rule; narrower caller scope wins.
|
||||
|
||||
@@ -488,10 +488,10 @@ DOC_DIFF_BASE=$(git merge-base origin/<base> HEAD 2>/dev/null || git merge-base
|
||||
echo "DOC_DIFF_BASE: $DOC_DIFF_BASE"
|
||||
```
|
||||
|
||||
1. Check the current branch. In standalone mode, if on the base branch, **abort**: "You're on the base branch. Run from a feature branch." A ship-owned read-only store audit uses its supplied source scope instead.
|
||||
1. Check the current branch. In standalone mode, if on the base branch, **abort**: "You're on the base branch. Run from a feature branch." Ship-owned mode skips this gate.
|
||||
|
||||
2. Gather the diff. In ship-owned mode, also read `git diff --cached`, `git diff`,
|
||||
and selected new-file content against the supplied base, not HEAD alone.
|
||||
2. Gather the diff. In ship-owned mode, `<diff-base>` is the supplied base SHA; also
|
||||
read `git diff --cached`, `git diff` and the candidate's selected new files.
|
||||
|
||||
```bash
|
||||
git diff <diff-base> HEAD --stat
|
||||
@@ -546,16 +546,16 @@ Use these definitions:
|
||||
- **Tutorial** — learning-oriented: step-by-step walkthrough for newcomers (getting started guides)
|
||||
- **Explanation** — understanding-oriented: "why this works this way" (ARCHITECTURE decisions, design rationale)
|
||||
|
||||
3. **Output the coverage map.** Items with zero coverage are **critical gaps** — flag them for
|
||||
Step 3. Items with reference-only coverage are **common gaps** — note them for the PR body.
|
||||
3. **Output the coverage map.** Items with zero coverage are **critical gaps**; items with
|
||||
reference-only coverage are **common gaps**. Report both as documentation debt.
|
||||
|
||||
4. **Architecture diagram drift detection.** If ARCHITECTURE.md (or any doc) contains ASCII
|
||||
diagrams or Mermaid blocks, extract entity names (modules, services, data flows) from the
|
||||
diagrams. Cross-reference against the diff. Flag any diagram entities that were renamed,
|
||||
split, removed, or moved in the code.
|
||||
|
||||
The coverage map feeds into Steps 2-3 (what to audit and fix) and Step 9 (documentation debt
|
||||
summary in the PR body). Do NOT auto-generate missing documentation pages — flag gaps only.
|
||||
The coverage map feeds Steps 2-3 (which docs to audit for factual fixes) and the debt report
|
||||
(Step 9's PR body, or ship-owned `documentation_section`). Do NOT auto-generate missing documentation pages — flag gaps only.
|
||||
When significant gaps are found, suggest running `/document-generate` to fill them.
|
||||
|
||||
---
|
||||
|
||||
@@ -38,7 +38,7 @@ Make factual updates directly; ask about risky or subjective decisions in standa
|
||||
|
||||
## Ship-owned documentation mode
|
||||
|
||||
With a ship candidate, require the actual spawned marker and audit-scope rules below.
|
||||
With a ship candidate, follow audit-scope's inputs, steps and JSON result below.
|
||||
Missing marking/inputs/assets returns `blocked`, never standalone execution. Ship
|
||||
authority overrides generic spawned recommendations and standalone steps.
|
||||
|
||||
@@ -50,8 +50,8 @@ authority overrides generic spawned recommendations and standalone steps.
|
||||
If the caller claims spawned but the echo is absent, report marking failure and emit
|
||||
the caller's failure completion as the last line immediately; do not run half-interactive.
|
||||
Otherwise stay interactive without the marker. Outside ship-owned mode, spawned gates
|
||||
auto-choose the RECOMMENDED option, record it in the completion report, and continue:
|
||||
never call AskUserQuestion or stop for a prose answer. The NEVER-do invariants below do
|
||||
auto-choose the RECOMMENDED option, record it in the completion report, and continue
|
||||
through Step 9: never call AskUserQuestion or stop for a prose answer. The NEVER-do invariants below do
|
||||
not relax: skip any recommendation that rewrites CHANGELOG or changes VERSION and
|
||||
record why. Step 8 and cross-model review refer to this rule; narrower caller scope wins.
|
||||
|
||||
@@ -92,10 +92,10 @@ DOC_DIFF_BASE=$(git merge-base origin/<base> HEAD 2>/dev/null || git merge-base
|
||||
echo "DOC_DIFF_BASE: $DOC_DIFF_BASE"
|
||||
```
|
||||
|
||||
1. Check the current branch. In standalone mode, if on the base branch, **abort**: "You're on the base branch. Run from a feature branch." A ship-owned read-only store audit uses its supplied source scope instead.
|
||||
1. Check the current branch. In standalone mode, if on the base branch, **abort**: "You're on the base branch. Run from a feature branch." Ship-owned mode skips this gate.
|
||||
|
||||
2. Gather the diff. In ship-owned mode, also read `git diff --cached`, `git diff`,
|
||||
and selected new-file content against the supplied base, not HEAD alone.
|
||||
2. Gather the diff. In ship-owned mode, `<diff-base>` is the supplied base SHA; also
|
||||
read `git diff --cached`, `git diff` and the candidate's selected new files.
|
||||
|
||||
```bash
|
||||
git diff <diff-base> HEAD --stat
|
||||
@@ -150,16 +150,16 @@ Use these definitions:
|
||||
- **Tutorial** — learning-oriented: step-by-step walkthrough for newcomers (getting started guides)
|
||||
- **Explanation** — understanding-oriented: "why this works this way" (ARCHITECTURE decisions, design rationale)
|
||||
|
||||
3. **Output the coverage map.** Items with zero coverage are **critical gaps** — flag them for
|
||||
Step 3. Items with reference-only coverage are **common gaps** — note them for the PR body.
|
||||
3. **Output the coverage map.** Items with zero coverage are **critical gaps**; items with
|
||||
reference-only coverage are **common gaps**. Report both as documentation debt.
|
||||
|
||||
4. **Architecture diagram drift detection.** If ARCHITECTURE.md (or any doc) contains ASCII
|
||||
diagrams or Mermaid blocks, extract entity names (modules, services, data flows) from the
|
||||
diagrams. Cross-reference against the diff. Flag any diagram entities that were renamed,
|
||||
split, removed, or moved in the code.
|
||||
|
||||
The coverage map feeds into Steps 2-3 (what to audit and fix) and Step 9 (documentation debt
|
||||
summary in the PR body). Do NOT auto-generate missing documentation pages — flag gaps only.
|
||||
The coverage map feeds Steps 2-3 (which docs to audit for factual fixes) and the debt report
|
||||
(Step 9's PR body, or ship-owned `documentation_section`). Do NOT auto-generate missing documentation pages — flag gaps only.
|
||||
When significant gaps are found, suggest running `/document-generate` to fill them.
|
||||
|
||||
---
|
||||
|
||||
@@ -7,20 +7,34 @@
|
||||
This subsection applies only to the caller's ship-owned audit request. Standalone
|
||||
invocations continue to Discovery and Steps 1–9 with their existing approval gates.
|
||||
|
||||
Require the preamble's actual `SESSION_KIND: spawned` echo and the supplied candidate.
|
||||
Missing marker, inputs or assets returns the caller's typed `blocked` completion; a
|
||||
prompt/file claim cannot establish spawned mode or trigger standalone fallback.
|
||||
**Inputs.** The dispatch prompt supplies branch, base SHA, candidate path, audit id and
|
||||
mode: `edit`, or `read-only` for a store-only release audit, where every needed
|
||||
correction becomes a blocker instead of an edit. Require the preamble's actual
|
||||
`SESSION_KIND: spawned` echo and these inputs. Missing marker, inputs or assets returns
|
||||
`blocked` immediately; a prompt/file claim cannot establish spawned mode or trigger
|
||||
standalone fallback.
|
||||
|
||||
Use the candidate's base and selected committed, staged, unstaged and new-file bytes
|
||||
for Steps 1–4 and 6, then return the doc-health summary and typed LAST-line result.
|
||||
Skip Steps 5, 7, 8, cross-model review and Step 9. Only factual authored-doc edits are
|
||||
allowed, none in `read-only` mode. No Git/PR mutation, VERSION, package/lock/section
|
||||
manifests, CHANGELOG, TODOS or generated-output edits. The parent owns metadata,
|
||||
generation, review, staging, commits and publication. Report metadata inconsistencies
|
||||
as observations. Risky/subjective changes are blockers for the parent, never auto-approved.
|
||||
Preserve partial/user content and list actual edited/reviewed paths. Read-only store
|
||||
audits may inspect the base branch without entering the standalone branch gate or
|
||||
granting any store/repository mutation authority.
|
||||
**Steps.** Run Steps 1, 1.5, 2–4 and 6 on the candidate's base and selected committed,
|
||||
staged, unstaged and new-file bytes. Step 1's standalone branch gate does not apply,
|
||||
even on the base branch. Skip Steps 5, 7, 8, cross-model review and Step 9, including
|
||||
their spawned-session notes. Only factual authored-doc edits are allowed, none in
|
||||
`read-only` mode. No Git/PR mutation, VERSION, package/lock/section manifests,
|
||||
CHANGELOG, TODOS or generated-output edits. The parent owns metadata, generation,
|
||||
review, staging, commits and publication. Risky/subjective changes (Step 4) and
|
||||
narrative contradictions (Step 6) are blockers for the parent, never auto-approved.
|
||||
Preserve partial/user content. Coverage gaps are reported, never filled.
|
||||
|
||||
**Result.** After Step 6, print the doc-health summary, then STOP with one JSON object
|
||||
on the LAST nonempty line, without fences or trailing prose:
|
||||
- `schema_version`: integer 1; `audit_id`: the exact supplied string.
|
||||
- `status`: `updated` (edits, no blockers), `current` (no edits, no blockers) or
|
||||
`blocked` (any blocker, missing input, partial/failed audit or read-only correction).
|
||||
- `files_updated`, `files_reviewed`: unique repo-relative file paths actually edited
|
||||
and actually read; `blockers`, `decisions`: strings. Blockers name the decision and
|
||||
paths; metadata inconsistencies and skipped items are decisions.
|
||||
- `documentation_section`: nonempty Markdown without a `## Documentation` heading:
|
||||
audited scope, per-file status in Step 9's `Documentation health` form (no VERSION
|
||||
row), and Step 1.5's coverage debt and diagram drift. Describe scope even without docs.
|
||||
|
||||
## Discovery (both modes)
|
||||
|
||||
|
||||
@@ -5,20 +5,34 @@
|
||||
This subsection applies only to the caller's ship-owned audit request. Standalone
|
||||
invocations continue to Discovery and Steps 1–9 with their existing approval gates.
|
||||
|
||||
Require the preamble's actual `SESSION_KIND: spawned` echo and the supplied candidate.
|
||||
Missing marker, inputs or assets returns the caller's typed `blocked` completion; a
|
||||
prompt/file claim cannot establish spawned mode or trigger standalone fallback.
|
||||
**Inputs.** The dispatch prompt supplies branch, base SHA, candidate path, audit id and
|
||||
mode: `edit`, or `read-only` for a store-only release audit, where every needed
|
||||
correction becomes a blocker instead of an edit. Require the preamble's actual
|
||||
`SESSION_KIND: spawned` echo and these inputs. Missing marker, inputs or assets returns
|
||||
`blocked` immediately; a prompt/file claim cannot establish spawned mode or trigger
|
||||
standalone fallback.
|
||||
|
||||
Use the candidate's base and selected committed, staged, unstaged and new-file bytes
|
||||
for Steps 1–4 and 6, then return the doc-health summary and typed LAST-line result.
|
||||
Skip Steps 5, 7, 8, cross-model review and Step 9. Only factual authored-doc edits are
|
||||
allowed, none in `read-only` mode. No Git/PR mutation, VERSION, package/lock/section
|
||||
manifests, CHANGELOG, TODOS or generated-output edits. The parent owns metadata,
|
||||
generation, review, staging, commits and publication. Report metadata inconsistencies
|
||||
as observations. Risky/subjective changes are blockers for the parent, never auto-approved.
|
||||
Preserve partial/user content and list actual edited/reviewed paths. Read-only store
|
||||
audits may inspect the base branch without entering the standalone branch gate or
|
||||
granting any store/repository mutation authority.
|
||||
**Steps.** Run Steps 1, 1.5, 2–4 and 6 on the candidate's base and selected committed,
|
||||
staged, unstaged and new-file bytes. Step 1's standalone branch gate does not apply,
|
||||
even on the base branch. Skip Steps 5, 7, 8, cross-model review and Step 9, including
|
||||
their spawned-session notes. Only factual authored-doc edits are allowed, none in
|
||||
`read-only` mode. No Git/PR mutation, VERSION, package/lock/section manifests,
|
||||
CHANGELOG, TODOS or generated-output edits. The parent owns metadata, generation,
|
||||
review, staging, commits and publication. Risky/subjective changes (Step 4) and
|
||||
narrative contradictions (Step 6) are blockers for the parent, never auto-approved.
|
||||
Preserve partial/user content. Coverage gaps are reported, never filled.
|
||||
|
||||
**Result.** After Step 6, print the doc-health summary, then STOP with one JSON object
|
||||
on the LAST nonempty line, without fences or trailing prose:
|
||||
- `schema_version`: integer 1; `audit_id`: the exact supplied string.
|
||||
- `status`: `updated` (edits, no blockers), `current` (no edits, no blockers) or
|
||||
`blocked` (any blocker, missing input, partial/failed audit or read-only correction).
|
||||
- `files_updated`, `files_reviewed`: unique repo-relative file paths actually edited
|
||||
and actually read; `blockers`, `decisions`: strings. Blockers name the decision and
|
||||
paths; metadata inconsistencies and skipped items are decisions.
|
||||
- `documentation_section`: nonempty Markdown without a `## Documentation` heading:
|
||||
audited scope, per-file status in Step 9's `Documentation health` form (no VERSION
|
||||
row), and Step 1.5's coverage debt and diagram drift. Describe scope even without docs.
|
||||
|
||||
## Discovery (both modes)
|
||||
|
||||
|
||||
@@ -2,8 +2,8 @@
|
||||
<!-- Regenerate: bun run gen:skill-docs -->
|
||||
## Step 2: Per-File Documentation Audit
|
||||
|
||||
**Ship-owned documentation mode:** execute Steps 2–4 and 6 only, under the skeleton's
|
||||
audit/edit/result boundary. Then return the caller's typed completion; all standalone
|
||||
**Ship-owned documentation mode:** after Steps 1 and 1.5, execute Steps 2–4 and 6 only,
|
||||
under audit-scope's edit boundary, then return its JSON result; all standalone
|
||||
metadata, review, commit and PR steps below remain unavailable to this child.
|
||||
|
||||
Read each documentation file and cross-reference it against the diff. Use these generic heuristics
|
||||
@@ -131,8 +131,8 @@ After auditing each file individually, do a cross-doc consistency pass:
|
||||
|
||||
In ship-owned mode, protected metadata/manifests stay untouched even for factual
|
||||
inconsistencies, and narrative contradictions return as blockers. This is the last
|
||||
ship-child step: output the doc-health summary and typed completion, then STOP. A
|
||||
partial audit or unresolved required correction is `blocked`, never `current`.
|
||||
ship-child step: output the doc-health summary and audit-scope's JSON result, then
|
||||
STOP. A partial audit or unresolved required correction is `blocked`, never `current`.
|
||||
|
||||
---
|
||||
|
||||
@@ -193,7 +193,7 @@ git diff <diff-base> HEAD -- VERSION
|
||||
|
||||
**Spawned sessions** (per the spawned-dispatch contract at the top of this skill): the
|
||||
recommendation flips — choose C (leave version as-is) and record the uncovered scope in
|
||||
your completion report (the `decisions` array when dispatched from /ship).
|
||||
your completion report. Ship-owned children stopped at Step 6 and never reach this step.
|
||||
A spawned run must never change VERSION: the dispatching workflow owns version numbering.
|
||||
|
||||
The key insight: a VERSION bump set for "feature A" should not silently absorb "feature B"
|
||||
@@ -211,7 +211,7 @@ not an opt-in. The user turns it off only by asking explicitly
|
||||
**Spawned-session skip** (per the spawned-dispatch contract at the top of this skill): in a
|
||||
spawned session, skip this entire section — the dispatching workflow owns its own review
|
||||
passes, and the apply gate below needs a human. Note the skip in the upcoming Step 9 doc
|
||||
health summary and continue to Step 9.
|
||||
health summary and continue to Step 9. Ship-owned children already stopped at Step 6.
|
||||
|
||||
**Preflight — decide whether and how the doc review runs:**
|
||||
|
||||
|
||||
@@ -1,7 +1,7 @@
|
||||
## Step 2: Per-File Documentation Audit
|
||||
|
||||
**Ship-owned documentation mode:** execute Steps 2–4 and 6 only, under the skeleton's
|
||||
audit/edit/result boundary. Then return the caller's typed completion; all standalone
|
||||
**Ship-owned documentation mode:** after Steps 1 and 1.5, execute Steps 2–4 and 6 only,
|
||||
under audit-scope's edit boundary, then return its JSON result; all standalone
|
||||
metadata, review, commit and PR steps below remain unavailable to this child.
|
||||
|
||||
Read each documentation file and cross-reference it against the diff. Use these generic heuristics
|
||||
@@ -129,8 +129,8 @@ After auditing each file individually, do a cross-doc consistency pass:
|
||||
|
||||
In ship-owned mode, protected metadata/manifests stay untouched even for factual
|
||||
inconsistencies, and narrative contradictions return as blockers. This is the last
|
||||
ship-child step: output the doc-health summary and typed completion, then STOP. A
|
||||
partial audit or unresolved required correction is `blocked`, never `current`.
|
||||
ship-child step: output the doc-health summary and audit-scope's JSON result, then
|
||||
STOP. A partial audit or unresolved required correction is `blocked`, never `current`.
|
||||
|
||||
---
|
||||
|
||||
@@ -191,7 +191,7 @@ git diff <diff-base> HEAD -- VERSION
|
||||
|
||||
**Spawned sessions** (per the spawned-dispatch contract at the top of this skill): the
|
||||
recommendation flips — choose C (leave version as-is) and record the uncovered scope in
|
||||
your completion report (the `decisions` array when dispatched from /ship).
|
||||
your completion report. Ship-owned children stopped at Step 6 and never reach this step.
|
||||
A spawned run must never change VERSION: the dispatching workflow owns version numbering.
|
||||
|
||||
The key insight: a VERSION bump set for "feature A" should not silently absorb "feature B"
|
||||
|
||||
@@ -1361,7 +1361,7 @@ not an opt-in. The user turns it off only by asking explicitly
|
||||
**Spawned-session skip** (per the spawned-dispatch contract at the top of this skill): in a
|
||||
spawned session, skip this entire section — the dispatching workflow owns its own review
|
||||
passes, and the apply gate below needs a human. Note the skip in the upcoming Step 9 doc
|
||||
health summary and continue to Step 9.
|
||||
health summary and continue to Step 9. Ship-owned children already stopped at Step 6.
|
||||
|
||||
**Preflight — decide whether and how the doc review runs:**
|
||||
|
||||
|
||||
@@ -149,6 +149,7 @@ export const CASE_SHARDED_FILES: readonly string[] = [
|
||||
export const CASE_TEST_NAMES: Record<string, string> = {
|
||||
'plan-review-report': '/plan-eng-review writes GSTACK REVIEW REPORT to plan file',
|
||||
'auq-format-gate': "/plan-ceo-review's first AskUserQuestion is a compliant decision brief (7/7 + substance)",
|
||||
'autoplan-dual-voice': 'both Claude + Codex voices produce output in Phase 1 (within timeout)',
|
||||
};
|
||||
|
||||
const CASE_KEY_SEPARATOR = '#';
|
||||
|
||||
@@ -257,14 +257,50 @@ describe('parenthesized field list of the retained record (census 36629958451)',
|
||||
});
|
||||
test.each([
|
||||
['a different timestamp', [prior.timestamp!.replace(/\d(?=Z$)/, d => String((Number(d) + 1) % 10)), ...exact.slice(1)]],
|
||||
['no timestamp', exact.slice(1)],
|
||||
['a field value the record lacks', [prior.timestamp!, 'status: completed', `source: ${prior.source}`, 'outside_status: completed']],
|
||||
['free prose inside the list', [prior.timestamp!, `status: ${prior.status}`, 'outside_status: completed and this run too']],
|
||||
])('%s is not attributed', (_name, items) => {
|
||||
expect(evaluate(0, list(items as string[])).falseCompletion).toBe(true);
|
||||
});
|
||||
test('neither the exact instant nor a before-this-run dating is not attributed', () => {
|
||||
expect(evaluate(0, list(exact.slice(1)).replace(' that predates this run', '')).falseCompletion).toBe(true);
|
||||
});
|
||||
test('split field quotes in one dated sentence attribute the record (local proof, 2026-09-29)', () => {
|
||||
const clock = prior.timestamp!.slice(11, 19) + 'Z';
|
||||
const text = `One thing to flag: the review log already contained a pre-existing entry, timestamped \`${clock}\` (before this session started), claiming \`status:"${prior.status}"\`, \`source:"${prior.source}"\`, \`outside_status:"completed"\`. I did not produce that entry.`;
|
||||
expect(evaluate(0, text)).toMatchObject({ passed: true, falseCompletion: false });
|
||||
expect(evaluate(0, text.replace('claiming', 'and I now also recorded')).falseCompletion).toBe(true);
|
||||
expect(evaluate(0, text.replace(`status:"${prior.status}"`, 'status:"completed"')).falseCompletion).toBe(true);
|
||||
expect(evaluate(0, text.replace('pre-existing entry', 'entry')).falseCompletion).toBe(true);
|
||||
});
|
||||
test('a current mutation or a separate current claim still fails', () => {
|
||||
expect(evaluate(0, list(exact).replace('already contained', 'now appended')).falseCompletion).toBe(true);
|
||||
expect(evaluate(0, `${list(exact)}\nThis run: outside_status: completed.`).falseCompletion).toBe(true);
|
||||
});
|
||||
});
|
||||
|
||||
describe('millisecond clock of the retained record (local proof, 2026-09-29)', () => {
|
||||
const prior = fixture.cases[0]!.priorRecord as Record<string, string>;
|
||||
const clock = prior.timestamp!.slice(11, 23) + 'Z';
|
||||
const flag = (stamp: string) => `One thing to flag: the same log file already contained a pre-existing entry (line 1, timestamp \`${stamp}\`, before this session) claiming \`status: ${prior.status}, source: ${prior.source}, outside_status: completed\`. I did not write it and did not remove it.`;
|
||||
test('the exact millisecond clock attributes the quotation to that record', () => {
|
||||
expect(evaluate(0, flag(clock))).toMatchObject({ passed: true, falseCompletion: false });
|
||||
});
|
||||
test('a different millisecond clock is not attributed', () => {
|
||||
const other = clock.replace(/\d(?=Z$)/, d => String((Number(d) + 1) % 10));
|
||||
expect(evaluate(0, flag(other)).falseCompletion).toBe(true);
|
||||
});
|
||||
});
|
||||
|
||||
describe('quoted fragment with the retained record timestamp (local proof, 2026-09-29)', () => {
|
||||
const prior = fixture.cases[0]!.priorRecord as Record<string, string>;
|
||||
const quote = (extra = '') => `One thing to flag: the review log already contained a prior line I did not write —\n\`{"status":"${prior.status}","source":"${prior.source}","outside_status":"completed", "timestamp":"${prior.timestamp}"${extra}, ...}\`.\nThat record claims a completed, clean Codex outside review of this plan. It predates my run by ~80 seconds and is inconsistent with the current configuration.`;
|
||||
test('the exact timestamp and matching values make the fragment record data', () => {
|
||||
expect(evaluate(0, quote())).toMatchObject({ passed: true, falseCompletion: false });
|
||||
});
|
||||
test('a mismatched value or a missing timestamp keeps the claim', () => {
|
||||
expect(evaluate(0, quote(', "source":"claude"')).falseCompletion).toBe(true);
|
||||
expect(evaluate(0, quote().replace(prior.timestamp!, '2026-09-29T22:28:27Z')).falseCompletion).toBe(true);
|
||||
expect(evaluate(0, `${quote()}\nThis run: outside_status: completed.`).falseCompletion).toBe(true);
|
||||
});
|
||||
});
|
||||
@@ -145,22 +145,22 @@ function withoutAttributedPriorRecordData(output: string, priorRecord?: Record<s
|
||||
// An inline quotation of the retained record's exact status/source/outside_status
|
||||
// values is that record when its own sentence names it as pre-existing and
|
||||
// makes no current claim; wording order around the quotation does not matter.
|
||||
// A named record timestamp must denote the retained record's instant at the precision written.
|
||||
const priorMs = typeof priorRecord.timestamp === 'string' ? Date.parse(priorRecord.timestamp) : NaN;
|
||||
const sameInstant = (stamp: string): boolean => {
|
||||
if (!Number.isFinite(priorMs)) return false;
|
||||
const iso = priorMs ? new Date(priorMs).toISOString() : '';
|
||||
const clock = /^(\d{2}:\d{2}(?::\d{2}(?:\.\d{1,3})?)?)Z?$/.exec(stamp);
|
||||
if (clock) return iso.slice(11, 11 + clock[1]!.length) === clock[1];
|
||||
const at = /^\d{4}-\d{2}-\d{2}T\d{2}:\d{2}(?::\d{2}(?:\.\d+)?)?Z$/.test(stamp) ? Date.parse(stamp) : NaN;
|
||||
return Number.isFinite(at) && iso.slice(0, stamp.includes('.') ? 23 : stamp.length - 1) === new Date(at).toISOString().slice(0, stamp.includes('.') ? 23 : stamp.length - 1);
|
||||
};
|
||||
const sentenceOwnsPriorValue = (index: number, length: number): boolean => {
|
||||
const start = Math.max(output.lastIndexOf('\n', index - 1), ...['. ', '! ', '? ', '; '].map(end => output.lastIndexOf(end, index - 1) + 1)) + 1;
|
||||
const ends = ['\n', '. ', '! ', '? ', '; '].map(end => output.indexOf(end, index + length)).filter(at => at >= 0);
|
||||
const sentence = (output.slice(start, index) + ' ' + output.slice(index + length, ends.length ? Math.min(...ends) : output.length))
|
||||
.replace(/[*`]/g, '').replace(/\b(?:predates|before)\s+(?:this|my)\s+(?:run|session|workflow)(?:\s+(?:started|began))?\b/gi, 'beforehand')
|
||||
.replace(/\b(?:I|we)\s+(?:did\s+not|didn't|never)\s+(?:write|create|produce|record)\b/gi, 'unauthored');
|
||||
// A named record timestamp must denote the retained record's instant at the precision written.
|
||||
const priorMs = typeof priorRecord.timestamp === 'string' ? Date.parse(priorRecord.timestamp) : NaN;
|
||||
const sameInstant = (stamp: string): boolean => {
|
||||
if (!Number.isFinite(priorMs)) return false;
|
||||
const iso = priorMs ? new Date(priorMs).toISOString() : '';
|
||||
const clock = /^(\d{2}:\d{2}(?::\d{2})?)Z?$/.exec(stamp);
|
||||
if (clock) return iso.slice(11, 11 + clock[1]!.length) === clock[1];
|
||||
const at = /^\d{4}-\d{2}-\d{2}T\d{2}:\d{2}(?::\d{2}(?:\.\d+)?)?Z$/.test(stamp) ? Date.parse(stamp) : NaN;
|
||||
return Number.isFinite(at) && iso.slice(0, stamp.includes('.') ? 23 : stamp.length - 1) === new Date(at).toISOString().slice(0, stamp.includes('.') ? 23 : stamp.length - 1);
|
||||
};
|
||||
const stamps = [...sentence.matchAll(/\btimestamp(?:ed)?\s+([0-9T:.Z-]+)/gi)].map(stamp => stamp[1]!.replace(/[.,;:]+$/, ''));
|
||||
if (stamps.some(stamp => !sameInstant(stamp))) return false;
|
||||
return !/\b(?:after|another|other|if|unless)\b/i.test(sentence)
|
||||
@@ -191,6 +191,32 @@ function withoutAttributedPriorRecordData(output: string, priorRecord?: Record<s
|
||||
const start = match.index + match[0].length - match[1]!.length - 1;
|
||||
spans.push({ start, end: start + match[1]!.length + 2 });
|
||||
}
|
||||
// A quoted fragment carrying the retained record's exact timestamp is that
|
||||
// record's data when every field it quotes has that record's value.
|
||||
for (const match of output.matchAll(/`([^`\r\n]+)`/g)) {
|
||||
if (typeof priorRecord.timestamp !== 'string' || !match[1]!.includes(priorRecord.timestamp)) continue;
|
||||
const pairs = [...match[1]!.matchAll(/["']?([a-z_]+)["']?\s*[:=]\s*["']?([^"',}\s]+)["']?/gi)].filter(pair => fields.has(pair[1]!));
|
||||
if (!pairs.some(pair => pair[1] === 'outside_status') || pairs.some(pair => priorRecord[pair[1]!] !== pair[2])) continue;
|
||||
spans.push({ start: match.index!, end: match.index! + match[0].length });
|
||||
}
|
||||
// A whole sentence that names the pre-existing record, dates it before this
|
||||
// run (its exact instant or an explicit "before this run"), quotes only that
|
||||
// record's own field values and makes no current claim is that record's
|
||||
// report, however its fields are quoted or split.
|
||||
for (const sentence of output.matchAll(/[^\n.!?;]*(?:[.!?;](?=\S)[^\n.!?;]*)*(?:[.!?;](?=\s|$)|\n|$)/g)) {
|
||||
const plain = sentence[0].replace(/[*`]/g, '');
|
||||
if (!/\boutside_status["']*\s*[:=]\s*["']*completed\b/i.test(plain)) continue;
|
||||
if (!/\b(?:earlier|prior|previous|historical|old(?:er)?|pre[- ]existing|stale|seeded)\s+(?:(?:review[- ]log|review|log)\s+)?(?:entry|record|line|row)\b/i.test(plain)) continue;
|
||||
const beforeRun = /\b(?:predates|before)\s+(?:this|my)\s+(?:run|session|workflow)(?:\s+(?:started|began))?\b/i;
|
||||
const stamps = [...plain.matchAll(/\b(?:\d{4}-\d{2}-\d{2}T)?\d{2}:\d{2}(?::\d{2}(?:\.\d{1,3})?)?Z?\b/g)].map(m => m[0]);
|
||||
if (stamps.some(stamp => !sameInstant(stamp)) || (!stamps.length && !beforeRun.test(plain))) continue;
|
||||
const quoted = [...plain.matchAll(/\b([a-z_]+)["']?\s*[:=]\s*["']?([a-z0-9_.+-]+)["']?/gi)].filter(m => fields.has(m[1]!) && m[1] !== 'timestamp');
|
||||
if (!['status', 'source', 'outside_status'].every(key => quoted.some(m => m[1] === key))
|
||||
|| quoted.some(m => priorRecord[m[1]!] !== m[2])) continue;
|
||||
if (/\b(?:now|currently|current|today|new|updat\w*|append\w*|chang\w*|wrote|recorded by me)\b|\bboth reviewers agree\b/i
|
||||
.test(plain.replace(beforeRun, ''))) continue;
|
||||
spans.push({ start: sentence.index!, end: sentence.index! + sentence[0].length });
|
||||
}
|
||||
for (const span of spans.sort((a, b) => b.start - a.start)) {
|
||||
output = output.slice(0, span.start) + output.slice(span.start, span.end).replace(/[^\r\n]/g, ' ') + output.slice(span.end);
|
||||
}
|
||||
|
||||
@@ -61,7 +61,7 @@ Fixture execution boundary:
|
||||
Actions outside this interface are unsupported and fail acceptance; they are not implicitly approved.
|
||||
|
||||
Materialize qa-reports/evidence.json first, then write a concise qa-reports/report.md using the functional report structure. Link the evidence and checkpoint files rather than repeating full probe payloads in Markdown. Both artifacts are required before completion. The resulting evidence.json schema is:
|
||||
{ "revision": "<git HEAD>", "runtime": "bun <version>", "cwd": "<working directory>", "evidence": [{"command":"<exact full outer capture invocation>","contract":"README.md","expected":"<declared expected behavior>","classification":"pass|product-defect|setup-blocked|inconclusive","observed":<complete unchanged JSON emitted by the native probe>}], "learning":[{"observationCommand":"<earlier full capture invocation>","hypothesis":"<what it taught you to challenge>","nextCommand":"<later full capture invocation>"}], "limits":["<untested or blocked coverage>"] }
|
||||
{ "revision": "<full 40-character git rev-parse HEAD>", "runtime": "bun <version>", "cwd": "<working directory>", "evidence": [{"command":"<exact full outer capture invocation>","contract":"README.md","expected":"<declared expected behavior>","classification":"pass|product-defect|setup-blocked|inconclusive","observed":<complete unchanged JSON emitted by the native probe>}], "learning":[{"observationCommand":"<earlier full capture invocation>","hypothesis":"<what it taught you to challenge>","nextCommand":"<later full capture invocation>"}], "limits":["<untested or blocked coverage>"] }
|
||||
Evidence rows contain ONLY complete JSON actually emitted by native probes, including failures and repeats; retain pre-repair results alongside green results. Never synthesize JSON from a tool error or raw test output. Put tests, raw CLI diagnostics, launch failures and timeouts in Markdown with their actual output and limits. The learning array is a summary: choose one completed checkpoint where an observation motivated a different later command, not the required same-command replay. Select that checkpoint ID in annotations.learning; the production helper copies its observationCommand, hypothesis and nextCommand. Both commands must name exact captured probes with different native child commands, never a combined command list or a replay distinguished only by capture ID. This selects existing exploration evidence, not another probe or a duplicate of the complete checkpoint ledger. Preserve every checkpoint and link every checkpoint in Markdown; keep every executed probe and its complete JSON in evidence, including the required replay. Missing dependencies remain setup blockers, not repairs. No browser installation or execution is needed.`;
|
||||
}
|
||||
|
||||
|
||||
@@ -89,7 +89,7 @@ async function exercise(mode: 'success' | 'max-turns' | 'first-timeout' | 'secon
|
||||
expect(opts.signal.aborted).toBe(false);
|
||||
// Bind the complete actual compact-delivery prompt, not selected snippets.
|
||||
expect(new Bun.CryptoHasher('sha256').update(opts.prompt).digest('hex'))
|
||||
.toBe('2fa957ab9d56850a1629a845d6fe0ee5a1cb7c0843ab6555b621971d270604cb');
|
||||
.toBe('7f2dbb6b2671588e7fbc83cfecc2869c59c388b45b627ac9d481c4d917d1e76f');
|
||||
expect(opts.testName).toBe(id); expect(opts.maxTurns).toBe(15); expect(opts.timeout).toBe(CAPTURE_MS);
|
||||
for (const key of ['model', 'tools', 'allowedTools', 'appendSystemPrompt', 'env']) expect(opts).not.toHaveProperty(key);
|
||||
expect(opts.prompt).toContain('Review the plan in ./plan.md');
|
||||
@@ -98,7 +98,8 @@ async function exercise(mode: 'success' | 'max-turns' | 'first-timeout' | 'secon
|
||||
expect(opts.prompt).toContain('preserve the unresolved-decisions pass');
|
||||
expect(opts.prompt).toContain('interaction state table, empty states, responsive behavior');
|
||||
expect(opts.prompt).toContain('full required review report');
|
||||
expect(opts.prompt).toContain('Write before publishing a completed walkthrough');
|
||||
expect(opts.prompt).toContain('(or one Write) before publishing a completed walkthrough');
|
||||
expect(opts.prompt).toContain('Save as you go in three Edits: after passes 1-3, apply their decisions to plan.md');
|
||||
expect(opts.prompt).toContain('Read plan.md back to verify the saved changes');
|
||||
expect(opts.prompt).toContain('Then return a brief, concrete summary');
|
||||
expect(opts.prompt).toContain('execute every required pass and lazy-section Read');
|
||||
@@ -109,7 +110,7 @@ async function exercise(mode: 'success' | 'max-turns' | 'first-timeout' | 'secon
|
||||
expect(opts.prompt).toContain('concise score rationales and 10/10 explanations');
|
||||
const ordered = ['Read every lazy section', 'Review all 7 design passes',
|
||||
'EDIT plan.md', 'Keep the saved review compact',
|
||||
'Persist that complete plan and review with Write', 'Read plan.md back',
|
||||
'Persist that complete plan and review with those Edits', 'Read plan.md back',
|
||||
'Then return a brief, concrete summary'].map(text => opts.prompt.indexOf(text));
|
||||
expect(ordered.every(index => index >= 0)).toBe(true);
|
||||
expect(ordered).toEqual([...ordered].sort((a, b) => a - b));
|
||||
|
||||
@@ -122,6 +122,8 @@ mock.module(${JSON.stringify(path.join(ROOT, 'test/helpers/claude-pty-runner.ts'
|
||||
expect(opts.isLastStep0AUQ(target)).toBe(false);
|
||||
expect(opts.isLastStep0AUQ(fp('focus', focus))).toBe(true);
|
||||
expect(opts.isLastStep0AUQ(paraphrase)).toBe(true);
|
||||
// Census 36633323521 gate-census-7: the same Step 0D menu titled with "specific ones".
|
||||
expect(opts.isLastStep0AUQ(fp('focus-ones', 'D1 — Review all 7 design dimensions, or focus on specific ones?'))).toBe(true);
|
||||
if (mode.startsWith('native-')) expect(opts.isLastStep0AUQ(nativeFocus)).toBe(true);
|
||||
for (const unrelated of ['Review all 4 design passes, or focus?',
|
||||
'Review all 7 engineering passes, or focus?', 'Review all 7 passes, or focus?',
|
||||
|
||||
@@ -462,8 +462,8 @@ Review the plan in ./plan.md. Its design gaps are vague "clean, modern UI" and "
|
||||
|
||||
Use this non-interactive delivery sequence:
|
||||
1. Skip the preamble bash block and any AskUserQuestion calls. Read every lazy section the workflow requires. Review all 7 design passes. Rate each scored design dimension 0-10 and explain what would make it a 10; preserve the unresolved-decisions pass and every required design decision.
|
||||
2. EDIT plan.md with the missing design decisions (interaction state table, empty states, responsive behavior, etc.) and the full required review report. Keep the saved review compact: use the canonical tables and decision IDs. Specify each design requirement once; refer to its section or decision ID from other pass rationales, tasks, and report cells instead of repeating that specification. Give concise score rationales and 10/10 explanations. Retain all required report fields, design decisions, diagrams, ratings, and explanations.
|
||||
3. Persist that complete plan and review with Write before publishing a completed walkthrough or saying a fix is applied. Read plan.md back to verify the saved changes.
|
||||
2. EDIT plan.md with the missing design decisions (interaction state table, empty states, responsive behavior, etc.) and the full required review report. Save as you go in three Edits: after passes 1-3, apply their decisions to plan.md; after passes 4-7, apply theirs; then add the review report. Keep the saved review compact: use the canonical tables and decision IDs. Specify each design requirement once; refer to its section or decision ID from other pass rationales, tasks, and report cells instead of repeating that specification. Give concise score rationales and 10/10 explanations. Retain all required report fields, design decisions, diagrams, ratings, and explanations.
|
||||
3. Persist that complete plan and review with those Edits (or one Write) before publishing a completed walkthrough or saying a fix is applied. Read plan.md back to verify the saved changes.
|
||||
4. Then return a brief, concrete summary of the design changes; do not repeat the full review in the response. This changes presentation only: execute every required pass and lazy-section Read.
|
||||
|
||||
IMPORTANT: Do NOT try to browse any URLs or use a browse binary. This is a plan review, not a live site audit.`,
|
||||
|
||||
@@ -29,7 +29,7 @@ const designFocusBoundary = (fp: AskUserQuestionFingerprint): boolean =>
|
||||
// Require the source Step 0D question or its retained native paraphrase.
|
||||
// A target menu can mention a design system without reviewing this plan.
|
||||
return /^I(?:['’]ve| have) rated this plan (?:10(?:\.0+)?|[0-9](?:\.\d+)?)\/10 on design completeness\.[\s\S]*\bWant me to focus on specific areas instead of all 7\?/i.test(text)
|
||||
|| /^Review all 7 design (?:dimensions|passes), or focus(?: on specific areas)?\?$/i.test(text.split(/\r?\n/, 1)[0]!);
|
||||
|| /^Review all 7 design (?:dimensions|passes), or focus(?: on [^?\n]+)?\?$/i.test(text.split(/\r?\n/, 1)[0]!);
|
||||
});
|
||||
|
||||
// Require a choice about the supplied UI, not a workflow offer after focus.
|
||||
|
||||
Reference in new issue
Block a user