diff --git a/qa-only/SKILL.md b/qa-only/SKILL.md index d82eb5b68..d3d957863 100644 --- a/qa-only/SKILL.md +++ b/qa-only/SKILL.md @@ -418,7 +418,7 @@ Read sections in full when directed; do not work from memory. | When | Read this section | |------|-------------------| -| running selected report-only baseline and exploratory probes without product or test writes | `sections/exploratory.md` relative to the installed `qa-only`/`gstack-qa-only` SKILL.md directory | +| selecting surfaces and later running the selected report-only probes without product or test writes (one Read covers both) | `sections/exploratory.md` relative to the installed `qa-only`/`gstack-qa-only` SKILL.md directory | | finalizing the report after probing stops | `sections/reporting.md` relative to the installed `qa-only`/`gstack-qa-only` SKILL.md directory | Start at Request Parameters, then follow the sections below in order. @@ -464,11 +464,11 @@ the current behavior. Reading old notes never requires writing new ones. ## Select Surfaces and Isolation -Load the shared preparation gate now (next section): complete its scope and selected-method Reads, +Load the shared preparation gate now (the exploratory STOP just below): complete its scope and selected-method Reads, await their results, and select the surfaces. Defer charters, clocks and probes to Run the Selected Checks, after report ownership and conditional browser setup below. -> **STOP.** Before running selected report-only baseline and exploratory probes without product or test writes, Read `sections/exploratory.md` relative to the installed `qa-only`/`gstack-qa-only` SKILL.md directory in full and follow it. +> **STOP.** Before selecting surfaces and later running the selected report-only probes without product or test writes (one Read covers both), Read `sections/exploratory.md` relative to the installed `qa-only`/`gstack-qa-only` SKILL.md directory in full and follow it. > Use this host's installed path, never the product working directory or another host's assets. > If missing or unreadable, report a QA setup blocker and its affected probes as blocked; continue other safe probes (independent functional/static checks). Missing/unreadable assets block required QA. @@ -478,7 +478,7 @@ For mixed Regression, the argument is the prior combined report. Resolve its fun replay evidence and browser baseline links first, then give each method its own baseline. A missing baseline blocks that surface's regression coverage, not independent checks. In mixed runs, use the user's surface order, defaulting to functional then browser. -Finish one surface's probes before starting the next surface's clock; any supplied +Each surface keeps its own clock in its own directory. Finish one surface's probes before starting the next surface's clock; any supplied absolute deadline still applies to both. Do not reset a clock when switching surfaces. ## Prepare Report Artifacts @@ -532,7 +532,7 @@ During browser discovery, observe behavior without reading source to diagnose it ### Assemble the report -After probing stops, load the finalization procedure below. Use retained evidence; +After probing stops, load the finalization procedure below. Order: exploratory §4 annotations and materialize (browser-only runs pass an empty evidence list), then this procedure, then the final report Write. Use retained evidence; this step does not authorize more probes or restart an expired clock. Do not preload reporting. To recover from an accidental early Read: If already read, issue another Read now and await its diff --git a/qa-only/SKILL.md.tmpl b/qa-only/SKILL.md.tmpl index 8ad77b792..af142b0f5 100644 --- a/qa-only/SKILL.md.tmpl +++ b/qa-only/SKILL.md.tmpl @@ -71,7 +71,7 @@ If neither exists, use git diff analysis. ## Select Surfaces and Isolation -Load the shared preparation gate now (next section): complete its scope and selected-method Reads, +Load the shared preparation gate now (the exploratory STOP just below): complete its scope and selected-method Reads, await their results, and select the surfaces. Defer charters, clocks and probes to Run the Selected Checks, after report ownership and conditional browser setup below. @@ -83,7 +83,7 @@ For mixed Regression, the argument is the prior combined report. Resolve its fun replay evidence and browser baseline links first, then give each method its own baseline. A missing baseline blocks that surface's regression coverage, not independent checks. In mixed runs, use the user's surface order, defaulting to functional then browser. -Finish one surface's probes before starting the next surface's clock; any supplied +Each surface keeps its own clock in its own directory. Finish one surface's probes before starting the next surface's clock; any supplied absolute deadline still applies to both. Do not reset a clock when switching surfaces. ## Prepare Report Artifacts @@ -137,7 +137,7 @@ During browser discovery, observe behavior without reading source to diagnose it ### Assemble the report -After probing stops, load the finalization procedure below. Use retained evidence; +After probing stops, load the finalization procedure below. Order: exploratory §4 annotations and materialize (browser-only runs pass an empty evidence list), then this procedure, then the final report Write. Use retained evidence; this step does not authorize more probes or restart an expired clock. Do not preload reporting. To recover from an accidental early Read: If already read, issue another Read now and await its diff --git a/qa-only/sections/manifest.json b/qa-only/sections/manifest.json index efe284cd6..9c32a5c2d 100644 --- a/qa-only/sections/manifest.json +++ b/qa-only/sections/manifest.json @@ -7,7 +7,7 @@ "id": "exploratory", "file": "exploratory.md", "title": "Report-only exploratory QA", - "trigger": "running selected report-only baseline and exploratory probes without product or test writes" + "trigger": "selecting surfaces and later running the selected report-only probes without product or test writes (one Read covers both)" }, { "id": "reporting", diff --git a/qa/references/issue-taxonomy.md b/qa/references/issue-taxonomy.md index 796815a20..a52b0990a 100644 --- a/qa/references/issue-taxonomy.md +++ b/qa/references/issue-taxonomy.md @@ -77,7 +77,7 @@ For each page visited during a QA session: 1. **Visual scan** — Take a screenshot (the Read-a-page script; `annotatedScreenshot(pg)` when you need ref labels). Look for layout issues, broken images, alignment. 2. **Interactive elements** — Click every button, link, and control. Does each do what it says? -3. **Forms** — Fill and submit (non-local target: consent first — rule 13). Test empty submission, invalid data, edge cases (long text, special characters). +3. **Forms** — Fill and submit (non-local target: consent first — browser rule 3). Test empty submission, invalid data, edge cases (long text, special characters). 4. **Navigation** — Check all paths in/out. Breadcrumbs, back button, deep links, mobile menu. 5. **States** — Check empty state, loading state, error state, full/overflow state. 6. **Console** — Print `CONSOLE_ERRORS=` after interactions. Any new JS errors or failed requests? diff --git a/test/skill-llm-eval.test.ts b/test/skill-llm-eval.test.ts index af9d1774c..065c6e700 100644 --- a/test/skill-llm-eval.test.ts +++ b/test/skill-llm-eval.test.ts @@ -203,7 +203,8 @@ describeIfSelected('QA skill quality evals', ['qa/SKILL.md workflow', 'qa/SKILL. const t0 = Date.now(); const section = readWorkflowJudgeInput({ root: ROOT, skillPath: 'qa/SKILL.md', startMarker: '# /qa: Test', endMarker: null, - references: ['qa/templates/functional-report-template.md'] }).text; + // qa-patterns.md loads both browser assets; judges penalized their absence. + references: ['qa/templates/functional-report-template.md', 'qa/templates/qa-report-template.md', 'qa/references/issue-taxonomy.md'] }).text; const samples = await judgePanel(() => callJudge(`You are evaluating the quality of a QA testing workflow document for an AI coding agent.