mirror of
https://github.com/garrytan/gstack.git
synced 2026-09-01 10:50:39 +02:00
* feat(session-kind): explicit GSTACK_SESSION_KIND override; skill-start spawned gates keyed on kind (#2733) Claude Code subagents inherit the parent env byte-for-byte, so ambient markers classify them as the parent's kind and the spawned classification was unreachable outside OpenClaw. GSTACK_SESSION_KIND=spawned (step 0, spawned-only by design) lets a dispatching skill mark its subagent per command. skill-start now keys SPAWNED_SESSION and the spawned-session instruction block on the resolved kind (was raw OPENCLAW_SESSION), suppresses CONDUCTOR_SESSION for spawned sessions, gates all 11 interactive-onboarding blocks plus their ack-at-emit marker writes on kind != spawned, and adds a destructive-gate carve-out to the spawned block (conservative-continue, never prose-STOP). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(hooks): spawned-session escape in Conductor AUQ deny; override coverage in AUQ-error fallback (#2733) Hooks inherit the harness env, so a per-command GSTACK_SESSION_KIND prefix inside a subagent's bash can never reach them. Levers added: a deterministic [conductor][spawned] auto-choose deny for env-level spawned sessions (OPENCLAW_SESSION or session-wide GSTACK_SESSION_KIND), and a spawned escape sentence appended to both hooks' prose directives so a marked subagent that slips and calls AUQ resolves to auto-choose instead of prose-STOP. The sentence lives in one shared constant (hosts/claude/hooks/spawned-directive.ts) so the two paths can never drift; destructive semantics are unified to conservative-continue. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(ship): Step 18 marks the document-release subagent spawned — env prefix + auto-choose prompt (#2733) The dispatch prompt now (1) frames the run as a SPAWNED subagent whose LAST line is machine-parsed, (2) instructs prefixing the preamble's gstack-skill-start invocation with GSTACK_SESSION_KIND=spawned on the same command line (template bash blocks don't share exports), and (3) resolves every AUQ gate to auto-choosing the recommended option, conservative on no-recommendation, never destructive. The JSON contract gains a required "decisions" array (auto-chosen gates, printed to the ship console — never embedded in the public PR body) and a placement clause so the skill's own doc-health summary stops competing with the LAST-line JSON. Tripwire pins added; codex/factory goldens refreshed. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat(auq-format): proactive SESSION_KIND=spawned rule ordered above the Conductor rule (#2733) The spawned classification previously existed only in the failure-fallback branch — a spawned session was invited to call AskUserQuestion and reach auto-choose via the deny/error detour, and a spawned session inside a Conductor workspace hit the Conductor prose-STOP rule first. The Tool resolution list now leads with the spawned rule (auto-choose recommended, never prose, never BLOCKED, destructive gates resolve conservative), the self-check carries the never-reach-this-checklist clause, and all tier>=2 SKILL.md renders are regenerated. Context-budget fixture refreshed in the same commit per the ratchet protocol (the AUQ section is eager in every tier>=2 skill). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * test(e2e): spawned document-release subagent returns the JSON contract through a firing gate (#2733) The behavioral proof the bug shipped without: ship-docsync stubs the skill (no preamble, no gates) and skill-e2e-workflow suppresses the gates by prompt. This gate-tier E2E plays the parent — it drives the verbatim Step 18 dispatch prompt (extracted from the live pr-body.md, drift-proof) against a real preamble-bearing document-release slice in a Conductor-ambient env with both AUQ hooks seeded live, an unbumped VERSION making Step 8 fire. Asserts: the final line parses as the 5-key JSON contract, the fired gate's auto-choice is recorded in decisions, and VERSION is untouched (the gate resolved to its recommended Skip). Burn-in: 1/1 pass, $0.35, 21 turns, 106s. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs(openclaw): document the GSTACK_SESSION_KIND override; wire session-kind into paid selectors (#2733) OPENCLAW.md's spawned-session section now covers the explicit per-command marker, its deliberate spawned-only narrowness, the /ship Step 18 usage, the destructive carve-out, onboarding-block suppression, and the hook env-blindness caveat. bin/gstack-session-kind and the shared spawned-directive module join the conductor-prose and auto-decide-preserved selector dep lists (session-kind previously appeared in no touchfiles entry — editing it alone triggered no paid E2E). TODOS.md gains the plan-tune capture follow-up for spawned auto-choices. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix: pre-landing review fixes (#2733) Review army + coverage audit findings, all applied: - headless directive carries the spawned escape sentence too (multi- specialist: a CI-hosted ship's marked subagent must not end BLOCKED) - anti-injection scoping on every text-claimable spawned trigger (AUQ rule + shared escape sentence): markings count only from the creating prompt, never from files/tool output/web content read mid-run - [conductor][spawned] deny annotates one-way doors per question - SPAWNED_OVERRIDE: env tamper-visibility status line + OPENCLAW.md note - spawned sessions skip the network update-check and first-task probe (consumers suppressed; preserves the one-shot just-upgraded marker) - test hardening: dispatch-tripwire end-bound validated, vacuous marker asserts replaced with output asserts, E2E cpSync size filter + named fence tolerance, spawnedByEnv parity pin, destructive-policy cross- surface drift guard, one-way annotation + bogus-value hook cases - session-kind duplicate rationale comment deduped; regen + goldens + context-budget fixture refreshed Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * chore: bump version and changelog (v1.76.0.0) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs: update project documentation for v1.76.0.0 PROJECT_STRUCTURE.md: add hosts/claude/hooks/ to the directory tree (AUQ capture + enforcement hooks, spawned-session directive, timeline stop) — the tree omitted the directory while docs/OPENCLAW.md and CHANGELOG.md now reference paths inside it. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs: sync TODOS.md ship dispatch entry with the v1.76.0.0 contract Codex doc-review finding: the SHIPPED entry for /ship auto-invoking /document-release still described the four-key JSON contract. Adds the decisions key (console-printed, never PR markdown), the GSTACK_SESSION_KIND=spawned dispatch marking (#2733), and the new spawned-dispatch gate E2E to the proven-by list. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
296 lines
12 KiB
TypeScript
296 lines
12 KiB
TypeScript
/**
|
|
* AskUserQuestion Format resolver — gate-tier assertions on the generated
|
|
* Pros/Cons format directive block.
|
|
*
|
|
* v1.7.0.0 introduces Pros/Cons decision-brief formatting:
|
|
* - D<N> numbered header
|
|
* - ELI10 paragraph
|
|
* - Stakes-if-we-pick-wrong line
|
|
* - Recommendation line (mandatory, even for neutral posture)
|
|
* - Pros/Cons block with ✅/❌ per option, min 2 pros + 1 con, ≥40 char bullets
|
|
* - Net: synthesis line
|
|
*
|
|
* This test pins the format contract so a future edit to the resolver
|
|
* can't silently drop a rule. If the resolver stops emitting one of
|
|
* these tokens, bun test catches it in milliseconds instead of waiting
|
|
* for the weekly periodic eval to notice.
|
|
*/
|
|
import { describe, test, expect } from 'bun:test';
|
|
import * as fs from 'fs';
|
|
import * as path from 'path';
|
|
import type { TemplateContext } from '../scripts/resolvers/types';
|
|
import { HOST_PATHS } from '../scripts/resolvers/types';
|
|
import { generateAskUserFormat } from '../scripts/resolvers/preamble/generate-ask-user-format';
|
|
|
|
function makeCtx(): TemplateContext {
|
|
return {
|
|
skillName: 'test-skill',
|
|
tmplPath: 'test.tmpl',
|
|
host: 'claude',
|
|
paths: HOST_PATHS.claude,
|
|
preambleTier: 2,
|
|
};
|
|
}
|
|
|
|
describe('generateAskUserFormat — v1.7.0.0 Pros/Cons format', () => {
|
|
const out = generateAskUserFormat(makeCtx());
|
|
|
|
test('includes AskUserQuestion Format header', () => {
|
|
expect(out).toContain('## AskUserQuestion Format');
|
|
});
|
|
|
|
test('documents D-numbered header requirement', () => {
|
|
expect(out).toContain('D<N>');
|
|
expect(out).toMatch(/first question in a skill invocation is `D1`/i);
|
|
});
|
|
|
|
test('documents ELI10 requirement', () => {
|
|
expect(out).toContain('ELI10');
|
|
expect(out).toMatch(/plain English.*16-year-old/);
|
|
});
|
|
|
|
test('documents Stakes-if-we-pick-wrong line', () => {
|
|
expect(out).toContain('Stakes if we pick wrong');
|
|
});
|
|
|
|
test('documents mandatory Recommendation line', () => {
|
|
expect(out).toContain('Recommendation: <choice>');
|
|
expect(out).toMatch(/Recommendation.*ALWAYS|Recommendation \(ALWAYS\)/);
|
|
});
|
|
|
|
test('documents Pros / cons block header', () => {
|
|
expect(out).toContain('Pros / cons:');
|
|
});
|
|
|
|
test('documents ✅ pro markers with min count + min length rule', () => {
|
|
expect(out).toContain('✅');
|
|
expect(out).toMatch(/[Mm]inimum 2 pros/);
|
|
expect(out).toMatch(/40 characters|≥40 chars/);
|
|
});
|
|
|
|
test('documents ❌ con markers with min count rule', () => {
|
|
expect(out).toContain('❌');
|
|
expect(out).toMatch(/1 con per option|minimum.*1 con/i);
|
|
});
|
|
|
|
test('documents hard-stop escape with exact phrase', () => {
|
|
// "No cons — this is a hard-stop choice" may span a line break in the
|
|
// rendered resolver text; match across whitespace collapses.
|
|
expect(out).toMatch(/No cons\s+—\s+this is a\s+hard-stop choice/);
|
|
});
|
|
|
|
test('documents neutral-posture escape preserving (recommended) label', () => {
|
|
// CT1 resolution: (recommended) label STAYS on default option to preserve
|
|
// AUTO_DECIDE contract. Neutrality expressed in prose only.
|
|
expect(out).toMatch(/taste call/i);
|
|
// `s` flag makes . match newlines — the label + STAYS phrase spans a line break
|
|
expect(out).toMatch(/\(recommended\)[\s\S]*STAYS|STAYS[\s\S]*\(recommended\)/);
|
|
expect(out).toMatch(/AUTO_DECIDE/);
|
|
});
|
|
|
|
test('documents Net line for closing synthesis', () => {
|
|
expect(out).toMatch(/^Net:/m);
|
|
expect(out).toMatch(/synthesis|tradeoff/i);
|
|
});
|
|
|
|
test('documents Completeness scoring rules (coverage vs kind)', () => {
|
|
expect(out).toContain('Completeness');
|
|
expect(out).toMatch(/10 = complete/);
|
|
expect(out).toMatch(/options differ in kind, not coverage/);
|
|
});
|
|
|
|
test('documents tool_use mandate (rule 11)', () => {
|
|
expect(out).toMatch(/tool_use/);
|
|
// "not a question" spans a newline in the rendered text
|
|
expect(out).toMatch(/not a[\s\S]*question|not[\s\S]*interactive/i);
|
|
});
|
|
|
|
test('includes self-check before emitting', () => {
|
|
expect(out).toContain('Self-check before emitting');
|
|
expect(out).toMatch(/D<N> header present/);
|
|
expect(out).toMatch(/Net line closes/);
|
|
});
|
|
|
|
test('documents D-numbering as model-level not runtime state', () => {
|
|
// Codex finding #4 caveat: D-numbering is a prompt wish, not a system
|
|
// guarantee. TemplateContext has no counter. This check pins the caveat.
|
|
expect(out).toMatch(/model-level instruction|not a runtime counter|count your own/i);
|
|
});
|
|
|
|
test('per-skill override guidance preserved', () => {
|
|
expect(out).toMatch(/Per-skill instructions may add/);
|
|
});
|
|
});
|
|
|
|
describe('generateAskUserFormat — 5+ option split rule (slim inline + docs pointer)', () => {
|
|
const out = generateAskUserFormat(makeCtx());
|
|
|
|
// 5 highest-signal pins. The full rule lives in
|
|
// docs/askuserquestion-split.md; this contract only checks what the
|
|
// inline subsection MUST surface so the agent can act without
|
|
// reading the docs file for routine 5-option splits.
|
|
|
|
test('forbids dropping options to fit the 4-option cap', () => {
|
|
expect(out).toMatch(/caps every call at \*\*4 options\*\*/);
|
|
expect(out).toMatch(/NEVER\s+drop, merge, or silently defer/);
|
|
});
|
|
|
|
test('names the Include / Defer / Cut / Hold buckets', () => {
|
|
expect(out).toMatch(/A\) Include/);
|
|
expect(out).toMatch(/B\) Defer/);
|
|
expect(out).toMatch(/C\) Cut/);
|
|
expect(out).toMatch(/D\) Hold/);
|
|
});
|
|
|
|
test('specifies D<N>.k child numbering and D<N>.final summary', () => {
|
|
expect(out).toContain('D<N>.k');
|
|
expect(out).toContain('D<N>.final');
|
|
});
|
|
|
|
test('AUTO_DECIDE is gated at runtime, not just collision-resistance', () => {
|
|
expect(out).toContain('bin/gstack-question-preference');
|
|
expect(out).toContain('*-split-*');
|
|
expect(out).toContain('never AUTO_DECIDE-eligible');
|
|
});
|
|
|
|
test('points to docs/askuserquestion-split.md for the full rule', () => {
|
|
expect(out).toContain('docs/askuserquestion-split.md');
|
|
expect(out).toMatch(/Read on demand when N>4/);
|
|
});
|
|
|
|
test('regression: orphan "12." prefix removed from CJK rule', () => {
|
|
expect(out).not.toContain('12. **Non-ASCII');
|
|
expect(out).toContain('**Non-ASCII characters');
|
|
});
|
|
});
|
|
|
|
describe('generateAskUserFormat — runtime-failure prose fallback', () => {
|
|
const out = generateAskUserFormat(makeCtx());
|
|
|
|
test('documents the unavailable/failed subsection', () => {
|
|
expect(out).toMatch(/When AskUserQuestion is unavailable or a call fails/i);
|
|
});
|
|
|
|
test('carves out the auto-decide denial as NOT a failure', () => {
|
|
expect(out).toContain('[plan-tune auto-decide]');
|
|
expect(out).toMatch(/NOT a failure/i);
|
|
// and explicitly: do not fall back to prose on an auto-decide denial
|
|
expect(out).toMatch(/Do NOT[\s\S]{0,40}fall back to prose|never prose/i);
|
|
});
|
|
|
|
test('retries the errored call exactly once before degrading', () => {
|
|
expect(out).toMatch(/retry the SAME call \*\*once\*\*|retry the same call.*once/i);
|
|
// idempotency guard against double-prompting
|
|
expect(out).toMatch(/double-prompt|no answer could have surfaced/i);
|
|
});
|
|
|
|
test('branches on SESSION_KIND: spawned / headless / interactive', () => {
|
|
expect(out).toContain('SESSION_KIND');
|
|
expect(out).toMatch(/`spawned`[\s\S]*auto-choose/);
|
|
expect(out).toMatch(/`headless`[\s\S]*BLOCKED/);
|
|
expect(out).toMatch(/`interactive`[\s\S]*prose fallback/);
|
|
// empty/absent SESSION_KIND degrades to interactive
|
|
expect(out).toMatch(/empty\/absent[\s\S]{0,40}interactive/i);
|
|
});
|
|
|
|
// The mandatory triad the user explicitly required for the plain-text output.
|
|
test('prose fallback mandates the triad: issue ELI10', () => {
|
|
expect(out).toMatch(/ELI10 of the issue itself/i);
|
|
});
|
|
|
|
test('prose fallback mandates the triad: per-choice Completeness score', () => {
|
|
expect(out).toMatch(/Completeness scores per choice/i);
|
|
expect(out).toMatch(/Completeness: X\/10.*EACH choice|on EACH choice/i);
|
|
});
|
|
|
|
test('prose fallback mandates the triad: recommendation + (recommended) marker', () => {
|
|
expect(out).toMatch(/Recommendation: <choice> because/);
|
|
expect(out).toMatch(/\(recommended\)`? marker on that choice/);
|
|
});
|
|
|
|
test('prose fallback is one paragraph per choice, not a bare bullet list', () => {
|
|
expect(out).toMatch(/ONE paragraph per choice/i);
|
|
expect(out).toMatch(/never a bare bullet list/i);
|
|
});
|
|
|
|
test('prose fallback tells the user to reply with a letter, then STOP', () => {
|
|
expect(out).toMatch(/reply with a letter/i);
|
|
expect(out).toMatch(/STOP and wait/i);
|
|
});
|
|
|
|
// OV2: the former "tool_use, not prose" assertions must carry the qualifier so the
|
|
// fallback is not self-contradicting. Guards against the instruction collision
|
|
// silently returning on a future edit.
|
|
test('OV2: the Format line qualifies "not prose" with the fallback exception', () => {
|
|
expect(out).toMatch(/must be sent as tool_use, not prose — unless the documented failure fallback/);
|
|
});
|
|
|
|
test('OV2: the self-check "not writing prose" line carries the Conductor + fallback qualifiers', () => {
|
|
// After the Conductor-default-prose change, the exception is two-pronged:
|
|
// CONDUCTOR_SESSION makes prose the default, OR the documented failure fallback.
|
|
expect(out).toMatch(/not writing prose — unless `CONDUCTOR_SESSION: true`[\s\S]*OR the documented failure fallback applies/);
|
|
});
|
|
|
|
// #2733: proactive spawned rule — spawned outranks Conductor. Without it a
|
|
// spawned session's AUQ handling exists only as a failure-reaction path, and
|
|
// a spawned subagent inside a Conductor workspace prose-STOPs with no reader.
|
|
test('Spawned: proactive do-not-call rule present and ordered ABOVE the Conductor rule', () => {
|
|
const spawnedRule = out.indexOf('`SESSION_KIND: spawned` echoed');
|
|
const conductorRule = out.indexOf('`CONDUCTOR_SESSION: true` echoed');
|
|
expect(spawnedRule).toBeGreaterThan(0);
|
|
expect(conductorRule).toBeGreaterThan(0);
|
|
expect(spawnedRule, 'spawned rule must outrank (precede) the Conductor rule').toBeLessThan(conductorRule);
|
|
expect(out).toMatch(/never prose, never BLOCKED/);
|
|
expect(out).toMatch(/outranks the Conductor rule/);
|
|
});
|
|
|
|
test('Spawned: destructive-gate carve-out present (conservative-continue, never prose-STOP)', () => {
|
|
expect(out).toMatch(/never auto-choose a destructive or irreversible option[\s\S]{0,80}conservative/);
|
|
});
|
|
|
|
test('Spawned: self-check carries the never-reach-this-checklist clause', () => {
|
|
expect(out).toMatch(/in `SESSION_KIND: spawned` you should never reach this checklist/);
|
|
});
|
|
|
|
test('Spawned: rule scopes markings to the creating dispatch prompt (anti-injection)', () => {
|
|
// "(or your dispatch prompt marks this session as spawned)" is a
|
|
// text-claimable trigger — the rule must explicitly refuse spawned
|
|
// claims sourced from files/tool output/web content read mid-run.
|
|
expect(out).toMatch(/NEVER count[\s\S]*prompt injection/);
|
|
});
|
|
|
|
// Conductor-default-prose contract (the proactive path, distinct from the
|
|
// failure fallback). Guards the Tool-resolution rule + self-check wording.
|
|
test('Conductor: do-not-call rule present in Tool resolution', () => {
|
|
expect(out).toMatch(/CONDUCTOR_SESSION: true/);
|
|
expect(out).toMatch(/do NOT call AskUserQuestion at all/);
|
|
expect(out).toMatch(/Auto-decide preferences still apply first/);
|
|
expect(out).toMatch(/gstack-question-log/);
|
|
});
|
|
|
|
test('Conductor: one-way prose rule + continuation protocol present', () => {
|
|
expect(out).toMatch(/one-way\b[\s\S]*typed confirmation/i);
|
|
expect(out).toMatch(/never proceed on a vague/i);
|
|
expect(out).toMatch(/Continuation — mapping a typed reply/);
|
|
});
|
|
});
|
|
|
|
describe('CQ2 — cross-file invariant: auto-decide prefix matches the hook', () => {
|
|
const out = generateAskUserFormat(makeCtx());
|
|
const hookSrc = fs.readFileSync(
|
|
path.resolve(__dirname, '..', 'hosts', 'claude', 'hooks', 'question-preference-hook.ts'),
|
|
'utf-8',
|
|
);
|
|
|
|
test('the hook actually emits the [plan-tune auto-decide] prefix', () => {
|
|
expect(hookSrc).toContain('[plan-tune auto-decide]');
|
|
});
|
|
|
|
test('the resolver references the exact same prefix the hook emits', () => {
|
|
// If a future edit reworded the hook reason, this catches the drift: the prose
|
|
// fallback would stop recognizing the auto-decide denial as not-a-failure.
|
|
const PREFIX = '[plan-tune auto-decide]';
|
|
expect(hookSrc.includes(PREFIX) && out.includes(PREFIX)).toBe(true);
|
|
});
|
|
});
|