mirror of
https://github.com/garrytan/gstack.git
synced 2026-09-09 06:28:59 +02:00
Burn-in calibration: run 1 (fence tail 'when unsure, ask') overshot the plan-ceo review band at reviewCount=8; run 2 (tail mentioning 'HOW MANY questions') undershot at 1. Any ask-count language in the fence anchors the model in one direction or the other. The tail now says only: classify as interactive, then follow the skill's own decision-point instructions exactly as written. Pins updated to forbid count language in either direction. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>