feat(skills): claimed limitations now require evidence, everywhere + wave follow-ups filed

Every tier-2+ skill's preamble gains one directive distilled from nine live
release failures in two days on the fork: a claimed limitation or
requirement ('the API can't do this', 'X requires a credential',
'impossible on this platform') is a material claim, stated only with the
verbatim error, the documented statement, or a live probe in hand —
pattern-matching a failure to a familiar story is not evidence, and a cheap
probe runs BEFORE asking the user or declaring a step blocked. ONE directive
adapted into the preamble resolver; the fork's full judgment contract is
deliberately not imported. Full regen (46 files), ship goldens refreshed,
parity guards bumped with the measured ~0.45KB/skill (investigate, autoplan,
plan-design-review, office-hours), Step 0.9 registered as an intentional
sub-step.

Approved deferrals filed: persona-fleet hostile-user harness + answer-key
methodology in TODOS; the fork's question-budget ACCOUNTING judgment (never
its 5/8/12 constants) folded into the V1.1 pacing design doc; the Apple
adapter added to #1882's coverage note.

Ported from time-attack/gstack (GStack 2).

Co-authored-by: Sina Matian <sina@time-attack.dev>
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
This commit is contained in:
Garry Tan
2026-08-14 13:30:21 -07:00
co-authored by Sina Matian Claude Fable 5
parent 15d1144fea
commit 81a2d48592
55 changed files with 326 additions and 30 deletions
+14
View File
@@ -93,3 +93,17 @@ V2 items remain deferred:
- Per-skill or per-topic explain levels
- Team profiles
- AST-based "delivered features" metric
## Fold-in from fork port wave 2 (2026-08-14)
The time-attack/gstack fork attacked the same question fatigue from a
complementary axis: build-scale classification (session/hobby/project/
product/venture) sizing the machinery, plus CHAIN-WIDE question budgets.
Approved decision (CEO review 2026-08-14): fold the fork's ACCOUNTING
judgment into this design round — the budget is chain-scoped (a chained
review deducts from what's left, never resets), handoffs carry
questions-already-spent, approval/mutation gates never count against it, and
the budget is spent on the hardest-to-reverse decisions first. Do NOT adopt
the fork's 5/8/12 numeric constants — the fork itself later replaced them
with a zero-default autonomy dial. Scale sizes the machinery and sets the
budget; pacing (this doc) ranks what the budget is spent on.