docs: apply cross-model doc-review fixes for v1.73.0.0

- README.md: host table gains the OpenClaw explainer arm row (setup has the
  arm; the table claimed to match setup)
- docs/skills.md: /review completeness-gaps section documents the
  gstack-shortcut(dec-<id>) acknowledged-debt suppression and orphan-marker
  flagging; /autoplan deep-dive states the recommended-option default with
  the 6 principles as tie-breakers
- CONTRIBUTING.md: host count 8 -> 10 (Hermes, GBrain), supported-hosts list
  completed
- docs/TESTING_INTERNALS.md: sandbox recipe says to source ~/.bashrc after
  the doctor seeds it; GSTACK_FREE_JOBS wording fixed from "caps" to
  "overrides in either direction" (matches the un-clamped runner)

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
This commit is contained in:
Garry Tan
2026-08-29 06:29:05 +00:00
co-authored by Claude Fable 5
parent 4217d12047
commit 74549cee57
4 changed files with 11 additions and 6 deletions
+4 -2
View File
@@ -49,14 +49,16 @@ run green: run `scripts/sandbox-doctor.sh` once per boot. It documents and
treats the full failure taxonomy (missing /dev/fd, 64M /dev/shm, spurious
access(2) EACCES from the seccomp supervisor under load, full-capability
processes defeating chmod-denial tests, no X server, no git identity, and
Conductor's git-shim exit-code laundering). Then:
Conductor's git-shim exit-code laundering). The doctor seeds `TMPDIR`,
`DISPLAY`, and the runner knobs into `~/.bashrc`, so open a new shell (or
`source ~/.bashrc`) before running the suite. Then:
```bash
setpriv --ambient-caps=-all --bounding-set=-all bun run test
```
Two runner knobs exist for these environments (both no-ops unless set):
`GSTACK_FREE_JOBS` caps shard concurrency (2 is the measured sweet spot — one
`GSTACK_FREE_JOBS` overrides the shard count in either direction (2 is the measured sweet spot — one
serial mega-shard and 6-way sharding both saturate the per-process syscall
supervisor), and `GSTACK_FREE_RETRY_FLAKY=1` re-runs attributed failures once
serially, downgrading a clean retry to a loud FLAKY-PASS (capped at 5 files so
+3 -1
View File
@@ -568,6 +568,8 @@ Findings get action, not just listed. Obvious mechanical fixes (dead code, stale
`/review` now flags shortcut implementations where the complete version costs less than 30 minutes of CC time. If you chose the 80% solution and the 100% solution is a lake, not an ocean, the review will call it out.
One exception: a shortcut you took deliberately and logged. A `gstack-shortcut(dec-<id>)` marker whose decision id resolves in the decision ledger downgrades the finding to acknowledged debt. An orphan marker — one with no ledger entry behind it — doesn't suppress anything; the gap is reported normally and the marker itself gets flagged.
### Example
Suppose the smart listing flow is implemented and the tests are green.
@@ -937,7 +939,7 @@ This is my **review autopilot mode**.
Running `/plan-ceo-review`, then `/plan-design-review`, then `/plan-eng-review` individually means answering 15-30 intermediate questions. Each question is valuable, but sometimes you want the gauntlet to run without stopping for every decision.
`/autoplan` reads the review skills from disk and runs them sequentially: CEO → Design (if UI scope) → DX (if developer-facing scope) → Eng, always last — the required shipping gate reviews the final amended plan, not a stale one. It makes decisions automatically using six encoded principles (prefer completeness, match existing patterns, choose reversible options, prefer the option the user chose for similar past decisions, defer ambiguous items, and escalate security). Taste decisions (close approaches, borderline scope expansions, cross-model disagreements) get saved and presented at a final approval gate.
`/autoplan` reads the review skills from disk and runs them sequentially: CEO → Design (if UI scope) → DX (if developer-facing scope) → Eng, always last — the required shipping gate reviews the final amended plan, not a stale one. It makes decisions automatically: each question resolves to its recommended option by default, with six encoded principles (prefer completeness, match existing patterns, choose reversible options, prefer the option the user chose for similar past decisions, defer ambiguous items, and escalate security) breaking ties and deciding questions that carry no recommendation. Taste decisions (close approaches, borderline scope expansions, cross-model disagreements) get saved and presented at a final approval gate.
One command, fully reviewed plan out.