mirror of
https://github.com/garrytan/gstack.git
synced 2026-08-24 06:52:31 +02:00
v1.67.2.0 feat: gpt-5.6-sol bounded-scope profile for Codex installs (#2633)
* feat: model taxonomy gains gpt-5.6-sol + per-host generation defaults
Adds 'gpt-5.6-sol' to the model taxonomy with exact-match-only resolution
(Terra/Luna/suffixed IDs deliberately fall back to generic gpt) and replaces
the hardcoded 'claude' generation default with a validated
HostConfig.defaultModel: codex renders the gpt profile when --model is
absent, every other host keeps claude. Codex ship golden regenerated
accordingly; ADDING_A_HOST documents the new field.
* feat: gpt-5.6-sol bounded-scope overlay + scope-aware resolvers
The Sol profile pins the explicit task as the lake: adjacent work is
report-only, investigation is bounded, runs terminate on one clean
verification pass, and the AskUserQuestion decision-brief format is never
trimmed. The overlay wrapper grants scope-interpretation precedence while
concrete workflow steps, gates, and skill-mandated re-verification loops
still win. Sol-specific Completeness Principle and first-run intro copy.
New SETUP_COMMAND resolver renders './setup --host <host>' for every
non-claude host so generated upgrade skills reinstall their own host.
* feat: setup reads the Codex model from config.toml
New resolve-codex-generation-model.ts reads the top-level model from
${CODEX_HOME:-~/.codex}/config.toml, validates against the model allowlist,
strips control characters from every config-derived string it surfaces,
guards against non-absolute config locations, and warns on Sol near-misses.
setup runs it on EVERY invocation (read-only TOML lookup) so a plain
./setup can never clobber a Sol user's rendered profile with the hardcoded
fallback; --model <id> overrides for one run and prints the persistence
hint. Kiro installs render the claude profile before copying (Kiro fronts
Claude-family models), rewrite the baked setup command to --host kiro, and
restore the resolved Codex profile after; the codex skills path honors
CODEX_HOME. Static pins cover the resolver wiring, fail-closed exit,
quoted argv, and the Kiro sandwich.
* feat: hermetic Codex runner hardening + Sol scope-termination E2E
The Codex E2E runner copies auth.json only (operator plugins, MCP servers,
rules, and skills no longer leak into hermetic evals), pins CODEX_HOME to
the temp dir, and supports per-run model, TOML overrides, and
--ignore-user-config. New periodic E2E installs the FULL generated
investigate skill on gpt-5.6-sol against a planted one-line bug with decoy
TODOs: the fix must land inside the boundary (untracked files counted via
git status --porcelain), decoys stay byte-identical, the regression oracle
survives unweakened, nothing gets committed, all within 30 tool calls.
The shared .agents tree is snapshotted and restored exactly in beforeAll;
fixture commits disable gpg signing. Wired into the periodic CI matrix,
paid-shard globs, eval scripts, touchfiles/E2E_TIERS
(codex-sol-scope-termination), and diff-based selection. Real-file
periodic-tier classification pins both codex E2Es out of the gate tier.
Free-tier test proves an explicit --model overrides the host default
through the real generation CLI.
* chore: bump version and changelog (v1.67.2.0)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* docs: post-ship documentation sync for v1.67.2.0
- README: Codex skills path is CODEX_HOME-aware; state that
--model overrides detection for one run only (persist via
the Codex config.toml model key)
- CONTRIBUTING: add the model-overlay axis to the per-host
config table (per-host defaultModel, override precedence)
- CLAUDE.md: eval results dir is ~/.gstack/projects/<slug>/evals/
(legacy fallback ~/.gstack-dev/evals/), matching eval-store.ts
and the eval:* CLI headers
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* docs: post-ship documentation sync (v1.67.2.0)
Sol exact-match and near-miss warning documented in README; CODEX_HOME-aware
uninstall and troubleshooting paths; hermetic auth.json-only detail and the
build-clobber gotcha in CLAUDE.md; eval-store location corrected in
ARCHITECTURE.md; defaultModel row in the ADDING_A_HOST field reference;
resolver test count corrected in the CHANGELOG entry.
---------
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
This commit is contained in:
co-authored by
Claude Fable 5
parent
c86e6472eb
commit
60e51342b5
+33
-11
@@ -113,7 +113,7 @@ if [ -d ".agents/skills/gstack" ] && [ ! -L ".agents/skills/gstack" ]; then
|
||||
fi
|
||||
fi
|
||||
echo "VENDORED_GSTACK: $_VENDORED"
|
||||
echo "MODEL_OVERLAY: claude"
|
||||
echo "MODEL_OVERLAY: gpt"
|
||||
_CHECKPOINT_MODE=$($GSTACK_BIN/gstack-config get checkpoint_mode 2>/dev/null || echo "explicit")
|
||||
_CHECKPOINT_PUSH=$($GSTACK_BIN/gstack-config get checkpoint_push 2>/dev/null || echo "false")
|
||||
echo "CHECKPOINT_MODE: $_CHECKPOINT_MODE"
|
||||
@@ -579,23 +579,45 @@ At skill END before telemetry:
|
||||
```
|
||||
|
||||
|
||||
## Model-Specific Behavioral Patch (claude)
|
||||
## Model-Specific Behavioral Patch (gpt)
|
||||
|
||||
The following nudges are tuned for the claude model family. They are
|
||||
The following nudges are tuned for the gpt model family. They are
|
||||
**subordinate** to skill workflow, STOP points, AskUserQuestion gates, plan-mode
|
||||
safety, and /ship review gates. If a nudge below conflicts with skill instructions,
|
||||
the skill wins. Treat these as preferences, not rules.
|
||||
|
||||
**Todo-list discipline.** When working through a multi-step plan, mark each task
|
||||
complete individually as you finish it. Do not batch-complete at the end. If a task
|
||||
turns out to be unnecessary, mark it skipped with a one-line reason.
|
||||
**Completion bias.** Do not end your turn with a partial solution when the full
|
||||
solution is reachable. If you encounter an error, debug it. If a test fails, fix it.
|
||||
If something is ambiguous, make your best judgment and proceed — don't stop and ask
|
||||
unless you're genuinely blocked.
|
||||
|
||||
**Think before heavy actions.** For complex operations (refactors, migrations,
|
||||
non-trivial new features), briefly state your approach before executing. This lets
|
||||
the user course-correct cheaply instead of mid-flight.
|
||||
**Prefer doing over listing.** When you'd be tempted to write "you could also try X,
|
||||
Y, or Z," try the best option yourself. Pick, execute, report results.
|
||||
|
||||
**Dedicated tools over Bash.** Prefer Read, Edit, Write, Glob, Grep over shell
|
||||
equivalents (cat, sed, find, grep). The dedicated tools are cheaper and clearer.
|
||||
**No preamble.** Skip "Great question!", "Let me help with that", and restating the
|
||||
user's request. Start with the work.
|
||||
|
||||
**AskUserQuestion is NOT preamble.** The "No preamble" and "Prefer doing over listing"
|
||||
rules above do NOT apply to AskUserQuestion content. When you invoke AskUserQuestion,
|
||||
the user is about to make a decision — they need context, not terseness. Always emit
|
||||
the full format from the preamble's AskUserQuestion Format section:
|
||||
|
||||
1. **Re-ground** (project + branch + task — 1-2 sentences).
|
||||
2. **Simplify (ELI10)** — explain what's happening in plain English a 16-year-old could
|
||||
follow. Concrete stakes, not abstract tradeoffs. Non-negotiable; this is NOT preamble.
|
||||
3. **Recommend** — `RECOMMENDATION: Choose [X] because [one-line reason]` on its own
|
||||
line. Never omit this line. Never collapse it into the options list.
|
||||
4. **Options** — lettered `A) B) C)` with Completeness scores (coverage-differentiated)
|
||||
or the "options differ in kind" note (kind-differentiated).
|
||||
|
||||
If you find yourself about to present an AskUserQuestion without the Simplify/ELI10
|
||||
paragraph, without a RECOMMENDATION line, or by just listing options and asking "which
|
||||
one?" — stop, back up, and emit the full format. The user will ask you to do it anyway,
|
||||
so do it the first time.
|
||||
|
||||
**Reminder: subordination applies.** When a skill workflow says STOP, stop. When the
|
||||
skill asks via AskUserQuestion, that is the wait-for-user gate, not an ambiguity.
|
||||
Completion bias does not override safety gates.
|
||||
|
||||
## Voice
|
||||
|
||||
|
||||
Reference in New Issue
Block a user