mirror of
https://github.com/garrytan/gstack.git
synced 2026-08-24 06:52:31 +02:00
* feat: model taxonomy gains gpt-5.6-sol + per-host generation defaults
Adds 'gpt-5.6-sol' to the model taxonomy with exact-match-only resolution
(Terra/Luna/suffixed IDs deliberately fall back to generic gpt) and replaces
the hardcoded 'claude' generation default with a validated
HostConfig.defaultModel: codex renders the gpt profile when --model is
absent, every other host keeps claude. Codex ship golden regenerated
accordingly; ADDING_A_HOST documents the new field.
* feat: gpt-5.6-sol bounded-scope overlay + scope-aware resolvers
The Sol profile pins the explicit task as the lake: adjacent work is
report-only, investigation is bounded, runs terminate on one clean
verification pass, and the AskUserQuestion decision-brief format is never
trimmed. The overlay wrapper grants scope-interpretation precedence while
concrete workflow steps, gates, and skill-mandated re-verification loops
still win. Sol-specific Completeness Principle and first-run intro copy.
New SETUP_COMMAND resolver renders './setup --host <host>' for every
non-claude host so generated upgrade skills reinstall their own host.
* feat: setup reads the Codex model from config.toml
New resolve-codex-generation-model.ts reads the top-level model from
${CODEX_HOME:-~/.codex}/config.toml, validates against the model allowlist,
strips control characters from every config-derived string it surfaces,
guards against non-absolute config locations, and warns on Sol near-misses.
setup runs it on EVERY invocation (read-only TOML lookup) so a plain
./setup can never clobber a Sol user's rendered profile with the hardcoded
fallback; --model <id> overrides for one run and prints the persistence
hint. Kiro installs render the claude profile before copying (Kiro fronts
Claude-family models), rewrite the baked setup command to --host kiro, and
restore the resolved Codex profile after; the codex skills path honors
CODEX_HOME. Static pins cover the resolver wiring, fail-closed exit,
quoted argv, and the Kiro sandwich.
* feat: hermetic Codex runner hardening + Sol scope-termination E2E
The Codex E2E runner copies auth.json only (operator plugins, MCP servers,
rules, and skills no longer leak into hermetic evals), pins CODEX_HOME to
the temp dir, and supports per-run model, TOML overrides, and
--ignore-user-config. New periodic E2E installs the FULL generated
investigate skill on gpt-5.6-sol against a planted one-line bug with decoy
TODOs: the fix must land inside the boundary (untracked files counted via
git status --porcelain), decoys stay byte-identical, the regression oracle
survives unweakened, nothing gets committed, all within 30 tool calls.
The shared .agents tree is snapshotted and restored exactly in beforeAll;
fixture commits disable gpg signing. Wired into the periodic CI matrix,
paid-shard globs, eval scripts, touchfiles/E2E_TIERS
(codex-sol-scope-termination), and diff-based selection. Real-file
periodic-tier classification pins both codex E2Es out of the gate tier.
Free-tier test proves an explicit --model overrides the host default
through the real generation CLI.
* chore: bump version and changelog (v1.67.2.0)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* docs: post-ship documentation sync for v1.67.2.0
- README: Codex skills path is CODEX_HOME-aware; state that
--model overrides detection for one run only (persist via
the Codex config.toml model key)
- CONTRIBUTING: add the model-overlay axis to the per-host
config table (per-host defaultModel, override precedence)
- CLAUDE.md: eval results dir is ~/.gstack/projects/<slug>/evals/
(legacy fallback ~/.gstack-dev/evals/), matching eval-store.ts
and the eval:* CLI headers
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* docs: post-ship documentation sync (v1.67.2.0)
Sol exact-match and near-miss warning documented in README; CODEX_HOME-aware
uninstall and troubleshooting paths; hermetic auth.json-only detail and the
build-clobber gotcha in CLAUDE.md; eval-store location corrected in
ARCHITECTURE.md; defaultModel row in the ADDING_A_HOST field reference;
resolver test count corrected in the CHANGELOG entry.
---------
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
71 lines
2.8 KiB
TypeScript
71 lines
2.8 KiB
TypeScript
/**
|
|
* Model overlay resolver — reads model-overlays/{model}.md and returns it
|
|
* wrapped in a subordinate behavioral-patch section.
|
|
*
|
|
* Precedence:
|
|
* 1. Exact match: ctx.model === 'gpt-5.4' → reads model-overlays/gpt-5.4.md
|
|
* 2. INHERIT directive: if the file's first non-whitespace line is
|
|
* `{{INHERIT:claude}}`, the resolver reads model-overlays/claude.md first
|
|
* and concatenates it ahead of the rest of this file's content.
|
|
* This lets `gpt-5.4.md` build on top of `gpt.md` without duplication.
|
|
* 3. Missing file: returns empty string (graceful degradation, no error).
|
|
* 4. No ctx.model set: returns empty string.
|
|
*
|
|
* The returned block is subordinate to skill workflow, safety gates, and
|
|
* AskUserQuestion instructions. The subordination language is part of the
|
|
* wrapper heading so it appears with every overlay regardless of file content.
|
|
*/
|
|
|
|
import * as fs from 'fs';
|
|
import * as path from 'path';
|
|
import type { TemplateContext } from './types';
|
|
|
|
const OVERLAY_DIR = path.resolve(import.meta.dir, '../../model-overlays');
|
|
|
|
const INHERIT_RE = /^\s*\{\{INHERIT:([a-z0-9-]+(?:\.[0-9]+)*)\}\}\s*\n/;
|
|
|
|
export function readOverlay(model: string, seen: Set<string> = new Set()): string {
|
|
if (seen.has(model)) return ''; // cycle guard
|
|
seen.add(model);
|
|
|
|
const filePath = path.join(OVERLAY_DIR, `${model}.md`);
|
|
if (!fs.existsSync(filePath)) return '';
|
|
|
|
const raw = fs.readFileSync(filePath, 'utf-8');
|
|
const match = raw.match(INHERIT_RE);
|
|
if (!match) return raw.trim();
|
|
|
|
const baseModel = match[1];
|
|
const base = readOverlay(baseModel, seen);
|
|
const rest = raw.replace(INHERIT_RE, '').trim();
|
|
|
|
if (!base) return rest;
|
|
return `${base}\n\n${rest}`;
|
|
}
|
|
|
|
export function generateModelOverlay(ctx: TemplateContext): string {
|
|
if (!ctx.model) return '';
|
|
|
|
const content = readOverlay(ctx.model);
|
|
if (!content) return '';
|
|
|
|
const precedence = ctx.model === 'gpt-5.6-sol'
|
|
? `The following instructions disambiguate scope for the ${ctx.model} model.
|
|
They govern ambiguous completeness words such as \`complete\`, \`full\`, \`every\`,
|
|
\`exhaustive\`, \`100%\`, and \`Boil the Ocean\`, and when to stop iterating on
|
|
work the user did not ask for. Concrete skill workflow steps, STOP points,
|
|
AskUserQuestion gates, plan-mode safety, required tests, skill-mandated
|
|
re-verification and re-review loops, and /ship review gates still win.
|
|
Never use this patch to skip a concrete requirement.`
|
|
: `The following nudges are tuned for the ${ctx.model} model family. They are
|
|
**subordinate** to skill workflow, STOP points, AskUserQuestion gates, plan-mode
|
|
safety, and /ship review gates. If a nudge below conflicts with skill instructions,
|
|
the skill wins. Treat these as preferences, not rules.`;
|
|
|
|
return `## Model-Specific Behavioral Patch (${ctx.model})
|
|
|
|
${precedence}
|
|
|
|
${content}`;
|
|
}
|