Skip to content

feat(server): start Codex threads on GPT-6.1-Sol at medium effort - #543

Merged
incognitojam merged 3 commits into
mainfrom
styal/gpt-6-1-sol
Sep 29, 2026
Merged

incognitojam merged 3 commits into
mainfrom
styal/gpt-6-1-sol

Conversation

@incognitojam

@incognitojam incognitojam commented Sep 29, 2026 •

Copy link
Copy Markdown
Owner

Note

New Codex threads start on GPT-6.1-Sol at medium effort. GPT-6.1-Sol is listed as a current model with a "New" badge, and GPT-6-Sol moves to legacy models. The effort picker now marks the effort that a turn without an explicit choice actually runs at.

Problem

OpenAI released GPT-6.1-Sol today. Codex CLI 0.159.0 lists it, but styal filed it under legacy models, because the model manifest's current Codex list only knew GPT-6-Sol. New Codex threads still defaulted to GPT-6-Sol.

The effort picker also misreported the effort a new thread runs at. It marked the effort Codex's catalog calls default, which is Low for GPT-6.1-Sol and GPT-5.6-Sol. But a turn without an explicit effort sends the collaboration mode's medium fallback, which overrides the catalog default. An untouched GPT-6.1-Sol thread showed Low and ran at Medium.

Medium is also the effort to start at. In the independent FrontierCode 1.1 run through Codex, GPT-6.1-Sol scores 45.5 at low and 50.2 at medium, with no gain above medium, for $0.36 per task at medium. OpenAI's DeepSWE numbers show the same step from low to medium (64.4% to 73.0%).

Change

  • Manifest: gpt-6.1-sol replaces gpt-6-sol in the current Codex models, and a providers.codex catalog entry gives it a new badge. classifyModels now copies a catalog entry's badge onto a model that Codex discovers. The entry can only badge or classify a model; it never adds one that Codex doesn't list.
  • Default model: DEFAULT_MODEL is gpt-6.1-sol. The preference order is now GPT-6.1-Sol, GPT-6-Sol, GPT-6-Astra, then older models, so Codex CLIs before 0.159.0 keep GPT-6-Sol.
  • Default effort: mapCodexModelCapabilities marks medium as styal's default whenever the model supports it, instead of Codex's catalog default. Today that matches what untouched turns run at, because the collaboration mode falls back to the same constant. fix(server): send Codex the effort shown as default when a turn names none #544 makes the server send the shown default explicitly. Only GPT-6.1-Sol and GPT-5.6-Sol show a different default than before; the other current models already defaulted to Medium in Codex's catalog.
  • Fork docs: updated the codex-default-model ledger entry and the styal differences page. Upstream still defaults to GPT-6-Astra at 2cbc24f.

This is written for the fork rather than imported. Upstream has an open PR that adds the same manifest entry and badge support: pingdotgg#14294. Upstream intake will meet that change in ModelManifest.ts and model-manifest.json.

Installed styal servers fetch the model manifest from main, so after merge they list GPT-6.1-Sol as current and GPT-6-Sol as legacy without an update. The badge, the new default and the effort default need the new server.

Validation

  • Live run on a dev server with isolated worktree state and Codex CLI 0.159.0, using a separate CODEX_HOME so the global Codex install and its cache were untouched:
    • A new draft opens on GPT-6.1-Sol · Medium. The picker lists GPT-6.1-Sol first with a New badge, and GPT-6-Sol under legacy models.
    • A thread created with gpt-6.1-sol and no effort option answered a real prompt. The server trace shows sendTurn with provider.model: gpt-6.1-sol, and Codex's thread settings recorded effort: "medium". That is the run where the composer, before this change, showed Low.
    • With Homebrew Codex 0.158.0, which doesn't list GPT-6.1-Sol, new threads fell back to GPT-6-Sol.
  • Before: the same setup on main opens new threads on GPT-6-Sol and files GPT-6.1-Sol under legacy models.
  • Tests: ModelManifest.test.ts covers badging discovered models. CodexProvider.test.ts covers the medium effort default and the GPT-6.1-Sol → GPT-6-Sol → GPT-6-Astra preference. The Codex provider, adapter, session runtime and manifest suites pass. Three npm-update tests in CodexDriver.test.ts fail locally on this machine; they fail the same way on main.
  • Checks: server and contracts typecheck; ledger check passes.
  • Not verified: mobile. It reads the same provider snapshot and was not run.
Before (main) After
Model picker on main: GPT-6-Sol default, GPT-6.1-Sol filed under legacy Model picker: GPT-6.1-Sol first with a New badge, GPT-6-Sol under legacy

Effort picker for GPT-6.1-Sol with Medium marked default


Written by an agent (Claude Code, claude-opus-5-5).

incognitojam added a commit that referenced this pull request Sep 29, 2026
… none (#544)

A Codex turn where the user never picked an effort ran at medium,
whatever the effort picker showed as default. The composer only sends
options the user chose explicitly, so the adapter sent no `effort` on
`turn/start`. The turn's collaboration mode then filled in a hardcoded
`medium` (`input.effort ?? "medium"` in `CodexSessionRuntime`). The
picker showed Codex's catalog default, which is Low for GPT-5.6-Sol and
GPT-6.1-Sol. Upstream has the same behaviour.

## Fix

When a turn's selection names no reasoning effort, the Codex adapter now
sends the effort the provider snapshot marks as that model's default.
Clients show the same value, so the label and the turn agree. The driver
builds the snapshot before the adapter and gives the adapter that
lookup. An explicit choice still wins. A model the snapshot has no
effort options for, such as a custom model, keeps the old fallback.

This is independent of #543,
which changes the default shown for Codex models to medium where
supported. With both, a new GPT-6.1-Sol thread shows Medium and sends
Medium explicitly.

Added the `codex-default-effort` fork ledger entry.

## Validation

- **Live run** on a dev server with isolated state and Codex CLI
0.158.0. I created a GPT-5.6-Sol thread with no effort option; the
effort picker shows Low for that model. Codex's thread settings for the
turn recorded `effort: "low"` and `reasoning_effort: "low"`, and the
agent replied. Before the change, an untouched GPT-6.1-Sol turn whose
picker showed Low recorded `medium`.
- **Tests:** a new `CodexAdapter.test.ts` case sends turns with no
effort, an explicit `high`, and a model with no default. The turns send
the default, `high`, and nothing, respectively. The `CodexAdapter`,
`CodexProvider` and `CodexSessionRuntime` suites pass. Three npm-update
tests in `CodexDriver.test.ts` fail locally on this machine and fail the
same way on `main`.
- **Checks:** server typecheck and ledger check pass.

---
Written by an agent (Claude Code, claude-opus-5-5).
GPT-6-Sol moves to legacy models. Manifest catalog entries can now badge
models that Codex discovers, so GPT-6.1-Sol is marked new.
GPT-6-Sol stays the first fallback for Codex CLIs that do not offer
GPT-6.1-Sol. A turn without an explicit effort runs at the collaboration
mode's medium fallback, so the effort picker now marks medium as the
default instead of the catalog's own default.
@incognitojam
incognitojam merged commit c396e5a into main Sep 29, 2026
24 checks passed
@incognitojam
incognitojam deleted the styal/gpt-6-1-sol branch September 29, 2026 21:09
incognitojam pushed a commit that referenced this pull request Sep 29, 2026
…ingdotgg#9921)

Fork adaptation: Fork #544 builds the Codex adapter after the provider snapshot so it can read each model's default effort, so the adapter and the text generation that now reads the snapshot's models are both created after it. The manifest's legacy check keeps fork #543's catalog lookup helper and adds the Codex family fallback to it.

(cherry picked from commit a92161a)
Upstream-PR: 9921
incognitojam pushed a commit that referenced this pull request Sep 29, 2026
…11347)

Co-authored-by: maria-rcks <254055478+maria-rcks@users.noreply.github.com>

Fork adaptation: Already present. The fork chooses its own Codex and Claude defaults (codex-default-model: GPT-6.1-Sol, then GPT-6-Sol and GPT-6-Astra; Claude Opus 5.5 from fork #405), fork #405 already gave Fable 5.1 a medium default effort, and fork #543 makes medium the default effort for every Codex model that offers it, which covers GPT-6-Astra. This commit records provenance only.

(cherry picked from commit 1f73a89)
Upstream-PR: 11347
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant