Skip to content

[Bug]: Stable 0.0.42 crash-loops on thread.message-sent rows with role: "reasoning" (#12514 fix not yet released) #13282

Description

@fabmnt

Before submitting

  • I searched existing issues and did not find a duplicate.
  • I included enough detail to reproduce or investigate the problem.

Area

apps/server / packages/contracts (persistence decode), release channel

Summary

Stable 0.0.42 cannot boot on any profile where a nightly persisted provider thinking traces as thread.message-sent events with payload_json.role = "reasoning". The backend child dies while replaying the event log during projectionPipeline.bootstrap, so the desktop shell never becomes ready and the window never appears.

The decoder failure is precisely what #12514 fixes ("fix(server): decode v1 message events persisted with a reasoning role", merged 2026-09-18T23:50Z). That fix has not shipped in a stable release. The newest stable tag is still v0.0.42 (published 2026-09-16T04:59Z, about two days before the fix merged); it is only present in the 0.0.43-nightly.* line.

So the people most likely to hold this data are nightly users, and the only channel they can fall back to is the one that cannot read it. t3.codes/download sends people to stable. #12514 also landed without an issue of its own, so there is no tracker item recording the user-visible impact or the repair, which is why I am filing this rather than commenting on the merged PR.

I am not asking to reopen #12514. I am asking for the fix to reach stable, and recording the impact plus a non-destructive repair for anyone already stuck.

Steps to reproduce

On an otherwise empty, fully migrated 0.0.42 database, insert one row:

INSERT INTO orchestration_events
 (event_id, aggregate_kind, stream_id, stream_version, event_type, occurred_at,
  command_id, causation_event_id, correlation_id, actor_kind, payload_json, metadata_json)
VALUES
 ('repro-reasoning-role', 'thread', '00000000-0000-4000-8000-000000000002', 0,
  'thread.message-sent', '2026-09-22T02:41:26.914Z', NULL, NULL, NULL, 'client',
  '{"threadId":"9d565b45-686d-4dec-9c0b-aa70f46ff053","messageId":"reasoning:raw:56432c3b-6684-4ad5-a14c-58d8047f78d2","role":"reasoning","text":"hello","turnId":"56432c3b-6684-4ad5-a14c-58d8047f78d2","streaming":true,"createdAt":"2026-09-22T02:41:26.914Z","updatedAt":"2026-09-22T02:41:26.914Z"}',
  '{}');

Then start the 0.0.42 desktop app.

Expected behavior

Either the row decodes (as #12514 now allows), or one unreadable row is skipped or quarantined with a warning. The backend becomes ready at http://127.0.0.1:3773/.well-known/t3/environment and the window opens.

Actual behavior

The backend child exits code=1 and the desktop shell respawns it roughly every 30 to 60 seconds, forever. Port 3773 is never listening. Clicking the app launcher appears to do nothing: the processes are already up (main, GPU, renderer) so a new launch just hands off to the stuck instance, and there is no window content to show.

Impact

Blocks work completely. The whole profile is unreachable: every thread, project and pull request is unreadable until the data is repaired by hand. There is no in-app recovery path.

Version or commit

T3 Code (Alpha) 0.0.42 desktop AppImage x86_64 on Linux (update channel latest).

The 237 offending rows were written on 2026-09-22 (occurred_at 2026-09-22T02:41Z onwards) by a build that persists thinking traces, consistent with #12514's "nightly builds (post-052c7ae53e)". This profile has been opened by both nightly and stable builds. The stable build is the one that cannot read what the nightly wrote.

Environment

  • Ubuntu 26.04.1 LTS, Linux x64
  • Desktop AppImage against the bundled local server on 127.0.0.1:3773
  • No systemd service
  • State: ~/.t3/userdata/state.sqlite, 328 MB, 42,907 rows in orchestration_events

Logs or stack traces

From ~/.t3/userdata/logs/server-child.log, repeating identically with a new pid each cycle:

ERROR (#3): PersistenceDecodeError: Decode error in OrchestrationEventStore.readFromSequence:rowToEvent: AnyOf(Composite(Pointer(Composite(Pointer(AnyOf())))))
    at PersistenceDecodeError.fromSchemaError (.../apps/server/dist/bin.mjs:56348:10)
    at (.../apps/server/dist/bin.mjs:56392:43)
    at (.../apps/server/dist/claudeHistoryWorker-CAn1PawV.mjs:10021:96)
    at ~effect/Utils/internal (.../apps/server/dist/claudeHistoryWorker-CAn1PawV.mjs:2558:10) {
  [cause]: SchemaError: Expected "user" | "assistant" | "system"
    at ["payload"]["role"]
}

backend child process failure output end | phase=END | details=pid=... code=1
backend child process failure output end | phase=END | details=Timed out after 60000ms waiting for desktop backend readiness at http://127.0.0.1:3773/.well-known/t3/environment.

DB finding: thread.message-sent = 4,736 rows, of which 237 carry role: "reasoning" (sequences 41379 to 42901, 4 threads). Every one has messageId prefixed reasoning:raw: and otherwise carries exactly the same payload keys as an assistant row:

createdAt, messageId, role, streaming, text, threadId, turnId, updatedAt

The role literal is the only difference. text is present on all 237.

Workaround

With the app closed, back up ~/.t3/userdata/state.sqlite, then normalise the role to an accepted value:

UPDATE orchestration_events
SET payload_json = json_set(payload_json, '$.role', 'assistant')
WHERE json_extract(payload_json, '$.role') = 'reasoning';
-- 237 rows, no deletions

assistant is the right target because reasoning rows share the assistant payload shape byte for byte on every key, and the reasoning:raw: messageId prefix still marks them as thinking traces.

Prefer UPDATE over DELETE: the log is append-only and carries a per-stream stream_version, so deleting rows leaves version gaps in a stream. Updating leaves sequence and stream_version untouched.

Verified on the affected profile: 237 rows updated, orchestration_events still 42,907 rows, all 117 threads and 16 projects intact, backend came up on the next start with serverVersion 0.0.42 and no further decode errors. Nothing was deleted and no re-login was needed.

Suggested fix

  1. Get fix(server): decode v1 message events persisted with a reasoning role #12514 into a stable release. This is the ask. Today the fix only exists on 0.0.43-nightly.*, so anyone holding this data has no safe channel to run.
  2. Backfill or repair at startup. The decider already collapses reasoning to assistant per fix(server): decode v1 message events persisted with a reasoning role #12514, so the same normalisation could be applied as a one-shot migration and the decode union would not need to stay widened forever.
  3. Do not make one undecodable row fatal. This is the same class as [Bug]: T3 writes origin.surface="cli" that its own event-store decoder rejects, making t3code unbootable #8789, Colliding migration IDs silently skip migrations; one undecodable event row bricks startup #8896 and fix(contracts): tolerate an unreadable snap-shot source on persisted image attachments #10795: an append-only log where a single row the current schema rejects makes the entire profile permanently unbootable, because the projector watermark only advances on a batch that decodes. Skipping or quarantining one bad row with a loud warning would have turned this into a cosmetic problem.
  4. Warn on the nightly to stable path. A nightly that writes data a stable decoder rejects is a downgrade hazard. A capability or schema-version check that fails with "this profile was written by a newer build" would be far clearer than a silent 30-second crash loop.

Related issues

Filed by Command Code (coding agent) on behalf of a user hit by this on Linux, after repairing their profile by hand with the query above.

Activity

  1. juliusmarminge commented on Sep 23, 2026

    @juliusmarminge
    Member

    Triage

    Confirmed on stable v0.0.42 (719a76ca). No open issue matches this row. The same database boots on current main (f5ef0ddb), which is the latest nightly, v0.0.43-nightly.20260923.2150. There is no stable tag after v0.0.42 (published 2026-09-16).

    v0.0.42 decodes thread.message-sent with role limited to "user" | "assistant" | "system":

    https://github.com/pingdotgg/t3code/blob/v0.0.42/packages/contracts/src/orchestration.ts#L506

    That is the schema error in the log (Expected "user" | "assistant" | "system" at ["payload"]["role"]). The operation name OrchestrationEventStore.readFromSequence:rowToEvent is the event-log replay. The sample row is a normal message payload aside from that literal: messageId is any non-empty string, and reasoning:raw:… is valid. text is present. The failure is the role.

    main accepts "reasoning" as well:

    /** `reasoning` carries a provider's thinking trace: a reasoning summary, or
     *  the raw chain of thought when the model exposes one. It is a sibling of the
     *  assistant text it precedes, not a replacement for it. */
    export const OrchestrationMessageRole = Schema.Literals([
      "user",
      "assistant",
      "system",
      "reasoning",
    ]);

    That literal arrived in #11784 (052c7ae53, “feat(chat): show provider thinking traces”), which is after the v0.0.42 tag. The first nightly that contains it is v0.0.43-nightly.20260916.1825. The previous nightly that day, v0.0.43-nightly.20260916.1811 (ccf220be), does not. Every nightly from 20260916.1825 on still has the literal; it was not removed on main. Rows written on 2026-09-22 were written by one of those nightlies.

    The decider on main still persists the role. A reasoning delta is thread.message-sent with role: "reasoning":

            type: "thread.message-sent",
            payload: {
              threadId: command.threadId,
              messageId: command.messageId,
              role: command.type === "thread.message.reasoning.delta" ? "reasoning" : "assistant",
              text: command.delta,
              turnId: command.turnId ?? null,
              streaming: true,
              createdAt: command.createdAt,
              updatedAt: command.createdAt,
            },

    The reasoning:raw: prefix is only how the live segmenter separates the raw chain of thought from the summary (reasoningSegmentBaseKeyFromEvent uses "raw", then assistantSegmentMessageId prefixes reasoning:). Replay stores event.payload.role (ProjectionPipeline.ts around the thread.message-sent upsert). It does not recover the role from the message id.

    #12514 merged on 2026-09-18 into t3code/codex-turn-mapping, commit b7acb3ee. That branch had dropped "reasoning" in 7279f61484 (2026-09-18, “Map orchestration turns to provider instances”). The PR put the literal back for v1 rows and left that branch’s v1 decider collapsed to assistant. main is a different history: compare shows b7acb3ee diverged (569 ahead, 168 behind, merge-base 82cd1d1a). Nightlies are cut from main. A stable release cut from current main already decodes these rows. Cherry-picking #12514 is the wrong patch, and a startup migration that rewrites reasoning to assistant would mislabel thinking traces on every nightly that still emits them.

    One undecodable row still aborts startup on main. readFromSequence fails the stream on the first bad row (OrchestrationEventStore.ts, readFromSequence:rowToEvent). projectionPipeline.bootstrap only catches SqlError. OrchestrationEngine yields that bootstrap before it serves (OrchestrationEngine.ts around the projectionPipeline.bootstrap line). The desktop window is created with show: false and revealed after /.well-known/t3/environment answers (DesktopWindow.ts). The readiness budget is 60 seconds (DEFAULT_BACKEND_READINESS_TIMEOUT in DesktopBackendManager.ts). Restart delay caps at 10 seconds; the 30–60 second cycle is the replay of ~41k events up to the first bad row, then that timeout. A second launch is not a new process: the primary holds the single-instance lock, and the secondary quits (DesktopClerk.ts). With no window yet, the reveal does nothing.

    Related, not duplicates:

    There is no check that this profile was written by a newer build. t3 update --allow-downgrade only guards replacing the installed binary.

    Workaround, with the app fully quit (otherwise the nightly AppImage hands off to the stuck v0.0.42 process): install a nightly AppImage from v0.0.43-nightly.20260916.1825 or later. Current is v0.0.43-nightly.20260923.2150. It uses the same ~/.t3/userdata database. Leave the rows as reasoning. Settings → General → Update track is unreachable until a window exists. Switching back to stable v0.0.42 crash-loops again on these rows, and on any new thinking trace a nightly writes.

    The UPDATE to assistant does make v0.0.42 decode, and it leaves sequence and stream_version alone. It also permanently stores those traces as assistant messages. The reasoning:raw: id is not replayed back into the role. Use it only to stay on v0.0.42. Back up state.sqlite first.

  2. added
    bugSomething is broken or behaving incorrectly.
    via-triageFiled through npx t3 triage
    on Sep 23, 2026
  3. t3dotgg commented on Oct 2, 2026

    @t3dotgg
    Member

    Note

    🤖 GPT-6.1-Sol responding on behalf of Theo

    This issue should be resolved in the next nightly build by Orchestrator v2 (#2829).

    Please try that nightly. If the problem still exists, open a new issue with the nightly version you tested and steps to reproduce it.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething is broken or behaving incorrectly.via-triageFiled through npx t3 triage

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions