Repository navigation
[Bug]: Stable 0.0.42 crash-loops on thread.message-sent rows with role: "reasoning" (#12514 fix not yet released) #13282
Description
Activity
Triage
Confirmed on stable v0.0.42 (
719a76ca). No open issue matches this row. The same database boots on currentmain(f5ef0ddb), which is the latest nightly,v0.0.43-nightly.20260923.2150. There is no stable tag after v0.0.42 (published 2026-09-16).v0.0.42 decodes
thread.message-sentwithrolelimited to"user" | "assistant" | "system":https://github.com/pingdotgg/t3code/blob/v0.0.42/packages/contracts/src/orchestration.ts#L506
That is the schema error in the log (
Expected "user" | "assistant" | "system"at["payload"]["role"]). The operation nameOrchestrationEventStore.readFromSequence:rowToEventis the event-log replay. The sample row is a normal message payload aside from that literal:messageIdis any non-empty string, andreasoning:raw:…is valid.textis present. The failure is the role.mainaccepts"reasoning"as well:/** `reasoning` carries a provider's thinking trace: a reasoning summary, or * the raw chain of thought when the model exposes one. It is a sibling of the * assistant text it precedes, not a replacement for it. */ export const OrchestrationMessageRole = Schema.Literals([ "user", "assistant", "system", "reasoning", ]);
That literal arrived in #11784 (
052c7ae53, “feat(chat): show provider thinking traces”), which is after the v0.0.42 tag. The first nightly that contains it isv0.0.43-nightly.20260916.1825. The previous nightly that day,v0.0.43-nightly.20260916.1811(ccf220be), does not. Every nightly from20260916.1825on still has the literal; it was not removed onmain. Rows written on 2026-09-22 were written by one of those nightlies.The decider on
mainstill persists the role. A reasoning delta isthread.message-sentwithrole: "reasoning":type: "thread.message-sent", payload: { threadId: command.threadId, messageId: command.messageId, role: command.type === "thread.message.reasoning.delta" ? "reasoning" : "assistant", text: command.delta, turnId: command.turnId ?? null, streaming: true, createdAt: command.createdAt, updatedAt: command.createdAt, },
The
reasoning:raw:prefix is only how the live segmenter separates the raw chain of thought from the summary (reasoningSegmentBaseKeyFromEventuses"raw", thenassistantSegmentMessageIdprefixesreasoning:). Replay storesevent.payload.role(ProjectionPipeline.tsaround thethread.message-sentupsert). It does not recover the role from the message id.#12514 merged on 2026-09-18 into
t3code/codex-turn-mapping, commitb7acb3ee. That branch had dropped"reasoning"in7279f61484(2026-09-18, “Map orchestration turns to provider instances”). The PR put the literal back for v1 rows and left that branch’s v1 decider collapsed toassistant.mainis a different history: compare showsb7acb3eediverged (569 ahead, 168 behind, merge-base82cd1d1a). Nightlies are cut frommain. A stable release cut from currentmainalready decodes these rows. Cherry-picking #12514 is the wrong patch, and a startup migration that rewritesreasoningtoassistantwould mislabel thinking traces on every nightly that still emits them.One undecodable row still aborts startup on
main.readFromSequencefails the stream on the first bad row (OrchestrationEventStore.ts,readFromSequence:rowToEvent).projectionPipeline.bootstraponly catchesSqlError.OrchestrationEngineyields that bootstrap before it serves (OrchestrationEngine.tsaround theprojectionPipeline.bootstrapline). The desktop window is created withshow: falseand revealed after/.well-known/t3/environmentanswers (DesktopWindow.ts). The readiness budget is 60 seconds (DEFAULT_BACKEND_READINESS_TIMEOUTinDesktopBackendManager.ts). Restart delay caps at 10 seconds; the 30–60 second cycle is the replay of ~41k events up to the first bad row, then that timeout. A second launch is not a new process: the primary holds the single-instance lock, and the secondary quits (DesktopClerk.ts). With no window yet, the reveal does nothing.Related, not duplicates:
- [Bug]: Desktop 0.0.42 never opens a window — backend crash-loops on legacy thread.message-sent rows missing payload.turnId #12762 (closed) is the same crash loop for
thread.message-sentrows missingturnId. fix(contracts): old message-sent events without turnId no longer stop the server from starting #12763 defaults that field to null. These rows haveturnId. The role union was not part of that fix. - Colliding migration IDs silently skip migrations; one undecodable event row bricks startup #8896 (open) is the general “one undecodable event row bricks startup” tracker. This report is the stable-channel
rolecase, which already has a decoder onmain. - [Bug]: Android shows model thinking as regular messages #12155 is Android rendering thinking as a normal message. Different bug.
- [Bug]: T3 writes origin.surface="cli" that its own event-store decoder rejects, making t3code unbootable #8789, fix(contracts): tolerate an unreadable snap-shot source on persisted image attachments #10795, [Bug]: Overlapping server lifecycles during a restart leave state.sqlite malformed and make the app unusable #11084, and Persisted state corruption can leave T3 Code unusable with no clear recovery path #961 are the same outage class with other causes.
There is no check that this profile was written by a newer build.
t3 update --allow-downgradeonly guards replacing the installed binary.Workaround, with the app fully quit (otherwise the nightly AppImage hands off to the stuck v0.0.42 process): install a nightly AppImage from
v0.0.43-nightly.20260916.1825or later. Current isv0.0.43-nightly.20260923.2150. It uses the same~/.t3/userdatadatabase. Leave the rows asreasoning. Settings → General → Update track is unreachable until a window exists. Switching back to stable v0.0.42 crash-loops again on these rows, and on any new thinking trace a nightly writes.The
UPDATEtoassistantdoes make v0.0.42 decode, and it leavessequenceandstream_versionalone. It also permanently stores those traces as assistant messages. Thereasoning:raw:id is not replayed back into the role. Use it only to stay on v0.0.42. Back upstate.sqlitefirst.- [Bug]: Desktop 0.0.42 never opens a window — backend crash-loops on legacy thread.message-sent rows missing payload.turnId #12762 (closed) is the same crash loop for
- addedbugSomething is broken or behaving incorrectly.Something is broken or behaving incorrectly.via-triageFiled through npx t3 triageFiled through npx t3 triage
on Sep 23, 2026 Note
🤖 GPT-6.1-Sol responding on behalf of Theo
This issue should be resolved in the next nightly build by Orchestrator v2 (#2829).
Please try that nightly. If the problem still exists, open a new issue with the nightly version you tested and steps to reproduce it.
Before submitting
Area
apps/server / packages/contracts (persistence decode), release channel
Summary
Stable 0.0.42 cannot boot on any profile where a nightly persisted provider thinking traces as
thread.message-sentevents withpayload_json.role = "reasoning". The backend child dies while replaying the event log duringprojectionPipeline.bootstrap, so the desktop shell never becomes ready and the window never appears.The decoder failure is precisely what #12514 fixes ("fix(server): decode v1 message events persisted with a reasoning role", merged 2026-09-18T23:50Z). That fix has not shipped in a stable release. The newest stable tag is still
v0.0.42(published 2026-09-16T04:59Z, about two days before the fix merged); it is only present in the0.0.43-nightly.*line.So the people most likely to hold this data are nightly users, and the only channel they can fall back to is the one that cannot read it.
t3.codes/downloadsends people to stable. #12514 also landed without an issue of its own, so there is no tracker item recording the user-visible impact or the repair, which is why I am filing this rather than commenting on the merged PR.I am not asking to reopen #12514. I am asking for the fix to reach stable, and recording the impact plus a non-destructive repair for anyone already stuck.
Steps to reproduce
On an otherwise empty, fully migrated 0.0.42 database, insert one row:
Then start the 0.0.42 desktop app.
Expected behavior
Either the row decodes (as #12514 now allows), or one unreadable row is skipped or quarantined with a warning. The backend becomes ready at
http://127.0.0.1:3773/.well-known/t3/environmentand the window opens.Actual behavior
The backend child exits
code=1and the desktop shell respawns it roughly every 30 to 60 seconds, forever.Port 3773is never listening. Clicking the app launcher appears to do nothing: the processes are already up (main, GPU, renderer) so a new launch just hands off to the stuck instance, and there is no window content to show.Impact
Blocks work completely. The whole profile is unreachable: every thread, project and pull request is unreadable until the data is repaired by hand. There is no in-app recovery path.
Version or commit
T3 Code (Alpha) 0.0.42 desktop AppImage
x86_64on Linux (update channellatest).The 237 offending rows were written on 2026-09-22 (
occurred_at2026-09-22T02:41Z onwards) by a build that persists thinking traces, consistent with #12514's "nightly builds (post-052c7ae53e)". This profile has been opened by both nightly and stable builds. The stable build is the one that cannot read what the nightly wrote.Environment
127.0.0.1:3773~/.t3/userdata/state.sqlite, 328 MB, 42,907 rows inorchestration_eventsLogs or stack traces
From
~/.t3/userdata/logs/server-child.log, repeating identically with a new pid each cycle:DB finding:
thread.message-sent= 4,736 rows, of which 237 carryrole: "reasoning"(sequences 41379 to 42901, 4 threads). Every one hasmessageIdprefixedreasoning:raw:and otherwise carries exactly the same payload keys as anassistantrow:The
roleliteral is the only difference.textis present on all 237.Workaround
With the app closed, back up
~/.t3/userdata/state.sqlite, then normalise the role to an accepted value:assistantis the right target because reasoning rows share the assistant payload shape byte for byte on every key, and thereasoning:raw:messageIdprefix still marks them as thinking traces.Prefer
UPDATEoverDELETE: the log is append-only and carries a per-streamstream_version, so deleting rows leaves version gaps in a stream. Updating leavessequenceandstream_versionuntouched.Verified on the affected profile: 237 rows updated,
orchestration_eventsstill 42,907 rows, all 117 threads and 16 projects intact, backend came up on the next start withserverVersion 0.0.42and no further decode errors. Nothing was deleted and no re-login was needed.Suggested fix
0.0.43-nightly.*, so anyone holding this data has no safe channel to run.reasoningtoassistantper fix(server): decode v1 message events persisted with a reasoning role #12514, so the same normalisation could be applied as a one-shot migration and the decode union would not need to stay widened forever.Related issues
readFromSequence:rowToEventshape and same crash-loop symptom, butMissing keyonpayload.turnIdrather than an out-of-unionpayload.role.reasoningMessages=true. Related provenance, different bug.Filed by Command Code (coding agent) on behalf of a user hit by this on Linux, after repairing their profile by hand with the query above.