Skip to content

[Bug]: Saved thread is reported missing and cannot be deleted, archived, or continued after two servers share its database #14353

Description

@praveenperera

Area

apps/server and apps/desktop

Problem

A saved thread remains visible, but its controls fail with a missing-thread invariant error. The user cannot delete, archive, settle, stop, or continue it.

Related: #6097 describes desktop and background servers using the same database. This report records the resulting stuck-thread state.

Observed sequence

  1. Run the background service and desktop app as the same user with the same default T3 home.
  2. Start the desktop server before creating a thread through the background server.
  3. Stop the background server and continue through the desktop server.
  4. Try to manage that thread from a remote client.

This sequence was reconstructed from an affected installation. A clean reproduction of every step has not been tested.

Expected behavior

A saved, active thread can be deleted, archived, settled, stopped, or continued.

Actual behavior

The thread is fully stuck. Repeated attempts fail with errors of this form:

Orchestration command invariant failed (thread.delete): Thread '<thread-id>' does not exist for command 'thread.delete'.

The same missing-thread error occurred for start, interrupt, settle, and session-stop commands.

Read-only checks confirmed that the thread and its creation event were present, with no deletion event. Rejected commands were saved in that same database. This was more than an old client entry.

Likely cause

Two independent servers used one T3 home. Each server has its own in-memory command read model. The desktop server started before the other server created the thread and did not load that thread into its command state.

OrchestrationEngine loads command state at startup. requireThread checks that state, while the client can read saved projections. These paths can disagree when another server writes to the database.

Two server processes were confirmed during recovery, even when both used the same version.

Verified workaround

After a database backup, restart the server and close the duplicate desktop server on the host. Keep one server owner for the T3 home.

The normal archive RPC then succeeded for the affected thread. The thread was archived and its conversation was preserved.

Suggested checks

Impact

Blocks work and all removal controls for the affected thread.

Version and environment

macOS, T3 Code Alpha 0.0.43, a background service initially using 0.0.42, and a remote desktop client using T3 Connect. Both servers used 0.0.43 during recovery.

Activity

  1. juliusmarminge commented on Sep 30, 2026

    @juliusmarminge
    Member

    Triage

    Confirmed on current main. This is a stuck-thread bug, separate from open issue #6097.

    Two servers can share one T3 home because desktop startup only checks ports. resolveDesktopBackendPort in apps/desktop/src/app/DesktopApp.ts walks up from 3773 until it finds a free port, then starts a backend on the default home anyway. Server startup writes server-runtime.json without checking whether another live process already owns that file (apps/server/src/server.ts). That split is #6097. This issue is about what the command engine does after it happens.

    OrchestrationEngine loads its command read model once, after projection bootstrap, from ProjectionSnapshotQuery.getCommandReadModel() (apps/server/src/orchestration/Layers/OrchestrationEngine.ts). thread.delete, thread.archive, thread.settle, thread.turn.start, thread.turn.interrupt, and thread.session.stop all go through requireThread, which checks that in-memory list (apps/server/src/orchestration/commandInvariants.ts). The error in the report is that invariant failing.

    The UI doesn't use that list. Shell and thread snapshots read the projection tables on every request (apps/server/src/orchestration/http.ts), so the other server's commits stay visible while every mutation is rejected and stored as a command receipt.

    Restarting works because startup reads the projection tables again. Stopping the extra server isn't enough, because the surviving process keeps its stale model.

    The failure path doesn't reliably repair it. After a rejection, reconcileReadModelAfterDispatchFailure replays orchestration_events rows whose sequence is greater than the in-memory snapshotSequence, up to 1000 rows. It doesn't reload projections and doesn't look backward. sequence is SQLite AUTOINCREMENT, so if this process appends any event after the other server inserted thread.created, projectEvent moves snapshotSequence past that row and the creation event is skipped. From then on, delete, archive, settle, start, interrupt, and session-stop keep failing until the process restarts. A projector decode error during replay is swallowed, which also leaves the model unchanged. And retrying with the same command id returns "previously rejected" without replaying at all.

    That replay is also the wrong recovery even when it does catch up: it publishes the other process's events onto this process's event bus. #6097 already has a duplicate provider-session report that comes from this path. When a thread is missing, catch-up should refresh the command model from the projection tables (the same read a restart does) and shouldn't feed those events to provider reactors.

    #6097 should still stop a second server from opening the database. This issue should stay open for the recovery gap: a thread that exists in the projection, with a creation event and no deletion, shouldn't stay permanently missing in the process that's still serving it.

  2. added
    bugSomething is broken or behaving incorrectly.
    via-triageFiled through npx t3 triage
    on Sep 30, 2026
  3. juliusmarminge commented on Oct 2, 2026

    @juliusmarminge
    Member

    Thanks for taking the time to report this and provide the details. We revisited it during the orchestrator V2 cleanup.

    The specific permanent missing-thread invariant depended on V1 startup-loaded command state and skipped event catch-up. V2 delete and mutation commands read ProjectionStore records when dispatched, removing that exact command-model split. Shared-home multi-process safety remains a separate concern; do not claim every two-server symptom fixed.

    I’m closing this report because the implementation it targets has been replaced. That does not mean every similar symptom is fixed.

    Source reviewed.

    Invite a fresh V2 reproduction. This retires the V1 stale in-memory command-state diagnosis only; no two-server runtime test performed.

    If you still hit this on a current build, please reply with the app/server versions and the steps that reproduce it. We can reopen this if the original problem is still there.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething is broken or behaving incorrectly.via-triageFiled through npx t3 triage

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions