Skip to content

Claude thread permanently fails with "No conversation found with session ID" after fresh-session fallback resumes a newly minted id #15103

Description

@areidyOTH

What happened

A long-running Claude thread (created via the T3 MCP t3_thread_* tools and driven by an orchestrator thread) suddenly started failing every turn with:

provider error. No conversation found with session ID: 13aeed5e-6b0e-4d18-bdf4-dca34e6572b0

That session ID never existed. The thread's real Claude session (9b7e437e-…) had been running for ~4 hours and its transcript is still intact in ~/.claude/projects/<worktree>/9b7e437e-….jsonl. The thread is now permanently stuck: every new message fails the same way.

Diagnosis

Sequence from server.trace.ndjson, the per-thread provider event log, and statev2.sqlite (times UTC):

  1. 04:20:55 – Claude query opened with sessionId: 9b7e437e-…; ~6,700 events of normal work follow, all with session_id: 9b7e437e-…. The agent starts a long-running background shell.
  2. 08:08:26 and 08:08:52 – Two message.dispatch turns (runs 13, 14; t3_thread_send with mode: auto then mode: queue) fail in ClaudeAdapterV2.startTurn with ClaudeBackgroundWorkBlocksQueryReplacementError ("Claude is still running background agents or commands…"). This is the guard in ClaudeAdapterV2.ts openQuery (~L6836–6843).
  3. 08:09:19 – The background task finishes (task_notification); T3 dispatches a provider-continuation turn (run 15).
  4. 08:09:24 – ProviderTurnStartService logs Provider resume failed; attempting a fresh native session with reason: "uncertain_history_delivery" (ProviderTurnStartService.ts ~L658–704). It never calls resumeThread; it fails fast because some context handoff to this provider thread has delivery.status === "pending" for the current native id, most likely left behind by the refused runs 13/14 (I could not confirm this from the DB). It then calls ensureThread with existingProviderThread: { ...providerThread, nativeThreadRef: null }, which mints a fresh native id 13aeed5e-….
  5. 08:09:27.871 – The live query is closed. 08:09:27.882 – A new query is opened with resume: "13aeed5e-…" rather than sessionId: "13aeed5e-…". The CLI answers immediately with result/error_during_execution, errors: ["No conversation found with session ID: 13aeed5e-…"]. Runs 16 and 17 repeat this.

Root cause of step 5, in ClaudeAdapterV2.ts openQuery (~L6861–6878 at 8ed276c2; L6887–6889 on current main):

const hasPersistedProviderTurn = turnInput.providerTurnOrdinal > 1;
const shouldResume =
  resumeSessionAt !== undefined || openedWithResume || hasPersistedProviderTurn;

providerTurnOrdinal > 1 is used as proof that the native session exists. But after the fresh-session fallback in ProviderTurnStartService, the provider thread keeps its turn history while getting a brand-new native id. So the adapter passes resume for a session the CLI has never seen. The fallback can never work for any thread with more than one provider turn.

Resulting persistent state: orchestration_v2_projection_provider_threads.payload_json.nativeThreadRef = { driver: "claudeAgent", nativeId: "13aeed5e-…", strength: "strong" }, so all later turns resume the non-existent id. Two context_handoffs rows (full_thread_summary for run 15, and delta_since_target_last_seen) were written with delivery.nativeThreadId: 13aeed5e-…, but the history never reached Claude.

Possible directions (not tested):

  • Have the fallback signal "fresh session" to the adapter, so openQuery uses sessionId, not resume, for a newly minted native id regardless of providerTurnOrdinal.
  • Don't persist the new nativeThreadRef as strong until the CLI confirms the session.
  • Turns refused by ClaudeBackgroundWorkBlocksQueryReplacementError probably shouldn't leave a pending handoff delivery that later forces the fresh-session path.

Steps to reproduce

Not reproduced deterministically; reconstructed from logs:

  1. Start a Claude (claudeAgent) thread and run several turns, so providerTurnOrdinal > 1.
  2. Have the agent start a long-running background shell (e.g. a until …; do sleep; done waiter run in the background).
  3. While it runs, send another message to the thread (UI or t3_thread_send). It is refused with "Claude is still running background agents or commands…". Sending a second one in queue mode is also refused.
  4. Let the background task finish, so T3 starts a provider-continuation turn.
  5. Observe Provider resume failed; attempting a fresh native session (uncertain_history_delivery) in the trace, followed by query.open with resume: <new uuid> and No conversation found with session ID: <new uuid>. Every later turn on the thread fails the same way.

A more direct check of the core bug: force the fallback path in ProviderTurnStartService (e.g. make resumeThread fail) on a Claude thread with ≥2 provider turns, and observe that the replacement query is opened with resume for the freshly minted id.

Version

Desktop AppImage 0.0.46-nightly.20261003.2610 (commit 8ed276c). Same code is present on main as of 2026-10-03.

Environment

Linux x64 (kernel 7.0.0), Node v26.8.2, Claude Code CLI 2.1.288, desktop app (server.mode: desktop, 127.0.0.1:3773)

Evidence

# server.trace.ndjson — runs 13/14 refused
08:08:26.922 ClaudeAdapterV2.startTurn Failure
  ProviderAdapterTurnStartError: Failed to start run run:thread:mcp%3Ac5639143-…:ordinal:13 …
  [cause]: ClaudeBackgroundWorkBlocksQueryReplacementError: Claude is still running background agents or commands, and this model or setting change would end them. …
08:08:55.278 ClaudeAdapterV2.startTurn Failure   (ordinal:14, same cause)

# run 15 — fallback
08:09:24.888 orchestrationV2.providerTurnStart.start
  "Provider resume failed; attempting a fresh native session"
  { driver: "claudeAgent", runId: "run:thread:mcp%3Ac5639143-…:ordinal:15",
    reason: "uncertain_history_delivery", errorTag: "ProviderAdapterTurnStartError" }

# provider event log (events.mcp-c5639143-….log)
04:20:55.366 outgoing query.open  { sessionId: "9b7e437e-d8ba-…", model: "claude-opus-5-5[1m]", ... }
   … 6746 events with session_id 9b7e437e-… …
08:09:19.648 incoming system/task_notification (background task finished)
08:09:27.871 outgoing query.close
08:09:27.882 outgoing query.open  { resume: "13aeed5e-6b0e-4d18-bdf4-dca34e6572b0", ... }
08:09:29.084 incoming result { subtype: "error_during_execution", num_turns: 0,
  errors: ["No conversation found with session ID: 13aeed5e-6b0e-4d18-bdf4-dca34e6572b0"] }
08:09:30.862 / 08:09:33.403  same query.open with resume 13aeed5e… → same error (runs 16, 17)

# statev2.sqlite (read-only)
orchestration_v2_projection_provider_threads.nativeThreadRef
  = {"driver":"claudeAgent","nativeId":"13aeed5e-6b0e-4d18-bdf4-dca34e6572b0","strength":"strong"}
# ~/.claude/projects/<worktree>/ contains 9b7e437e-….jsonl; no 13aeed5e-* file exists anywhere

Related issues

#2336 (closed): a thread becomes permanently unusable because T3 resumes a Claude session id that has no transcript (V1 resume_cursor_json, CLI killed before first write). Same symptom and the same "permanently stuck" outcome, but a different code path. This one is the V2 ProviderTurnStartService fresh-session fallback combined with ClaudeAdapterV2 deciding resume from providerTurnOrdinal > 1.

Fix applied or workaround

Nothing was changed on the machine. Workaround: the original conversation is intact, so it can be continued outside T3 with claude --resume <original session id> in the thread's worktree, or the work can be handed to a new T3 thread. The stuck thread itself cannot recover without editing statev2.sqlite.

Filed by

Claude Code (claude-opus-5-5) via t3 triage

Activity

  1. juliusmarminge commented on Oct 3, 2026

    @juliusmarminge
    Member

    Note

    Grok responding on behalf of Julius.

    Triage

    Thanks for the detailed diagnosis, @areidyOTH! This reproduces on current main (fed41fa88b, which includes 8ed276c2). It's a real bug, and it isn't a duplicate of #2336 (closed). That one was the V1 path, where a Claude session id was resumed before the CLI had written a transcript. This one is the V2 fresh-session fallback opening a newly created id with resume.

    Why the replacement query uses resume

    ClaudeAdapterV2.openQuery treats providerTurnOrdinal > 1 as proof that the native session already exists, so makeClaudeQueryOptions sends resume instead of sessionId. That's correct after an idle release of a session the CLI actually created, because reopening with sessionId fails with "already in use".

    It's wrong after ProviderTurnStartService drops nativeThreadRef and calls ensureThread. For Claude, ensureThread only allocates a new session id and keeps the same provider-thread row, so the next ordinal is still above 1. The replacement query ends up as resume: <new uuid> for a conversation the CLI has never seen.

    Why it stays stuck

    The new ref is written as strength: "strong" before startTurn runs. openedNativeThreads is only updated after a successful open, and the next turn doesn't check it anyway, because the ordinal check forces resume again. Retrying can't recover the thread. As you noted, the original transcript is still under the earlier session id. Continuing that session outside T3, or starting a new thread, is the workaround for now.

    Where uncertain_history_delivery comes from

    That branch only runs when a handoff for this provider thread is still pending on the current native id, and it never calls resumeThread.

    Claude has no history injection, so text handoffs are saved as pending before startTurn and only marked inline once the turn is accepted. ClaudeBackgroundWorkBlocksQueryReplacementError is thrown in openQuery before the prompt goes out, so a refused turn that had something to deliver leaves a pending marker behind. Run 14 fits that. Run 13 failed before a provider turn was recorded, so run 14 treated it as missed history, wrote the marker, and then failed the same way. Run 15 then abandoned the live session. A later delivery rewrite onto the new native id would replace that marker, which explains why you didn't find a pending row for the original session afterward.

    What needs to change

    Two things:

    • A newly created native id needs to be opened with sessionId, whatever the provider-turn ordinal is.
    • A turn start that never reaches the CLI shouldn't leave a pending delivery that later forces this fallback.

    A maintainer will decide on the fix direction.

  2. added
    bugSomething is broken or behaving incorrectly.
    via-triageFiled through npx t3 triage
    on Oct 3, 2026
  3. added 2 commits that reference this issue on Oct 4, 2026
    5d73e3a
    aaa260e
  4. CoreInfusion commented on Oct 4, 2026

    @CoreInfusion

    Hit the same on 0.0.46-nightly.20261003.2632 (Linux server, claudeAgent, claude-opus-5-5).

    Sequence:

    1. Messages were refused by the background-work block ([Bug]: Claude refuses every message as a "model or setting change" when a thread moves between mobile and web while background work runs #15471, which this followed directly).
    2. The user stopped the background task.
    3. On the next message, T3 recorded a provider_handoff context transfer and a full_thread_summary context handoff, set the provider thread's nativeThreadRef to a new id, and opened the Claude query with resume: <new id>.
    4. Result: No conversation found with session ID: <new id> (error_during_execution). Every later message resumes the same missing id, so the thread can't be used anymore. The previous native session's transcript was intact on disk.

    In openQuery, shouldResume becomes true through hasPersistedProviderTurn = providerTurnOrdinal > 1. That seems to count turns of the provider thread rather than of the newly created native session, so a brand-new session after a handoff is resumed instead of created.

  5. AigarsMurins commented on Oct 5, 2026

    @AigarsMurins

    This still reproduces on 0.0.46-nightly.20261005.2667 (37de6cbde65c), which includes #15770 (efecd3cf8b). It's the "rejected starts" path that #15770 left open, and the newly allocated id is still opened with resume.

    Sequence on macOS, claudeAgent, claude-opus-5-5, on a long thread that had an earlier context handoff:

    1. Run 48 wedged on an AskUserQuestion with a blank option description ([Bug]: Claude AskUserQuestion with a blank option description fails the run and wedges the thread #15316).
    2. Runs 49, 50 and 51 were rejected in ClaudeAdapterV2.startTurn with ProviderAdapterProtocolError: Claude provider turn … is still active.
    3. I restarted the app, which also updated it to 20261005.2667, and sent a message.
    4. At 06:19:34.524Z, a provider-thread.updated event (seq 25332) set a new native id, 5f16b04e-b3bb-4952-b090-05e584cd706b, and context-handoff.updated events followed.
    5. At 06:19:35Z, the provider log shows:
    {"type":"query.close"}
    {"type":"query.open","options":{"model":"claude-opus-5-5[1m]","permissionMode":"bypassPermissions","resume":"5f16b04e-b3bb-4952-b090-05e584cd706b", ...}}
    {"type":"prompt.offer","message":{... "content":"Context handoff (full_thread_summary, delta_since_target_last_seen). 5 handoff records; ..."}}
    {"type":"result","subtype":"error_during_execution","duration_ms":0,"is_error":true,"num_turns":0,"session_id":"5f16b04e-b3bb-4952-b090-05e584cd706b"}

    The UI showed No conversation found with session ID: 5f16b04e-…. The thread's real session (ac7ceaa1-…) is intact on disk and resumes fine with claude --resume.

    So two things are still wrong on a build that includes #15770:

    • The rejected starts forced a fresh session.
    • The fresh id was still opened with resume rather than sessionId.

    The thread can't be used now: every message resumes the same missing id. I can share the provider event log and the orchestration_events rows if that helps.

  6. tallboxdesign commented on Oct 5, 2026

    @tallboxdesign

    Additional occurrence on macOS; installed version at investigation: 0.0.46-nightly.20261005.2667.

    A long-running Claude conversation stopped accepting messages after working throughout the day and overnight. Local logs show that T3 replaced a working session reference during recovery, then tried to resume the replacement before it existed. The original transcript remained on disk, but subsequent resume and fork attempts failed. Please investigate session recovery and provide a supported way to reconnect the existing conversation.

    Environment

    • macOS, T3 Code Nightly.
    • Installed version at investigation: 0.0.46-nightly.20261005.2667. The exact build running at every earlier failure is unverified.
    • Claude Agent SDK with Claude Opus 5.5; the successful runtime used the 1M-context variant.
    • Long-running conversation with background commands.

    Observed sequence

    1. Claude worked throughout the day and overnight.

    2. Ordinary follow-up messages began failing with:

      Claude is still running background agents or commands, and this model or setting change would end them. Wait for them to finish, or press Stop, then send the message again.

    3. A history handoff was left marked as pending after a blocked message.

    4. On a later message, T3 logged a fresh-session fallback with the reason uncertain_history_delivery.

    5. T3 replaced the working native session reference with a newly allocated reference.

    6. It closed the previous query and opened another using resume for the new reference, rather than creating the replacement session.

    7. Claude returned “No conversation found with session ID: [redacted]”. The result reported zero turns, zero tokens, and zero API time.

    8. Later retries reused the invalid replacement reference. Fork attempts also failed.

    The original local transcript and T3 conversation history remained available. This is a failure of session continuity, not demonstrated deletion of the conversation. The survival of individual background jobs was not established.

    Separate blocker when switching providers

    Switching to Codex allowed replies. Switching back to Claude then produced:

    Insufficient context allowance for the provider handoff. Compact the target conversation or use a larger-context model; the current request has not been truncated.

    The handoff-budget calculation has not been independently reproduced, so this report does not claim that calculation is incorrect. However, it creates a recovery dead end when the target Claude session is already unavailable.

    Expected behavior

    • A message rejected before delivery should not leave history state that unnecessarily abandons a valid session.
    • A newly allocated session must be created before it can be resumed.
    • An unsuccessful replacement must not permanently displace the last working session reference.
    • Resume and fork failures should offer a supported recovery action, without manual database edits.
    • Context-allowance failures should provide an actionable compaction or concise-handoff recovery path.

    Evidence and limits

    Local provider logs, server traces, and read-only database inspection support the sequence above. The server explicitly recorded an uncertain-history-delivery fallback, followed by a resume request for the replacement session and a missing-session response. The exact internal reason the creation-versus-resume decision selected resume has not been reproduced.

    No app, session, or database files were changed during diagnosis. This report contains no work titles, project names, account identifiers, session/thread IDs, transcript excerpts, local paths, or private-work screenshots.

    Prepared by Codex (GPT-6-Astra), using local read-only investigation; posted with the user’s authorization.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething is broken or behaving incorrectly.via-triageFiled through npx t3 triage

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions