Skip to content

Windows Desktop: turns initiated while session window is unfocused (cross-session send_message delivery) never start; watchdog-killed after ~980s — regression 2.1.222 -> 2.1.227 #86088

Description

@007ArunSharma

Summary

After Claude Desktop auto-updated its bundled Claude Code runtime 2.1.222 → 2.1.227 (Windows), any session turn initiated while the session's window is not focused — most notably the delivery of a cross-session message sent via the ccd_session_mgmt send_message tool — never starts. The spawned/resumed session produces no output at all (seconds_since_stderr=never, hadFirstResponse=false) and the app watchdog kills the cycle after ~980s (~16 minutes), surfacing as an error in the UI. Turns initiated with the window focused (user typing directly, new-session launches) work normally on the same runtime.

This breaks multi-session coordination workflows (sessions messaging each other) that worked reliably on 2.1.222 and earlier.

Environment

  • OS: Windows 11 Pro 10.0.26200
  • App: Claude Desktop (Cowork/CCD), Electron updater version 1.28929.0
  • Claude Code runtime: 2.1.227 (auto-installed by the app on 2026-08-12, ~05:55 UTC), previous 2.1.222
  • Sessions run in git worktrees; 12 MCP servers configured (issue reproduces regardless — the child emits nothing before MCP init would even log)

Regression evidence (from %APPDATA%\Claude\logs\main.log)

  • On 2.1.222 (Aug 9–11): 103 session spawns, zero timed out ... hadFirstResponse=false events; cross-session message delivery to idle sessions worked for weeks.
  • 2.1.227 installed:
    2026-08-12 11:25:59 [info] [CCD] Downloading from https://downloads.claude.ai/claude-code-releases/2.1.227/win32-x64/claude.exe.zst
    2026-08-12 11:26:24 [info] [CCD] Installed at C:\Users\<user>\AppData\Roaming\Claude\claude-code\2.1.227\claude.exe
    
  • First failure begins ~40 min later. 5/5 background message deliveries after the update were killed with the identical signature; examples:
    2026-08-12 12:22:59 [warn] [CCD] Session local_8af63ba7-... timed out after 1010s of inactivity (hadFirstResponse=false, last_message_type=user, last_tool_name=none, seconds_since_stderr=never)
    2026-08-12 16:15:26 [warn] [CCD] Session local_1ec13f4b-... timed out after 980s of inactivity (hadFirstResponse=false, last_message_type=user, last_tool_name=none, seconds_since_stderr=never)
    2026-08-12 16:26:26 [warn] [CCD] Session local_d385c4d8-... timed out after 979s of inactivity (hadFirstResponse=false, last_message_type=user, last_tool_name=none, seconds_since_stderr=never)
    2026-08-12 16:15:26 [info] [CCD CycleHealth] unhealthy cycle for local_1ec13f4b-... (980s, hadFirstResponse=false, reason=no_response)
    
  • A representative hung delivery (cold resume path) — after Loaded N transcript messages the child never logs anything again until the watchdog kill:
    2026-08-12 15:59:05 [info] Resuming session local_1ec13f4b-... in C:\Users\<user>\...\worktrees\...
    2026-08-12 15:59:05 [info] Starting local session local_1ec13f4b-...
    2026-08-12 15:59:06 [info] [CCD] Passing 4 plugin(s) to SDK (skills: 1, remote: 3, local: 0)
    2026-08-12 15:59:06 [info] Using Claude Code binary at: C:\Users\<user>\AppData\Roaming\Claude\claude-code\2.1.227\claude.exe
    2026-08-12 15:59:06 [info] Loaded 72 transcript messages for session local_1ec13f4b-...
    ... (total silence from this session until) ...
    2026-08-12 16:15:26 [warn] [CCD] Session local_1ec13f4b-... timed out after 980s ...
    

Analysis — focus at turn start is the discriminator

  • Not context size: both large (6,000+ message) and freshly created sessions fail identically.
  • Not warm-vs-cold: one hung delivery went to a warm process that had completed a healthy cycle 10 minutes earlier ([CCD CycleHealth] healthy cycle ... 178s, hadFirstResponse=true); it hung the same way on the unfocused delivery.
  • Focused deliveries succeed: a cross-session message delivered while the user happened to be actively focused in the receiving session's window ran a healthy 218s cycle on 2.1.227. All direct user-typed turns ran healthy all day on 2.1.227.
  • Focusing after the hang starts does not rescue it: the app pushed replaceEnabledMcpTools / replaceRemoteMcpServers to the hung session when its window was focused mid-hang; it was still killed at 980s.
  • The undelivered message is not written to the session's .jsonl transcript, but survives as a pending user message in the UI; the next focused turn processes it.

It looks like the 2.1.227 spawn/resume path waits on something only provided when a renderer/window is attached at turn start (config handshake? stdin? permission-mode negotiation?), and deadlocks silently when the turn is initiated headlessly/in background.

Repro steps

  1. Claude Desktop on Windows with runtime 2.1.227, two local sessions A and B.
  2. Leave B idle/unfocused (or even warm but unfocused).
  3. From A, send B a message via the session-management send_message tool.
  4. B's window shows the incoming message with a spinner; B's process emits nothing; ~16 min later the turn errors (timed out after ~980s ... hadFirstResponse=false ... seconds_since_stderr=never).
  5. Same message typed directly into B's focused window runs immediately.

Expected

Turns triggered by cross-session message delivery run headlessly (as on ≤2.1.222).

Workaround in use

After each cross-session send, the user manually focuses the receiving session, stops the spinner, and types a nudge — the pending message is then processed.

Activity

  1. skinner commented on Aug 12, 2026

    @skinner

    possible dupe of #86012 but I haven't seen the "focus at turn start" discriminator in the other possible dupes I've looked at

  2. arthurmoraesfernandes-afk commented on Aug 14, 2026

    @arthurmoraesfernandes-afk

    This looks like the "wedge" mechanism documented in the follow-up comment on #86298, with app-log evidence: an injected cross-session message that the CLI holds (consent gate) gets drained into the CLI at a turn boundary and is never "echoed" back as a user turn; the app then logs isRunning held by unechoed input at result … pendingEchoUuids: […] and treats the session as busy indefinitely — phantom turn, watchdog kills, and the user's own typed prompts queue/stall behind a turn that never ends until a priority=next send (typing another message force-flushes) or a restart reclaims it. The watchdog-killed unfocused-window turns match the wedged phantom turn.

  3. mimydrug commented on Aug 16, 2026

    @mimydrug

    Two more triggers (neither is cross-session), a precursor signal ~16 min before the kill, and a recovery that works

    #86557 and #86629 already establish that this persists on 2.1.229 and quantify the cross-session delivery failure, so I won't repeat that. Everything below is from a third Windows machine and covers ground I couldn't find in those three threads.

    Environment: Windows 11 Pro for Workstations 10.0.26100 · Desktop v1.26832.0 → v1.28929.0 (auto-installed 2026-08-12 10:10:23 local) → v1.30096.5 · runtime 2.1.229.

    Scale here: 14 reason=no_response events across 8 distinct sessions in 4.47 days (≈3.1/day); one session hit 3 times. Zero events in the 29 days before the update — and that zero is not an instrumentation gap: [CCD CycleHealth] and hadFirstResponse have been emitting since 2026-07-14 (2,245 cycle records), with reason=api_error verdicts appearing from 2026-07-17. Only the no_response value is new.


    1. Two triggers that are not cross-session

    All three existing threads describe cross-session send_message. I get byte-identical timeouts from two other paths:

    Trigger drained … deferred send(s) first?
    (a) cross-session send_message — as reported yes
    (b) user types while a turn is still running yes
    (c) resuming an old session no

    (b) is plain interactive use — typing the next instruction while the assistant is still working, in a single session, no MCP involved. The message is queued, drained when the turn ends, and then never mapped. If the working hypothesis is "unfocused turn initiation", (b) is worth checking against it, because the window is focused and the user is sitting there.

    (c) happens on the first action after opening a session that had completed normally earlier:

    08-14 18:06:06  [Stop hook] Query completed for session H
    08-15 13:14:21  Resuming session H
    08-15 13:14:21  Loaded N transcript messages for session H
    08-15 13:31:21  [warn] [CCD] Session H timed out after 1020s of inactivity
                    (hadFirstResponse=false, last_message_type=user,
                     last_tool_name=none, seconds_since_stderr=never)
    

    No deferred send precedes (c), so it may be a separate defect that lands in the same watchdog.


    2. Precursor signal — visible ~16 minutes before the kill

    For (a) and (b) the doomed turn is always preceded by this pair, emitted when the previous turn completes successfully:

    [LocalSessionManager] drained N deferred send(s) for <session>
    [LocalSessionManager] isRunning held by unechoed input at result for <session> { resultUuid: …, pendingEchoUuids: … }
    

    Correlation on this machine:

    drained … deferred send(s) followed by no_response within 35 min
    before v1.28929.0 10 0 / 10
    after 9 8 / 9

    The same log line existed before the regression and was harmless then. So drained looks like dequeued, not delivered, and isRunning held by unechoed input at result looks like the state that never clears. This might be the cheapest place to look in the code, and it is also usable as an early detector.


    3. The log line that is missing

    Healthy turn:

    sendMessage → Sending message to session X → Mapping internal session X to CLI session <uuid>
    

    Failed turn:

    sendMessage → Sending message to session X → (nothing)
    

    Mapping internal session … to CLI session … is simply absent — consistent with seconds_since_stderr=never and last_tool_name=none in all 14 timeouts. The child is never wired up, so nothing reaches stderr.


    4. Recovery — no restart needed

    #86557 lists the watchdog as the only exit. There is a cheaper one: send any new message to the frozen session and it maps immediately and resumes normally. A 9-character message was enough, twice. The queued message that was lost is not recovered, but the session does not have to be restarted and you do not have to wait out the ~16 minutes.

    Practical consequence: the sidebar warning icon means "the last turn failed and is still failing" — one keystroke clears it.


    5. Watchdog timing

    All 14 events cluster at 962–1020 s (mean ≈ 991 s), which is the 900 s [WarmLifecycle:session] idle timeout plus about a minute. The tight clustering is the watchdog's signature, not a property of the hang — the hang itself appears to be indefinite.


    Session identifiers redacted. Happy to provide sanitized log excerpts if that would help.

  4. pierred commented on Aug 16, 2026

    @pierred

    Bug report: cross-session messages are never delivered, and they wedge the target session in a permanent "running" state

    Known-issue note (added after research): this matches an open regression cluster on
    github.com/anthropics/claude-code — #86088, #86012, #86344, #86557, #86965, #87110
    (#86385 and #86398 closed as duplicates). Reporters pin the regression to engine
    2.1.222 → 2.1.227; it persists through 2.1.229 (this machine) and reproduces as of
    2026-08-16. No fix appears in changelogs through 2.1.233. This report is intended as a
    corroborating comment on #86088 or #86557, not a new issue.

    Environment

    • Claude desktop app: Claude 1.30096.5.0 (Windows Store install)
    • Claude Code CLI: 2.1.229 (%APPDATA%\Roaming\Claude\claude-code\2.1.229\claude.exe)
    • OS: Windows 11 Home, 10.0.26200
    • Date observed: 2026-08-16, reproduced repeatedly over ~13 hours
    • Setup: 4 concurrent local sessions in one project folder (E:\Clearwater-Construction),
      titled Engineer / Designer / Architect / Webmaster. They message each other with
      mcp__ccd_session_mgmt__send_message.

    Summary

    Two linked faults in the cross-session messaging path.

    1. Messages sent with mcp__ccd_session_mgmt__send_message are never delivered to the
      target session process.
      They appear in the app UI and in list_events, but they never
      appear in the target session's transcript, so the target model never sees them.
    2. The arrival of such a message puts the target session into a "running" state that never
      clears.
      The sidebar dot pulses, the in-chat spinner starts and counts up indefinitely,
      and every later input — including text the user types — queues behind a turn that never
      started. After a long wait the session shows a warning flag.

    The session is not computing. No turn is in flight. The app is waiting for the end of a turn
    that never began.

    Evidence

    1. Zero delivered messages, across three sessions and a full day

    Searched every raw transcript under
    C:\Users\pierr\.claude\projects\E--Clearwater-Construction\ for delivered
    cross-session-message records:

    6982849f-36d0-41ad-adc1-6450f5f4941a.jsonl (Architect)  ->  0
    e470afcf-96b1-459a-81ca-6a2426b92e02.jsonl (Engineer)   ->  0
    93546942-10ad-4a4b-9a47-5861c4939aca.jsonl (Designer)   ->  0
    

    Each file covers ~12h40m and 596–1152 records. Dozens of messages were sent between these
    sessions during that window; send_message returned success each time, with two distinct
    result strings observed: "Message sent" (target idle) and
    "Message queued for session <id>" (target flagged running).

    mcp__ccd_session_mgmt__list_events renders those same messages as pending [user] records
    for the target session. So the app-side store holds them. The session process never gets them.

    2. The app registers "activity" for a session that writes nothing and runs nothing

    Engineer (e470afcf-…), measured at 16:32Z:

    Signal Value
    Last transcript record assistant, stop_reason: end_turn, 16:19:49.506Z
    App lastActivityAt 16:31:41.726Z
    App isRunning true
    Transcript growth between those two times none
    Unmatched tool_use blocks in file 0

    The 16:31:41 "activity" is the arrival of a ping from the Designer. It is 12 minutes after the
    session's last real work. It produced no transcript record and no turn.

    Architect (6982849f-…) at the same moment: last record end_turn at 16:17:24.756Z,
    isRunning: true, no transcript growth for 15 minutes, in-chat spinner counting past 10
    minutes with two received message cards displayed above it.

    3. The stall is not a slow model turn, a tool call, an error, or a hook

    Checked across all three transcripts:

    • Every file ends on a completed assistant turn (stop_reason: end_turn).
    • 0 tool_use blocks anywhere lack a matching tool_result.
    • 0 records matching isApiErrorMessage, overloaded_error, rate_limit_error,
      Retrying, socket hang up, AbortError, context_length_exceeded, prompt is too long.
    • 0 compaction events (isCompactSummary / compactMetadata / compact_boundary).
    • No hooks are configured at any level. (A third-party plugin with a Stop hook was fully
      removed mid-investigation; the fault survived unchanged, before and after.)
    • Sampled CPU of each session process over a 30-second window during a stall: 0.34s–1.11s,
      i.e. near-idle. Note: this is weak evidence on its own, since a client waiting on the API is
      also near-idle. It is listed for completeness, not as proof.

    4. User-typed input queues behind the phantom turn

    Once a session is wedged, typing into it does not recover it. The typed message is displayed
    with an Interrupt control and a running timer, and is not processed. Observed with a typed
    check briefs sitting unprocessed for 10m58s behind two received message cards.

    Steps to reproduce

    1. Open two local sessions, A and B, in the same project folder.
    2. Let A finish a turn so it is idle.
    3. From B, call mcp__ccd_session_mgmt__send_message targeting A.
    4. Observe A: the message card appears in A's chat, the sidebar dot pulses, and a spinner
      starts counting.
    5. Observe A's transcript file under ~/.claude/projects/<project>/<sessionId>.jsonl: no new
      record is written, and no record of the message ever appears.
    6. Type anything into A. It queues behind the spinner and is not processed.
    7. The spinner runs indefinitely. A warning flag appears on the session after a long wait.

    Reproduces reliably with 3+ concurrent sessions that message each other.

    Impact

    Multi-session workflows that rely on inter-session messaging do not work at all. In our case
    three role sessions coordinate through this channel; every message silently failed for a full
    working day. The sessions appeared busy while doing nothing, which hides the failure — the
    user assumes work is in progress. Recovering each session costs a manual Interrupt.

    Secondary cost: because the wedge is only cleared by a manual Interrupt and, in our case, a
    restart, sessions get restarted often. Each restart invalidates the prompt cache. With large
    contexts this is expensive: on resume we measured cache_read_input_tokens dropping to
    35,731 while cache_creation_input_tokens spiked to 320,792, i.e. a full cold
    reprocess of ~320k tokens on the first turn after every recovery.

    Expected behaviour

    A message sent with send_message should be delivered to the target session as a user turn.
    If the target is mid-turn, it should be queued and delivered when that turn ends. Arrival of a
    message should not by itself mark a session as running.

    Workaround in use

    1. Press Interrupt on the phantom spinner to clear the running state.
    2. Do not rely on the messages. Write the content to a file in the repo, and type a prompt to
      the target session telling it to read the file.

    Root cause narrowed (from the app's own log)

    %APPDATA%\Roaming\Claude\logs\main.log shows the full mechanism. All timestamps local
    (UTC-4).

    1. Delivery spawns a fresh CLI process that never produces a first response:
      13:05:45 [info] Resuming session local_515b9419-… in E:\Clearwater-Construction
      13:05:45 [info] Loaded 76 transcript messages for session local_515b9419-…
      (no [CCD start-timing] line ever follows)
      
      A healthy start one minute earlier logged
      [CCD start-timing] … init=3913ms first_assistant=3983ms.
    2. The watchdog then kills it, always at ~1000s:
      12:19:35 [info] Sending message to session local_c6022935-…
      12:36:30 [warn] [CCD] Session local_c6022935-… timed out after 1015s of inactivity
               (hadFirstResponse=false, last_message_type=user, last_tool_name=none,
                seconds_since_stderr=never)
      12:36:30 [info] [CCD CycleHealth] unhealthy cycle … reason=no_response
      
      The session state JSON then carries errorCategory: "timeout" with the same timestamp
      — that is the UI's "stopped responding" flag.
    3. The stale busy flag is held by an echo that never arrives:
      10:49:44 [info] [LocalSessionManager] isRunning held by unechoed input at result for
               local_c6022935-… { pendingEchoUuids: [ '9d993390…' ], … }
      
    4. The hung delivery process is in a silent retry loop against the API: observed live with
      39 established TCP connections to 160.79.104.10:443 (a healthy session holds 2), CPU
      slowly climbing, zero bytes of output, zero stderr, transcript file frozen.
    5. The launch command of the hung delivery process is identical to a healthy
      user-driven resume except for the --resume=<id> value. User-typed messages through
      the same resume path always work. Only injected cross-session deliveries fail.
    6. This has been true since the first ping ever sent on this machine
      (main.log 2026-08-16 00:10:18 onward — the same timeout fingerprint repeats all day).
      The same project ran on another Windows PC where pings reportedly worked.

    Versions

    • Desktop app 1.30096.5.0, installed 2026-08-15 20:17 local. This was a re-install
      after the previous app install crashed
      earlier that day. The re-install kept all app
      data (%APPDATA%\Roaming\Claude, transcripts back to 2026-05-30) and replaced the
      binaries. The first failed ping followed ~4 hours later, at first use.
    • Bundled CLI 2.1.215 replaced by 2.1.229 92 seconds later during first-run update;
      all failures occurred on 2.1.229. npm latest at time of writing: 2.1.233.
    • Cross-session messaging worked for this user before the crash (possibly on another
      machine). No transcript on THIS machine has ever recorded a delivered message, so on
      this install the feature was broken from first use — either a regression in the
      current build or an interaction with app data carried across the re-install.

    Notes

    Nothing user-configurable appears to be involved. The fault was present with a third-party
    plugin installed and after that plugin, its files, and all of its configuration entries were
    removed; with no hooks configured at any level; and across a full app restart with fresh
    session processes. Reproduced on demand: a one-line test message to a clean idle session
    (no error state, 76-message transcript) hung its delivery process exactly as above.

  5. mimydrug commented on Aug 16, 2026

    @mimydrug

    Correction to my "recovery" claim above — and it reconciles with @pierred's opposite finding

    @pierred writes "Once a session is wedged, typing into it does not recover it", which directly contradicts point 4 of my earlier comment. I re-measured. We are both right; the difference is timing, and I stated my claim without the condition.

    The condition I left out

    When the message is sent Result
    While the spinner is running (phantom turn in flight, before the watchdog fires) swallowed — queues behind a turn that never started
    After the watchdog kill (~980s, once the session shows the warning flag) maps immediately and resumes normally

    Both of my recovery observations were the second case — the kill had already happened 10 and 23 minutes earlier:

    20:46:38  Session X timed out after 993s ...           (watchdog kill)
    20:56:16  sendMessage (9 chars) → Mapping internal session X to CLI session <uuid>   ✅ resumed
    

    So my "no restart needed" holds, but only after the ~16 minutes have already elapsed — it does not let you skip them. I should have said that. @pierred's case is the first row, and it matches trigger (b) in my earlier comment (user typing while a turn is in flight): that input is exactly what gets drained and never mapped.

    Net: there is still no known way to skip the ~16 min wait. @pierred's Interrupt on the phantom spinner is the only candidate I've seen for that; I haven't been able to test it yet, since it has to be pressed during the hang.

    One possible distinction worth checking

    @pierred's hung deliveries do spawn a process:

    Resuming session ... → Loaded 76 transcript messages → (no [CCD start-timing] line ever follows)
    

    …with 39 established TCP connections to the API and zero output.

    My (b) and (c) cases look different one step earlier — there is no process start at all:

    sendMessage → Sending message to session X → (no "Mapping internal session ... to CLI session" line)
    

    No Resuming, no Loaded N transcript messages, no [CCD start-timing], seconds_since_stderr=never.

    That may mean two distinct failure points landing in the same watchdog: one where the child spawns and then wedges against the API, and one where the child is never wired up. Worth separating if anyone is bisecting, because a fix for the retry-loop wedge would not necessarily touch the never-mapped path.

    Status here

    Still 14 events over 4.9 days on this machine (≈2.9/day) — no new ones in the ~10 hours since my previous comment, though most of that was overnight with little session use, so I would not read anything into it yet. Still no fix in changelogs through 2.1.233, consistent with @pierred's note.

  6. codebytere-ant commented on Aug 25, 2026

    @codebytere-ant

    duplicate of #86012

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions