Skip to content

Session resume silently creates new session despite valid session file and correct parameters #555

Description

@RomanLuo523

When attempting to resume a session with a valid resume parameter and an existing session file, the SDK silently creates a new session instead of resuming the expected one. No error or warning is logged
by the SDK.

Environment

  • SDK Version: 0.1.29 (also tested concept with 0.1.31)
  • Python Version: 3.12
  • Platform: Linux (Kubernetes)
  • Storage: Shared storage (NFS) mounted at ~/.claude

Steps to Reproduce

  1. Create a new session, SDK returns session_id = "61568b3d-8671-487f-a553-7647fbffff5e"
  2. Save the session_id for later resume
  3. Wait ~27 seconds (not a concurrency issue)
  4. Resume the session with the following options:
    options = ClaudeAgentOptions(
        resume="61568b3d-8671-487f-a553-7647fbffff5e",
        fork_session=False,
        continue_conversation=True,
        # ... other options
    )
  5. SDK returns a different session_id = "220a71ec-c933-4091-ac05-dc2eb8366c00"

Expected Behavior

SDK should resume the existing session and return the same session_id (61568b3d-...), or raise an error/warning if resume fails.

Actual Behavior

SDK silently creates a new session with a different session_id (220a71ec-...) without any error or warning.

Verification Done

Before reporting, we verified:

  1. Session file exists and is valid:
    ~/.claude/projects/-app-project/61568b3d-8671-487f-a553-7647fbffff5e.jsonl
    Size: 21KB, valid JSONL format
  2. Parameters are correctly passed:
    - resume = "61568b3d-8671-487f-a553-7647fbffff5e"
    - fork_session = False
    - continue_conversation = True
  3. CWD is consistent between requests: /app/project
  4. Not a concurrency issue: 27 seconds between requests, different trace IDs

Logs

First Request (08:44:57) - New Session

08:44:57.327 | SessionManager: Creating new session: app_session=e5705931...
08:44:59.882 | Captured SDK session ID: 61568b3d-8671-487f-a553-7647fbffff5e
08:44:59.883 | SDK callback on_system_init: is_resuming=False
08:44:59.886 | Saved session mapping: app_session -> sdk_session=61568b3d-...

Second Request (08:45:24) - Resume Attempt

08:45:24.553 | Redis query: is_resuming=True, sdk_session_id=61568b3d-... (correct)
08:45:24.559 | SessionManager resume session: resume=61568b3d-... (correct)
08:45:28.263 | SDK returned SystemMessage: session_id='220a71ec-...' (NEW ID!)
08:45:28.263 | SDK callback on_system_init: sdk_sid=220a71ec..., is_resuming=True, expected=61568b3d...

Impact

  • Users lose conversation context unexpectedly
  • No way to detect resume failure before it happens
  • Silent failure makes debugging difficult

Suggested Improvement

  1. Log a warning when resume fails and a new session is created
  2. Or raise an exception if fork_session=False but resume fails
  3. Or provide a callback/event to notify the application of resume failure

Workaround

We added detection logic in our application:
if is_resuming and expected_session_id and actual_session_id != expected_session_id:
logger.warning(f"Session resume failed: expected={expected_session_id}, actual={actual_session_id}")

Activity

  1. khalido commented on Feb 8, 2026

    @khalido

    Adding to this - session resume was working correctly before but now its not working despite the code being the same.

    Session resumes fails silently:

    requested=0cd7e350... got=a9489974...
    requested=a9489974... got=02a058e6...

  2. n0isy commented on Feb 10, 2026

    @n0isy

    @lld98523 @khalido Check this:

          284      await client1.connect()
          285      await client1.query("What is 2+2?")
          286      events_initial = await collect_events(client1, "initial")
          287 +    # Wait for CLI to flush write queue (enqueueWrite uses setTimeout)                                                           
          288 +    await asyncio.sleep(3)                                                                                                       
          289      await client1.disconnect()
          290  
          291      # Extract session_id
    

    We need to PR here:

      # Close stdin stream                                                                                                                             
      async with self._write_lock:                      
          self._ready = False                                                                                                                          
          if self._stdin_stream:                                                                                                                       
              await self._stdin_stream.aclose()    # 1. Close stdin
              self._stdin_stream = None
    
      # Terminate and wait for process
      if self._process.returncode is None:
          self._process.terminate()               # 2. SIGTERM immediately
          await self._process.wait()              # 3. Reap
    

    It closes stdin, then immediately sends SIGTERM — no grace period. The CLI could potentially detect stdin EOF and begin flushing, but terminate() fires right after without waiting for a clean exit.

    That's why the 3-second sleep before disconnect() is necessary — it gives the CLI's setTimeout-based scheduleDrain time to actually flush the write queue to disk before the process gets killed.

  3. n0isy commented on Feb 10, 2026

    @n0isy

    https://github.com/anthropics/claude-agent-sdk-python/blob/main/src/claude_agent_sdk/_internal/transport/subprocess_cli.py#L427-L466

    Key lines:

    • L445 — closes stdin
    • L456 — immediately calls self._process.terminate() (SIGTERM) with no grace period
  4. added a commit that references this issue on Feb 10, 2026
    86b9b1f
  5. msaah-cleric commented on Feb 18, 2026

    @msaah-cleric

    I'm seeing something similar - when using continue_conversation=True w/o resume, what was previously reliable session rediscovery is now sometimes broken.

  6. n0isy commented on Mar 15, 2026

    @n0isy

    @chrislloyd Can you check this?

  7. maliang-agent commented on Mar 22, 2026

    @maliang-agent

    I ran into the same problem. A lot of the pain here is not really prompt-related — it’s that Claude Code is being used like a short-lived command instead of a managed runtime.

    I built a small project called claude-node that directly controls the local Claude CLI as a persistent subprocess, so an external Python app can explicitly start, send, resume, and stop sessions while keeping the native CLI behavior.

    If what you want is a stable session/controller layer around Claude Code, this might be relevant:
    https://github.com/claw-army/claude-node

  8. qing-ant commented on Mar 24, 2026

    @qing-ant
    Contributor

    This has been fixed by PR #642 (merged March 19, 2026), which adds a 5-second grace period for the CLI process to exit gracefully before sending SIGTERM. Previously, terminate() was called immediately after closing stdin, which could interrupt session file writes and cause the session file to be incomplete on disk — exactly the root cause identified in this thread.

    The fix is included in SDK v0.1.50+ (which bundles CLI 2.1.81). Please upgrade and confirm that session resume works reliably now. Closing as resolved.

  9. n0isy commented on Apr 8, 2026

    @n0isy

    @qing-ant :

    However, there is no early stop after SIGTERM. In the timeout branch:

    self._process.terminate()
    with suppress(Exception):
    await self._process.wait() # unbounded

    If the CLI ignores SIGTERM (stuck in uninterruptible I/O, signal handler hang, etc.), close() will block forever. There's no second fail_after and no SIGKILL escalation. That's a latent hang risk the PR didn't address — worth a follow-up if you care about robustness under misbehaving CLI builds.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions