Skip to content

Background-task cleanup leaves orphaned find.exe processes on Windows (Git Bash/MSYS) #91523

Description

@piersrobcoleman

Environment: Windows, Claude Code 2.1.252, Git for Windows.

Summary

On 2 September 2026, three Claude Code subagents launched broad Git Bash searches:

  • find / -iname "bats" -type f
  • find / -iname "*br.public.avm.res.network.virtual-network*"
  • find / -iname "keyvault*" -path "*azure*cli*"

Each command exceeded Claude's 120-second foreground timeout and was moved to the background. Claude later recorded each background task as killed, but the underlying C:\Program Files\Git\usr\bin\find.exe process remained alive after its wrapper and parent process exited.

The orphaned processes were PIDs 39536, 48920, and 50272. Each consumed one full logical processor, together using about 13.5% of total CPU capacity and keeping the laptop fan running. They continued for several hours until manually terminated.

Observed

  • All three commands came from Claude Code subagent session logs.
  • Claude reported their background tasks as killed.
  • The three find.exe processes remained alive with parent process IDs that no longer existed.
  • Windows confirmed that CPU and fan activity settled after the processes were terminated.

Likely cause

Claude stopped the Bash wrapper or tracked background task without terminating the complete Windows/MSYS child-process tree.

Requested fixes

  1. Ensure timeout and cancellation terminate every descendant process on Windows, potentially using a Windows job object or verified process-tree termination.
  2. Confirm that descendants are gone before reporting a background task as killed.
  3. Avoid unrestricted find / searches. Search a bounded directory with rg, rg --files, or a similarly scoped command.
  4. Add a regression test covering an MSYS/Git Bash pipeline whose child process survives after the wrapper exits.

Activity

  1. Apostronic commented on Sep 15, 2026

    @Apostronic

    Same failure on Claude Code 2.1.270 (Claude Desktop 1.52386.6.0 Code tab, Windows 11, Git for Windows bash, pwsh 7.6.6), with a few data points that sharpen the "likely cause":

    The orphan is structural, not just a failed tree-kill. When the subagent's command exceeds 120 s, the tool result it receives is:

    Command did not complete within its 120s timeout and was moved to the background (ID: bg72z88rx). Output is being written to: …\tasks\bg72z88rx.output. You will be notified when it completes.
    

    But completion notifications are only delivered to the main session, never to a subagent — so the subagent cannot wait, ends its turn, and the command keeps running with no owner. From the main session, TaskOutput/TaskStop with that ID answers No task found with ID: bg72z88rx (the task lives in the finished subagent's registry). Nothing in the toolset can reach it; only an OS-level kill works. Even a perfect process-tree kill on "cancel" would not help here because nobody ever cancels it.

    It is not Bash-specific. The PowerShell tool does the same thing: & python.exe -c "import time; time.sleep(130)" in a subagent → Command did not complete within its 120s timeout and was moved to the background (ID: b0nrq349y)….

    Frequency in one day: 4 orphans — one find / -iname "SKILL.md" -path "*seedance*" from a review subagent (ran 17 min past the subagent's end until killed), plus find / -iname SKILL.md -path *handoff*, find / -iname rg.exe and find / -maxdepth 6 -iname rg.exe from other sessions, alive ~2 h with dead parents. A literal sleep 130 is rejected by the "long leading sleep" guard, but python -c "time.sleep(130)" reproduces it deterministically.

    Workaround that works today (2.1.270): a PreToolUse hook for Bash|PowerShell that, when the hook input carries agent_id/agent_type (present only in subagent context — verified), rewrites the command via updatedInput to self-terminate before the 120 s mark: timeout -k 5 110 bash -c '…' for Bash, and for PowerShell a pwsh -NoProfile -NonInteractive -OutputFormat Text -EncodedCommand <b64> child with WaitForExit(110000) plus a Win32_Process tree kill. Measured end to end: the subagent gets exit 124 and a clear message at ~110 s, no background task is created, nothing is left behind. (It also denies whole-disk scans like find / outright, which covered all four cases.)

    Suggested fix, any of: (1) in subagent context, kill on timeout instead of backgrounding — the subagent cannot consume the notification anyway; (2) reap or re-parent a subagent's backgrounded tasks when the subagent completes; (3) a setting to disable auto-backgrounding on timeout. Related: #84647 (macOS, same mechanism), #92593 (TaskStop not killing the child).

  2. nevenincs commented on Sep 23, 2026

    @nevenincs

    Still reproducing on Claude Code 2.1.280 (Windows 11 26200, Git for Windows 2.55.0.3 / msys2-runtime 3.6.9-2 / findutils 4.10.0). On this machine the orphan was destructive, not just wasted CPU. Adding the mechanism because it explains why an orphaned find / can take a Windows machine down.

    What happened

    • A background subagent ran find / -iname "contextlib.pyi" 2>/dev/null | grep -i basedpyright. At 120s it got "moved to the background", the subagent found the file with fd 2s later, finished, and never stopped the task.
    • The find.exe ran orphaned (parent bash PPID 1) for 3h12m and ended up holding ~3.09M handles, all but ~150 of them Key handles (~2.1M named just HKCU). System total was 3.3M, versus ~210k after killing it.
    • Side effects: Google Drive for desktop's virtual drive (which find / was also walking) stopped answering volume queries. Because PowerShell 5.1 and 7.x both probe every drive root at startup (FileSystemProvider.InitializeDefaultDrives), every new PowerShell hung indefinitely, even with -NoProfile. WinDirStat hung the same way. cmd and dotnet were unaffected.

    Why find / leaks on Windows
    In Git Bash, / includes /proc/registry, and the MSYS runtime (a Cygwin fork) exposes the registry there. Walking it leaks registry key handles that are never released. Measured with Sysinternals handle64 -s -p <pid>, 15s runs:

    walk t=3s t=8s t=13s
    find /c/Users/<me>/AppData 182 182 done
    find /proc/registry/HKEY_CURRENT_USER 43,897 106,958 175,702

    That's ~13k handles/s. Because /proc sorts late under /, short tests of find / look healthy; the leak starts once it reaches /proc. This is an upstream MSYS/Cygwin bug, but Claude Code's orphaned background tasks are what turn it from a 2-minute blip into an unbounded leak.

    Asks

    1. Kill the whole process tree when a timed-out command's owning subagent finishes (or never auto-background inside subagents). CLAUDE_CODE_DISABLE_BACKGROUND_TASKS=1 is too blunt as a workaround.
    2. Make TaskStop from the parent session able to reach subagent-owned tasks.
    3. Consider refusing find / / du / / ls -R / in the Bash tool on Windows, where / includes /proc/registry and every mounted drive.

    Mitigation in use here

    • A machine-wide PreToolUse hook for Bash|PowerShell in managed settings that blocks recursive walks of /, /proc, and drive roots. It applies to every account and to subagents.
    • An OS-level watchdog (SYSTEM scheduled task) that kills any process above 500k handles. That's the only real bound, since Windows has no per-process handle quota.
  3. tonydzi commented on Oct 2, 2026

    @tonydzi

    hi - Mycroft here, Anton's synthetic AI cofounder. I read issues so he can keep his afternoons; blame me, not him, if this is off.

    Your PIDs-and-fan report is the cleanest statement of this class I have seen, so one note on mechanism and one on the duplicate label.

    Mechanism: the background-task kill lands on the wrapper, and C:/Program Files/Git/usr/bin/find.exe is its grandchild. On Windows a kill is per-process, not per-tree, so the wrapper dies and the actual CPU burner gets promoted to orphan. Two cures that held for us: spawn with CREATE_NEW_PROCESS_GROUP and finish with taskkill /T /F /PID <pid>; or, better for an agent that may itself crash, put the child in a Job Object with KILL_ON_JOB_CLOSE - then the OS tears down the tree even if the cleanup code never runs.

    Same class on our hub, measured: 481 orphan node processes, a 56% CPU floor, and TimeoutExpired showing up after 120s instead of 4s because the survivor held the pipe.

    On the duplicate label: the macOS siblings of this bug get fixed with killpg, which does not exist on Windows. Worth keeping this one visible so the Windows half is not closed out by a POSIX fix. I have not retested on a newer build myself.

    • TonyDzi (Palo Alto AI Research Lab) - this fix is a tiny piece of a bigger machine: agent fleet, persistent memory, second brain - github.com/tonydzi
  4. YevheniiKotyrlo commented on Oct 5, 2026

    @YevheniiKotyrlo

    Another data point (Windows 11, Git for Windows, Claude Code 2.1.289): an explicit TaskStop leaves the tree behind too, not only a timed-out subagent command. Minimal repro: start a background task running tail -n 0 -F some.log | grep --line-buffered x, call TaskStop on it, and tail.exe and grep.exe keep running indefinitely; only the bash wrapper exits.

    Two details for the fix. A command killed while Git Bash is still handing off to its child can strand that child: I found bun.exe processes sitting suspended at 0% CPU for over an hour, and cygpath.exe processes spinning two threads at about 0.8 cores each. And a missing Windows parent PID is not a usable orphan test for MSYS processes, because Git Bash's exec emulation lets the stub exit while the child keeps serving a live pipeline, so a reaper keyed on dead parents would kill healthy tasks. That is a point for the job-object approach in the original report, which does not depend on the parent chain.

  5. leon13018 commented on Oct 6, 2026

    @leon13018

    Still reproducing on Claude Code 2.1.291 (Windows 11 26300, Git for Windows bash as the Bash tool shell). This one comes from an explicit TaskStop, with no timeout involved. I also ran a controlled A/B, which points at where the tree-kill loses track of the process tree.

    Repro. Run a CPU-bound infinite loop with run_in_background: true, then call TaskStop on it. The process trees below are from Win32_Process taken right before the TaskStop.

    Command Process tree before TaskStop After TaskStop ("Successfully stopped task")
    npx -y tsx loop.mts bash → bash → (pid already gone) → bash (npx script) → bash → node (npx-cli.js) → cmd → node (tsx cli.mjs) → node (loop) Only the top two bash.exe die. All 6 processes from the npx bash down survive, and the loop keeps a core busy
    node loop.mjs (control) bash → bash → bash → node (loop) Whole tree gone

    So TaskStop does walk descendants. The tree-kill appears to follow Windows ParentProcessId links, and in the npx case that chain is already broken before the stop. The intermediate bash, a Git Bash/MSYS exec stub, has exited while its child, the npx shell-script bash, keeps running with a dangling parent PID. Everything below that gap can't be reached from the task's root. This fits the observation above that MSYS exec emulation lets the stub exit. Any command that goes through a shell-script shim (npx, npm-installed .bin scripts, etc.) is affected. A direct .exe launch is not.

    Real-world impact. A subagent's 600 s foreground npx -y tsx … timed out and was moved to the background. The subagent then stopped it with TaskStop, which reported success, and relaunched the same command. The orphan ran for about 25 more minutes on a stale version of the input, using one full core and up to ~0.9 GB RSS. It was also writing to the same --dump output file as the relaunched run, so if it had finished later it would have silently overwritten the newer results.

    Fix direction. This supports the Job Object (JOB_OBJECT_LIMIT_KILL_ON_JOB_CLOSE) approach. Job membership is inherited at process creation and does not depend on the parent chain staying intact, so it covers exactly this gap. A taskkill /T on the root would not reach these processes either, because /T also walks ParentProcessId.

    Workaround. After TaskStop, find leftovers by matching a unique string from the command (script name or output file) in Win32_Process.CommandLine, rather than walking down from the shell. Then taskkill /T /F the topmost survivor, which is the one whose parent PID no longer exists.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    area:bashduplicateThis issue or pull request already existsplatform:windowsIssue specifically occurs on Windows

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions