Skip to content

Bundled ugrep balloons to 9–14 GB RSS compiling a bounded-interval BRE (plain grep is transparently routed to it) #83342

Description

@developerinlondon

Environment

  • Claude Code 2.1.220, Linux x86_64
  • Bundled ugrep 7.5.0 x86_64-pc-linux-gnu +sse2; -P:pcre2jit
  • Claude Code's shell integration installs a grep() function that re-execs the claude binary as ugrep, so ordinary grep calls made by agent Bash tool calls run this bundled ugrep.

Repro

Zero input needed — the blowup is in pattern compilation:

bash -c 'exec -a ugrep ~/.local/bin/claude -G -o "\"content\":\"[^\"]*coder[^\"]\{0,300\}" /dev/null'

RSS climbs past 9.3 GB within 60 s (still growing when killed). Against a real 11 MB file (longest line 28 KB) it reached 14.3 GB max RSS in ~95 s.

GNU grep, same pattern, same 11 MB file: 8 MB RSS, <10 ms.

Tool Input Max RSS Time
bundled ugrep -G -o /dev/null >9.3 GB killed at 60 s
bundled ugrep -G -o 11 MB jsonl 14.3 GB ~95 s
GNU grep -o 11 MB jsonl 8 MB <10 ms

The trigger appears to be the bounded interval \{0,300\} after a negated class under -G — presumably DFA/state-set construction exploding on the counted repetition.

Curiously, under ulimit -v 2097152 the same invocation exits 0 almost immediately, so an internal allocation-failure path already avoids the explosion — a default cap on pattern-compile memory would likely be a small fix.

Impact

An agent session issued an innocuous-looking grep -o '"content":"[^"]*coder[^"]\{0,300\}' session.jsonl via the Bash tool; the shell integration routed it to the bundled ugrep, which grew to 11 GB and pushed a 125 GB host into memory-pressure thrash (swap exhausted, sibling Claude Code sessions frozen in D-state reclaim for ~15 minutes until the process was killed).

Since agents generate grep patterns freely, a pattern-compile memory/complexity guard (or falling back to system grep for interval-heavy BREs) would prevent a single tool call from taking out the host.

Activity

  1. developerinlondon commented on Aug 2, 2026

    @developerinlondon
    Author

    Root-caused and fixed upstream. The blowup is in ugrep's DFA construction: a counted repetition overlapping a preceding repeat ([^"]*x[^"]\{0,300\}) legitimately explodes the DFA, but ugrep's exceeds complexity limits guard only counts states, and each state carries ~2.5 KB of iteration-tagged positions — so the guard fires only after ~14 GB / 95 s. Reproduced on ugrep 7.8.3 (latest).

    Filed Genivia/ugrep#555 with measurements; fix submitted as Genivia/ugrep#556 — a position-count cap feeding the existing error path, which turns this case into a clean exceeds complexity limits error in 0.59 s at 155 MB. Once merged upstream, bumping the bundled ugrep picks it up. Until then, a session-level cgroup memory cap on the CLI is an effective host-side mitigation.

  2. developerinlondon commented on Aug 2, 2026

    @developerinlondon
    Author

    Cross-referencing prior sightings of the same underlying defect that were filed as platform-specific symptoms: #64133 (macOS, 8 GB+ on bounded repetition over minified input) and #54394 (WSL2, embedded ugrep wrapper OOM freezing the host). Root cause and upstream fix are in this issue: Genivia/ugrep#555 / PR Genivia/ugrep#556.

  3. ErikEremenko commented on Aug 4, 2026

    @ErikEremenko

    Corroborating on aarch64, still present in 2.1.221.

    Same shape as your repro — bundled ugrep under -G -o with a \{0,300\} bounded interval — reached 7.0 GB RSS and was still growing 775 s later on an 8 GB Raspberry Pi 5.

    (My sampler truncated argv at 200 chars, so I can't show the full pattern. The visible portion is -G … -o .\{0,300\}worklist. and it was almost certainly followed by a second interval, consistent with #67021.)

    Why I'm adding a data point rather than another duplicate: on this host there was no OOM kill at all.

    Raspberry Pi OS ships cgroup_disable=memory on the kernel cmdline, and this box also runs vm.overcommit_memory=1. With no memory cgroup controller, systemd-oomd and Docker --memory limits are silently inert, and reclaim always "succeeds" by evicting page cache — so no allocation ever fails and the kernel OOM killer never fires. Zero oom-kill events across every occurrence.

    The consequence is worse than an OOM kill:

    • the runaway is never terminated (7.0 GB, still alive at 775 s)
    • page cache collapses (3.1 GB → 24 MB), swap fills to 99%, and the box enters refault thrash at ~182 MB/s with 45+ processes in D-state
    • load reaches 50–90 and a watchdog resets the machine

    Because it dies by watchdog reset there is no shutdown sequence, pstore is empty, and nothing is logged — it presents as spontaneous hardware resets. It took several weeks and a purpose-built per-process RSS sampler to attribute it to ugrep at all, because the crash destroys its own evidence.

    So "the OOM killer will contain it" doesn't hold universally. I think that strengthens the case for the pattern-compile memory cap you suggest — on a host like this it's the only layer that works.

    Worth noting your ulimit -v observation held here too: an address-space cap makes ugrep's allocation-failure path exit cleanly rather than explode, so it's an effective interim mitigation for anyone hitting this.

    Environment: Claude Code 2.1.221 (npm install, Node v22.22.2), Raspberry Pi 5, aarch64, 8 GB RAM + 2 GB swap, Debian 12 (bookworm), kernel 6.12 series.

  4. JulianPscheid commented on Aug 6, 2026

    @JulianPscheid

    Confirmed on macOS arm64 in Claude Code 2.1.223.

    A two-line, 417-byte synthetic fixture is enough to trigger it:

    printf '%s\n' \
      'This synthetic segment is surrounded by ordinary words to exercise the bounded context expression.' \
      'Another synthetic segment is surrounded by ordinary words to exercise the bounded context expression.' \
      > /tmp/ugrep-fixture.txt
    
    grep -o -i '.\{0,70\}segment.\{0,70\}' /tmp/ugrep-fixture.txt

    The shell snapshot routed plain grep to the embedded ugrep through ARGV0=ugrep and -G. I ran the embedded binary under a watchdog that sampled proc_pid_rusage every 10 ms and killed only the test child at 512 MiB physical footprint. It reached the cap in 1.312 seconds. /usr/bin/grep completed with the same fixture and pattern in under 0.01 seconds.

    The original incident used the same pattern against a 35,068-byte, 104-line file. The detached ugrep process reached 16.4 GB physical footprint before a separate local monitor paused it. The parent shell had already detached, so the process would otherwise have remained alive.

    Environment: macOS 26.5.1 (25F80), Apple Silicon, 36 GiB RAM, native Claude Code 2.1.223.

    This confirms that the current macOS native build still has the affected embedded ugrep behavior. No large or private input is required.

  5. MG674 commented on Aug 7, 2026

    @MG674

    Two things from a WSL2 hit today (2.1.224) that I don't think are in the thread yet. Not a
    "me too" — the root cause and the upstream fix are settled above; these are about diagnosing
    it and about mitigating it client-side while the bundled ugrep is unbumped.

    1. timeout N grep … silently does NOT reproduce this, and reports success

    timeout is an external binary: it execvps its argument, so the grep() shell function
    the snapshot installs is invisible to it and real /usr/bin/grep runs instead.

    timeout 10 grep -o '[^"]\{0,120\}X[^"]\{0,120\}' file    # -> exit 0, instant. NOT ugrep.
    

    I used exactly that to "verify" the command was harmless and got a clean pass while the real
    invocation was still eating the box. Anyone reaching for timeout to bound the blast radius
    while testing will get a false negative and may close this as unreproducible. Same trap for
    rg, also a function — timeout … rg … exits 127, which reads as "ripgrep not installed"
    rather than "your harness bypassed the shim".

    Reproducing needs the function invoked directly, or the binary called with argv[0]=ugrep as
    in the repros above.

    2. Killing it is not enough when the command was chained

    Ours was grep -o '…£150…' file ; grep -o '…<cyrillic>…' file. Killing the first ugrep let the
    shell advance to the second, which started its own blowup; killing the parent shell then
    orphaned that child rather than reaping it, and it kept climbing to ~5 GB on its own. Memory
    went 6.4 GB → recovered → 6.1 GB while I was writing up the first kill. Two separate kills were
    needed, several minutes apart.

    So a pgrep/free check immediately after the kill reports all-clear and is wrong. Worth
    re-checking a minute later. (Overlaps the orphan reports in #77230 / #80230, but the
    "kill it and the shell launches another" sequence seems undescribed.)

    3. Client-side mitigation that doesn't need root

    The cgroup cap suggested above works but needs privilege. A PreToolUse/Bash hook is a
    Claude-Code-native alternative that blocks the shape before the process is ever spawned —
    useful for hosts where cgroup_disable=memory makes limits inert, as @ErikEremenko found.

    Keying on two or more bounded intervals in a non--P, non--F Bash grep matches the
    trigger as characterised in #67021, which I re-confirmed here against a one-byte file:

    invocation (1-byte input) result
    -G two intervals, with or without -o SIGSEGV at a 500 MB ulimit -v
    -G one interval fine
    -E two intervals SIGSEGV at the cap
    -P two intervals fine

    Confirms -o is not required and that a single interval is safe — so a guard can be narrow
    enough not to interfere with everyday grep -rn / grep -c.

    Environment

    Claude Code 2.1.224, WSL2 Ubuntu 24.04 (kernel 6.6.87.2-microsoft-standard-WSL2), 8.9 GB.
    Real-world trigger was a £ context-window pattern over a 192 KB HTML file, then the identical
    pattern over its Ukrainian translation — so any repo with non-ASCII content is doubly exposed.

  6. HoschiS commented on Aug 11, 2026

    @HoschiS

    Another data point, this time a real OOM-kill (not the silent-thrash variant) — aarch64, Claude Code 2.1.226.

    Environment: Lima VM on Apple Silicon (M3), Ubuntu 25.10, kernel 6.17, 3.8 GB total RAM — small enough that the kernel OOM killer fires quickly instead of thrashing for minutes first.

    Trigger: an agent Bash-tool call ran an entirely ordinary

    grep -o '<w:rPr>.\{0,200\}Baloo 2.\{0,100\}</w:rPr>' document.xml
    

    against a 79 KB file (extracted Word document XML). Confirmed via dmesg/journalctl -k:

    Out of memory: Killed process 339430 (2.1.226) total-vm:6596196kB, anon-rss:3318272kB, ...
    

    3.3 GB anon-RSS, killed within seconds of the call. Same file/pattern through real GNU grep (command grep, bypassing the shell shim, ulimit -v 300000 as a safety cap): 0.17s, no measurable memory pressure.

    Corroborates the root cause already nailed down above (Genivia/ugrep#555 / #556) — bounded interval (.{0,200}) under -G blows up DFA construction regardless of platform or input size. Confirms MG674's point too: timeout grep ... would have given a false negative here since it bypasses the shell function shim entirely.

    Mitigation in the meantime: command grep (bypasses the shim, uses real system grep) for any pattern with interval quantifiers + -o. On memory-constrained hosts specifically, this isn't a "slow but survivable" bug — it takes the whole session (and anything else sharing the VM's RAM) down with it.

  7. bcherny commented on Aug 20, 2026

    @bcherny
    Collaborator

    Fixed in 2.1.235 — changelog: "Improved the embedded grep in native macOS/Linux builds: pathological patterns now fail fast instead of exhausting memory, and -m N with -A/-C prints correct context" (https://github.com/anthropics/claude-code/blob/main/CHANGELOG.md). Closing — reply to reopen if it still happens on the latest version.

    🤖 Generated with Claude Code

  8. JulianPscheid commented on Aug 20, 2026

    @JulianPscheid

    Confirmed by the boss himself 🙌

  9. github-actions commented on Oct 11, 2026

    @github-actions

    This issue has been automatically locked since it was closed and has not had any activity for 7 days. If you're experiencing a similar issue, please file a new issue and reference this one if it's relevant.

  10. locked as resolved and limited conversation to collaborators on Oct 11, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions