Skip to content

epoll: dispatch loop invokes callbacks for completions the owning queue no longer holds, crashing queueWrite's callback on head.? (the epoll counterpart of #169/#224/#227) #234

Description

@musingfox

Summary

On the epoll backend, a Completion can have its callback invoked by an event
that the queue owning it no longer considers outstanding. When that completion
belongs to a queued write, the callback generated by queueWrite asserts an
invariant libxev cannot guarantee:

// src/watcher/stream.zig:812-818 (main @ 9ce8e8e6; 811-817 in 34fa5087)
// The queue MUST have a request because a completion
// can only be added if the queue is not empty, and
// nothing else should be popping!.
const req_inner: *xev.WriteRequest = q_inner.head.?;

In ReleaseFast that .? is unchecked, so it becomes a null dereference. We
traced a reproducible SIGSEGV in Ghostty 1.3.1 to exactly this.

This looks like the epoll member of the completion-reuse class already tracked
in #169 (generic), #224 (kqueue, open PR) and #227 (IOCP) — see "Relation to
existing issues" below.

The crash is provably an illegitimate delivery

This part needs no reproducer; it follows from the core dump plus libxev's own
design.

Ghostty 1.3.1 (Arch, ReleaseFast, libxev 34fa5087), SIGSEGV on the
per-surface IO thread:

kernel: io[59183]: segfault at 108 ip ... error 4 in ghostty[...]

The faulting instruction is the inlined queue.Intrusive(T).pop():

mov  (%rdi),%r14          ; r14 = q.head
cmp  0x8(%rdi),%r14       ; if (head == tail)
jne  +8
movq $0x0,0x8(%rdi)       ;     tail = null
mov  0x108(%r14),%rcx     ; <-- SIGSEGV, r14 = 0   (next = head.next)

Read out of the core:

what address / value
WriteQueue (in rdi) head = 0, tail = 0 — cleanly empty
firing completion 0x7fbb90013400
completion.callback (+0x80) the crashing function itself — queueWrite's generated callback
completion.userdata (+0x78) the same WriteQueue that is in rdi
op tag (+0x74) 0x07 = .write
flags (+0xc0) 0x6d0 → dup = true, dup_fd = 54
WriteRequest.next (+0x108) 0

Field offsets identify the backend unambiguously: next at 0x108, userdata
at 0xd0, flags at 0xc0 are the epoll layout (io_uring's are 0xd0,
0x98, 0x90).

Now the design argument. In queueWrite, a completion is only ever handed to
loop.add() when its request is — or is about to become — the queue head:

  • queueWrite: if (q.empty()) loop.add(&req.completion); q.push(req);
    → added only when the queue was empty, so that request becomes head.
  • end of the callback: if (q_inner.head) |req_next| l_inner.add(&req_next.completion);
    → adds the new head.
  • partial-write path: re-adds req_inner, which has not been popped and is
    still head.

So for any legitimate delivery, q.head equals the firing request. At the
crash q.head == 0 while the firing request is 0x7fbb90013400. They differ,
therefore this delivery was not one the queue was expecting. The .? then
turns that into a null dereference.

Reproducer

  • Ghostty 1.3.1 built with -Doptimize=ReleaseSafe (Zig 0.15.2), with
    async-backend = epoll in Ghostty's config.
  • In a freshly opened surface, flood the terminal with queries that each produce
    a reply written back to the pty: 800 rounds of
    DA1 / DA2 / DSR / CPR / XTVERSION / kitty-keyboard / CSI 14t / CSI 18t / OSC 10 / OSC 11.
  • Background CPU and fd-churn load.

Hit rate ~1 in 13 rounds under load; 0 in 10 rounds on an idle machine.

ReleaseSafe dies earlier than production does, in the same dispatch loop:

thread 149503 panic: attempt to unwrap error: FileDescriptorIncompatibleWithEpoll
  src/backend/epoll.zig:461:37  in tick     ) catch unreachable;   (34fa5087; :505 on main)
  src/backend/epoll.zig:103:62  in run

epoll_ctl(CTL_DEL) returning EPERM means that fd number no longer refers to
anything epoll-compatible — it had been closed and reused. In ReleaseFast
catch unreachable is UB, so execution continues and the following
posix.close(v) closes an unrelated descriptor while self.active -= 1
miscounts.

Instrumented observations

Logging at the top of the dispatch loop, behaviour otherwise unchanged:

[DETECT] stale delivery c=0x7fb8ac122838 op=write state=dead dup_fd=0
  • op=write — the re-delivered completion is a write, i.e. one whose callback
    performs q_inner.head.?.
  • state=dead on entry — the loop is about to invoke a callback on a completion
    that is not armed. submit guards on state
    (if (c.flags.state != .adding) continue;); the dispatch loop does not.
  • dup_fd=0 together with state=dead — start sets dup_fd via
    fd_maybe_dup and only then sets state = .active, so this completion was
    never started in its current incarnation, yet epoll delivered an event
    pointing at it.

The reading most consistent with that is that the registration belongs to a
previous incarnation of the same memory: Ghostty pools WriteRequests in a
ring and resets them wholesale (req.* = .{ .full_write_buffer = buf }), so a
registration that was never removed keeps delivering into recycled objects. We
did not instrument the pool itself, so that specific step is inference rather
than observation.

What is proven and what is not

Proven:

  • The production crash is queueWrite's callback running against a queue that
    does not contain the firing request, and head.? dereferencing null.
  • stream.zig's head.? / pop().? assert an invariant the library does not
    enforce.
  • Under the workload above, write completions do get dispatched while not armed.

Not proven:

  • That the path reproduced above is the same path the production crash took.
    There is a concrete discrepancy: the production completion had
    dup_fd = 54 (it had been started and disarmed at some point), whereas the
    reproduced stale deliveries had dup_fd = 0 (recycled memory, never started
    in that incarnation). Same class — a delivery to a completion the queue no
    longer owns — but not demonstrably the same sub-path.

Two fixes that do NOT work

Recording these so nobody repeats them.

1. Guard the dispatch loop on state, mirroring submit:

if (c.flags.state != .active) continue;

The predicate is right: over 100 rounds it fired 0 times in healthy rounds, so
it never rejects a legitimate delivery. But epoll is level-triggered, and
continue skips the event without removing the registration, so the same event
re-fires immediately and forever. In the rounds where the bug triggered a single
completion spun 4.5M–10.7M times. This trades a rare crash for a livelock.

2. Preserve flags.dup_fd across writeInit.

Fails at round 10 of 100. A detector for "writeInit called on a completion
whose state == .active" fired 0 times, so writeInit is not where the
bookkeeping is lost and there is nothing there to preserve.

Both point the same way: by the time the event is dispatched, the information
needed to unregister is already gone. A fix has to be on the retirement paths
that fail to remove the registration in the first place — which is #231's
territory and a design call we did not want to guess at.

Relation to existing issues

Note on backend selection

The affected machine sets async-backend = epoll in Ghostty's config — a
long-standing workaround for ghostty-org/ghostty discussion #3224 ("general
slowness on Hyprland"). So these processes are on epoll deliberately, not as
a fallback. Ghostty's default is auto, which picks io_uring on a typical Linux
box, and the io_uring path does not have this bug. That one config line is all
that is needed to reproduce.

Separately, IO_Uring.available() creates a real 256-entry ring
(linux.IoUring.init(256, 0) catch return false), so a process that loses that
probe at startup drops to epoll for its whole lifetime even under auto. We
confirmed this by launching with RLIMIT_MEMLOCK=64K, which reproduces the same
fd profile (0 io_uring fds, epoll loops). That is a second, unconfigured route
onto this code path.

Environment

  • libxev 34fa50878aec6e5fa8f532867001ab3c36fae23e (the rev Ghostty 1.3.1
    pins). The same code is present in the rev Ghostty main pins and in libxev
    main (9ce8e8e6ff89e583258a7f8e7adeeeaeae8611bf).
  • Zig 0.15.2, Linux 7.1.8, x86_64.
  • Reproduced against Ghostty 1.3.1 built from source; originally observed three
    times over three days on the distro ReleaseFast build.

Investigated and written by Claude Opus 5 via Claude Code, posted by the machine
owner.

Activity

  1. rsteenwyk commented on Aug 27, 2026

    @rsteenwyk

    This is a report I had my clanker produce - can confirm I am seeing this issue intermittently, adding it here in hopes that it will be useful.


    Independent confirmation of this crash from Ghostty 1.3.1 on Arch/Omarchy.

    I have now seen five crashes with the same SIGSEGV at address 0x108 over four days. The two most recent crashes happened about ten seconds apart; repeating the same workload afterward succeeded, so the failure is intermittent.

    The practical trigger is an SSH session to a Linux development host followed by starting herdr, a TUI that queries terminal palette colors. One symbolized core contained this in-flight PTY response:

    ESC ] 4;126;rgb:afaf/0000/8787 ESC \
    

    The remote program and SSH connection remain usable after restarting Ghostty. This is not yet a minimal reproducer, but it independently supports terminal-query replies as a way to exercise the faulty queued-write path. SSH/TCP timing may explain why identical launches do not always crash.

    I rebuilt the distro's Ghostty 1.3.1 package with debug information while retaining ReleaseFast. The latest core resolves to:

    #0 queue.Intrusive(...WriteRequest).pop
       libxev/src/queue.zig:48
       self.head = next.next;
    #1 queueWrite callback
       libxev/src/watcher/stream.zig:844
    #2 xev.Loop.tick
       libxev/src/backend/epoll.zig:450
    #3 termio.Thread.threadMain
       ghostty/src/termio/Thread.zig:278
    

    GDB shows the same invalid state described in this issue:

    next = 0x0
    req_inner = 0x0
    epoll result = .write(28)
    

    The 28-byte successful write result exactly matches the retained OSC response buffer. The queue is empty and pop dereferences next.next, producing the NULL + 0x108 fault. Earlier crashes from the unmodified distro package have the same kernel segfault at 108 fingerprint.

    The affected configuration explicitly selects epoll:

    # Workaround for Ghostty discussion #3224 on Hyprland
    async-backend = epoll

    Environment:

    Ghostty 1.3.1-arch2
    Zig 0.15.2
    Build mode ReleaseFast
    libxev revision 34fa50878aec6e5fa8f532867001ab3c36fae23e
    Linux 7.1.9-arch1-2 x86_64
    Omarchy 4.0.1-1 / Hyprland
    

    I can retain the private core for additional targeted GDB queries, but will not upload it because it contains terminal session data.

  2. lcorneliussen commented on Sep 7, 2026

    @lcorneliussen

    I have a variant reproducer that fires reliably on an idle machine, which may be
    useful for evaluating candidate fixes.

    The only real difference from the workload in the issue is that it never reads the
    replies, so the pty input buffer fills and writes go partial/EAGAIN continuously:

    #!/usr/bin/env bash
    while :; do
      for _ in $(seq 1 200); do
        printf '\e[c'; printf '\e[>c'; printf '\e[>q'
        printf '\e]10;?\e\\'; printf '\e]11;?\e\\'
        printf '\e[?u'; printf '\e[5n'; printf '\e[6n'
        printf '\eP+q544e\e\\'
      done
    done
    ghostty --config-default-files=false --async-backend=epoll -e bash ./repro.sh
    

    No background CPU or fd-churn load:

    build result
    Arch ghostty 1.3.1-arch2 (ReleaseFast, stock) 9/9 SIGSEGV at 0x108, within 2–6 s
    local ReleaseSafe 1.3.1 19/20 abort at epoll.zig:461
    either, with --async-backend=io_uring 0/5

    The ReleaseFast row is the production queue.zig:48 queued-write null dereference,
    which the testing notes in #237 say that workload could not reproduce.

    Evaluating #237

    I backported the three code changes from #237 (not the new tests, which need Zig
    0.16) onto 34fa5087, rebuilt Ghostty 1.3.1 ReleaseSafe against it, and ran 20 runs
    per arm with the script above:

    arm aborts
    unpatched 19/20
    #237 backported 16/20

    Fisher exact, two-sided: p = 0.34 — not a significant difference. Both arms abort in
    epoll_ctl(EPOLL_CTL_DEL) with FileDescriptorIncompatibleWithEpoll.

    That agrees with what @rsteenwyk already concluded in the PR itself, just with a
    larger sample: the cancellation and duplicate-fd fixes are not sufficient for the
    stale-registration problem tracked here.

    Environment: Ghostty 1.3.1, libxev 34fa5087, Zig 0.15.2, Linux 7.1.9-arch1-2,
    Omarchy (which ships async-backend = epoll by default per ghostty#3224, so this
    backend is not an unusual configuration in practice).


    Investigation assisted by Claude Code (Claude Opus). The reproducer and the run counts
    above I verified directly.

  3. rot13maxi commented on Oct 8, 2026

    @rot13maxi

    Independent Omarchy confirmation (see also omarchy#6868) of this crash, plus an
    alternative to the continue guard in #239. Reproduced on stock 1.3.1 and tested
    the #239 guard against my reproducer.

    Reproducer (idle box, no background load)

    async-backend = epoll, flood terminal-query replies with nobody reading them
    so pty writes go partial/EAGAIN continuously:

    # repro.sh
    while :; do
      for _ in $(seq 1 200); do
        printf '\e[c'; printf '\e[>c'; printf '\e[>q'
        printf '\e]10;?\e\\'; printf '\e]11;?\e\\'
        printf '\e[?u'; printf '\e[5n'; printf '\e[6n'
        printf '\eP+q544e\e\\'
      done
    done
    # ghostty --config-default-files=false --async-backend=epoll -e bash repro.sh

    Stock 1.3.1 ReleaseFast (libxev 34fa5087, Arch build-id 36aeb61c…):
    5/5 SIGSEGV @ 0x108 in 2-6s. Same workload on --async-backend=io_uring:
    0/5. Matches this issue's field table exactly.

    Symbolized (ReleaseFast+symbols build)

    #0 queue.Intrusive(...Shared(epoll).WriteRequest).pop  queue.zig:116   (self.head = next.next, next = r14 = 0)
    #1 queueWrite callback                                  stream.zig:844  (q_inner.pop().?)
    #2 backend.epoll.Loop.tick                              epoll.zig:450   (c.callback(...))
    #3 termio.Thread.threadMain_                            ghostty Thread.zig:278
    

    Core readout at the crash: q_inner.head == tail == 0; the firing completion's
    userdata points at that same queue; op = .write; r = .write = <len>. i.e.
    a write completion dispatched by epoll for a request the queue already popped.

    On the continue guard (#239)

    The predicate is right (only .active completions are legitimately outstanding).
    On my reproducer the #239 build stayed alive with no crash, but its io thread ran
    very hot under the flood; I could not cleanly separate a livelock spin from the
    flood's own write work, so I can't independently confirm the livelock this issue
    warns about - only that continue leaves the stale event registered.

    Variant: retire the registration instead of skipping it

    The only completion this layer ever arms is the queue head's (empty-push,
    partial-rearm, next-readd and the initial arm all imply firing == head), so a
    delivery whose completion isn't the head is stale. Guard the generated
    queueWrite callback (src/watcher/stream.zig) before the head.? / pop():

     const q_inner = @as(?*xev.WriteQueue, @ptrCast(@alignCast(ud))).?;
    +
    +// STALE-DELIVERY GUARD: the only completion this layer arms is the one
    +// belonging to the queue head, so a firing completion that isn't the head's
    +// is a stale epoll registration from a prior incarnation of this pooled
    +// (wholesale-reset) request. Return .disarm so the dispatcher CTL_DELs the
    +// fd (the stale OUT event can't re-fire -> no livelock) and we never pop an
    +// empty queue.
    +const firing: *xev.WriteRequest = @fieldParentPtr("completion", c_inner);
    +if (q_inner.head != firing) return .disarm;
    
     const req_inner: *xev.WriteRequest = q_inner.head.?;

    Results on the same reproducer: 0/20 crashes, with ~4.7 stale deliveries
    caught per run (all firing != head, r = .write). No dropped output on a
    pty fast-reader conservation harness: 3/3 runs, 20000/20000 callbacks,
    649488/649488 bytes exact.
    (A second harness driving sustained backpressure
    via a pool-reuse ring did not finish cleanly and is inconclusive; I am not
    claiming byte-conservation-under-backpressure yet.)

    Caveats


    Investigated and drafted by Claude (Opus 5) via omp; the reproducer, cores,
    CPU samples, and conservation runs were verified by me on my own machine.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions