Skip to content

Restart interrupted syscalls when replay resumes from a syscall stop - #4103

Open
Keno wants to merge 1 commit into
rr-debugger:masterfrom
ChronitonAI:replay-restart-without-breakpoint
Open

Keno wants to merge 1 commit into
rr-debugger:masterfrom
ChronitonAI:replay-restart-without-breakpoint

Conversation

@Keno

@Keno Keno commented Oct 1, 2026

Copy link
Copy Markdown
Member

[Encountered, debugged and patch by AI, but the mechanism is plausible to me and the fix seems right]

On x86, when a syscall returns -ERESTARTSYS, -ERESTARTNOINTR, -ERESTARTNOHAND or -ERESTART_RESTARTBLOCK and no signal handler runs (e.g. another thread takes the signal, or it's a SIGSTOP/SIGCONT), the kernel restarts it on the way back to userspace:
arch_do_signal_or_restart() sets ip -= 2 and ax = orig_ax (or the restart_syscall number). That code only looks at the registers, and it only runs when the task goes through signal processing on the way out, i.e. when it resumes from a stop inside get_signal() (signal-delivery stop, group stop, PTRACE_EVENT_STOP) or has a signal pending. A task that resumes from a syscall stop with no signal pending returns to userspace with ax unchanged.

rr records the -ERESTART* exit, and replay relies on the kernel doing the same restart when the tracee resumes. That works when replay reached the syscall via its internal breakpoint, because the tracee then resumes from the breakpoint's SIGTRAP stop, so TIF_SIGPENDING is set. It doesn't when the tracee resumes from a syscall stop instead:

  • After rr ran a remote syscall in the task at that point. Creating a checkpoint does a remote fork, so e.g. rr replay -g <event of the -ERESTART* exit> followed by continue fails with "expecting tracee signal or trap", for ordinary code in read-only memory too.
  • When the syscall instruction is in writable or MAP_SHARED memory (e.g. JIT code). enter_syscall() can't use a breakpoint there, so it enters the syscall with PTRACE_SYSEMU.

The tracee then sees -ERESTART* as the syscall's result, and replay diverges if its next event is at the restarted syscall instruction, e.g. a signal that arrived after the restart.

AutoRemoteSyscalls already re-enables TIF_SIGPENDING with PTRACE_INTERRUPT after its remote syscalls, but only for a task at a real syscall-exit stop with a restartable result, i.e. during recording. During replay the task's status at such an exit is the breakpoint's SIGTRAP, or nothing after finish_emulated_syscall() on the PTRACE_SYSEMU path, so that never triggers, and there is no pending signal to preserve anyway.

So apply the restart ourselves in ReplayTask::will_resume_execution() when the registers are at such a syscall exit (orig_ax >= 0 and syscall_may_restart()). On aarch64 the kernel adjusts pc and x0 for the restart before the signal stop, and replay already applies that when processing the syscall exit, so this is x86 only.

On x86, when a syscall returns -ERESTARTSYS, -ERESTARTNOINTR,
-ERESTARTNOHAND or -ERESTART_RESTARTBLOCK and no signal handler runs
(e.g. another thread takes the signal, or it's a SIGSTOP/SIGCONT), the
kernel restarts it on the way back to userspace:
arch_do_signal_or_restart() sets ip -= 2 and ax = orig_ax (or the
restart_syscall number). That code only looks at the registers, and it
only runs when the task goes through signal processing on the way out,
i.e. when it resumes from a stop inside get_signal() (signal-delivery
stop, group stop, PTRACE_EVENT_STOP) or has a signal pending. A task
that resumes from a syscall stop with no signal pending returns to
userspace with ax unchanged.

rr records the -ERESTART* exit, and replay relies on the kernel doing
the same restart when the tracee resumes. That works when replay reached
the syscall via its internal breakpoint, because the tracee then resumes
from the breakpoint's SIGTRAP stop. It doesn't when the tracee resumes
from a syscall stop instead:

- After rr ran a remote syscall in the task at that point. Creating a
  checkpoint does a remote fork, so e.g. `rr replay -g <event of the
  -ERESTART* exit>` followed by `continue` fails with "expecting tracee
  signal or trap", for ordinary code in read-only memory too.
- When the syscall instruction is in writable or MAP_SHARED memory (e.g.
  JIT code). enter_syscall() can't use a breakpoint there, so it enters
  the syscall with PTRACE_SYSEMU.

The tracee then sees -ERESTART* as the syscall's result, and replay
diverges if its next event is at the restarted syscall instruction, e.g.
a signal that arrived after the restart.

AutoRemoteSyscalls already re-enables TIF_SIGPENDING with
PTRACE_INTERRUPT after its remote syscalls, but only for a task at a
real syscall-exit stop with a restartable result, i.e. during recording.
During replay the task's status at such an exit is the breakpoint's
SIGTRAP, or nothing after finish_emulated_syscall() on the PTRACE_SYSEMU
path, so that never triggers, and there is no pending signal to preserve
anyway.

So apply the restart ourselves in ReplayTask::will_resume_execution()
when the registers are at such a syscall exit (orig_ax >= 0 and
syscall_may_restart()):
- Doing it when the task resumes, rather than when the exit is
  processed, keeps an event recorded exactly at the exit point matching
  the recorded registers.
- If a signal handler frame was set up in between, ax no longer holds
  -ERESTART*, and nothing happens.
- After the adjustment ax no longer holds -ERESTART* either, so on the
  breakpoint path the kernel doesn't restart the syscall a second time.
- Changing the registers gives the same result as making the kernel do
  it with a PTRACE_INTERRUPT, without an extra stop on the next resume.
- On aarch64 the kernel adjusts pc and x0 for the restart before the
  signal stop, and replay already applies that when processing the
  syscall exit, so this is x86 only.

The new test syscall_restart_in_writable_mem has the main thread block
in read() through a syscall instruction in an RWX page. Another thread
then overwrites that instruction with ud2 and has a child process send
the process SIGSTOP and SIGCONT. The main thread blocks SIGCONT and
SIGCHLD, so read() is restarted without a handler and the main thread
gets SIGILL at the syscall instruction. Without this change, replay
fails in all four x86 test variants (80 of 80 runs); with it, 200 of 200
runs pass. The full test suite on x86_64 (Linux 7.0) shows no new
failures. On aarch64 the test only checks the existing behavior, since
the problem and the fix are x86-only. There, with the syscall buffer, rr
replaces the svc in the page with a branch to its syscall hook, so the
read() would block in rr's code and the restart would never reach the
page; the test detects that and exits early.

Known issues this doesn't address (all predate it):
- In the debugger, stepi from such an exit onto a signal at the
  restarted instruction still aborts: emulate_async_signal() decides how
  to resume before the restart moves ip onto its internal breakpoint.
  continue and reverse-stepi work.
- A tracee that ptraces another process and writes an -ERESTART* value
  at its syscall-exit stop, then resumes it without a signal: the kernel
  doesn't restart that syscall, but replay does (as the breakpoint path
  already did).
- On aarch64 the kernel sets x8 to the restart_syscall number for
  -ERESTART_RESTARTBLOCK after the stop; replay doesn't reproduce that.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@Keno
Keno force-pushed the replay-restart-without-breakpoint branch from b05cda9 to 3d734e4 Compare October 1, 2026 05:36

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant