Skip to content

A client cancel during SuspendActor lets a retry crash the actor mid-checkpoint #2394

Description

@ygao-g

If a client cancels SuspendActor while atelet is still checkpointing, ateapi releases the actor's lease.
A retry then starts a second checkpoint on a sandbox that has just saved and exited. ateapi marks the actor
CRASHED, and the first checkpoint completes a second later.

Seen with 1000 actors suspended at once by a client with a 30 s per-attempt deadline:

  1. The client calls SuspendActor. atelet's Checkpoint takes 30.85 s: 4.24 s checkpoint, 2.75 s upload,
    and about 24 s queued behind 125 other checkpoints on the node.
  2. At 30 s the client cancels. ateapi logs workflow failed at step CallAteletSuspend: context canceled
    and releases the lease.
  3. The client retries 50 ms later, and ateapi starts a second suspend.
  4. The second runsc checkpoint fails with exit status 128 on the exited sandbox. ateapi logs
    Setting Actor to crashed.
  5. 1.1 s later the first checkpoint finishes and its snapshot is persisted. The actor is already CRASHED.

In one run whose suspend took 30.1 s, 45 of 1000 actors ended CRASHED this way.

Fix

ateapi must not start a second checkpoint while one is in flight for the actor. A client cancel should not
release the lease while the atelet call continues.

Activity

  1. doniacld commented on Oct 9, 2026

    @doniacld
    Contributor

    Hi! I know that it's not yet triaged yet but I am happy to pick it up if that's welcome.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions