Skip to content

Replication: a session restart request is consumed before the restart runs #925

Description

@vharseko

LDAPReplicationDomain.runRequestedSessionRestarts() takes the request before it runs it:

while (sessionRestartRequested.getAndSet(false))
{
  restartSession(wait);
}

restartSession() stops the session in its first synchronized block, waits out the backoff, and starts it again in the second. If enableService() throws there - broker.start() on an unreachable replication server, a configuration or registration failure - the exception leaves the loop with the request already cleared. Nothing asks for the restart again: the domain is left with no broker and no listener, out of the topology until the server is restarted or a configuration change happens to call restartService(), while the change which opened the recovery waits for a delivery which cannot come.

Clearing the flag only once the restart has run - or setting it back in a catch - closes it.

Two related asymmetries on the same path, worth deciding together:

  • runRequestedSessionRestarts(wait) applies the wait of the thread which entered it to a request another thread made, so a hand-back which does not wait can run a recovery which was owed the backoff;
  • abandonReplay() restarts with wait == false by design (the backend is not what is going away), which means a ds-cfg-num-update-replay-threads change can produce one stop/reconnect per abandoning thread with no delay between them.

Found while reviewing #892.

Activity

  1. added 20 commits that reference this issue on Sep 9, 2026
  2. added 6 commits that reference this issue on Sep 16, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions