LDAPReplicationDomain.runRequestedSessionRestarts() takes the request before it runs it:
while (sessionRestartRequested.getAndSet(false))
{
restartSession(wait);
}
restartSession() stops the session in its first synchronized block, waits out the backoff, and starts it again in the second. If enableService() throws there - broker.start() on an unreachable replication server, a configuration or registration failure - the exception leaves the loop with the request already cleared. Nothing asks for the restart again: the domain is left with no broker and no listener, out of the topology until the server is restarted or a configuration change happens to call restartService(), while the change which opened the recovery waits for a delivery which cannot come.
Clearing the flag only once the restart has run - or setting it back in a catch - closes it.
Two related asymmetries on the same path, worth deciding together:
runRequestedSessionRestarts(wait) applies the wait of the thread which entered it to a request another thread made, so a hand-back which does not wait can run a recovery which was owed the backoff;
abandonReplay() restarts with wait == false by design (the backend is not what is going away), which means a ds-cfg-num-update-replay-threads change can produce one stop/reconnect per abandoning thread with no delay between them.
Found while reviewing #892.
LDAPReplicationDomain.runRequestedSessionRestarts()takes the request before it runs it:restartSession()stops the session in its first synchronized block, waits out the backoff, and starts it again in the second. IfenableService()throws there -broker.start()on an unreachable replication server, a configuration or registration failure - the exception leaves the loop with the request already cleared. Nothing asks for the restart again: the domain is left with no broker and no listener, out of the topology until the server is restarted or a configuration change happens to callrestartService(), while the change which opened the recovery waits for a delivery which cannot come.Clearing the flag only once the restart has run - or setting it back in a catch - closes it.
Two related asymmetries on the same path, worth deciding together:
runRequestedSessionRestarts(wait)applies thewaitof the thread which entered it to a request another thread made, so a hand-back which does not wait can run a recovery which was owed the backoff;abandonReplay()restarts withwait == falseby design (the backend is not what is going away), which means ads-cfg-num-update-replay-threadschange can produce one stop/reconnect per abandoning thread with no delay between them.Found while reviewing #892.