How long a replica retries a change its backend could not apply, and how long it waits between
attempts, is decided by constants in LDAPReplicationDomain. An operator whose maintenance windows
are longer than the built-in budget cannot raise it, and the only way to exercise the policy in a
test is a public setter on a production singleton.
Today
opendj-server-legacy/src/main/java/org/opends/server/replication/plugin/LDAPReplicationDomain.java
| constant |
value |
what it decides |
IN_PLACE_REPLAY_ATTEMPTS |
10 |
how many times a replay is retried before the session is restarted |
REPLAY_GIVE_UP_DELAY_IN_MS |
5 min |
how long a change is retried before the replica gives up on it and diverges |
REPLAY_RETRY_DELAY_IN_MS |
1 s |
the backoff, multiplied by the number of restarts in a row |
MAX_REPLAY_RETRY_DELAY_IN_MS |
10 s |
the longest the session is left down |
(Edited: a fifth row listed MAX_FAILED_REPLAY_ATTEMPTS_TRACKED = 1000, "how many failing changes
are remembered". There is no such constant. It described an intermediate revision of #892, whose
bounded ReplayFailures map was removed under review - the failures live on the PendingChange
which is already the barrier, so they are bounded by the pending changes rather than by a number of
their own.)
Giving up means recording a change which was never applied as replayed: the replica diverges, raises
the org.opends.server.replication.UnreplayedChange alert and has to be reinitialized. Five minutes
is a reasonable default, but an import-ldif or a rebuild-index on a large backend outlasts it, and
the administrator who knows that has no way to say so.
What it would look like
The domain already carries this kind of knob in ReplicationDomainCfg - replay-thread-number,
heartbeat-interval, changetime-heartbeat-interval, assured-timeout - so the natural shape is a
property next to them, in
opendj-maven-plugin/src/main/resources/config/xml/org/forgerock/opendj/server/config/ReplicationDomainConfiguration.xml,
with the admin guide entry which comes with it. changeConfig() already re-reads that configuration
when it changes.
At least replay-give-up-delay is worth exposing; the backoff and the in-place attempt count could
stay as they are, or follow it.
Decided while implementing this
Only replay-give-up-delay becomes a property. The other three constants stay:
REPLAY_RETRY_DELAY_IN_MS and MAX_REPLAY_RETRY_DELAY_IN_MS are waited out on the replay threads,
which are shared by every domain of the server - that is why the ceiling is deliberately short. A
per-domain knob over a wait which runs on a shared pool lets one sick domain starve the healthy
ones; if they are ever exposed, their place is ReplicationSynchronizationProviderConfiguration,
next to num-update-replay-threads, not the domain.
IN_PLACE_REPLAY_ATTEMPTS is the micro-retry over a lock or a storage which is busy for a moment
(OPENDJ-885). There is nothing for an operator to decide there, and the tests read it as the known
boundary between two deliveries.
The property takes unlimited as well as a duration, for the operator who would rather have the
replication of a domain stop than have it diverge silently. Two things that value makes visible are
filed on their own: #942 (the retry warning is logged once per delivery, and only the five-minute
constant used to bound that flood) and #943 (applyConfigurationChange() publishes a configuration
it may then report as failed).
Why this is separate
Doing it well means a generated configuration property, its documentation, and a decision about which
of the five knobs deserve to be public. #892 introduced the budget as a constant with a
@VisibleForTesting setter (setReplayGiveUpDelay()), which the tests use because they cannot wait
five minutes. A real property would remove that setter, remove the constant the test duplicates, and
answer the operator question at the same time.
Raised from the review of #892.
How long a replica retries a change its backend could not apply, and how long it waits between
attempts, is decided by constants in
LDAPReplicationDomain. An operator whose maintenance windowsare longer than the built-in budget cannot raise it, and the only way to exercise the policy in a
test is a public setter on a production singleton.
Today
opendj-server-legacy/src/main/java/org/opends/server/replication/plugin/LDAPReplicationDomain.javaIN_PLACE_REPLAY_ATTEMPTSREPLAY_GIVE_UP_DELAY_IN_MSREPLAY_RETRY_DELAY_IN_MSMAX_REPLAY_RETRY_DELAY_IN_MS(Edited: a fifth row listed
MAX_FAILED_REPLAY_ATTEMPTS_TRACKED = 1000, "how many failing changesare remembered". There is no such constant. It described an intermediate revision of #892, whose
bounded
ReplayFailuresmap was removed under review - the failures live on thePendingChangewhich is already the barrier, so they are bounded by the pending changes rather than by a number of
their own.)
Giving up means recording a change which was never applied as replayed: the replica diverges, raises
the
org.opends.server.replication.UnreplayedChangealert and has to be reinitialized. Five minutesis a reasonable default, but an
import-ldifor arebuild-indexon a large backend outlasts it, andthe administrator who knows that has no way to say so.
What it would look like
The domain already carries this kind of knob in
ReplicationDomainCfg-replay-thread-number,heartbeat-interval,changetime-heartbeat-interval,assured-timeout- so the natural shape is aproperty next to them, in
opendj-maven-plugin/src/main/resources/config/xml/org/forgerock/opendj/server/config/ReplicationDomainConfiguration.xml,with the admin guide entry which comes with it.
changeConfig()already re-reads that configurationwhen it changes.
At least
replay-give-up-delayis worth exposing; the backoff and the in-place attempt count couldstay as they are, or follow it.
Decided while implementing this
Only
replay-give-up-delaybecomes a property. The other three constants stay:REPLAY_RETRY_DELAY_IN_MSandMAX_REPLAY_RETRY_DELAY_IN_MSare waited out on the replay threads,which are shared by every domain of the server - that is why the ceiling is deliberately short. A
per-domain knob over a wait which runs on a shared pool lets one sick domain starve the healthy
ones; if they are ever exposed, their place is
ReplicationSynchronizationProviderConfiguration,next to
num-update-replay-threads, not the domain.IN_PLACE_REPLAY_ATTEMPTSis the micro-retry over a lock or a storage which is busy for a moment(OPENDJ-885). There is nothing for an operator to decide there, and the tests read it as the known
boundary between two deliveries.
The property takes
unlimitedas well as a duration, for the operator who would rather have thereplication of a domain stop than have it diverge silently. Two things that value makes visible are
filed on their own: #942 (the retry warning is logged once per delivery, and only the five-minute
constant used to bound that flood) and #943 (
applyConfigurationChange()publishes a configurationit may then report as failed).
Why this is separate
Doing it well means a generated configuration property, its documentation, and a decision about which
of the five knobs deserve to be public. #892 introduced the budget as a constant with a
@VisibleForTestingsetter (setReplayGiveUpDelay()), which the tests use because they cannot waitfive minutes. A real property would remove that setter, remove the constant the test duplicates, and
answer the operator question at the same time.
Raised from the review of #892.