Replicated objectClass value additions are not written to the objectClass equality index
Summary
In a three-way replicated topology, when a modify operation adds a value to
objectClass, the replica that receives the operation directly indexes it
correctly, but replicas that receive the same change via replication apply it to
the stored entry without updating their local objectClass equality index.
Within the same replayed operation, objectClass value deletions are indexed
correctly on all replicas. Only additions are skipped.
The result is that (objectClass=X) searches silently return incomplete results
on the replicas that did not receive the write directly. No error is logged and
no error is returned to the client.
Environment
- OpenDJ Server 5.0.2, build 20251125082946, revision
16fb4328facfa08d2d6beb68a8d607bdf0ef5e15
- Java 11.0.25+9-LTS (Red Hat)
- RHEL 9, kernel 5.14.0
- Three servers, each both a directory server and a replication server,
fully meshed
- Backend:
userRoot, JE, ~1927 entries
objectClass index: equality only, index-entry-limit: 4000
- The affected objectClass is a site-defined AUXILIARY class with ~400 entries,
well under the entry limit
Steps to reproduce
-
Three-replica topology, objectClass index in a verified-clean state on all
three (verify-index -i objectClass --countErrors returns 0).
-
Confirm all replicas agree — index candidate count equals the number of
entries carrying the class:
ldapsearch -h <host> -x -D <rootdn> -w <pw> \
-b "<baseDN>" -s sub '(objectClass=MyAuxClass)' debugsearchindex
All three report [INDEX:objectClass.equality][COUNT:406] ... final=[COUNT:406].
-
Issue a single modify against one replica that adds one objectClass value
and deletes four others:
dn: <entry>
changetype: modify
add: objectClass
objectClass: MyAuxClass
-
delete: objectClass
objectClass: posixAccount
objectClass: shadowAccount
objectClass: sambaSamAccount
objectClass: inetLocalMailRecipient
-
Allow replication to settle, then re-run the debugsearchindex query on all
three replicas.
Expected
All three replicas report an index candidate count of 407.
Actual
- Replica that received the write:
[COUNT:407] ... final=[COUNT:407]
- Other two replicas:
[COUNT:406] ... final=[COUNT:406]
The entry itself is correct everywhere. A base-scoped lookup on the replicas that
failed to index returns the entry with objectClass: MyAuxClass present. Only
the index differs.
The state persists indefinitely — measured again several hours later, unchanged.
Replication metadata is identical on all three
entryUUID: ce3a9241-8ee5-4ec2-9d2a-f2762de40615
modifyTimestamp: 20260909134015Z
ds-sync-hist: objectClass:000001a08665d4391c040000042f:add:MyAuxClass
ds-sync-hist: objectClass:000001a08665d4391c040000042f:del:inetLocalMailRecipient
ds-sync-hist: objectClass:000001a08665d4391c040000042f:del:sambaSamAccount
ds-sync-hist: objectClass:000001a08665d4391c040000042f:del:shadowAccount
ds-sync-hist: objectClass:000001a08665d4391c040000042f:del:posixAccount
Same CSN, same UUID, same history on every replica. Replication delivered the
change correctly; only the local index write differs.
Deletions from the same operation index correctly
On the replicas that failed to index the addition, all four deleted
objectClasses show index candidate counts equal to their final result counts,
and those counts dropped to reflect the deletion within minutes of the replay.
So within a single CSN, the del: entries update the index on the receiving
replica and the add: entry does not.
Impact
The failure is silent. Nothing appears in logs/errors or logs/replication on
the affected replicas. Clients receive a successful search response with an
incomplete result set.
In our deployment, writes are distributed across replicas by a round-robin DNS
alias. Each affected entry ends up indexed on exactly one replica and missing
from the other two. Over time this accumulated to 81 entries with no
symptom other than an unrelated consistency check happening to intersect two of
them.
Confirmation via verify-index
Offline on one replica before repair:
verify-index -b <baseDN> -i objectClass --countErrors # exit code 54
54 matches exactly the difference between the entries carrying the objectClass
(413) and the index candidate count on that replica (359) at the time.
After rebuild-index --offline -b <baseDN> -i objectClass, the same command
returns 0.
Workaround
rebuild-index -i objectClass restores consistency. This must be repeated after
every batch of writes that add objectClass values, since each new addition
reintroduces the problem on two of three replicas.
Notes
verify-index -i objectClass -c (the cleanliness check) produced no output and
exited 255 in our environment, both with the backend disabled and with the server
fully stopped. The completeness check in the same conditions returned a correct
count. This may be a separate issue.
Replicated objectClass value additions are not written to the objectClass equality index
Summary
In a three-way replicated topology, when a modify operation adds a value to
objectClass, the replica that receives the operation directly indexes itcorrectly, but replicas that receive the same change via replication apply it to
the stored entry without updating their local
objectClassequality index.Within the same replayed operation, objectClass value deletions are indexed
correctly on all replicas. Only additions are skipped.
The result is that
(objectClass=X)searches silently return incomplete resultson the replicas that did not receive the write directly. No error is logged and
no error is returned to the client.
Environment
16fb4328facfa08d2d6beb68a8d607bdf0ef5e15fully meshed
userRoot, JE, ~1927 entriesobjectClassindex: equality only,index-entry-limit: 4000well under the entry limit
Steps to reproduce
Three-replica topology,
objectClassindex in a verified-clean state on allthree (
verify-index -i objectClass --countErrorsreturns 0).Confirm all replicas agree — index candidate count equals the number of
entries carrying the class:
All three report
[INDEX:objectClass.equality][COUNT:406] ... final=[COUNT:406].Issue a single modify against one replica that adds one objectClass value
and deletes four others:
Allow replication to settle, then re-run the
debugsearchindexquery on allthree replicas.
Expected
All three replicas report an index candidate count of 407.
Actual
[COUNT:407] ... final=[COUNT:407][COUNT:406] ... final=[COUNT:406]The entry itself is correct everywhere. A base-scoped lookup on the replicas that
failed to index returns the entry with
objectClass: MyAuxClasspresent. Onlythe index differs.
The state persists indefinitely — measured again several hours later, unchanged.
Replication metadata is identical on all three
Same CSN, same UUID, same history on every replica. Replication delivered the
change correctly; only the local index write differs.
Deletions from the same operation index correctly
On the replicas that failed to index the addition, all four deleted
objectClasses show index candidate counts equal to their final result counts,
and those counts dropped to reflect the deletion within minutes of the replay.
So within a single CSN, the
del:entries update the index on the receivingreplica and the
add:entry does not.Impact
The failure is silent. Nothing appears in
logs/errorsorlogs/replicationonthe affected replicas. Clients receive a successful search response with an
incomplete result set.
In our deployment, writes are distributed across replicas by a round-robin DNS
alias. Each affected entry ends up indexed on exactly one replica and missing
from the other two. Over time this accumulated to 81 entries with no
symptom other than an unrelated consistency check happening to intersect two of
them.
Confirmation via verify-index
Offline on one replica before repair:
54 matches exactly the difference between the entries carrying the objectClass
(413) and the index candidate count on that replica (359) at the time.
After
rebuild-index --offline -b <baseDN> -i objectClass, the same commandreturns 0.
Workaround
rebuild-index -i objectClassrestores consistency. This must be repeated afterevery batch of writes that add objectClass values, since each new addition
reintroduces the problem on two of three replicas.
Notes
verify-index -i objectClass -c(the cleanliness check) produced no output andexited 255 in our environment, both with the backend disabled and with the server
fully stopped. The completeness check in the same conditions returned a correct
count. This may be a separate issue.