Skip to content

Raft snapshot install of a page-memory partition hangs forever: OutgoingSnapshot.rowEntry NPEs on a committed tombstone and the sender never replies #13676

Description

@kagrawal-tibco

h3. Problem
If a page-memory partition (aimem or aipersist) contains a committed tombstone that GC has not yet removed, a full-state transfer (Raft InstallSnapshot) of that partition never completes. The receiving replica stays in snapshot install indefinitely. Its safe time never advances, so read-only transactions mapped to it (implicit RO SQL prefers the local replica) fail with IGN-REP-3 "replication timeout", even though the group has a healthy primary and majority.

h3. Root cause

{{ScanVersionsCursor.rowVersionToReadResult}} (modules/storage-page-memory, 3.1.0 line 108-109) returns {{ReadResult.empty(rowId)}} for a committed tombstone. {{ReadResult.empty}} has a null {{commitTimestamp}}.

#* {{RocksDbMvPartitionStorage}} returns the same version as {{ReadResult.createFromCommitted(rowId, null, rowCommitTimestamp)}}, so the two engines disagree on the {{scanVersions}} contract for tombstones.

{{PartitionMvStorageAccessImpl.getAllRowVersions}} passes every {{scanVersions}} result through unchanged.

{{OutgoingSnapshot.rowEntry}} (3.1.0 line 456; main line ~479) treats every non-write-intent version as committed and runs {{commitTimestamps[j++] = version.commitTimestamp().longValue();}}, which throws a NullPointerException for the tombstone. That covers a tombstone at the head of the chain (a deleted row) and one in the middle (deleted, then re-inserted).

{{OutgoingSnapshotsManager.respond}} (line 207-214) passes the throwable to {{failureProcessor.process(new FailureContext(throwable, "Something went wrong while handling a request"))}} and returns without sending any response.

{{IncomingSnapshotCopier}} sends its requests with {{NETWORK_TIMEOUT = Long.MAX_VALUE}} (line 93), so the copier waits forever.

#* IGNITE-27497 (shorter copier timeout) would limit the wait, but every retry would hit the same NPE as long as the tombstone exists on the sender.

h3. When it happens
All of these must be true:

  • A replica needs a full-state transfer: it restarted, or fell behind, after the leader's Raft log had been truncated by a snapshot. With hourly snapshots, that is roughly 1-2 h of uptime.
  • The partition uses a page-memory storage engine (aimem or aipersist).
  • The partition holds rows deleted by a committed transaction whose tombstones are still above the low watermark. A longer {{gc-data-availability-time}} keeps them longer.

h3. Observed
Ignite 3.1.0, embedded in TIBCO BusinessEvents, 5-node cluster. The receiver is node cache4, the sender is node cache1, and the zone uses the aimem profile. Receiver log, 2026-10-08:
{noformat}
16:36:33Z ZonePartitionKey [zoneId=21, partitionId=0] "Copier has loaded the snapshot meta" ... (cfgIndex=26341)
no "finished loading multi-versioned data" afterwards; the install never completes
17:26:39Z same partition, new sender incarnation "Copier has loaded the snapshot meta" ... (cfgIndex=29373)
again no completion
{noformat}
A later install of the same partition from another node (2026-10-09, about 17 h later) completed normally.
The sender's log for that window had already been rotated, so the NPE itself was not captured. The diagnosis comes from the code above and the bytecode of the 3.1.0 jars ({{ReadResult.empty}} passes null for both timestamps).

h3. Suggested fix

Storage: make page memory return a committed tombstone the way RocksDB does, i.e. {{ReadResult.createFromCommitted(rowId, null, rowVersion.timestamp())}} in {{ScanVersionsCursor}} (and anywhere else {{scanVersions}} maps tombstones to {{ReadResult.empty}}).

Also harden {{OutgoingSnapshotsManager.respond}}: on an unexpected exception, send an error response (or close the snapshot) so the receiver fails the install and retries, instead of waiting forever.

Test: an {{AbstractMvPartitionStorageTest}} case that commits a tombstone and checks that {{scanVersions}} returns it with a non-null {{commitTimestamp}}. We expect it to fail on page memory and pass on RocksDB; it has not been run yet.

h3. Workaround
Unverified. Once the deletes are older than {{gc-data-availability-time}} and GC has removed the tombstones, restarting the stuck node should let a new install complete.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions