Skip to content

A resumed subscription can skip a transaction that commits after its cursor with a lower key #2928

Description

@dawsontoth

A resume cursor is a transaction-log key, and resuming replays everything after it (auditStore.getRange({ start: cursor, exclusiveStart: true })). But on RocksDB a transaction's key is assigned when the transaction is created, not when it commits. Transaction.getTimestamp() in @harperfast/rocksdb-js defaults to "a process-wide monotonic value assigned when the transaction was created", and some write paths set it from JS time instead (DatabaseTransaction calls transaction.setTimestamp(txnTime)). So a transaction can commit after a cursor has already moved past its key, and then it is never replayed:

  1. Transaction A gets key 100 and stays open.
  2. Transaction B gets key 200, commits, and is delivered and acknowledged. The consumer persists 200.
  3. A commits. The live broadcast, which dispatches in log order, delivers it, but the consumer disconnects before acknowledging it.
  4. The consumer resumes from 200. A's key is below the cursor, so the replay never includes it, and nothing reports the loss.

Who is affected:

Candidate fixes:

  • Per-log continuation positions. Resume each transaction log from an anchor transaction in append order, using the machinery replication already has (getRange({ startByLog, exactStart, resumeAfterExactStart }) and the endTxn markers). A resume position becomes a per-log map instead of a scalar. Single-record history walks, which follow version links by key, need their own answer.
  • Cap cursors at the oldest open local write transaction. Scalar cursors stay, but a persisted cursor never passes the key of a local write transaction that is still open. This is only as safe as the tracking's coverage of write paths, and it does not help replicated writes.

Found in the planning review of #2448's MQTT durable-session change, which keeps scalar cursors and documents this gap.

Activity

  1. added theissue type on Sep 30, 2026
  2. kriszyp commented on Oct 5, 2026

    @kriszyp
    Member

    This affects live delivery as well as resume. includeSuperseded subscriptions (every MQTT durable QoS>0 subscription, server/DurableSubscriptionsSession.ts:308) and previousCount subscriptions still drop a transaction that commits late with a lower key, while the subscription is open:

    • Both branches of Table.subscribe set subscription.startTime to the newest key their initial read saw (cursorMaxTime). For previousCount that happens at resources/Table.ts:6763-6770. For the current-state branch it is :6806-6814, which the subscriber fix PR below keeps for includeSuperseded only.
    • The broadcaster skips any entry whose key is at or below that time (resources/transactionBroadcast.ts:296).
    • A RocksDB transaction's key is assigned when the transaction is created. So a transaction that started before the read but commits after it is skipped live, not just on resume. For an at-least-once QoS>0 subscriber that means a message or update is never sent.

    For default subscriptions this is fixed in the subscriber PR for #3024 and #2933, which gates by record version instead of time. That approach does not carry over to includeSuperseded and previousCount, because their events are historical records rather than current state. They need the per-log continuation positions (or the open-transaction cap) proposed here. Run a database's async commits on up to four concurrent commit threads (HarperFast/rocksdb-js#902) makes late commits routine. Related: #3024.

    🤖 Generated by Claude Code (Claude Opus 5.5); posted via @kriszyp.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    Fields

    Priority

    P2

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions