Skip to content

Replication W9: Exclusive distributed locking (replication integration) #438

Description

@kriszyp

Workstream W9 of #430 · exclusive distributed locking — replication integration

Summary

Implement table.lock(primaryKey): Promise<WritableRecord> distributed exclusive locking. The core API and semantics are tracked in harper#483 (Table.ts lock() is currently a stub: throw new Error('Not yet implemented'), line 1735). This workstream tracks the replication-protocol integration and the coordination with the core change. Locking is the explicit-serialization escape hatch from wall-clock LWW lost-updates (epic Theme G).

Design direction (fits the real substrate)

  • LOCKED metadata bit — a free bit in the RecordEncoder uint32 metadata bitmap (alongside INVALIDATED/EVICTED/LOCAL_ONLY). Because it lives in the already-decoded metadata integer, the write/send paths can test "locked / by whom" without decoding the record value (the same throughput trick LOCAL_ONLY uses).
  • LOCK/UNLOCK audit types — a lock transition is a version-changing audit entry that replicates through the existing pipe for free (aftercommit → subscription → SUBSCRIPTION_UPDATE). Peers applying it set their local LOCKED bit.
  • Write-gating — a non-holder write checks the LOCKED bit before LWW resolution and defers/rejects, converting silent lost-updates into explicit serialization. Unlock (a version change) lets queued writes proceed in version order.
  • Grant set follows residency — only resident nodes can hold/serve the record, so only they need to grant (far cheaper than whole-cluster).
  • Grant handshake rides the existing OPERATION_REQUEST/RESPONSE (136/137) RPC; durable lock state rides the audit stream.
  • Leases + fencing by lock version — lease via the existing expiresAt machinery; a write must present a lock version ≥ the record's current lock version, so a stale holder's write is rejected by precedesExistingVersion regardless of clock skew (mitigates Theme G).

Phasing

  • Phase 0 — local-only lock. Over the existing cross-thread store.tryLock + the LOCAL_ONLY metadata bit. No replication, no protocol change. Validates the metadata/version-change wiring and the WritableRecord API surface.
  • Phase 1 — single-owner delegation. Replicated LOCK/UNLOCK; if another node holds it, request via OPERATION_REQUEST from that one node; write-gating on the LOCKED bit; leases + fencing. Covers the common "one node owns the hot key" case with point-to-point RPC.
  • Phase 2 — broadcast / residency-quorum grant. Shared→exclusive upgrade: fan out grant requests to the residency holder set, Promise.all. Deadlock avoidance via global key ordering + Wound-Wait by lock timestamp. Highest-risk piece.

Hardest correctness problems

  • Clock-skew-safe lease expiry — conservative lease margins + version fencing (a stale holder's write is fenced even if its clock disagrees).
  • Split-brain during an all-nodes grant — require durable-commit confirmation of the lock entry; version fencing ensures two exclusive holders can never both pass write-gating.
  • Holder crash — lease auto-expire sweep emits UNLOCK (itself a version change) to unblock waiters.
  • Node-owned vs transaction-owned locks — an option on lock(); transaction-owned auto-unlock at commit, node-owned persist.

Dependencies

Core metadata/audit wiring (harper#483). Independent of W1/W4 — Phase 0/1 can start in parallel with Phase 2 of the epic once the metadata bit + audit type land.

Effort / risk

Phase 0 M / low · Phase 1 L / medium · Phase 2 L / high.

Acceptance criteria

  • lock() serializes concurrent multi-node writes to a key with no LWW lost update.
  • A crashed lock-holder's lock auto-expires and waiters proceed.
  • A stale (lease-expired) holder's write is fenced.

🤖 Filed by Claude on behalf of Kris.

Activity

  1. added this to the v5.3 milestone on Jun 20, 2026
  2. added theissue type on Aug 7, 2026
  3. kriszyp commented on Sep 3, 2026

    @kriszyp
    MemberAuthor

    W9 Phase 1 protocol design (2026-09-03): cluster-wide lock() as Ricart–Agrawala over replicated control entries

    Premise change. Phase 0 landed differently from the design direction above: harper#2462 makes the rocksdb-js in-memory key lock the sole authority — no LOCKED metadata bit, no LOCK/UNLOCK record writes, no write-gating of plain writes, and no version movement for a lock. The bullets above about the LOCKED bit, write-gating and precedesExistingVersion fencing are superseded. Phase 1 therefore carries lock coordination as control transaction-log entries (the reload marker precedent), never as record rewrites, and mutual exclusion is Ricart–Agrawala with total order (timestamp, nodeId).

    Contract (unchanged from Phase 0, made cluster-wide)

    • lock() is mutually exclusive only with other lock() calls on the same key — now across every participating node. Plain writes are never gated.
    • Holder writes are stamped with the lock's acquisition timestamp, which in cluster mode is the LOCK_REQUEST entry's timestamp ts_R. Any write committed anywhere after the request was durable carries a later stamp and wins by LWW, so "a write during a lock period is treated as happening after it" holds cluster-wide, and a stale holder (lease expired) can never beat a newer holder: the newer holder's ts_R' > ts_R. Fencing falls out of the stamp; no precedesExistingVersion change. Phase 0's 409-on-expired-handle stays as the local guard.

    Participants and the mixed-version rule

    • Participant set for a (database, table, key): nodes that replicate that database (hdb_nodes replicates:true, subscription desired) and advertise recordLocks in the W11 capability registry (Replication W11: Code organization, protocol versioning & decode-loop safety #440, task hp-440-protocol-capabilities; the design there explicitly left recordLocks for W9).
    • Fail-closed by default: if any desired peer for the database lacks recordLocks, lock() rejects with a retryable 503 (LockUnavailable) unless the caller asked for { scope: 'node' } (Phase 0 semantics). lock() is new in 5.3, so nothing depends on it during the 5.2→5.3 rolling window; fail-closed is the only choice that never hands two nodes the same key.
    • Layering: the rocksdb-js key lock remains the intra-node authority. The cluster round only starts once the local key is held, so each node has at most one outstanding request per key.

    Entries (harper core)

    Three new action-nibble values in auditStore.ts — LOCK_REQUEST = 9, LOCK_GRANT = 10, LOCK_RELEASE = 12 (11 is REMOTE_SEQUENCE_UPDATE, 13 spare, 14/15 are the width flags). Each is written to the writer's own origin log in its own transaction exactly like Table.writeReloadMarker (recordId = the key so per-key locality holds; no record value; tableToTrack: null), and is not LOCAL_ONLY — it must replicate.

    entry written by fields (entry header extension, no value decode needed)
    LOCK_REQUEST requester R ts_R (= the entry's own txn timestamp / log key = identity), leaseMs, waitMs
    LOCK_GRANT participant P request identity (R, ts_R)
    LOCK_RELEASE holder or requester R request identity (R, ts_R) — also the withdraw for a request that timed out (423) before acquiring

    Receiver routing: the replicated-event consumer in core (Table.ts writeUpdate switch, the sink of harper-pro's tableSubscriptionToReplicator.send(event)) dispatches the three types to the table's LockCoordinator, never to _writeUpdate; the subscriber fan-out filter (Table.ts:4353, where end_txn/reload are skipped) treats them the same way, so MQTT/SSE subscribers never see them.

    Protocol per key (Ricart–Agrawala)

    1. Request. R holds the local key, writes LOCK_REQUEST(ts_R), then waits for a LOCK_GRANT(R, ts_R) from every participant P ≠ R.
    2. On LOCK_REQUEST(Q, ts_Q) at P. If P currently holds the key (cluster-granted), or P has its own pending request with (ts_P, P) < (ts_Q, Q), P defers; otherwise P writes LOCK_GRANT(Q, ts_Q). Deferred grants are written, in (ts, nodeId) order, when P releases or withdraws.
    3. Release. On unlock(), transaction end (scoped lock), or lease expiry, R writes LOCK_RELEASE(R, ts_R). On applying it, P drops the hold and grants its deferred queue.
    4. Acquired when all grants are in. Exclusion: a holder never grants until it releases. Deadlock-free per key by the total order; multi-key deadlocks across nodes resolve by waitMs timeout (423) — wound-wait stays Phase 2.

    Leases, crashes, partitions

    • Lease. Every participant computes a hold's expiry from its own apply time + leaseMs + LOCK_LEASE_SKEW_MS (default 5 s); the holder uses acquiredAt + leaseMs with no margin, so the holder always expires first (its writes get 409) before any participant treats the hold as released and grants a deferred request. Expiry needs no RELEASE; a late RELEASE is a no-op.
    • Pending requests expire too: at ts_R + waitMs + skew, a request nobody acquired is treated as withdrawn (covers a requester that crashed before writing its withdraw).
    • Peer down (W1 truth DOWN). A requester blocks (up to waitMs) while a participant is down for less than maxLease + skew. Once a node has been down longer than that, every hold or request it could have had has expired, so it is excluded from the grant set for new requests. Holds always expire before exclusion, so a two-holder window never opens; asymmetric partitions get the same rule from each side.
    • Joining node / base copy. A joiner participates for requests with ts > copyStartTime. Safety only needs holders to defer, and a holder is by definition a node that participated, so a joiner granting cannot break exclusion.
    • Replay / duplicates. All three entries are idempotent by identity (R, ts_R); a replayed REQUEST older than now − (maxLease + waitMs + skew) is ignored on arrival, so cursor replay after reconnect is harmless. Control entries are tiny and go through normal audit cleanup.
    • Topology. Entries flow like data through relays, so hub-and-spoke works without new plumbing; grant latency is two replication hops. Identity per origin log is exactly what the dual-clock decision (harper#2412) makes durable.

    ⚠ Dual-clock interaction (needs Kris's sequencing call)

    Phase 0 pins the transaction clock to acquiredAt for holder writes. Under the settled dual-clock model (harper#2412 / rocksdb-js#811: first word = txn timestamp = origin log key = identity; only the replication receiver and crash replay may setTimestamp), the lock-ordered stamp belongs in the distinct-version second word (HAS_DISTINCT_VERSION_FLAG), not the first. If #2412 lands first, Phase 1 stamps there from the start; otherwise Phase 1 keeps the Phase 0 mechanism and the move is fenced to #2412.

    Observability

    cluster_status per database: locks: { held, pending, deferred, expiredHolds, fencedWrites }; each expiry/exclusion event logs with the same truth snapshot the recovery nets use.

    Where the code lives, and dispatch split

    1. harper (core) — harper-483-phase1-core: action types + entry encode/decode; LockCoordinator per table (holders, pending, deferred; pure state machine with injected clock, unit-tested with a 3-node in-process fake transport for every case above); a ClusterLockTransport interface (participants(), writeControl(entry), onControlEntry); lock() integration (scope: 'cluster' | 'node', cluster default when a transport is registered; handle carries ts_R). Independent of the registry; starts now.
    2. harper-pro — hp-438-phase1-transport: transport over the audit stream (send-path gate: skip control entries to a peer without recordLocks, next to the LOCAL_ONLY skip at sendAuditRecord), participant set = hdb_nodes ∩ peerCapabilities.recordLocks ∩ truth, recordLocks key in protocolCapabilities.ts, cluster tests (3-node exclusion with no LWW lost update; holder crash → lease expiry → waiter proceeds; stale holder fenced; mixed-version fail-closed). Starts once the registry PR from hp-440 exists.

    Rejected: grant RPC over OPERATION_REQUEST/RESPONSE (not durable, not ordered with the data stream, does not relay, needs its own retry/cursor story); record-descriptor locking (the Phase 0 v1 design — leaked the descriptor onto the wire and moved the version under peers, see the #2462 thread); best-effort node-local fallback on mixed versions (silently hands two nodes the key).

  4. self-assigned this
    on Sep 7, 2026
  5. kriszyp commented on Sep 9, 2026

    @kriszyp
    MemberAuthor

    Phase 1 lands close to this issue's original plan, not to what harper-pro#822 currently implements

    The core arbitration rule on harper#2498 —
    Ricart–Agrawala over replicated control entries — is being replaced before it ships. Design note:
    docs/record-lock-ownership.md.

    The replacement is essentially this issue's "Phase 1 — single-owner delegation", generalized:

    • Home node per key, derived by rendezvous hash from the epoch's members[], rather than from
      residency. Same "ask the one node that owns it" shape as the OPERATION_REQUEST handshake sketched
      here, but with no dependence on operator-configured sharding or residency — only one customer runs
      sharding today, so a design that requires it is not shippable.
    • Delegations are volatile and the epoch is durable. A delegation is the exclusive right to admit
      critical sections on one key for a bounded time; while it is live, lock()/unlock() are the local
      Phase 0 key lock with zero cluster messages. Releasing the application lock does not release the
      delegation, so repeated locking by one node amortizes to nothing.
    • Leases + fencing as this issue anticipated, with one correction: the fencing token has to be
      ordered, (epochNumber, homeIncarnation, delegationCounter), with homeIncarnation a durably
      persisted monotonic counter. A random incarnation makes a stale reply identifiable but not
      orderable, and a home that restarts and re-issues generation 1 after having issued generation 50
      would let a delayed generation-50 write defeat its successor.
    • Phase 2's broadcast / residency-quorum grant is not needed. A single arbiter per key is
      trivially exclusive, so the deadlock-avoidance work listed here as "highest-risk piece" —
      Wound-Wait, global key ordering, split-vote resolution — is deleted rather than simplified.

    What this asks of harper-pro

    The membership epoch (note §4) is harper-pro's, because harper-pro owns topology. hdb_nodes cannot
    back it directly: that table is LWW-replicated and therefore not agreed. It is single-decree agreement
    per epoch number over a majority of the current members, with promises and accepted values persisted
    before they are acknowledged
    — durable, at membership-change frequency, O(1) per node, never on an
    acquisition path. §4.0 of the note carries the counterexample that killed the stateless version (a
    restarted acceptor that forgets its acceptance forks the configuration into two live epochs renewing
    on disjoint majorities), and it is worth reading before anyone proposes the cheaper variant again.

    That epoch is also what finally lets harper-pro assert agreedDown, the gap
    #822 records as a core follow-up.

    #822 keeps its capability, participant-set,
    coordination-ownership and replication.recordLocks switch work; the transport implementation is what
    changes, from broadcast control entries to unicast delegation request/grant/recall over the existing
    replication connections. The recordLocks capability becomes versioned and the versions mutually
    exclusive — a cluster running both would have two independent arbiters for one key.

    Before any of it

    The measurement gate comes first. Nothing in the note is a benchmark; they are message counts. The
    numbers needed are acquisition latency on a real cluster, control-entry and byte growth per lock,
    hot-key handoff throughput, and throughput with the feature disabled.

    The open decision — what lock() promises for conflict ordering and for freshness after a holder
    crash — is in harper#483 and sizes the work.

    🤖 Claude Opus 5

  6. kriszyp commented on Sep 9, 2026

    @kriszyp
    MemberAuthor

    Two sub-issues filed, and one bullet in this issue's design direction is now superseded

    The core guarantee decision behind the Phase 1 redesign was taken on 2026-09-09:
    lock() ships exclusion-only — exclusive admission, plus successor freshness after a clean
    handoff, and explicitly not on the recovery path — where the predecessor crashed, is unreachable,
    or simply had a native commit settle after the barrier was measured. It also does not promise that a
    predecessor's write cannot outrank its successor's under LWW, which is reachable on a clean handoff
    via a future context.timestamp. Details in
    harper#483 and §10 of the design note.

    That directly supersedes this issue's "Leases + fencing by lock version" bullet, and the
    "stale (lease-expired) holder's write is fenced" acceptance criterion with it. There is no fencing
    token in conflict resolution: a lease-expired holder is fenced at admission and at commit
    submission — Phase 0 already checks the lease synchronously on every staged write and again
    immediately before the native commit submits — but a write that escaped before expiry resolves by
    last-write-wins like any other. Fencing by generation was considered and deferred to
    harper#2540; it changes _writeUpdate's
    resolution rule for every write in the database, which is a cost deployments that never call
    lock() would pay.

    The harper-pro work is now two tracked pieces:

    Core carries harper#2541 (home ring, delegations,
    drain/recall, caps) and harper#2542 (successor
    freshness).

    🤖 Claude Opus 5

  7. modified the milestones: v5.3, v5.4 on Oct 10, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

area:replicationReplication, cluster sync, peer connectionsenhancementNew feature or request

Fields

Priority

P3

Projects

No projects

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions