Repository navigation
Replication W9: Exclusive distributed locking (replication integration) #438
Description
Activity
- addedenhancementNew feature or requestNew feature or requestarea:replicationReplication, cluster sync, peer connectionsReplication, cluster sync, peer connections
on Jun 20, 2026 W9 Phase 1 protocol design (2026-09-03): cluster-wide
lock()as Ricart–Agrawala over replicated control entriesPremise change. Phase 0 landed differently from the design direction above: harper#2462 makes the rocksdb-js in-memory key lock the sole authority — no
LOCKEDmetadata bit, noLOCK/UNLOCKrecord writes, no write-gating of plain writes, and no version movement for a lock. The bullets above about theLOCKEDbit, write-gating andprecedesExistingVersionfencing are superseded. Phase 1 therefore carries lock coordination as control transaction-log entries (thereloadmarker precedent), never as record rewrites, and mutual exclusion is Ricart–Agrawala with total order(timestamp, nodeId).Contract (unchanged from Phase 0, made cluster-wide)
lock()is mutually exclusive only with otherlock()calls on the same key — now across every participating node. Plain writes are never gated.- Holder writes are stamped with the lock's acquisition timestamp, which in cluster mode is the
LOCK_REQUESTentry's timestampts_R. Any write committed anywhere after the request was durable carries a later stamp and wins by LWW, so "a write during a lock period is treated as happening after it" holds cluster-wide, and a stale holder (lease expired) can never beat a newer holder: the newer holder'sts_R'>ts_R. Fencing falls out of the stamp; noprecedesExistingVersionchange. Phase 0's 409-on-expired-handle stays as the local guard.
Participants and the mixed-version rule
- Participant set for a
(database, table, key): nodes that replicate that database (hdb_nodesreplicates:true, subscription desired) and advertiserecordLocksin the W11 capability registry (Replication W11: Code organization, protocol versioning & decode-loop safety #440, taskhp-440-protocol-capabilities; the design there explicitly leftrecordLocksfor W9). - Fail-closed by default: if any desired peer for the database lacks
recordLocks,lock()rejects with a retryable 503 (LockUnavailable) unless the caller asked for{ scope: 'node' }(Phase 0 semantics).lock()is new in 5.3, so nothing depends on it during the 5.2→5.3 rolling window; fail-closed is the only choice that never hands two nodes the same key. - Layering: the rocksdb-js key lock remains the intra-node authority. The cluster round only starts once the local key is held, so each node has at most one outstanding request per key.
Entries (harper core)
Three new action-nibble values in
auditStore.ts—LOCK_REQUEST = 9,LOCK_GRANT = 10,LOCK_RELEASE = 12(11 isREMOTE_SEQUENCE_UPDATE, 13 spare, 14/15 are the width flags). Each is written to the writer's own origin log in its own transaction exactly likeTable.writeReloadMarker(recordId= the key so per-key locality holds; no record value;tableToTrack: null), and is notLOCAL_ONLY— it must replicate.entry written by fields (entry header extension, no value decode needed) LOCK_REQUESTrequester R ts_R(= the entry's own txn timestamp / log key = identity),leaseMs,waitMsLOCK_GRANTparticipant P request identity (R, ts_R)LOCK_RELEASEholder or requester R request identity (R, ts_R)— also the withdraw for a request that timed out (423) before acquiringReceiver routing: the replicated-event consumer in core (
Table.tswriteUpdateswitch, the sink of harper-pro'stableSubscriptionToReplicator.send(event)) dispatches the three types to the table'sLockCoordinator, never to_writeUpdate; the subscriber fan-out filter (Table.ts:4353, whereend_txn/reloadare skipped) treats them the same way, so MQTT/SSE subscribers never see them.Protocol per key (Ricart–Agrawala)
- Request. R holds the local key, writes
LOCK_REQUEST(ts_R), then waits for aLOCK_GRANT(R, ts_R)from every participantP ≠ R. - On
LOCK_REQUEST(Q, ts_Q)at P. If P currently holds the key (cluster-granted), or P has its own pending request with(ts_P, P) < (ts_Q, Q), P defers; otherwise P writesLOCK_GRANT(Q, ts_Q). Deferred grants are written, in(ts, nodeId)order, when P releases or withdraws. - Release. On
unlock(), transaction end (scoped lock), or lease expiry, R writesLOCK_RELEASE(R, ts_R). On applying it, P drops the hold and grants its deferred queue. - Acquired when all grants are in. Exclusion: a holder never grants until it releases. Deadlock-free per key by the total order; multi-key deadlocks across nodes resolve by
waitMstimeout (423) — wound-wait stays Phase 2.
Leases, crashes, partitions
- Lease. Every participant computes a hold's expiry from its own apply time +
leaseMs+LOCK_LEASE_SKEW_MS(default 5 s); the holder usesacquiredAt + leaseMswith no margin, so the holder always expires first (its writes get 409) before any participant treats the hold as released and grants a deferred request. Expiry needs noRELEASE; a lateRELEASEis a no-op. - Pending requests expire too: at
ts_R + waitMs + skew, a request nobody acquired is treated as withdrawn (covers a requester that crashed before writing its withdraw). - Peer down (W1 truth DOWN). A requester blocks (up to
waitMs) while a participant is down for less thanmaxLease + skew. Once a node has been down longer than that, every hold or request it could have had has expired, so it is excluded from the grant set for new requests. Holds always expire before exclusion, so a two-holder window never opens; asymmetric partitions get the same rule from each side. - Joining node / base copy. A joiner participates for requests with
ts > copyStartTime. Safety only needs holders to defer, and a holder is by definition a node that participated, so a joiner granting cannot break exclusion. - Replay / duplicates. All three entries are idempotent by identity
(R, ts_R); a replayedREQUESTolder thannow − (maxLease + waitMs + skew)is ignored on arrival, so cursor replay after reconnect is harmless. Control entries are tiny and go through normal audit cleanup. - Topology. Entries flow like data through relays, so hub-and-spoke works without new plumbing; grant latency is two replication hops. Identity per origin log is exactly what the dual-clock decision (harper#2412) makes durable.
⚠ Dual-clock interaction (needs Kris's sequencing call)
Phase 0 pins the transaction clock to
acquiredAtfor holder writes. Under the settled dual-clock model (harper#2412 / rocksdb-js#811: first word = txn timestamp = origin log key = identity; only the replication receiver and crash replay maysetTimestamp), the lock-ordered stamp belongs in the distinct-version second word (HAS_DISTINCT_VERSION_FLAG), not the first. If #2412 lands first, Phase 1 stamps there from the start; otherwise Phase 1 keeps the Phase 0 mechanism and the move is fenced to #2412.Observability
cluster_statusper database:locks: { held, pending, deferred, expiredHolds, fencedWrites }; each expiry/exclusion event logs with the same truth snapshot the recovery nets use.Where the code lives, and dispatch split
- harper (core) —
harper-483-phase1-core: action types + entry encode/decode;LockCoordinatorper table (holders, pending, deferred; pure state machine with injected clock, unit-tested with a 3-node in-process fake transport for every case above); aClusterLockTransportinterface (participants(),writeControl(entry),onControlEntry);lock()integration (scope: 'cluster' | 'node', cluster default when a transport is registered; handle carriests_R). Independent of the registry; starts now. - harper-pro —
hp-438-phase1-transport: transport over the audit stream (send-path gate: skip control entries to a peer withoutrecordLocks, next to theLOCAL_ONLYskip atsendAuditRecord), participant set =hdb_nodes∩peerCapabilities.recordLocks∩ truth,recordLockskey inprotocolCapabilities.ts, cluster tests (3-node exclusion with no LWW lost update; holder crash → lease expiry → waiter proceeds; stale holder fenced; mixed-version fail-closed). Starts once the registry PR fromhp-440exists.
Rejected: grant RPC over
OPERATION_REQUEST/RESPONSE(not durable, not ordered with the data stream, does not relay, needs its own retry/cursor story); record-descriptor locking (the Phase 0 v1 design — leaked the descriptor onto the wire and moved the version under peers, see the #2462 thread); best-effort node-local fallback on mixed versions (silently hands two nodes the key).Phase 1 lands close to this issue's original plan, not to what harper-pro#822 currently implements
The core arbitration rule on harper#2498 —
Ricart–Agrawala over replicated control entries — is being replaced before it ships. Design note:
docs/record-lock-ownership.md.The replacement is essentially this issue's "Phase 1 — single-owner delegation", generalized:
- Home node per key, derived by rendezvous hash from the epoch's
members[], rather than from
residency. Same "ask the one node that owns it" shape as theOPERATION_REQUESThandshake sketched
here, but with no dependence on operator-configured sharding or residency — only one customer runs
sharding today, so a design that requires it is not shippable. - Delegations are volatile and the epoch is durable. A delegation is the exclusive right to admit
critical sections on one key for a bounded time; while it is live,lock()/unlock()are the local
Phase 0 key lock with zero cluster messages. Releasing the application lock does not release the
delegation, so repeated locking by one node amortizes to nothing. - Leases + fencing as this issue anticipated, with one correction: the fencing token has to be
ordered,(epochNumber, homeIncarnation, delegationCounter), withhomeIncarnationa durably
persisted monotonic counter. A random incarnation makes a stale reply identifiable but not
orderable, and a home that restarts and re-issues generation 1 after having issued generation 50
would let a delayed generation-50 write defeat its successor. - Phase 2's broadcast / residency-quorum grant is not needed. A single arbiter per key is
trivially exclusive, so the deadlock-avoidance work listed here as "highest-risk piece" —
Wound-Wait, global key ordering, split-vote resolution — is deleted rather than simplified.
What this asks of harper-pro
The membership epoch (note §4) is harper-pro's, because harper-pro owns topology.
hdb_nodescannot
back it directly: that table is LWW-replicated and therefore not agreed. It is single-decree agreement
per epoch number over a majority of the current members, with promises and accepted values persisted
before they are acknowledged — durable, at membership-change frequency, O(1) per node, never on an
acquisition path. §4.0 of the note carries the counterexample that killed the stateless version (a
restarted acceptor that forgets its acceptance forks the configuration into two live epochs renewing
on disjoint majorities), and it is worth reading before anyone proposes the cheaper variant again.That epoch is also what finally lets harper-pro assert
agreedDown, the gap
#822 records as a core follow-up.#822 keeps its capability, participant-set,
coordination-ownership andreplication.recordLocksswitch work; the transport implementation is what
changes, from broadcast control entries to unicast delegation request/grant/recall over the existing
replication connections. TherecordLockscapability becomes versioned and the versions mutually
exclusive — a cluster running both would have two independent arbiters for one key.Before any of it
The measurement gate comes first. Nothing in the note is a benchmark; they are message counts. The
numbers needed are acquisition latency on a real cluster, control-entry and byte growth per lock,
hot-key handoff throughput, and throughput with the feature disabled.The open decision — what
lock()promises for conflict ordering and for freshness after a holder
crash — is in harper#483 and sizes the work.🤖 Claude Opus 5
- Home node per key, derived by rendezvous hash from the epoch's
Two sub-issues filed, and one bullet in this issue's design direction is now superseded
The core guarantee decision behind the Phase 1 redesign was taken on 2026-09-09:
lock()ships exclusion-only — exclusive admission, plus successor freshness after a clean
handoff, and explicitly not on the recovery path — where the predecessor crashed, is unreachable,
or simply had a native commit settle after the barrier was measured. It also does not promise that a
predecessor's write cannot outrank its successor's under LWW, which is reachable on a clean handoff
via a futurecontext.timestamp. Details in
harper#483 and §10 of the design note.That directly supersedes this issue's "Leases + fencing by lock version" bullet, and the
"stale (lease-expired) holder's write is fenced" acceptance criterion with it. There is no fencing
token in conflict resolution: a lease-expired holder is fenced at admission and at commit
submission — Phase 0 already checks the lease synchronously on every staged write and again
immediately before the native commit submits — but a write that escaped before expiry resolves by
last-write-wins like any other. Fencing by generation was considered and deferred to
harper#2540; it changes_writeUpdate's
resolution rule for every write in the database, which is a cost deployments that never call
lock()would pay.The harper-pro work is now two tracked pieces:
- Record locks: measure the Phase 1 cost baseline before the protocol change #824 — the measurement gate. Five numbers on the 3-node harness Cluster record locks: operator-agreed home map transport and successor-freshness barriers (harper-pro#825, harper#2542 inside #822) #822 already has. It is the
only part of this workstream unblocked today, and the design note is explicit that everything in it
is a message count rather than a benchmark. - Record locks Phase 1: durable membership epoch protocol (single-decree agreement per database) #825 — the durable membership epoch protocol. Single-decree agreement per database over a
majority of current members, with acceptor state persisted before it is acknowledged. This is the
only place consensus appears in the whole design, and it is also what finally lets harper-pro assert
agreedDown.
Core carries harper#2541 (home ring, delegations,
drain/recall, caps) and harper#2542 (successor
freshness).🤖 Claude Opus 5
- Record locks: measure the Phase 1 cost baseline before the protocol change #824 — the measurement gate. Five numbers on the 3-node harness Cluster record locks: operator-agreed home map transport and successor-freshness barriers (harper-pro#825, harper#2542 inside #822) #822 already has. It is the
- added sub-issues
on Sep 18, 2026 - added a commit that references this issue
on Sep 18, 2026
Metadata
Metadata
Assignees
Labels
Type
Fields
Priority
Workstream W9 of #430 · exclusive distributed locking — replication integration
Summary
Implement
table.lock(primaryKey): Promise<WritableRecord>distributed exclusive locking. The core API and semantics are tracked in harper#483 (Table.tslock()is currently a stub:throw new Error('Not yet implemented'), line 1735). This workstream tracks the replication-protocol integration and the coordination with the core change. Locking is the explicit-serialization escape hatch from wall-clock LWW lost-updates (epic Theme G).Design direction (fits the real substrate)
LOCKEDmetadata bit — a free bit in theRecordEncoderuint32metadata bitmap (alongsideINVALIDATED/EVICTED/LOCAL_ONLY). Because it lives in the already-decoded metadata integer, the write/send paths can test "locked / by whom" without decoding the record value (the same throughput trickLOCAL_ONLYuses).LOCK/UNLOCKaudit types — a lock transition is a version-changing audit entry that replicates through the existing pipe for free (aftercommit→ subscription →SUBSCRIPTION_UPDATE). Peers applying it set their localLOCKEDbit.LOCKEDbit before LWW resolution and defers/rejects, converting silent lost-updates into explicit serialization. Unlock (a version change) lets queued writes proceed in version order.OPERATION_REQUEST/RESPONSE (136/137)RPC; durable lock state rides the audit stream.expiresAtmachinery; a write must present a lock version ≥ the record's current lock version, so a stale holder's write is rejected byprecedesExistingVersionregardless of clock skew (mitigates Theme G).Phasing
store.tryLock+ theLOCAL_ONLYmetadata bit. No replication, no protocol change. Validates the metadata/version-change wiring and theWritableRecordAPI surface.LOCK/UNLOCK; if another node holds it, request viaOPERATION_REQUESTfrom that one node; write-gating on theLOCKEDbit; leases + fencing. Covers the common "one node owns the hot key" case with point-to-point RPC.Promise.all. Deadlock avoidance via global key ordering + Wound-Wait by lock timestamp. Highest-risk piece.Hardest correctness problems
UNLOCK(itself a version change) to unblock waiters.lock(); transaction-owned auto-unlock at commit, node-owned persist.Dependencies
Core metadata/audit wiring (harper#483). Independent of W1/W4 — Phase 0/1 can start in parallel with Phase 2 of the epic once the metadata bit + audit type land.
Effort / risk
Phase 0 M / low · Phase 1 L / medium · Phase 2 L / high.
Acceptance criteria
lock()serializes concurrent multi-node writes to a key with no LWW lost update.🤖 Filed by Claude on behalf of Kris.