Part of the Phase 1 record-lock redesign, and the piece the 2026-09-09 guarantee decision sizes.
Design note:
docs/record-lock-ownership.md
§7.1, §7.2. Sibling of #2541 (the arbitration rule itself).
The property being replaced
The Ricart–Agrawala implementation on #2498 gets successor freshness for free, and it is worth
naming why before deleting it: a grant is a control entry on the grantor's own transaction log and
rides the same per-origin replication stream as the grantor's record writes, so applying a peer's
grant implies having already applied that peer's earlier data writes to the key. Unicast delegation
messages lose that. Every off-stream design loses it — a Raft coordinator would too.
So the replacement has to re-establish it explicitly, and this issue is that work.
Clean handoff: an inherited dependency set
After draining, the delegate writes a LOCK_RELEASE control entry to the table's transaction log —
the same non-LOCAL_ONLY entry the branch already builds — carrying a dependency set
{(originNodeName → position)}, one entry per origin that has written the key since the last
quiesce, bounded by |members|. The entry is ordered behind the delegate's own data writes on its
own stream.
The fence: admit only when, for every (origin → position) in the set, this node has applied
and made visible that origin's stream to at least that position. Applied and visible, not
received or queued.
A scalar record version is not an applied-history fence, and neither is a single predecessor
position. Three concrete reasons, all of which the set handles:
A writes K; B acquires, never writes, releases; C has applied B but not A. A
B-position check passes and C reads stale — so the set is inherited: a delegate that did
not write the key passes it through unchanged, one that did merges its own (origin, position) in.
- Core breaks equal-
version conflicts by node name, so a replica can hold a losing value at the
same timestamp as the winner and still pass a version ≥ V test.
- Receiving
B's later patch does not prove receipt of A's earlier change to other fields — the
same reason the set cannot be collapsed to its newest member.
Deletes and tombstones are writes and carry positions like any other. An origin whose stream has
been purged past the listed position, or is otherwise unsatisfiable, is a recovery case and must
fail closed.
Recovery: the barrier, and where the guarantee is explicitly weaker
The home holds the dependency set in memory, so a home restart, an epoch change, or a delegate crash
with no clean release loses it. First: check whether a durable copy is recoverable from the
transaction log within retention — if it is, that is a cheaper recovery than the barrier and should
be preferred.
Without it, the acquirer drains its inbound replication streams from every reachable member to the
position each held at grant time before admitting. That is the strongest condition available without
synchronous replication, and it is weaker in two ways that must be documented rather than
implied: an unreachable member's committed writes may not be visible, and a predecessor's native
commit submitted before expiry can settle after the barrier was measured.
Fail-closed is a fallback, not a fix. A barrier that succeeds still admits the forbidden history —
A commits K=1, becomes unreachable before replication, every reachable member's position is
satisfied, B is admitted and reads K=0. Nothing failed, so rejecting on failure does not detect
it. Closing that gap is #2540 and is deliberately out of scope here.
Cost, and the row that can dominate
The recovery path is not amortized: a scan touching many distinct cold keys after a recovery pays
a barrier per key. Barriers must be shared or batched across keys without weakening their ordering,
and the cost measured rather than assumed (HarperFast/harper-pro#824).
Retention is a separate lifetime from delegation expiry
A home that has evicted a key's dependency set cannot distinguish it from a key never delegated at
all, and would take the barrier on both — so a workload cycling through more keys than the cap pays
the recovery cost during normal operation, not only after a restart. The home therefore keeps a
compact ever-delegated-in-this-epoch filter (add-only, so no false negatives): outside it, a key has
no predecessor and needs no barrier; inside it with the set evicted, the barrier is required. Retain
dependency sets longer than delegations, and size the filter as part of the cap budget.
Wire format
The LOCK_RELEASE payload is today a fixed three-field tuple validated on exact length, so growing
it is a compatibility break — it gains a leading version field plus the dependency set. A historical
entry replayed from the log must still decode safely after its producer is gone, and a delayed old
release must not clear a newer delegation.
Refs #483
🤖 Filed by Claude Opus 5 on behalf of Kris.
Part of the Phase 1 record-lock redesign, and the piece the 2026-09-09 guarantee decision sizes.
Design note:
docs/record-lock-ownership.md§7.1, §7.2. Sibling of #2541 (the arbitration rule itself).
The property being replaced
The Ricart–Agrawala implementation on #2498 gets successor freshness for free, and it is worth
naming why before deleting it: a grant is a control entry on the grantor's own transaction log and
rides the same per-origin replication stream as the grantor's record writes, so applying a peer's
grant implies having already applied that peer's earlier data writes to the key. Unicast delegation
messages lose that. Every off-stream design loses it — a Raft coordinator would too.
So the replacement has to re-establish it explicitly, and this issue is that work.
Clean handoff: an inherited dependency set
After draining, the delegate writes a
LOCK_RELEASEcontrol entry to the table's transaction log —the same non-
LOCAL_ONLYentry the branch already builds — carrying a dependency set{(originNodeName → position)}, one entry per origin that has written the key since the lastquiesce, bounded by
|members|. The entry is ordered behind the delegate's own data writes on itsown stream.
A scalar record version is not an applied-history fence, and neither is a single predecessor
position. Three concrete reasons, all of which the set handles:
AwritesK;Bacquires, never writes, releases;Chas appliedBbut notA. AB-position check passes andCreads stale — so the set is inherited: a delegate that didnot write the key passes it through unchanged, one that did merges its own
(origin, position)in.versionconflicts by node name, so a replica can hold a losing value at thesame timestamp as the winner and still pass a
version ≥ Vtest.B's later patch does not prove receipt ofA's earlier change to other fields — thesame reason the set cannot be collapsed to its newest member.
Deletes and tombstones are writes and carry positions like any other. An origin whose stream has
been purged past the listed position, or is otherwise unsatisfiable, is a recovery case and must
fail closed.
Recovery: the barrier, and where the guarantee is explicitly weaker
The home holds the dependency set in memory, so a home restart, an epoch change, or a delegate crash
with no clean release loses it. First: check whether a durable copy is recoverable from the
transaction log within retention — if it is, that is a cheaper recovery than the barrier and should
be preferred.
Without it, the acquirer drains its inbound replication streams from every reachable member to the
position each held at grant time before admitting. That is the strongest condition available without
synchronous replication, and it is weaker in two ways that must be documented rather than
implied: an unreachable member's committed writes may not be visible, and a predecessor's native
commit submitted before expiry can settle after the barrier was measured.
Fail-closed is a fallback, not a fix. A barrier that succeeds still admits the forbidden history —
AcommitsK=1, becomes unreachable before replication, every reachable member's position issatisfied,
Bis admitted and readsK=0. Nothing failed, so rejecting on failure does not detectit. Closing that gap is #2540 and is deliberately out of scope here.
Cost, and the row that can dominate
The recovery path is not amortized: a scan touching many distinct cold keys after a recovery pays
a barrier per key. Barriers must be shared or batched across keys without weakening their ordering,
and the cost measured rather than assumed (HarperFast/harper-pro#824).
Retention is a separate lifetime from delegation expiry
A home that has evicted a key's dependency set cannot distinguish it from a key never delegated at
all, and would take the barrier on both — so a workload cycling through more keys than the cap pays
the recovery cost during normal operation, not only after a restart. The home therefore keeps a
compact ever-delegated-in-this-epoch filter (add-only, so no false negatives): outside it, a key has
no predecessor and needs no barrier; inside it with the set evicted, the barrier is required. Retain
dependency sets longer than delegations, and size the filter as part of the cap budget.
Wire format
The
LOCK_RELEASEpayload is today a fixed three-field tuple validated on exact length, so growingit is a compatibility break — it gains a leading version field plus the dependency set. A historical
entry replayed from the log must still decode safely after its producer is gone, and a delayed old
release must not clear a newer delegation.
Refs #483
🤖 Filed by Claude Opus 5 on behalf of Kris.