Skip to content

Record locks Phase 1: successor freshness — inherited dependency sets and the recovery barrier #2542

Description

@kriszyp

Part of the Phase 1 record-lock redesign, and the piece the 2026-09-09 guarantee decision sizes.
Design note:
docs/record-lock-ownership.md
§7.1, §7.2. Sibling of #2541 (the arbitration rule itself).

The property being replaced

The Ricart–Agrawala implementation on #2498 gets successor freshness for free, and it is worth
naming why before deleting it: a grant is a control entry on the grantor's own transaction log and
rides the same per-origin replication stream as the grantor's record writes, so applying a peer's
grant implies having already applied that peer's earlier data writes to the key. Unicast delegation
messages lose that. Every off-stream design loses it — a Raft coordinator would too.

So the replacement has to re-establish it explicitly, and this issue is that work.

Clean handoff: an inherited dependency set

After draining, the delegate writes a LOCK_RELEASE control entry to the table's transaction log —
the same non-LOCAL_ONLY entry the branch already builds — carrying a dependency set
{(originNodeName → position)}, one entry per origin that has written the key since the last
quiesce, bounded by |members|. The entry is ordered behind the delegate's own data writes on its
own stream.

The fence: admit only when, for every (origin → position) in the set, this node has applied
and made visible
that origin's stream to at least that position. Applied and visible, not
received or queued.

A scalar record version is not an applied-history fence, and neither is a single predecessor
position. Three concrete reasons, all of which the set handles:

  • A writes K; B acquires, never writes, releases; C has applied B but not A. A
    B-position check passes and C reads stale — so the set is inherited: a delegate that did
    not write the key passes it through unchanged, one that did merges its own (origin, position) in.
  • Core breaks equal-version conflicts by node name, so a replica can hold a losing value at the
    same timestamp as the winner and still pass a version ≥ V test.
  • Receiving B's later patch does not prove receipt of A's earlier change to other fields — the
    same reason the set cannot be collapsed to its newest member.

Deletes and tombstones are writes and carry positions like any other. An origin whose stream has
been purged past the listed position, or is otherwise unsatisfiable, is a recovery case and must
fail closed.

Recovery: the barrier, and where the guarantee is explicitly weaker

The home holds the dependency set in memory, so a home restart, an epoch change, or a delegate crash
with no clean release loses it. First: check whether a durable copy is recoverable from the
transaction log within retention — if it is, that is a cheaper recovery than the barrier and should
be preferred.

Without it, the acquirer drains its inbound replication streams from every reachable member to the
position each held at grant time before admitting. That is the strongest condition available without
synchronous replication, and it is weaker in two ways that must be documented rather than
implied
: an unreachable member's committed writes may not be visible, and a predecessor's native
commit submitted before expiry can settle after the barrier was measured.

Fail-closed is a fallback, not a fix. A barrier that succeeds still admits the forbidden history —
A commits K=1, becomes unreachable before replication, every reachable member's position is
satisfied, B is admitted and reads K=0. Nothing failed, so rejecting on failure does not detect
it. Closing that gap is #2540 and is deliberately out of scope here.

Cost, and the row that can dominate

The recovery path is not amortized: a scan touching many distinct cold keys after a recovery pays
a barrier per key. Barriers must be shared or batched across keys without weakening their ordering,
and the cost measured rather than assumed (HarperFast/harper-pro#824).

Retention is a separate lifetime from delegation expiry

A home that has evicted a key's dependency set cannot distinguish it from a key never delegated at
all, and would take the barrier on both — so a workload cycling through more keys than the cap pays
the recovery cost during normal operation, not only after a restart. The home therefore keeps a
compact ever-delegated-in-this-epoch filter (add-only, so no false negatives): outside it, a key has
no predecessor and needs no barrier; inside it with the set evicted, the barrier is required. Retain
dependency sets longer than delegations, and size the filter as part of the cap budget.

Wire format

The LOCK_RELEASE payload is today a fixed three-field tuple validated on exact length, so growing
it is a compatibility break — it gains a leading version field plus the dependency set. A historical
entry replayed from the log must still decode safely after its producer is gone, and a delayed old
release must not clear a newer delegation.

Refs #483

🤖 Filed by Claude Opus 5 on behalf of Kris.

Activity

  1. added theissue type on Sep 9, 2026
  2. self-assigned this
    on Sep 13, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

No labels
No labels

Fields

Priority

P2

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions