Skip to content

Record locks: measure the Phase 1 cost baseline before the protocol change #824

Description

@kriszyp

The design note for the Phase 1 record-lock redesign
(docs/record-lock-ownership.md
§10, §14) states this plainly: no number in it is a benchmark — they are message counts. The
protocol change should not land before the numbers exist, because the numbers are also the baseline
the new design has to beat.

#822 already carries an enablement gate and a
3-node integration harness, which is where this runs.

What to measure

Against the Ricart–Agrawala implementation currently on
HarperFast/harper#2498 + #822, with
replication.recordLocks enabled:

  1. Uncontended acquisition latency — distribution, not a mean, at 3 nodes. The theoretical cost
    is P+1 durable commits and up to P²−1 frame deliveries per lock; this says what that is in
    milliseconds.
  2. Repeat-lock latency on one key from one node. This is the before figure the delegation path
    has to collapse to a local key lock, so it is the single most load-bearing number here.
  3. Hot-key handoff throughput — alternating locks on one key from two nodes, which is the case
    both designs pay a full network round for.
  4. Transaction-log bytes and control entries per lock, including the audit entries. Audit growth
    per lock is currently unknown and it bounds how usable the feature is at rate.
  5. Write throughput with the feature off versus absent, to prove the ungated ordinary-write path
    is untouched. §8 of the note requires this and the branch currently violates it — the commit-time
    fence walks the whole write set on every commit even with no transport registered.

What it gates

The protocol change itself: #825 (epoch protocol) and
HarperFast/harper#2541 (home ring and delegations). It is independent of every open design
question, so it can start now.

Also worth capturing while the harness is up, since the redesign's steady-state cost is bounded
below by it: the delegation re-acquisition rate at candidate lease durations, and the
cold-key recovery path cost — the row of the note's cost table that can dominate and is not
amortized.

Refs #438, HarperFast/harper#483

🤖 Filed by Claude Opus 5 on behalf of Kris.

Activity

  1. added theissue type on Sep 9, 2026
  2. self-assigned this
    on Sep 13, 2026
  3. kriszyp commented on Sep 13, 2026

    @kriszyp
    MemberAuthor

    Measurement gate: results

    Ran the AFTER half against the Ricart–Agrawala before-figures in RECORD_LOCK_COST_BASELINE.md.
    Three full runs on the same box and bench, plus a fourth past the lock timeout. Full write-up and
    raw JSON are in the PR; the short version:

    §10 row BEFORE AFTER verdict
    First lock, key homed elsewhere — 1 RTT 1.07 ms p50 0.37 ms p50 as predicted
    First lock, key homed here — 0 RTT 1.07 ms p50 0.04 ms p50 as predicted
    Steady state — local key lock, 0 messages 0.68 ms p50 0.01–0.03 ms p50 as predicted
    Durable commits per uncontended acquisition 4 (~310 B) 0 as predicted
    Durable commits per contended section 4 0.023–0.054 release entries as predicted
    Hot-key handoff throughput 877–911 sections/s 1 031–1 707 sections/s faster, see below
    Hot-key exact convergence exact, every run 0.05–0.15 % lost at 3 contenders not met
    Hot-key fairness shared by turn-taking loser starved in 2 of 3 runs found here
    Write throughput, off vs absent inside noise inside noise as predicted
    Delegation re-acquisition rate n/a gap / (360 s − lease) exact at the 60 s window
    Cold-key recovery path n/a unmeasurable harper#2542 not implemented

    The amortization rows land where the design says they should — a repeat lock is a local key lock at
    0.01–0.03 ms with zero control entries, against 0.68 ms and four durable commits per acquisition.

    Two results do not match what this issue assumed, and both are about the handoff.

    The counter no longer converges exactly

    Every section's written value is audited. The shortfall is entirely duplicate values written by two
    different nodes
    , with no holes below the maximum and no failed requests — a successor reading a
    value its predecessor committed and had not yet replicated.

    That is disclosed, not new: resources/recordLockCoordinator.ts:43 states the §7 successor-freshness
    fence is deliberately unimplemented and is harper#2542, and §14 adds that §6 step 3 settlement is
    unimplemented too. The consequence worth acting on is that harper#2542's blast radius is wider than
    this issue assumed
    — it does not only make the cold-key recovery row unmeasurable, it costs
    correctness on the ordinary contended path, at 0.05–0.15 % of sections.

    Because of this the bench now records convergence rather than asserting it; an assertion would fail
    every run while testing a guarantee this phase does not offer.

    A contended key can be monopolized, and the loser starved to a 423

    At two contenders the losing node completed no section at all while the key was held, in two runs
    of three — its one recorded section carries the cluster's maximum written value, so it landed after
    the winner's loop ended rather than during the round. The winner meanwhile wrote ~597 lockRelease
    entries, so the recall machinery ran six hundred times and moved the key to the asking node once.
    Taken past DEFAULT_LOCK_TIMEOUT_MS (30 s), the starved contender's first request fails with
    423 Record is locked and was not released in time while the holder is still running: zero sections
    and one user-visible failure over forty seconds. The third run of the identical workload split 58/42,
    so the outcome is bistable rather than consistent.

    This is not a §10 row and not a stated non-goal — it is a fairness/liveness property the gate found,
    and it is a decision for harper-pro#825's enablement rather than a patch.

    Also

    • Re-acquisition rate follows gap / (DELEGATION_LEASE_MS − lease); at a 60 s window the measured
      rate is 0.0833 against a predicted 0.0833 in all three runs. At a 120 s window there is one
      extra lapse per run that the lease does not explain and that is worth understanding before any lease
      is tuned on these numbers.
    • Delegation retention, for harper#2581 sizing: exactly 121 delegations per node after 120
      distinct keys — one per distinct key touched per lease window, not per lock in flight.
    • One product defect fell out of the measurement: createDisabledRecordLockTransport
      (replication/recordLockTransport.ts:248) raises a plain Error, so a gated-off cluster lock()
      answers 500 rather than the retryable 503 its own docstring and replication/DESIGN.md both
      promise. Not fixed here.

    Results document and raw JSON for all four runs are in
    Record locks: the delegation cost measurement, and the two handoff gaps it exposes
    (draft, based on #822's branch). Measured on #822's head with core re-pinned to the current
    harper#2498 head — #822's own pin was left unreachable by a force-push of that branch and was four
    commits behind, two of which change grant and delegation holding.

    🤖 Generated with Claude Code

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

No labels
No labels

Type

Fields

Priority

P2

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions