Skip to content

Record locks: drain by recall so a membership change is not a ~6 minute lock outage - #861

Merged
kriszyp merged 7 commits into
feat/record-lock-cluster-transportfrom
kris/856-drain-by-recall
Sep 17, 2026
Merged

kriszyp merged 7 commits into
feat/record-lock-cluster-transportfrom
kris/856-drain-by-recall

Conversation

@kriszyp

@kriszyp kriszyp commented Sep 16, 2026

Copy link
Copy Markdown
Member

A record-lock membership change no longer has to wait out the delegation lease before the next generation can activate. record_lock_stage_generation now drains and reports what, if anything, is still live.

Closes #856. Core half: HarperFast/harper#2663 (kris/856-quiesce-delegations) — merge that first; this PR pins core at it. Branched from feat/record-lock-cluster-transport (#822), so it merges after that.

Why

Staging retracts active in the same durable write that records the staged generation — that is what makes a successful stage be quiescence evidence. But homeMap() needs an active generation, so from that write until record_lock_activate_generation every cluster-scoped lock() on that database answers 503, and §4.3 has the operator wait DELEGATION_LEASE_MS + LOCK_LEASE_SKEW_MS (365 s) in between. Adding one node is a ~6 minute cluster-lock outage for that database, and rendezvous hashing re-homes roughly 1/(n+1) of the keyspace onto a new node, so additions are as exposed as removals.

What changes

stageGeneration calls core's quiesceDelegations after its durable write — the write stops new grants, this drains the existing ones — and returns the result as quiesced. An orchestrator that sees a proven-clean drain on every node in homes(g) ∪ homes(g+1) can activate immediately; anything else falls back to the interval, or record_lock_fence_external, for that node only.

Three things it deliberately does:

  • Drains on the coordinating thread. The operations API answers on whichever thread took the request, and coordinator state is per-thread — calling core directly would sweep an empty registry and report "nothing outstanding" while the owner thread held every grant. It relays through the existing record-lock relay, and a relay that does not arrive is an error, never a clean drain.
  • Runs outside withRow. The drain waits on live critical sections; holding the database's transition queue for that would block every other stage, fence and activate behind it. The row write has already retracted active, so nothing new can be granted while it runs.
  • Never fails the stage. A stage that durably landed must not report failure because the drain did, or the operator retries a transition that already happened. A drain that cannot run is reported as { error }, which reads to an orchestrator exactly like outstanding work.

provesQuiescence() is the single predicate an orchestrator may skip the interval on: reached the coordinating thread, and complete, and an empty outstanding. Every field is checked positively, because the value crosses a worker boundary.

Verification

unitTests/replication/recordLockHomes.test.mjs — 96 record-lock unit tests passing, new ones driving stageGeneration itself: the coordinating thread's result passes through unchanged; a throwing drain becomes { error } without failing a stage that landed; an unwired drain refuses rather than reading as empty; and the proof predicate rejects a malformed reply. Core: 110 passing (HarperFast/harper#2663).

Cross-model pre-push review: six rounds. What they found, in order: the drain ran on the wrong thread; an empty sweep was treated as a proof; ownership continuity, not uptime, is the fact that proof rests on; the first-incarnation waiver cannot stand in for it; retired coordinators' grants outlive their coordinator; and a parameter name plus a non-array guard.

For the human reviewer

🤖 Generated with Claude Code

kriszyp and others added 7 commits September 16, 2026 12:41
stage already retracts active in its durable write, which stops NEW
grants; authority already outstanding still admits until its lease runs
out, and that is the whole reason §4.3 waits ~6 minutes before activate.
Drain it instead: stage now calls core's quiesceDelegations and returns
the result, so an orchestrator that sees empty outstanding on every node
in homes(g) union homes(g+1) can activate immediately, and falls back to
the timer only for a node that reports something left.

The drain never fails the stage — a stage that durably landed must not
report failure because the drain did, or the operator retries a
transition that already happened.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SaBbXbyR851pvCiaj6xbaL
Review findings on the first cut:

The drain ran on whichever thread answered the operation, while
coordinator state is per-thread — so it swept an empty registry and
reported "nothing outstanding" while the owner thread still held every
grant. It now relays to the owner through the existing record-lock relay
and treats a relay that does not arrive as an error, never as a clean
drain.

It also ran inside withRow, holding the database's transition queue for
the whole drain; it now runs after the row write, which has already
retracted active, so nothing new can be granted while it works and a
concurrent transition is not blocked behind it.

Tests now drive stageGeneration itself: the coordinating thread's result
passes through unchanged, a throwing drain becomes {error} without
failing a stage that already landed, and an unwired drain refuses rather
than reading as an empty outstanding list.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SaBbXbyR851pvCiaj6xbaL
provesQuiescence() is the single place that says what counts: a drain
that reached the coordinating thread, completed, and found nothing. An
empty outstanding list on its own does not, because core cannot sweep a
coordinator that was never built.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SaBbXbyR851pvCiaj6xbaL
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SaBbXbyR851pvCiaj6xbaL
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SaBbXbyR851pvCiaj6xbaL
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SaBbXbyR851pvCiaj6xbaL
Review finding: provesQuiescence accepted { complete: true,
outstanding: {} } because `outstanding?.length ?? 0` is 0 for a
non-array. The value crosses a worker boundary, so a malformed reply has
to read as "not proven" rather than slipping through on a missing
length. Every field is now checked positively, with tests.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SaBbXbyR851pvCiaj6xbaL
@kriszyp kriszyp added this to the v5.3 milestone Sep 16, 2026
@kriszyp
kriszyp marked this pull request as ready for review September 17, 2026 04:03
@kriszyp
kriszyp requested a review from a team as a code owner September 17, 2026 04:03
@kriszyp
kriszyp merged commit a6cb5a4 into feat/record-lock-cluster-transport Sep 17, 2026
34 of 35 checks passed
@kriszyp
kriszyp deleted the kris/856-drain-by-recall branch September 17, 2026 04:03
kriszyp added a commit that referenced this pull request Sep 17, 2026
#861 pinned core at kris/856-quiesce-delegations' head, which the squash
merge left off main — so deleting that branch would make this PR's
submodule pointer unreachable, which has already happened once on this
branch. Point it at main, which carries the same work as 72d65fa42.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SaBbXbyR851pvCiaj6xbaL
* which is the one thing this must never do. A relay that cannot reach the owner reports that as an
* error rather than an empty result, so the operator falls back to the drain interval.
*/
export async function quiesceOnOwner(database: string, deadlineMs: number): Promise<any> {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Missing test: quiesceOnOwner's cross-thread relay path — the actual "drain by recall" mechanism this PR adds — has no coverage at all.

quiesceOnOwner is the new public entry point this PR wires in as the drain reader (recordLockTransport.ts: setHomesDrainReader((database, deadlineMs) => quiesceOnOwner(database, deadlineMs));). It has two branches that matter: answering locally when ownership.ownsDatabase(database), and relaying to the owner worker via the new 'quiesce' RelayKind otherwise (with fail-closed handling when the relay times out, there's no owner, or the reply isn't a proper { outstanding: [...] } shape).

The only tests added in this PR (unitTests/replication/recordLockHomes.test.mjs) exercise stageGeneration/drainForStage/provesQuiescence entirely through setHomesDrainReader(mock), which replaces quiesceOnOwner wholesale — none of them call quiesceOnOwner itself. unitTests/replication/recordLockRpc.test.mjs doesn't exist, and integrationTests/cluster/recordLockCluster.test.mjs (unchanged by this PR) has no reference to quiesce/outstanding/provesQuiescence anywhere.

So the code that this PR's own title is about — reaching the actual coordinating worker for a database during a membership change, across the new 'quiesce' relay hop — is exercised by nothing: not the happy path (relay reaches the owner and gets a real QuiesceResult), not "no owner yet," not a relay timeout, not a malformed reply. Given the sibling function provesQuiescence needed a follow-up fix in this same PR for a shape it didn't handle correctly, this relay path (relay()/executeLocally()/routeFromMain() for kind: 'quiesce') is exactly the kind of new branch that benefits from a direct test — e.g. a unit test with two fake worker "threads" (mirroring the existing recordLockTransport.test.mjs style) verifying quiesceOnOwner returns the owner's real result when relayed, and throws the fail-closed 503 when the relay times out or no owner is registered.

@claude

claude Bot commented Sep 17, 2026

Copy link
Copy Markdown
Contributor

Reviewed the actual diff for this PR (cb2340d..de3f39a: recordLockHomes.ts, recordLockRpc.ts, recordLockTransport.ts wiring, and the core submodule bump) — the drain-by-recall mechanism itself is well-built and the PR's own last commit already fixed the one correctness gap a reviewer would look for (provesQuiescence checking every field positively). One blocker: quiesceOnOwner's cross-thread relay path, the actual mechanism this PR adds, has no test coverage anywhere (unit or integration) — see inline comment on replication/recordLockRpc.ts.

kriszyp added a commit that referenced this pull request Sep 18, 2026
…r-freshness barriers (harper-pro#825, harper#2542 inside #822) (#822)

* Pin core to the record-lock phase-1 branch head (harper#2498)

Temporary: re-bump to the merge commit once harper#2498 lands on main.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DXhaW9MrWxeZr5HXysyFBM

* Cluster-wide record locks: the harper-pro transport for core's LockCoordinator (harper-pro#438, W9 Phase 1)

Core (harper#2498) owns the Ricart-Agrawala protocol, writes its control entries to the
table's own transaction log and applies received ones from its replicated-event sink, in
order with the data of the batch. This adds what core cannot know:

- recordLocks capability level in the protocol registry, advertised only while
  replication.recordLocks is on; the send path skips lock control entries to a peer that
  has not advertised it (the registry's first gated frame).
- The participant set: every member of the database's replication group (explicit
  subscriptions included, direction ignored), each with the capability its own NODE_NAME
  bag asserted, kept in slot 13 of the per-(database, peer) shared status buffer; never
  learned reads as not capable.
- Per-database coordination ownership conferred by the main thread, moved only after the
  owner worker exits, with every subscription for the database placed on that worker while
  the feature is on so the coordinator applies the database's inbound entries.
- replication.recordLocks (default off): placement unchanged when off, and a cluster-scoped
  lock() on a replicated database fails closed naming the switch.
- cluster_status.recordLocks per database, with a correlated request/response so overlapping
  status calls cannot strand each other.

Depends on harper#2498 (core pinned to its head) and harper-pro#813 (merged into this branch).

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Vcn5gxtbSvWXLGk4ZWFRNf

* Read replication.recordLocks from the config tree, and fix two test-helper body reads

env.get resolves only keys registered in core's CONFIG_PARAM_MAP, so the harper-pro-only
replication.recordLocks switch read as undefined and the feature never armed; read it from
getConfigObj() instead (no core change). The cluster test's counter/controlEntries helpers
consumed the response body in an assertion message and then again via json(), and a received
control entry is applied to the coordinator rather than persisted in the receiver's log, so
the bag-less-peer test now asserts only the grant this node wrote.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Vcn5gxtbSvWXLGk4ZWFRNf

* Address the pre-push review: test lifecycle, config warning, slot doc

- recordLockCluster.test.mjs: start nodes with allSettled so one failed start does not
  orphan the nodes that came up; thread the wait deadline's abort signal into the mesh probe.
- recordLockConfig.ts: warn when replication.recordLocks is a truthy non-boolean (a YAML 1,
  a quoted "true"), which leaves the node fail-closed, so an operator is not left believing
  the switch is on.
- knownNodes.ts: record that shared-status slot 13 now holds the record-lock capability so a
  future slot taker does not overwrite it.
- Drop one reviewer-addressing comment.

The blob-gap/analytics/schema-merge findings the review surfaced are on code inherited
through this branch's base (harper-pro#432), not this change.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Vcn5gxtbSvWXLGk4ZWFRNf

* Install the connection-down reader in start(), not at module load

knownNodes -> replicator -> recordLockTransport is an import cycle; assigning
recordLockTransport's downSinceReader while that module is mid-evaluation hit its temporal
dead zone under the unit-test import order (adding the config module's imports shifted
evaluation order enough to expose it). start() runs after every module has loaded.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Vcn5gxtbSvWXLGk4ZWFRNf

* Record-lock cost baseline: the numbers the enablement gate calls for

`npm run bench:record-locks` (integrationTests/cluster/recordLockCost.bench.mjs, not part of
test:integration:cluster) boots the same 3-node mesh as recordLockCluster.test.mjs and measures
uncontended and repeat-lock acquisition latency, hot-key handoff throughput with 2 and 3 contending
nodes, control entries and bytes per acquisition from each node's transaction log, and unlocked write
throughput with the feature off, unregistered, and on. Timing is in-process (fixture-record-lock-bench).

replication/RECORD_LOCK_COST_BASELINE.md records one run with its distributions, sample counts,
machine class, and which figures are noisy. No change to the lock protocol; core is not moved.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013sPwr2JJbAcbHxnhoNr5qQ

* Address the pre-push review: node-side section timing, pooled hot-key distributions, release boundary

The hot-key section time was the client's wall clock around fetch; LockedIncrement now reports the
node's own lock and lock-through-save times, and the bench pools them across contenders beside the
per-node distributions and the client round trip. Log-cost ratios divide by rounds started, so a
timed-out round's request and withdraw cannot inflate them; the after-snapshot waits until every
started round's release is in its node's log; the convergence probe threads the wait's signal.
Baseline re-recorded from a run on the updated harness.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013sPwr2JJbAcbHxnhoNr5qQ

* Baseline: the request-minus-section gap is the commit plus request overhead, not HTTP alone

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013sPwr2JJbAcbHxnhoNr5qQ

* Record locks: delegation transport for the amortized-ownership protocol (static epoch)

harper#2498 replaced Ricart-Agrawala with amortized per-record ownership, and its
ClusterLockTransport contract changed with it: core now needs epoch(), requestDelegation()
and recallDelegation(), and registerClusterLockTransport throws on a transport without
them. Against that core this branch's transport did not register at all. This is the
harper-pro half, pinned to core 1deac506d.

epoch() is STATIC in this tranche - number 1, never advanced, not agreed. Members are the
database's replication group filtered to peers that advertised the delegation level of
recordLocks, plus this node, sorted; ringVersion hashes the sorted list so two nodes with
the same set agree without a deep compare. That is the design note's section 9 "static
owner" step: enough for one arbiter per key, not enough for section 4. What a static epoch
cannot do is advance across a restart to invalidate a previous incarnation's delegations,
so core's obligation on ClusterLockTransport.epoch is met the blunt way: epoch() returns
undefined for DELEGATION_LEASE_MS + LOCK_LEASE_SKEW_MS after process start, which blocks
every cluster lock on this node for that window. That is the cost of a static epoch, and
harper-pro#825 removes it. HARPER_TEST_RECORD_LOCK_RESTART_HOLD_MS lifts it for tests.

homeIncarnation is durable and monotonic - core orders fencing tokens on it, and a random
value is identifiable but not orderable. The main thread bumps recordLockIncarnation on
this node's own hdb_nodes row once per process start (merged via ensureNode); workers read
the mirror, and epoch() withholds while it still reads 0.

Request and recall are two registered operations, record_lock_delegate and
record_lock_recall (recordLockRpc.ts), sent over this worker's live outbound subscription
session to the home when it has one - its inbound end is on the home's coordinating
worker, so the request lands where the coordinator lives - and over sendOperationToNode
otherwise. An operation that arrives on a non-owner thread is relayed through main, which
mints its own hop id (worker-minted ids collide across workers), under a 5 s bound; a
timed-out relay answers not-home, never a grant. The requester is the authenticated node
principal of the connection, never the payload; a caller that is not a known node gets
403, so a super_user cannot mint or clear a delegation through the operations API.

The recordLocks capability is now level 2 and mutually exclusive: peerSupportsRecordLocks
requires the level exactly. Level 1 was Ricart-Agrawala and never shipped enabled; a peer
still advertising it is a different arbiter, not a slower one.

cluster_status.recordLocks reports { delegations, granted, admitted, droppedOffOwner,
members } per database; members is the epoch as the owner sees it, or absent while it is
withheld.

Sync-Core cost carried by the pointer bump, stated so it is not mistaken for a lock
change: four AuditRecord.localTime reads in replicationConnection.ts follow core's rename
to txnLogKey. The branch already pinned @harperfast/rocksdb-js 2.8.0, which core now
hard-requires at load (RecordEncoder throws below it); a checkout installed before that
pin has to reinstall before any core import loads.

The crash-recovery integration case is skipped with its reason: a crashed delegate holds
its keys for up to DELEGATION_LEASE_MS (six minutes), which does not fit a test, and
whether that lease is configurable is an open question on harper#2498. The property it
covered - a home never re-grants before the delegate's deadline plus skew, on independent
clocks - is asserted in core's coordinator suite.

Two defects the first cluster run caught, both now covered by tests that fail without
the fix: the transport read its home incarnation from server.nodes, which excludes the
local node on every path, so epoch() was withheld for the life of the process
(readOwnIncarnation reads the own hdb_nodes row); and a node whose bag was suppressed
still built a ring including itself while every peer excluded it - two arbiters for one
key. epoch() now withholds unless the bag this node actually sends claims the level. The
home-side half of that guard (refuse a requester outside the member set) is filed on
harper#2541 rather than reopened in harper#2498 mid-review.

The first pre-push round (full coverage: codex, gemini, cursor-grok, domain) returned BLOCK.
Its design-level finding stands and is put to the human on the PR: with a static epoch and
locally derived membership, two nodes can hold different rings for one key during a
membership transition and each self-home it - two arbiters - and nothing short of the
agreed epoch (harper-pro#825) closes that. Its concrete findings are fixed here, each
with a test where one applies: principalNodeName trusted a payload-supplied `user.name`
as a fallback (now hdb_user only); the main-thread rpc handler was unguarded; resolveLevel
min-clamped recordLocks so a future level-3 peer resolved to 2 and passed the equality
gate (now an exact, unclamped level); epoch() rebuilt the ring on every acquisition (now
memoized for 250 ms on the injected clock); a node that never joined a mesh had no self
row so the incarnation bump spun forever and every cluster lock 503'd for the life of
the process (the counter now goes on a LOCAL_ONLY self row); the bench omitted the
restart-hold override; executeRecall reported success after a timed-out relay (now a
503); the status-view cache was keyed on `auditStore && peer`; and three DESIGN.md
statements described the previous protocol.

Round 2 was degraded (Codex and the domain leg timed out on this box), but Gemini's two
majors were real and are fixed: the recall acknowledgement compared object identity across
a postMessage structured clone, so every relayed recall would have 503'd (structural check
now); and a worker that read its home incarnation from the table before main's bump landed
would cache the previous process's value for the life of the process. Workers now never
read the table: main broadcasts the bumped value (record-lock-incarnation) the way it
confers ownership, a late-registering worker asks for it, and the worker-side setter never
moves backwards (unit test). The 5830 audit-key fallback chain also matches its sibling.

Verification: 80 unit tests (recordLockTransport + protocolCapabilities); the 3-node
cluster integration suite 7 passing, 1 skipped as above - delegation request over the
live subscription session, recall handover, 24 concurrent increments landing exactly 24
on every node, ten repeat locks writing no release entries, the LWW/409 fence, and the
bag-less peer excluded from the ring and failing its own cluster lock closed with 503.
Typecheck: 31 errors, all pre-existing environment drift (harper-pro main has 32); none
in the changed files.

Refs #438, #822, #824, #825, HarperFast/harper#483, HarperFast/harper#2498,
HarperFast/harper#2541, HarperFast/harper#2542

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01M6aY2ERiYM8294P2f3aoUq

* Fix two more stale slot-13 references and a schema-list omission in DESIGN.md

The epoch() paragraph still said the recordLocks capability is read from
slot 13 (moved to 29 during the main rebase); homeIncarnation was described
as something workers read off their own hdb_nodes row, when they only ever
adopt what main pushes over record-lock-incarnation. Also note that
recordLockIncarnation is written via ensureNode without being a declared
table attribute, so the schema list right below doesn't omit it by mistake.

Surfaced by the independent pre-push review (Cursor Grok + Harper domain
adjudication) on 3729b81f.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* Fail closed instead of crashing when recordLockConfig loads before boot

getConfigObj() throws when no boot properties file exists yet. Every
other getConfigObj() call site in the codebase defers the call into a
function body for exactly this reason; recordLockConfig.ts read it as a
module-scoped constant at import time, so a bare mocha process (no
harperdb boot, unlike CI's own server-driven tests) crashed the whole
unit-test run the moment anything imported replicator.ts. This branch's
Unit Tests workflow never ran on GitHub Actions before this rebase (the
PR was mergeable_state: dirty, so CI skipped it) and local runs on this
box succeed only because an inherited HDB_ROOT happens to point at a
real properties file (dispatch-session leakage), masking the crash.

Catch the throw and treat it the same as "not configured": fail closed,
matching this module's own stated default.

Verified with `env -u HDB_ROOT npm run test:unit` (1094 passing) to
reproduce a from-scratch environment with no boot properties file.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* Pin core to harper#2498 head (operator-agreed home map)

Companion PR HarperFast/harper#2498 is open at 71d32bf6, per the
dispatch's explicit companion-PR instruction. Not a rebase merge of
"both sides" of the gitlink -- the instruction is to take neither side
and set the pointer to this exact sha.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* Record locks: the delegation cost measurement, and the two handoff gaps it exposes (#837)

* Measure record-lock cost under the delegation protocol, and record the handoff gaps it exposes

The AFTER half of the §10 measurement gate (harper-pro#824), against the Ricart-Agrawala
baseline in RECORD_LOCK_COST_BASELINE.md. Three full runs plus a fourth past the lock
timeout; raw JSON under replication/record-lock-cost-runs.

Where §10's predictions hold, they hold clearly: a first lock splits into 0.51 ms with the
home elsewhere and 0.05 ms when this node homes the key, repeat locks collapse from 0.68 ms
to 0.01-0.02 ms, and control entries per uncontended acquisition go from 4 cluster-wide to
zero.

Two results do not fit the task's stated expectations, and both are about the handoff:

- The counter no longer converges exactly. Auditing every written value shows the shortfall
  is entirely duplicate values written by two different nodes, with no holes and no failed
  requests - a successor reading a predecessor's unreplicated commit. That is the disclosed
  position: recordLockCoordinator.ts:43 states the §7 freshness fence is unimplemented
  (harper#2542), and §14 adds that §6 step 3 settlement is too. 0.03-0.12% of sections at
  three contenders.
- A contended key is monopolized rather than shared. At two contenders the losing node
  completed one section in fifteen seconds in all three runs, and past the 30 s lock timeout
  it fails with 423.

The bench therefore records convergence instead of asserting it: an assertion here would fail
every run while testing a guarantee this phase deliberately does not offer, and the bench
measures rather than gates.

core moves to the current harper#2498 head. #822's pin was left unreachable by a force-push
of that branch and is four commits behind, two of which change grant and delegation holding.

Refs #824

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LWmasaL9vLBSSKfGnG81kj

* Correct the starvation finding: the starved contender got zero sections, not one

The pre-push review traced `lastN` in the committed runs and found the write-up had the
order backwards. The loser's single recorded section carries the cluster's MAXIMUM written
value, so it landed after the winner's loop ended and dropped the key - not during the
contended window. In the 40 s run its first request had already failed with 423 at
DEFAULT_LOCK_TIMEOUT_MS while the holder was still running.

So over the contended window the starved contender completed zero critical sections and one
user-visible failure. That is worse than what the document claimed, and the document now says
it with the evidence.

Two harness bugs found in the same round:

- waitForAgreedCounter returned the agreed value straight to waitForCondition, which discards
  a falsy probe, so a round where every lock() answered 423 would settle on 0, be discarded,
  time out after 90 s and record agreedCounter: undefined for a cluster that did agree.
- distribution([]) produced NaN percentiles that serialize as null.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LWmasaL9vLBSSKfGnG81kj

* Claim only what the audit measures, and guard two paths it could crash on

The written-value audit shows two nodes computing n+1 from the same n. That rules out a lost
commit, but it does not by itself separate a successor admitted before applying its
predecessor's write from two nodes admitted at once - both produce the same signature, and
telling them apart needs holder intervals this bench does not record. The document now says
so and attributes the reading to core's own statement that the freshness fence is
unimplemented, rather than to these numbers.

Also: measurement 6 asserts it has a remote-home reference instead of dereferencing an absent
one, and LockStats reports unavailable coordinator stats as such rather than as an empty
object.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LWmasaL9vLBSSKfGnG81kj

* Clarify delegation benchmark conclusions

Co-Authored-By: GPT-5 Codex <noreply@openai.com>

* Address the post-rebase review: fix a real audit bug, correct three factual overclaims

- writtenValueAudit: track every node that has written a value, not just the first,
  so a node repeating its own write after another node already wrote it is no longer
  misattributed as a cross-node duplicate. Verified against all eight committed hot-key
  rounds: every one already had repeatedCount === repeatedAcrossNodes (no value was ever
  written by the same node twice), so this does not change any reported number.
- Note the threshold/delta unit mismatch in measurement 6 (an absolute latency floor
  compared against a delta with the local lock already subtracted) and the ~500ms
  convergence-poll window's limits, without changing either's behavior blind (the
  currently-pinned core is incompatible with harper-pro's transport, so the cluster
  bench cannot actually be run right now to validate a behavior change - see below).
- RECORD_LOCK_COST_DELEGATIONS.md: record the core sha the numbers were actually
  measured at (729aefd2) and disclose that the base's own further core re-pin
  (71d32bf6) removed `epoch()` in favor of `homeMap()`, which harper-pro's transport
  does not yet implement - the committed numbers are not currently re-runnable.
  Narrow the 120s-window "real rounds (0.35ms+)" claim: runs 2 and 3's off-window
  lapses (0.196-0.239ms) are far closer to their own local-reference noise than run
  1's clean case. Replace the "abandoned instrument" paragraph with a home-per-round
  table derived from lockStats snapshots already in the committed JSON (no new
  instrumentation) - it settles most of the "why does the holder win" question.
- DESIGN.md: the disabled-transport 503 claim is wrong (a plain Error, so 500) and the
  repeat-lock range was narrower than the document it now points to actually measured.
- README.md: note that run 4's committed JSON has only one of the two expected
  measurement-6 entries.

Cosmetic, from the same round: fix the quiet-poll count in a docblock (two vs three),
drop a no-op multiplier constant, trim narrated history from two docblocks, remove an
orphaned comment, and de-duplicate a restated comment in the fixture.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* Trim two new comments to Harper's zero-narration default, record the harper-pro sha

- Move the reacquisition threshold's absolute-vs-delta bias explanation into the results
  document's §6 (where the other measurement-6 methodology notes already live) instead of
  narrating it in code; same for the agreed-counter wait's convergence caveat.
- Record the harper-pro sha the runs were measured at (4cd0b9d8), not just core's; the
  prior commit only fixed the moving core reference.
- Drop the "dispatch task" tracker reference in the document - that addresses this
  session's tooling, not a reader of the PR.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* Fix the §6 undercount rationale: reclassification is offline, not a re-run

The prior commit said fixing the threshold would need the bench re-run against the
now-unstartable core pin. Wrong: every tick's deltaMs is already in the committed JSON,
so reclassifying is an offline check. Verified that check against run 1's 300s row -
the naive fix (floor minus local reference) pulls two known-local ticks across the cut
as false lapses (neither on a window multiple, both inside that row's own local-reference
range) - so the real obstacle is that the floor needs a better basis than a lower number,
not that it can't be checked without a cluster. Also trims a comment that narrated the
fix it sits next to.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Co-authored-by: GPT-5 Codex <noreply@openai.com>

* Design note: operator-agreed home map (harper-pro#825 inside #822)

Core's home-map interface (harper#2498, merged) replaced the static
epoch with an operator-agreed, digest-checked, immutable-per-generation
map. This note scopes the harper-pro-side implementation to what #825's
issue text still covers after the design's round-7 revision.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011BeTkLNGebWSxyc2hiF6xg

* Design note round 2: replace ack/activate with a per-node drain timer

Round 1 of the planning review found the acknowledge/canonical-artifact
mechanism unsafe (local-only evidence can't support a cross-node check,
among other blockers). This revision replaces it with a purely local
wall-clock timer anchored at stage-receipt, removing the need for
cross-node evidence collection entirely, and fixes the storage,
digest, handshake, and incarnation-ordering issues round 1 found.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011BeTkLNGebWSxyc2hiF6xg

* Design note round 3: operator-timed activation, immediate quiesce-on-stage

Round 2 of the planning review rejected the local-timer mechanism with
concrete two-holder counterexamples (staging didn't quiesce immediately;
per-node deadlines aren't a last-node barrier; Date.now() isn't a safe
elapsed-time proof across a restart). This revision adopts round 2's own
stated resolution: stage atomically retracts the active generation (real
quiescence, not an inferred one), and a separate operator-issued activate
call, timed externally by the operator's own wait from the last stage/
fence event, promotes it. Also fixes the digest-mismatch ring-shrink bug,
centralizes the incarnation-bump gate in recordLockOwnerFor for both call
sites, and moves the digest off NODE_NAME onto a dedicated message.

No third planning round: this converges on both rounds' own explicit
recommendations rather than introducing a new mechanism.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011BeTkLNGebWSxyc2hiF6xg

* Cluster record locks: implement harper-pro#825's operator-agreed home map

harper#2498 merged with a round-7 revision that deleted the durable
membership-epoch consensus protocol #825 was originally scoped to build,
replacing it with an operator-agreed, digest-checked, immutable-per-
generation home map. This branch's static-epoch transport could not even
compile against core's new homeMap() interface. Implements what remains
of #825 after that redefinition, per two rounds of planning review:

- replication/recordLockHomes.ts: a dedicated LOCAL_ONLY system table
  (hdb_record_lock_homes), three super_user-gated operations
  (record_lock_stage_generation, record_lock_fence_external,
  record_lock_activate_generation), and pure planStage/planActivate
  decision functions. Staging atomically retracts any active generation
  in the same durable write (real quiescence, not inferred); activation
  is a separate, operator-timed call — the operator's own external wait
  is the safety mechanism, not anything the code measures, after two
  planning rounds showed a node-timed alternative reopens the two-holder
  bug the design exists to prevent. A small backstop timer is defense in
  depth only.
- replication/recordLockTransport.ts: epoch() -> homeMap(), sourced from
  a frozen per-thread cache of the durable generation, gated on both
  peer digest agreement and protocol capability (independent checks — a
  matching digest from the wrong protocol level is not enough). A digest
  mismatch fails the whole map closed, not a shrunk ring (excluding only
  the disagreeing peer is itself a two-arbiter bug). homeIncarnation now
  advances per coordination incarnation via a central gate in
  recordLockOwnerFor, not only at process start. grantableAfterMono
  (core's restart-quarantine waiver) needs the transport rebuilt at three
  independent points -- first-incarnation known, ownership conferred, and
  the active generation newly available -- since core's coordinator is
  built lazily and reads that field once, at construction.
- replication/replicationConnection.ts: a dedicated RECORD_LOCK_HOMES_DIGEST
  wire message (not a NODE_NAME resend, which has untested side effects),
  and peer-digest reconciliation centralized by (database, peer) rather
  than per connection object -- the mesh keeps a separate connection per
  direction, each with independent local state.
- protocolCapabilities.ts/recordLockRpc.ts: capability bump 2->3 and the
  epoch->generation wire rename for the interface change.
- core re-pinned to origin/main post-merge (4c353f880) per the task
  owner's instruction.

Verification: full unit suite 1121/1121 passing throughout. Integration
(recordLockCluster.test.mjs, rewritten for the new bootstrap model) --
a full clean end-to-end run was not obtained; blocked repeatedly by two
confirmed pre-existing, external causes on the machine this ran on (this
session's background processes killed by the harness's own memory-
pressure policy, and integration-testing's shared loopback-address pool
file corrupted by a non-atomic write raced with a concurrent process --
neither touched by this change). What is confirmed: a partial run after
all fixes passed all 4 real tests in the hardest suite, including 24-way
concurrent contention across 3 nodes, before being killed moving into
the next suite. See RECORD_LOCK_HOMES_DESIGN.md's "For the human
reviewer" section for the full account.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011BeTkLNGebWSxyc2hiF6xg

* Address round-1 pre-push review: 3 blockers, sender gating, digest truncation

- recordLockConfig.ts (via recordLockTransport.ts startup warning): enabling
  replication.recordLocks now logs an operator-visible warning that a
  handoff carries exclusion but not freshness until harper#2542, with the
  measured lost-update rate -- previously disclosed only in design docs
  nobody reads at enablement time.
- recordLockHomes.ts: stage/activate no longer return before this node's
  local cache refresh AND every other thread's has been confirmed via a
  bounded cross-thread ack (recordLockTransport.ts) -- a durable write
  landing was not the same fact as every grant-capable thread having
  actually stopped serving the old generation.
- recordLockHomes.ts: stage/fence/activate now serialize behind a
  per-database in-process queue (withRow), closing a read-outside-
  transaction race where a concurrent stage and fenceExternal (or two
  concurrent stages) could each read the same prior state and the second
  write silently discard the first's outcome, both reporting success.
- recordLockHomes.ts: digestOf no longer truncates generation to 32 bits
  (generation >>> 0 made generation 1 and 2**32+1 hash identically over
  the same homes).
- replicationConnection.ts: RECORD_LOCK_HOMES_DIGEST is now sender-gated
  on the peer's advertised recordLocks capability, matching this file's
  own stated sender-side gating discipline; sent once handshake actually
  establishes that, not only at the point databaseName was first known.
- recordLockTransport.ts/recordLockRpc.ts: disabled-transport refusal is
  now a real 503 (ClientError, not a plain Error with no statusCode --
  independently found by the delegation-cost bench); guarded a previously
  unguarded worker-side postMessage reply.

Unit suite 1121/1121 passing throughout. Integration re-confirmation of
this specific commit was blocked by the same pre-existing, external
causes as before (this session's OOM-killed background processes;
integration-testing's shared loopback-address pool file corrupted by a
concurrent process) -- every attempt reproduced the identical error
signature, never a functional failure. The fixes are narrow and
logically self-contained (an in-process serialization queue, an
await/ack chain on an existing message pair, a sender-side gate check,
an encoding width fix); nothing here touches the mechanism the prior
partial integration run already exercised successfully (peer digest
reconciliation, the three grantableAfterMono recreate triggers).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011BeTkLNGebWSxyc2hiF6xg

* Fix round-2 pre-push blocker: ack timeout must fail a stage/activate, not fake quiescence

The homes-changed ack protocol added in the round-1 fix resolved success on a
missing ack (dead worker, full mailbox, or a live worker whose event loop was
too busy to have applied the change yet) — indistinguishable from the exact
failure the protocol exists to catch: a thread still granting under the
retracted generation. It also gave the origin-to-main relay the same budget
as main's own per-sibling fan-out, so main could never answer in time
whenever a sibling was slow.

Make the ack fail-closed (timeout and a false ack both reject, propagating a
503 up through stageGeneration/activateGeneration) and give the relay leg a
longer budget than the leaf fan-out it waits on. Also fix stageGeneration's
and activateGeneration's noop (idempotent-retry) path to still re-notify: a
retry after a failed relay was previously silent, leaving the gap the retry
exists to close unreconciled.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011BeTkLNGebWSxyc2hiF6xg

* Fix round-3 pre-push finding: settle the armed ack entry on a postMessage throw, don't abandon it

sendHomesChangedAndWaitAck and the origin-to-main relay both already register a
pending-ack timer via waitForHomesChangedAck before attempting postMessage. On a
synchronous postMessage throw, the previous code deleted that map entry and
manufactured a second, unrelated rejected promise for the caller — leaving the
original promise's own timeout still armed with nothing subscribed to it. That
promise rejects unhandled ~2-3.5s later, and Node's default policy terminates
the process.

Settle the SAME already-armed entry (ok=false) instead: its resolver clears the
timer and rejects the one promise the caller is actually awaiting.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011BeTkLNGebWSxyc2hiF6xg

* Fix round-4 finding: the concurrent-increment test asserted freshness this phase doesn't provide

recordLockCluster.test.mjs asserted seen === [1..24] and exact convergence to
N — a successor-freshness guarantee harper#2542 has not landed yet.
RECORD_LOCK_COST_DELEGATIONS.md measures 0.05-0.13% of sections losing an
update at 3 contenders for exactly this reason, and the bench already
dropped its own equivalent assertion citing it; the integration test kept
it, so it could flake on the same race the code openly documents as
outstanding.

Assert what this phase actually guarantees instead: every admitted
increment lands in range, and every node converges to the same final value
— not that the value is N. Corrected the module docstring's matching claim
("never loses an update").

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011BeTkLNGebWSxyc2hiF6xg

* Fix round-5 finding: require near-total distinctness so the test still catches exclusion failures

The prior fix (7a2ad19d) replaced the flaky exact-[1..N] assertion with a
pure range check, but a range check alone can't tell "exclusion works, one
rare freshness race" from "no exclusion at all" (24 requests all reading
the unwritten n=0 and writing 1 pass equally). Require the written values
to be almost all distinct (tolerating up to 2 collisions, well above the
documented 0.05-0.13% rate) so a genuine serialization regression still
fails the test.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011BeTkLNGebWSxyc2hiF6xg

* Format: wrap the distinctness assert.ok to satisfy prettier

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011BeTkLNGebWSxyc2hiF6xg

* Docs: fix two accuracy findings from pre-push review (capability level, test-coverage overclaim)

DESIGN.md said recordLocks advertises as level 2; protocolCapabilities.ts
has been at 3 since the homeMap() redesign. Both DESIGN.md's test index and
RECORD_LOCK_HOMES_DESIGN.md's Testing section claimed integration coverage
(digest mismatch, record_lock_fence_external, restart incarnation
ordering, mid-bump persistence failure, holder-crash lease hand-over) that
recordLockCluster.test.mjs does not contain — corrected both to describe
what the suite actually covers and what's still worth adding.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011BeTkLNGebWSxyc2hiF6xg

* Fix round-7 nit: crash-recovery skip cites harper#2498, not harper#2542

The test's own comment ties the six-minute lease/configurability question
to harper#2498; harper#2542 is the separate successor-freshness fence.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011BeTkLNGebWSxyc2hiF6xg

* Design note: the harper-pro successor-freshness barrier (harper#2542)

Bump core to harper#2613 and record the transport-side design for
ClusterLockTransport.establishLockFreshness() before implementing it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SaBbXbyR851pvCiaj6xbaL

* Design note round 2: adopt the planning findings, fail recovery closed

Clean-handoff barriers evaluate against a published apply-visible
watermark; the null recovery marker rejects with 503 until an
append-order source head exists (the open decision, with options).

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SaBbXbyR851pvCiaj6xbaL

* Design note round 3: origin-keyed atomic publication, event-driven waiters

Adopt round 2: publish apply-visible progress per authenticated origin
in one Atomics-addressed 64-bit slot, wake waiters from the publication
instead of polling, record exact capability levels, select the disabled
transport on LMDB, and recommend a marker write for the recovery fence.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SaBbXbyR851pvCiaj6xbaL

* Design note round 4: fences only from direct full-coverage streams

Adopt round 3: publish per-origin fences only from the origin's own
direct, full-database-coverage stream (the exclusion-origin set), in a
generation-keyed CAS-advanced atomic word bootstrapped from the durable
resume cursor; relayed and selectively-routed origins fail closed.
Recovery fence tracked as harper#2625.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SaBbXbyR851pvCiaj6xbaL

* Design note round 5: both halves against harper#2627's barrier entry

Pin core at harper#2627 (lockBarrier, writeLockBarrier, deadlineMs).
Adopt round 4: zero-exclusion publishers poisoned on any drop, no cursor
bootstrap, barrier probes for recovery and for a clean dependency the
fence has not reached, self-origin clone baseline, capped waiters.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SaBbXbyR851pvCiaj6xbaL

* Design note round 6: exact nonce barriers only, durable poison

Adopt round 5: drop every numeric fence (a restart reissues log keys),
prove each unsatisfied cross-origin dependency with a same-table
lockBarrier entry matched on (origin, position, nonce), poison an
(origin, table) durably on any drop or on core's terminal-apply-failure
hook (harper#2628), decide self-origin by incarnation start or an
ever-recloned flag, and bound/authorize/coalesce barrier requests.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SaBbXbyR851pvCiaj6xbaL

* Design note round 7: record the append-order probe, adopt round 6's rest

Round 6's headline blocker (replay reorders a log by key) is disproved
by a probe on rocksdb-js 2.9.0: per-log range reads are append-ordered.
The residual resume skip is harper#2629. Adopt everything else: poison
is permanent, every self-origin dependency rejects after a reclone,
client-only coalescing, relay instead of a shared ring, failed poison
writes hold the frame, checks before state, an executable cost gate.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SaBbXbyR851pvCiaj6xbaL

* Successor freshness: prove every cross-origin handoff with a lockBarrier

Implement ClusterLockTransport.establishLockFreshness() (harper#2613,
harper#2625/#2627): each inherited (origin, position) dependency and
each recovery marker is proven only by a lockBarrier entry the origin
commits on request, matched on (origin, position, nonce) once applied
here. Durable per-(origin, table) poison on every dropped record, an
ever-recloned flag written at clone start, a node-principal barrier
operation restricted to current members at the exact level, capability
level 4, the LMDB gate, and the exact-convergence cluster assertion.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SaBbXbyR851pvCiaj6xbaL

* Docs: successor freshness landed, slot map, cost rows

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SaBbXbyR851pvCiaj6xbaL

* Consume core's apply-failure listener when present; wildcard poison

harper#2628's registerReplicatedApplyFailureListener is looked up at
registration and, when present, poisons the failed origin durably; an
origin poisoned with table '*' fails every table. The design note now
states the residual (terminal apply failures stay unrecorded against a
core without the hook) instead of a refusal the code did not implement.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SaBbXbyR851pvCiaj6xbaL

* Cluster test: assert a barrier was applied and nothing was poisoned

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SaBbXbyR851pvCiaj6xbaL

* Fix pre-push review majors: poison durability, cross-thread poison, settle recheck

A failed poison write is retried on the next report instead of being
cached as recorded; a durable row is announced to every thread so the
coordinating thread's barrier refuses it; a matching barrier rechecks
poison at settle time; transport replacement settles the old barrier's
waits; release forgets the poison cache; the hole-recording closure is
one per connection rather than one per frame; JSON pair keys replace a
NUL separator.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SaBbXbyR851pvCiaj6xbaL

* Read poison and reclone state from the store, not a per-thread cache

The hole is recorded on the socket's thread while the barrier waits on
the coordinating thread, and an announcement without acknowledgement
left a window; the drop completes only after the row is durable, so
the cold-path checks read the dbis store directly and the broadcast is
gone. A pair whose row could not be written stays marked on its thread
and is retried on the next report. The LOCAL_ONLY defense drop poisons
like the other drop kinds.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SaBbXbyR851pvCiaj6xbaL

* Fail closed when poison state cannot be read; poison unknown-table drops

A database with no readable dbis store, or a read that throws, reports
as poisoned and recloned; the barrier answers a throwing check with 503
instead of letting it escape a settlement callback; a local-only drop
whose table is unknown poisons the whole origin.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SaBbXbyR851pvCiaj6xbaL

* Pin core at main (harper#2627 + #2630) and consume the apply-failure listener directly

registerReplicatedApplyFailureListener lives in
core/resources/replicatedApplyFailure.ts, so the feature-detected lookup
on Table.ts would never have found it; import it, register per database,
unregister on release. The one hole class the notes called unrecorded
is now recorded.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SaBbXbyR851pvCiaj6xbaL

* Design note: record_lock_bootstrap_generation

A local, fresh-state-only generation-1 write that derives homes from
hdb_nodes by default. The fresh-state guard (no active, no staged,
highestActedOn 0) is what makes the stage/drain/activate sequence
unnecessary: homeMap() has never answered on that node, so core has
never granted there.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SaBbXbyR851pvCiaj6xbaL

* Add record_lock_propose_homes, a read-only home-set proposal

Planning review rejected the mutating bootstrap this replaces, on two
counterexamples now recorded in the design note: homeMap() iterates its
OWN active.homes, so a node that derived [A] checks no peers and serves
immediately — a digest cannot detect a participant omitted from the set
being digested; and a node newly added to a cluster already active at
generation 2 has an untouched row, so any "never acted here" guard
passes and it would activate its own generation 1.

What lands instead writes nothing: it returns the canonical home set
this node's hdb_nodes view suggests, the generation one past its own
floor, the digest, the current state and warnings, so an operator can
capture one list instead of typing it and still stage and activate it
everywhere through the operations that already exist. planProposal is
the pure decision, unit-tested like planStage/planActivate; membership
readers are injected from recordLockTransport because a static import of
knownNodes here is an initialization cycle.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SaBbXbyR851pvCiaj6xbaL

* Return the §4.3 quiesce union, not just the new ring

Review finding, and a real defect in the first draft: the proposal's
warning said to stage "every node named in homes", which on a shrink
omits the node being removed — it is never staged, keeps its old active
generation, and keeps granting while the new ring grants too. The union
of the proposal with the current active and staged rings is now a
returned field, with a warning naming the leaving nodes.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SaBbXbyR851pvCiaj6xbaL

* Say what the proposal cannot know, and what a departing node becomes

Two review findings on the advisory response: quiesce is built from the
answering node's own rings, so a node with no active generation cannot
name the ring the cluster is serving — it now says so and tells the
operator to ask a node that holds it. And a departing node, once staged,
is never activated, so it stays unable to lock; that is what leaving the
ring means, and saying it keeps it from reading as a stuck transition.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SaBbXbyR851pvCiaj6xbaL

* Docs: the design note now states both limits the response warns about

Review finding: the note described the proposal's response before the
completeness caveat and the departing-node state were added, so the
canonical note contradicted the code.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SaBbXbyR851pvCiaj6xbaL

* Record locks: drain by recall so a membership change is not a ~6 minute lock outage (#861)

* Drain on stage, so a membership change need not wait out the lease

stage already retracts active in its durable write, which stops NEW
grants; authority already outstanding still admits until its lease runs
out, and that is the whole reason §4.3 waits ~6 minutes before activate.
Drain it instead: stage now calls core's quiesceDelegations and returns
the result, so an orchestrator that sees empty outstanding on every node
in homes(g) union homes(g+1) can activate immediately, and falls back to
the timer only for a node that reports something left.

The drain never fails the stage — a stage that durably landed must not
report failure because the drain did, or the operator retries a
transition that already happened.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SaBbXbyR851pvCiaj6xbaL

* Drain on the coordinating thread, outside the transition queue

Review findings on the first cut:

The drain ran on whichever thread answered the operation, while
coordinator state is per-thread — so it swept an empty registry and
reported "nothing outstanding" while the owner thread still held every
grant. It now relays to the owner through the existing record-lock relay
and treats a relay that does not arrive as an error, never as a clean
drain.

It also ran inside withRow, holding the database's transition queue for
the whole drain; it now runs after the row write, which has already
retracted active, so nothing new can be granted while it works and a
concurrent transition is not blocked behind it.

Tests now drive stageGeneration itself: the coordinating thread's result
passes through unchanged, a throwing drain becomes {error} without
failing a stage that already landed, and an unwired drain refuses rather
than reading as an empty outstanding list.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SaBbXbyR851pvCiaj6xbaL

* Expose the one predicate an orchestrator may skip the interval on

provesQuiescence() is the single place that says what counts: a drain
that reached the coordinating thread, completed, and found nothing. An
empty outstanding list on its own does not, because core cannot sweep a
coordinator that was never built.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SaBbXbyR851pvCiaj6xbaL

* Pin core at the ownership-based quiescence proof

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SaBbXbyR851pvCiaj6xbaL

* Pin core at the ownership-only quiescence proof

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SaBbXbyR851pvCiaj6xbaL

* Pin core at the retired-coordinator reporting fix

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SaBbXbyR851pvCiaj6xbaL

* Check every field of a drain reply positively

Review finding: provesQuiescence accepted { complete: true,
outstanding: {} } because `outstanding?.length ?? 0` is 0 for a
non-array. The value crosses a worker boundary, so a malformed reply has
to read as "not proven" rather than slipping through on a missing
length. Every field is now checked positively, with tests.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SaBbXbyR851pvCiaj6xbaL

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>

* Pin core at harper main now that #2663 has merged

#861 pinned core at kris/856-quiesce-delegations' head, which the squash
merge left off main — so deleting that branch would make this PR's
submodule pointer unreachable, which has already happened once on this
branch. Point it at main, which carries the same work as 72d65fa42.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SaBbXbyR851pvCiaj6xbaL

* Record locks: one operator call applies a home map across the cluster from an explicit node list (#863)

* Design note: record_lock_apply_homes, one call to apply a home map cluster-wide

Refs #862

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016EVzk62HUFq2v4RCtG13Lq
Dispatch-Task: harper-pro-862-apply-homes

* Add record_lock_apply_homes: one call applies a home map across the cluster

Survey every node in the operator's explicit list, refuse before staging on an
unreachable node, an unlisted ring member, a digest disagreement or a stage the
node would reject; stage everywhere over a node-principal hop that re-validates
locally; activate immediately when every node proves its drain, otherwise
report per node with a relative wait for an attested second call that covers
only what was already staged. The stage persists the quiesce set so a retry
cannot lose the old ring once active is retracted.

Refs #862

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016EVzk62HUFq2v4RCtG13Lq
Dispatch-Task: harper-pro-862-apply-homes

* Cluster test: judge an untouched node by its row, not by a lock a staged peer homes

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016EVzk62HUFq2v4RCtG13Lq
Dispatch-Task: harper-pro-862-apply-homes

* Cluster test: open generation 2 explicitly before injecting the failure

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016EVzk62HUFq2v4RCtG13Lq
Dispatch-Task: harper-pro-862-apply-homes

* Apply review round 1: required quiesce, initiator row, staged-race recovery, hop cancellation

- record_lock_stage_generation now requires quiesce and persists it; a matching
  re-stage backfills a row that lacks one, and the survey refuses a staged row
  with none, so a manually staged node cannot hide the ring it stopped serving.
- The node taking the call reads its own row too when it is not in quiesce; a
  ring it serves that the list omits is refused like a peer's.
- Only an active disagreement is fatal; two sets staged by racing operators are
  judged per node against the target, so an explicit higher generation proceeds.
- Every hop's deadline retires the request on the wire: the live session drops
  its pending entry and sendOperationToNode closes its socket.
- The per-database apply queue releases its entry when idle.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016EVzk62HUFq2v4RCtG13Lq
Dispatch-Task: harper-pro-862-apply-homes

* Apply review round 2: a stage must name every ring its row remembers; cancel the one-shot hop

planStage now refuses a quiesce that omits a member of the active ring, the
staged ring or the previous staged transition's participants, since the write
erases them from the row and a later survey could not see the node still
serving them. sendOperationToNode passes its timeout into the session so the
socket closes when a peer accepts and never answers. DESIGN.md index updated.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016EVzk62HUFq2v4RCtG13Lq
Dispatch-Task: harper-pro-862-apply-homes

* DESIGN.md: four operations, not three

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016EVzk62HUFq2v4RCtG13Lq
Dispatch-Task: harper-pro-862-apply-homes

* Clear the operation timeout timer when the response arrives

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016EVzk62HUFq2v4RCtG13Lq
Dispatch-Task: harper-pro-862-apply-homes

---------

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>

* Cluster record locks: serve lock() on every worker at threads.count > 1 (#865)

* Cluster record locks: serve lock() on every worker at threads.count > 1 (#852)

Cluster-scoped lock() only worked with one http worker: a request landing on
a worker that does not coordinate the database answered 503, and the row that
drives the operator-agreed home map was guarded per worker isolate. This makes
the feature usable at the default worker count.

- recordLockRpc.ts: relay a local lock() acquire/release to the coordinating
  worker over the worker-to-worker port mesh (main only broadcasts which
  thread owns each database). The admission crosses the boundary, not the
  handle; a recall on the owner fences the caller's handle over the mesh and
  the owner awaits that ack (or the handle's lease) before writing the
  release. Admissions are bound to an owner-session nonce and the
  harness-stamped origin thread, so a stale release after an ownership handoff
  cannot address another handle. Caller identity is the sender port, never a
  payload field; the messages are internal, never registered operations.
- recordLockHomes.ts: withRow now takes a process-wide node-scoped lock on the
  hdb_record_lock_homes row and writes through that locked handle, so a stage
  racing a fence_external across workers can no longer restore a retracted
  generation, and a lease lost to a storage stall fails the write.
- recordLockTransport.ts: wire acquireOnOwner/releaseOnOwner, broadcast the
  owner thread id to every worker, sum relayedAdmissions across workers (and
  main) in cluster_status, and remove the "run one http worker" warning. The
  restart-quarantine waiver (which lets a genuinely fresh node grant a key
  without waiting out a departed incarnation's lease) is cleared on the first
  handoff bump, so a successor coordinator built after ownership has already
  changed hands cannot grant while a departed worker's relayed handle can
  still commit.
- A caller worker that EXITS during an ownerless handoff counts as fenced. The
  restart quarantine does not back that up (it is read only on the home's grant
  path, so a peer-homed key renews straight back here); the process-wide native
  key lock does, since lock() takes it before the cluster admission and both the
  departed worker and any new caller are on this node. The residual teardown
  question is tracked as HarperFast/rocksdb-js#865. Recorded at the call site
  and in replication/DESIGN.md.
- Bump the core submodule to the matching harper change.

Test: new threads.count: 3 integration suite proves a lock() served on a
non-owner worker relays and succeeds, and concurrent increments stay
exclusive; core unit tests cover the remote-admission lifecycle including
fence-before-release; a transport unit test covers the waiver clearing on the
first handoff bump.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* Pin the relay handoff guards, and drop a retracted safety claim

Review adjudication on #865.

- `broadcastOwnerlessAndWait`'s JSDoc still credited the successor's restart
  quarantine for making a worker exit safe to treat as fenced. The inline
  comment in the same function, DESIGN.md and the PR ledger all retract that in
  favour of the process-wide native key lock; a later change that trusted the
  JSDoc would reopen the two-writer window believing the gate still covered it.

- `recordLockRpc.ts` had no unit coverage at all, so the caller-side relay
  lifecycle is now pinned: an in-flight acquire fails retryably when the
  coordinating thread goes away, a grant that lands after that is handed back to
  the thread that minted it, and a release carries the session its admission was
  minted under. Both guards were verified to fail without the code that provides
  them. Naming the ACQUIRE_REPLY handler is what lets a test deliver one.

- The existing `owner-c` assertion depended on the round-robin counter starting
  at zero, which only held while this file was the first to assign an owner; it
  now asserts the invariant it is named for.

- `key` is `unknown` throughout harper-pro's relay signatures, matching
  `recordLockRpc.ts` and `establishLockFreshness` in the same file.

Dispatch-Task: fix-kriszyp_harper-pro_865-1f3b384a
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* Do not confer record lock ownership on a successor that exited

Pre-push review finding on the fence wait this PR adds. `recordLockOwnerFor`
picks the successor before awaiting the incarnation bump and
`broadcastOwnerlessAndWait`, which runs as long as OWNER_FENCE_ACK_TIMEOUT_MS
and resolves a worker's own exit as a completed fence. So the wait can resolve
*because* the successor died, and `assignOwner` then confers on it.

Nothing recovers from that: `watchOwnerExit` attaches its listener inside
`assignOwner`, after the exit event it needs has already fired, so the entry is
never cleared and every relayed `lock()` for the database is routed to a dead
thread until the process restarts. Fail closed into the existing retry instead,
which re-derives over a fresh live set once the worker is back.

Before this PR the same window existed but spanned only the durable bump; the
fence wait widened it to ten seconds and made the successor's own death one of
the ways it completes.

Dispatch-Task: fix-kriszyp_harper-pro_865-1f3b384a
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* Check the successor by its captured thread id, not its post-exit one

Round-2 pre-push review, confirmed here on Node v26.2.0: a Worker reports
`threadId` -1 from before its `exit` listener runs, so the previous check asked
the thread tombstone about -1, never matched, and still conferred ownership on
the dead successor. `manageThreads.addPort` captures the id for the same reason.

The unit test masked it — its fake worker kept its id after exiting. It now
models a real Worker (tombstone keyed by the live id, `threadId` already -1) and
fails against the previous check.

Also corrects the `recordLocks` capability level in DESIGN.md, which still
documented 3 while `protocolCapabilities.ts` advertises 4.

Dispatch-Task: fix-kriszyp_harper-pro_865-1f3b384a
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* The caller relay's exit is not a fence; say so where the claim was made

cb1kenobi and the review bot both flagged the header comment added in 73fc424d:
it says the owner writes the delegation release when "this worker exits", which
the same file contradicts twice. `revokeRemoteHandle` settles only on the
caller's REVOKE_ACK or the handle's lease timer, and `onThreadExit` keeps a
departed caller's admission to its lease precisely so a write already handed to
the engine cannot be overtaken.

The borrowed native-key-lock justification does not reach this path either: the
next holder after a recall is a PEER node, so a process-wide key lock on this
node proves nothing. Exit-counts-as-fenced belongs only to main's ownerless
handoff, where both threads are on this node — the sibling claim dropped from
`broadcastOwnerlessAndWait` earlier in this branch.

Dispatch-Task: fix-kriszyp_harper-pro_865-1f3b384a
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* Prove withRow serializes on the row lock, not on its per-isolate queue

The review bot's remaining thread was right that nothing exercised the property
`withRow` was changed for: `recordLockHomes.test.mjs` scopes itself to pure
decision logic, and the threads.count: 3 cluster suite only drives the relay. It
asked for an integration case racing stage/fence/activate across workers, which
is not what this proves -- that race is probabilistic, and a test that cannot be
shown to fail against the old code is not coverage.

The discriminating fact IS testable in one isolate: hold the same node-scoped
lock on the `hdb_record_lock_homes` row that `withRow` takes, from outside
`withRow`'s own queue, and a stage must wait for it. Verified to fail against the
pre-#852 implementation (the per-isolate promise queue with a `transaction()`
write), where the stage runs straight through and settles while the row is held.
That the lock ALSO excludes across threads is core's property; one isolate cannot
demonstrate it, and the test does not claim to.

The end-to-end multi-worker race remains a follow-up, alongside the operator
operations' integration coverage generally.

Dispatch-Task: fix-kriszyp_harper-pro_865-c2d9fa46
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* Keep the handoff's two claims honest: the fence ack, and the attempt it belongs to

Three things the pre-push review found, all on the ownership handoff path.

A worker that could not fence one of its tables still acked the fence to main.
The ack meant "every relayed handle here is dead", main confers the successor on
it, and the successor may grant a key whose old handle can still commit.
`fenceRelayedAdmissionsForDatabase` now answers whether every table fenced, and
the worker withholds the ack when one did not, so main's wait times out and the
handoff fails closed -- which is the behaviour the gate was already built for.
Main's own fence is held to the same rule inside `broadcastOwnerlessAndWait`.
Nothing reachable throws there today (the resolver is a field read and core
swallows a resolver throw before this code sees it), so this enforces the
invariant the comment was arguing for rather than fixing a live defect.

A handoff settling late could act on another attempt's state. PENDING_BUMP is not
an identity: `releaseRecordLockOwner` clears it and the next attempt re-sets it,
so a rejection arriving after a release deleted the NEW attempt's marker and
scheduled a retry that re-assigned an owner to a database ownership had been
given up on. Each attempt now carries a token, checked on both settlement paths,
and a release bumps it. Covered by a test proven to fail without it.

The relayed acquire reserved a flat 250ms of the caller's wait for the two thread
hops, so `lock(id, { timeout: 200 })` -- or any lock that spent most of a longer
timeout on the native key first -- reached the owner with `waitMs: 0` and failed
on the first contention it met. Off-owner only, which is the uniformity the relay
exists to provide. The margin is no…
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant