Skip to content

Pace replication retries on one jittered schedule and bound subscription setup - #800

Merged
kriszyp merged 25 commits into
mainfrom
fix/replication-uniform-backoff
Sep 29, 2026
Merged

kriszyp merged 25 commits into
mainfrom
fix/replication-uniform-backoff

Conversation

@kriszyp

@kriszyp kriszyp commented Sep 2, 2026 •

Copy link
Copy Markdown
Member

One backoff schedule for every replication retry site, and admission control on the one surface where a retry could amplify: subscription setup.

The defect (#327)

A transient boot-time DNS failure kept onNodeUpdate → onDatabase firing. Each qualifying event turned straight into a retained (not unref'd) 200 ms setTimeout, a subscribe-to-node message, worker-side WebSocket/TLS setup, and one Setting up subscription with leader … warn — no dedup, no cap, no jitter. Whatever re-drove the event amplified 1:1: ~1,400 log lines/s/node in the field, ending in an OOM kill. Pacing alone would not have fixed it — the amplification is per event, not per delay — so the storm surface gets admission control as well as a schedule.

The utility

createBackoff in replication/backoff.ts — {initialMs, maxMs, minMs?, factor?, jitter?, budgetMs?, maxAttempts?, random?, now?}: exponential ceiling, full jitter, optional fixed floor, optional wall-clock/attempt budget, injectable RNG and monotonic clock. Two properties worth knowing:

  • an exhausted schedule returns undefined, never 0, so a missed exhaustion check cannot become a busy loop;
  • budgetMs is a deadline read off the clock, not a sum of requested sleeps, so a resolver hang or an event-loop stall cannot stretch a grace period past what it advertises.

Before this there was no jitter anywhere in production code: every exponential in the repo was lockstep across a fleet.

Admission control

Subscription setup (createSubscribeSetupScheduler) holds at most one armed setup per (peer URL, database), carrying the newest payload, on a 200 ms floor / 400 ms → 30 s jittered ceiling that resets when the pair connects. It lives in its own map rather than on the connectionReplicationMap entry because onDatabase's stale-worker path deletes and recreates that entry — per-entry state would be wiped on exactly the path that most needs the dedup. Within a stale-worker sweep, each setup's fire time, on a monotonic clock, slides past any the sweep already armed within RECONNECT_STAGGER_MS, so the 50 ms spacing that bounds concurrent TLS setup (#446) holds across independent jitter draws, and a pair whose own backoff has escalated is placed at its own delay without dragging the rest of the sweep out to its ceiling. Setups are cancelled on unsubscribe, node deletion and same-name URL migration — all reachable once a pending timer can live 30 s instead of 200 ms.

An armed setup never falls onto the main thread. If its worker exited while it waited, dispatchSubscriptionRequest defers it (with a warn) to the stale-worker reconcile, which rebinds the entry; main would have run subscribeToNode on the orchestrator. Only single-thread mode, where the main thread is the worker, dispatches there.

Self-catchup is claimed by the first dispatch for one connection entry, which keeps the rider for its life. Until the owning primary connection opens it rides on every recovery re-drive; after that, re-drives to that worker leave it off (its old startTime would make the leader rescan from there to now) and a replacement worker is sent it again, since the main thread cannot see catchup finish. The global claim is consumed only once the worker message is accepted, so cancellation, a long offline retry, or a synchronous postMessage throw cannot lose it.

The worker readiness boundary used to attach a continuation to the component-readiness promise per message (createWorkerSubscriptionAdmission). Before readiness, one insertion-ordered map retains the latest action per connection key and database behind a single continuation, with exactly one armed re-attempt if that readiness rejects; after readiness the handlers run inline. Subscribe, unsubscribe, force-reconnect and this map all derive identity through one getSubscriptionConnectionKey in replicator.ts, which replaces three hand-built url + '-' + … keys, including the missing nested-URL fallback.

Recovery timers are owned. One entry.reDriveTimer per entry, unref'd, disarmed on unsubscribe, node deletion, worker exit and entry replacement, with the sweep's jitter drawn once and used as a common base. A connect report resets the escalated setup delay and cancels nothing; each timer re-checks live state when it fires. The stall kick's half of that is shouldFireStallKick, which claims its entry through connectGeneration and the worker it captured.

Review-thread fixes in this round

  • Leadership is recomputed, not carried forward. onDatabase derives it through deriveEffectiveLeader before the existing-entry fast path and assigns the result, so an explicit persisted isLeader: false demotes immediately. Recomputing exposed a second way to get it wrong: a component reload cleared routes before rebuilding it, so a recomputation could briefly see "no configured leader" while the previous hdb_nodes watcher stayed live. The reload now builds the list off to the side and publishes it in one splice, keeping the array identity.
  • Only the monitored connection's open cancels a stall kick. connectReportAdvancesGeneration advances connectGeneration only for a socket-open edge (newSocket) from the owning worker's thread whose subscription URL (reported with the open) is the entry's primary. That last check matters because one worker can carry both an entry's primary connection and a proxied failover subscription over the same URL. Pongs, metadata posts, superseded workers' sockets and failover opens no longer move it. A shared-memory truth up-correction does, since it stands in for an open edge that never arrived. Look hardest here.

Sites adopting the schedule

scheduleReconnect (full jitter under the unchanged 500 ms → 30 s ceiling, keeping the hard 500 ms floor from #339; onFrameSent is still the only reset), the wedge and receive-stall re-drives, the hdb_nodes watcher restart, the send-auth reprobe (a 30 s wall-clock deadline in place of 60 fixed sleeps), the copy-cursor flush retry, the two clone probes (monitorSync's leader version probe, 1 s → 4 s; fetchJWTKeyWithRetry, 250 ms → 1 s; attempt counts unchanged), the blob-repair per-record pause, and the sendBlobs in-place 503 re-read. DESIGN.md carries the per-site table and the list of what is deliberately excluded.

One definition per schedule. The subscription-setup parameters (200 ms floor, 400 ms → 30 s ceiling) live in createSubscribeBackoff, which the setup scheduler, the wedge/stall re-drive draw and the worker readiness re-attempt all call. The re-drive only ever takes the first draw, so its window is still [200, 400) ms. The scheduler's timing overrides (initialMs/maxMs/minMs/staggerMs) are removed; nothing passed them. The blob-send array [250, 500, 1000, 2000] is now BLOB_SEND_RETRY_BACKOFF (maxAttempts: 4, jitter: 'none'). It produces the same four waits, and the attempt bound now lives in the schedule instead of being split between shouldRetrySourceBlobRead and an array index. sendBlobs creates the backoff lazily on the first eligible 503, so a successful send allocates nothing. The stored-body, partial-send, closed-socket and draining guards are unchanged. It stays unjittered: the retried read is this node's own blob, so there is no fleet to decorrelate, and full jitter would halve the expected time the PENDING placeholder has to heal (DESIGN.md row). Planning review for this consolidation: Framing-Verdict: chosen-approach-sound (a4fea92b7e93).

The hdb_nodes watcher reset bug: iteratedSuccessfully was set the instant subscribe() resolved, so a subscription that resolved and immediately threw counted as a success every pass and pinned the restart delay at 1 s forever. Progress is now an iteration that stayed live past NODE_WATCHER_HEALTHY_UPTIME_MS, measured from when the subscription actually came up.

Merged with main

The branch was 76 commits behind and conflicting, so origin/main is merged in rather than rebased (no force-push; the conflict resolution is one commit). Resolutions that change behaviour: the socket-open report now uses main's threadId/newSocket fields plus this branch's subscriptionUrl, instead of two parallel sets; scheduled setups carry main's send-time exclusionOrigins; and main's hasDeadOwner note is corrected, since an unowned entry now defers instead of subscribing on the main thread.

For the human reviewer

Decisions worth disagreeing with:

  • Explicit persisted isLeader: false beats a configured leader (HDB_LEADER_URL / cli / routes[0]); main ignored a persisted false. The owner chose this. One consequence: update_node { isLeader: false } is written LOCAL_ONLY only when true (setNode.ts), so a demotion issued on one node replicates to every peer and overrides their configuration too. In practice the derived flag only gates shouldReplicateFromNode for a database that does not exist locally, and the bootstrap loop that reaches that case already keys on the raw persisted flag. So the reach is narrow, but it is a policy change.
  • Full jitter at the reconnect site, resolved by the task owner in favour of decorrelation. It halves the expected dial interval the Harper OOM-killed (SIGKILL) during high-throughput bulk upsert writes (10 GB) #339 fix installed (mean ~15 s against a deterministic 30 s at the cap), so a node facing a permanently-dead peer dials ~4×/min instead of 2×/min — far below the incident rate, and the 500 ms floor keeps the hard minimum. Equal jitter would restore the old mean and is a one-line change.
  • Subscription setup escalates to a 30 s ceiling. This path is also the recovery path, so worst-case time-to-resubscribe for a never-connecting pair grows from 200 ms to 30 s in order to bound the storm.
  • Self-catchup is re-sent per worker replacement, not per re-drive. One historical scan per worker replacement buys not losing the range when a worker dies mid-catchup. A same-worker wedge re-drive after the open still replaces the worker's subscription list without the rider, as main always did; closing that needs a worker→main "catchup complete" signal.
  • The first reconnect is exactly 500 ms on every link (floor = initial ceiling), so after a shared blip where the peer stays up, the first redial is not decorrelated. Jittering it is one constant (initial 1000 ms, floor 500 ms).
  • The hdb_nodes watcher backoff resets only after 10 s of live iteration. A stream that keeps ending sooner stops watching for up to 30 s at a time; the collection subscribe replays current rows on restart, so the cost is latency, not lost updates.
  • A setup whose worker exited waits for the reconcile instead of running on the main thread — up to one reconcile interval with no subscription for that pair, in exchange for never doing replication work on the orchestrator.
  • The setup schedule outlives its entry, deliberately — but that is a second lifecycle every future entry path has to remember to cancel.
  • The armed setup carries its own payload plus a refreshPending side channel, rather than fixing onDatabase's early-return path so entry.nodes is always the enriched array.
  • NodeReplicationConnection.retryTime is removed. On main it was a mutable field; nothing outside the unit tests read it, and the tests now assert retryBackoff.ceiling.
  • The blob-repair single-flight guard was removed: getRecord() has no timeout, so one wedged-but-open peer would hold the flag for the life of the thread and take away repair_blob_data until restart. Concurrent sweeps stack again, as on main.
  • A hopeless blob-repair sweep runs unpaced after 60 s of consecutive failures instead of stopping: a cluster-wide-lost prefix must not make the repairable records behind it unreachable. Both the per-record and per-peer warns are sampled once unpaced. A reviewer proposed keeping a capped delay instead; that holds the cursor open ~1 s per unrepairable record for the rest of the sweep, which is what the budget exists to end.

Still open, tracked or stated:

  • A same-name URL migration can still force-reconnect the address the node left. The migration cancels the setup scheduler for the old URL, but the old connectionReplicationMap entry survives with its owned reDriveTimer. The entry leak predates this branch; the timer riding on it does not.
  • Pre-readiness unsubscribes are retained as tombstones (Pre-readiness unsubscribes are retained as tombstones, so the worker-subscription admission map can grow unbounded #807): with a persistently rejecting readiness promise plus peer/database churn, the admission map keeps one entry per historic connection key.
  • Pre-existing, surfaced by review, not fixed here: connect() checks intentionallyUnsubscribed only before the createWebSocket await, so an unsubscribe landing inside a TLS handshake can leave an orphan live session; and the revive path (unsubscribed = false) does not reset receiveStallReconnectAt, so an unsubscribe/re-add can hide a stalled connection from the stall net until real progress.
  • Coverage gap, blob-send retry: no fixture shows a 503 healing during the in-place retry. blobGapEscalationBudget covers the 503-forever path (retry, then forward), and the unit test pins the exact waits and exhaustion.
  • Coverage gap: startOnMainThread encloses onNodeUpdate, onDatabase, connectedToNode and reconcileWorkers, so every unit test targets an extracted helper. The end-to-end evidence is the cluster and stress suites below.

Review findings declined, with the evidence:

  • "A deferred setup waits on the reconcile tick." findStaleNodeUrls flags a worker-less entry whenever live workers exist, and a worker exit triggers an immediate reconcile, so the wait is one tick at most — the same bound main had when its bare setTimeout posted to an exited worker.
  • "Dropping !isWedged adds TLS churn." A URL that is both stale and wedged now also gets the full onNodeUpdate pass, but its wedged live-worker entries take the refresh-only early return (forceResubscribe is false there), so no extra dial happens; the pass only re-posts unsubscribe-from-node for databases not replicated from that peer, until the stale entries are rebound on that tick.
  • "Drop the arithmetic comments in forceReconnect.test.mjs." They name the pinned draw's delays, which is what makes tick(500) / tick(998) readable.
  • "canClearCapabilitiesForNewSocket ignores subscriptionUrl" and "pin the explicit-leader bootstrap gate with a test" — both pre-existing on main, recorded for follow-up rather than folded in here.

Verification

On the merged head (main at c2ad3351), then on the consolidation head bacd6b6c:

  • Build — npm run build exits 0 with no tsc errors. npm run lint:required and prettier --check on changed files clean.
  • Unit — npm run test:unit: 1412 passing, 1 pending; 1410 on bacd6b6c (the two removed tests were the retryTime getter's and the array-length bound's; blobSendRetry.test.mjs now pins [250, 500, 1000, 2000, undefined], and reconnectJitter/forceReconnect/connectReschedulesOnRejection read retryBackoff.ceiling). Covers the backoff matrix, the setup scheduler (a 60,000-event / 60 s storm collapsing to 9 dispatches with one pending timer, hasRef() === false, sweep spacing by fire time under colliding draws, staggered calls, and an escalated pair, deferral when the worker exits before fire), leadership derivation and atomic route publication, socket-open ownership (pong, superseded worker, proxied failover, single-thread), the rider re-sent to a replacement worker, worker readiness admission, and the stall-kick decision matrix.
  • connectedBitRestartChurn (QA-587), locally — 5 of 6 runs green across this round's heads (1/1 on the final head), including the new rolling HTTP-worker replacement test (restart_service http_workers, then a write per database must reach the follower) in every run. The one red run was chaos cycle 3 missing its 25 s budget: one leg connected ~100 ms before the SIGKILL, never sent a frame, so its reconnect ceiling was not reset and it drew 29.6 s at the 30 s cap. main has the same reset-on-first-sent-frame rule and waits the full ceiling, so this is the test's budget against that rule, not this change.
  • Test files. New: backoff.test.mjs (ceiling growth, factor, full-jitter range, floor, budget/attempt exhaustion on an injected clock), reconnectJitter.test.mjs (decorrelation, the 500 ms → 30 s ceilings, the 500 ms floor, onFrameSent reset), subscribeSetupScheduler.test.mjs (storm bound, dedup, sweep spacing, deferral), workerSubscriptionAdmission.test.mjs (connection-key derivation, latest-action retention, rejected-readiness re-attempt), subscriptionRoutingOwnership.test.mjs (atomic route publication, persisted isLeader precedence, socket-open ownership), shouldFireStallKick.test.mjs (the stall-kick decision matrix) and blobRepairPacing.test.mjs (the budget starts at the first unrepairable record). Extended: nodeUpdateWatcher.test.mjs (escalation when iteration throws at once, healthy-uptime reset, subscribe time not counted as uptime), shouldCloseSendAuthWatch.test.mjs (fail closed past the deadline, including a read that crosses it), findStalledReceivingNodeUrls.test.mjs (a fresh socket's grace period), and forceReconnect.test.mjs / connectReschedulesOnRejection.test.mjs (pinned draws against the jittered ceiling, read from retryBackoff.ceiling).
  • Integration files. connectedBitRestartChurn.test.mjs gains the rolling HTTP-worker replacement test (and a comment now names the reconnect ceiling instead of the removed retryTime); fixture-blob-pending-source/resources.js changes only a comment, to the new constant's name.
  • Consolidation head bacd6b6c, cluster — blobGapEscalationBudget (exercises the in-place 503 re-read, then the forward) and connectedBitRestartChurn: 6/6 pass.
  • The replication: no-backoff subscription-setup retry storm on transient boot-time DNS failure ends in OOM #327 scenario end to end — integrationTests/stress/wedgedPeerSubscribeStorm.test.mjs under HARPER_RUN_STRESS_TESTS=1: passes on the merged head and again after the round-1 review fixes (bounded Setting up subscription with leader rate, no listener leak, no OOM marker, RSS under cap, reconverging mesh with a peer permanently wedged).

Pre-push review of this round (prepush-review.mjs --author claude --legs auto, codex + gemini + cursor-composer + harper-domain every round): a full round on the merged head (adjudicated minor), then delta rounds — nit, nit (COMMENTS), and nit (COMMENTS, on the pushed head) after a post-push human review showed the sweep compared relative delays from calls made at different instants (now absolute fire times). The full-round findings — the sweep floor dragged by an escalated pair, self-catchup lost when a worker dies mid-catchup, unbounded integration probes, the unsampled per-peer repair warn — are all fixed above, as is round 2's spacing gap next to a just-escalated pair. The consolidation commits got two delta rounds (all four lenses): nit (COMMENTS; a test title and a field comment, both fixed), then LGTM on bacd6b6c.

Refs #327
Fixes #805

🤖 Generated with Claude Code

https://claude.ai/code/session_017CTUbhGiW2AXhogHZLg8eo

— Claude Opus 5.5

Origin — the dispatch brief this PR was written from

Pace replication retries on one jittered schedule and bound subscription setup

LIVE CONVERSATION about #800.

You are answering a person, in a thread, one turn at a time. Every turn:

  1. Read the whole thread in this dispatch file's # Log — it is the conversation so far, and
    each of your previous turns is in it. Read the PR/issue and the code as needed.
  2. Answer the LAST message. Append your answer to # Log as your turn. Prose, not a report:
    they are talking to you, and a status template is not an answer.
  3. Set status: needs-input and stop. The thread stays open; their next message resumes it.

Each turn arrives as ASK (answer it, change nothing) or PERFORM (do it, then say what you did) —
the person chose which when they sent it, and the run's own prompt tells you which one this is.
Never infer it from the wording: an unrequested commit in the middle of a discussion and a polite
description of work that was supposed to happen are the two failures this exists to prevent.

Never mark a PR ready and never merge from this conversation.

Dispatch: task chat-pr-harper-pro-800-kriszyp · queued by unknown · ran by claude/opus/xhigh · worker kzyp-xps-1

Review-Coverage: authored=claude; ran=codex,cursor-composer,gemini; adjudicated=domain; declined=cursor-grok,cursor-kimi,cursor-muse; rounds=18; full=1 @ bacd6b6

Human-Review-Need: 4 (decisions: self-catchup-retained-for-entry-life, reconnect-first-attempt-unjittered, full-jitter-at-cap, send-auth-fail-closed-at-deadline, blob-repair-budget-then-unpaced, watcher-healthy-uptime-reset) @ bacd6b6

kriszyp and others added 12 commits September 1, 2026 18:35
Replication had no jitter anywhere in production code and a different
retry policy at every site, from a flat 200ms subscription-setup delay to
jitterless exponentials. The subscription-setup site had no dedup, no
cap, and a non-unref'd timer, so whatever re-drove onNodeUpdate amplified
1:1 into main-thread timers, worker-side WebSocket/TLS setup, and
"Setting up subscription with leader" warns — ~1,400 lines/s/node in the
field, ending in an OOM kill (harper-pro#327).

Add replication/backoff.ts: one createBackoff schedule with an
exponential ceiling, full jitter, an optional per-site floor, an optional
wall-clock budget, and injectable RNG/clock. Adopt it at the
subscription-setup scheduler, scheduleReconnect, the wedge and
receive-stall re-drives, the hdb_nodes watcher restart, the send-auth
reprobe, the clone JWT/version retries, and the blob-repair sweep.

The subscription-setup scheduler additionally enforces at most one
pending setup per (peer URL, database), keyed in its own map because the
stale-worker path deletes and recreates the connectionReplicationMap
entry. Its dispatch re-reads the live entry at fire time rather than
capturing a payload, so a deduped update is never lost and no closure is
allocated per suppressed event, and the setup is cancelled on connect,
unsubscribe, and node deletion.

Also fix the hdb_nodes watcher's premature backoff reset: the success
marker was set the instant subscribe() resolved, so a subscribe-then-
immediately-throw cycle reset the delay on every pass and never
escalated. It now resets only after an iteration survives a healthy-
uptime threshold.

Refs #327

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LHXY4wfz8qJ4ryA5gKPAme
The deduped scheduling path read `entry.nodes` at fire time, but
`onDatabase` replaces that array on its early-return path *without*
running the leader/url enrichment — so a node update arriving between
arming and firing left the setup posting a request with no url, and the
worker logged "Failed to create web socket to undefined" and never
connected (caught by selectiveTableSubscription.test.mjs). The schedule
now carries the payload of the call that armed or last refreshed it;
only calls that did enrich reach the scheduler, so "newest payload wins"
holds without the clobber.

Also resolve `nodes[0].url` where the payload is built rather than at the
scheduling site, so the wedge re-drive — which posts from `entry.nodes`
— cannot inherit the same urlless array.

Refs #327

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LHXY4wfz8qJ4ryA5gKPAme
…covery timers

From the independent pre-push review of the previous two commits:

- send-auth reprobe (blocker): checking the budget only at the top of the
  loop let a row that became decodable *after* an event-loop stall pushed
  us past the 30s deadline still authorize the peer — the loop exits on a
  non-UNCHANGED row without consulting the deadline. It now fails closed
  on the elapsed time before trusting the next read. The backoff is also
  built lazily, so an ordinary decodable authorization event allocates
  nothing and never reads the clock.
- scheduleReconnect: a fixed 100ms floor under full jitter did not
  preserve the #339 dial-rate guard, which is ceiling-relative by nature.
  Added `jitter: 'equal'` to the utility (draw from the top half of the
  window) and used it here, so the minimum dial interval stays
  proportional as the ceiling escalates. Full jitter stays the default
  everywhere the invariant is latency rather than rate.
- subscription setup: `onDatabase`'s early-return path builds a fresh
  payload but never reaches the scheduler, so an armed setup could still
  dispatch pre-update routing state. It now calls `refreshPending`, and
  carries `isLeader` forward onto the replacement array so neither that
  path nor the wedge re-drive loses the leader decision.
- wedge / receive-stall re-drives: each entry now owns a single
  `reDriveTimer` that a later decision replaces and connect/unsubscribe/
  delete disarm, so a sweep staggered across many databases cannot stack
  waves or fire after an unsubscribe. The stall kick additionally
  re-reads the receive watermark at fire time, so a copy that resumed
  during the delay is not reconnected.
- a same-name node moving to a new URL now cancels any setup armed for
  the address it left.

Refs #327

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LHXY4wfz8qJ4ryA5gKPAme
Bound worker-side pre-readiness subscription work, retain self-catchup state until dispatch, and make exhausted backoffs fail closed. Align reconnect and watcher jitter with the full-jitter policy and extend regression coverage.

Co-Authored-By: GPT-5 Codex <noreply@openai.com>
Draw the wedge/stall re-drive jitter once per reconcile sweep instead of once per
entry: a per-entry draw varied consecutive delays by up to ±200ms, letting several
dials share a 50ms instant and weakening the concurrency bound RECONNECT_STAGGER_MS
exists to hold (#446). Only the worker that currently owns an entry may retire its
pending setup on connect, so a stale worker's still-open connection cannot cancel the
setup armed for the entry's replacement. End a blob-repair sweep once the schedule is
exhausted rather than pacing an unrepairable backlog at 1s/record with the cursor open,
and record why the subscription connection-key helper keeps its two teardown callers'
argument order.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HEgBi5zh8QKMbFmeGwFv7w
A component-readiness rejection retained the pre-readiness subscribe/unsubscribe
actions but never re-attempted them, so nothing applied until another parentPort
message arrived or the wedge reconcile noticed ~30s later — the empty-subscription
window this gate exists to close. Re-attempt on the same capped, jittered, unref'd
schedule the rest of the discipline uses, resetting once readiness succeeds.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HEgBi5zh8QKMbFmeGwFv7w
…ownership

- shouldCloseSendAuthWatch: the row read is synchronous and can itself carry the clock
  past the advertised grace period, so re-check the deadline after it and fail closed.
- repairBlobs: bound how long a failure run may spend *pausing* instead of ending the
  sweep on a consecutive-failure count — a cluster-wide-lost prefix must not make the
  repairable records behind it permanently unreachable through the operation. Drop the
  per-thread single-flight guard with it: `getRecord()` has no timeout, so one wedged
  peer would hold the flag for the life of the thread and remove the operator's lever.
- runNodeUpdateWatcher: stamp the health clock once the subscription is live, so a
  subscribe() that blocks past the threshold and then throws no longer reads as uptime.
- Worker admission: own the readiness re-attempt so a pre-readiness burst against a
  rejecting import cannot multiply timers, imports and warn lines 1:1 with messages.
- Connect ownership: tag the report with the sending thread id rather than relying on
  core's undocumented onMessageByType listener arity, behind a tested pure helper.
- Disarm an entry's recovery timer when its worker exits or the entry is replaced, and
  adopt the shared schedule for the copy-cursor flush retry.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HEgBi5zh8QKMbFmeGwFv7w
The backoff table described the blob-repair cutoff the previous commit replaced and
listed neither the copy-cursor flush retry nor the worker readiness re-attempt; the
exclusions paragraph now also names the copy-finalize timeout bound, which is a wait
bound rather than a retry and must stay unjittered. Give the readiness re-attempt the
same floor every other adopting site has, so a zero draw cannot re-probe on the next
macrotask. Sample the per-record blob-repair warn once the sweep goes unpaced, and put
clearWorkerFromEntries' description back over clearWorkerFromEntries.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HEgBi5zh8QKMbFmeGwFv7w
Gating the connect edge on "is this the entry's current worker" is not answerable on the
main thread: connectToNextWorker subscribes a failover peer on a worker that is not
entry.worker, and a superseded worker's hung-but-open connection reports for the same
(url, database) as its replacement. Gating on it meant the escalated setup delay stopped
resetting, and connectedBitRestartChurn's chaos cycles converged in 5.3s, 12.0s, then not
within 25s — a backoff compounding across reconnects. Reset the delay unconditionally and
leave the armed setup and recovery timers to their own fire-time guards, which is what
this path did before the schedule existed. That suite is green again.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HEgBi5zh8QKMbFmeGwFv7w
Removing the connect-edge cancellation left the receive-stall kick with no guard that
observes a reconnect: entry identity holds, receiveStallReconnectAt is untouched by
connectedToNode, and the watermark cannot move on a socket that just opened — so a leg
that drops and reconnects inside the kick's stagger window was force-reconnected on the
strength of the old socket's watermark. Claim it through a connect generation, the way
the wedge kick claims its entry through disconnectedAt (which a stalled, connected:true
entry does not have). The design table, its prose, and dispatchSubscribeSetup's JSDoc
still described the cancel-on-connect behavior that commit removed.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HEgBi5zh8QKMbFmeGwFv7w
entry.receiveStallReconnectAt both claims the armed kick and throttles re-detection,
which needs lastReceivedTime past the stamp. Skipping the kick because the leg
reconnected inside the stagger window therefore spent an epoch on a kick that never
happened: if the fresh socket stalled too it never moved the watermark past the stamp
and the receive-stall net never re-armed for that (peer, database) again. The decision
moves into shouldFireStallKick — a pure helper with the arm/bump/fire cases under test,
which is what the guard shipped without.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HEgBi5zh8QKMbFmeGwFv7w

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request introduces a unified, jittered backoff discipline across the replication system to mitigate thundering herd issues and prevent OOM conditions during retries. Key changes include the addition of a generic createBackoff utility, a SubscribeSetupScheduler to dedup and pace subscription setups, and a createWorkerSubscriptionAdmission gate to manage worker readiness. The changes are well-documented in DESIGN.md and supported by extensive new unit tests. I have kept the review comment regarding Map lookup optimization in subscriptionManager.ts as it provides a valid improvement opportunity.

Comment thread replication/subscriptionManager.ts Outdated
Addresses the Gemini review comment on #800.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HEgBi5zh8QKMbFmeGwFv7w
@kriszyp
kriszyp requested review from cb1kenobi and heskew September 2, 2026 23:01
@kriszyp
kriszyp marked this pull request as ready for review September 2, 2026 23:01
@kriszyp
kriszyp requested a review from a team as a code owner September 2, 2026 23:01
@claude

claude Bot commented Sep 2, 2026 •

Copy link
Copy Markdown
Contributor

Reviewed; no blockers found.

Comment thread replication/subscriptionManager.ts Outdated
Comment thread replication/subscriptionManager.ts Outdated
kriszyp and others added 7 commits September 6, 2026 22:47
Recompute effective leader state before subscription reuse and publish route reloads atomically. Scope stall generations and delayed kicks to the worker that owns the monitored connection.

Co-Authored-By: GPT-5 Codex <noreply@openai.com>
Retain the self-catchup rider across recovery dispatches and chain jittered setup delays through a shared floor so reassignment handshakes remain spaced. Use deterministic injected timers for scheduler coverage.

Co-Authored-By: GPT-5 Codex <noreply@openai.com>
Co-Authored-By: GPT-5 Codex <noreply@openai.com>
… the orchestrator

A worker can carry both an entry's primary connection and a proxied failover subscription, so a
socket-open report now carries the subscription half of its connection key; only the owning worker's
primary open advances connectGeneration and retires the main-thread self-catchup rider. An armed setup
whose worker exited before it fired is deferred to the stale-worker reconcile instead of subscribing on
the main thread, and that reconcile now reassigns stale entries even when a sibling database is wedged.

Co-Authored-By: GPT-5 Codex <noreply@openai.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017CTUbhGiW2AXhogHZLg8eo
Dispatch-Task: fix-kriszyp_harper-pro_800-16768da5
…rm-backoff

Conflict resolutions that change behaviour on either side:
- Socket-open reports use main's `threadId` and `newSocket` fields plus this branch's `subscriptionUrl`;
  the branch's duplicate `reportingThreadId`/`opened` fields are gone. A truth up-correction advances
  the connect generation directly in the reconcile, since it must not carry `newSocket`.
- Scheduled subscription setups carry main's send-time `exclusionOrigins`.
- The unsubscribe teardown keeps main's retire-before-close order on the shared connection key helper.
- This branch's DESIGN note on atomic route publication is renumbered 23 after main's 18-22.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017CTUbhGiW2AXhogHZLg8eo
Dispatch-Task: fix-kriszyp_harper-pro_800-16768da5
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017CTUbhGiW2AXhogHZLg8eo
Dispatch-Task: fix-kriszyp_harper-pro_800-16768da5
…hup to a replacement worker

The stale-worker sweep chained its spacing floor through every setup, so one pair whose own backoff had
escalated to ~30 s dragged every later database in the sweep out to its ceiling. The scheduler now owns
the sweep rule: a setup chains the floor only when it lands inside the fresh-draw window above it.

The self-catchup rider was dropped from the entry when the owning primary connection opened, so a worker
that exited mid-catchup took the range with it. The entry now keeps the rider and records which one the
current worker opened with: recovery re-drives to that worker still skip it, a replacement worker is sent
it again.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017CTUbhGiW2AXhogHZLg8eo
Dispatch-Task: fix-kriszyp_harper-pro_800-16768da5
@kriszyp
kriszyp merged commit 505ff33 into main Sep 29, 2026
45 checks passed
@kriszyp
kriszyp deleted the fix/replication-uniform-backoff branch September 29, 2026 11:06
kriszyp added a commit that referenced this pull request Sep 29, 2026
… only inline blobs can earn one

Round-12 rebase review finding (post-#800 rebase onto main): settleReceiptRequests
reused storedBlobsAreComplete, the duplicate-skip verifier, to gate a receipt on
blob durability. That function only counts FILE-backed blobs and returns false
whenever none are found, including when a HAS_BLOBS row's blobs are all inline
(below core's file-storage threshold) -- so a transition image whose blobs are
all inline could never earn a receipt: every request waited out its TTL and the
sender's sweep re-sent the same image forever, pinning the image permanently.

Added receiptBlobsComplete: it still finds every reachable blob to catch a
genuinely unreachable HAS_BLOBS claim (the same conservative case
storedBlobsAreComplete declines), but an inline blob needs no file check --
its bytes are already inside the row's own encoded value, which the caller
already confirmed durable. Only file-backed blobs still go through
blobFileMissingOrIncompleteAsync. storedBlobsAreComplete/isDurableIdentityTie
are untouched, so the duplicate-skip's conservative behavior for inline blobs
is unchanged.

New unit coverage in durableIdentityTie.test.mjs exercises real stored blobs
(inline and file-backed) through receiptBlobsComplete, the same fixture
pattern the neighboring isDurableIdentityTie file-blob tests use.

Also: correct an inbound-batch comment that overclaimed uniqueness (the sender
guarantees no repeated record per batch; a violation would only race
recordHandoffReceipt to a lower stored version, harmlessly), and fix a missing
line break in DESIGN.md that ran the Redelivery bullet into Release.

Declined, with reasons (framing recheck cleared this change at
chosen-approach-sound before implementation; both declines recorded in the PR
body's decision ledger): a same-version-different-origin receipt-proof gap
that needs a wire/capability change for a narrow, compound scenario; a
resweep-timer/wait-site nit whose cost is zero on any core that exists today
and whose fix touches a hot loop already hardened over many liveness rounds.

Dispatch-Task: pr-maint-0e208afb66ea625b57ed77ba834daf66
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
cb1kenobi added a commit that referenced this pull request Sep 29, 2026
One content conflict, in replication/DESIGN.md: both sides appended a
note numbered 23 and a row to the "where is X" table. Kept main's note 23
(config-route reload publishes atomically), renumbered this branch's to
24, and merged both cheat-sheet tables so main's three new rows and this
branch's replicate:false row all survive.

Checked rather than assumed, since #800 touches the same subsystem: main's
replicationConnection.ts hunks are the copy-flush retry backoff, a
send-auth log constant and the sendBlobs read-retry backoff, none of which
touch the outbound table gate, the GET_RECORD refusal, the schema builders
or the receive drop. The core bump still exposes the three APIs this branch
reads — schemaDescribe's replicate field, HAS_BLOBS and Table.replicate.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
kriszyp added a commit that referenced this pull request Sep 30, 2026
…ress, so it no longer carries across outages (#948)

* Reset a receiving leg's reconnect backoff on durable progress, not only on a sent frame

An outbound replication leg's reconnect backoff reset only when that leg sent a transaction
frame. In a two-way mesh the peer serves our subscription from its own server session, so a
subscriber's outbound leg never sends one: its ceiling carried across unrelated outages until
every reconnect waited out the 30 s cap. #800's full jitter reached the cap within three
outages, which is how connectedBitRestartChurn's chaos cycle 3 started missing its 25 s bound.

The receiving side now credits its durable watermark where it advances (a commit, and the
last in-flight blob's drain) through NodeReplicationConnection.onDurableProgress. Only an
advance past the highest value the connection has already credited counts, and only from the
session that still owns the socket, so a session that re-receives the same undurable frames
after every reconnect keeps escalating (harper-pro#339). The churn test now asserts that every
database logs a fresh disconnect in every cycle, which only a reset backoff does.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Dispatch-Task: main-red-kriszyp_harper-pro_039c07d6c_f0efc59e

* Credit receive progress only when it is durable and the leg's own sends work

From the pre-push review. Copy-apply rows are not durable until the copy's flush, so the
watermark they move is not progress yet. A sender-loop failure on a leg that also serves a
subscription now vetoes receive credit until a frame is sent again, so a leg whose own sends
fail every time (#713) does not redial at the floor on the strength of what it received. A
non-finite or out-of-range sequence can no longer pin or disable the credited maximum, and the
send-side reset now checks socket ownership the way the receive side does.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Dispatch-Task: main-red-kriszyp_harper-pro_039c07d6c_f0efc59e

* Clear the send-failure veto once the sender catches up, and credit a finished copy

From the second pre-push review round. The veto previously cleared only on a sent frame, so a
leg whose sender threw once and then had nothing to send kept escalating across outages. It
now also clears when the sender loop reaches its idle wait, which a loop stuck on an
oversized frame never does. A copy-apply copy's completion now counts as progress, since its
final flush has run by the time copy mode ends.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Dispatch-Task: main-red-kriszyp_harper-pro_039c07d6c_f0efc59e

* Apply receive progress the send-failure veto held back once the sender catches up

From the third pre-push review round: an advance credited during the veto was dropped, so a
leg whose whole catch-up landed before its sender caught up entered the next outage escalated.
onSenderCaughtUp now applies it, and replaces the direct write to the veto field.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Dispatch-Task: main-red-kriszyp_harper-pro_039c07d6c_f0efc59e

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
kriszyp added a commit that referenced this pull request Sep 30, 2026
… only inline blobs can earn one

Round-12 rebase review finding (post-#800 rebase onto main): settleReceiptRequests
reused storedBlobsAreComplete, the duplicate-skip verifier, to gate a receipt on
blob durability. That function only counts FILE-backed blobs and returns false
whenever none are found, including when a HAS_BLOBS row's blobs are all inline
(below core's file-storage threshold) -- so a transition image whose blobs are
all inline could never earn a receipt: every request waited out its TTL and the
sender's sweep re-sent the same image forever, pinning the image permanently.

Added receiptBlobsComplete: it still finds every reachable blob to catch a
genuinely unreachable HAS_BLOBS claim (the same conservative case
storedBlobsAreComplete declines), but an inline blob needs no file check --
its bytes are already inside the row's own encoded value, which the caller
already confirmed durable. Only file-backed blobs still go through
blobFileMissingOrIncompleteAsync. storedBlobsAreComplete/isDurableIdentityTie
are untouched, so the duplicate-skip's conservative behavior for inline blobs
is unchanged.

New unit coverage in durableIdentityTie.test.mjs exercises real stored blobs
(inline and file-backed) through receiptBlobsComplete, the same fixture
pattern the neighboring isDurableIdentityTie file-blob tests use.

Also: correct an inbound-batch comment that overclaimed uniqueness (the sender
guarantees no repeated record per batch; a violation would only race
recordHandoffReceipt to a lower stored version, harmlessly), and fix a missing
line break in DESIGN.md that ran the Redelivery bullet into Release.

Declined, with reasons (framing recheck cleared this change at
chosen-approach-sound before implementation; both declines recorded in the PR
body's decision ledger): a same-version-different-origin receipt-proof gap
that needs a wire/capability change for a narrow, compound scenario; a
resweep-timer/wait-site nit whose cost is zero on any core that exists today
and whose fix touches a hot loop already hardened over many liveness rounds.

Dispatch-Task: pr-maint-0e208afb66ea625b57ed77ba834daf66
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
kriszyp added a commit that referenced this pull request Oct 6, 2026
… only inline blobs can earn one

Round-12 rebase review finding (post-#800 rebase onto main): settleReceiptRequests
reused storedBlobsAreComplete, the duplicate-skip verifier, to gate a receipt on
blob durability. That function only counts FILE-backed blobs and returns false
whenever none are found, including when a HAS_BLOBS row's blobs are all inline
(below core's file-storage threshold) -- so a transition image whose blobs are
all inline could never earn a receipt: every request waited out its TTL and the
sender's sweep re-sent the same image forever, pinning the image permanently.

Added receiptBlobsComplete: it still finds every reachable blob to catch a
genuinely unreachable HAS_BLOBS claim (the same conservative case
storedBlobsAreComplete declines), but an inline blob needs no file check --
its bytes are already inside the row's own encoded value, which the caller
already confirmed durable. Only file-backed blobs still go through
blobFileMissingOrIncompleteAsync. storedBlobsAreComplete/isDurableIdentityTie
are untouched, so the duplicate-skip's conservative behavior for inline blobs
is unchanged.

New unit coverage in durableIdentityTie.test.mjs exercises real stored blobs
(inline and file-backed) through receiptBlobsComplete, the same fixture
pattern the neighboring isDurableIdentityTie file-blob tests use.

Also: correct an inbound-batch comment that overclaimed uniqueness (the sender
guarantees no repeated record per batch; a violation would only race
recordHandoffReceipt to a lower stored version, harmlessly), and fix a missing
line break in DESIGN.md that ran the Redelivery bullet into Release.

Declined, with reasons (framing recheck cleared this change at
chosen-approach-sound before implementation; both declines recorded in the PR
body's decision ledger): a same-version-different-origin receipt-proof gap
that needs a wire/capability change for a narrow, compound scenario; a
resweep-timer/wait-site nit whose cost is zero on any core that exists today
and whose fix touches a hot loop already hardened over many liveness rounds.

Dispatch-Task: pr-maint-0e208afb66ea625b57ed77ba834daf66
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
kriszyp added a commit that referenced this pull request Oct 6, 2026
… only inline blobs can earn one

Round-12 rebase review finding (post-#800 rebase onto main): settleReceiptRequests
reused storedBlobsAreComplete, the duplicate-skip verifier, to gate a receipt on
blob durability. That function only counts FILE-backed blobs and returns false
whenever none are found, including when a HAS_BLOBS row's blobs are all inline
(below core's file-storage threshold) -- so a transition image whose blobs are
all inline could never earn a receipt: every request waited out its TTL and the
sender's sweep re-sent the same image forever, pinning the image permanently.

Added receiptBlobsComplete: it still finds every reachable blob to catch a
genuinely unreachable HAS_BLOBS claim (the same conservative case
storedBlobsAreComplete declines), but an inline blob needs no file check --
its bytes are already inside the row's own encoded value, which the caller
already confirmed durable. Only file-backed blobs still go through
blobFileMissingOrIncompleteAsync. storedBlobsAreComplete/isDurableIdentityTie
are untouched, so the duplicate-skip's conservative behavior for inline blobs
is unchanged.

New unit coverage in durableIdentityTie.test.mjs exercises real stored blobs
(inline and file-backed) through receiptBlobsComplete, the same fixture
pattern the neighboring isDurableIdentityTie file-blob tests use.

Also: correct an inbound-batch comment that overclaimed uniqueness (the sender
guarantees no repeated record per batch; a violation would only race
recordHandoffReceipt to a lower stored version, harmlessly), and fix a missing
line break in DESIGN.md that ran the Redelivery bullet into Release.

Declined, with reasons (framing recheck cleared this change at
chosen-approach-sound before implementation; both declines recorded in the PR
body's decision ledger): a same-version-different-origin receipt-proof gap
that needs a wire/capability change for a narrow, compound scenario; a
resweep-timer/wait-site nit whose cost is zero on any core that exists today
and whose fix touches a hot loop already hardened over many liveness rounds.

Dispatch-Task: pr-maint-0e208afb66ea625b57ed77ba834daf66
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
github-actions Bot pushed a commit that referenced this pull request Oct 6, 2026
… only inline blobs can earn one

Round-12 rebase review finding (post-#800 rebase onto main): settleReceiptRequests
reused storedBlobsAreComplete, the duplicate-skip verifier, to gate a receipt on
blob durability. That function only counts FILE-backed blobs and returns false
whenever none are found, including when a HAS_BLOBS row's blobs are all inline
(below core's file-storage threshold) -- so a transition image whose blobs are
all inline could never earn a receipt: every request waited out its TTL and the
sender's sweep re-sent the same image forever, pinning the image permanently.

Added receiptBlobsComplete: it still finds every reachable blob to catch a
genuinely unreachable HAS_BLOBS claim (the same conservative case
storedBlobsAreComplete declines), but an inline blob needs no file check --
its bytes are already inside the row's own encoded value, which the caller
already confirmed durable. Only file-backed blobs still go through
blobFileMissingOrIncompleteAsync. storedBlobsAreComplete/isDurableIdentityTie
are untouched, so the duplicate-skip's conservative behavior for inline blobs
is unchanged.

New unit coverage in durableIdentityTie.test.mjs exercises real stored blobs
(inline and file-backed) through receiptBlobsComplete, the same fixture
pattern the neighboring isDurableIdentityTie file-blob tests use.

Also: correct an inbound-batch comment that overclaimed uniqueness (the sender
guarantees no repeated record per batch; a violation would only race
recordHandoffReceipt to a lower stored version, harmlessly), and fix a missing
line break in DESIGN.md that ran the Redelivery bullet into Release.

Declined, with reasons (framing recheck cleared this change at
chosen-approach-sound before implementation; both declines recorded in the PR
body's decision ledger): a same-version-different-origin receipt-proof gap
that needs a wire/capability change for a narrow, compound scenario; a
resweep-timer/wait-site nit whose cost is zero on any core that exists today
and whose fix touches a hot loop already hardened over many liveness rounds.

Dispatch-Task: pr-maint-0e208afb66ea625b57ed77ba834daf66
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
kriszyp added a commit that referenced this pull request Oct 6, 2026
… only inline blobs can earn one

Round-12 rebase review finding (post-#800 rebase onto main): settleReceiptRequests
reused storedBlobsAreComplete, the duplicate-skip verifier, to gate a receipt on
blob durability. That function only counts FILE-backed blobs and returns false
whenever none are found, including when a HAS_BLOBS row's blobs are all inline
(below core's file-storage threshold) -- so a transition image whose blobs are
all inline could never earn a receipt: every request waited out its TTL and the
sender's sweep re-sent the same image forever, pinning the image permanently.

Added receiptBlobsComplete: it still finds every reachable blob to catch a
genuinely unreachable HAS_BLOBS claim (the same conservative case
storedBlobsAreComplete declines), but an inline blob needs no file check --
its bytes are already inside the row's own encoded value, which the caller
already confirmed durable. Only file-backed blobs still go through
blobFileMissingOrIncompleteAsync. storedBlobsAreComplete/isDurableIdentityTie
are untouched, so the duplicate-skip's conservative behavior for inline blobs
is unchanged.

New unit coverage in durableIdentityTie.test.mjs exercises real stored blobs
(inline and file-backed) through receiptBlobsComplete, the same fixture
pattern the neighboring isDurableIdentityTie file-blob tests use.

Also: correct an inbound-batch comment that overclaimed uniqueness (the sender
guarantees no repeated record per batch; a violation would only race
recordHandoffReceipt to a lower stored version, harmlessly), and fix a missing
line break in DESIGN.md that ran the Redelivery bullet into Release.

Declined, with reasons (framing recheck cleared this change at
chosen-approach-sound before implementation; both declines recorded in the PR
body's decision ledger): a same-version-different-origin receipt-proof gap
that needs a wire/capability change for a narrow, compound scenario; a
resweep-timer/wait-site nit whose cost is zero on any core that exists today
and whose fix touches a hot loop already hardened over many liveness rounds.

Dispatch-Task: pr-maint-0e208afb66ea625b57ed77ba834daf66
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
github-actions Bot pushed a commit that referenced this pull request Oct 6, 2026
… only inline blobs can earn one

Round-12 rebase review finding (post-#800 rebase onto main): settleReceiptRequests
reused storedBlobsAreComplete, the duplicate-skip verifier, to gate a receipt on
blob durability. That function only counts FILE-backed blobs and returns false
whenever none are found, including when a HAS_BLOBS row's blobs are all inline
(below core's file-storage threshold) -- so a transition image whose blobs are
all inline could never earn a receipt: every request waited out its TTL and the
sender's sweep re-sent the same image forever, pinning the image permanently.

Added receiptBlobsComplete: it still finds every reachable blob to catch a
genuinely unreachable HAS_BLOBS claim (the same conservative case
storedBlobsAreComplete declines), but an inline blob needs no file check --
its bytes are already inside the row's own encoded value, which the caller
already confirmed durable. Only file-backed blobs still go through
blobFileMissingOrIncompleteAsync. storedBlobsAreComplete/isDurableIdentityTie
are untouched, so the duplicate-skip's conservative behavior for inline blobs
is unchanged.

New unit coverage in durableIdentityTie.test.mjs exercises real stored blobs
(inline and file-backed) through receiptBlobsComplete, the same fixture
pattern the neighboring isDurableIdentityTie file-blob tests use.

Also: correct an inbound-batch comment that overclaimed uniqueness (the sender
guarantees no repeated record per batch; a violation would only race
recordHandoffReceipt to a lower stored version, harmlessly), and fix a missing
line break in DESIGN.md that ran the Redelivery bullet into Release.

Declined, with reasons (framing recheck cleared this change at
chosen-approach-sound before implementation; both declines recorded in the PR
body's decision ledger): a same-version-different-origin receipt-proof gap
that needs a wire/capability change for a narrow, compound scenario; a
resweep-timer/wait-site nit whose cost is zero on any core that exists today
and whose fix touches a hot loop already hardened over many liveness rounds.

Dispatch-Task: pr-maint-0e208afb66ea625b57ed77ba834daf66
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
github-actions Bot pushed a commit that referenced this pull request Oct 6, 2026
… only inline blobs can earn one

Round-12 rebase review finding (post-#800 rebase onto main): settleReceiptRequests
reused storedBlobsAreComplete, the duplicate-skip verifier, to gate a receipt on
blob durability. That function only counts FILE-backed blobs and returns false
whenever none are found, including when a HAS_BLOBS row's blobs are all inline
(below core's file-storage threshold) -- so a transition image whose blobs are
all inline could never earn a receipt: every request waited out its TTL and the
sender's sweep re-sent the same image forever, pinning the image permanently.

Added receiptBlobsComplete: it still finds every reachable blob to catch a
genuinely unreachable HAS_BLOBS claim (the same conservative case
storedBlobsAreComplete declines), but an inline blob needs no file check --
its bytes are already inside the row's own encoded value, which the caller
already confirmed durable. Only file-backed blobs still go through
blobFileMissingOrIncompleteAsync. storedBlobsAreComplete/isDurableIdentityTie
are untouched, so the duplicate-skip's conservative behavior for inline blobs
is unchanged.

New unit coverage in durableIdentityTie.test.mjs exercises real stored blobs
(inline and file-backed) through receiptBlobsComplete, the same fixture
pattern the neighboring isDurableIdentityTie file-blob tests use.

Also: correct an inbound-batch comment that overclaimed uniqueness (the sender
guarantees no repeated record per batch; a violation would only race
recordHandoffReceipt to a lower stored version, harmlessly), and fix a missing
line break in DESIGN.md that ran the Redelivery bullet into Release.

Declined, with reasons (framing recheck cleared this change at
chosen-approach-sound before implementation; both declines recorded in the PR
body's decision ledger): a same-version-different-origin receipt-proof gap
that needs a wire/capability change for a narrow, compound scenario; a
resweep-timer/wait-site nit whose cost is zero on any core that exists today
and whose fix touches a hot loop already hardened over many liveness rounds.

Dispatch-Task: pr-maint-0e208afb66ea625b57ed77ba834daf66
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
github-actions Bot pushed a commit that referenced this pull request Oct 6, 2026
… only inline blobs can earn one

Round-12 rebase review finding (post-#800 rebase onto main): settleReceiptRequests
reused storedBlobsAreComplete, the duplicate-skip verifier, to gate a receipt on
blob durability. That function only counts FILE-backed blobs and returns false
whenever none are found, including when a HAS_BLOBS row's blobs are all inline
(below core's file-storage threshold) -- so a transition image whose blobs are
all inline could never earn a receipt: every request waited out its TTL and the
sender's sweep re-sent the same image forever, pinning the image permanently.

Added receiptBlobsComplete: it still finds every reachable blob to catch a
genuinely unreachable HAS_BLOBS claim (the same conservative case
storedBlobsAreComplete declines), but an inline blob needs no file check --
its bytes are already inside the row's own encoded value, which the caller
already confirmed durable. Only file-backed blobs still go through
blobFileMissingOrIncompleteAsync. storedBlobsAreComplete/isDurableIdentityTie
are untouched, so the duplicate-skip's conservative behavior for inline blobs
is unchanged.

New unit coverage in durableIdentityTie.test.mjs exercises real stored blobs
(inline and file-backed) through receiptBlobsComplete, the same fixture
pattern the neighboring isDurableIdentityTie file-blob tests use.

Also: correct an inbound-batch comment that overclaimed uniqueness (the sender
guarantees no repeated record per batch; a violation would only race
recordHandoffReceipt to a lower stored version, harmlessly), and fix a missing
line break in DESIGN.md that ran the Redelivery bullet into Release.

Declined, with reasons (framing recheck cleared this change at
chosen-approach-sound before implementation; both declines recorded in the PR
body's decision ledger): a same-version-different-origin receipt-proof gap
that needs a wire/capability change for a narrow, compound scenario; a
resweep-timer/wait-site nit whose cost is zero on any core that exists today
and whose fix touches a hot loop already hardened over many liveness rounds.

Dispatch-Task: pr-maint-0e208afb66ea625b57ed77ba834daf66
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
kriszyp added a commit that referenced this pull request Oct 7, 2026
… only inline blobs can earn one

Round-12 rebase review finding (post-#800 rebase onto main): settleReceiptRequests
reused storedBlobsAreComplete, the duplicate-skip verifier, to gate a receipt on
blob durability. That function only counts FILE-backed blobs and returns false
whenever none are found, including when a HAS_BLOBS row's blobs are all inline
(below core's file-storage threshold) -- so a transition image whose blobs are
all inline could never earn a receipt: every request waited out its TTL and the
sender's sweep re-sent the same image forever, pinning the image permanently.

Added receiptBlobsComplete: it still finds every reachable blob to catch a
genuinely unreachable HAS_BLOBS claim (the same conservative case
storedBlobsAreComplete declines), but an inline blob needs no file check --
its bytes are already inside the row's own encoded value, which the caller
already confirmed durable. Only file-backed blobs still go through
blobFileMissingOrIncompleteAsync. storedBlobsAreComplete/isDurableIdentityTie
are untouched, so the duplicate-skip's conservative behavior for inline blobs
is unchanged.

New unit coverage in durableIdentityTie.test.mjs exercises real stored blobs
(inline and file-backed) through receiptBlobsComplete, the same fixture
pattern the neighboring isDurableIdentityTie file-blob tests use.

Also: correct an inbound-batch comment that overclaimed uniqueness (the sender
guarantees no repeated record per batch; a violation would only race
recordHandoffReceipt to a lower stored version, harmlessly), and fix a missing
line break in DESIGN.md that ran the Redelivery bullet into Release.

Declined, with reasons (framing recheck cleared this change at
chosen-approach-sound before implementation; both declines recorded in the PR
body's decision ledger): a same-version-different-origin receipt-proof gap
that needs a wire/capability change for a narrow, compound scenario; a
resweep-timer/wait-site nit whose cost is zero on any core that exists today
and whose fix touches a hot loop already hardened over many liveness rounds.

Dispatch-Task: pr-maint-0e208afb66ea625b57ed77ba834daf66
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
github-actions Bot pushed a commit that referenced this pull request Oct 7, 2026
… only inline blobs can earn one

Round-12 rebase review finding (post-#800 rebase onto main): settleReceiptRequests
reused storedBlobsAreComplete, the duplicate-skip verifier, to gate a receipt on
blob durability. That function only counts FILE-backed blobs and returns false
whenever none are found, including when a HAS_BLOBS row's blobs are all inline
(below core's file-storage threshold) -- so a transition image whose blobs are
all inline could never earn a receipt: every request waited out its TTL and the
sender's sweep re-sent the same image forever, pinning the image permanently.

Added receiptBlobsComplete: it still finds every reachable blob to catch a
genuinely unreachable HAS_BLOBS claim (the same conservative case
storedBlobsAreComplete declines), but an inline blob needs no file check --
its bytes are already inside the row's own encoded value, which the caller
already confirmed durable. Only file-backed blobs still go through
blobFileMissingOrIncompleteAsync. storedBlobsAreComplete/isDurableIdentityTie
are untouched, so the duplicate-skip's conservative behavior for inline blobs
is unchanged.

New unit coverage in durableIdentityTie.test.mjs exercises real stored blobs
(inline and file-backed) through receiptBlobsComplete, the same fixture
pattern the neighboring isDurableIdentityTie file-blob tests use.

Also: correct an inbound-batch comment that overclaimed uniqueness (the sender
guarantees no repeated record per batch; a violation would only race
recordHandoffReceipt to a lower stored version, harmlessly), and fix a missing
line break in DESIGN.md that ran the Redelivery bullet into Release.

Declined, with reasons (framing recheck cleared this change at
chosen-approach-sound before implementation; both declines recorded in the PR
body's decision ledger): a same-version-different-origin receipt-proof gap
that needs a wire/capability change for a narrow, compound scenario; a
resweep-timer/wait-site nit whose cost is zero on any core that exists today
and whose fix touches a hot loop already hardened over many liveness rounds.

Dispatch-Task: pr-maint-0e208afb66ea625b57ed77ba834daf66
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
kriszyp added a commit that referenced this pull request Oct 7, 2026
… only inline blobs can earn one

Round-12 rebase review finding (post-#800 rebase onto main): settleReceiptRequests
reused storedBlobsAreComplete, the duplicate-skip verifier, to gate a receipt on
blob durability. That function only counts FILE-backed blobs and returns false
whenever none are found, including when a HAS_BLOBS row's blobs are all inline
(below core's file-storage threshold) -- so a transition image whose blobs are
all inline could never earn a receipt: every request waited out its TTL and the
sender's sweep re-sent the same image forever, pinning the image permanently.

Added receiptBlobsComplete: it still finds every reachable blob to catch a
genuinely unreachable HAS_BLOBS claim (the same conservative case
storedBlobsAreComplete declines), but an inline blob needs no file check --
its bytes are already inside the row's own encoded value, which the caller
already confirmed durable. Only file-backed blobs still go through
blobFileMissingOrIncompleteAsync. storedBlobsAreComplete/isDurableIdentityTie
are untouched, so the duplicate-skip's conservative behavior for inline blobs
is unchanged.

New unit coverage in durableIdentityTie.test.mjs exercises real stored blobs
(inline and file-backed) through receiptBlobsComplete, the same fixture
pattern the neighboring isDurableIdentityTie file-blob tests use.

Also: correct an inbound-batch comment that overclaimed uniqueness (the sender
guarantees no repeated record per batch; a violation would only race
recordHandoffReceipt to a lower stored version, harmlessly), and fix a missing
line break in DESIGN.md that ran the Redelivery bullet into Release.

Declined, with reasons (framing recheck cleared this change at
chosen-approach-sound before implementation; both declines recorded in the PR
body's decision ledger): a same-version-different-origin receipt-proof gap
that needs a wire/capability change for a narrow, compound scenario; a
resweep-timer/wait-site nit whose cost is zero on any core that exists today
and whose fix touches a hot loop already hardened over many liveness rounds.

Dispatch-Task: pr-maint-0e208afb66ea625b57ed77ba834daf66
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
github-actions Bot pushed a commit that referenced this pull request Oct 7, 2026
… only inline blobs can earn one

Round-12 rebase review finding (post-#800 rebase onto main): settleReceiptRequests
reused storedBlobsAreComplete, the duplicate-skip verifier, to gate a receipt on
blob durability. That function only counts FILE-backed blobs and returns false
whenever none are found, including when a HAS_BLOBS row's blobs are all inline
(below core's file-storage threshold) -- so a transition image whose blobs are
all inline could never earn a receipt: every request waited out its TTL and the
sender's sweep re-sent the same image forever, pinning the image permanently.

Added receiptBlobsComplete: it still finds every reachable blob to catch a
genuinely unreachable HAS_BLOBS claim (the same conservative case
storedBlobsAreComplete declines), but an inline blob needs no file check --
its bytes are already inside the row's own encoded value, which the caller
already confirmed durable. Only file-backed blobs still go through
blobFileMissingOrIncompleteAsync. storedBlobsAreComplete/isDurableIdentityTie
are untouched, so the duplicate-skip's conservative behavior for inline blobs
is unchanged.

New unit coverage in durableIdentityTie.test.mjs exercises real stored blobs
(inline and file-backed) through receiptBlobsComplete, the same fixture
pattern the neighboring isDurableIdentityTie file-blob tests use.

Also: correct an inbound-batch comment that overclaimed uniqueness (the sender
guarantees no repeated record per batch; a violation would only race
recordHandoffReceipt to a lower stored version, harmlessly), and fix a missing
line break in DESIGN.md that ran the Redelivery bullet into Release.

Declined, with reasons (framing recheck cleared this change at
chosen-approach-sound before implementation; both declines recorded in the PR
body's decision ledger): a same-version-different-origin receipt-proof gap
that needs a wire/capability change for a narrow, compound scenario; a
resweep-timer/wait-site nit whose cost is zero on any core that exists today
and whose fix touches a hot loop already hardened over many liveness rounds.

Dispatch-Task: pr-maint-0e208afb66ea625b57ed77ba834daf66
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
kriszyp added a commit that referenced this pull request Oct 9, 2026
… only inline blobs can earn one

Round-12 rebase review finding (post-#800 rebase onto main): settleReceiptRequests
reused storedBlobsAreComplete, the duplicate-skip verifier, to gate a receipt on
blob durability. That function only counts FILE-backed blobs and returns false
whenever none are found, including when a HAS_BLOBS row's blobs are all inline
(below core's file-storage threshold) -- so a transition image whose blobs are
all inline could never earn a receipt: every request waited out its TTL and the
sender's sweep re-sent the same image forever, pinning the image permanently.

Added receiptBlobsComplete: it still finds every reachable blob to catch a
genuinely unreachable HAS_BLOBS claim (the same conservative case
storedBlobsAreComplete declines), but an inline blob needs no file check --
its bytes are already inside the row's own encoded value, which the caller
already confirmed durable. Only file-backed blobs still go through
blobFileMissingOrIncompleteAsync. storedBlobsAreComplete/isDurableIdentityTie
are untouched, so the duplicate-skip's conservative behavior for inline blobs
is unchanged.

New unit coverage in durableIdentityTie.test.mjs exercises real stored blobs
(inline and file-backed) through receiptBlobsComplete, the same fixture
pattern the neighboring isDurableIdentityTie file-blob tests use.

Also: correct an inbound-batch comment that overclaimed uniqueness (the sender
guarantees no repeated record per batch; a violation would only race
recordHandoffReceipt to a lower stored version, harmlessly), and fix a missing
line break in DESIGN.md that ran the Redelivery bullet into Release.

Declined, with reasons (framing recheck cleared this change at
chosen-approach-sound before implementation; both declines recorded in the PR
body's decision ledger): a same-version-different-origin receipt-proof gap
that needs a wire/capability change for a narrow, compound scenario; a
resweep-timer/wait-site nit whose cost is zero on any core that exists today
and whose fix touches a hot loop already hardened over many liveness rounds.

Dispatch-Task: pr-maint-0e208afb66ea625b57ed77ba834daf66
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
github-actions Bot pushed a commit that referenced this pull request Oct 9, 2026
… only inline blobs can earn one

Round-12 rebase review finding (post-#800 rebase onto main): settleReceiptRequests
reused storedBlobsAreComplete, the duplicate-skip verifier, to gate a receipt on
blob durability. That function only counts FILE-backed blobs and returns false
whenever none are found, including when a HAS_BLOBS row's blobs are all inline
(below core's file-storage threshold) -- so a transition image whose blobs are
all inline could never earn a receipt: every request waited out its TTL and the
sender's sweep re-sent the same image forever, pinning the image permanently.

Added receiptBlobsComplete: it still finds every reachable blob to catch a
genuinely unreachable HAS_BLOBS claim (the same conservative case
storedBlobsAreComplete declines), but an inline blob needs no file check --
its bytes are already inside the row's own encoded value, which the caller
already confirmed durable. Only file-backed blobs still go through
blobFileMissingOrIncompleteAsync. storedBlobsAreComplete/isDurableIdentityTie
are untouched, so the duplicate-skip's conservative behavior for inline blobs
is unchanged.

New unit coverage in durableIdentityTie.test.mjs exercises real stored blobs
(inline and file-backed) through receiptBlobsComplete, the same fixture
pattern the neighboring isDurableIdentityTie file-blob tests use.

Also: correct an inbound-batch comment that overclaimed uniqueness (the sender
guarantees no repeated record per batch; a violation would only race
recordHandoffReceipt to a lower stored version, harmlessly), and fix a missing
line break in DESIGN.md that ran the Redelivery bullet into Release.

Declined, with reasons (framing recheck cleared this change at
chosen-approach-sound before implementation; both declines recorded in the PR
body's decision ledger): a same-version-different-origin receipt-proof gap
that needs a wire/capability change for a narrow, compound scenario; a
resweep-timer/wait-site nit whose cost is zero on any core that exists today
and whose fix touches a hot loop already hardened over many liveness rounds.

Dispatch-Task: pr-maint-0e208afb66ea625b57ed77ba834daf66
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
kriszyp added a commit that referenced this pull request Oct 10, 2026
… only inline blobs can earn one

Round-12 rebase review finding (post-#800 rebase onto main): settleReceiptRequests
reused storedBlobsAreComplete, the duplicate-skip verifier, to gate a receipt on
blob durability. That function only counts FILE-backed blobs and returns false
whenever none are found, including when a HAS_BLOBS row's blobs are all inline
(below core's file-storage threshold) -- so a transition image whose blobs are
all inline could never earn a receipt: every request waited out its TTL and the
sender's sweep re-sent the same image forever, pinning the image permanently.

Added receiptBlobsComplete: it still finds every reachable blob to catch a
genuinely unreachable HAS_BLOBS claim (the same conservative case
storedBlobsAreComplete declines), but an inline blob needs no file check --
its bytes are already inside the row's own encoded value, which the caller
already confirmed durable. Only file-backed blobs still go through
blobFileMissingOrIncompleteAsync. storedBlobsAreComplete/isDurableIdentityTie
are untouched, so the duplicate-skip's conservative behavior for inline blobs
is unchanged.

New unit coverage in durableIdentityTie.test.mjs exercises real stored blobs
(inline and file-backed) through receiptBlobsComplete, the same fixture
pattern the neighboring isDurableIdentityTie file-blob tests use.

Also: correct an inbound-batch comment that overclaimed uniqueness (the sender
guarantees no repeated record per batch; a violation would only race
recordHandoffReceipt to a lower stored version, harmlessly), and fix a missing
line break in DESIGN.md that ran the Redelivery bullet into Release.

Declined, with reasons (framing recheck cleared this change at
chosen-approach-sound before implementation; both declines recorded in the PR
body's decision ledger): a same-version-different-origin receipt-proof gap
that needs a wire/capability change for a narrow, compound scenario; a
resweep-timer/wait-site nit whose cost is zero on any core that exists today
and whose fix touches a hot loop already hardened over many liveness rounds.

Dispatch-Task: pr-maint-0e208afb66ea625b57ed77ba834daf66
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
github-actions Bot pushed a commit that referenced this pull request Oct 10, 2026
… only inline blobs can earn one

Round-12 rebase review finding (post-#800 rebase onto main): settleReceiptRequests
reused storedBlobsAreComplete, the duplicate-skip verifier, to gate a receipt on
blob durability. That function only counts FILE-backed blobs and returns false
whenever none are found, including when a HAS_BLOBS row's blobs are all inline
(below core's file-storage threshold) -- so a transition image whose blobs are
all inline could never earn a receipt: every request waited out its TTL and the
sender's sweep re-sent the same image forever, pinning the image permanently.

Added receiptBlobsComplete: it still finds every reachable blob to catch a
genuinely unreachable HAS_BLOBS claim (the same conservative case
storedBlobsAreComplete declines), but an inline blob needs no file check --
its bytes are already inside the row's own encoded value, which the caller
already confirmed durable. Only file-backed blobs still go through
blobFileMissingOrIncompleteAsync. storedBlobsAreComplete/isDurableIdentityTie
are untouched, so the duplicate-skip's conservative behavior for inline blobs
is unchanged.

New unit coverage in durableIdentityTie.test.mjs exercises real stored blobs
(inline and file-backed) through receiptBlobsComplete, the same fixture
pattern the neighboring isDurableIdentityTie file-blob tests use.

Also: correct an inbound-batch comment that overclaimed uniqueness (the sender
guarantees no repeated record per batch; a violation would only race
recordHandoffReceipt to a lower stored version, harmlessly), and fix a missing
line break in DESIGN.md that ran the Redelivery bullet into Release.

Declined, with reasons (framing recheck cleared this change at
chosen-approach-sound before implementation; both declines recorded in the PR
body's decision ledger): a same-version-different-origin receipt-proof gap
that needs a wire/capability change for a narrow, compound scenario; a
resweep-timer/wait-site nit whose cost is zero on any core that exists today
and whose fix touches a hot loop already hardened over many liveness rounds.

Dispatch-Task: pr-maint-0e208afb66ea625b57ed77ba834daf66
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
github-actions Bot pushed a commit that referenced this pull request Oct 10, 2026
… only inline blobs can earn one

Round-12 rebase review finding (post-#800 rebase onto main): settleReceiptRequests
reused storedBlobsAreComplete, the duplicate-skip verifier, to gate a receipt on
blob durability. That function only counts FILE-backed blobs and returns false
whenever none are found, including when a HAS_BLOBS row's blobs are all inline
(below core's file-storage threshold) -- so a transition image whose blobs are
all inline could never earn a receipt: every request waited out its TTL and the
sender's sweep re-sent the same image forever, pinning the image permanently.

Added receiptBlobsComplete: it still finds every reachable blob to catch a
genuinely unreachable HAS_BLOBS claim (the same conservative case
storedBlobsAreComplete declines), but an inline blob needs no file check --
its bytes are already inside the row's own encoded value, which the caller
already confirmed durable. Only file-backed blobs still go through
blobFileMissingOrIncompleteAsync. storedBlobsAreComplete/isDurableIdentityTie
are untouched, so the duplicate-skip's conservative behavior for inline blobs
is unchanged.

New unit coverage in durableIdentityTie.test.mjs exercises real stored blobs
(inline and file-backed) through receiptBlobsComplete, the same fixture
pattern the neighboring isDurableIdentityTie file-blob tests use.

Also: correct an inbound-batch comment that overclaimed uniqueness (the sender
guarantees no repeated record per batch; a violation would only race
recordHandoffReceipt to a lower stored version, harmlessly), and fix a missing
line break in DESIGN.md that ran the Redelivery bullet into Release.

Declined, with reasons (framing recheck cleared this change at
chosen-approach-sound before implementation; both declines recorded in the PR
body's decision ledger): a same-version-different-origin receipt-proof gap
that needs a wire/capability change for a narrow, compound scenario; a
resweep-timer/wait-site nit whose cost is zero on any core that exists today
and whose fix touches a hot loop already hardened over many liveness rounds.

Dispatch-Task: pr-maint-0e208afb66ea625b57ed77ba834daf66
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Subscription-setup stagger is added on top of independent jitter draws, so the #446 spacing does not hold on stale-worker reassignment

2 participants