Repository navigation
Pace replication retries on one jittered schedule and bound subscription setup - #800
Merged
Merged
Conversation
Replication had no jitter anywhere in production code and a different retry policy at every site, from a flat 200ms subscription-setup delay to jitterless exponentials. The subscription-setup site had no dedup, no cap, and a non-unref'd timer, so whatever re-drove onNodeUpdate amplified 1:1 into main-thread timers, worker-side WebSocket/TLS setup, and "Setting up subscription with leader" warns — ~1,400 lines/s/node in the field, ending in an OOM kill (harper-pro#327). Add replication/backoff.ts: one createBackoff schedule with an exponential ceiling, full jitter, an optional per-site floor, an optional wall-clock budget, and injectable RNG/clock. Adopt it at the subscription-setup scheduler, scheduleReconnect, the wedge and receive-stall re-drives, the hdb_nodes watcher restart, the send-auth reprobe, the clone JWT/version retries, and the blob-repair sweep. The subscription-setup scheduler additionally enforces at most one pending setup per (peer URL, database), keyed in its own map because the stale-worker path deletes and recreates the connectionReplicationMap entry. Its dispatch re-reads the live entry at fire time rather than capturing a payload, so a deduped update is never lost and no closure is allocated per suppressed event, and the setup is cancelled on connect, unsubscribe, and node deletion. Also fix the hdb_nodes watcher's premature backoff reset: the success marker was set the instant subscribe() resolved, so a subscribe-then- immediately-throw cycle reset the delay on every pass and never escalated. It now resets only after an iteration survives a healthy- uptime threshold. Refs #327 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01LHXY4wfz8qJ4ryA5gKPAme
The deduped scheduling path read `entry.nodes` at fire time, but `onDatabase` replaces that array on its early-return path *without* running the leader/url enrichment — so a node update arriving between arming and firing left the setup posting a request with no url, and the worker logged "Failed to create web socket to undefined" and never connected (caught by selectiveTableSubscription.test.mjs). The schedule now carries the payload of the call that armed or last refreshed it; only calls that did enrich reach the scheduler, so "newest payload wins" holds without the clobber. Also resolve `nodes[0].url` where the payload is built rather than at the scheduling site, so the wedge re-drive — which posts from `entry.nodes` — cannot inherit the same urlless array. Refs #327 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01LHXY4wfz8qJ4ryA5gKPAme
…covery timers From the independent pre-push review of the previous two commits: - send-auth reprobe (blocker): checking the budget only at the top of the loop let a row that became decodable *after* an event-loop stall pushed us past the 30s deadline still authorize the peer — the loop exits on a non-UNCHANGED row without consulting the deadline. It now fails closed on the elapsed time before trusting the next read. The backoff is also built lazily, so an ordinary decodable authorization event allocates nothing and never reads the clock. - scheduleReconnect: a fixed 100ms floor under full jitter did not preserve the #339 dial-rate guard, which is ceiling-relative by nature. Added `jitter: 'equal'` to the utility (draw from the top half of the window) and used it here, so the minimum dial interval stays proportional as the ceiling escalates. Full jitter stays the default everywhere the invariant is latency rather than rate. - subscription setup: `onDatabase`'s early-return path builds a fresh payload but never reaches the scheduler, so an armed setup could still dispatch pre-update routing state. It now calls `refreshPending`, and carries `isLeader` forward onto the replacement array so neither that path nor the wedge re-drive loses the leader decision. - wedge / receive-stall re-drives: each entry now owns a single `reDriveTimer` that a later decision replaces and connect/unsubscribe/ delete disarm, so a sweep staggered across many databases cannot stack waves or fire after an unsubscribe. The stall kick additionally re-reads the receive watermark at fire time, so a copy that resumed during the delay is not reconnected. - a same-name node moving to a new URL now cancels any setup armed for the address it left. Refs #327 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01LHXY4wfz8qJ4ryA5gKPAme
Bound worker-side pre-readiness subscription work, retain self-catchup state until dispatch, and make exhausted backoffs fail closed. Align reconnect and watcher jitter with the full-jitter policy and extend regression coverage. Co-Authored-By: GPT-5 Codex <noreply@openai.com>
Draw the wedge/stall re-drive jitter once per reconcile sweep instead of once per entry: a per-entry draw varied consecutive delays by up to ±200ms, letting several dials share a 50ms instant and weakening the concurrency bound RECONNECT_STAGGER_MS exists to hold (#446). Only the worker that currently owns an entry may retire its pending setup on connect, so a stale worker's still-open connection cannot cancel the setup armed for the entry's replacement. End a blob-repair sweep once the schedule is exhausted rather than pacing an unrepairable backlog at 1s/record with the cursor open, and record why the subscription connection-key helper keeps its two teardown callers' argument order. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01HEgBi5zh8QKMbFmeGwFv7w
A component-readiness rejection retained the pre-readiness subscribe/unsubscribe actions but never re-attempted them, so nothing applied until another parentPort message arrived or the wedge reconcile noticed ~30s later — the empty-subscription window this gate exists to close. Re-attempt on the same capped, jittered, unref'd schedule the rest of the discipline uses, resetting once readiness succeeds. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01HEgBi5zh8QKMbFmeGwFv7w
…ownership - shouldCloseSendAuthWatch: the row read is synchronous and can itself carry the clock past the advertised grace period, so re-check the deadline after it and fail closed. - repairBlobs: bound how long a failure run may spend *pausing* instead of ending the sweep on a consecutive-failure count — a cluster-wide-lost prefix must not make the repairable records behind it permanently unreachable through the operation. Drop the per-thread single-flight guard with it: `getRecord()` has no timeout, so one wedged peer would hold the flag for the life of the thread and remove the operator's lever. - runNodeUpdateWatcher: stamp the health clock once the subscription is live, so a subscribe() that blocks past the threshold and then throws no longer reads as uptime. - Worker admission: own the readiness re-attempt so a pre-readiness burst against a rejecting import cannot multiply timers, imports and warn lines 1:1 with messages. - Connect ownership: tag the report with the sending thread id rather than relying on core's undocumented onMessageByType listener arity, behind a tested pure helper. - Disarm an entry's recovery timer when its worker exits or the entry is replaced, and adopt the shared schedule for the copy-cursor flush retry. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01HEgBi5zh8QKMbFmeGwFv7w
The backoff table described the blob-repair cutoff the previous commit replaced and listed neither the copy-cursor flush retry nor the worker readiness re-attempt; the exclusions paragraph now also names the copy-finalize timeout bound, which is a wait bound rather than a retry and must stay unjittered. Give the readiness re-attempt the same floor every other adopting site has, so a zero draw cannot re-probe on the next macrotask. Sample the per-record blob-repair warn once the sweep goes unpaced, and put clearWorkerFromEntries' description back over clearWorkerFromEntries. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01HEgBi5zh8QKMbFmeGwFv7w
Gating the connect edge on "is this the entry's current worker" is not answerable on the main thread: connectToNextWorker subscribes a failover peer on a worker that is not entry.worker, and a superseded worker's hung-but-open connection reports for the same (url, database) as its replacement. Gating on it meant the escalated setup delay stopped resetting, and connectedBitRestartChurn's chaos cycles converged in 5.3s, 12.0s, then not within 25s — a backoff compounding across reconnects. Reset the delay unconditionally and leave the armed setup and recovery timers to their own fire-time guards, which is what this path did before the schedule existed. That suite is green again. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01HEgBi5zh8QKMbFmeGwFv7w
Removing the connect-edge cancellation left the receive-stall kick with no guard that observes a reconnect: entry identity holds, receiveStallReconnectAt is untouched by connectedToNode, and the watermark cannot move on a socket that just opened — so a leg that drops and reconnects inside the kick's stagger window was force-reconnected on the strength of the old socket's watermark. Claim it through a connect generation, the way the wedge kick claims its entry through disconnectedAt (which a stalled, connected:true entry does not have). The design table, its prose, and dispatchSubscribeSetup's JSDoc still described the cancel-on-connect behavior that commit removed. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01HEgBi5zh8QKMbFmeGwFv7w
entry.receiveStallReconnectAt both claims the armed kick and throttles re-detection, which needs lastReceivedTime past the stamp. Skipping the kick because the leg reconnected inside the stagger window therefore spent an epoch on a kick that never happened: if the fresh socket stalled too it never moved the watermark past the stamp and the receive-stall net never re-armed for that (peer, database) again. The decision moves into shouldFireStallKick — a pure helper with the arm/bump/fire cases under test, which is what the guard shipped without. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01HEgBi5zh8QKMbFmeGwFv7w
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01HEgBi5zh8QKMbFmeGwFv7w
There was a problem hiding this comment.
Code Review
This pull request introduces a unified, jittered backoff discipline across the replication system to mitigate thundering herd issues and prevent OOM conditions during retries. Key changes include the addition of a generic createBackoff utility, a SubscribeSetupScheduler to dedup and pace subscription setups, and a createWorkerSubscriptionAdmission gate to manage worker readiness. The changes are well-documented in DESIGN.md and supported by extensive new unit tests. I have kept the review comment regarding Map lookup optimization in subscriptionManager.ts as it provides a valid improvement opportunity.
Addresses the Gemini review comment on #800. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01HEgBi5zh8QKMbFmeGwFv7w
Contributor
|
Reviewed; no blockers found. |
cb1kenobi
reviewed
Sep 2, 2026
This was referenced Sep 3, 2026
Recompute effective leader state before subscription reuse and publish route reloads atomically. Scope stall generations and delayed kicks to the worker that owns the monitored connection. Co-Authored-By: GPT-5 Codex <noreply@openai.com>
Retain the self-catchup rider across recovery dispatches and chain jittered setup delays through a shared floor so reassignment handshakes remain spaced. Use deterministic injected timers for scheduler coverage. Co-Authored-By: GPT-5 Codex <noreply@openai.com>
Co-Authored-By: GPT-5 Codex <noreply@openai.com>
… the orchestrator A worker can carry both an entry's primary connection and a proxied failover subscription, so a socket-open report now carries the subscription half of its connection key; only the owning worker's primary open advances connectGeneration and retires the main-thread self-catchup rider. An armed setup whose worker exited before it fired is deferred to the stale-worker reconcile instead of subscribing on the main thread, and that reconcile now reassigns stale entries even when a sibling database is wedged. Co-Authored-By: GPT-5 Codex <noreply@openai.com> Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017CTUbhGiW2AXhogHZLg8eo Dispatch-Task: fix-kriszyp_harper-pro_800-16768da5
…rm-backoff Conflict resolutions that change behaviour on either side: - Socket-open reports use main's `threadId` and `newSocket` fields plus this branch's `subscriptionUrl`; the branch's duplicate `reportingThreadId`/`opened` fields are gone. A truth up-correction advances the connect generation directly in the reconcile, since it must not carry `newSocket`. - Scheduled subscription setups carry main's send-time `exclusionOrigins`. - The unsubscribe teardown keeps main's retire-before-close order on the shared connection key helper. - This branch's DESIGN note on atomic route publication is renumbered 23 after main's 18-22. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017CTUbhGiW2AXhogHZLg8eo Dispatch-Task: fix-kriszyp_harper-pro_800-16768da5
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017CTUbhGiW2AXhogHZLg8eo Dispatch-Task: fix-kriszyp_harper-pro_800-16768da5
…hup to a replacement worker The stale-worker sweep chained its spacing floor through every setup, so one pair whose own backoff had escalated to ~30 s dragged every later database in the sweep out to its ceiling. The scheduler now owns the sweep rule: a setup chains the floor only when it lands inside the fresh-draw window above it. The self-catchup rider was dropped from the entry when the owning primary connection opened, so a worker that exited mid-catchup took the range with it. The entry now keeps the rider and records which one the current worker opened with: recovery re-drives to that worker still skip it, a replacement worker is sent it again. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017CTUbhGiW2AXhogHZLg8eo Dispatch-Task: fix-kriszyp_harper-pro_800-16768da5
kriszyp
added a commit
that referenced
this pull request
Sep 29, 2026
… only inline blobs can earn one Round-12 rebase review finding (post-#800 rebase onto main): settleReceiptRequests reused storedBlobsAreComplete, the duplicate-skip verifier, to gate a receipt on blob durability. That function only counts FILE-backed blobs and returns false whenever none are found, including when a HAS_BLOBS row's blobs are all inline (below core's file-storage threshold) -- so a transition image whose blobs are all inline could never earn a receipt: every request waited out its TTL and the sender's sweep re-sent the same image forever, pinning the image permanently. Added receiptBlobsComplete: it still finds every reachable blob to catch a genuinely unreachable HAS_BLOBS claim (the same conservative case storedBlobsAreComplete declines), but an inline blob needs no file check -- its bytes are already inside the row's own encoded value, which the caller already confirmed durable. Only file-backed blobs still go through blobFileMissingOrIncompleteAsync. storedBlobsAreComplete/isDurableIdentityTie are untouched, so the duplicate-skip's conservative behavior for inline blobs is unchanged. New unit coverage in durableIdentityTie.test.mjs exercises real stored blobs (inline and file-backed) through receiptBlobsComplete, the same fixture pattern the neighboring isDurableIdentityTie file-blob tests use. Also: correct an inbound-batch comment that overclaimed uniqueness (the sender guarantees no repeated record per batch; a violation would only race recordHandoffReceipt to a lower stored version, harmlessly), and fix a missing line break in DESIGN.md that ran the Redelivery bullet into Release. Declined, with reasons (framing recheck cleared this change at chosen-approach-sound before implementation; both declines recorded in the PR body's decision ledger): a same-version-different-origin receipt-proof gap that needs a wire/capability change for a narrow, compound scenario; a resweep-timer/wait-site nit whose cost is zero on any core that exists today and whose fix touches a hot loop already hardened over many liveness rounds. Dispatch-Task: pr-maint-0e208afb66ea625b57ed77ba834daf66 Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
cb1kenobi
added a commit
that referenced
this pull request
Sep 29, 2026
One content conflict, in replication/DESIGN.md: both sides appended a note numbered 23 and a row to the "where is X" table. Kept main's note 23 (config-route reload publishes atomically), renumbered this branch's to 24, and merged both cheat-sheet tables so main's three new rows and this branch's replicate:false row all survive. Checked rather than assumed, since #800 touches the same subsystem: main's replicationConnection.ts hunks are the copy-flush retry backoff, a send-auth log constant and the sendBlobs read-retry backoff, none of which touch the outbound table gate, the GET_RECORD refusal, the schema builders or the receive drop. The core bump still exposes the three APIs this branch reads — schemaDescribe's replicate field, HAS_BLOBS and Table.replicate. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
This was referenced Sep 29, 2026
kriszyp
added a commit
that referenced
this pull request
Sep 30, 2026
…ress, so it no longer carries across outages (#948) * Reset a receiving leg's reconnect backoff on durable progress, not only on a sent frame An outbound replication leg's reconnect backoff reset only when that leg sent a transaction frame. In a two-way mesh the peer serves our subscription from its own server session, so a subscriber's outbound leg never sends one: its ceiling carried across unrelated outages until every reconnect waited out the 30 s cap. #800's full jitter reached the cap within three outages, which is how connectedBitRestartChurn's chaos cycle 3 started missing its 25 s bound. The receiving side now credits its durable watermark where it advances (a commit, and the last in-flight blob's drain) through NodeReplicationConnection.onDurableProgress. Only an advance past the highest value the connection has already credited counts, and only from the session that still owns the socket, so a session that re-receives the same undurable frames after every reconnect keeps escalating (harper-pro#339). The churn test now asserts that every database logs a fresh disconnect in every cycle, which only a reset backoff does. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Dispatch-Task: main-red-kriszyp_harper-pro_039c07d6c_f0efc59e * Credit receive progress only when it is durable and the leg's own sends work From the pre-push review. Copy-apply rows are not durable until the copy's flush, so the watermark they move is not progress yet. A sender-loop failure on a leg that also serves a subscription now vetoes receive credit until a frame is sent again, so a leg whose own sends fail every time (#713) does not redial at the floor on the strength of what it received. A non-finite or out-of-range sequence can no longer pin or disable the credited maximum, and the send-side reset now checks socket ownership the way the receive side does. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Dispatch-Task: main-red-kriszyp_harper-pro_039c07d6c_f0efc59e * Clear the send-failure veto once the sender catches up, and credit a finished copy From the second pre-push review round. The veto previously cleared only on a sent frame, so a leg whose sender threw once and then had nothing to send kept escalating across outages. It now also clears when the sender loop reaches its idle wait, which a loop stuck on an oversized frame never does. A copy-apply copy's completion now counts as progress, since its final flush has run by the time copy mode ends. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Dispatch-Task: main-red-kriszyp_harper-pro_039c07d6c_f0efc59e * Apply receive progress the send-failure veto held back once the sender catches up From the third pre-push review round: an advance credited during the veto was dropped, so a leg whose whole catch-up landed before its sender caught up entered the next outage escalated. onSenderCaughtUp now applies it, and replaces the direct write to the veto field. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Dispatch-Task: main-red-kriszyp_harper-pro_039c07d6c_f0efc59e --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
kriszyp
added a commit
that referenced
this pull request
Sep 30, 2026
… only inline blobs can earn one Round-12 rebase review finding (post-#800 rebase onto main): settleReceiptRequests reused storedBlobsAreComplete, the duplicate-skip verifier, to gate a receipt on blob durability. That function only counts FILE-backed blobs and returns false whenever none are found, including when a HAS_BLOBS row's blobs are all inline (below core's file-storage threshold) -- so a transition image whose blobs are all inline could never earn a receipt: every request waited out its TTL and the sender's sweep re-sent the same image forever, pinning the image permanently. Added receiptBlobsComplete: it still finds every reachable blob to catch a genuinely unreachable HAS_BLOBS claim (the same conservative case storedBlobsAreComplete declines), but an inline blob needs no file check -- its bytes are already inside the row's own encoded value, which the caller already confirmed durable. Only file-backed blobs still go through blobFileMissingOrIncompleteAsync. storedBlobsAreComplete/isDurableIdentityTie are untouched, so the duplicate-skip's conservative behavior for inline blobs is unchanged. New unit coverage in durableIdentityTie.test.mjs exercises real stored blobs (inline and file-backed) through receiptBlobsComplete, the same fixture pattern the neighboring isDurableIdentityTie file-blob tests use. Also: correct an inbound-batch comment that overclaimed uniqueness (the sender guarantees no repeated record per batch; a violation would only race recordHandoffReceipt to a lower stored version, harmlessly), and fix a missing line break in DESIGN.md that ran the Redelivery bullet into Release. Declined, with reasons (framing recheck cleared this change at chosen-approach-sound before implementation; both declines recorded in the PR body's decision ledger): a same-version-different-origin receipt-proof gap that needs a wire/capability change for a narrow, compound scenario; a resweep-timer/wait-site nit whose cost is zero on any core that exists today and whose fix touches a hot loop already hardened over many liveness rounds. Dispatch-Task: pr-maint-0e208afb66ea625b57ed77ba834daf66 Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
kriszyp
added a commit
that referenced
this pull request
Oct 6, 2026
… only inline blobs can earn one Round-12 rebase review finding (post-#800 rebase onto main): settleReceiptRequests reused storedBlobsAreComplete, the duplicate-skip verifier, to gate a receipt on blob durability. That function only counts FILE-backed blobs and returns false whenever none are found, including when a HAS_BLOBS row's blobs are all inline (below core's file-storage threshold) -- so a transition image whose blobs are all inline could never earn a receipt: every request waited out its TTL and the sender's sweep re-sent the same image forever, pinning the image permanently. Added receiptBlobsComplete: it still finds every reachable blob to catch a genuinely unreachable HAS_BLOBS claim (the same conservative case storedBlobsAreComplete declines), but an inline blob needs no file check -- its bytes are already inside the row's own encoded value, which the caller already confirmed durable. Only file-backed blobs still go through blobFileMissingOrIncompleteAsync. storedBlobsAreComplete/isDurableIdentityTie are untouched, so the duplicate-skip's conservative behavior for inline blobs is unchanged. New unit coverage in durableIdentityTie.test.mjs exercises real stored blobs (inline and file-backed) through receiptBlobsComplete, the same fixture pattern the neighboring isDurableIdentityTie file-blob tests use. Also: correct an inbound-batch comment that overclaimed uniqueness (the sender guarantees no repeated record per batch; a violation would only race recordHandoffReceipt to a lower stored version, harmlessly), and fix a missing line break in DESIGN.md that ran the Redelivery bullet into Release. Declined, with reasons (framing recheck cleared this change at chosen-approach-sound before implementation; both declines recorded in the PR body's decision ledger): a same-version-different-origin receipt-proof gap that needs a wire/capability change for a narrow, compound scenario; a resweep-timer/wait-site nit whose cost is zero on any core that exists today and whose fix touches a hot loop already hardened over many liveness rounds. Dispatch-Task: pr-maint-0e208afb66ea625b57ed77ba834daf66 Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
kriszyp
added a commit
that referenced
this pull request
Oct 6, 2026
… only inline blobs can earn one Round-12 rebase review finding (post-#800 rebase onto main): settleReceiptRequests reused storedBlobsAreComplete, the duplicate-skip verifier, to gate a receipt on blob durability. That function only counts FILE-backed blobs and returns false whenever none are found, including when a HAS_BLOBS row's blobs are all inline (below core's file-storage threshold) -- so a transition image whose blobs are all inline could never earn a receipt: every request waited out its TTL and the sender's sweep re-sent the same image forever, pinning the image permanently. Added receiptBlobsComplete: it still finds every reachable blob to catch a genuinely unreachable HAS_BLOBS claim (the same conservative case storedBlobsAreComplete declines), but an inline blob needs no file check -- its bytes are already inside the row's own encoded value, which the caller already confirmed durable. Only file-backed blobs still go through blobFileMissingOrIncompleteAsync. storedBlobsAreComplete/isDurableIdentityTie are untouched, so the duplicate-skip's conservative behavior for inline blobs is unchanged. New unit coverage in durableIdentityTie.test.mjs exercises real stored blobs (inline and file-backed) through receiptBlobsComplete, the same fixture pattern the neighboring isDurableIdentityTie file-blob tests use. Also: correct an inbound-batch comment that overclaimed uniqueness (the sender guarantees no repeated record per batch; a violation would only race recordHandoffReceipt to a lower stored version, harmlessly), and fix a missing line break in DESIGN.md that ran the Redelivery bullet into Release. Declined, with reasons (framing recheck cleared this change at chosen-approach-sound before implementation; both declines recorded in the PR body's decision ledger): a same-version-different-origin receipt-proof gap that needs a wire/capability change for a narrow, compound scenario; a resweep-timer/wait-site nit whose cost is zero on any core that exists today and whose fix touches a hot loop already hardened over many liveness rounds. Dispatch-Task: pr-maint-0e208afb66ea625b57ed77ba834daf66 Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
github-actions Bot
pushed a commit
that referenced
this pull request
Oct 6, 2026
… only inline blobs can earn one Round-12 rebase review finding (post-#800 rebase onto main): settleReceiptRequests reused storedBlobsAreComplete, the duplicate-skip verifier, to gate a receipt on blob durability. That function only counts FILE-backed blobs and returns false whenever none are found, including when a HAS_BLOBS row's blobs are all inline (below core's file-storage threshold) -- so a transition image whose blobs are all inline could never earn a receipt: every request waited out its TTL and the sender's sweep re-sent the same image forever, pinning the image permanently. Added receiptBlobsComplete: it still finds every reachable blob to catch a genuinely unreachable HAS_BLOBS claim (the same conservative case storedBlobsAreComplete declines), but an inline blob needs no file check -- its bytes are already inside the row's own encoded value, which the caller already confirmed durable. Only file-backed blobs still go through blobFileMissingOrIncompleteAsync. storedBlobsAreComplete/isDurableIdentityTie are untouched, so the duplicate-skip's conservative behavior for inline blobs is unchanged. New unit coverage in durableIdentityTie.test.mjs exercises real stored blobs (inline and file-backed) through receiptBlobsComplete, the same fixture pattern the neighboring isDurableIdentityTie file-blob tests use. Also: correct an inbound-batch comment that overclaimed uniqueness (the sender guarantees no repeated record per batch; a violation would only race recordHandoffReceipt to a lower stored version, harmlessly), and fix a missing line break in DESIGN.md that ran the Redelivery bullet into Release. Declined, with reasons (framing recheck cleared this change at chosen-approach-sound before implementation; both declines recorded in the PR body's decision ledger): a same-version-different-origin receipt-proof gap that needs a wire/capability change for a narrow, compound scenario; a resweep-timer/wait-site nit whose cost is zero on any core that exists today and whose fix touches a hot loop already hardened over many liveness rounds. Dispatch-Task: pr-maint-0e208afb66ea625b57ed77ba834daf66 Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
kriszyp
added a commit
that referenced
this pull request
Oct 6, 2026
… only inline blobs can earn one Round-12 rebase review finding (post-#800 rebase onto main): settleReceiptRequests reused storedBlobsAreComplete, the duplicate-skip verifier, to gate a receipt on blob durability. That function only counts FILE-backed blobs and returns false whenever none are found, including when a HAS_BLOBS row's blobs are all inline (below core's file-storage threshold) -- so a transition image whose blobs are all inline could never earn a receipt: every request waited out its TTL and the sender's sweep re-sent the same image forever, pinning the image permanently. Added receiptBlobsComplete: it still finds every reachable blob to catch a genuinely unreachable HAS_BLOBS claim (the same conservative case storedBlobsAreComplete declines), but an inline blob needs no file check -- its bytes are already inside the row's own encoded value, which the caller already confirmed durable. Only file-backed blobs still go through blobFileMissingOrIncompleteAsync. storedBlobsAreComplete/isDurableIdentityTie are untouched, so the duplicate-skip's conservative behavior for inline blobs is unchanged. New unit coverage in durableIdentityTie.test.mjs exercises real stored blobs (inline and file-backed) through receiptBlobsComplete, the same fixture pattern the neighboring isDurableIdentityTie file-blob tests use. Also: correct an inbound-batch comment that overclaimed uniqueness (the sender guarantees no repeated record per batch; a violation would only race recordHandoffReceipt to a lower stored version, harmlessly), and fix a missing line break in DESIGN.md that ran the Redelivery bullet into Release. Declined, with reasons (framing recheck cleared this change at chosen-approach-sound before implementation; both declines recorded in the PR body's decision ledger): a same-version-different-origin receipt-proof gap that needs a wire/capability change for a narrow, compound scenario; a resweep-timer/wait-site nit whose cost is zero on any core that exists today and whose fix touches a hot loop already hardened over many liveness rounds. Dispatch-Task: pr-maint-0e208afb66ea625b57ed77ba834daf66 Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
github-actions Bot
pushed a commit
that referenced
this pull request
Oct 6, 2026
… only inline blobs can earn one Round-12 rebase review finding (post-#800 rebase onto main): settleReceiptRequests reused storedBlobsAreComplete, the duplicate-skip verifier, to gate a receipt on blob durability. That function only counts FILE-backed blobs and returns false whenever none are found, including when a HAS_BLOBS row's blobs are all inline (below core's file-storage threshold) -- so a transition image whose blobs are all inline could never earn a receipt: every request waited out its TTL and the sender's sweep re-sent the same image forever, pinning the image permanently. Added receiptBlobsComplete: it still finds every reachable blob to catch a genuinely unreachable HAS_BLOBS claim (the same conservative case storedBlobsAreComplete declines), but an inline blob needs no file check -- its bytes are already inside the row's own encoded value, which the caller already confirmed durable. Only file-backed blobs still go through blobFileMissingOrIncompleteAsync. storedBlobsAreComplete/isDurableIdentityTie are untouched, so the duplicate-skip's conservative behavior for inline blobs is unchanged. New unit coverage in durableIdentityTie.test.mjs exercises real stored blobs (inline and file-backed) through receiptBlobsComplete, the same fixture pattern the neighboring isDurableIdentityTie file-blob tests use. Also: correct an inbound-batch comment that overclaimed uniqueness (the sender guarantees no repeated record per batch; a violation would only race recordHandoffReceipt to a lower stored version, harmlessly), and fix a missing line break in DESIGN.md that ran the Redelivery bullet into Release. Declined, with reasons (framing recheck cleared this change at chosen-approach-sound before implementation; both declines recorded in the PR body's decision ledger): a same-version-different-origin receipt-proof gap that needs a wire/capability change for a narrow, compound scenario; a resweep-timer/wait-site nit whose cost is zero on any core that exists today and whose fix touches a hot loop already hardened over many liveness rounds. Dispatch-Task: pr-maint-0e208afb66ea625b57ed77ba834daf66 Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
github-actions Bot
pushed a commit
that referenced
this pull request
Oct 6, 2026
… only inline blobs can earn one Round-12 rebase review finding (post-#800 rebase onto main): settleReceiptRequests reused storedBlobsAreComplete, the duplicate-skip verifier, to gate a receipt on blob durability. That function only counts FILE-backed blobs and returns false whenever none are found, including when a HAS_BLOBS row's blobs are all inline (below core's file-storage threshold) -- so a transition image whose blobs are all inline could never earn a receipt: every request waited out its TTL and the sender's sweep re-sent the same image forever, pinning the image permanently. Added receiptBlobsComplete: it still finds every reachable blob to catch a genuinely unreachable HAS_BLOBS claim (the same conservative case storedBlobsAreComplete declines), but an inline blob needs no file check -- its bytes are already inside the row's own encoded value, which the caller already confirmed durable. Only file-backed blobs still go through blobFileMissingOrIncompleteAsync. storedBlobsAreComplete/isDurableIdentityTie are untouched, so the duplicate-skip's conservative behavior for inline blobs is unchanged. New unit coverage in durableIdentityTie.test.mjs exercises real stored blobs (inline and file-backed) through receiptBlobsComplete, the same fixture pattern the neighboring isDurableIdentityTie file-blob tests use. Also: correct an inbound-batch comment that overclaimed uniqueness (the sender guarantees no repeated record per batch; a violation would only race recordHandoffReceipt to a lower stored version, harmlessly), and fix a missing line break in DESIGN.md that ran the Redelivery bullet into Release. Declined, with reasons (framing recheck cleared this change at chosen-approach-sound before implementation; both declines recorded in the PR body's decision ledger): a same-version-different-origin receipt-proof gap that needs a wire/capability change for a narrow, compound scenario; a resweep-timer/wait-site nit whose cost is zero on any core that exists today and whose fix touches a hot loop already hardened over many liveness rounds. Dispatch-Task: pr-maint-0e208afb66ea625b57ed77ba834daf66 Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
github-actions Bot
pushed a commit
that referenced
this pull request
Oct 6, 2026
… only inline blobs can earn one Round-12 rebase review finding (post-#800 rebase onto main): settleReceiptRequests reused storedBlobsAreComplete, the duplicate-skip verifier, to gate a receipt on blob durability. That function only counts FILE-backed blobs and returns false whenever none are found, including when a HAS_BLOBS row's blobs are all inline (below core's file-storage threshold) -- so a transition image whose blobs are all inline could never earn a receipt: every request waited out its TTL and the sender's sweep re-sent the same image forever, pinning the image permanently. Added receiptBlobsComplete: it still finds every reachable blob to catch a genuinely unreachable HAS_BLOBS claim (the same conservative case storedBlobsAreComplete declines), but an inline blob needs no file check -- its bytes are already inside the row's own encoded value, which the caller already confirmed durable. Only file-backed blobs still go through blobFileMissingOrIncompleteAsync. storedBlobsAreComplete/isDurableIdentityTie are untouched, so the duplicate-skip's conservative behavior for inline blobs is unchanged. New unit coverage in durableIdentityTie.test.mjs exercises real stored blobs (inline and file-backed) through receiptBlobsComplete, the same fixture pattern the neighboring isDurableIdentityTie file-blob tests use. Also: correct an inbound-batch comment that overclaimed uniqueness (the sender guarantees no repeated record per batch; a violation would only race recordHandoffReceipt to a lower stored version, harmlessly), and fix a missing line break in DESIGN.md that ran the Redelivery bullet into Release. Declined, with reasons (framing recheck cleared this change at chosen-approach-sound before implementation; both declines recorded in the PR body's decision ledger): a same-version-different-origin receipt-proof gap that needs a wire/capability change for a narrow, compound scenario; a resweep-timer/wait-site nit whose cost is zero on any core that exists today and whose fix touches a hot loop already hardened over many liveness rounds. Dispatch-Task: pr-maint-0e208afb66ea625b57ed77ba834daf66 Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
This was referenced Oct 6, 2026
kriszyp
added a commit
that referenced
this pull request
Oct 7, 2026
… only inline blobs can earn one Round-12 rebase review finding (post-#800 rebase onto main): settleReceiptRequests reused storedBlobsAreComplete, the duplicate-skip verifier, to gate a receipt on blob durability. That function only counts FILE-backed blobs and returns false whenever none are found, including when a HAS_BLOBS row's blobs are all inline (below core's file-storage threshold) -- so a transition image whose blobs are all inline could never earn a receipt: every request waited out its TTL and the sender's sweep re-sent the same image forever, pinning the image permanently. Added receiptBlobsComplete: it still finds every reachable blob to catch a genuinely unreachable HAS_BLOBS claim (the same conservative case storedBlobsAreComplete declines), but an inline blob needs no file check -- its bytes are already inside the row's own encoded value, which the caller already confirmed durable. Only file-backed blobs still go through blobFileMissingOrIncompleteAsync. storedBlobsAreComplete/isDurableIdentityTie are untouched, so the duplicate-skip's conservative behavior for inline blobs is unchanged. New unit coverage in durableIdentityTie.test.mjs exercises real stored blobs (inline and file-backed) through receiptBlobsComplete, the same fixture pattern the neighboring isDurableIdentityTie file-blob tests use. Also: correct an inbound-batch comment that overclaimed uniqueness (the sender guarantees no repeated record per batch; a violation would only race recordHandoffReceipt to a lower stored version, harmlessly), and fix a missing line break in DESIGN.md that ran the Redelivery bullet into Release. Declined, with reasons (framing recheck cleared this change at chosen-approach-sound before implementation; both declines recorded in the PR body's decision ledger): a same-version-different-origin receipt-proof gap that needs a wire/capability change for a narrow, compound scenario; a resweep-timer/wait-site nit whose cost is zero on any core that exists today and whose fix touches a hot loop already hardened over many liveness rounds. Dispatch-Task: pr-maint-0e208afb66ea625b57ed77ba834daf66 Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
github-actions Bot
pushed a commit
that referenced
this pull request
Oct 7, 2026
… only inline blobs can earn one Round-12 rebase review finding (post-#800 rebase onto main): settleReceiptRequests reused storedBlobsAreComplete, the duplicate-skip verifier, to gate a receipt on blob durability. That function only counts FILE-backed blobs and returns false whenever none are found, including when a HAS_BLOBS row's blobs are all inline (below core's file-storage threshold) -- so a transition image whose blobs are all inline could never earn a receipt: every request waited out its TTL and the sender's sweep re-sent the same image forever, pinning the image permanently. Added receiptBlobsComplete: it still finds every reachable blob to catch a genuinely unreachable HAS_BLOBS claim (the same conservative case storedBlobsAreComplete declines), but an inline blob needs no file check -- its bytes are already inside the row's own encoded value, which the caller already confirmed durable. Only file-backed blobs still go through blobFileMissingOrIncompleteAsync. storedBlobsAreComplete/isDurableIdentityTie are untouched, so the duplicate-skip's conservative behavior for inline blobs is unchanged. New unit coverage in durableIdentityTie.test.mjs exercises real stored blobs (inline and file-backed) through receiptBlobsComplete, the same fixture pattern the neighboring isDurableIdentityTie file-blob tests use. Also: correct an inbound-batch comment that overclaimed uniqueness (the sender guarantees no repeated record per batch; a violation would only race recordHandoffReceipt to a lower stored version, harmlessly), and fix a missing line break in DESIGN.md that ran the Redelivery bullet into Release. Declined, with reasons (framing recheck cleared this change at chosen-approach-sound before implementation; both declines recorded in the PR body's decision ledger): a same-version-different-origin receipt-proof gap that needs a wire/capability change for a narrow, compound scenario; a resweep-timer/wait-site nit whose cost is zero on any core that exists today and whose fix touches a hot loop already hardened over many liveness rounds. Dispatch-Task: pr-maint-0e208afb66ea625b57ed77ba834daf66 Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
kriszyp
added a commit
that referenced
this pull request
Oct 7, 2026
… only inline blobs can earn one Round-12 rebase review finding (post-#800 rebase onto main): settleReceiptRequests reused storedBlobsAreComplete, the duplicate-skip verifier, to gate a receipt on blob durability. That function only counts FILE-backed blobs and returns false whenever none are found, including when a HAS_BLOBS row's blobs are all inline (below core's file-storage threshold) -- so a transition image whose blobs are all inline could never earn a receipt: every request waited out its TTL and the sender's sweep re-sent the same image forever, pinning the image permanently. Added receiptBlobsComplete: it still finds every reachable blob to catch a genuinely unreachable HAS_BLOBS claim (the same conservative case storedBlobsAreComplete declines), but an inline blob needs no file check -- its bytes are already inside the row's own encoded value, which the caller already confirmed durable. Only file-backed blobs still go through blobFileMissingOrIncompleteAsync. storedBlobsAreComplete/isDurableIdentityTie are untouched, so the duplicate-skip's conservative behavior for inline blobs is unchanged. New unit coverage in durableIdentityTie.test.mjs exercises real stored blobs (inline and file-backed) through receiptBlobsComplete, the same fixture pattern the neighboring isDurableIdentityTie file-blob tests use. Also: correct an inbound-batch comment that overclaimed uniqueness (the sender guarantees no repeated record per batch; a violation would only race recordHandoffReceipt to a lower stored version, harmlessly), and fix a missing line break in DESIGN.md that ran the Redelivery bullet into Release. Declined, with reasons (framing recheck cleared this change at chosen-approach-sound before implementation; both declines recorded in the PR body's decision ledger): a same-version-different-origin receipt-proof gap that needs a wire/capability change for a narrow, compound scenario; a resweep-timer/wait-site nit whose cost is zero on any core that exists today and whose fix touches a hot loop already hardened over many liveness rounds. Dispatch-Task: pr-maint-0e208afb66ea625b57ed77ba834daf66 Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
github-actions Bot
pushed a commit
that referenced
this pull request
Oct 7, 2026
… only inline blobs can earn one Round-12 rebase review finding (post-#800 rebase onto main): settleReceiptRequests reused storedBlobsAreComplete, the duplicate-skip verifier, to gate a receipt on blob durability. That function only counts FILE-backed blobs and returns false whenever none are found, including when a HAS_BLOBS row's blobs are all inline (below core's file-storage threshold) -- so a transition image whose blobs are all inline could never earn a receipt: every request waited out its TTL and the sender's sweep re-sent the same image forever, pinning the image permanently. Added receiptBlobsComplete: it still finds every reachable blob to catch a genuinely unreachable HAS_BLOBS claim (the same conservative case storedBlobsAreComplete declines), but an inline blob needs no file check -- its bytes are already inside the row's own encoded value, which the caller already confirmed durable. Only file-backed blobs still go through blobFileMissingOrIncompleteAsync. storedBlobsAreComplete/isDurableIdentityTie are untouched, so the duplicate-skip's conservative behavior for inline blobs is unchanged. New unit coverage in durableIdentityTie.test.mjs exercises real stored blobs (inline and file-backed) through receiptBlobsComplete, the same fixture pattern the neighboring isDurableIdentityTie file-blob tests use. Also: correct an inbound-batch comment that overclaimed uniqueness (the sender guarantees no repeated record per batch; a violation would only race recordHandoffReceipt to a lower stored version, harmlessly), and fix a missing line break in DESIGN.md that ran the Redelivery bullet into Release. Declined, with reasons (framing recheck cleared this change at chosen-approach-sound before implementation; both declines recorded in the PR body's decision ledger): a same-version-different-origin receipt-proof gap that needs a wire/capability change for a narrow, compound scenario; a resweep-timer/wait-site nit whose cost is zero on any core that exists today and whose fix touches a hot loop already hardened over many liveness rounds. Dispatch-Task: pr-maint-0e208afb66ea625b57ed77ba834daf66 Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
kriszyp
added a commit
that referenced
this pull request
Oct 9, 2026
… only inline blobs can earn one Round-12 rebase review finding (post-#800 rebase onto main): settleReceiptRequests reused storedBlobsAreComplete, the duplicate-skip verifier, to gate a receipt on blob durability. That function only counts FILE-backed blobs and returns false whenever none are found, including when a HAS_BLOBS row's blobs are all inline (below core's file-storage threshold) -- so a transition image whose blobs are all inline could never earn a receipt: every request waited out its TTL and the sender's sweep re-sent the same image forever, pinning the image permanently. Added receiptBlobsComplete: it still finds every reachable blob to catch a genuinely unreachable HAS_BLOBS claim (the same conservative case storedBlobsAreComplete declines), but an inline blob needs no file check -- its bytes are already inside the row's own encoded value, which the caller already confirmed durable. Only file-backed blobs still go through blobFileMissingOrIncompleteAsync. storedBlobsAreComplete/isDurableIdentityTie are untouched, so the duplicate-skip's conservative behavior for inline blobs is unchanged. New unit coverage in durableIdentityTie.test.mjs exercises real stored blobs (inline and file-backed) through receiptBlobsComplete, the same fixture pattern the neighboring isDurableIdentityTie file-blob tests use. Also: correct an inbound-batch comment that overclaimed uniqueness (the sender guarantees no repeated record per batch; a violation would only race recordHandoffReceipt to a lower stored version, harmlessly), and fix a missing line break in DESIGN.md that ran the Redelivery bullet into Release. Declined, with reasons (framing recheck cleared this change at chosen-approach-sound before implementation; both declines recorded in the PR body's decision ledger): a same-version-different-origin receipt-proof gap that needs a wire/capability change for a narrow, compound scenario; a resweep-timer/wait-site nit whose cost is zero on any core that exists today and whose fix touches a hot loop already hardened over many liveness rounds. Dispatch-Task: pr-maint-0e208afb66ea625b57ed77ba834daf66 Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
github-actions Bot
pushed a commit
that referenced
this pull request
Oct 9, 2026
… only inline blobs can earn one Round-12 rebase review finding (post-#800 rebase onto main): settleReceiptRequests reused storedBlobsAreComplete, the duplicate-skip verifier, to gate a receipt on blob durability. That function only counts FILE-backed blobs and returns false whenever none are found, including when a HAS_BLOBS row's blobs are all inline (below core's file-storage threshold) -- so a transition image whose blobs are all inline could never earn a receipt: every request waited out its TTL and the sender's sweep re-sent the same image forever, pinning the image permanently. Added receiptBlobsComplete: it still finds every reachable blob to catch a genuinely unreachable HAS_BLOBS claim (the same conservative case storedBlobsAreComplete declines), but an inline blob needs no file check -- its bytes are already inside the row's own encoded value, which the caller already confirmed durable. Only file-backed blobs still go through blobFileMissingOrIncompleteAsync. storedBlobsAreComplete/isDurableIdentityTie are untouched, so the duplicate-skip's conservative behavior for inline blobs is unchanged. New unit coverage in durableIdentityTie.test.mjs exercises real stored blobs (inline and file-backed) through receiptBlobsComplete, the same fixture pattern the neighboring isDurableIdentityTie file-blob tests use. Also: correct an inbound-batch comment that overclaimed uniqueness (the sender guarantees no repeated record per batch; a violation would only race recordHandoffReceipt to a lower stored version, harmlessly), and fix a missing line break in DESIGN.md that ran the Redelivery bullet into Release. Declined, with reasons (framing recheck cleared this change at chosen-approach-sound before implementation; both declines recorded in the PR body's decision ledger): a same-version-different-origin receipt-proof gap that needs a wire/capability change for a narrow, compound scenario; a resweep-timer/wait-site nit whose cost is zero on any core that exists today and whose fix touches a hot loop already hardened over many liveness rounds. Dispatch-Task: pr-maint-0e208afb66ea625b57ed77ba834daf66 Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
kriszyp
added a commit
that referenced
this pull request
Oct 10, 2026
… only inline blobs can earn one Round-12 rebase review finding (post-#800 rebase onto main): settleReceiptRequests reused storedBlobsAreComplete, the duplicate-skip verifier, to gate a receipt on blob durability. That function only counts FILE-backed blobs and returns false whenever none are found, including when a HAS_BLOBS row's blobs are all inline (below core's file-storage threshold) -- so a transition image whose blobs are all inline could never earn a receipt: every request waited out its TTL and the sender's sweep re-sent the same image forever, pinning the image permanently. Added receiptBlobsComplete: it still finds every reachable blob to catch a genuinely unreachable HAS_BLOBS claim (the same conservative case storedBlobsAreComplete declines), but an inline blob needs no file check -- its bytes are already inside the row's own encoded value, which the caller already confirmed durable. Only file-backed blobs still go through blobFileMissingOrIncompleteAsync. storedBlobsAreComplete/isDurableIdentityTie are untouched, so the duplicate-skip's conservative behavior for inline blobs is unchanged. New unit coverage in durableIdentityTie.test.mjs exercises real stored blobs (inline and file-backed) through receiptBlobsComplete, the same fixture pattern the neighboring isDurableIdentityTie file-blob tests use. Also: correct an inbound-batch comment that overclaimed uniqueness (the sender guarantees no repeated record per batch; a violation would only race recordHandoffReceipt to a lower stored version, harmlessly), and fix a missing line break in DESIGN.md that ran the Redelivery bullet into Release. Declined, with reasons (framing recheck cleared this change at chosen-approach-sound before implementation; both declines recorded in the PR body's decision ledger): a same-version-different-origin receipt-proof gap that needs a wire/capability change for a narrow, compound scenario; a resweep-timer/wait-site nit whose cost is zero on any core that exists today and whose fix touches a hot loop already hardened over many liveness rounds. Dispatch-Task: pr-maint-0e208afb66ea625b57ed77ba834daf66 Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
github-actions Bot
pushed a commit
that referenced
this pull request
Oct 10, 2026
… only inline blobs can earn one Round-12 rebase review finding (post-#800 rebase onto main): settleReceiptRequests reused storedBlobsAreComplete, the duplicate-skip verifier, to gate a receipt on blob durability. That function only counts FILE-backed blobs and returns false whenever none are found, including when a HAS_BLOBS row's blobs are all inline (below core's file-storage threshold) -- so a transition image whose blobs are all inline could never earn a receipt: every request waited out its TTL and the sender's sweep re-sent the same image forever, pinning the image permanently. Added receiptBlobsComplete: it still finds every reachable blob to catch a genuinely unreachable HAS_BLOBS claim (the same conservative case storedBlobsAreComplete declines), but an inline blob needs no file check -- its bytes are already inside the row's own encoded value, which the caller already confirmed durable. Only file-backed blobs still go through blobFileMissingOrIncompleteAsync. storedBlobsAreComplete/isDurableIdentityTie are untouched, so the duplicate-skip's conservative behavior for inline blobs is unchanged. New unit coverage in durableIdentityTie.test.mjs exercises real stored blobs (inline and file-backed) through receiptBlobsComplete, the same fixture pattern the neighboring isDurableIdentityTie file-blob tests use. Also: correct an inbound-batch comment that overclaimed uniqueness (the sender guarantees no repeated record per batch; a violation would only race recordHandoffReceipt to a lower stored version, harmlessly), and fix a missing line break in DESIGN.md that ran the Redelivery bullet into Release. Declined, with reasons (framing recheck cleared this change at chosen-approach-sound before implementation; both declines recorded in the PR body's decision ledger): a same-version-different-origin receipt-proof gap that needs a wire/capability change for a narrow, compound scenario; a resweep-timer/wait-site nit whose cost is zero on any core that exists today and whose fix touches a hot loop already hardened over many liveness rounds. Dispatch-Task: pr-maint-0e208afb66ea625b57ed77ba834daf66 Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
github-actions Bot
pushed a commit
that referenced
this pull request
Oct 10, 2026
… only inline blobs can earn one Round-12 rebase review finding (post-#800 rebase onto main): settleReceiptRequests reused storedBlobsAreComplete, the duplicate-skip verifier, to gate a receipt on blob durability. That function only counts FILE-backed blobs and returns false whenever none are found, including when a HAS_BLOBS row's blobs are all inline (below core's file-storage threshold) -- so a transition image whose blobs are all inline could never earn a receipt: every request waited out its TTL and the sender's sweep re-sent the same image forever, pinning the image permanently. Added receiptBlobsComplete: it still finds every reachable blob to catch a genuinely unreachable HAS_BLOBS claim (the same conservative case storedBlobsAreComplete declines), but an inline blob needs no file check -- its bytes are already inside the row's own encoded value, which the caller already confirmed durable. Only file-backed blobs still go through blobFileMissingOrIncompleteAsync. storedBlobsAreComplete/isDurableIdentityTie are untouched, so the duplicate-skip's conservative behavior for inline blobs is unchanged. New unit coverage in durableIdentityTie.test.mjs exercises real stored blobs (inline and file-backed) through receiptBlobsComplete, the same fixture pattern the neighboring isDurableIdentityTie file-blob tests use. Also: correct an inbound-batch comment that overclaimed uniqueness (the sender guarantees no repeated record per batch; a violation would only race recordHandoffReceipt to a lower stored version, harmlessly), and fix a missing line break in DESIGN.md that ran the Redelivery bullet into Release. Declined, with reasons (framing recheck cleared this change at chosen-approach-sound before implementation; both declines recorded in the PR body's decision ledger): a same-version-different-origin receipt-proof gap that needs a wire/capability change for a narrow, compound scenario; a resweep-timer/wait-site nit whose cost is zero on any core that exists today and whose fix touches a hot loop already hardened over many liveness rounds. Dispatch-Task: pr-maint-0e208afb66ea625b57ed77ba834daf66 Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
One backoff schedule for every replication retry site, and admission control on the one surface where a retry could amplify: subscription setup.
The defect (#327)
A transient boot-time DNS failure kept
onNodeUpdate → onDatabasefiring. Each qualifying event turned straight into a retained (notunref'd) 200 mssetTimeout, asubscribe-to-nodemessage, worker-side WebSocket/TLS setup, and oneSetting up subscription with leader …warn — no dedup, no cap, no jitter. Whatever re-drove the event amplified 1:1: ~1,400 log lines/s/node in the field, ending in an OOM kill. Pacing alone would not have fixed it — the amplification is per event, not per delay — so the storm surface gets admission control as well as a schedule.The utility
createBackoffinreplication/backoff.ts—{initialMs, maxMs, minMs?, factor?, jitter?, budgetMs?, maxAttempts?, random?, now?}: exponential ceiling, full jitter, optional fixed floor, optional wall-clock/attempt budget, injectable RNG and monotonic clock. Two properties worth knowing:undefined, never0, so a missed exhaustion check cannot become a busy loop;budgetMsis a deadline read off the clock, not a sum of requested sleeps, so a resolver hang or an event-loop stall cannot stretch a grace period past what it advertises.Before this there was no jitter anywhere in production code: every exponential in the repo was lockstep across a fleet.
Admission control
Subscription setup (
createSubscribeSetupScheduler) holds at most one armed setup per (peer URL, database), carrying the newest payload, on a 200 ms floor / 400 ms → 30 s jittered ceiling that resets when the pair connects. It lives in its own map rather than on theconnectionReplicationMapentry becauseonDatabase's stale-worker path deletes and recreates that entry — per-entry state would be wiped on exactly the path that most needs the dedup. Within a stale-worker sweep, each setup's fire time, on a monotonic clock, slides past any the sweep already armed withinRECONNECT_STAGGER_MS, so the 50 ms spacing that bounds concurrent TLS setup (#446) holds across independent jitter draws, and a pair whose own backoff has escalated is placed at its own delay without dragging the rest of the sweep out to its ceiling. Setups are cancelled on unsubscribe, node deletion and same-name URL migration — all reachable once a pending timer can live 30 s instead of 200 ms.An armed setup never falls onto the main thread. If its worker exited while it waited,
dispatchSubscriptionRequestdefers it (with a warn) to the stale-worker reconcile, which rebinds the entry;mainwould have runsubscribeToNodeon the orchestrator. Only single-thread mode, where the main thread is the worker, dispatches there.Self-catchup is claimed by the first dispatch for one connection entry, which keeps the rider for its life. Until the owning primary connection opens it rides on every recovery re-drive; after that, re-drives to that worker leave it off (its old
startTimewould make the leader rescan from there to now) and a replacement worker is sent it again, since the main thread cannot see catchup finish. The global claim is consumed only once the worker message is accepted, so cancellation, a long offline retry, or a synchronouspostMessagethrow cannot lose it.The worker readiness boundary used to attach a continuation to the component-readiness promise per message (
createWorkerSubscriptionAdmission). Before readiness, one insertion-ordered map retains the latest action per connection key and database behind a single continuation, with exactly one armed re-attempt if that readiness rejects; after readiness the handlers run inline. Subscribe, unsubscribe, force-reconnect and this map all derive identity through onegetSubscriptionConnectionKeyinreplicator.ts, which replaces three hand-builturl + '-' + …keys, including the missing nested-URL fallback.Recovery timers are owned. One
entry.reDriveTimerper entry,unref'd, disarmed on unsubscribe, node deletion, worker exit and entry replacement, with the sweep's jitter drawn once and used as a common base. A connect report resets the escalated setup delay and cancels nothing; each timer re-checks live state when it fires. The stall kick's half of that isshouldFireStallKick, which claims its entry throughconnectGenerationand the worker it captured.Review-thread fixes in this round
onDatabasederives it throughderiveEffectiveLeaderbefore the existing-entry fast path and assigns the result, so an explicit persistedisLeader: falsedemotes immediately. Recomputing exposed a second way to get it wrong: a component reload clearedroutesbefore rebuilding it, so a recomputation could briefly see "no configured leader" while the previoushdb_nodeswatcher stayed live. The reload now builds the list off to the side and publishes it in one splice, keeping the array identity.connectReportAdvancesGenerationadvancesconnectGenerationonly for a socket-open edge (newSocket) from the owning worker's thread whose subscription URL (reported with the open) is the entry's primary. That last check matters because one worker can carry both an entry's primary connection and a proxied failover subscription over the same URL. Pongs, metadata posts, superseded workers' sockets and failover opens no longer move it. A shared-memory truth up-correction does, since it stands in for an open edge that never arrived. Look hardest here.Sites adopting the schedule
scheduleReconnect(full jitter under the unchanged 500 ms → 30 s ceiling, keeping the hard 500 ms floor from #339;onFrameSentis still the only reset), the wedge and receive-stall re-drives, thehdb_nodeswatcher restart, the send-auth reprobe (a 30 s wall-clock deadline in place of 60 fixed sleeps), the copy-cursor flush retry, the two clone probes (monitorSync's leader version probe, 1 s → 4 s;fetchJWTKeyWithRetry, 250 ms → 1 s; attempt counts unchanged), the blob-repair per-record pause, and thesendBlobsin-place 503 re-read.DESIGN.mdcarries the per-site table and the list of what is deliberately excluded.One definition per schedule. The subscription-setup parameters (200 ms floor, 400 ms → 30 s ceiling) live in
createSubscribeBackoff, which the setup scheduler, the wedge/stall re-drive draw and the worker readiness re-attempt all call. The re-drive only ever takes the first draw, so its window is still [200, 400) ms. The scheduler's timing overrides (initialMs/maxMs/minMs/staggerMs) are removed; nothing passed them. The blob-send array[250, 500, 1000, 2000]is nowBLOB_SEND_RETRY_BACKOFF(maxAttempts: 4,jitter: 'none'). It produces the same four waits, and the attempt bound now lives in the schedule instead of being split betweenshouldRetrySourceBlobReadand an array index.sendBlobscreates the backoff lazily on the first eligible 503, so a successful send allocates nothing. The stored-body, partial-send, closed-socket and draining guards are unchanged. It stays unjittered: the retried read is this node's own blob, so there is no fleet to decorrelate, and full jitter would halve the expected time the PENDING placeholder has to heal (DESIGN.mdrow). Planning review for this consolidation:Framing-Verdict: chosen-approach-sound (a4fea92b7e93).The
hdb_nodeswatcher reset bug:iteratedSuccessfullywas set the instantsubscribe()resolved, so a subscription that resolved and immediately threw counted as a success every pass and pinned the restart delay at 1 s forever. Progress is now an iteration that stayed live pastNODE_WATCHER_HEALTHY_UPTIME_MS, measured from when the subscription actually came up.Merged with
mainThe branch was 76 commits behind and conflicting, so
origin/mainis merged in rather than rebased (no force-push; the conflict resolution is one commit). Resolutions that change behaviour: the socket-open report now usesmain'sthreadId/newSocketfields plus this branch'ssubscriptionUrl, instead of two parallel sets; scheduled setups carrymain's send-timeexclusionOrigins; andmain'shasDeadOwnernote is corrected, since an unowned entry now defers instead of subscribing on the main thread.For the human reviewer
Decisions worth disagreeing with:
isLeader: falsebeats a configured leader (HDB_LEADER_URL/ cli /routes[0]);mainignored a persistedfalse. The owner chose this. One consequence:update_node { isLeader: false }is written LOCAL_ONLY only whentrue(setNode.ts), so a demotion issued on one node replicates to every peer and overrides their configuration too. In practice the derived flag only gatesshouldReplicateFromNodefor a database that does not exist locally, and the bootstrap loop that reaches that case already keys on the raw persisted flag. So the reach is narrow, but it is a policy change.mainalways did; closing that needs a worker→main "catchup complete" signal.hdb_nodeswatcher backoff resets only after 10 s of live iteration. A stream that keeps ending sooner stops watching for up to 30 s at a time; the collection subscribe replays current rows on restart, so the cost is latency, not lost updates.refreshPendingside channel, rather than fixingonDatabase's early-return path soentry.nodesis always the enriched array.NodeReplicationConnection.retryTimeis removed. Onmainit was a mutable field; nothing outside the unit tests read it, and the tests now assertretryBackoff.ceiling.getRecord()has no timeout, so one wedged-but-open peer would hold the flag for the life of the thread and take awayrepair_blob_datauntil restart. Concurrent sweeps stack again, as onmain.Still open, tracked or stated:
connectionReplicationMapentry survives with its ownedreDriveTimer. The entry leak predates this branch; the timer riding on it does not.connect()checksintentionallyUnsubscribedonly before thecreateWebSocketawait, so an unsubscribe landing inside a TLS handshake can leave an orphan live session; and the revive path (unsubscribed = false) does not resetreceiveStallReconnectAt, so an unsubscribe/re-add can hide a stalled connection from the stall net until real progress.blobGapEscalationBudgetcovers the 503-forever path (retry, then forward), and the unit test pins the exact waits and exhaustion.startOnMainThreadenclosesonNodeUpdate,onDatabase,connectedToNodeandreconcileWorkers, so every unit test targets an extracted helper. The end-to-end evidence is the cluster and stress suites below.Review findings declined, with the evidence:
findStaleNodeUrlsflags a worker-less entry whenever live workers exist, and a worker exit triggers an immediate reconcile, so the wait is one tick at most — the same boundmainhad when its baresetTimeoutposted to an exited worker.!isWedgedadds TLS churn." A URL that is both stale and wedged now also gets the fullonNodeUpdatepass, but its wedged live-worker entries take the refresh-only early return (forceResubscribeis false there), so no extra dial happens; the pass only re-postsunsubscribe-from-nodefor databases not replicated from that peer, until the stale entries are rebound on that tick.forceReconnect.test.mjs." They name the pinned draw's delays, which is what makestick(500)/tick(998)readable.canClearCapabilitiesForNewSocketignoressubscriptionUrl" and "pin the explicit-leader bootstrap gate with a test" — both pre-existing onmain, recorded for follow-up rather than folded in here.Verification
On the merged head (
mainatc2ad3351), then on the consolidation headbacd6b6c:npm run buildexits 0 with notscerrors.npm run lint:requiredandprettier --checkon changed files clean.npm run test:unit: 1412 passing, 1 pending; 1410 onbacd6b6c(the two removed tests were theretryTimegetter's and the array-length bound's;blobSendRetry.test.mjsnow pins[250, 500, 1000, 2000, undefined], andreconnectJitter/forceReconnect/connectReschedulesOnRejectionreadretryBackoff.ceiling). Covers the backoff matrix, the setup scheduler (a 60,000-event / 60 s storm collapsing to 9 dispatches with one pending timer,hasRef() === false, sweep spacing by fire time under colliding draws, staggered calls, and an escalated pair, deferral when the worker exits before fire), leadership derivation and atomic route publication, socket-open ownership (pong, superseded worker, proxied failover, single-thread), the rider re-sent to a replacement worker, worker readiness admission, and the stall-kick decision matrix.connectedBitRestartChurn(QA-587), locally — 5 of 6 runs green across this round's heads (1/1 on the final head), including the new rolling HTTP-worker replacement test (restart_service http_workers, then a write per database must reach the follower) in every run. The one red run was chaos cycle 3 missing its 25 s budget: one leg connected ~100 ms before the SIGKILL, never sent a frame, so its reconnect ceiling was not reset and it drew 29.6 s at the 30 s cap.mainhas the same reset-on-first-sent-frame rule and waits the full ceiling, so this is the test's budget against that rule, not this change.backoff.test.mjs(ceiling growth, factor, full-jitter range, floor, budget/attempt exhaustion on an injected clock),reconnectJitter.test.mjs(decorrelation, the 500 ms → 30 s ceilings, the 500 ms floor,onFrameSentreset),subscribeSetupScheduler.test.mjs(storm bound, dedup, sweep spacing, deferral),workerSubscriptionAdmission.test.mjs(connection-key derivation, latest-action retention, rejected-readiness re-attempt),subscriptionRoutingOwnership.test.mjs(atomic route publication, persistedisLeaderprecedence, socket-open ownership),shouldFireStallKick.test.mjs(the stall-kick decision matrix) andblobRepairPacing.test.mjs(the budget starts at the first unrepairable record). Extended:nodeUpdateWatcher.test.mjs(escalation when iteration throws at once, healthy-uptime reset, subscribe time not counted as uptime),shouldCloseSendAuthWatch.test.mjs(fail closed past the deadline, including a read that crosses it),findStalledReceivingNodeUrls.test.mjs(a fresh socket's grace period), andforceReconnect.test.mjs/connectReschedulesOnRejection.test.mjs(pinned draws against the jittered ceiling, read fromretryBackoff.ceiling).connectedBitRestartChurn.test.mjsgains the rolling HTTP-worker replacement test (and a comment now names the reconnect ceiling instead of the removedretryTime);fixture-blob-pending-source/resources.jschanges only a comment, to the new constant's name.bacd6b6c, cluster —blobGapEscalationBudget(exercises the in-place 503 re-read, then the forward) andconnectedBitRestartChurn: 6/6 pass.integrationTests/stress/wedgedPeerSubscribeStorm.test.mjsunderHARPER_RUN_STRESS_TESTS=1: passes on the merged head and again after the round-1 review fixes (boundedSetting up subscription with leaderrate, no listener leak, no OOM marker, RSS under cap, reconverging mesh with a peer permanently wedged).Pre-push review of this round (
prepush-review.mjs --author claude --legs auto, codex + gemini + cursor-composer + harper-domain every round): a full round on the merged head (adjudicated minor), then delta rounds — nit, nit (COMMENTS), and nit (COMMENTS, on the pushed head) after a post-push human review showed the sweep compared relative delays from calls made at different instants (now absolute fire times). The full-round findings — the sweep floor dragged by an escalated pair, self-catchup lost when a worker dies mid-catchup, unbounded integration probes, the unsampled per-peer repair warn — are all fixed above, as is round 2's spacing gap next to a just-escalated pair. The consolidation commits got two delta rounds (all four lenses): nit (COMMENTS; a test title and a field comment, both fixed), then LGTM onbacd6b6c.Refs #327
Fixes #805
🤖 Generated with Claude Code
https://claude.ai/code/session_017CTUbhGiW2AXhogHZLg8eo
— Claude Opus 5.5
Origin — the dispatch brief this PR was written from
Pace replication retries on one jittered schedule and bound subscription setup
LIVE CONVERSATION about #800.
You are answering a person, in a thread, one turn at a time. Every turn:
each of your previous turns is in it. Read the PR/issue and the code as needed.
they are talking to you, and a status template is not an answer.
Each turn arrives as ASK (answer it, change nothing) or PERFORM (do it, then say what you did) —
the person chose which when they sent it, and the run's own prompt tells you which one this is.
Never infer it from the wording: an unrequested commit in the middle of a discussion and a polite
description of work that was supposed to happen are the two failures this exists to prevent.
Never mark a PR ready and never merge from this conversation.
Dispatch: task
chat-pr-harper-pro-800-kriszyp· queued by unknown · ran by claude/opus/xhigh · worker kzyp-xps-1Review-Coverage: authored=claude; ran=codex,cursor-composer,gemini; adjudicated=domain; declined=cursor-grok,cursor-kimi,cursor-muse; rounds=18; full=1 @ bacd6b6
Human-Review-Need: 4 (decisions: self-catchup-retained-for-entry-life, reconnect-first-attempt-unjittered, full-jitter-at-cap, send-auth-fail-closed-at-deadline, blob-repair-budget-then-unpaced, watcher-healthy-uptime-reset) @ bacd6b6