Skip to content

fix(replication): recover open-but-idle wedged subscriptions via watchdog-driven reconnect - #424

Merged
kriszyp merged 1 commit into
mainfrom
kris/420-replication-wedge-recovery
Jun 19, 2026
Merged

kriszyp merged 1 commit into
mainfrom
kris/420-replication-wedge-recovery

Conversation

@kriszyp

@kriszyp kriszyp commented Jun 19, 2026

Copy link
Copy Markdown
Member

Closes #420.

Summary

After a cluster-wide simultaneous restart, a per-DB receive socket can connect and then go open-but-idle — copy stalls, no bytes flow, and no close event ever fires. Because the close handler only retries on close, and the wedge reconciler only re-drives entries already marked connected:false, the (peer, db) pair makes zero further connection attempts and stays wedged until a manual staggered restart. When system is among the wedged sockets, replicated deploys then fail with the 120s hdb_deployment row did not replicate timeout. Observed live on JJill preprod (5.1.5, 4-node, RocksDB).

The root gap: an open-but-idle socket never flips its per-DB entry to connected:false (no close), and even if it did, the cached connection is still "reusable" (replicator.isReusableConnection) so a reconciler re-subscribe just hands back the same dead socket. So recovery has to be driven from the one thing that does fire on silence — the receive watchdog.

What changed

  • NodeReplicationConnection.forceReconnect() — the receive watchdog's onSilence now calls this on the client (subscription) side instead of a bare ws.terminate(). It notifies disconnectedFromNode (flipping the entry to connected:false for the reconciler backstop), tears the socket down, and schedules one fresh connect() — independent of whether close ever fires. Server-accepted connections (no connection object) keep terminate(); the remote client reconnects.
  • reconnectScheduled flag — set by forceReconnect, cleared in connect()'s finally once the new socket is installed. The close handler returns early if it's set, so the close path and forceReconnect never both arm a connect for the same drop.
  • Socket-identity guard in the close handler (if (this.socket !== socket) return) — a late close from a socket forceReconnect already replaced can no longer tear down the live connection.
  • Tests — a unit test pinning forceReconnect's contract (single reconnect, double-arm guard, no-revive on teardown, backoff, stale-listener cleanup), plus an env-gated one-shot fault-injection hook and a cluster integration regression test that wedges a socket open-but-idle and asserts automatic recovery with no restart (validated against a negative control: pre-change code times out).

Scope notes (per the issue's investigator handoff)

Where to look

  • forceReconnect() and the two close-handler guards in replication/replicationConnection.ts — the reconnect lifecycle is the subtle part; the interleavings (fast close, late close during the createWebSocket await, concurrent normal close) are what to scrutinize.
  • The test-only fault-injection hook armReplicationWedgeForTest lives in replicationConnection.ts (env-gated on HARPER_TEST_REPLICATION_WEDGE_DB, one-shot, no-ops in production). It's there because there's no clean way to inject an open-but-idle wedge from outside replicateOverWS; flag it if you'd prefer it factored out.
  • Pre-existing and not introduced here (considered during review): during the wedge window the sender can briefly hold two accepted subscription loops for one (peer, db); the receiver de-dupes by sequence id and the sender's own watchdog reaps the idle side.

A multi-model review (Codex + a Harper-domain replication/concurrency pass) ran on this change; its findings — a stale-close race, a subscriptions-updated listener leak, an await-window double-schedule, and a test-hook production-arming footgun — are all addressed in the current diff.

🤖 Generated by Claude (Opus 4.8, 1M context).

…hdog-driven reconnect

A per-DB receive socket that connects and then goes idle (copy stalls, no
transport close) never fires 'close', so the close-handler retry never arms and
the wedge reconciler skips the still-connected:true entry — the (peer, db) pair
makes zero further connection attempts (harper-pro#420; blocks replicated
deploys when 'system' is among the wedged sockets).

Drive recovery from the receive watchdog instead: NodeReplicationConnection
gains forceReconnect(), invoked from onSilence on the client side. It flips the
node entry to connected:false (reconciler backstop), tears the socket down, and
schedules one fresh connect() independent of whether 'close' ever fires. A
reconnectScheduled flag (cleared in connect()'s finally once the new socket is
installed) plus a socket-identity guard in the close handler keep a late close
from the superseded socket from double-scheduling or tearing down the live
connection.

Adds a unit test for forceReconnect's contract and an env-gated, one-shot
fault-injection hook + cluster integration regression test that wedges a socket
open-but-idle and asserts automatic recovery with no restart.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
@kriszyp
kriszyp requested a review from cb1kenobi June 19, 2026 05:42
@gemini-code-assist

This comment has been minimized.

@claude

claude Bot commented Jun 19, 2026

Copy link
Copy Markdown
Contributor

Reviewed; no blockers found.

@kriszyp
kriszyp marked this pull request as ready for review June 19, 2026 05:49
@kriszyp
kriszyp requested a review from a team as a code owner June 19, 2026 05:49
@gemini-code-assist

Copy link
Copy Markdown

Warning

You have reached your daily quota limit. Please wait up to 24 hours and I will start processing your requests again!

@kriszyp
kriszyp merged commit 6f5668e into main Jun 19, 2026
30 checks passed
@kriszyp
kriszyp deleted the kris/420-replication-wedge-recovery branch June 19, 2026 19:38
kriszyp added a commit that referenced this pull request Jun 25, 2026
)

Adds the deferred third recovery layer from #466 / PR #467. While a receive
leg is paused for back-pressure the byte-silence receiveWatchdog is stopped
(ws.pause() freezes bytesRead) and the active sendPing is exempt, so a leg
that dies mid-pause — e.g. a system base copy stalled at ~100% back-pressure
whose peer restarted — had no recovery driver and could wedge connected:false
forever, removing the #424 forceReconnect path for exactly that case.

A pause-stall watchdog (createPauseStallWatchdog, a thin wrapper over the
existing createReceiveWatchdog stall-timer) now guards the paused window,
keyed on a local consumerProgress counter that advances on signals surviving
ws.pause(): onCommit (apply loop drained a queued batch) and blob-stream
drains. Armed on pause, stopped on resume; exactly one of {receiveWatchdog,
pauseStallWatchdog} is armed at a time. It fires forceReconnect only after a
sustained window of ZERO consumer progress, so a pause that is legitimately
making progress re-arms every window and never trips.

Cross-model review (Codex + agy): also makes resetPingTimer pause-aware so a
frame handler can't re-arm the byte watchdog while paused; threshold defaults
to max(PING_TIMEOUT*2, blobTimeout*2) and the slow-single-operation residual
(benign: re-streams from the durable cursor) is documented.

Tests: unitTests/replication/pauseStallWatchdog.test.mjs. Replication unit
suite green (187 passing). Single-repo (no core change).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Replication wedges permanently after simultaneous cluster restart (reconciler skips open-but-idle sockets) → blocks replicated deploys

1 participant